Coding methods, decoding methods, coders, decoders, code stream and storage medium
By applying conditions in OBMC technology, determining whether to perform overlapping block motion compensation and selecting appropriate fusion mode, the problem of OBMC technology reducing the encoding and decoding quality in video encoding and decoding is solved, and a more efficient encoding and decoding effect is achieved.
Patent Information
- Application Number
- PCT/CN2023/137989
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-19
AI Technical Summary
When used in video encoding and decoding, OBMC technology may reduce the encoding and decoding quality, especially in screen content video scenarios.
By applying certain conditions, OBMC technology is used, for example, based on the motion vector of the current block and the information of adjacent blocks, determine whether to perform motion compensation of overlapping blocks, and select appropriate fusion mode and weight.
The encoding and decoding quality is improved, especially in screen content video encoding, and the compression efficiency of the code stream is improved by reducing block effects and prediction errors.
Smart Images

Figure CN2023137989_19062025_PF_FP_ABST
Abstract
Description
Coding and decoding method, codec, code stream and storage medium Technical Field
[0001] The present application relates to the field of video coding and decoding technology, and in particular to a coding and decoding method, a codec, a bit stream, and a storage medium. Background Art
[0002] Overlapped block motion compensation (OBMC) is a commonly used motion compensation technique. OBMC can be used to eliminate blocking artifacts and reduce prediction errors. However, using OBMC can sometimes reduce codec quality.
[0003] Summary of the Invention
[0004] The present application provides a coding and decoding method, a codec, a bit stream, and a storage medium. The following introduces various aspects of the present application.
[0005] In a first aspect, a decoding method is provided, which is applied to a decoder, and includes: parsing a code stream to determine a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; and performing overlapped block motion compensation on the first prediction block if a first condition is met.
[0006] In a second aspect, a decoding method is provided, which is applied to a decoder, and the decoding method includes: parsing a code stream to determine a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; determining a co-located prediction block of the current block based on the motion vector of a neighboring block of the current block; if a first condition is met, fusing the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponds to a first fusion area, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; if the first condition is not met, fusing the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponds to a second fusion area, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein the first fusion area and the second fusion area are different; and / or the first weight is different from the second weight.
[0007] In a third aspect, a coding method is provided, which is applied to an encoder and includes: determining a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; and performing overlapping block motion compensation on the first prediction block if a first condition is met.
[0008] In a fourth aspect, a coding method is provided, which is applied to an encoder, and the coding method includes: determining a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; determining a co-located prediction block of the current block based on the motion vector of a neighboring block of the current block; if a first condition is met, fusing the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponds to a first fusion area, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; if the first condition is not met, fusing the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponds to a second fusion area, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein the first fusion area and the second fusion area are different; and / or the first weight is different from the second weight.
[0009] In a fifth aspect, a decoder is provided, comprising: a first determination unit configured to parse a code stream and determine a motion vector of a current block; a second determination unit configured to determine a first prediction block of the current block based on the motion vector; and a third determination unit configured to perform overlapping block motion compensation on the first prediction block if a first condition is met.
[0010] In a sixth aspect, a decoder is provided, comprising: a first determination unit configured to parse a code stream and determine a motion vector of a current block; a second determination unit configured to determine a first prediction block of the current block based on the motion vector; a third determination unit configured to determine a co-located prediction block of the current block based on the motion vector of an adjacent block of the current block; a fourth determination unit configured to, if a first condition is satisfied, fuse the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponding to a first fusion area, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; if the first condition is not satisfied, fuse the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponding to a second fusion area, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein the first fusion area and the second fusion area are different; and / or the first weight is different from the second weight.
[0011] In a seventh aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the first aspect or the second aspect when running the computer program.
[0012] In an eighth aspect, an encoder is provided, comprising: a first determination unit configured to determine a motion vector of a current block; a second determination unit configured to determine a first prediction block of the current block based on the motion vector; and a third determination unit configured to perform overlapping block motion compensation on the first prediction block if a first condition is met.
[0013] In a ninth aspect, an encoder is provided, comprising: a first determination unit configured to determine a motion vector of a current block; a second determination unit configured to determine a first prediction block of the current block based on the motion vector; a third determination unit configured to determine a co-located prediction block of the current block based on the motion vector of an adjacent block of the current block; a fourth determination unit configured to, if a first condition is satisfied, fuse the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponds to a first fusion area, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; if the first condition is not satisfied, fuse the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponds to a second fusion area, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein the first fusion area and the second fusion area are different; and / or the first weight is different from the second weight.
[0014] In a tenth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the third aspect or the fourth aspect when running the computer program.
[0015] In the eleventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method described in any one of the first to fourth aspects is implemented.
[0016] In a twelfth aspect, a computer program product is provided, comprising a computer program, which implements the method described in any one of aspects 1 to 4 when the computer program is executed.
[0017] In a thirteenth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method using an encoder, or the bit stream is decoded by a decoding method using a decoder, wherein the decoding method is the method described in the first aspect or the second aspect, and the encoding method is the method described in the third aspect or the fourth aspect.
[0018] In a fourteenth aspect, a code stream is provided, comprising a code stream generated according to the method described in the first aspect or the second aspect; or comprising a code stream generated according to the method described in the third aspect or the fourth aspect.
[0019] The embodiments of the present application impose certain conditions on the use of OBMC, thereby helping to improve the encoding and decoding quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG1 is a structural diagram illustrating an example of a video encoder to which an embodiment of the present application may be applied.
[0021] FIG2 is a diagram showing an example structure of a video decoder to which an embodiment of the present application can be applied.
[0022] FIG3 is a schematic diagram of an overlapping block motion compensation method.
[0023] FIG4 a is a schematic diagram illustrating the use of overlapped block motion compensation in the encoding process.
[0024] FIG4 b is a schematic diagram illustrating the use of overlapped block motion compensation in the decoding process.
[0025] FIG5 is an example diagram of a screen content image.
[0026] FIG6 is a flowchart of a decoding method provided in one embodiment of the present application.
[0027] FIG7 is a flowchart of a decoding method provided in another embodiment of the present application.
[0028] FIG8 is a flow chart of an encoding method provided in one embodiment of the present application.
[0029] FIG9 is a flow chart of an encoding method provided in another embodiment of the present application.
[0030] FIG10 is a schematic diagram of the structure of a decoder provided in one embodiment of the present application.
[0031] FIG11 is a schematic diagram of the structure of a decoder provided in another embodiment of the present application.
[0032] FIG12 is a schematic structural diagram of a decoder provided in yet another embodiment of the present application.
[0033] FIG13 is a schematic diagram of the structure of an encoder provided in one embodiment of the present application.
[0034] FIG14 is a schematic diagram of the structure of an encoder provided in another embodiment of the present application.
[0035] FIG15 is a schematic diagram of the structure of an encoder provided in yet another embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solution in this application will be described below with reference to the accompanying drawings.
[0037] FIG1 is a schematic block diagram of a video encoder according to an embodiment of the present application.
[0038] It should be understood that the video encoder 100 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless compression or mathematically lossless compression.
[0039] The video encoder 100 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0040] For example, the video encoder 100 reads video data and, for each image in the video data, divides the image into a number of coding tree units (CTUs). In some examples, a CTU may be referred to as a "tree block," "largest coding unit" (LCU) or "coding tree block" (CTB). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to one luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with one luminance sample block and two chrominance sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU may be further divided into a number of coding units (CUs) for encoding. A CU may be a rectangular block or a square block. A CU may correspond to a prediction unit (PU) and a transform unit (TU).
[0041] In some embodiments, as shown in FIG1 , the video encoder 100 may include a prediction module 110, a residual module 120, a transform / quantization module 130, an inverse transform / quantization module 140, a reconstruction module 150, a loop filter module 160, a decoded image buffer 170, and an entropy coding module 180. It should be noted that the video encoder 100 may include more, fewer, or different functional components.
[0042] Optionally, in this application, the current block may be referred to as the current coding unit (CU). The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstruction block. Due to the need for parallel processing, an image may be divided into slices. Slices in the same image may be processed in parallel, meaning that there is no data dependency between them. The term "frame" is commonly used, and it can generally be understood that a frame is an image. The term "frame" herein may also be replaced by "image" or "slice," etc.
[0043] In some embodiments, the prediction module 110 includes an inter-frame prediction module 111 and an intra-frame prediction module 112. Because there is a strong correlation between adjacent pixels in a video image, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent images in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.
[0044] The inter-frame prediction module 111 can be used for inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation. It can refer to image information from different images. Inter-frame prediction uses motion information to find a reference block from a reference image and generate a prediction block based on the reference block to eliminate temporal redundancy. Inter-frame prediction uses motion information to find a reference block from a reference image and generate a prediction block based on the reference block. Motion information includes the reference image list in which the reference image is located, the reference image index, and the motion vector. The motion vector can be integer pixel or fractional pixel. If the motion vector is fractional pixel, interpolation filtering is required to generate the required fractional pixel block in the reference image. Here, the integer pixel or fractional pixel block in the reference image found based on the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, while others further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.
[0045] The intra-frame prediction module 112 only refers to information of the same image to predict pixel information within the current code image block to eliminate spatial redundancy.
[0046] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. High-efficiency video coding (HEVC) uses planar, direct current (DC), and 33 angular modes for a total of 35 intra-frame prediction modes. Versatile video coding (VVC) uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.
[0047] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of high-definition and ultra-high-definition digital video development.
[0048] Residual module 120 may generate a residual block for a CU based on the pixel block of the CU and the prediction block of the CU. For example, residual module 120 may generate a residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the CU.
[0049] The transform / quantization module 130 may quantize the transform coefficients. The transform / quantization module 130 may quantize the transform coefficients associated with the CU based on a quantization parameter (QP) value associated with the CU. The video encoder 100 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0050] The inverse transform / quantization module 140 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.
[0051] Reconstruction module 150 can add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction module 110 to generate a reconstructed image block associated with the CU. By reconstructing each sample block of the CU in this manner, video encoder 100 can reconstruct the pixel blocks of the CU.
[0052] The loop filter module 160 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent pixel encoding. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.
[0053] In some embodiments, the loop filtering module 160 includes a deblocking filtering module and a sample adaptive offset / adaptive loop filtering (SAO / ALF) module, wherein the deblocking filtering module is used to remove blocking effects, and the SAO / ALF module is used to remove ringing effects.
[0054] The decoded image buffer 170 may store the reconstructed pixel blocks. The inter prediction module 111 may use a reference image containing the reconstructed pixel blocks to perform inter prediction on PUs of other images. In addition, the intra prediction module 112 may use the reconstructed pixel blocks in the decoded image buffer 170 to perform intra prediction on other PUs in the same image as the CU.
[0055] The entropy encoding module 180 may receive the quantized transform coefficients from the transform / quantization module 130. The entropy encoding module 180 may perform one or more entropy encoding operations on the quantized transform coefficients to generate entropy-encoded data.
[0056] FIG2 is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0057] 2 , video decoder 200 includes an entropy decoding module 210, a prediction module 220, an inverse quantization / transformation module 230, a reconstruction module 240, a loop filter module 250, and a decoded image buffer 260. It should be noted that video decoder 200 may include more, fewer, or different functional components.
[0058] The video decoder 200 may receive a bitstream. The entropy decoding module 210 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding module 210 may parse the entropy-encoded syntax elements in the bitstream. The prediction module 220, the inverse quantization / transformation module 230, the reconstruction module 240, and the loop filter module 250 may decode the video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.
[0059] In some embodiments, the prediction module 220 includes an intra-frame prediction module 222 and an inter-frame prediction module 221 .
[0060] The intra prediction module 222 may perform intra prediction to generate a prediction block for a PU. The intra prediction module 222 may use an intra prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs. The intra prediction module 222 may also determine the intra prediction mode for the PU based on one or more syntax elements parsed from the codestream.
[0061] The inter-frame prediction module 221 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. Furthermore, if a PU is encoded using inter-frame prediction, the entropy decoding module 210 may parse the motion information of the PU. The inter-frame prediction module 221 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter-frame prediction module 221 may generate a prediction block for the PU based on the one or more reference blocks of the PU.
[0062] The inverse quantization / transform module 230 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform module 230 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0063] After inverse quantizing the transform coefficients, inverse quantization / transform module 230 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.
[0064] Reconstruction module 240 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, reconstruction module 240 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.
[0065] The loop filtering module 250 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.
[0066] The video decoder 200 may store the reconstructed image of the CU in the decoded image buffer 260. The video decoder 200 may use the reconstructed image in the decoded image buffer 260 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0067] The basic process of video encoding and decoding is as follows: At the encoder end, an image is divided into blocks. For the current block, the prediction module 110 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual module 120 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block can also be referred to as residual information. This residual block undergoes transformation and quantization by the transform / quantization module 130, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization module 130 can be referred to as a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization module 130 can be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding module 180 receives the quantized change coefficients output by the transform and quantization module 130 and performs entropy coding on these quantized change coefficients to output a bitstream. For example, the entropy coding module 180 can eliminate character redundancy based on the target context model and the probability information of the binary bitstream.
[0068] At the decoding end, the entropy decoding module 210 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction module 220 uses intra-frame prediction or inter-frame prediction on the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation module 230 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction module 240 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering module 250 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference image for inter-frame prediction of subsequent images.
[0069] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.
[0070] It is understandable that the "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text. The "transformation" and "inverse transformation" in the embodiments of the present application correspond to two opposite processes. For example, if the "transformation" converts the numerical values in the spatial domain to the coefficients in the frequency domain, then the "inverse transformation" converts the coefficients in the frequency domain to the numerical values in the spatial domain. If the standard only stipulates decoding, then the "transformation" in the standard text is the decoding part, which refers to the "inverse transformation" in this article. The "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text.
[0071] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0072] The preceding text describes in detail the codec framework provided by the embodiments of this application. This application relates to OBMC, which can be applied to the codec end. OBMC is described in detail below.
[0073] In traditional video codecs, inter-frame prediction is typically based on motion search and motion compensation to eliminate inter-frame information redundancy. Motion search uses block matching to find a reference block in an already encoded reference frame that has the smallest difference from the current block as the best match. The displacement from the reference block to the current block is called a motion vector (MV). Motion compensation uses the reference block as the prediction block for the current block and can determine whether motion interpolation is necessary. The motion search and motion compensation mentioned above are both performed at the coding block level. Therefore, under normal circumstances, there will be problems such as blocking artifacts and boundary errors between the predicted block and the adjacent reconstructed blocks. OBMC acts on such prediction blocks, helping to reduce blocking artifacts and improve prediction accuracy.
[0074] OBMC is typically applied to inter-frame prediction blocks and has certain restrictions. For example, the inter-frame prediction block cannot be predicted using affine techniques; the inter-frame prediction block cannot be predicted using angular weighted prediction (AWP) techniques; the reference frames of the inter-frame prediction block cannot all be forward reference frames; and the inter-frame prediction block cannot be smaller than 64 pixels. The following describes the implementation of OBMC in detail with reference to Figure 3.
[0075] OBMC determines the co-located prediction block of the current block based on the MV of the surrounding adjacent blocks. Then, the co-located prediction block is blended with the current prediction block to offset the block effect and prediction error. For example, as shown in Figure 3, A and B in Figure 3 are adjacent prediction blocks of the current prediction block, where A can be called the upper adjacent block and B can be called the left adjacent block. Based on the MV of the upper adjacent block, the reference block of the upper adjacent block is obtained, and then the co-located prediction block Ptop corresponding to the fusion area is found based on the reference block. Based on the MV of the left adjacent block, the reference block of the left adjacent block is obtained, and then the co-located prediction block Pleft corresponding to the fusion area is found based on the reference block. The prediction block of the current block is Pcur.
[0076] The current prediction block and the co-located prediction block are fused based on the fusion formula. Taking the fusion of Pcur and Ptop as an example, the fusion formula is shown in formula (1).
[0077] in, is the weight coefficient used for fusion, The value of can be calculated according to the following formula (2):
[0078] In formula (2), l=0,1,…..,k-1,ω * is a predefined constant. *Can be used to control the dynamic range of the weight coefficient. In related art, ω * The value of can be 1 / 8 or 3 / 8. k is the height of the fusion area of the prediction block, indicating that the k rows of pixels adjacent to the upper boundary of the current prediction block are fused with the k rows of pixels adjacent to the upper boundary of the same-position prediction block.
[0079] The fusion method of Pleft and Pcur can be the same as the fusion method of Ptop and Pcur mentioned above, that is: the first step is to generate the co-located prediction block Pleft of the current block according to the MV of B; the second step is to fuse Pleft (the k columns of pixels adjacent to the left boundary of the co-located prediction block) with Pcur.
[0080] It should be noted that if the current block is not a sub-block, the fusion area formed by the k rows of pixels is 1 / 2 of the height of the current block; if the current block is a sub-block, the fusion area formed by the k rows of pixels is the same as the height of the current block. In the related art, prediction modes with sub-block modes include one or more of the following: Affine, motion vector angle prediction (MVAP), enhanced temporal motion vector prediction (ETMVP), and subblock-based temporal motion vector prediction (SbTMVP).
[0081] In the specific implementation process of OBMC, there is a sequence when the co-located prediction blocks at different positions are fused with the current prediction block. That is to say, Ptop and Pleft are not fused with Pcur at the same time. Specifically, in the implementation process of OBMC, the prediction mode of the sample points of the upper adjacent block of the current block can be checked with 4 sample points (or pixels) as a unit. Then, all non-inter-frame prediction sample points are skipped, the continuous inter-frame prediction sample points are recorded, and OBMC is only implemented on these inter-frame prediction sample points. The upper adjacent block of the current block is determined based on the above inter-frame prediction sample points, and the co-located prediction block Ptop is obtained based on the MV of the upper adjacent block. Finally, the updated prediction block Pcur is obtained by fusion with Pcur according to the formula (1) above. top .
[0082] Next, the prediction mode of the sample points of the left adjacent block is checked in units of 4 sample points, all non-inter-frame prediction sample points are skipped, and the continuous inter-frame prediction sample points are recorded. OBMC is performed only on these inter-frame prediction sample points. The left adjacent block of the current block is determined based on the above inter-frame prediction sample points, and the co-located prediction block Pleft is obtained based on the MV of the left adjacent block. Then, according to the above formula (1) and Pcur top Fusion is performed to obtain the final output prediction block Pcur final .
[0083] OBMC can be used on both the encoder and decoder sides. It operates on the inter-frame coding portion of the hybrid video coding framework. As shown in Figure 4a, on the encoder side, OBMC can be applied to the motion compensation and motion estimation modules. On the decoder side, as shown in Figure 4b, OBMC can be applied to the motion compensation module.
[0084] As mentioned previously, OBMC is a commonly used motion compensation technique. Effective application of OBMC can eliminate blocking artifacts and reduce prediction errors. However, its use can sometimes degrade codec quality. This is because OBMC isn't suitable for all video codec scenarios. For example, OBMC isn't ideal for video with screen content. This is discussed in detail below.
[0085] Traditional general-purpose videos are primarily captured in nature. These videos inherently contain white noise, and the presence of various complex objects in the natural world makes the video content transition more natural. OBMC can be applied effectively to these videos.
[0086] Unlike traditional general-purpose videos, screen content videos are mostly computer desktop recordings or gaming videos. These videos are characterized by pixel-level sharpness and very pronounced contrast. Therefore, applying OBMC to the encoding and decoding of such videos may be counterproductive. For example, consider the screen content image shown in Figure 5. Assume that the rectangular area containing the word "colour" (boxed in Figure 5) represents the current block to be encoded. The word "e same" (boxed above it) represents the block above the current block. The word "e" (boxed to its left) represents the block to the left of the current block. If motion compensation is performed on the word "colour" within the box using OBMC, it will be merged with the co-located predicted block derived from the MV of "e same." This may blur the previously clear "colour," resulting in degraded encoding and decoding quality.
[0087] To address the above problem, an embodiment of the present application provides an encoding method, comprising: determining a motion vector of a current block; determining a first prediction block of the current block based on the motion vector of the current block; and performing overlapped block motion compensation on the first prediction block if a first condition is met.
[0088] In addition, an embodiment of the present application also provides a decoding method, including: parsing a code stream to determine a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; and performing overlapping block motion compensation on the first prediction block if a first condition is met.
[0089] The embodiments of the present application impose certain conditions on the use of OBMC, thereby helping to improve the quality of encoding and decoding. Taking screen content video as an example, if the current video is determined to be screen content video, overlapping block motion compensation can be omitted, thereby helping to improve the quality of encoding and decoding.
[0090] The following is a detailed explanation of the decoding method provided in the embodiment of the present application with reference to FIG6 .
[0091] Figure 6 is a flowchart of a decoding method provided by an embodiment of the present application. The method of Figure 6 can be applied to a decoder.
[0092] Referring to Figure 6, in step S610, the code stream is parsed to determine the motion vector of the current block. The current block may refer to the current block to be decoded. In some implementations, the current block is a luma block. In other implementations, the current block may also be a chroma block. The current block may be an inter-frame predicted block, and the motion vector of the current block may be parsed from the code stream.
[0093] In step S620, a first prediction block of the current block is determined based on the motion vector. For example, a reference block (or matching block) of the current block in a reference frame can be determined based on the motion vector, and the reference block is used as the first prediction block of the current block.
[0094] In step S630, if the first condition is satisfied, overlapping block motion compensation is performed on the first prediction block. Alternatively, if the first condition is not satisfied, overlapping block motion compensation is not performed on the first prediction block. For example, a first co-located prediction block of the current block can be determined based on the motion vector of the upper adjacent block of the current block; a second prediction block can be determined based on the first co-located prediction block and the first prediction block (or, the first co-located prediction block and the first prediction block can be fused, and the specific fusion method can be referred to the formula (1) in the previous text); a second co-located prediction block of the current block is determined based on the motion vector of the left adjacent block of the current block; a third prediction block is determined based on the second co-located prediction block and the second prediction block (or, the second co-located prediction block and the second prediction block can be fused, and the fusion method of the two is similar to the fusion method of the first co-located prediction block and the first prediction block). Further, in some implementations, after obtaining the third prediction block, a reconstructed block of the current block can be determined based on the residual block of the third prediction block and the current block.
[0095] The embodiments of the present application impose certain conditions on the use of OBMC, thereby helping to improve the encoding and decoding quality.
[0096] The embodiment of the present application does not specifically limit the type of the first condition, and can be set according to actual needs. Several possible implementations of the first condition are given below.
[0097] Implementation method 1: The first condition is: associated with the first identification information
[0098] The first identification information can be obtained based on the parsed code stream. The first identification information can be used to indicate that the current frame is a screen content coding frame; and / or, the first identification information can be used to indicate whether the current block uses overlapping block motion compensation. In some implementations, the first identification information may include a first value and a second value. The first value indicates that the current frame is a screen content coding frame and / or the current frame or the current block does not use overlapping block motion compensation; the second value indicates that the current frame is not a screen content coding frame and / or the current frame or the current block uses overlapping block motion compensation. The first condition may be that the value of the first identification information is not the first value. That is, if the value of the first indication information is not the first value, overlapping block motion compensation is performed on the first prediction block.
[0099] As discussed above, performing overlapping block motion compensation on the prediction blocks of a screen content coding frame may degrade the prediction performance of the codec. Therefore, if the value of the first identification information indicates that the current frame is a screen content coding frame, the first condition is not met, and overlapping block motion compensation is not performed on the first prediction block. This approach helps improve prediction performance during the decoding process.
[0100] The first identification information may be sequence-level identification information, for example, may be represented by ph_obmc_scc_flag (of course, the first identification information may also be represented by any other letters and / or numbers).
[0101] In some implementations, the value of ph_obmc_scc_flag can be true or false. For example, if the value of ph_obmc_scc_flag is true, it can indicate that the first condition is not satisfied and overlapped block motion compensation is not performed on the first predicted block of the current block; if the value of ph_obmc_scc_flag is false, it can indicate that the first condition is satisfied and overlapped block motion compensation is performed on the first predicted block of the current block.
[0102] Implementation 2: The first condition is associated with the information of the adjacent blocks of the current block
[0103] The first condition can be associated with various information about neighboring blocks of the current block. For example, the first condition can be associated with the prediction mode of the neighboring block. The neighboring block may contain multiple sub-blocks, and these sub-blocks may correspond to different coding units. In other words, there may be multiple prediction modes in these sub-blocks. Assuming that the prediction mode of M sub-blocks in the neighboring block is the target prediction mode, then the first condition is: the number of target prediction modes of the M sub-blocks is less than or equal to the first threshold. The size of the M sub-blocks mentioned here can be 4×4.
[0104] The target prediction modes mentioned above may include one or more of the following: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
[0105] For another example, the first condition can be associated with the co-located prediction block of the current block. As can be seen from the above description, the co-located prediction block can be determined based on the motion vector of the neighboring block of the current block. There may be an error between the first prediction block of the current block and the co-located prediction block. If the error between the two is too large, for example, exceeding a certain threshold (third threshold), it can be said that the similarity between the first prediction block and the co-located prediction block is low, or in other words, the correlation between the current block and the neighboring block is low. The first prediction block can include multiple sub-blocks, and the co-located prediction block can also include multiple sub-blocks. Therefore, if the error between the sub-blocks of the first prediction block and the sub-blocks of the co-located prediction block at the same corresponding position is large, for example, exceeding a certain threshold (second threshold), it can also be said that the correlation between the current block and the neighboring block is low. If overlapping block motion compensation is performed when the correlation between the current block and the neighboring block is low, the prediction accuracy may be reduced.
[0106] The first condition may be: the error between at least one pair of corresponding sub-blocks in the first prediction block and the co-located prediction block is less than or equal to a second threshold. Alternatively, the first condition may be: the error between the first prediction block and the co-located prediction block is less than or equal to a third threshold. Based on the above analysis, if the correlation between the current block and the adjacent blocks is low, then the prediction accuracy of the first prediction block may be higher when overlapping block motion compensation is not performed.
[0107] There are many ways to determine the error described above. Here are two ways to calculate the error, including the sum of absolute difference (SAD) and the sum of absolute transformed difference (SATD).
[0108] Implementation 3: The first condition is associated with the reference frame information of the current block
[0109] The first prediction block can be a bidirectional prediction block, which can include different reference frame settings. The first condition can be that one reference frame corresponding to the bidirectional prediction block is located before the current frame in the temporal domain, and the other reference frame is located after the current frame in the temporal domain. Compared to other reference frame settings, the co-located prediction block obtained using the reference frame in the first condition may be more informative and help improve the prediction accuracy of the current frame during the decoding process.
[0110] Implementation 4: The first condition is associated with the temporal level of the current frame
[0111] The current frame corresponding to the current block may be located at a different time domain level. The first condition may be that the time domain level at which the current frame is located is lower than the time domain level at which other frames are located. Assuming that the first time domain level at which the current frame is located is lower than the second time domain level at which other frames are located, the probability that the current frame serves as a reference frame for other frames will increase. Conversely, the probability that the current frame serves as a reference frame for other frames will decrease. Performing overlapping block motion compensation on these current frames at the first time domain level, while improving their prediction accuracy, also helps improve the prediction accuracy of other frames that use the current frame as a reference frame during the decoding process.
[0112] Implementation 5: The first condition is associated with the size of the current block
[0113] For different types of coded frames, different sizes can be set in the first condition. The first condition can be: when the current frame is a screen content coded frame, the size of the current block is greater than or equal to a fourth threshold. The fourth threshold mentioned here can be, for example, 16. The first condition can also be: when the current block is not a screen content coded frame, the size of the current block is greater than or equal to a fifth threshold. The fifth threshold mentioned here is different from the fourth threshold and can be, for example, 64.
[0114] In the embodiment of the present application, the first condition can be used not only to determine whether overlapped block motion compensation is performed on the first prediction block, but also to determine the overlapped block motion compensation mode when the first prediction block is fused. The following describes another decoding method according to the embodiment of the present application in detail with reference to FIG7 .
[0115] Figure 7 is a flowchart of a decoding method provided by an embodiment of the present application. The method of Figure 7 can be applied to a decoder.
[0116] Referring to FIG. 7 , in step S710 , the bitstream is parsed to determine the motion vector of the current block. The current block may refer to the current block to be decoded. In some implementations, the current block is a luma block. In other implementations, the current block may be a chroma block. The current block may be an inter-frame predicted block, and the motion vector of the current block may be parsed from the bitstream.
[0117] In step S720, a first prediction block of the current block is determined based on the motion vector. For example, a reference block (or matching block) of the current block in a reference frame can be determined based on the motion vector, and the reference block is used as the first prediction block of the current block.
[0118] In step S730, a co-located prediction block for the current block is determined based on the motion vectors of the neighboring blocks of the current block. The method for determining the co-located prediction block for the current block may include, for example, obtaining a reference block of the neighboring block in a reference frame based on the motion vectors of the neighboring blocks of the current block. Then, a reconstructed block corresponding to the position of the current block is found around the reference block, and the reference block is used as the co-located prediction block for the current block.
[0119] The first condition can be used to determine the overlapped block motion compensation mode. Different overlapped block motion compensation modes can be set based on whether the first condition is met.
[0120] In step 740, if the first condition is met, the first prediction block is fused with the co-located prediction block based on a first overlapped block motion compensation mode. The first overlapped block motion compensation mode may correspond to a first fusion region. In the first overlapped block motion compensation mode, a fusion weight of the first prediction block may be set to a first weight.
[0121] Alternatively, in step 740, if the first condition is not satisfied, the first prediction block is fused with the co-located prediction block based on a second overlapped block motion compensation mode. The second overlapped block motion compensation mode may correspond to a second fusion region. In the second overlapped block motion compensation mode, a fusion weight of the first prediction block may be set to the second weight.
[0122] The method of fusing the first prediction block with the co-located prediction block may include, for example, performing weighted average calculation on the pixel values of the first prediction block and the pixel values of the co-located prediction block. The weighted average calculation formula may be, for example, formula (3).
[0123] The blending region mentioned above may refer to the region within the first prediction block where motion compensation is performed. The blending region may be associated with the coefficient k in formula (4). Alternatively, the blending region may refer to the k rows or k columns within the first prediction block where overlapped block motion compensation is performed. For example, the blending region may be used to determine the coefficient k in formula (4).
[0124] The fusion weight mentioned above may refer to the weight obtained by weighted averaging the first prediction block and the co-located prediction block. The fusion weight may be combined with the coefficient in formula (3) For example, the fusion weights can be used to determine the coefficients in formula (3)
[0125] The fusion regions and weights mentioned above can be used in combination. For example, the first fusion region and the second fusion region can be set to different sizes. In another example, the first weight and the second weight can be set to different values. In another example, the first fusion region and the second fusion region can be set to different sizes, and the first weight and the second weight can be set to different values.
[0126] It should be noted that the first condition shown in FIG7 may include all of the first conditions shown in FIG6 above, or may also include part of the first conditions shown in FIG6 above. The first condition introduced in FIG6 will not be repeated here. In addition, the first condition shown in FIG7 may also be: the first identification information can be used to indicate that the current frame uses a default overlapping block motion compensation fusion area or fusion weight. The default overlapping area mentioned here can be, for example, the size of the current block or half the size of the current block; the default fusion weight can be, for example, the weight coefficient shown in formula (2) above.
[0127] In the decoding method provided in the embodiment of the present application, different overlapping block motion compensation modes can be set according to whether the first condition is met, so that the first prediction blocks can be flexibly merged according to different situations to improve decoding performance.
[0128] The fusion area and weight mentioned above can be set according to actual needs. The following are several possible implementations of the fusion area and weight.
[0129] Implementation method 1: The first fusion area is smaller than the second fusion area, and the first weight is the same as the second weight.
[0130] In the decoding method shown in FIG7 , when the first condition is not satisfied, the first blending area corresponding to the first overlapped block motion compensation mode may be set to be smaller than the second blending area corresponding to the second overlapped block motion compensation mode.
[0131] As an example, let's take Formula (3) and Formula (4) above as examples. The second fusion area can correspond to the k rows in Formula (3), and the first fusion area can correspond to the k-1 rows. Although for Formula (3), the result of the weighted average calculation of Formula (3) does not change when the first condition is met or not met. However, since the fusion area is reduced from k rows to k-1 rows, the proportion of the first prediction block in the final fused prediction block is correspondingly increased.
[0132] If the first condition is not met, the correlation between the current block and the adjacent blocks may be relatively low. Therefore, when fusing the first prediction block and the co-located prediction block, the proportion of the first prediction block in the final fused prediction block can be increased by reducing the fusion area, which helps improve the prediction accuracy during the decoding process.
[0133] Implementation method two: the first weight is smaller than the second weight, and the first fusion area is the same as the second fusion area.
[0134] In the decoding method shown in FIG. 7 , when the first condition is not satisfied, the first weight corresponding to the first overlapped block motion compensation mode may be set to be smaller than the second weight corresponding to the second overlapped block motion compensation mode.
[0135] As an example, take the above formula (3) and formula (4) as an example. The second weight can correspond to the formula (3) The first weight can correspond to and Less than Although for formula (3), the fusion area does not change when the first condition is met or not. However, the fusion weight is determined by Reduce to The result of the weighted average calculation based on formula (3) will decrease, and the proportion of the first prediction block in the final fused prediction block will increase accordingly.
[0136] If the first condition is not met, the correlation between the current block and the adjacent blocks may be relatively low. Therefore, when fusing the first prediction block and the co-located prediction block, the fusion weight can be reduced to increase the proportion of the first prediction block in the final fused prediction block, which helps improve the prediction accuracy during the decoding process.
[0137] Implementation method three: the first fusion area is smaller than the second fusion area, and the first weight is smaller than the second weight.
[0138] In the decoding method shown in Figure 7, when the first condition is not met, the first fusion area corresponding to the first overlapping block motion compensation mode can be set to be smaller than the second fusion area corresponding to the second overlapping block motion compensation mode, and the first weight corresponding to the first overlapping block motion compensation mode can be set to be smaller than the second weight corresponding to the second overlapping block motion compensation mode.
[0139] As an example, take the above formula (3) and formula (4) as an example. The second fusion area can correspond to the k rows in formula (4), and the first fusion area can correspond to the k-1 rows. The second weight can correspond to the The first weight can correspond to and Less than The fusion area is reduced from k rows to k-1 rows, and the fusion weight is reduced from Reduce to The result of the weighted average calculation based on formula (3) will be reduced, which can increase the proportion of the first prediction block in the final fused prediction block.
[0140] If the first condition is not met, the correlation between the current block and the adjacent blocks may be relatively low. Therefore, when fusing the first prediction block and the co-located prediction block, the proportion of the first prediction block in the final fused prediction block can be increased by changing the fusion area and fusion weight, which helps improve the prediction accuracy during the decoding process.
[0141] The decoding method provided by the embodiment of the present application is described in detail above in conjunction with Figures 6 and 7. The encoding method provided by the embodiment of the present application is described in detail below in conjunction with Figures 8 and 9.
[0142] Figure 8 is a flow chart of an encoding method provided by an embodiment of the present application. The method of Figure 8 can be applied to an encoder.
[0143] Referring to FIG8 , in step S8010, the motion vector of the current block is determined. The current block may refer to the current block to be encoded. In some implementations, the current block is a luma block. In other implementations, the current block may be a chroma block. The current block may be an inter-frame predicted block, and the motion vector of the current block may be parsed from the bitstream.
[0144] In step S8020, a first prediction block of the current block is determined based on the motion vector. For example, a reference block (or matching block) of the current block in a reference frame can be determined based on the motion vector, and the reference block is used as the first prediction block of the current block.
[0145] In step S8030, if the first condition is satisfied, overlapping block motion compensation is performed on the first prediction block. Alternatively, if the first condition is not satisfied, overlapping block motion compensation is not performed on the first prediction block. For example, a first co-located prediction block of the current block can be determined based on the motion vector of the upper adjacent block of the current block; a second prediction block can be determined based on the first co-located prediction block and the first prediction block (or, the first co-located prediction block and the first prediction block can be fused, and the specific fusion method can be referred to the formula (3) in the previous text); a second co-located prediction block of the current block is determined based on the motion vector of the left adjacent block of the current block; a third prediction block is determined based on the second co-located prediction block and the second prediction block (or, the second co-located prediction block and the second prediction block can be fused, and the fusion method of the two is similar to the fusion method of the first co-located prediction block and the first prediction block). Further, in some implementations, after obtaining the third prediction block, a reconstructed block of the current block can be determined based on the residual block of the third prediction block and the current block.
[0146] The embodiments of the present application impose certain conditions on the use of OBMC, thereby helping to improve the encoding and decoding quality.
[0147] The embodiment of the present application does not specifically limit the type of the first condition, and can be set according to actual needs. Several possible implementations of the first condition are given below.
[0148] Implementation method 1: The first condition is: associated with the first identification information
[0149] The first identification information can be obtained based on the parsed code stream. The first identification information can be used to indicate that the current frame is a screen content coding frame; and / or, the first identification information can be used to indicate whether the current block uses overlapping block motion compensation. In some implementations, the first identification information may include a first value and a second value. The first value indicates that the current frame is a screen content coding frame and / or the current frame or the current block does not use overlapping block motion compensation; the second value indicates that the current frame is not a screen content coding frame and / or the current frame or the current block uses overlapping block motion compensation. The first condition may be that the value of the first identification information is not the first value. That is, if the value of the first indication information is not the first value, overlapping block motion compensation is performed on the first prediction block.
[0150] As previously explained, performing overlapping block motion compensation on the prediction blocks of a screen content coding frame may reduce prediction performance. Therefore, if the value of the first identification information indicates that the current frame is a screen content coding frame, the first condition is not met, and overlapping block motion compensation is not performed on the first prediction block. This approach helps improve prediction performance during the encoding process.
[0151] The first identification information may be sequence-level identification information, for example, may be represented by ph_obmc_scc_flag (of course, the first identification information may also be represented by any other letters and / or numbers). The first identification information may be written into the bitstream.
[0152] In some implementations, the value of ph_obmc_scc_flag can be true or false. For example, if the value of ph_obmc_scc_flag is true, it can indicate that the first condition is not satisfied and overlapped block motion compensation is not performed on the first predicted block of the current block; if the value of ph_obmc_scc_flag is false, it can indicate that the first condition is satisfied and overlapped block motion compensation is performed on the first predicted block of the current block.
[0153] Implementation method 2: The first condition is associated with information of adjacent blocks of the current block.
[0154] The first condition can be associated with various information about neighboring blocks of the current block. For example, the first condition can be associated with the prediction mode of the neighboring block. The neighboring block may contain multiple sub-blocks, and these sub-blocks may correspond to different coding units. In other words, there may be multiple prediction modes in these sub-blocks. Assuming that the prediction mode of M sub-blocks in the neighboring block is the target prediction mode, then the first condition is: the number of target prediction modes of the M sub-blocks is less than or equal to the first threshold. The size of the M sub-blocks mentioned here can be 4×4.
[0155] The target prediction modes mentioned above may include one or more of the following: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
[0156] For another example, the first condition can be associated with the co-located prediction block of the current block. As can be seen from the above description, the co-located prediction block can be determined based on the motion vector of the neighboring block of the current block. There may be an error between the first prediction block of the current block and the co-located prediction block. If the error between the two is too large, for example, exceeding a certain threshold (third threshold), it can be said that the similarity between the first prediction block and the co-located prediction block is low, or in other words, the correlation between the current block and the neighboring block is low. The first prediction block can include multiple sub-blocks, and the co-located prediction block can also include multiple sub-blocks. Therefore, if the error between the sub-blocks of the first prediction block and the sub-blocks of the co-located prediction block at the same corresponding position is large, for example, exceeding a certain threshold (second threshold), it can also be said that the correlation between the current block and the neighboring block is low. If overlapping block motion compensation is performed when the correlation between the current block and the neighboring block is low, the prediction accuracy may be reduced.
[0157] The first condition may be: the error between at least one pair of corresponding sub-blocks in the first prediction block and the co-located prediction block is less than or equal to a second threshold. Alternatively, the first condition may be: the error between the first prediction block and the co-located prediction block is less than or equal to a third threshold. Based on the above analysis, if the correlation between the current block and the adjacent blocks is low, then the prediction accuracy of the first prediction block may be higher when overlapping block motion compensation is not performed.
[0158] There are many ways to determine the error described above. Here are two ways to calculate the error: sum of absolute difference (SAD) and sum of absolute transformed difference (SATD).
[0159] Implementation method 3: The first condition is associated with the reference frame information of the current block.
[0160] The first prediction block can be a bidirectional prediction block, which can include different reference frame settings. The first condition can be that one reference frame corresponding to the bidirectional prediction block is located before the current frame in the temporal domain, and the other reference frame is located after the current frame in the temporal domain. Compared to other reference frame settings, the co-located prediction block obtained using the reference frame in the first condition may be more informative and help improve the prediction accuracy of the current frame during the encoding process.
[0161] Implementation method 4: The first condition is associated with the time domain level of the current frame.
[0162] The current frame corresponding to the current block may be located at different time domain levels. The first condition may be that the time domain level at which the current frame is located is lower than the time domain level at which other frames are located. Assuming that the first time domain level at which the current frame is located is lower than the second time domain level at which other frames are located, the probability that the current frame serves as a reference frame for other frames will increase. Conversely, the probability that the current frame serves as a reference frame for other frames will decrease. Performing overlapping block motion compensation on these current frames at the first time domain level, while improving their prediction accuracy, also helps improve the prediction accuracy of other frames that use the current frame as a reference frame during the encoding process.
[0163] Implementation 5: The first condition is associated with the size of the current block.
[0164] For different types of coded frames, different sizes can be set in the first condition. The first condition can be: when the current frame is a screen content coded frame, the size of the current block is greater than or equal to a fourth threshold. The fourth threshold mentioned here can be, for example, 16. The first condition can also be: when the current block is not a screen content coded frame, the size of the current block is greater than or equal to a fifth threshold. The fifth threshold mentioned here is different from the fourth threshold and can be, for example, 64.
[0165] In the embodiment of the present application, the first condition can be used not only to determine whether overlapped block motion compensation is performed on the first prediction block, but also to determine the overlapped block motion compensation mode when the first prediction block is merged. The following describes another encoding method according to the embodiment of the present application in detail with reference to FIG9 .
[0166] Figure 9 is a flow chart of an encoding method provided by an embodiment of the present application. The method of Figure 9 can be applied to an encoder.
[0167] Referring to FIG. 9 , in step S9010, the motion vector of the current block is determined. The current block may refer to the current block to be encoded. In some implementations, the current block is a luma block. In other implementations, the current block may be a chroma block. The current block may be an inter-frame predicted block, and the motion vector of the current block may be parsed from the bitstream.
[0168] In step S9020, a first prediction block of the current block is determined based on the motion vector. For example, a reference block (or matching block) of the current block in a reference frame can be determined based on the motion vector, and the reference block is used as the first prediction block of the current block.
[0169] In step S9030, a co-located prediction block for the current block is determined based on the motion vectors of the neighboring blocks of the current block. The method for determining the co-located prediction block for the current block may include, for example, obtaining a reference block of the neighboring block in a reference frame based on the motion vectors of the neighboring blocks of the current block. Then, a reconstructed block corresponding to the position of the current block is found around the reference block, and the reference block is used as the co-located prediction block for the current block.
[0170] The first condition can be used to determine the overlapped block motion compensation mode. Different overlapped block motion compensation modes can be set based on whether the first condition is met.
[0171] In step 9040, if the first condition is met, the first prediction block is fused with the co-located prediction block based on a first overlapped block motion compensation mode. The first overlapped block motion compensation mode may correspond to a first fusion region. In the first overlapped block motion compensation mode, a fusion weight of the first prediction block may be set to a first weight.
[0172] Alternatively, in step 9040, if the first condition is not satisfied, the first prediction block is fused with the co-located prediction block based on a second overlapped block motion compensation mode. The second overlapped block motion compensation mode may correspond to a second fusion region. In the second overlapped block motion compensation mode, a fusion weight of the first prediction block may be set to the second weight.
[0173] The method of fusing the first prediction block with the co-located prediction block may include, for example, performing weighted average calculation on the pixel values of the first prediction block and the pixel values of the co-located prediction block. The weighted average calculation formula may be, for example, formula (5).
[0174] The blending region mentioned above may refer to the region within the first prediction block where motion compensation is performed. The blending region may be associated with the coefficient k in formula (6). Alternatively, the blending region may refer to the k rows or k columns within the first prediction block where overlapped block motion compensation is performed. For example, the blending region may be used to determine the coefficient k in formula (6).
[0175] The fusion weight mentioned above may refer to the weight obtained by weighted averaging the first prediction block and the co-located prediction block. The fusion weight may be combined with the coefficient in formula (5) For example, the fusion weights can be used to determine the coefficients in formula (5)
[0176] The fusion regions and weights mentioned above can be used in combination. For example, the first fusion region and the second fusion region can be set to different sizes. In another example, the first weight and the second weight can be set to different values. In another example, the first fusion region and the second fusion region can be set to different sizes, and the first weight and the second weight can be set to different values.
[0177] It should be noted that the first condition shown in FIG9 may include all of the first conditions shown in FIG8 above, or may also include part of the first conditions shown in FIG8 above. The first condition introduced in FIG8 will not be repeated here. In addition, the first condition shown in FIG9 may also be: the first identification information can be used to indicate that the current frame uses a default overlapping block motion compensation fusion area or fusion weight. The default overlapping area mentioned here can be, for example, the size of the current block or half the size of the current block; the default fusion weight can be, for example, the weight coefficient shown in formula (2) above.
[0178] In the encoding method provided in the embodiment of the present application, different overlapping block motion compensation modes can be set according to whether the first condition is met, so that the first prediction blocks can be flexibly merged according to different situations to improve encoding performance.
[0179] The fusion area and weight mentioned above can be set according to actual needs. The following are several possible implementations of the fusion area and weight.
[0180] Implementation method 1: The first fusion area is smaller than the second fusion area, and the first weight is the same as the second weight.
[0181] In the encoding method shown in FIG9 , when the first condition is not satisfied, the first blending area corresponding to the first overlapped block motion compensation mode may be set to be smaller than the second blending area corresponding to the second overlapped block motion compensation mode.
[0182] As an example, let's take Formula (5) and Formula (6) above as examples. The second fusion area can correspond to the k rows in Formula (6), and the first fusion area can correspond to the k-1 rows. Although for Formula (5), the result of the weighted average calculation of Formula (5) does not change when the first condition is met or not met. However, since the fusion area is reduced from k rows to k-1 rows, the proportion of the first prediction block in the final fused prediction block is correspondingly increased.
[0183] If the first condition is not met, the correlation between the current block and the adjacent blocks may be relatively low. Therefore, when fusing the first prediction block and the co-located prediction block, the proportion of the first prediction block in the final fused prediction block can be increased by reducing the fusion area, which helps improve the prediction accuracy during the encoding process.
[0184] Implementation method two: the first weight is smaller than the second weight, and the first fusion area is the same as the second fusion area.
[0185] In the encoding method shown in FIG9 , when the first condition is not satisfied, the first weight corresponding to the first overlapped block motion compensation mode may be set to be smaller than the second weight corresponding to the second overlapped block motion compensation mode.
[0186] As an example, take the above formula (5) and formula (6) as an example. The second weight can correspond to the formula (5) The first weight can correspond to and Less than Although for formula (5), the fusion area does not change when the first condition is met or not. However, the fusion weight is determined by Reduce to The result of the weighted average calculation based on formula (5) will decrease, and the proportion of the first prediction block in the final fused prediction block will increase accordingly.
[0187] If the first condition is not met, the correlation between the current block and the adjacent blocks may be relatively low. Therefore, when fusing the first prediction block and the co-located prediction block, the proportion of the first prediction block in the final fused prediction block can be increased by reducing the fusion weight, which helps improve the prediction accuracy during the encoding process.
[0188] Implementation method three: the first fusion area is smaller than the second fusion area, and the first weight is smaller than the second weight.
[0189] In the encoding method shown in Figure 9, when the first condition is not met, the first fusion area corresponding to the first overlapping block motion compensation mode can be set to be smaller than the second fusion area corresponding to the second overlapping block motion compensation mode, and the first weight corresponding to the first overlapping block motion compensation mode can be set to be smaller than the second weight corresponding to the second overlapping block motion compensation mode.
[0190] As an example, take the above formula (5) and formula (6) as an example. The second fusion area can correspond to the k rows in formula (6), and the first fusion area can correspond to the k-1 rows. The second weight can correspond to the The first weight can correspond to and Less than The fusion area is reduced from k rows to k-1 rows, and the fusion weight is reduced from Reduce to The result of the weighted average calculation based on formula (5) will be reduced, which can increase the proportion of the first prediction block in the final fused prediction block.
[0191] If the first condition is not met, the correlation between the current block and the adjacent blocks may be relatively low. Therefore, when fusing the first prediction block and the co-located prediction block, the proportion of the first prediction block in the final fused prediction block can be increased by changing the fusion area and fusion weight, which helps improve the prediction accuracy during the encoding process.
[0192] The above describes the decoding method and encoding method provided by the embodiments of the present application from the perspectives of the decoding end and the encoding end, respectively. The following describes two more specific embodiments of the present application in conjunction with the encoding and decoding end.
[0193] Example 1
[0194] Encoding end
[0195] In this embodiment, at the encoding end, the encoder traverses the prediction mode. If the current prediction mode type is the inter-frame prediction mode, a flag is obtained that allows use. This flag is a sequence-level flag, indicating that the current encoder allows the use of OBMC technology, which can be in the form of sqh_obmc_enable_flag.
[0196] Step 1: If the OBMC usage permission flag is true, calculate whether the current frame is a screen content coded image. If so, set the frame-level OBMC screen content coding flag (ph_obmc_scc_flag) to true, otherwise set it to false. If the area of the current coding unit is greater than threshold 1, and the current coding unit is not predicted using Affine technology, and the current coding unit is not predicted using AWP technology, and the current frame is not a P frame, then use OBMC technology, that is, execute step 2; if the OBMC usage permission flag or other conditions such as area conditions are not met, then skip OBMC technology, that is, skip step 2 and directly execute step 3.
[0197] Step 2: If ph_obmc_scc_flag is true, the current coding unit does not use the OBMC technology; if ph_obmc_scc_flag is false, the prediction block of the current coding block is obtained according to the prediction mode information obtained by parsing.
[0198] Check the mode attributes of the upper side of the current block with 4 sample points as a unit, skip all non-inter-frame prediction sample points, record the continuous inter-frame prediction sample points, and only perform OBMC overlap compensation on these inter-frame prediction sample points. According to the motion vector of the adjacent minimum unit inter-frame prediction mode part, the upper side co-located prediction block Ptop is obtained, and the updated Pcur is obtained by merging it with Pcur according to the weight coefficient. topSimilarly, check the mode attributes of the left side of the current block with 4 sample points as a unit, skip all non-inter-frame prediction sample points, record the continuous inter-frame prediction sample points, and only perform OBMC overlap compensation on these inter-frame prediction sample points. According to the motion vector of the adjacent minimum unit inter-frame prediction mode part, the left co-located prediction block Pleft is obtained, and the weight coefficient and Pcur are used to calculate the overlap compensation. top Fusion is performed to obtain the final Pcur final .
[0199] In step 3, the encoder continues to traverse other inter-frame prediction technologies, calculates the rate-distortion cost value corresponding to each technology, and selects the prediction mode corresponding to the minimum cost value as the optimal prediction mode for the current coding unit.
[0200] Step 4: After traversing all coding units, the code stream is output after loop filtering, entropy coding and other technologies.
[0201] Decoding end
[0202] Step 1: The decoder parses and obtains the OBMC enable flag, which is a sequence-level flag (sqh_obmc_enable_flag), indicating that the current decoder allows the use of OBMC technology. If sqh_obmc_enable_flag is true, ph_obmc_scc_flag is parsed and step 2 is executed. If sqh_obmc_enable_flag is false, the decoder does not allow the use of OBMC. By default, all coding units in the bitstream do not use OBMC technology, and step 3 is executed.
[0203] Step 2: If the current block is in inter prediction mode, not Affine prediction mode, and not AWP prediction mode, and the area of the current coding unit is greater than the threshold 1, and the current frame is not a P frame, and ph_obmc_scc_flag is false, then the current coding unit uses OBMC technology for overlap compensation.
[0204] Otherwise, the current coding unit does not use the OBMC technology, skip the remaining content, and execute step 3.
[0205] According to the prediction mode information obtained by parsing, the prediction block of the current coding block is obtained. The mode attributes of the adjacent blocks above the current block are checked in units of 4 sample points, all non-inter-frame prediction sample points are skipped, and the continuous inter-frame prediction sample points are recorded. OBMC is performed only on these inter-frame prediction sample points. The upper co-located prediction block Ptop is obtained based on the motion vector of the above inter-frame prediction sample point, and the updated Pcur is obtained by fusing it with Pcur according to the weight coefficient. topSimilarly, check the mode attributes of the left side of the current block with 4 sample points as a unit, skip all non-inter-frame prediction sample points, record the continuous inter-frame prediction sample points, and only perform OBMC on these inter-frame prediction sample points. According to the motion vector of the above inter-frame prediction sample point, the left co-located prediction block Pleft is obtained, and the weight coefficient and Pcur are used to calculate the left co-located prediction block Pleft. top Fusion is performed to obtain the final Pcur final .
[0206] Step 3: The decoder continues to traverse other prediction modes and predicts the final prediction block for the current coding unit. If there is no inactive prediction mode, the current prediction block is the final prediction block.
[0207] Step 4: parse the bitstream and obtain residual information, and obtain time domain residual information through inverse quantization and inverse transformation. The final prediction block is superimposed with the time domain residual information to obtain a reconstructed sample block.
[0208] In step 5, all reconstructed sample blocks are processed through loop filtering and other techniques to obtain the final reconstructed image, which can be used as both video output and as a reference for subsequent decoding.
[0209] Example 2
[0210] Encoding end
[0211] The encoder traverses the prediction mode. If the current prediction mode type is the inter-frame prediction mode, it obtains the flag bit that allows use. This flag bit is a sequence-level flag bit, indicating that the current encoder allows the use of OBMC technology, which can be in the form of sqh_obmc_enable_flag.
[0212] In step 1, if the OBMC usage permission flag is true, the area of the current coding unit is greater than threshold 1, the current coding unit is not predicted by Affine technology, the current coding unit is not predicted by AWP technology, and the current frame is not a P frame, then the OBMC technology is used, that is, step 2 is executed; if the OBMC usage permission flag or other conditions such as area area are not met, the OBMC technology is skipped, that is, step 2 is skipped and step 3 is executed directly.
[0213] Step 2: Based on the prediction mode information obtained from the analysis, the prediction block of the current coding block is obtained. A prediction mode check is performed on the upper adjacent block, using four sample points as a basic unit. If the number of intra block copy or string matching prediction modes in the upper adjacent block exceeds half of the total number of basic units on the upper side, OBMC technology is not allowed to be used in the upper part of the current coding block; otherwise, OBMC technology is allowed on the upper side.
[0214] Similarly, a prediction mode check is performed on the left adjacent block with 4 sample points as a basic unit. If the number of intra-block copy or string matching prediction modes in the left adjacent block exceeds half of the total number of basic units on the left, the OBMC technology is not allowed to be used in the left part of the current coding block; otherwise, the OBMC technology is allowed to be used on the left.
[0215] If the upper side allows the use of OBMC technology, the mode attributes of the upper side of the current block are checked in units of 4 sample points, all non-inter-frame prediction sample points are skipped, and continuous inter-frame prediction sample points are recorded. OBMC overlap compensation is performed only on these inter-frame prediction sample points. The upper side co-located prediction block Ptop is obtained based on the motion vector of the adjacent minimum unit inter-frame prediction mode part, and the updated Pcur is obtained by merging it with Pcur according to the weight coefficient. top ;
[0216] If the OBMC technology is allowed on the left side, the mode attributes of the left side of the current block are also checked with 4 sample points as a unit, all non-inter-frame prediction sample points are skipped, and continuous inter-frame prediction sample points are recorded. OBMC overlap compensation is performed only on these inter-frame prediction sample points. The left co-located prediction block Pleft is obtained according to the motion vector of the adjacent minimum unit inter-frame prediction mode part, and the weight coefficient and Pcur are used to calculate the overlap compensation. top Fusion is performed to obtain the final Pcur final .
[0217] In step 3, the encoder continues to traverse other inter-frame prediction technologies, calculates the rate-distortion cost value corresponding to each technology, and selects the prediction mode corresponding to the minimum cost value as the optimal prediction mode for the current coding unit.
[0218] Step 4: After traversing all coding units, the code stream is output after loop filtering, entropy coding and other technologies.
[0219] Decoding end
[0220] The decoder parses and obtains the OBMC permission flag, which is a sequence-level flag (sqh_obmc_enable_flag), indicating that the current decoder allows the use of OBMC technology.
[0221] In step 1, if sqh_obmc_enable_flag is true, parse ph_obmc_scc_flag and proceed to step 2. If sqh_obmc_enable_flag is false, OBMC is not allowed. By default, all coding units in the code stream do not use OBMC technology. Then proceed to step 3.
[0222] Step 2: If the current block is in inter prediction mode, not Affine prediction mode, and not AWP prediction mode, and the area of the current coding unit is greater than threshold 1, and the current frame is not a P frame, then the current coding unit allows OBMC technology to perform overlap compensation.
[0223] Otherwise, the current coding unit does not use the OBMC technology, skip the remaining content, and execute step 3.
[0224] Based on the prediction mode information obtained from the analysis, the prediction block of the current coding block is obtained. The prediction mode of the upper adjacent block is checked with 4 sample points as a basic unit. If the number of intra block copy or string matching prediction modes in the upper adjacent block exceeds half of the total number of basic units in the upper part, the OBMC technology is not allowed to be used in the upper part of the current coding block; otherwise, the OBMC technology is allowed in the upper part.
[0225] Similarly, a prediction mode check is performed on the left adjacent block with 4 sample points as a basic unit. If the number of intra-block copy or string matching prediction modes in the left adjacent block exceeds half of the total number of basic units on the left, the OBMC technology is not allowed to be used in the left part of the current coding block; otherwise, the OBMC technology is allowed to be used on the left.
[0226] If the upper side allows the use of OBMC technology, the mode attributes of the upper side of the current block are checked in units of 4 sample points, all non-inter-frame prediction sample points are skipped, and continuous inter-frame prediction sample points are recorded. OBMC overlap compensation is performed only on these inter-frame prediction sample points. The left co-located prediction block Ptop is obtained based on the motion vector of the adjacent minimum unit inter-frame prediction mode part, and the updated Pcur is obtained by merging it with Pcur according to the weight coefficient. top ;
[0227] If the OBMC technology is allowed on the left side, the mode attributes of the left side of the current block are also checked with 4 sample points as a unit, all non-inter-frame prediction sample points are skipped, and continuous inter-frame prediction sample points are recorded. OBMC overlap compensation is performed only on these inter-frame prediction sample points. The left co-located prediction block Pleft is obtained according to the motion vector of the adjacent minimum unit inter-frame prediction mode part, and the weight coefficient and Pcur are used to calculate the overlap compensation. top Fusion is performed to obtain the final Pcur final .
[0228] Step 3: The decoder continues to traverse other prediction modes and predicts the current coding unit to obtain the final prediction block. If there is no unused prediction mode, the current prediction block is the final prediction block.
[0229] Step 4: parse the bitstream and obtain residual information, and obtain time domain residual information through inverse quantization and inverse transformation. The final prediction block is superimposed with the time domain residual information to obtain a reconstructed sample block.
[0230] In step 5, all reconstructed sample blocks are processed through loop filtering and other techniques to obtain the final reconstructed image, which can be used as both video output and as a reference for subsequent decoding.
[0231] The above describes the encoding and decoding method provided in the embodiment of the present application from the encoding and decoding end, and the quality of the encoding and decoding performance can be verified through testing. The following introduces the test parameters obtained by testing the encoding and decoding method provided in the embodiment of the present application.
[0232] The object of this codec test is screen content coding based on inter-frame prediction. The test results under the test conditions of random access (RA) and low delay B (LDB) are shown in Table 1 and Table 2. In the test parameters in Table 1 and Table 2, negative numbers represent performance gains, the number of bits is reduced and the compression rate is improved under the same quality. From the test results in Table 1, there is a small amount of performance improvement under the test condition RA. However, from the test results in Table 2, there is a good compression performance gain under the low delay test condition LDB. Overall, on TGM, which is the most important for screen content coding, the codec method provided by the embodiment of this application can bring performance gains of more than 0.4% for RA and more than 1.1% for LDB.
[0233] Table 1
[0234] Table 2
[0235] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 9 . The device embodiment of the present application is described in detail below in conjunction with Figures 10 to 15 . It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.
[0236] FIG10 is a schematic diagram of the structure of a decoder provided by an embodiment of the present application. As shown in FIG10 , the decoder 1000 includes: a first determining unit 1010 , a second determining unit 1020 , and a third determining unit 1030 .
[0237] The first determining unit 1010 is configured to parse the code stream and determine the motion vector of the current block.
[0238] The second determining unit 1020 is configured to determine a first prediction block of the current block according to the motion vector.
[0239] The third determining unit 1030 is configured to perform overlapped block motion compensation on the first prediction block if the first condition is met.
[0240] In some implementations, the first condition includes that the value of the first identification information parsed from the code stream is not a first value, and the first value is used to indicate one or more of the following: the current frame is a screen content coding frame; the current frame or the current block does not use overlapping block motion compensation.
[0241] In some implementations, the first condition is associated with information of neighboring blocks of the current block.
[0242] In some implementations, the first condition is associated with whether the prediction mode of the neighboring block is a target prediction mode.
[0243] In some implementations, the target prediction mode includes one or more of: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
[0244] In some implementations, the adjacent block includes M sub-blocks, and the first condition includes: the number of prediction modes corresponding to the M sub-blocks that are target prediction modes is less than or equal to a first threshold.
[0245] In some implementations, each of the M sub-blocks has a size of 4×4.
[0246] In some implementations, the neighboring blocks include a left neighboring block and / or an upper neighboring block of the current block.
[0247] In some implementations, the first condition is associated with a co-located prediction block of the current block, and the co-located prediction block is determined based on a motion vector of a neighboring block of the current block.
[0248] In some implementations, the first condition is associated with one or more of the following: an error between the first prediction block and the co-located prediction block; an error between corresponding sub-block pairs in the first prediction block and the co-located prediction block.
[0249] In some implementations, the first condition includes: an error between at least one pair of the corresponding sub-blocks is less than or equal to a second threshold.
[0250] In some implementations, the first condition includes: an error between the first prediction block and the co-located prediction block is less than or equal to a third threshold.
[0251] In some implementations, the error includes an absolute error and / or an absolute transformation error.
[0252] In some implementations, the first condition includes: the current block is a bidirectional prediction block, the first reference frame corresponding to the bidirectional prediction block is located before the current frame in the time domain, and the second reference frame corresponding to the bidirectional prediction block is located after the current frame in the time domain.
[0253] In some implementations, the first condition is associated with a temporal level of the current frame.
[0254] In some implementations, the first condition includes: the current frame belongs to a first time domain level; or, the current frame does not belong to a second time domain level; wherein the first time domain level is lower than the second time domain level.
[0255] In some implementations, the first condition includes: when the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or when the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein the fourth threshold is different from the fifth threshold.
[0256] In some implementations, the fourth threshold is 16, and / or the fifth threshold is 64.
[0257] In some implementations, the third determination unit 1030 is configured to determine a first co-located prediction block of the current block based on a motion vector of an upper neighboring block of the current block; determine a second prediction block based on the first co-located prediction block and the first prediction block; determine a second co-located prediction block of the current block based on a motion vector of a left neighboring block of the current block; and determine a third prediction block based on the second co-located prediction block and the second prediction block.
[0258] In some implementations, the decoder 1000 further includes a fourth determining unit 1040 configured to determine a reconstructed block of the current block according to the third prediction block and the residual block of the current block.
[0259] FIG11 is a schematic diagram of the structure of a decoder provided by another embodiment of the present application. As shown in FIG11 , the decoder 1100 includes: a first determining unit 1110 , a second determining unit 1120 , a third determining unit 1130 , and a fourth determining unit 1140 .
[0260] The first determining unit 1110 is configured to parse the code stream and determine the motion vector of the current block.
[0261] The second determining unit 1120 is configured to determine a first prediction block of the current block according to the motion vector.
[0262] The third determining unit 1130 is configured to determine a co-located prediction block of the current block according to the motion vector of the neighboring block of the current block.
[0263] The fourth determining unit 1140 is configured to, if a first condition is satisfied, fuse the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponding to a first fusion area, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; if the first condition is not satisfied, fuse the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponding to a second fusion area, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein the first fusion area and the second fusion area are different; and / or the first weight is different from the second weight.
[0264] In some implementations, the first fusion region is smaller than the second fusion region.
[0265] In some implementations, the first weight is less than the second weight.
[0266] In some implementations, the first condition includes that the value of the first identification information parsed from the code stream is not a first value, and the first value is used to indicate one or more of the following: the current frame is a screen content coding frame; the current frame or the current block does not use overlapping block motion compensation.
[0267] In some implementations, the first condition is associated with information of neighboring blocks of the current block.
[0268] In some implementations, the first condition is associated with whether the prediction mode of the neighboring block is a target prediction mode.
[0269] In some implementations, the target prediction mode includes one or more of: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
[0270] In some implementations, the adjacent block includes M sub-blocks, and the first condition includes: the number of prediction modes corresponding to the M sub-blocks that are target prediction modes is less than or equal to a first threshold.
[0271] In some implementations, each of the M sub-blocks has a size of 4×4.
[0272] In some implementations, the neighboring blocks include a left neighboring block and / or an upper neighboring block of the current block.
[0273] In some implementations, the first condition is associated with a co-located prediction block of the current block, and the co-located prediction block is determined based on a motion vector of a neighboring block of the current block.
[0274] In some implementations, the first condition is associated with one or more of the following: an error between the first prediction block and the co-located prediction block; an error between corresponding sub-block pairs in the first prediction block and the co-located prediction block.
[0275] In some implementations, the first condition includes: an error between at least one pair of the corresponding sub-blocks is less than or equal to a second threshold.
[0276] In some implementations, the first condition includes: an error between the first prediction block and the co-located prediction block is less than or equal to a third threshold.
[0277] In some implementations, the error includes an absolute error and / or an absolute transformation error.
[0278] In some implementations, the first condition includes: the current block is a bidirectional prediction block, the first reference frame corresponding to the bidirectional prediction block is located before the current frame in the time domain, and the second reference frame corresponding to the bidirectional prediction block is located after the current frame in the time domain.
[0279] In some implementations, the first condition is associated with a temporal level of the current frame.
[0280] In some implementations, the first condition includes: the current frame belongs to a first time domain level; or, the current frame does not belong to a second time domain level; wherein the first time domain level is lower than the second time domain level.
[0281] In some implementations, the first condition includes: when the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or when the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein the fourth threshold is different from the fifth threshold.
[0282] In some implementations, the fourth threshold is 16, and / or the fifth threshold is 64.
[0283] In some implementations, the third determination unit 1130 is configured to determine a first co-located prediction block of the current block based on a motion vector of an upper neighboring block of the current block; determine a second prediction block based on the first co-located prediction block and the first prediction block; and determine the second co-located prediction block of the current block based on a motion vector of a left neighboring block of the current block.
[0284] In some implementations, the fourth determining unit 1140 is configured to determine a third prediction block according to the second co-located prediction block and the second prediction block.
[0285] In some implementations, the decoder 1100 further includes a fifth determining unit 1150 configured to determine a reconstructed block of the current block according to the third prediction block and the residual block of the current block.
[0286] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0287] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0288] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to decoder 1000 or decoder 1100. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method in the aforementioned first embodiment.
[0289] Based on the composition of the above-mentioned decoder and the computer-readable storage medium, refer to Figure 12, which shows a specific hardware structure diagram of the decoder 1200 provided in an embodiment of the present application. As shown in Figure 12, the decoder 1200 may include: a communication interface 1210, a memory 1220 and a processor 1230; each component is coupled together through a bus system 1240. It can be understood that the bus system 1240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1240 in Figure 12. Among them,
[0290] The communication interface 1210 is used to receive and send signals when sending and receiving information with other external network elements.
[0291] The memory 1220 is used to store computer programs.
[0292] The processor 1230 is configured to, when running the computer program, execute the following steps: parsing a bitstream to determine a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; and performing overlapped block motion compensation on the first prediction block if a first condition is met.
[0293] Alternatively, the following steps are performed: parsing a code stream to determine a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; determining a co-located prediction block of the current block based on motion vectors of neighboring blocks of the current block; if a first condition is satisfied, fusing the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, wherein the first overlapping block motion compensation mode corresponds to a first fusion region, and in the first overlapping block motion compensation mode, a fusion weight of the first prediction block is a first weight; if the first condition is not satisfied, fusing the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, wherein the second overlapping block motion compensation mode corresponds to a second fusion region, and in the second overlapping block motion compensation mode, a fusion weight of the second prediction block is a second weight; wherein the first fusion region and the second fusion region are different; and / or the first weight is different from the second weight.
[0294] It is understood that the memory 1220 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1220 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0295] The processor 1230 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1230. The above-mentioned processor 1230 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1220 , and the processor 1230 reads the information in the memory 1220 and completes the steps of the above method in combination with its hardware.
[0296] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0297] Optionally, as another embodiment, the processor 1230 is further configured to execute the decoding method described in the above embodiment when running the computer program.
[0298] FIG13 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. As shown in FIG13 , the encoder 1300 includes: a first determining unit 1310 , a second determining unit 1320 , and a third determining unit 1330 .
[0299] The first determining unit 1310 is configured to determine a motion vector of a current block.
[0300] The second determining unit 1320 is configured to determine a first prediction block of the current block according to the motion vector.
[0301] The third determining unit 1330 is configured to perform overlapped block motion compensation on the first prediction block if the first condition is satisfied.
[0302] In some implementations, the first condition includes that the current frame is a screen content coding frame.
[0303] In some implementations, the encoder 1300 further includes a writing unit 1340 configured to write first identification information into a bitstream, wherein the first identification information is used to indicate one or more of the following: whether the current frame is a screen content coding frame; whether the current frame or the current block uses overlapping block motion compensation.
[0304] In some implementations, the first condition is associated with information of neighboring blocks of the current block.
[0305] In some implementations, the first condition is associated with whether the prediction mode of the neighboring block is a target prediction mode.
[0306] In some implementations, the target prediction mode includes one or more of: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
[0307] In some implementations, the adjacent block includes M sub-blocks, and the first condition includes: the number of prediction modes corresponding to the M sub-blocks that are target prediction modes is less than or equal to a first threshold.
[0308] In some implementations, each of the M sub-blocks has a size of 4×4.
[0309] In some implementations, the neighboring blocks include a left neighboring block and / or an upper neighboring block of the current block.
[0310] In some implementations, the first condition is associated with a co-located prediction block of the current block, and the co-located prediction block is determined based on a motion vector of a neighboring block of the current block.
[0311] In some implementations, the first condition is associated with one or more of the following: an error between the first prediction block and the co-located prediction block; an error between corresponding sub-block pairs in the first prediction block and the co-located prediction block.
[0312] In some implementations, the first condition includes: an error between at least one pair of the corresponding sub-blocks is less than or equal to a second threshold.
[0313] In some implementations, the first condition includes: an error between the first prediction block and the co-located prediction block is less than or equal to a third threshold.
[0314] In some implementations, the error includes an absolute error and / or an absolute transformation error.
[0315] In some implementations, the first condition includes: the current block is a bidirectional prediction block, the first reference frame corresponding to the bidirectional prediction block is located before the current frame in the time domain, and the second reference frame corresponding to the bidirectional prediction block is located after the current frame in the time domain.
[0316] In some implementations, the first condition is associated with a temporal level of the current frame.
[0317] In some implementations, the first condition includes: the current frame belongs to a first time domain level; or, the current frame does not belong to a second time domain level; wherein the first time domain level is lower than the second time domain level.
[0318] In some implementations, the first condition includes: when the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or when the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein the fourth threshold is different from the fifth threshold.
[0319] In some implementations, the fourth threshold is 16, and / or the fifth threshold is 64.
[0320] In some implementations, the third determination unit 1330 is configured to determine a first co-located prediction block of the current block based on a motion vector of an upper neighboring block of the current block; determine a second prediction block based on the first co-located prediction block and the first prediction block; determine a second co-located prediction block of the current block based on a motion vector of a left neighboring block of the current block; and determine a third prediction block based on the second co-located prediction block and the second prediction block.
[0321] In some implementations, the encoder 1300 further includes a fourth determining unit 1350 configured to determine a reconstructed block of the current block according to the third prediction block and the residual block of the current block.
[0322] FIG14 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. As shown in FIG14 , the encoder 1400 includes: a first determining unit 1410 , a second determining unit 1420 , a third determining unit 1430 , and a fourth determining unit 1440 .
[0323] The first determining unit 1410 is configured to determine a motion vector of a current block.
[0324] The second determining unit 1420 is configured to determine a first prediction block of the current block according to the motion vector.
[0325] The third determining unit 1430 is configured to determine a co-located prediction block of the current block according to the motion vector of the neighboring block of the current block.
[0326] The fourth determining unit 1440 is configured to, if a first condition is satisfied, fuse the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponding to a first fusion area, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; if the first condition is not satisfied, fuse the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponding to a second fusion area, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein the first fusion area and the second fusion area are different; and / or the first weight is different from the second weight.
[0327] In some implementations, the first fusion region is smaller than the second fusion region.
[0328] In some implementations, the first weight is less than the second weight.
[0329] In some implementations, the first condition includes that the current frame is a screen content coding frame.
[0330] In some implementations, the encoder 1400 further includes a writing unit 1450 configured to write first identification information into a bitstream, wherein the first identification information is used to indicate one or more of the following: whether the current frame is a screen content coding frame; whether the current frame or the current block uses overlapped block motion compensation.
[0331] In some implementations, the first condition is associated with information of neighboring blocks of the current block.
[0332] In some implementations, the first condition is associated with whether the prediction mode of the neighboring block is a target prediction mode.
[0333] In some implementations, the target prediction mode includes one or more of: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
[0334] In some implementations, the adjacent block includes M sub-blocks, and the first condition includes: the number of prediction modes corresponding to the M sub-blocks that are target prediction modes is less than or equal to a first threshold.
[0335] In some implementations, each of the M sub-blocks has a size of 4×4.
[0336] In some implementations, the neighboring blocks include a left neighboring block and / or an upper neighboring block of the current block.
[0337] In some implementations, the first condition is associated with a co-located prediction block of the current block, and the co-located prediction block is determined based on a motion vector of a neighboring block of the current block.
[0338] In some implementations, the first condition is associated with one or more of the following: an error between the first prediction block and the co-located prediction block; an error between corresponding sub-block pairs in the first prediction block and the co-located prediction block.
[0339] In some implementations, the first condition includes: an error between at least one pair of the corresponding sub-blocks is less than or equal to a second threshold.
[0340] In some implementations, the first condition includes: an error between the first prediction block and the co-located prediction block is less than or equal to a third threshold.
[0341] In some implementations, the error includes an absolute error and / or an absolute transformation error.
[0342] In some implementations, the first condition includes: the current block is a bidirectional prediction block, the first reference frame corresponding to the bidirectional prediction block is located before the current frame in the time domain, and the second reference frame corresponding to the bidirectional prediction block is located after the current frame in the time domain.
[0343] In some implementations, the first condition is associated with a temporal level of the current frame.
[0344] In some implementations, the first condition includes: the current frame belongs to a first time domain level; or, the current frame does not belong to a second time domain level; wherein the first time domain level is lower than the second time domain level.
[0345] In some implementations, the first condition includes: when the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or when the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein the fourth threshold is different from the fifth threshold.
[0346] In some implementations, the fourth threshold is 16, and / or the fifth threshold is 64.
[0347] In some implementations, the third determination unit 1430 is configured to determine a first co-located prediction block of the current block based on a motion vector of an upper neighboring block of the current block; determine a second prediction block based on the first co-located prediction block and the first prediction block; and determine the second co-located prediction block of the current block based on a motion vector of a left neighboring block of the current block.
[0348] In some implementations, the fourth determining unit 1440 is configured to determine a third prediction block according to the second co-located prediction block and the second prediction block.
[0349] In some implementations, the encoder 1400 further includes a fifth determining unit 1460 configured to determine a reconstructed block of the current block according to the third prediction block and the residual block of the current block.
[0350] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0351] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0352] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 1300 or the encoder 1400. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the decoding method in the aforementioned first embodiment.
[0353] Based on the composition of the above-mentioned encoder and the computer-readable storage medium, refer to Figure 15, which shows a specific hardware structure diagram of the encoder 1500 provided in an embodiment of the present application. As shown in Figure 15, the encoder 1500 may include: a communication interface 1510, a memory 1520 and a processor 1530; each component is coupled together through a bus system 1540. It can be understood that the bus system 1540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1540 in Figure 15. Among them,
[0354] The communication interface 1510 is used to send and receive signals when sending and receiving information with other external network elements.
[0355] The memory 1520 is used to store computer programs.
[0356] The processor 1530 is configured to, when running the computer program, execute the following: determining a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; and performing overlapped block motion compensation on the first prediction block if a first condition is met.
[0357] Alternatively, the following steps are performed: determining a motion vector of a current block; determining a first prediction block of the current block based on the motion vector; determining a co-located prediction block of the current block based on motion vectors of neighboring blocks of the current block; if a first condition is satisfied, fusing the first prediction block with the co-located prediction block based on a first overlapping block motion compensation mode, wherein the first overlapping block motion compensation mode corresponds to a first fusion region, and in the first overlapping block motion compensation mode, a fusion weight of the first prediction block is a first weight; if the first condition is not satisfied, fusing the first prediction block with the co-located prediction block based on a second overlapping block motion compensation mode, wherein the second overlapping block motion compensation mode corresponds to a second fusion region, and in the second overlapping block motion compensation mode, a fusion weight of the second prediction block is a second weight; wherein the first fusion region and the second fusion region are different; and / or the first weight is different from the second weight.
[0358] It is understood that the memory 1520 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1520 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0359] The processor 1530 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1530. The above-mentioned processor 1530 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1520 , and the processor 1530 reads the information in the memory 1520 and completes the steps of the above method in combination with its hardware.
[0360] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0361] Optionally, as another embodiment, the processor 1530 is further configured to execute the encoding method described in the above embodiment when running the computer program.
[0362] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0363] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0364] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0365] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0366] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0367] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, applied to a decoder, comprising: Parse the bitstream to determine the motion vector of the current block; Determine a first prediction block of the current block according to the motion vector; If a first condition is satisfied, perform overlapping block motion compensation on the first prediction block.
2. The method according to claim 1, wherein, The first condition includes that the value of first identification information parsed from the bitstream is not a first value, and the first value is used to indicate one or more of the following: The current frame is a screen content coding frame; The current frame or the current block does not use overlapping block motion compensation.
3. The method according to claim 1, wherein, The first condition is associated with the information of adjacent blocks of the current block.
4. The method according to claim 3, wherein, The first condition is associated with whether the prediction mode of the adjacent blocks is a target prediction mode.
5. The method according to claim 4, wherein, The target prediction mode includes one or more of the following: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
6. The method according to claim 4 or 5, wherein, The adjacent block includes M sub-blocks, and the first condition includes: the number of sub-blocks corresponding to the M sub-blocks whose prediction modes are target prediction modes is less than or equal to a first threshold.
7. The method according to claim 6, wherein, The size of each sub-block in the M sub-blocks is 4×4.
8. The method according to any one of claims 3 to 6, wherein, The adjacent block includes the left adjacent block and / or the upper adjacent block of the current block.
9. The method according to claim 1, wherein, The first condition is associated with the co-located prediction block of the current block, and the co-located prediction block is determined based on the motion vectors of adjacent blocks of the current block.
10. The method according to claim 9, wherein, The first condition is associated with one or more of the following: The error between the first prediction block and the co-located prediction block; The error between corresponding sub-block pairs in the first prediction block and the co-located prediction block.
11. The method according to claim 10, wherein, The first condition includes: the error between at least one pair of corresponding sub-blocks is less than or equal to a second threshold.
12. The method according to claim 10, wherein, The first condition includes: the error between the first prediction block and the co-located prediction block is less than or equal to a third threshold.
13. The method according to any one of claims 9 to 12, wherein, The error includes absolute error and / or absolute transform error.
14. The method according to any one of claims 1 to 10, wherein, The first condition includes: The current block is a bi-predicted block, the first reference frame corresponding to the bi-predicted block is before the current frame in the time domain, and the second reference frame corresponding to the bi-predicted block is after the current frame in the time domain.
15. The method according to any one of claims 1 to 14, wherein, The first condition is associated with the temporal level of the current frame.
16. The method according to claim 15, wherein, The first condition includes: The current frame belongs to a first temporal level; or, The current frame does not belong to a second temporal level; wherein, the first temporal level is lower than the second temporal level.
17. The method according to any one of claims 1 to 16, wherein, The first condition includes: When the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or When the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein, the fourth threshold is different from the fifth threshold.
18. The method according to claim 17, wherein, The fourth threshold is 16, and / or, the fifth threshold is 64.
19. The method according to any one of claims 1 to 18, wherein, Performing overlapping block motion compensation on the first prediction block includes: Determine a first co-located prediction block of the current block based on the motion vector of the upper adjacent block of the current block; Determine a second prediction block according to the first co-located prediction block and the first prediction block; Determine a second co-located prediction block of the current block based on the motion vector of the left adjacent block of the current block; Determine a third prediction block according to the second co-location prediction block and the second prediction block.
20. The method according to claim 19, wherein, The method further includes: Determine a reconstructed block of the current block according to the third prediction block and the residual block of the current block.
21. A coding method, applied to an encoder, comprising: Determine a motion vector of the current block; Determine a first prediction block of the current block according to the motion vector; If a first condition is satisfied, perform overlapping block motion compensation on the first prediction block.
22. The method according to claim 21, wherein, The first condition includes: the current frame is a screen content coding frame.
23. The method according to claim 22, wherein, The method further includes: Write first identification information into a bitstream, where the first identification information is used to indicate one or more of the following: Whether the current frame is a screen content coding frame; Whether the current frame or the current block uses overlapping block motion compensation.
24. The method according to claim 21, wherein, The first condition is associated with information of adjacent blocks of the current block.
25. The method according to claim 24, wherein, The first condition is associated with whether the prediction mode of the adjacent block is a target prediction mode.
26. The method according to claim 25, wherein, The target prediction mode includes one or more of the following: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and DC prediction mode.
27. The method according to claim 25 or 26, wherein, The adjacent block includes M sub-blocks, and the first condition includes: the number of sub-blocks corresponding to the M sub-blocks whose prediction modes are target prediction modes is less than or equal to a first threshold.
28. The method according to claim 27, wherein, The size of each sub-block in the M sub-blocks is 4×4.
29. The method according to any one of claims 23 to 26, wherein, The adjacent block includes the left adjacent block and / or the upper adjacent block of the current block.
30. The method according to claim 21, wherein, The first condition is associated with a co-location prediction block of the current block, and the co-location prediction block is determined based on motion vectors of adjacent blocks of the current block.
31. The method according to claim 30, wherein, The first condition is associated with one or more of the following: The error between the first prediction block and the co-location prediction block; The error between corresponding sub-block pairs in the first prediction block and the co-location prediction block.
32. The method according to claim 31, wherein, The first condition includes: the error between at least one of the corresponding sub-blocks is less than or equal to a second threshold.
33. The method according to claim 31, wherein, The first condition includes: the error between the first prediction block and the co-location prediction block is less than or equal to a third threshold.
34. The method according to any one of claims 30 to 33, wherein, The error includes absolute error and / or absolute transform error.
35. The method according to any one of claims 21 to 34, wherein, The first condition includes: The current block is a bi-prediction block, the first reference frame corresponding to the bi-prediction block is before the current frame in the time domain, and the second reference frame corresponding to the bi-prediction block is after the current frame in the time domain.
36. The method according to any one of claims 21 to 35, wherein, The first condition is associated with the temporal level of the current frame.
37. The method according to claim 36, wherein, The first condition includes: The current frame belongs to the coding frame of the first temporal level; or, The current frame does not belong to the coding frame of the second temporal level; wherein, the first temporal level is lower than the second temporal level.
38. The method according to any one of claims 21 to 37, wherein, The first condition includes: When the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or, When the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein, the fourth threshold is different from the fifth threshold.
39. The method according to claim 38, wherein, The fourth threshold is 16, and / or the fifth threshold is 64.
40. The method according to any one of claims 21 to 39, wherein, Performing the overlapping block motion compensation on the first prediction block includes: Determine a first co-located prediction block of the current block based on a motion vector of an upper adjacent block of the current block; Determine a second prediction block according to the first co-located prediction block and the first prediction block; Determine a second co-located prediction block of the current block based on a motion vector of a left adjacent block of the current block; Determine a third prediction block according to the second co-located prediction block and the second prediction block.
41. The method according to claim 40, wherein, The method further includes: Determine a residual block of the current block according to the third prediction block.
42. A decoding method, applied to a decoder, comprising: Parse a bitstream to determine a motion vector of the current block; Determine a first prediction block of the current block according to the motion vector; Determine a co-located prediction block of the current block according to motion vectors of adjacent blocks of the current block; If a first condition is satisfied, fuse the first prediction block and the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponding to a first fusion region, and in the first overlapping block motion compensation mode, a fusion weight of the first prediction block being a first weight; If the first condition is not satisfied, fuse the first prediction block and the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponding to a second fusion region, and in the second overlapping block motion compensation mode, a fusion weight of the second prediction block being a second weight; Wherein, the first fusion region and the second fusion region are different; and / or, the first weight and the second weight are different.
43. The method according to claim 42, wherein, The first fusion region is smaller than the second fusion region.
44. The method according to claim 42, wherein, The first weight is smaller than the second weight.
45. The method according to any one of claims 42 to 44, wherein, The first condition includes that a value of first identification information parsed from the bitstream is not a first value, and the first value is used to indicate one or more of the following: The current frame is a screen content coding frame; The current frame uses a default overlapping block motion compensation fusion region / fusion weight; The current frame or the current block does not use overlapping block motion compensation.
46. The method according to any one of claims 42 to 44, wherein, The first condition is associated with information of adjacent blocks of the current block.
47. The method according to any one of claims 42 to 44, wherein, The first condition is associated with whether a prediction mode of the adjacent blocks is a target prediction mode.
48. The method according to claim 47, wherein, The target prediction mode includes one or more of the following: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
49. The method according to claim 47 or 48, wherein, The adjacent block includes M sub-blocks, and the first condition includes: the number of sub-blocks corresponding to the M sub-blocks whose prediction modes are target prediction modes is less than or equal to a first threshold.
50. The method according to claim 49, wherein, Each sub-block of the M sub-blocks has a size of 4×4.
51. The method according to any one of claims 46 to 49, wherein, The adjacent block includes a left adjacent block and / or an upper adjacent block of the current block.
52. The method according to any one of claims 42 to 44, wherein, The first condition is associated with the co-located prediction block of the current block, and the co-located prediction block is determined based on motion vectors of adjacent blocks of the current block.
53. The method according to claim 52, wherein, The first condition is associated with one or more of the following: An error between the first prediction block and the co-located prediction block; An error between corresponding sub-block pairs in the first prediction block and the co-located prediction block.
54. The method according to claim 53, wherein, The first condition includes: an error between at least one pair of corresponding sub-blocks is less than or equal to a second threshold.
55. The method according to claim 53, wherein, The first condition includes that the error between the first prediction block and the co-located prediction block is less than or equal to a third threshold.
56. The method according to any one of claims 52 to 55, wherein, The error includes an absolute error and / or an absolute transform error.
57. The method according to any one of claims 42 to 56, wherein, The first condition includes: The current block is a bi-prediction block, the first reference frame corresponding to the bi-prediction block is temporally before the current frame, and the second reference frame corresponding to the bi-prediction block is temporally after the current frame.
58. The method according to any one of claims 42 to 57, wherein The first condition is associated with the temporal level of the current frame.
59. The method according to claim 58, wherein The first condition includes: The current frame belongs to a first temporal level; or, The current frame does not belong to a second temporal level; wherein, the first temporal level is lower than the second temporal level.
60. The method according to any one of claims 42 to 59, wherein The first condition includes: When the current frame is a screen content coding frame, the size of the current block is greater than or equal to a fourth threshold; and / or When the current block is not a screen content coding frame, the size of the current block is greater than or equal to a fifth threshold; wherein, the fourth threshold is different from the fifth threshold.
61. The method according to claim 60, wherein The fourth threshold is 16, and / or, the fifth threshold is 64.
62. The method according to any one of claims 42 to 61, wherein Determining the co-located prediction block of the current block according to the motion vectors of the adjacent blocks of the current block includes: Determining a first co-located prediction block of the current block based on the motion vector of the upper adjacent block of the current block; Determining a second prediction block according to the first co-located prediction block and the first prediction block; Determining a second co-located prediction block of the current block based on the motion vector of the left adjacent block of the current block.
63. The method according to claim 62, wherein Fusing the first prediction block and the co-located prediction block includes: Determining a third prediction block according to the second co-located prediction block and the second prediction block.
64. The method according to claim 63, wherein The method further includes: Determining the reconstructed block of the current block according to the third prediction block and the residual block of the current block.
65. A coding method, applied to an encoder, comprising: Determining the motion vector of the current block; Determining the first prediction block of the current block according to the motion vector; Determining the co-located prediction block of the current block according to the motion vectors of the adjacent blocks of the current block; If the first condition is satisfied, fusing the first prediction block and the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponds to a first fusion region, and in the first overlapping block motion compensation mode, the fusion weight of the first prediction block is a first weight; If the first condition is not satisfied, fusing the first prediction block and the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponds to a second fusion region, and in the second overlapping block motion compensation mode, the fusion weight of the second prediction block is a second weight; wherein, the first fusion region is different from the second fusion region; and / or, the first weight is different from the second weight.
66. The method according to claim 65, wherein The first fusion region is smaller than the second fusion region.
67. The method according to claim 65, wherein The first weight is smaller than the second weight.
68. The method according to any one of claims 65 to 67, wherein The first condition includes that the current frame is a screen content coding frame.
69. The method according to claim 68, wherein, The method further includes: Writing first identification information into the bitstream, the first identification information is used to indicate one or more of the following: Whether the current frame is a screen content coding frame; Whether the current frame or the current block uses overlapping block motion compensation.
70. The method according to any one of claims 65 to 67, wherein, The first condition is associated with information of adjacent blocks of the current block.
71. The method according to claim 70, wherein, The first condition is associated with whether the prediction mode of the adjacent block is a target prediction mode.
72. The method according to claim 71, wherein, The target prediction mode includes one or more of the following: intra block copy mode, string matching prediction mode, palette mode, pulse code modulation mode, horizontal prediction mode, vertical prediction mode, and direct current prediction mode.
73. The method according to claim 71 or 72, wherein, The adjacent block includes M sub - blocks, and the first condition includes: the number of sub - blocks corresponding to the M sub - blocks whose prediction mode is the target prediction mode is less than or equal to a first threshold.
74. The method according to claim 73, wherein, The size of each of the M sub - blocks is 4×4.
75. The method according to any one of claims 69 to 72, wherein, The adjacent block includes the left adjacent block and / or the upper adjacent block of the current block.
76. The method according to any one of claims 65 to 67, wherein, The first condition is associated with the co - located prediction block of the current block, and the co - located prediction block is determined based on the motion vectors of the adjacent blocks of the current block.
77. The method according to claim 76, wherein, The first condition is associated with one or more of the following: The error between the first prediction block and the co - located prediction block; The error between corresponding sub - block pairs in the first prediction block and the co - located prediction block.
78. The method according to claim 77, wherein, The first condition includes: the error between at least one of the corresponding sub - blocks is less than or equal to a second threshold.
79. The method according to claim 77, wherein, The first condition includes: the error between the first prediction block and the co - located prediction block is less than or equal to a third threshold.
80. The method according to any one of claims 76 to 79, wherein, The error includes absolute error and / or absolute transform error.
81. The method according to any one of claims 65 to 80, wherein, The first condition includes: The current block is a bi - directional prediction block, the first reference frame corresponding to the bi - directional prediction block is temporally before the current frame, and the second reference frame corresponding to the bi - directional prediction block is temporally after the current frame.
82. The method according to any one of claims 65 to 81, wherein, The first condition is associated with the temporal level of the current frame.
83. The method according to claim 82, wherein, The first condition includes: The current frame belongs to the coded frames of the first temporal level; or, The current frame does not belong to the coded frames of the second temporal level; wherein, the first temporal level is lower than the second temporal level.
84. The method according to any one of claims 65 to 83, wherein,The first condition includes: When the current frame is a screen content coded frame, the size of the current block is greater than or equal to a fourth threshold; and / or When the current block is not a screen content coded frame, the size of the current block is greater than or equal to a fifth threshold; wherein, the fourth threshold is different from the fifth threshold.
85. The method according to claim 84, wherein, The fourth threshold is 16, and / or, the fifth threshold is 64.
86. The method according to any one of claims 65 to 85, wherein, Determining the co - located prediction block of the current block according to the motion vectors of the adjacent blocks of the current block includes: Determining a first co - located prediction block of the current block based on the motion vector of the upper adjacent block of the current block; Determining a second prediction block according to the first co - located prediction block and the first prediction block; Determining a second co - located prediction block of the current block based on the motion vector of the left adjacent block of the current block.
87. The method according to claim 86, wherein, Fusing the first prediction block and the co - located prediction block includes: Determining a third prediction block according to the second co - located prediction block and the second prediction block.
88. The method according to claim 87, wherein, The method further includes: Determining the residual block of the current block according to the third prediction block.
89. A decoder, comprising: A first determination unit, configured to parse the code stream and determine the motion vector of the current block; A second determination unit, configured to determine the first prediction block of the current block according to the motion vector; A third determination unit configured to perform overlapping block motion compensation on the first prediction block if a first condition is satisfied.
90. A decoder, comprising: A memory for storing a computer program; A processor configured to execute the method according to any one of claims 1 to 20 when running the computer program.
91. An encoder, comprising: A first determination unit configured to determine a motion vector of a current block; A second determination unit configured to determine a first prediction block of the current block according to the motion vector; A third determination unit configured to perform overlapping block motion compensation on the first prediction block if a first condition is satisfied.
92. An encoder, comprising: A memory for storing a computer program; A processor configured to execute the method according to any one of claims 21 to 41 when running the computer program.
93. A decoder, comprising: A first determination unit configured to parse a bitstream and determine a motion vector of a current block; A second determination unit configured to determine a first prediction block of the current block according to the motion vector; A third determination unit configured to determine a co-located prediction block of the current block according to motion vectors of adjacent blocks of the current block; A fourth determination unit configured to, if the first condition is satisfied, fuse the first prediction block and the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponding to a first fusion region, and in the first overlapping block motion compensation mode, a fusion weight of the first prediction block being a first weight; If the first condition is not satisfied, fuse the first prediction block and the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponding to a second fusion region, and in the second overlapping block motion compensation mode, a fusion weight of the second prediction block being a second weight; Wherein, the first fusion region and the second fusion region are different; and / or, the first weight and the second weight are different.
94. A decoder, comprising: A memory for storing a computer program; A processor configured to execute the method according to any one of claims 42 to 64 when running the computer program.
95. An encoder, comprising: A first determination unit configured to determine a motion vector of a current block; A second determination unit configured to determine a first prediction block of the current block according to the motion vector; A third determination unit configured to determine a co-located prediction block of the current block according to motion vectors of adjacent blocks of the current block; A fourth determination unit configured to, if the first condition is satisfied, fuse the first prediction block and the co-located prediction block based on a first overlapping block motion compensation mode, the first overlapping block motion compensation mode corresponding to a first fusion region, and in the first overlapping block motion compensation mode, a fusion weight of the first prediction block being a first weight; If the first condition is not satisfied, fuse the first prediction block and the co-located prediction block based on a second overlapping block motion compensation mode, the second overlapping block motion compensation mode corresponding to a second fusion region, and in the second overlapping block motion compensation mode, a fusion weight of the second prediction block being a second weight; Wherein, the first fusion region and the second fusion region are different; and / or, the first weight and the second weight are different.
96. An encoder, comprising: A memory for storing a computer program; A processor, configured to execute the method according to any one of claims 65 to 88 when running the computer program.
97. A non - volatile computer - readable storage medium storing a bitstream, the bitstream being generated by using an encoding method of an encoder or decoded by using a decoding method of a decoder, wherein, The decoding method is the method according to any one of claims 1 to 20 or 42 to 64, and the encoding method is the method according to any one of claims 21 to 41 or 65 to 88.
98. A bitstream, the bitstream comprising a bitstream generated by the method according to any one of claims 1 to 20 or 42 to 64, or a bitstream generated by the method according to any one of claims 21 to 41 or 65 to 88.
Citation Information
Patent Citations
Method and apparatus for encoding and decoding an image by using consecutive motion estimation
CN101960855A
Method and apparatus for encoding / decoding image, and recording medium in which bit stream is stored
CN110024394A
Video encoding and decoding method and related device
CN116896640A
Systems and methods for combining subblock motion compensation and overlapped block motion compensation
US20230388535A1