Corresponding methods for boundary strength derivation of encoders, decoders, and deblocking filters
By determining the boundary strength of deblocking filters based on CIIP prediction in video coding, the method addresses the challenge of efficient video compression, enhancing compression ratio and picture quality through optimized deblocking filter adaptation.
Patent Information
- Application Number
- JP2024110997
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-14
- Filing Date
- 2024-07-10
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2039-12-07
AI Technical Summary
Existing video coding technologies face challenges in efficiently compressing video data while maintaining high quality, particularly in scenarios with limited bandwidth and memory constraints.
The method involves determining the boundary strength of deblocking filters based on whether adjacent blocks are predicted using combined inter-intra prediction (CIIP), and adjusting the boundary strength accordingly to enhance deblocking filter effectiveness for block edges predicted by CIIP.
This approach improves the compression ratio and picture quality by optimizing deblocking filter adaptation, specifically enhancing the opportunity for deblocking filters at block edges predicted by CIIP, thereby reducing artifacts and improving subjective and objective quality.
Smart Images

Figure 0007692096000006 
Figure 0007692096000007 
Figure 0007692096000008
Abstract
Description
Technical Field
[0001] [Related Applications] This application claims the benefit of U.S. Provisional Application No. 62 / 776,491, filed Dec. 7, 2018, entitled “AN ENCODER, A DECODER AND CORRESPONDING METHODS OF DEBLOCKING FILTER ADAPTATION”, and U.S. Provisional Application No. 62 / 792,380, filed Jan. 14, 2019, entitled “AN ENCODER, A DECODER AND CORRESPONDING METHODS OF DEBLOCKING FILTER ADAPTATION”, both of which are hereby incorporated by reference. [Technical Field] Embodiments of the present disclosure generally relate to the field of picture processing, and more particularly, to an encoder, a decoder, and corresponding methods for deriving boundary strengths of a deblocking filter.
Background Art
[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications such as broadcast digital television, video transmission over the Internet and mobile networks, real-time conferencing applications such as video chat, video conferencing, DVDs and Blu-ray discs, video content acquisition and editing systems, and camcorders for security applications.
[0003] The amount of video data required to depict even relatively short videos can be substantial. This can pose difficulties when the data is streamed across a communication network with limited bandwidth capabilities or communicated in other scenarios. Thus, video data is typically compressed before being communicated across today's telecommunications networks. When videos are stored on storage devices, the size of the videos can also be an issue since memory resources may be limited. Video compression devices often use software and / or hardware to encode video data at the source prior to transmission or storage, thereby reducing the amount of data required to represent digital video images. The compressed data is then received at the destination by a video decompression device that decodes the video data. With limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve the compression ratio without sacrificing picture quality significantly or at all are desirable. SUMMARY OF THE INVENTION
[0004] Embodiments of the present application provide apparatuses and methods for encoding and decoding in accordance with the independent claims.
[0005] The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementations are apparent from the dependent claims, the description, and the drawings.
[0006] According to a first aspect, the present invention relates to a method of encoding, said encoding including decoding or encoding, said method comprising the step of determining whether at least one of two blocks is a block by combined inter-intra prediction (CIIP) (or MH) prediction, said two blocks including a first block (block Q) and a second block (block P), said two blocks being associated with a boundary. When at least one of said two blocks is a block by CIIP, the method further comprises the step of setting the boundary strength (Bs) of said boundary to a first value, or when neither of said two blocks is a block by CIIP, setting the boundary strength (Bs) of said boundary to a second value.
[0007] The method according to the first aspect of the present invention can be executed by a device according to a second aspect of the present invention. The device according to the second aspect of the present invention includes a determination unit configured to determine whether at least one of two blocks is predicted by the application of combined inter-intra prediction (CIIP), said two blocks including a first block (block Q) and a second block (block P), said two blocks being associated with a boundary. The device according to the second aspect of the present invention further includes a setting unit configured to set the boundary strength (Bs) of said boundary to a first value when at least one of said two blocks is predicted by the application of CIIP, and to set the boundary strength (Bs) of said boundary to a second value when neither of said two blocks is predicted by the application of CIIP.
[0008] Further features and implementation forms of the method according to the second aspect of the present invention correspond to the features and implementation forms of the device according to the first aspect of the present invention.
[0009] According to a third aspect, the present invention relates to an apparatus for decoding a video stream, including a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the first aspect.
[0010] According to a fourth aspect, the present invention relates to an apparatus for encoding a video stream, including a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the first aspect.
[0011] According to a fifth aspect, there is provided a computer-readable storage medium storing instructions that, when executed, cause one or more processors to be configured to encode video data. The instructions cause the one or more processors to execute the method according to the first or second aspect or any possible embodiment of the first aspect.
[0012] According to a sixth aspect, the present invention relates to a computer program including program code for executing the method according to the first aspect or any possible embodiment of the first aspect when executed on a computer.
[0013] According to an embodiment of the present invention, when at least one of the two blocks is a block by CIIP, by setting the boundary strength to the first value (for example, set to 2), the opportunity of the deblocking filter for the block edge predicted by the application of CIIP prediction is increased.
[0014] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0015] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings and figures.
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
[0016] Unless otherwise expressly stated, the same reference numerals represent the same or at least functionally equivalent features.
DETAILED DESCRIPTION OF THE INVENTION
[0017] In the following description, reference is made to the accompanying drawings which form a part of the present disclosure and which illustrate specific aspects of embodiments of the invention or specific aspects in which embodiments of the invention may be used. It is understood that embodiments of the invention may be used in other aspects and may include structural or logical changes not shown in the figures. The following detailed description should not, therefore, be considered in a limiting sense, and the scope of the invention is defined by the appended claims.
[0018] For example, the disclosure related to the described method also applies to the corresponding apparatus or system configured to execute the method, and vice versa. For example, when one or more specific method steps are described, the corresponding apparatus may include one or more units, such as functional units, to execute the one or more method steps described, even if one or more units are not explicitly described or illustrated (e.g., one unit executes one or more steps, or each of a plurality of units executes one or more of a plurality of steps). On the other hand, for example, when a specific device is described based on one or more units, such as functional units, the corresponding method may include one step for executing the functions of the one or more units, even if one or more steps are not explicitly described or illustrated (e.g., one step executes the functions of one or more units, or each of a plurality of steps executes the functions of one or more of a plurality of units). Further, it is understood that the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.
[0019] Video encoding typically represents the processing of a sequence of pictures that form a video or video sequence. Instead of the term "picture", the terms "frame" or "image" may be used synonymously in the field of video encoding. Video encoding (or generally encoding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves the reverse process compared to the encoder to reconstruct the video picture. Embodiments that refer to the "encoding" of a video picture (or generally a picture) should be understood to be related to the "encoding" or "decoding" of the video picture or each video sequence. The combination of the encoding part and the decoding part is also called a CODEC (Coding and Decoding).
[0020] In the case of lossless video encoding, the original video picture is reconstructible. That is, the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video encoding, further compression, e.g., by quantization, is performed to reduce the amount of data representing the video picture. This cannot be fully reconstructed by the decoder. That is, the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.
[0021] Several video coding standards belong to the group of "lossy hybrid video coders" (i.e., combining spatial and temporal prediction in the sample domain with 2D transform coding applying quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and the coding is typically performed at the block level. In other words, in the encoder, for example, prediction blocks are generated using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, the prediction blocks are subtracted from the current block (the block being currently processed / to be processed) to obtain a residual block, the residual block is transformed, and the residual block is quantized in the transform domain to reduce (compress) the amount of data to be transmitted. Thus, the video is typically processed, i.e., coded, at the block (video block) level. On the other hand, in the decoder, the reverse process compared to the encoder is applied to the coded or compressed blocks to reconstruct the current block for presentation. Further, the encoder duplicates the decoder processing loop to generate the same prediction (e.g., intra and inter prediction) and / or reconstruction for both to process, i.e., code, subsequent blocks.
[0022] In the following embodiments of the video coding system 10, the video encoder 20 and the video decoder 30 will be described with reference to FIGS. 1 - 3.
[0023] FIG. 1A is a schematic block diagram showing an exemplary coding system 10 that can utilize the technology of the present application, e.g., a video coding system 10 (or simply coding system 10). The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) of the video coding system 10 represent examples of devices that can be configured to perform techniques according to various examples described in the present application.
[0024] As shown in FIG. 1A, the coding system 10 includes a source device 12 configured to provide coded picture data 21 to a destination device 14 that decodes, for example, coded picture data 13.
[0025] The source device 12 includes an encoder 20 and may additionally or optionally include a picture source 16, a preprocessor (or preprocessing unit) 18, for example a picture preprocessor 18, and a communication interface or communication unit 22.
[0026] The picture source 16 may include or be any kind of picture capture device, for example a camera that captures real pictures, and / or any kind of picture generation device, for example a computer graphics processor that generates computer animation pictures, or any other device that acquires and / or provides real-world pictures, computer-generated pictures (for example, screen content, virtual reality (VR) pictures) and / or any combination thereof (for example, augmented reality (AR) pictures). The picture source may be any kind of memory or storage that stores any of the aforementioned pictures.
[0027] In contrast to the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.
[0028] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, trimming, color format conversion (for example, from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component.
[0029] The video encoder 20 is configured to receive the preprocessed picture data 19 and provide encoded picture data 21 (further details will be described later, for example, based on FIG. 2).
[0030] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0031] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and additionally, optionally, may include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0032] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof) from, for example, directly from the source device 12 or any other source, such as a storage device, e.g., an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.
[0033] The communication interface 22 and the communication interface 28 are configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or any type of network, such as a wired or wireless network, or any combination thereof, or any type of private and public network, or any combination of any type thereof.
[0034] The communication interface 22 may be configured to package the encoded picture data 21 into an appropriate format, such as a packet, and / or process the encoded picture data using any type of transmission encoding or processing for transmission via a communication link or communication network.
[0035] The communication interface 28 forms the counterpart of the communication interface 22 and may be configured to, for example, receive the transmitted data, process the transmitted data using any kind of corresponding transmission decoding or processing, and / or unpack to obtain the encoded picture data 21.
[0036] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface, as indicated by the arrow of the communication channel 13 pointing from the source device 12 to the destination device 14 in FIG. 1A. For example, to establish a connection, to send an affirmative response and to exchange any other information related to the communication link and / or data transmission, for example the encoded picture data transmission, it may be configured to send and receive messages, for example.
[0037] The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or a decoded picture 31 (further details will be described later, for example based on FIG. 3 or FIG. 5).
[0038] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example the decoded picture 31, to obtain post-processed picture data 33, for example a post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31, for example, for display by a display device 34.
[0039] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33, for example, to display a picture to a user or viewer. The display device 34 may be or include any type of display that presents a reconstructed picture, such as a built-in or external display or monitor. The display may include, for example, liquid crystal displays (LCDs), organic light emitting diodes (OLED) displays, plasma displays, projectors, micro LED displays, liquid crystal on silicon (LCoS), digital light processor (DLP), or any other type of display.
[0040] FIG. 1A shows the source device 12 and the destination device 14 as separate devices, but embodiments of the devices may include both the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality, or both. In such embodiments, the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or separate hardware and / or software, or any combination thereof.
[0041] As will be apparent to those skilled in the art based on the description, the functions or the presence and (exact) partitioning of different units within the source device 12 and / or destination device 14 as shown in FIG. 1A may vary depending on the actual device and application.
[0042] The encoder 20 (e.g., video encoder 20), or the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30 may be implemented by a processing circuit as shown in FIG. 1B, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding specific, or any combination thereof. The encoder 20 may be implemented by the processing circuit 46 to implement various modules as discussed with respect to the encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by the processing circuit 46 to implement various modules as discussed with respect to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations as discussed later. As shown in FIG. 5, if the technology is implemented partially in software, the apparatus may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder 20 or the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) within a single device, for example, as shown in FIG. 1B.
[0043] The source device 12 and the destination device 14 may include any of a wide range of devices, such as any type of handheld or stationary device, e.g., a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiving device, a broadcast transmitting device, etc., and may or may not use any type of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0044] In some cases, the video encoding system 10 shown in FIG. 1A is merely an example, and the technology of the present application may be applicable to video encoding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between the encoding device and the decoding device. In other examples, the data is read from local memory, streamed over a network, etc. The video encoding device may encode the data and store it in memory, and / or the video decoding device may read the data from memory and decode it. In some examples, the encoding and decoding are performed by devices that do not communicate with each other but merely encode data to and / or read and decode data from memory.
[0045] For the sake of convenience of explanation, embodiments of the present invention will be described herein with reference to, for example, the reference software of High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), and the next-generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0046] Encoder and encoding method FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction processing unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder according to a hybrid video codec.
[0047] The residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming the forward signal path of the encoder 20. On the other hand, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction processing unit 244, and the intra prediction unit 254 may be referred to as forming the reverse signal path of the video encoder 20, and the reverse signal path of the video encoder 20 corresponds to the signal path of the decoder (see the video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also represented as forming the "built-in decoder" of the video encoder 20.
[0048] Picture and picture partition (picture and block) The encoder 20 may be configured to receive, for example, via the input 201, a picture 17 (or picture data 17), for example, a sequence of pictures forming a video or a video sequence. The received picture or picture data may be preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 may also be referred to as the current picture or the picture to be encoded (especially in video coding, to distinguish the current picture from other pictures, for example, pictures encoded and / or decoded before the same video sequence, i.e., the video sequence including the current picture).
[0049] (Digital) A picture is or can be considered to be a two-dimensional array or matrix of samples having intensity values. Samples in the array may also be called pixels (abbreviation of picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are typically used. That is, a picture may be represented by or include three sample arrays. In the RGB format or color space, a picture includes corresponding red, green, and blue sample arrays. However, in video encoding, each pixel is typically represented in YCbCr, which includes a luminance component indicated by Y (sometimes L is used instead) and two chrominance components indicated by Cb and Cr. The luminance (or simply luma) component Y represents brightness or gray-level intensity (such as in a grayscale picture). On the other hand, the two chrominance (or simply chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted to or from the YCbCr format, and the process is also known as color conversion or transformation. If a picture is monochromatic, the picture may include only a luminance sample array. Thus, a picture may be, for example, an array of luma samples in a monochromatic format or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0050] An embodiment of video encoder 20 may include a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may use the same block size for all pictures in a video sequence and for the corresponding grid that defines the block size, or may vary the block size between pictures or subsets or groups of pictures and be configured to partition each picture into corresponding blocks.
[0051] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, such as one, some, or all of the blocks that form picture 17. Picture blocks 203 may also be referred to as current picture blocks or blocks being encoded.
[0052] Similar to picture 17, picture blocks 203 are also here a two-dimensional array or matrix of samples having intensity values (sample values) or can be considered as such, but of smaller dimensions than picture 17. In other words, block 203 may include, for example, one sample array (e.g., the luma array in the case of a monochrome picture 17, or the luma or chroma arrays in the case of a color picture), or three sample arrays (e.g., in the case of a color picture 17, the luma and two chroma arrays), or any other number and / or kind of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of transform coefficients.
[0053] An embodiment of the video encoder 20 as shown in FIG. 2 may be configured to encode the picture 17 block by block. For example, encoding and prediction may be performed for each block 203.
[0054] Residual calculation The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as residual 205) based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 will be provided later). For example, the sample values of the prediction block 265 are subtracted from the sample values of the picture block 203 sample by sample (pixel by pixel) to obtain the residual block 205 in the sample domain.
[0055] Transformation The transformation processing unit 206 may be configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transformation coefficients 207 in the transform domain. The transformation coefficients 207, also referred to as transform residual coefficients, may represent the residual block 205 in the transform domain.
[0056] The transformation processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the transformation specified for H.265 / HEVC. Such integer approximations are typically scaled by a particular factor compared to orthogonal DCT transforms. To maintain the norm of the residual blocks processed by the forward and inverse transforms, an additional scaling factor is applied as part of the transformation process. The scaling factor is typically selected based on specific constraints such as the bit depth of the transform coefficients, the trade-off between accuracy and implementation cost, where the scaling factor is a power of two for shift operations, etc. A particular scaling factor may be specified, for example, for the inverse transform by the inverse transform processing unit 212 (and the corresponding inverse transform by the inverse transform processing unit 312 in the video decoder 30, for example), and the corresponding scaling factor for the forward transform by the transformation processing unit 206 in the encoder 20 may be correspondingly specified.
[0057] Embodiments of the video encoder 20 (each, the transformation processing unit 206) may be configured to output transformation parameters, such as the type of transformation or transformations, for example, to be encoded or compressed by the entropy encoding unit 270. As a result, for example, the video decoder 30 may receive and use the transformation parameters for decoding.
[0058] Quantization The quantization unit 208 may be configured to quantize the transform coefficients 207 to obtain quantized coefficients 209, for example, by applying scalar quantization or vector quantization. The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209.
[0059] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be truncated to an m-bit transform coefficient during quantization. Here, n is greater than m. The degree of quantization may be changed by adjusting a quantization parameter (QP). For example, in scalar quantization, different scalings may be applied to achieve finer or coarser quantization. The smaller the quantization step size, the more it corresponds to fine quantization. On the other hand, the larger the quantization step size, the more it corresponds to coarse quantization. The applicable quantization step size may be indicated by a quantization parameter (QP). The quantization parameter may be, for example, an index for a predetermined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), and a large quantization parameter may correspond to coarse quantization (large quantization step size). The reverse is also true. Quantization may include division by the quantization step size. For example, the corresponding and / or inverse inverse quantization by the inverse quantization unit 210 may include multiplication by the quantization step size. Some standards, such as embodiments compliant with HEVC, may be configured to use a quantization parameter to determine the quantization step size. Usually, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation including division. Additional scaling factors for quantization and inverse quantization may be introduced to restore the norm of the residual block that can be changed for the scaling used in the fixed-point approximation of the equation of the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and the inverse quantization may be combined. Alternatively, a customized quantization table may be used and signaled, for example, in a bitstream, from the encoder to the decoder. Quantization is a lossy operation, and the loss increases with an increase in the quantization step size.
[0060] Embodiments of the video encoder 20 (each, quantization unit 208) may be configured to output quantization parameters (QP), for example, encoded by the entropy encoding unit 270 directly or otherwise. As a result, for example, the video decoder 30 may receive and apply the quantization parameters for decoding.
[0061] Inverse quantization The inverse quantization unit 210 is configured to apply inverse quantization of the quantization unit 208 to the quantized coefficients, for example, by applying the inverse of the quantization method applied by the quantization unit 208 based on or using the same quantization step size as the quantization unit 208, to obtain the inverse quantized coefficients 211. The inverse quantized coefficients 211, also referred to as inverse quantized residual coefficients 211, are not typically the same as the transform coefficients due to quantization loss, but may correspond to the transform coefficients 207.
[0062] Inverse transform The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding inverse quantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as the transform block 213.
[0063] Reconstruction The reconstruction unit 214 (e.g., an adder or adder 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 for each sample, to obtain a reconstructed block 215 in the sample domain.
[0064] Filtering The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstruction block 215 to obtain a filtered block 221, or typically to filter the reconstruction samples to obtain filtered samples. The loop filter unit is configured to, for example, smooth pixel transitions or in other cases improve video quality. The loop filter unit 220 may include a deblocking filter, a sample-adaptive offset (SAO) filter or one or more other filters, such as a bidirectional filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter or a joint filter, or any combination thereof. The loop filter unit 220 is shown in FIG. 2 as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may be referred to as a filtered reconstruction block 221.
[0065] Embodiments of the video encoder 20 (each, the loop filter unit 220) may be configured to output loop filter parameters (such as sample-adaptive offset information) encoded, for example, directly or by the entropy encoding unit 270. As a result, for example, the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.
[0066] Decoded Picture Buffer The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures or generally reference picture data for encoding of video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices such as a dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previous filtered blocks of the same current picture or different pictures, e.g., previous reconstructed pictures, e.g., previous reconstructed and filtered blocks 221, and may provide a fully previously reconstructed, i.e., decoded, picture (and corresponding reference blocks and samples), and / or a partially reconstructed current picture (and corresponding reference blocks and samples) for, e.g., inter prediction. The decoded picture buffer (DPB) 230 may be configured to store one or more unfiltered reconstruction blocks 215, or generally unfiltered reconstruction samples, or any other further processed version of a reconstruction block or sample, e.g., if the reconstruction block 215 is not filtered by the loop filter unit 220.
[0067] Mode Selection (Partitioning and Prediction) The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, for example, the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data, for example, filtered and / or unfiltered reconstructed samples or blocks from the same (current) picture and / or one or more previous decoded pictures, for example, from the decoded picture buffer 230 or other buffers (for example, a line buffer not shown). The reconstructed picture data is used as reference picture data for prediction, for example, inter prediction or intra prediction, in order to obtain the prediction block 265 or predictor 265.
[0068] The mode selection unit 260 may be configured to determine or select a partition of the current block prediction mode (including not partitioning), and a prediction mode (for example, an intra or inter prediction mode), and generate a corresponding prediction block 265 used for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.
[0069] Embodiments of the mode selection unit 260 may be configured to select partition and prediction modes that provide the best match, or in other words the minimum residual (the minimum residual implies better compression for transmission or storage) or the minimum signaling overhead (the minimum signaling overhead implies better compression for transmission or storage), or consider or balance both, from, for example, those supported or available by the mode selection unit 260. The mode selection unit 260 may be configured to determine partition and prediction modes based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion. Terms such as "best", "minimum", "optimal", etc. in this context do not necessarily represent the overall "best", "minimum", "optimal", etc., but may represent an end or selection criterion such as a value above or below a threshold, or may result in a "quasi-optimal selection" but also represent the satisfaction of other constraints that reduce complexity and processing time.
[0070] In other words, the partitioning unit 262 may be configured to repeatedly use, for example, a quad-tree (QT) partition, a binary (BT) partition, or a triple-tree (TT) partition, or any combination thereof, to further partition block 203 into smaller block partitions or sub-blocks (which also form blocks), and to perform prediction, for example, on each block partition or sub-block. Here, mode selection includes selection of the tree structure of the partitioned block 203 and the prediction mode applied to each of the block partitions or sub-blocks.
[0071] The partitioning and prediction processing (e.g., by the partitioning unit 260) performed by the exemplary video encoder 20 is described in more detail below.
[0072] Partition Partition unit 262 may now partition (or divide) block 203 into even smaller partitions, for example, smaller blocks of square or rectangular size. These even smaller blocks (which may also be called sub-blocks) may be further partitioned into even smaller partitions. This is also called a tree partition or a hierarchical tree partition. Here, for example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned into, for example, two or more blocks at the next lower tree level, for example, nodes at tree level 1 (hierarchical level 1, depth 1). Here, for example, until the termination criteria are met, for example, until the maximum tree depth or the minimum block size is reached and the partition ends, these blocks may again be partitioned into two or more blocks at the next lower level, for example, tree level 2 (hierarchical level 2, depth 2), and so on. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree using a partition into two partitions is called a binary-tree (BT), a tree using a partition into three partitions is called a ternary-tree (TT), and a tree using a partition into four partitions is called a quad-tree (QT).
[0073] As described above, the term "block", as used herein, may be a portion of a picture, particularly a square or rectangular portion. For example, referring to HEVC and VVC, a block may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).
[0074] For example, a coding tree unit (CTU) may be or may include a picture encoded using a syntax structure used to encode a CTB of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a CTB of samples of a monochrome picture, or a picture encoded using a syntax structure used to encode three separate color planes and samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N. As a result, the division of a component into CTBs is a partition. A coding unit (CU) may be or may include a picture encoded using a syntax structure used to encode a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture having three sample arrays, or a coding block of samples of a monochrome picture, or a picture encoded using a syntax structure used to encode three separate color planes and samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N. As a result, the division of a CTB into coding blocks is a partition.
[0075] For example, in an embodiment according to HEVC, a coding tree unit may be divided into CUs using a quadtree structure shown as a coding tree. The decision of whether to encode a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into 1, 2, or 4 PUs according to the PU partition type. Within one PU, the same prediction process is applied, and related information is sent to the decoder for each PU. By applying a prediction process based on the PU partition type, after obtaining a residual block, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU.
[0076] For example, in an embodiment according to the currently under development latest video coding standard called Versatile Video Coding (VVC), quad-tree and binary tree (QTBT) partitions are used to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. The leaf nodes of the quadtree are further partitioned by a binary tree or a ternary (or triple) tree structure. The leaf nodes of the partition tree are called coding units (CUs) and are segmented for prediction and transform processing without any further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. At the same time, multiple partitions, such as ternary tree partitions, have also been proposed for use with the QTBT block structure.
[0077] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0078] As described above, the video encoder 20 is configured to determine the best or optimal prediction mode or select from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0079] Intra prediction The set of intra prediction modes may include 35 different intra prediction modes, such as non - directional modes like the DC (or average) mode and the planar mode, or directional modes as defined, for example, in HEVC, or may include 67 different intra prediction modes, such as non - directional modes like the DC (or average) mode and the planar mode, or directional modes as defined, for example, for VCC.
[0080] The intra prediction unit 254 is configured to use the reconstructed samples of neighboring blocks of the same current picture to generate an intra prediction block 265 according to an intra prediction mode from the set of intra prediction modes.
[0081] The intra prediction unit 254 (or generally the mode selection unit 260) is further configured to output an intra prediction parameter (or generally information indicating the intra prediction mode selected for a block) in the form of a syntax element 266 to the entropy encoding unit 270 for inclusion in the encoded picture data 21. As a result, for example, the video decoder 30 may receive and use the prediction parameter for decoding.
[0082] Inter prediction The set of inter prediction modes (or possibilities) depends on available reference pictures (i.e., at least previously partially decoded pictures stored in, e.g., DBP230) and other inter prediction parameters, e.g., whether only the whole or a part of the reference picture is used to search for the best matching reference block in the search window area around the area of the current block of the reference picture, and / or, e.g., whether pixel interpolation, e.g., half / semi pixel and / or quarter pixel interpolation, is applied or not.
[0083] In addition to the prediction modes described above, a skip mode and / or a direct mode may be applied.
[0084] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or acquire, for motion estimation, the picture block 203 (the current picture block 203 of the current picture 17) and at least one or more of the decoded picture 231 or the previous reconstructed blocks, e.g., one or more reconstructed blocks of one or more other / different previously decoded pictures 231. For example, the video sequence may include the current picture and the previous decoded picture 231, or in other words, the current picture and the previous decoded picture 231 may be part of or form a sequence of pictures forming the video sequence.
[0085] The encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different pictures of a plurality of other pictures, and provide an offset (spatial offset) between the reference picture (or reference picture index) and / or the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as an inter prediction parameter. This offset is also called a motion vector (MV).
[0086] The motion compensation unit is configured to obtain, for example receive, an inter prediction parameter and perform an inter prediction based on or using the inter prediction parameter to obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit may include fetching or generating a prediction block based on the motion / block vector determined by motion estimation and, optionally, performing interpolation to sub-pixel accuracy. Interpolation filtering may generate additional pixel samples from known pixel samples and thus increase the number of candidate prediction blocks that can be used to encode a picture block. Upon receiving the motion vector of the PU of the current picture block, the motion compensation unit may identify the position of the prediction block pointed to by the motion vector within one of the reference picture lists.
[0087] The motion compensation unit may also generate syntax elements related to the block and video slice for use by the video decoder 30 when decoding a picture block of a video slice.
[0088] Entropy coding The entropy encoding unit 270 applies or bypasses (without compressing), for example, an entropy encoding algorithm or method (e.g., variable length coding (VLC) method, context adaptive VLC (CAVLC), arithmetic coding method, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or another entropy coding method or technique) to the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, loop filter parameters and / or other syntax elements, and is configured to obtain encoded picture data 21 that can be output, for example, in the form of an encoded bitstream 21 by an output 272. As a result, for example, the video decoder 30 may receive and use the parameters for decoding. The encoded bitstream 21 may be transmitted to the video decoder 30 or stored in a memory for later transmission or reading by the video decoder 30.
[0089] Other structural variations of the video encoder 20 may be used to encode a video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal for a particular block or frame without having a transform processing unit 206. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 coupled to a single unit.
[0090] Decoder and Decoding Method FIG. 3 shows an example of a video decoder 30 configured to implement the technology of the present application. The video decoder 30 is configured to receive, for example, encoded picture data 21 (e.g., an encoded bitstream 21) encoded by an encoder 20 in order to obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, such as data indicating picture blocks of an encoded video slice and related syntax elements.
[0091] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding path that is generally reciprocal to the encoding path described with respect to the video encoder 100 in FIG. 2.
[0092] As described with respect to the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 are also represented as forming the "built-in decoder" of the video encoder 20. Accordingly, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Accordingly, the description provided for each unit and the function of the video 20 encoder applies correspondingly to each unit and function of the video decoder 30.
[0093] Entropy decoding The entropy decoding unit 304 parses the bitstream 21 (or generally the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain, for example, the quantized coefficients 309 and / or decoded encoding parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference picture index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or method corresponding to the encoding method as described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0094] Inverse quantization The inverse quantization unit 310 receives a quantization parameter (quantization parameter (QP)) (or generally information regarding inverse quantization) and the quantized coefficients from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304 for example), and applies inverse quantization to the decoded quantized coefficients 309 based on the quantization parameter to obtain inverse quantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include determining, for each video block within the video slice, the degree of quantization and likewise the degree of inverse quantization to be applied, using the quantization parameter determined by the video encoder 20.
[0095] Inverse transform The inverse transformation processing unit 312 may be configured to receive the inverse quantized coefficients 311, also referred to as transformation coefficients 311, and apply a transformation to the inverse quantized coefficients 311 to obtain a reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may also be referred to as a transformation block 313. The transformation may be an inverse transformation, for example, an inverse DCT, an inverse DST, an inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may be further configured to receive transformation parameters or corresponding information from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304 for example) to determine the transformation to be applied to the inverse quantized coefficients 311.
[0096] Reconstruction The reconstruction unit 314 (e.g., an adder or adder 314) may be configured to add the reconstructed residual block 313 to the prediction block 365 to obtain a reconstruction block 315 in the sample domain, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365.
[0097] Filtering The loop filter unit 320 (either within the encoding loop or after the encoding loop) is configured to filter the reconstruction block 315 to obtain a filtered block 321, for example, to smooth pixel transitions or improve video quality in other cases. The loop filter unit 320 may include one or more loop filters such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, for example, a bidirectional filter, an adaptive loop filter (ALF), a sharpening, a smoothing filter, or a joint filter, or any combination thereof. The loop filter unit 320 is shown in FIG. 3 as an in-loop filter, but in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0098] Decoded picture buffer The decoded video block 321 of the picture is then stored in a decoded picture buffer 330 that stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for the display output respectively.
[0099] The decoder 30 is configured to output the decoded picture 311, for example, via the output 312 for presentation or browsing to the user.
[0100] Prediction The inter prediction unit 344 may be identical to the inter prediction unit 244 (especially the motion compensation unit), the intra prediction unit 354 may be functionally identical to the inter prediction unit 254, and based on each piece of information received from the partition and / or prediction parameters or the encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoder unit 304), perform partitioning or partition determination and prediction. The mode selection unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the reconstructed picture, block, or each (filtered or unfiltered) sample to obtain the prediction block 365.
[0101] When a video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the mode selection unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from a decoded block of the previous picture of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode selection unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 304. In inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame lists: list 0 and list 1 using a prescribed construction technique based on the reference pictures stored in the DPB 330.
[0102] The mode selection unit 360 is configured to determine prediction information for a video block of the current video slice by parsing motion vectors and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, the mode selection unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to code a video block of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists of the slice, the motion vector of each inter-coded video block of the slice, the inter prediction state of each inter-coded video block of the slice, and other information for decoding a video block within the current video slice.
[0103] Other variations of the video decoder 30 can be used to decode the encoded picture data 21. For example, the decoder 30 can generate an output video stream without having a loop filter unit 320. For example, a non-conversion based decoder 30 can directly inverse quantize the residual signal for a particular block or frame without having an inverse transform processing unit 312. In another implementation, the video decoder 30 can have an inverse quantization unit 310 and an inverse transform processing unit 312 coupled to a single unit.
[0104] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of the interpolation filtering, motion vector derivation, or loop filtering.
[0105] Note that further operations may be applied to the derived motion vectors of the current block (including, but not limited to, affine mode control point motion vectors, affine, planar, sub-block motion vectors in ATMVP mode, temporal motion vectors, etc.). For example, the value of the motion vector is constrained to a predetermined range according to its representation bits. When the representation bits of the motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where " ^ " means exponentiation. For example, when bitDepth is equal to 16, the range is -32768 to 32767, and when bitDepth is equal to 18, the range is -131072 to 131071. For example, the value of the derived motion vector (e.g., the MV of 4 4×4 sub-blocks in one 8×8 block) is constrained such that the maximum difference between the integer parts of the MVs of the 4×4 sub-blocks is not greater than N pixels, such as not greater than 1 pixel.
[0106] Here, two methods of constraining the motion vector according to bitDepth are provided.
[0107] Method 1: Remove the overflow most significant bit (MSB) by the following operation.
Number
[0108]
Number
[0109] Method 2: Remove the overflow MSB by clipping the following values.
Number
Number
[0110] The video encoding device 400 includes an ingress port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an egress port 450 (or output port 460) for transmitting data, and a memory 460 for storing data. The video encoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components for ingress or egress of optical or electrical signals connected to the ingress port 410, the receiver unit 420, the transmitter unit 440, and the egress port 450.
[0111] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes an encoding module 470. The encoding module 470 implements the embodiments of the above disclosure. For example, the encoding module 470 implements, processes, prepares, or provides various encoding operations. What is included in the encoding module 470 thus provides a substantial improvement to the functions of the video encoding device 400 and results in the conversion of the video encoding device 400 to different states. Alternatively, the encoding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0112] Memory 460 may include one or more disks, tape drives, and solid state drives, and may be used as an overflow data storage device for storing a program when the program is selected for execution and for storing instructions and data read during execution of the program. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0113] FIG. 5 is a simplified block diagram of a device 500 that may be used as one or both of the source device 12 and the destination device 14 from FIG. 1 according to an exemplary embodiment.
[0114] The processor 502 within device 500 may be a central processing unit. Alternatively, processor 502 may be any other type of device or devices capable of manipulating or processing information that exists currently or that may be developed in the future. Although the disclosed implementation may be carried out by a single processor, such as processor 502 as illustrated, benefits in terms of speed and efficiency may be achieved using more than one processor.
[0115] In one implementation, the memory 504 within the machine 500 can be a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and an application program 510. The application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 can include applications 1 to N that further include a video encoding application that executes the methods described herein. The machine 500 can also include one or more output devices such as a display 518. In one example, the display 518 can be a touch-sensitive display that combines a touch-sensitive element that operates to sense touch inputs with the display. The display 518 can be coupled to the processor 502 via the bus 512.
[0116] Although shown here as a single bus, the bus 512 of the machine 500 can be composed of multiple buses. Further, the secondary storage 514 can be directly connected to other components of the machine 500 or accessed via a network and can include a single integrated unit such as a memory card or multiple units such as multiple memory cards. The machine 500 can thus be implemented in various configurations.
[0117] Combined Inter-Intra Prediction (CIIP) Conventionally, a coding unit is either intra-predicted (i.e., using reference samples within the same picture) or inter-predicted (i.e., using reference samples in other pictures). Multiple hypothesis prediction combines these two prediction approaches. Thus, it is sometimes also called combined inter-intra prediction (CIIP). When combined inter-intra prediction is enabled, the intra-predicted and inter-predicted samples are applied with weights, and the final prediction is derived as a weighted average sample.
[0118] A flag, the multiple hypothesis prediction (CIIP) flag, is used to indicate when a block is applied with combined inter-intra prediction.
[0119] A block to which CIIP is applied can be further divided into several sub-blocks as shown in FIG. 6A. In one example, the sub-blocks are derived by dividing the block horizontally, and each sub-block has the same width as the original block but a height that is 1 / 4 of the original block.
[0120] In one example, the sub-blocks are derived by dividing the block vertically, and each sub-block has the same height as the original block but a width that is 1 / 4 of the original block.
[0121] Due to CIIP prediction, block artifacts may be introduced because it typically includes the results of intra-prediction which usually has more residual signals. Block artifacts occur not only at the boundaries of CIIP blocks but also at the sub-block boundaries inside the CIIP blocks, such as the vertical sub-block edges A, B, C in FIG. 6A. The horizontal sub-block edges can be correspondingly identified.
[0122] Block artifacts can occur at both the CIIP boundaries and the sub-block boundaries inside the CIIP blocks, but the distortions caused by these two boundaries may be different, and different boundary strengths may be required.
[0123] In the remainder of this application, the following terms are used.
[0124] CIIP block: An encoded block predicted by applying multi-hypothesis prediction (CIIP).
[0125] Intra block: An encoded block predicted by applying intra prediction rather than CIIP prediction.
[0126] Inter block: An encoded block predicted by applying inter prediction rather than CIIP prediction.
[0127] Deblocking filter and boundary strength Video coding schemes such as HEVC and VVC are designed in accordance with the successful principle of block-based hybrid video coding. Using this principle, a picture is first partitioned into blocks, and then each block is predicted using intra-picture or inter-picture prediction. These blocks are encoded from relatively neighboring blocks and approximate the original signal with a certain degree of similarity. Since the encoded blocks only approximate the original signal, the differences between the approximations can cause discontinuities at the prediction and transform block boundaries. These discontinuities are attenuated by a deblocking filter.
[0128] The decision of whether to filter the block boundary uses bitstream information such as the prediction mode and motion vector. Some coding conditions are more likely to generate strong block artifacts, which are represented by the so-called boundary strength (Bs or BS) variable that is assigned to all block boundaries and determined as shown in Table 1.
Table 1
[0129] Normally, two adjacent blocks at the boundary are labeled P and Q as shown in FIG. 6B. FIG. 6B shows the case of a vertical boundary. When a horizontal boundary is considered, FIG. 6B should be rotated 90 degrees clockwise, with P on the upper side and Q on the lower side.
[0130] Most Probable Mode List Construction The Most Probable Mode (MPM) list is used in intra mode coding to improve coding efficiency. Due to a large number of intra modes (e.g., 35 in HEVC and 67 in VVC), the intra mode of the current block is not directly signaled. Instead, the MPM list of the current block is constructed based on the intra prediction modes of its neighboring blocks. Since the intra mode of the current block is related to those of its neighbors, the MPM list usually provides a good prediction as its name (Most Probable Mode list) indicates. Therefore, there is a high possibility that the intra mode of the current block is included in its MPM list. In this way, only the index of the MPM list is signaled to derive the intra mode of the current block. Compared with the number of all intra modes, the length of the MPM list is much smaller (e.g., a 3-MPM list is used in HEVC and a 6-MPM list is used in VVC), and thus only a few bits are needed to code the intra mode. A flag (mpm_flag) is used to indicate whether the intra mode of the current block belongs to its MPM list. If it belongs, the intra mode of the current block can be indexed using the MPM list. In other cases, the intra mode is directly signaled using a binary code. In both VVC and HEVC, the MPM list is constructed based on its left and upper neighboring blocks. When the left and upper neighboring blocks of the current block are not available for prediction, a predefined mode list is used.
[0131] Motion Vector Prediction Motion vector prediction is a technique used in motion data encoding. A motion vector usually has two components x and y that represent motion in the horizontal and vertical directions respectively. The motion vector of the current block is usually correlated with the motion vectors of neighboring blocks within the current picture or in the previous encoded picture. This is because neighboring blocks may correspond to the same moving object having similar motion, and the motion of the object is unlikely to change abruptly over time. Therefore, using the motion vectors among neighboring blocks as predictors reduces the size of the motion vector differences to be signaled. A Motion Vector Predictor (MVP) is usually derived from the already decoded motion vectors from spatial neighboring blocks or from temporal neighboring blocks in the picture at the same position.
[0132] If a block is determined to be predicted by CIIP prediction, its finally predicted samples are partially based on intra prediction samples. Since intra prediction is also included, usually, the residuals and transform coefficients are more than those of INTER blocks (mvd, merge, skip). Therefore, when these MH blocks are adjacent to other blocks, there are more discontinuities across the boundaries. In HEVC and VVC, when either of the two adjacent blocks at the boundary is intra predicted, a strong deblocking filter is applied to this boundary, and the parameter of the Boundary Strength (BS) is set to 2 (the strongest).
[0133] However, in VTM3.0, the possible block artifacts caused by the blocks predicted by CIIP prediction are not considered. The boundary strength derivation still considers the blocks by CIIP prediction as inter blocks. Under certain circumstances, such a processing approach can result in low subjective and objective quality.
[0134] Embodiments of the present invention provide several alternatives for incorporating MH blocks to improve deblocking filters. Here, the boundary strength derivation for a particular boundary is affected by the MH block.
[0135] The reference Versatile Video Coding (Draft 3) is defined as VVC Draft 3.0 and can be found at the following link.
[0136] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 12_Macao / wg11 / JVET-L1001-v3.zip
[0137] Embodiment 1 For a boundary with two sides (where the spatially adjacent blocks on each side are shown as P block and Q block), the boundary strength is determined as follows.
[0138] 1. As shown in FIG. 7, if at least one of the P and Q blocks is a block by CIIP prediction, the boundary strength parameter of this boundary is set to a first value, for example, the first value may be equal to 2.
[0139] 2. If both the P and Q blocks are not predicted by the application of CIIP prediction, and at least one of block Q or block P is predicted by the application of intra prediction, the boundary strength is determined to be equal to 2.
[0140] 3. If both the P and block Q are not predicted by the application of CIIP prediction, and both block Q and P are predicted by the application of inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation), and the derivation of the boundary strength of this boundary is shown in FIG. 7.
[0141] 4. For comparison, the method specified in the VVC or ITU-H.265 video coding standard is provided in FIG. 8.
[0142] 5. The pixel samples included in block Q and block P are filtered by applying a deblocking filter according to the determined boundary strength.
[0143] Embodiment 2 As shown in FIG. 9, for a boundary having two sides (where the spatially adjacent blocks on each side are shown as P-blocks and Q-blocks), the boundary strength is derived as follows.
[0144] 6. If at least one of the blocks P and Q is a block by intra prediction, the boundary strength is set to 2.
[0145] 7. In other cases, if at least one of the blocks P and Q is a block by CIIP prediction, the boundary strength parameter of this boundary is set to a first value, for example, 1 or 2.
[0146] 8. In other cases, if at least one of the adjacent blocks P and Q has a non-zero transform coefficient, the boundary strength parameter of this boundary is set to a second value, for example, 1.
[0147] 9. In other cases, if the absolute difference between the motion vectors belonging to blocks P and Q is greater than or equal to one integer luma sample, the boundary strength parameter of this boundary is set to the second value, for example, 1.
[0148] 10. In other cases, if the motion prediction in the adjacent blocks refers to different reference pictures or the number of motion vectors is different, the boundary strength parameter of this boundary is set to 1.
[0149] 11. In other cases, the boundary strength parameter of this boundary is set to 0.
[0150] 12. The pixel samples included in block Q and block P are filtered by applying a deblocking filter according to the determined boundary strength.
[0151] Embodiment 3 As shown in FIG. 10, in the case of a boundary having two sides (where the spatially adjacent blocks on each side are shown as P-blocks and Q-blocks), the boundary strength parameter of this boundary is set as follows.
[0152] 13. When at least one of the blocks P and Q is predicted by applying intra prediction instead of applying CIIP prediction (the possibilities include that the P-block is predicted by intra prediction rather than by multi-hypothesis prediction and the Q-block is predicted by any prediction function, and vice versa), the boundary strength is set equal to 2.
[0153] 14. When both blocks Q and P are predicted by applying inter prediction or by applying CIIP prediction (the possibilities include that the P-block is an inter-block, the Q-block is an inter-block, or alternatively, the P-block is an inter-block, the Q-block is an MH-block, or alternatively, the P-block is an HM-block, the Q-block is an inter-block, or alternatively, the P-block is an MH-block, the Q-block is an MH-block).
[0154] 1. When at least one of the blocks P and Q has a non-zero transform coefficient, the boundary strength parameter of this boundary is set equal to 1. 2. In other cases (when the blocks Q and P do not have non-zero transform coefficients), when the absolute difference between the motion vectors used to predict the blocks P and Q is greater than or equal to one integer sample, the boundary strength parameter of this boundary is set equal to 1.
[0155] 3. In other cases (when blocks Q and P do not have non-zero transform coefficients and the absolute difference between the motion vectors is less than 1 sample), if blocks P and Q are predicted based on different reference pictures or the number of motion vectors used to predict block Q and block P is not equal, the boundary strength parameter of this boundary is set equal to 1.
[0156] 4. In other cases (when the above three conditions are evaluated as false), the boundary strength parameter of this boundary is set equal to 0.
[0157] 15. If at least one of the blocks P and Q is a block by CIIP prediction, the boundary strength is changed as follows.
[0158] 1. When the boundary strength is not equal to a predetermined first value (in one example, the predetermined first value is equal to 2), the boundary strength is incremented by a predetermined second value (in one example, the predetermined second value is equal to 1).
[0159] 16. The pixel samples included in block Q and block P are filtered by applying a deblocking filter according to the determined boundary strength.
[0160] Embodiment 4 For a boundary with two sides (P and Q as defined in VVC Draft 3.0), the boundary strength is derived as follows.
[0161] 17. When this boundary is a horizontal boundary and P and Q belong to different CTUs, 1. When block Q is a block by CIIP prediction, the boundary strength is set to 2.
[0162] 2. In other cases, the boundary strength is derived as defined in VVC Draft 3.0.
[0163] 18. In other cases, 1. When at least one of the blocks of P and Q is a block by CIIP prediction, the boundary strength parameter of this boundary is set to 2.
[0164] In other cases, derive the boundary strength of this boundary as defined in VVC Draft 3.0.
[0165] Embodiment 5 For a boundary with two sides (where the spatially adjacent blocks on each side are shown as P-block and Q-block), the boundary strength is determined as follows.
[0166] 19. When at least one of the P-block or Q-block is predicted by intra prediction rather than by the application of CIIP prediction (the possibilities include that the P-block is predicted by intra prediction rather than by multi-hypothesis prediction and the Q-block is predicted by any prediction function, and vice versa), the boundary strength is set equal to 2.
[0167] 20. When both blocks are predicted by the application of inter prediction or CIIP prediction (the possibilities include that the P-block is an inter-block, the Q-block is an inter-block, or alternatively, the P-block is an inter-block, the Q-block is an MH-block, or alternatively, the P-block is an HM-block, the Q-block is an inter-block, or alternatively, the P-block is an MH-block, the Q-block is an MH-block). 1. When this boundary is a horizontal boundary and P and Q are located in two different CTUs. 1. When the block Q (the Q-block is shown as the block located downward compared to the P-block) is predicted by the application of CIIP prediction, the boundary strength parameter of this boundary is set equal to 1.
[0168] 2. In other cases (when block Q is not predicted by applying CIIP prediction), if at least one of the adjacent blocks P and Q has a non-zero transform coefficient, the boundary strength parameter of the boundary is set equal to 1.
[0169] 3. In other cases, if the absolute difference between the motion vectors used to predict blocks P and Q is greater than or equal to one integer luma sample, the boundary strength parameter of the boundary is set equal to 1.
[0170] 4. In other cases, if the motion compensation prediction in adjacent blocks P and Q is performed based on different reference pictures, or if the number of motion vectors used to predict block Q and block P is not equal, the boundary strength parameter of the boundary is set equal to 1.
[0171] 5. In other cases, the boundary strength parameter of the boundary is set equal to 0.
[0172] 2. In other cases (when the boundary is a vertical boundary, or when blocks Q and P are included in the same CTU), 1. If at least one of blocks P and Q is predicted by applying CIIP prediction, the boundary strength parameter of the boundary is set equal to 1.
[0173] 2. In other cases, if at least one of the adjacent blocks P and Q has a non-zero transform coefficient, the boundary strength parameter of the boundary is set equal to 1.
[0174] 3. In other cases, if the absolute difference between the motion vectors used to predict blocks P and Q is greater than or equal to one integer luma sample, the boundary strength parameter of the boundary is set equal to 1.
[0175] 4. In other cases, when the motion compensation prediction in adjacent blocks P and Q is performed based on different reference pictures, or when the number of motion vectors used to predict blocks Q and P is not equal, the boundary strength parameter of the boundary is set to 1.
[0176] 5. In other cases, the boundary strength parameter of this boundary is set to 0.
[0177] 21. The pixel samples included in block Q and block P are filtered by applying a deblocking filter according to the determined boundary strength.
[0178] Embodiment 6 For a boundary with two sides (where the spatially adjacent blocks on each side are shown as P-block and Q-block), the boundary strength is determined as follows.
[0179] 22. First, determine the boundary strength of the boundary according to the method specified in the VVC or ITU-H.265 video coding standard.
[0180] 23. When the boundary is a horizontal boundary and P and Q are located in two different CTUs, 1. When block Q is predicted by applying CIIP prediction, the boundary strength is changed as follows.
[0181] 1. When the boundary strength is not equal to 2, the boundary strength is incremented by 1.
[0182] 24. In other cases (when the boundary is a vertical boundary, or when blocks Q and P are included in the same CTU), 1. When at least one of block P or block Q is predicted by applying CIIP prediction, the boundary strength parameter of the boundary is adjusted as follows.
[0183] 1. When the boundary strength is not equal to 2, the boundary strength is incremented by 1.
[0184] 25. The pixel samples included in block Q and block P are filtered by applying a deblocking filter according to the determined boundary strength.
[0185] Embodiment 7 For a boundary with two sides (where the spatially adjacent blocks on each side are shown as a P block and a Q block), the boundary strength is derived as follows.
[0186] 26. When the boundary is a horizontal boundary and blocks P and Q are located in different CTUs, 1. When block Q (block Q is shown as the block located downward compared to block P) is predicted by applying CIIP prediction, the boundary strength is set equal to 2.
[0187] 2. When block Q is not predicted by applying CIIP prediction, and at least one of block Q or block P is predicted by applying intra prediction, the boundary strength is determined to be equal to 2.
[0188] 3. When block Q is not predicted by applying CIIP prediction, and both block Q and P are predicted by applying inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation).
[0189] 27. In other cases (when the boundary is a vertical boundary, or when blocks Q and P are included in the same CTU), 1. When at least one of block P or Q is predicted by applying CIIP prediction, the boundary strength parameter of the boundary is set equal to 2.
[0190] 2. When both P and Q blocks are not predicted by applying CIIP prediction, and at least one of block Q or block P is predicted by applying intra prediction, the boundary strength is determined to be equal to 2.
[0191] 3. When both block P and block Q are not predicted by the application of CIIP prediction and both block Q and P are predicted by the application of inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation).
[0192] 28. The pixel samples included in block Q and block P are filtered by the application of a deblocking filter according to the determined boundary strength.
[0193] Embodiment 8 For a boundary having two sides (where the spatially adjacent blocks on each side are shown as P blocks and Q blocks), the boundary strength is determined as follows.
[0194] 29. When at least one of the blocks P and Q is predicted by the application of intra prediction rather than the application of CIIP prediction (the possibilities include that the P block is predicted by intra prediction rather than by multi-hypothesis prediction and the Q block is predicted by any prediction function, and vice versa), the boundary strength is set equal to 2.
[0195] 30. When both block Q and block P are predicted by the application of inter prediction or by the application of CIIP prediction (the possibilities include that the P block is an inter-block, the Q block is an inter-block, or alternatively, the P block is an inter-block, the Q block is an MH block, or alternatively, the P block is an HM block, the Q block is an inter-block, or alternatively, the P block is an MH block, the Q block is an MH block).
[0196] 1. When at least one of blocks P and Q has a non-zero transform coefficient, the boundary strength parameter of the boundary is set equal to 1. 2. In other cases (when blocks Q and P do not have non-zero transform coefficients), if the absolute difference between the motion vectors used to predict blocks P and Q is greater than or equal to one integer sample, the boundary strength parameter of this boundary is set equal to 1.
[0197] 3. In other cases (when blocks Q and P do not have non-zero transform coefficients and the absolute difference between the motion vectors is less than one sample), if blocks P and Q are predicted based on different reference pictures, or if the number of motion vectors used to predict blocks Q and P is not equal, the boundary strength parameter of this boundary is set equal to 1.
[0198] 4. In other cases (when the above three conditions are evaluated as false), the boundary strength parameter of this boundary is set equal to 0.
[0199] 31. When this boundary is a horizontal boundary and P and Q are located in two different CTUs, 1. When block Q is predicted by applying CIIP prediction, the determined boundary strength is changed as follows.
[0200] 1. If the boundary strength is not equal to 2, the boundary strength is incremented by 1.
[0201] 32. When this boundary is a vertical boundary, or when blocks Q and P are included in the same CTU, 1. When at least one of blocks P and Q is predicted by applying CIIP prediction, the boundary strength parameter of this boundary is adjusted as follows.
[0202] 1. If the boundary strength is not equal to 2, the boundary strength is incremented by 1.
[0203] 33. The pixel samples included in blocks Q and P are filtered by applying a deblocking filter according to the determined boundary strength.
[0204] Embodiment 9 In one example, the boundary strength (Bs) of the boundary of the CIIP block is set to a value of 2, while the boundary strength of the boundary of the sub-blocks within the CIIP is set to a value of 1. When the boundaries of the sub-blocks are not aligned with the 8×8 sample grid, the boundary strength of such edges is set to a value of 0. The 8×8 grid is shown in FIG. 11.
[0205] In another example, the boundary strength of the edge is determined as follows.
[0206] For a boundary having two sides (where the spatially adjacent blocks on each side are shown as a P-block and a Q-block), the boundary strength is derived as follows.
[0207] 34. When the boundary is a horizontal boundary and blocks P and Q are located in different CTUs, 1. When block Q (block Q is shown as the block located downward compared to block P) is predicted by the application of CIIP prediction, the boundary strength is set equal to 2.
[0208] 2. When block Q is not predicted by the application of CIIP prediction and at least one of block Q or block P is predicted by the application of intra prediction, the boundary strength is determined to be equal to 2.
[0209] 3. When block Q is not predicted by the application of CIIP prediction and both blocks Q and P are predicted by the application of inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation).
[0210] 35. In other cases (P and Q corresponding to two sub-blocks within the CIIP block, i.e., when the target boundary is the boundary of the sub-blocks within the CIIP block), 1. When the sub-block boundary is aligned with the 8×8 grid, set the boundary strength to a value of 1.
[0211] 2. In other cases (where the sub-block boundaries are not aligned with the 8×8 grid), set the boundary strength to a value of 0.
[0212] 36. In other cases (when the boundary is a vertical boundary, or when blocks Q and P are included in the same CTU and P and Q are not within the same CIIP block), 1. If at least one of blocks P or Q is predicted by applying CIIP prediction, the boundary strength parameter of the boundary is set equal to 2.
[0213] 2. If neither P nor Q blocks are predicted by applying CIIP prediction, and at least one of block Q or block P is predicted by applying intra prediction, the boundary strength is determined to be equal to 2.
[0214] 3. If neither P nor block Q is predicted by applying CIIP prediction, and both block Q and P are predicted by applying inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation).
[0215] 37. The pixel samples included in blocks Q and P are filtered by applying a deblocking filter according to the determined boundary strength.
[0216] Embodiment 10 In one example, set the boundary strength (Bs) of the boundary of the CIIP block to a value of 2, but set the boundary strength of the boundary of the sub-blocks within the CIIP to a value of 1. When the boundaries of the sub-blocks are not aligned with the 4×4 sample grid, set the boundary strength of such edges to a value of 0. The 4×4 grid is shown in FIG. 12.
[0217] In another example, determine the boundary strength of the edge as follows.
[0218] For a boundary with two sides (where the spatially adjacent blocks on each side are shown as a P-block and a Q-block), the boundary strength is derived as follows.
[0219] 38. When the boundary is a horizontal boundary and blocks P and Q are located in different CTUs, 1. If block Q (the Q-block is shown as the block located downward compared to the P-block) is predicted by the application of CIIP prediction, the boundary strength is set equal to 2.
[0220] 2. If the Q-block is not predicted by the application of CIIP prediction and at least one of block Q or block P is predicted by the application of intra prediction, the boundary strength is determined to be equal to 2.
[0221] 3. If the Q-block is not predicted by the application of CIIP prediction and both blocks Q and P are predicted by the application of inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation).
[0222] 39. In other cases (P and Q corresponding to two sub-blocks within a CIIP block, that is, when the target boundary is a sub-block boundary within a CIIP block), 1. When the sub-block boundary is aligned with a 4×4 grid, set the boundary strength to a value of 1.
[0223] 2. In other cases (the sub-block boundary is not aligned with a 4×4 grid), set the boundary strength to a value of 0.
[0224] 40. In other cases (when the boundary is a vertical boundary, or when blocks Q and P are included in the same CTU and P and Q are not within the same CIIP block), 1. If at least one of block P or Q is predicted by the application of CIIP prediction, the boundary strength parameter of the boundary is set equal to 2.
[0225] 2. When both block P and block Q are not predicted by the application of CIIP prediction, and at least one of block Q or block P is predicted by the application of intra prediction, the boundary strength is determined to be equal to 2.
[0226] 3. When both block P and block Q are not predicted by the application of CIIP prediction, and both block Q and P are predicted by the application of inter prediction, the boundary strength is determined to be less than 2 (the exact value of the boundary strength is determined according to further condition evaluation).
[0227] 41. The pixel samples included in block Q and block P are filtered by the application of a deblocking filter according to the determined boundary strength.
[0228] The present invention further provides the following embodiments.
[0229] Embodiment 1. A method of encoding, wherein the encoding includes decoding or encoding, and the method includes: determining whether a current encoding unit (or encoding block) is predicted by the application of combined inter-intra prediction; when the current encoding unit is predicted by the application of combined inter-intra prediction, setting the boundary strength (Bs) of the boundary of the current encoding unit to a first value; setting the boundary strength (Bs) of the boundary of a sub-encoding unit to a second value, wherein the current encoding unit includes at least two sub-encoding units, and the boundary of the sub-encoding unit is the boundary between the at least two sub-encoding units.
[0230] Embodiment 2. The method further includes a step of performing deblocking when the value of the Bs is greater than zero for the luma component or performing deblocking when the value of the Bs is greater than 1 for the chroma component, where the value of the Bs is one of the first value or the second value, the method of Embodiment 1.
[0231] Embodiment 3. When the current coding unit (or block) is predicted by applying combined inter-intra prediction, the current coding unit is considered to be a unit by intra prediction when performing deblocking, the method of Embodiment 1 or 2.
[0232] FIG. 13 is a block diagram showing an exemplary structure of a device 1300 for the boundary strength derivation process. The device 1300 is configured to execute the above-described method, a determination unit 1302 configured to determine whether at least one of two blocks is predicted by applying combined inter-intra prediction (CIIP), where the two blocks include a first block (block Q) and a second block (block P), and the two blocks are associated with a boundary, the determination unit 1302; a setting unit 1304 configured to set the boundary strength (Bs) of the boundary to a first value when at least one of the two blocks is predicted by applying CIIP, and set the boundary strength (Bs) of the boundary to a second value when none of the two blocks is predicted by applying CIIP, may be included.
[0233] As an example, the setting unit is configured to set the Bs of the boundary to the first value when the first block is predicted by applying CIIP or when the second block is predicted by applying CIIP.
[0234] The machine 2500 may further include a parsing unit (not shown in FIG. 13) configured to parse a bitstream to obtain a flag, where the flag is used to indicate whether at least one of two blocks is predicted by application of CIIP.
[0235] The determination unit 1302 may be further configured to determine whether at least one of two blocks is predicted by application of intra prediction. The setting unit 1304 is configured to set the boundary Bs to a first value when neither of the two blocks is predicted by application of intra prediction and when at least one of the two blocks is predicted by application of CIIP. For example, the first value may be 1 or 2.
[0236] The setting unit 1304 is configured to set the boundary Bs to a second value when neither of the two blocks is predicted by application of intra prediction and when neither of the two blocks is predicted by application of CIIP. For example, when at least one of two blocks (P and Q) has a transform coefficient that is not zero, the second value may be 1.
[0237] Advantages of the Embodiment A deblocking filter for a block predicted by application of multi-hypothesis prediction with a deblocking filter having a medium strength (boundary strength equal to 1).
[0238] When a block is predicted by applying CIIP prediction, the first prediction is obtained by applying inter prediction, the second prediction is obtained by applying intra prediction, and they are later combined. Since the final prediction includes the intra prediction part, more block artifacts are typically observed. Therefore, block artifacts may also exist at the boundaries of the blocks predicted by CIIP prediction. To reduce this problem, according to an embodiment of the present invention, the boundary strength is set to 2, and the opportunity of the deblocking filter for the block edges predicted by applying CIIP prediction is improved.
[0239] A further embodiment of the present invention reduces the required line memory as follows. The line memory is defined as the memory required to store information corresponding to the upper CTU row and the memory required during the processing of the neighboring lower CTU row. For example, to filter the horizontal boundary between two CTU rows, the prediction mode information (intra prediction / inter prediction / multiple hypothesis prediction) of the upper CTU needs to be stored in the list memory. Since three states (intra prediction / inter prediction / multiple hypothesis prediction) are possible to describe the prediction mode of a block, the line memory requirement can be defined as 2 bits per block.
[0240] According to an embodiment of the present invention, when a block (a P block in the embodiment) belongs to the upper CTU row, the deblocking operation only requires information regarding whether the block is predicted by inter prediction or intra prediction (therefore, only two states, which can be stored using 1 bit per block). The reason is as follows.
[0241] When the boundary between the P block and the Q block is a horizontal boundary and the Q block and the P block belong to two different CTUs (in all embodiments, the Q block is below the P block with respect to the P block), the information on whether the P block is predicted by the application of CIIP prediction is not used in the determination of the boundary strength parameter. Therefore, it does not need to be stored. With the help of the embodiments of the present invention, in a hardware implementation, the prediction mode of the P block can be temporarily changed to an inter prediction (when the P block is predicted by CIIP prediction), and the determination of the boundary strength can be performed according to the changed prediction mode. After that (after the determination of the boundary strength), the prediction mode can be changed back to the CIIP prediction. The hardware implementation is not limited to the method described here (changing the prediction mode of the P block at the CTU boundary), and this is merely presented as an example to illustrate that according to the embodiments of the present invention, the information on whether the P block is predicted by CIIP prediction is not essential in the determination of the boundary strength (at the horizontal CTU boundary).
[0242] Therefore, according to the embodiments of the present invention, the required line memory is reduced from 2 bits per block to 1 bit per block. Note that the total line memory that needs to be implemented in hardware is proportional to the picture width and inversely proportional to the minimum block width.
[0243] Note that according to the embodiments, in all the above embodiments, when a block is predicted by the application of CIIP prediction, the first prediction is obtained by the application of inter prediction, the second prediction is obtained by the application of intra prediction, and they are later combined.
[0244] The above embodiments show that when executing the deblocking filter, the CIIP block is considered as an intra block in different ranges. Embodiments 1, 2, and 3 use three different strategies to adjust the boundary strength of the boundary. Embodiment 1 considers the MH block as a complete intra block. Therefore, the condition for setting Bs to 2 is the same as in Table 1.
[0245] Embodiment 2 also assumes that the distortion caused by the MH block is not as high as that within the block. Therefore, the boundary strength condition is first checked for the intra-block and then for the CIIP block. However, when the CIIP block is detected, Bs is still considered to be 2.
[0246] In Embodiment 3, the MH block is partially regarded as an intra-block. When at least one adjacent block of the boundary is an MH block, Bs is increased by only 1. Using the conventional derivation guideline, when Bs is already 2, Bs is not changed.
[0247] FIG. 8 shows the derivation of Bs in VVC Draft 3.0. FIGS. 7, 9, and 10 show the changes to the Bs derivation of Embodiments 1, 2, and 3, respectively.
[0248] It is worth noting that in Embodiments 1 and 2, not only the potential distortion but also the processing logic is reduced. In Embodiments 1 and 2, as long as the P or Q block is an MH block, the checks for the coefficients and motion vectors are no longer necessary, and thus the delay of the conditional check is reduced.
[0249] Embodiments 4, 5, and 6 are modifications of Embodiments 1, 2, and 3, respectively, and the line buffer memory is considered. The core changes to them compared to Embodiments 1, 2, and 3 are that when the two side Ps and Qs are located in different CTUs and the edge is horizontal, the check of the MH block is performed asymmetrically. That is, the P-side block (i.e., the upper side) is not checked, but only the Q-side (i.e., the lower side) is checked. In this way, no additional line buffer memory is allocated to store the CIIP flag of the P-side block located in another CTU.
[0250] In addition to the six embodiments described above, one further feature of the MH block may be that the MH block does not necessarily have to be consistently considered as an intra block. In one example, when searching for the motion vector predictor of the current block, if its neighboring blocks are MH blocks, the motion vectors of these MH blocks can be considered as motion vector predictors. In this case, the inter prediction information of the HM block is used, and thus the HM block is no longer considered an intra block. In another example, when constructing the MPM list of an intra block, the neighboring MH blocks of the current block can be considered not to contain intra information. Therefore, when checking the availability of those MH blocks for constructing the MPM list of the current block, they are labeled as unavailable. Note that the MH blocks mentioned in this paragraph are not limited only to the MH blocks used to determine the Bs value of the deblocking filter.
[0251] In addition to the six embodiments described above, one further feature of the MH block may be that the MH block has to be consistently considered as an intra block. In one example, when searching for the motion vector predictor of the current block, if its neighboring blocks are MH blocks, the motion vectors of these MH blocks are excluded from the motion vector predictor. In this case, the inter prediction information of the HM block is not used, and thus the HM block is considered an intra block. In another example, when constructing the MPM list of an intra block, the neighboring MH blocks of the current block can be considered to contain intra information. Therefore, when checking the availability of those MH blocks for constructing the MPM list of the current block, they are labeled as available. Note that the MH blocks mentioned in this paragraph are not limited only to the MH blocks used to determine the Bs value of the deblocking filter.
[0252] The following is an explanation of the application of the encoding method and decoding method as shown in the above embodiments, and the systems using them.
[0253] FIG. 14 is a block diagram showing a content supply system 3100 that realizes a content delivery service. This content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types, etc.
[0254] The capture device 3102 may generate data and encode the data by an encoding method as shown in the above embodiments. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown in the figure), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a Pad, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination of these, etc. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video camera system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0255] In the content supply system 3100, the terminal device 310 receives and plays back the encoded data. The terminal device 3106 can be a device with data reception and restoration capabilities, such as a smartphone or Pad 3108 having the ability to decode the above-mentioned encoded data, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0256] In a terminal device equipped with a display, a smartphone, or Pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its own display. In a terminal device without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is connected to receive and display the decoded data.
[0257] When each device in this system performs encoding or decoding, as shown in the above-described embodiments, a picture encoding device or a picture decoding device can be used.
[0258] FIG. 15 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol progress unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, the Real Time Streaming Protocol (RTSP), the Hyper Text Transfer Protocol (HTTP), the HTTP Live Streaming protocol (HLS), MPEG-DASH, the Real-time Transport protocol (RTP), the Real Time Messaging Protocol (RTMP), or any combination of those of any kind, etc.
[0259] After the protocol progress unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0260] By inverse multiplexing processing, a video elementary stream (ES), an audio ES, and any subtitles are generated. A video decoder 3206 including a video decoder 30 as described in the above embodiments decodes the video ES by the decoding method as shown in the above embodiments to generate video frames, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in FIG. Y) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in FIG. Y) before being supplied to the synchronization unit 3212.
[0261] The synchronization unit 3212 synchronizes the video frames and the audio frames, and supplies the video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be encoded within the syntax using time stamps related to the presentation of the encoded audio and visual data, and time stamps related to the delivery of the data stream itself.
[0262] When subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frames and the audio frames, and supplies the video / audio / subtitle to the video / audio / subtitle display 3216.
[0263] The present invention is not limited to the above-described system, and any of the picture encoding device or the picture decoding device in the above embodiments can be incorporated into other systems, such as vehicle systems.
[0264] Embodiments of the present invention have been mainly described based on video coding. It should be noted, however, that embodiments of the coding system 10, the encoder 20, and the decoder 30 (and correspondingly the system 10), as well as other embodiments described herein, may be configured for still image processing or coding, i.e., processing or coding of individual pictures independent of any preceding or consecutive pictures as in video coding. Generally, when picture processing coding is limited to a single picture 17, it is not necessary that only the inter prediction units 244 (encoder) and 344 (decoder) be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 may be equally used for still image processing, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filters 220, 320, and entropy coding 270 and entropy decoding 304.
[0265] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described herein with reference to, for example, the encoder 20 and the decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored on a computer-readable medium as one or more instructions or code or transmitted via a communication medium and executed by a processing unit based on hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that enables transfer of a computer program from one place to another, for example, in accordance with a communication protocol. In this way, the computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium accessible by one or more computers or one or more processors for reading instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0266] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be accessed by a computer and is capable of storing the required program code in the form of instructions or data structures. Also, any connection can be properly called a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transient tangible storage media. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data and disc optically reproduces data by laser. The foregoing combinations should also be included within the scope of computer-readable media.
[0267] The commands may be executed by one or more processors such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the term "processor" as used herein may represent any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. Further, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured to be incorporated in an encoder and decoder or combined codec. Further, the techniques may be implemented entirely with one or more circuits or logic elements.
[0268] The techniques of the present disclosure may be implemented in a variety of apparatuses or devices including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). The various components, modules, or units have been described in this disclosure to emphasize functional aspects of an apparatus configured to execute the disclosed techniques, but implementation by different hardware units is not necessarily required. Rather, as described above, the various units may be combined within a codec hardware unit in combination with appropriate software and / or firmware, or provided by a collection of interoperable hardware units including one or more processors as described above.
Description of the Signs
[0269] 12 Source device 14 Destination device 18 Video source 20 Video encoder 22 Output interface 28 Input interface 30 Video decoder 32 - display device
Claims
1. 1. A method of video encoding, comprising: determining whether at least one of two blocks of an image of a video is predicted by combined inter-intra prediction (CIIP), the two blocks including a first block and a second block, and a boundary exists between the first block and the second block; setting a boundary strength (Bs) of the boundary to a first value when at least one of the two blocks is predicted by CIIP, the first value being 2; encoding a flag into a video bitstream, the flag being used to indicate whether at least one of the two blocks is predicted by CIIP; The method includes:
2. setting the boundary strength (Bs) of the boundary to a second value when neither of the two blocks is predicted by CIIP; The method of claim 1 further comprising:
3. determining whether at least one of the two blocks is predicted by intra prediction; 3. The method of claim 1 or 2, wherein the boundary strength (Bs) of the boundary is set based on determining whether at least one of the two blocks is predicted by CIIP and based on determining whether at least one of the two blocks is predicted by intra prediction.
4. The step of setting the Bs of the boundary to the first value includes:
4. The method of claim 3, further comprising setting the Bs of the boundary to the first value when neither of the two blocks is predicted by intra prediction and when at least one of the two blocks is predicted by CIIP.
5. 4. The method of claim 3, further comprising setting the Bs of the boundary to a second value when neither of the two blocks is predicted by intra prediction and neither of the two blocks is predicted by CIIP.
6. 3. The method of claim 2, wherein the second value is one when at least one of the two blocks has non-zero transform coefficients and when neither of the two blocks is predicted by CIIP.
7. The method comprises: performing a deblocking filter on a boundary between the luma components of the first block and the second block; The method of any one of claims 1 to 6, further comprising:
8. The method comprises: when the value of Bs is greater than 1, performing a deblocking filter on the boundary between the chroma components of the first block and the second block; The method of any one of claims 1 to 7, further comprising:
9. 1. A method of video decoding, comprising: Parsing a video bitstream to obtain a flag, the flag being used to indicate whether at least one of two blocks of an image in the video is predicted by combined inter-intra prediction (CIIP), the two blocks including a first block and a second block, and a boundary exists between the first block and the second block; determining whether at least one of the two blocks is predicted by CIIP; setting a boundary strength (Bs) of the boundary to a first value when at least one of the two blocks is predicted by CIIP, the first value being 2; The method includes:
10. setting the boundary strength (Bs) of the boundary to a second value when neither of the two blocks is predicted by CIIP; 10. The method of claim 9 further comprising:
11. determining whether at least one of the two blocks is predicted by intra prediction; 11. The method according to claim 9 or 10, wherein the boundary strength (Bs) of the boundary is set based on determining whether at least one of the two blocks is predicted by CIIP and based on determining whether at least one of the two blocks is predicted by intra prediction.
12. The step of setting the Bs of the boundary to the first value includes:
12. The method of claim 11, comprising setting the Bs of the boundary to the first value when neither of the two blocks is predicted by intra prediction and when at least one of the two blocks is predicted by CIIP.
13. 12. The method of claim 11, wherein when neither of the two blocks is predicted by intra prediction and neither of the two blocks is predicted by CIIP, the Bs of the boundary is set to a second value.
14. 11. The method of claim 10, wherein the second value is one when at least one of the two blocks has non-zero transform coefficients and when neither of the two blocks is predicted by CIIP.
15. The method comprises: performing a deblocking filter on a boundary between the luma components of the first block and the second block; The method according to any one of claims 9 to 14, further comprising:
16. The method comprises: when the value of Bs is greater than 1, performing a deblocking filter on the boundary between the chroma components of the first block and the second block; The method according to any one of claims 9 to 15, further comprising:
17. 1. An encoder comprising: one or more processors; a computer-readable medium communicatively coupled to the one or more processors; Including, The computer readable medium stores instructions which, when executed by the one or more processors, cause the encoder to perform the method of any one of claims 1 to 8.
18. A decoder comprising: one or more processors; a computer-readable medium communicatively coupled to the one or more processors; Including, The computer readable medium storing instructions which, when executed by the one or more processors, cause the decoder to perform the method of any one of claims 9 to 16.
19. An encoder comprising processing circuitry for carrying out the method according to any one of claims 1 to 8.
20. A decoder including processing circuitry for carrying out the method according to any one of claims 9 to 16.
21. A computer program comprising a program code for carrying out the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Method and apparatus for encoding hybrid intra-inter coded blocks
JP2007503775A
Method and apparatus for encoding hybrid intra-inter coded blocks
US8085845B2
Video coding using combined inter and intra predictors
US9374578B1
Determining application of deblocking filtering to palette coded blocks in video coding
WO2015191834A1