Harmonization of separate merge lists for sub-block merge candidates and intra-inter techniques for video coding
By employing separate merge lists and multiple hypothesis prediction techniques with conditional flag signaling, the method optimizes video encoding and decoding, enhancing compression efficiency and maintaining quality in limited bandwidth and memory-constrained environments.
Patent Information
- Application Number
- JP2025146315
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-10-27
- Filing Date
- 2025-09-03
- Publication Date
- 2026-01-08
AI Technical Summary
Existing video coding technologies face challenges in achieving efficient compression ratios with minimal quality sacrifice, particularly in limited bandwidth communication networks and memory-constrained systems, necessitating improved methods for intra and inter-mode video encoding and decoding.
The implementation of separate merge lists for sub-block merging candidates and multiple hypothesis prediction techniques, with conditional signaling of control flags to determine the use of these methods, allowing for optimized video encoding and decoding processes.
Enhances video compression efficiency by allowing decoders to make informed decisions on using separate merge lists and multiple hypothesis prediction, thereby improving compression ratios without compromising picture quality.
Smart Images

Figure 2026002851000001_ABST
Abstract
Description
[Technical Field]
[0001] The embodiments of the present disclosure relate generally to the field of picture processing, and more specifically to the interplay of two methods. More specifically, the embodiments propose a method for the cooperation and signaling of separate merge lists for sub-block merging candidates and multiple hypothesis prediction for intra- and inter-mode techniques. [Background technology]
[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat, video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems, and camcorders for security applications.
[0003] The amount of video data required to represent even a relatively short video can be considerable, which can cause difficulties when the data is to be streamed or otherwise communicated over communication networks with limited bandwidth capacity. Thus, video data is typically compressed before being communicated over modern telecommunications networks. Because memory resources may be limited, the size of the video can also be an issue when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. Due to limited network resources and an ever-increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little or no sacrifice in picture quality are desirable. Summary of the Invention
[0004] SUMMARY The present disclosure provides apparatus and methods for encoding and decoding video.
[0005] The present invention aims to harmonize the use and signaling of multiple hypothesis prediction for intra and inter modes with separate merge lists for sub-block merging candidates.
[0006] These and other objects are achieved by the subject matter of the independent claims. Further embodiments are evident from the dependent claims, the description and the figures.
[0007] According to a first aspect of the present invention, there is provided a method for encoding video data into a bitstream. The method comprises using a first technique and / or a second technique. The first technique comprises using separate merge lists for sub-block merging candidates. The second technique comprises multiple hypothesis prediction for intra and inter modes. The method comprises transmitting a first control flag in the bitstream for a coding block, and transmitting or not transmitting a second control flag in the bitstream depending on whether the first technique is used for the coding block. The first control flag indicates whether the first technique should be used. The second control flag indicates whether the second technique should be used.
[0008] According to a second aspect of the present invention, there is provided a method for decoding video data received in a bitstream. The method includes using a first technique and / or a second technique. The first technique includes using a separate merge list for sub-block merging candidates. The second technique includes multiple hypothesis prediction for intra and inter modes. The method includes receiving a first control flag from the bitstream for a coding block, the first control flag indicating whether the first technique should be used, and receiving a second control flag from the bitstream depending on whether the first technique is used for the coding block. The second control flag indicates whether the second technique should be used.
[0009] It is a particular approach of the present invention to conditionally generate and transmit a second control flag indicating whether or not to use multiple hypothesis prediction for intra and inter modes once a decision is made as to whether or not to use the separate merge list technique for sub-block merging candidates. On the other hand, even if the second control flag is only conditionally transmitted, the decoder is still capable of deciding regarding the use of multiple hypothesis prediction for intra and inter modes and the separate merge list technique for sub-block merging candidates.
[0010] In a possible embodiment of the method according to the first aspect per se, the second control flag is transmitted if and only if the first technique is not used for said coding block.
[0011] In possible embodiments of the method according to the aforementioned implementations or the first aspect per se, the second control flag is transmitted if the coding block is coded in merge mode. In possible embodiments of the method according to the aforementioned implementations of the first aspect per se, the second control flag is not transmitted if the coding block is not coded in merge mode. In possible embodiments of the method according to the aforementioned implementations of the first aspect or the first aspect per se, the second control flag is transmitted if the coding block is coded in skip mode. In possible embodiments of the method according to the aforementioned implementations of the first aspect per se, the second control flag is not transmitted if the coding block is not coded in skip mode or merge mode.
[0012] Thus, signaling according to the particular approach of the present invention is applicable in either merge mode or skip mode, or in both merge and skip modes.
[0013] In a possible embodiment of the method according to the second aspect per se, the second control flag is received only if the first technique is not used.
[0014] Therefore, a decoder can directly infer from the presence of the second control flag in the received bitstream that the separate merge list technique will not be used for the current coding block. Thus, evaluation of the first control flag indicating whether to use the separate merge list technique for a sub-block candidate is only necessary if the second control flag is not included in the received bitstream.
[0015] In possible embodiments of the method according to the aforementioned implementations of the second aspect or the second aspect itself, the second control flag is received if the coding block is coded in merge mode. In possible embodiments of the method according to the aforementioned implementations of the second aspect, the second control flag is not received if the coding block is not coded in merge mode. In possible embodiments of the method according to the aforementioned embodiments of the second aspect or the second aspect itself, the second control flag is received if the coding block is coded in skip mode. In possible embodiments of the method according to the aforementioned embodiments of the second aspect, the second control flag is not received if the coding block is not coded in skip mode.
[0016] The encoding and decoding methods defined in the claims, the description and the figures can be performed by an encoding device and a decoding device, respectively.
[0017] According to a third aspect, the present invention relates to an encoder comprising processing circuitry for carrying out the method according to the first aspect per se or any of its embodiments.
[0018] According to a fourth aspect, the present invention relates to a decoder comprising processing circuitry for carrying out the method according to the second aspect per se or any of its embodiments.
[0019] According to a fifth aspect, the present invention relates to an encoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processors, configuring the encoder to perform a method according to the first aspect itself or any of its embodiments.
[0020] According to a sixth aspect, the present invention relates to a decoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processors, configuring the decoder to perform the method of the second aspect itself or any of its embodiments.
[0021] According to a seventh aspect, the present invention relates to an encoder for video encoding of video data into a bitstream, the encoder comprising means for performing a first technique and / or a second technique. The first technique comprises using separate merge lists for sub-block merging candidates. The second technique comprises multiple hypothesis prediction for intra and inter modes. The encoder further comprises means for transmitting a first control flag in the bitstream for a coding block, and means for transmitting or not transmitting a second control flag in the bitstream depending on whether the first technique is used for the coding block. The first control flag indicates whether the first technique should be used. The second control flag indicates whether the second technique should be used.
[0022] According to an eighth aspect, the present invention relates to a decoder for video decoding of video data received in a bitstream. The decoder comprises means for performing a first technique and / or a second technique. The first technique comprises using separate merge lists for sub-block merging candidates. The second technique comprises multiple hypothesis prediction for intra and inter modes. The decoder further comprises means for receiving, for a coding block, a first control flag from the bitstream, the first control flag indicating whether the first technique should be used, and means for receiving, from the bitstream, a second control flag depending on whether the first technique is used for the coding block. The second control flag indicates whether the second technique should be used.
[0023] According to a further aspect, the invention relates to a non-transitory computer readable medium carrying program code which, when executed by a computing device, causes the computing device to perform a method according to the first or second aspect.
[0024] Possible embodiments of the encoder and decoder according to the seventh and eighth aspects correspond to possible embodiments of the methods according to the first and second aspects.
[0025] An apparatus for encoding or decoding a video stream may include a processor and a memory, the memory storing instructions that cause the processor to perform the encoding or decoding method.
[0026] For each of the encoding or decoding methods disclosed herein, a computer-readable storage medium is proposed, the storage medium storing instructions that, when executed, cause one or more processors to encode or decode video data, the instructions causing the one or more processors to perform the respective encoding or decoding method.
[0027] Furthermore, for each of the encoding or decoding methods disclosed herein, a computer program product is proposed, the computer program product comprising a program code for carrying out the respective method.
[0028] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0029] In addition, the present invention further provides the following embodiments. The encoded bitstream of video data has a plurality of syntax elements including a first control flag and a second control flag conditionally signaled based on the first control flag, the first control flag indicating whether a first technique is used, the first technique (S101) including using a separate merge list for sub-block merging candidates, the second control flag indicating whether a second technique is used, the second technique (S103) including multiple hypothesis prediction for intra and inter modes. A computing storage medium is provided that stores a bitstream to be decoded by a video decoding device, the bitstream having a number of coding blocks of an image or video signal and a number of syntax elements including a first control flag and a second control flag that is conditionally signaled based on the first control flag, the first control flag indicating whether a first technique is used, the first technique (S101) including using a separate merge list for sub-block merging candidates, the second control flag indicating whether a second technique is used, the second technique (S103) including multiple hypothesis prediction for intra and inter modes. A non-transitory computer readable storage medium is provided that stores video information generated by using any of the encoding methods set forth in claims 1 to 6 of the pending claims.
[0030] In the following, embodiments of the invention will be described in more detail with reference to the accompanying figures and drawings. [Brief explanation of the drawings]
[0031] [Figure 1A] 1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Figure 1B] FIG. 2 is a block diagram illustrating another example of a video coding system configured to implement embodiments of the present invention. [Figure 2] 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present invention; [Figure 3] 1 is a block diagram illustrating an exemplary structure of a video decoder configured to implement embodiments of the present invention. [Figure 4] 1 is a block diagram illustrating an example of an encoding device or a decoding device. [Figure 5] FIG. 10 is a block diagram illustrating another example of an encoding device or a decoding device. [Figure 6] 3 is a flowchart illustrating an exemplary encoding method according to an embodiment of the present invention. [Figure 7] 4 is a flowchart illustrating an exemplary decoding method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0032] In the following, the same reference signs refer to the same or at least functionally equivalent features, unless expressly stated otherwise.
[0033] In the following description, reference is made to the accompanying figures which form part of this disclosure and which show, by way of illustration, specific aspects of embodiments of the invention or in which embodiments of the invention may be used. It is understood that embodiments of the invention may be used in other ways and may include structural or logical changes not depicted in the figures. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0034] For example, it is understood that disclosure related to a described method also applies to a corresponding device or system configured to perform the method, and vice versa. For example, when one or more specific method steps are described, a corresponding device may include one or more units, e.g., functional units, that perform the described one or more method steps (e.g., one unit that performs one or more steps, or multiple units that each perform one or more of the steps), even if such one or more units are not explicitly described or shown. On the other hand, for example, when a specific apparatus is described based on one or more units, e.g., functional units, a corresponding method may include one or more steps that perform the function of one or more units (e.g., one step that performs the function of one or more units, or multiple steps that each perform the function of one or more of the units), even if such one or more steps are not explicitly described or shown. Furthermore, it is understood that features of various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.
[0035] Video coding typically refers to processing a sequence of pictures that form a video or a video sequence. Instead of the term "picture," the terms "frame" or "image" are sometimes used synonymously in the field of video coding. Video coding (or coding in general) has two parts: video encoding and video decoding. Video encoding occurs at the source side and typically involves processing original video pictures (e.g., by compression) to reduce the amount of data needed to represent the video picture (for more efficient storage and / or transmission). Video decoding occurs at the destination side and typically involves the reverse process compared to the encoder to reconstruct the video picture. Embodiments referring to "coding" a video picture (or pictures in general) should be understood to relate to "encoding" or "decoding" the video picture or each video sequence. The combination of the encoding and decoding parts is also called CODEC (Coding and Decoding).
[0036] In the case of lossless video coding, the original video picture can be reconstructed, i.e., the reconstructed video picture has the same quality as the original video picture (assuming no transmission or other data loss during storage or transmission). In the case of lossy video coding, further compression, for example by quantization, is performed to reduce the amount of data representing the video picture, and the video picture cannot be perfectly reconstructed at a decoder, i.e., the quality of the reconstructed video picture is reduced or deteriorated compared to the quality of the original video picture.
[0037] Some video coding standards belong to the group of "lossy hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding that applies quantization in the transform domain). Each picture of a video sequence is usually divided into a set of non-overlapping blocks, and coding is usually performed at the block level. That is, at the encoder, video is usually processed, i.e., encoded, at the block (video block) level, for example, by generating a predictive block using spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, subtracting the predictive block from a current block (the block currently being processed / to be processed) to obtain a residual block, transforming the residual block, and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression); whereas at the decoder, an inverse process compared to the encoder is applied to the encoded or compressed block to reconstruct the current block for display. Additionally, the encoder replicates the decoder processing loop so that both generate the same predictions (e.g., intra and inter predictions) and / or reconstructions for processing, i.e., coding, subsequent blocks.
[0038] In the following, embodiments of a video coding system 10, a video encoder 20, and a video decoder 30 are described based on FIGS.
[0039] 1A is a schematic block diagram of an example coding system 10, e.g., video coding system 10 (or, for short, coding system 10), that may employ techniques of the present application. A video encoder 20 (or, for short, encoder 20) and a video decoder 30 (or, for short, decoder 30) of video coding system 10 represent examples of devices that may be configured to perform techniques according to various examples described herein.
[0040] As shown in FIG. 1A, coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 that decodes the encoded picture data 21.
[0041] The source device 12 comprises an encoder 20 and may further comprise, namely, optionally, a picture source 16 , a pre-processor (or pre-processing unit) 18 , for example a picture pre-processor 18 , and a communication interface or unit 22 .
[0042] Picture source 16 may comprise or be any kind of picture capture device, e.g., a camera that captures real-world pictures, and / or any kind of picture generation device, e.g., a computer graphics processor that generates computer-animated pictures, or any other device that acquires and / or provides real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). Picture source may be any kind of memory or storage that stores any of the above pictures.
[0043] To distinguish from the preprocessor 18 and the processing performed by the preprocessing unit 18 , the pictures or picture data 17 may also be referred to as raw pictures or raw picture data 17 .
[0044] The pre-processor 18 is configured to receive the (raw) picture data 17 and perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. The pre-processing performed by the pre-processor 18 may comprise, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It may be understood that the pre-processing unit 18 may be any component.
[0045] Video encoder 20 is configured to receive pre-processed picture data 19 and to provide encoded picture data 21 (further details are described below, eg, with reference to FIG. 2).
[0046] The communications interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and to transmit the encoded picture data 21 (or any further processed version thereof) to another device, e.g., the destination device 14 or some other device, via the communications channel 13 for storage or direct reconstruction.
[0047] The destination device 14 comprises a decoder 30 (eg, a video decoder 30), and may further comprise, optionally, a communications interface or communications unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0048] The communications interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof), for example directly from the source device 12 or from some other source, for example a storage device, for example an encoding picture data storage device, and to provide the encoded picture data 21 to the decoder 30.
[0049] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g., a direct wired or wireless connection, or via any type of network, e.g., a wired or wireless network or any combination thereof, or any type of private and public network, or any combination thereof.
[0050] The communications interface 22 may be configured, for example, to package the encoded picture data 21 into a suitable format, e.g., packets, and / or process the encoded picture data using any type of transmission encoding or processing for transmission over a communications link or network.
[0051] The communications interface 28, which forms the counterpart of the communications interface 22, may for example be configured to receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or unpackaging to obtain the encoded picture data 21.
[0052] Both communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces, as indicated by the arrow for communication channel 13 pointing from source device 12 to destination device 14 in FIG. 1A, or as bidirectional communication interfaces, e.g., configured to send and receive messages, e.g., to set up connections, acknowledge and exchange any other information related to the communication link and / or data transmission, e.g., encoded picture data transmission.
[0053] The decoder 30 is configured to receive the encoded picture data 21 and to provide decoded picture data 31 or decoded pictures 31 (further details are described below, for example, based on Figure 3 or Figure 5).
[0054] Post-processor 32 of destination device 14 is configured to post-process decoded picture data 31 (also called reconstructed picture data), e.g., decoded picture 31, to obtain post-processed picture data 33, e.g., post-processed picture 33. The post-processing performed by post-processing unit 32 may include, e.g., color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing that, e.g., prepares decoded picture data 31 for display, e.g., by display device 34.
[0055] A display device 34 of destination device 14 is configured to receive the post-processed picture data 33 for displaying the picture, e.g., to a user or viewer. Display device 34 may be or include any type of display, e.g., an internal or external display or monitor, that presents the reconstructed picture. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, an LCD (Light Cosine Scattering) display, a digital light processor (DLP), or any other type of display.
[0056] 1A depicts source device 12 and sending device 14 as separate devices, an embodiment of the device may have both or both functions, i.e., source device 12 or corresponding functions and destination device 14 or corresponding functions. In such an embodiment, source device 12 or corresponding functions and destination device 14 or corresponding functions may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.
[0057] As will be clear to those skilled in the art based on the description, the functionality of the different units, or the presence and (exact) division of functionality within the source device 12 and / or destination device 14 shown in FIG. 1A, may vary depending on the actual device and application.
[0058] Encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder), or both encoder 20 and decoder 30, may be implemented using processing circuitry shown in FIG. 1B , such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated to video coding, or any combination thereof. Encoder 20 may be implemented using processing circuitry 46 to implement various modules discussed with respect to encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented using processing circuitry 46 to implement various modules discussed with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuitry may be configured to perform various operations discussed below. 5, where the techniques are implemented partially in software, a device may store instructions for the software in a suitable, non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either video encoder 20 and video decoder 30 may be incorporated as part of a combined encoder / decoder (CODEC) in a single device, for example, as shown in FIG. 1B.
[0059] The source device 12 and the destination device 14 may comprise any of a wide range of devices, including any type of portable or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television set, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiver device, a broadcast transmitter device, etc., and may or may not use any type of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0060] 1A is merely an example, and the techniques herein may be applied to video coding settings (e.g., video encoding or video decoding) that do not necessarily involve any data communication between encoding and decoding devices. In other examples, data may be retrieved from local memory, streamed over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data and store it in memory and / or read data from memory and decode it.
[0061] For convenience of description, embodiments of the present invention are described herein with reference to, for example, High-Efficiency Video Coding (HEVC) or reference software for Versatile Video coding (VVC), a next-generation video coding standard developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG) Joint Collaboration Team on Video Coding (JCT-VC); those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.
[0062] Encoder and encoding method FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes an input unit 201 (or input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoding picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output unit 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder conforming to a hybrid video codec.
[0063] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 are sometimes said to form a forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoding picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are sometimes said to form a backward signal path of the video encoder 20, which corresponds to the signal path of a decoder (see video decoder 30 of FIG. 3). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoding picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also said to form a “built-in decoder” of the video encoder 20.
[0064] Picture and Picture Partitioning (Picture and Block) The encoder 20 may be configured to receive, for example, via an input 202, a picture 17 (or picture data 17), for example a picture in a sequence of pictures forming a video or a video sequence. The received picture or picture data may be a preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to the picture 17. The picture 17 may also be called the current picture or the picture to be coded (particularly in video coding, to distinguish the current picture from other pictures, for example previously encoded and / or decoded pictures of the same video sequence, i.e., the video sequence that also includes the current picture).
[0065] A (digital) picture is, or may be considered as, a two-dimensional array or matrix of samples with intensity values. The samples in the array may also be called pixels (short for picture element) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are usually used; that is, a picture may be represented or contain three sample arrays. In an RBG format or color space, a picture has corresponding red, green, and blue sample arrays. However, in video coding, each pixel is usually represented in a luminance and chrominance format or color space, e.g., YCbCr, with a luminance component denoted by Y (sometimes L is used instead) and two chrominance components denoted by Cb and Cr. The luminance (or luma for short) component Y represents brightness or gray-level intensity (e.g., as in a grayscale picture), while the two chrominance (or chroma for short) components Cb and Cr represent chromaticity or color information components. Thus, a picture in YCbCr format has a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in RGB format may be converted or transformed to YCbCr format, and vice versa, a process also known as color conversion or transformation. If a picture is monochrome, the picture may only have a luminance sample array. Thus, a picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0066] Embodiments of video encoder 20 may include a picture partitioning unit (not shown in FIG. 2) configured to partition picture 17 into multiple (usually non-overlapping) picture blocks 203. These blocks may also be called root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size for all pictures of a video sequence and a corresponding grid defining the block size, or to vary the block size among pictures or subsets or groups of pictures and divide each picture into corresponding blocks.
[0067] In further embodiments, the video encoder may be configured to directly receive blocks 203 of picture 17, e.g., one, some, or all of the blocks forming picture 17. Picture blocks 203 may also be referred to as current picture blocks or the picture to be coded.
[0068] Like picture 17, picture block 203 may also be, or may be considered to be, a two-dimensional array or matrix of samples having intensity values (sample values), albeit with smaller dimensions than picture 17. That is, block 203 may have, for example, one sample array (e.g., a luma array in the case of a monochrome picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma and two chroma arrays in the case of a color picture 17), or any other number and / or type of array depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 define the size of block 203. Thus, a block may be, for example, an M×N (N rows and M columns) array of samples, or an M×N array of transform coefficients.
[0069] The embodiment of video encoder 20 shown in FIG. 2 may be configured to encode picture 17 block by block, eg, encoding and prediction is performed for each block 203.
[0070] Residual calculation The residual calculation unit 204 may be configured to calculate the residual block 205 (also referred to as the residual 205) based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 will be described later), for example, by subtracting sample values of the prediction block 265 from sample values of the picture block 203 on a sample-by-sample (pixel-by-pixel) basis to obtain the residual block 205 in the sample domain.
[0071] conversion The transform processing unit 206 may be configured to apply a transform, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207, also called transform residual coefficients, may represent the residual block 205 in the transform domain.
[0072] Transform processing unit 206 may be configured to apply an integer approximation of a DCT / DST, such as the transform defined for H.265 / HEVC. Compared to an orthogonal DCT transform, such an integer approximation is typically scaled by a specific factor. To preserve the norm of the residual blocks processed by the forward and inverse transforms, additional scaling factors are applied as part of the transform process. The scaling factors are typically selected based on specific constraints, such as scaling factors being powers of two for shift operations, bit depths of the transform coefficients, a trade-off between accuracy and implementation cost, etc. For example, specific scaling factors may be defined for the inverse transform, e.g., by inverse transform processing unit 212 (and a corresponding inverse transform, e.g., by inverse transform processing unit 312 in video decoder 30), and corresponding scaling factors for the forward transform, e.g., by transform processing unit 206 in encoder 20, may be defined accordingly.
[0073] An embodiment of video encoder 20 (respectively, transform processing unit 206) may be configured to output transform parameters, e.g., a type of transform or multiple transforms, e.g., directly or encoded or compressed by entropy encoding unit 270, so that video decoder 30 may receive the transform parameters and use them for decoding.
[0074] Quantization The quantization unit 208 may be configured to quantize the transform coefficients 207, for example by applying scalar quantization or vector quantization, to obtain quantized coefficients 209. The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209.
[0075] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be changed by adjusting a quantization parameter (QP). For example, for scalar quantization, different scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may, for example, be an index into a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to finer quantization (small quantization step size) and a large quantization parameter may correspond to coarser quantization (large quantization step size), or vice versa. Quantization may involve division by a quantization step size, and corresponding and / or inverse dequantization, e.g., by the inverse quantization unit 210, may involve multiplication by the quantization step size. Embodiments according to some standards, e.g., HEVC, may be configured to use a quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of a formula involving division. Additional scaling factors may be introduced for quantization and dequantization to restore norms of the residual block that may be changed due to scaling used in the fixed-point approximation of the formula for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, customized quantization tables may be used and conveyed from the encoder to the decoder, e.g., in the bitstream. Quantization is a lossy operation, and loss increases with increasing quantization step size.
[0076] Embodiments of video encoder 20 (individually, quantization unit 208) may be configured to output a quantization parameter (QP), e.g., directly or encoded by entropy encoding unit 270, such that video decoder 30 may receive and apply the quantization parameter for decoding.
[0077] inverse quantization Inverse quantization unit 210 is configured to apply the inverse quantization of quantization unit 208 to the quantized coefficients to obtain inverse quantized coefficients 211, e.g., by applying the inverse of the quantization scheme applied by quantization unit 208, based on or using the same quantization step size as quantization unit 208. The inverse quantized coefficients 211, also referred to as inverse quantized residual coefficients 211, may correspond to transform coefficients 207, although they are typically not the same as transform coefficients due to loss in quantization.
[0078] Inverse transformation The inverse transform processing unit 212 may be configured to apply an inverse transform of the transform applied by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213.
[0079] Reconstruction The reconstruction unit 214 (e.g., adder or summator 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the sample domain, for example, by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265 sample by sample.
[0080] filtering The loop filter unit 220 (or "loop filter" 220 for short) is configured to filter the reconstructed block 215 to obtain a filtered block 221, or generally, to filter reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although the loop filter unit 220 is illustrated in FIG. 2 as being an in-loop filter, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as a filtered reconstructed block 221.
[0081] Embodiments of video encoder 20 (individually, loop filter unit 220) may be configured to output loop filter parameters (e.g., sample adaptive offset information), e.g., directly or encoded by entropy encoding unit 270, such that decoder 30 may receive and apply the same loop filter parameters or each loop filter for decoding.
[0082] Decoding Picture Buffer The decoding picture buffer (DPB) 230 may be a memory that stores reference pictures, or generally reference picture data, for encoding video data by the video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoding picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoding picture buffer 230 may further be configured to store other previously filtered blocks, e.g., previously reconstructed filtered blocks 221, of the same current picture or of a different picture, e.g., a previously reconstructed picture, to provide a complete, previously reconstructed, i.e., decoded, picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples), e.g., for inter-prediction. The decoding picture buffer (DPB) 230 may also be configured to store one or more unfiltered reconstructed blocks 215, for example, if the reconstructed blocks 215 have not been filtered by the loop filter unit 220, or generally, unfiltered reconstructed samples, or any other further processed version of the reconstructed blocks or samples.
[0083] Mode Selection (Partitioning and Prediction) The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, e.g., the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data, e.g., filtered and / or unfiltered reconstructed samples or blocks of the same (current) picture and / or from one or more previously decoded pictures, e.g., from the decoding picture buffer 230 or another buffer (e.g., a line buffer, not shown). The reconstructed picture data is used as reference picture data for prediction, e.g., inter prediction or intra prediction, to obtain a prediction block 265 or predictor 265.
[0084] The mode selection unit 260 may be configured to determine or select a partitioning for the current block prediction mode (not including the partitioning) and a prediction mode (e.g., intra or inter prediction mode) and generate a corresponding prediction block 265 that is used for calculating the residual block 205 and for reconstructing the reconstructed block 215.
[0085] Embodiments of the mode selection unit 260 may be configured to select partitioning and prediction modes (e.g., from those supported by or available for the mode selection unit 260) that result in the best match, i.e., in other words, the smallest residual (minimum residual means better compression for transmission or storage), or the smallest signaling overhead (minimum signaling overhead means better compression for transmission or storage), or that consider or balance both. The mode selection unit 260 may be configured to determine the partitioning and prediction modes based on rate-distortion optimization (RDO), i.e., to select the prediction mode that results in the lowest rate-distortion. Terms such as “best,” “minimum,” and “optimized” in this context do not necessarily refer to the overall “best,” “minimum,” “optimized,” etc., but may refer to the achievement of termination or selection criteria, such as values above or below a threshold, or other constraints that may lead to a “suboptimal selection,” but reduce complexity and processing time.
[0086] That is, the partitioning unit 262 may be configured to divide the block 203 into smaller block partitions or sub-blocks (which again form blocks), for example, using quad-tree partitioning (QT), binary tree partitioning (BT) or triple-tree partitioning (TT), or any combination thereof repeatedly, and to perform prediction for each of the block partitions or sub-blocks, for example, wherein the mode selection comprises selecting a tree structure of the divided block 203, and a prediction mode is applied to each of the block partitions or sub-blocks.
[0087] Below, the partitioning (eg, by partitioning unit 262) and prediction processes (by inter prediction unit 244 and intra prediction unit 254) performed by example video encoder 20 are described in more detail.
[0088] Partitioning The partitioning unit 262 may divide (or segment) the current block 203 into smaller partitions, e.g., smaller blocks of square or rectangular size. These smaller blocks (also called sub-blocks) may be further divided into even smaller partitions. This is also called tree partitioning or hierarchical tree partitioning; for example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively divided, e.g., into two or more blocks at the next lower tree level, e.g., a node at tree level 1 (hierarchical level 1, depth 1), which may then be further divided into two or more blocks at the next lower level, e.g., tree level 2 (hierarchical level 2, depth 2), etc., until the partitioning terminates, e.g., by a termination criterion being satisfied, e.g., a maximum tree depth or a minimum block size being reached. Blocks that are not further divided are also called leaf blocks or leaf nodes of the tree. A tree resulting from a division into two partitions is called a binary tree (BT), a tree resulting from a division into three partitions is called a ternary tree (TT), and a tree resulting from a division into four partitions is called a quad tree (QT).
[0089] As mentioned above, the term "block" as used herein may refer to a portion of a picture, in particular a square or rectangular portion. For example, with reference to HEVC and VVC, a block may be or correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).
[0090] For example, a coding tree unit (CTU) may be or include a CTB of luma samples for a picture having three sample arrays, two corresponding CTBs of chroma samples, or a CTB of samples for a monochrome picture or a picture coded using three separate color planes and a syntax structure used to code the samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N, such that the division of components into CTBs is a partitioning. A coding unit (CU) may be or include a coding block of luma samples for a picture having three sample arrays, two corresponding coding blocks of chroma samples, or a coding block of samples for a monochrome picture or a picture coded using three separate color planes and a syntax structure used to code the samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N, such that the division of a CTB into coding blocks is a partitioning.
[0091] In an embodiment, for example, according to HEVC, a coding tree unit (CTU) may be divided into CUs using a quadtree structure represented as a coding tree. The decision of whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four PUs according to a PU partition type. Within one PU, the same prediction process is applied, and related information is sent to the decoder on a PU-by-PU basis. After obtaining residual blocks by applying a prediction process based on the PU partition type, the CU may be divided into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.
[0092] In an embodiment, quad-tree and binary-tree (QTBT) partitioning is used to divide coding blocks, for example, in accordance with the latest video coding standard currently under development, called Versatile Video Coding (VVC). In the QTBT block structure, CUs can have either square or rectangular shapes. For example, coding tree units (CTUs) are first divided by a quad-tree structure. The quad-tree leaf nodes are further divided by a binary tree or a ternary (or triple) tree structure. The partitioning tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transform processing without further partitioning. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. At the same time, multiple partitions, for example, triple-tree partitions, have also been proposed to be used with the QTBT block structure.
[0093] In one example, mode select unit 260 of video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0094] As described above, video encoder 20 is configured to determine or select a best or optimal prediction mode from a (predetermined) set of prediction modes, which may include, for example, intra-prediction modes and / or inter-prediction modes.
[0095] Intra prediction The set of intra prediction modes may include 35 different intra prediction modes, for example, omnidirectional modes such as DC (or average) mode and planar mode, or directional modes, for example, as defined in HEVC, or may include 67 different intra prediction modes, for example, omnidirectional modes such as DC (or average) mode and planar mode, or directional modes, for example, as defined in VVC.
[0096] The intra prediction unit 254 is configured to use reconstructed samples of neighboring blocks of the same current picture to generate an intra prediction block 265 according to an intra prediction mode in a set of intra prediction modes.
[0097] The intra prediction unit 254 (or generally, the mode selection unit 260) is further configured to output the intra prediction parameters (or generally, information indicating the selected intra prediction mode for the block) to the entropy encoding unit 270 in the form of a syntax element 266 for inclusion in the encoded picture data 21, e.g., so that the video decoder 30 may receive and use the prediction parameters for decoding.
[0098] Inter Prediction The set of (possible) inter prediction modes depends on the available reference pictures (i.e., previous, at least partially decoded pictures, e.g., stored in DPB230) and other inter prediction parameters, such as whether the entire reference picture or only a portion of the reference picture, e.g., a search window area around the area of the current block, is used to find the reference block that shows the best match, and / or whether pixel interpolation, e.g., half / semi-pel and / or quarter-pel interpolation, is applied.
[0099] In addition to the above prediction modes, skip mode and / or direct mode may also be applied.
[0100] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit may be configured to receive or obtain the picture block 203 (current picture block 203 of current picture 17) and the decoded picture 231, or at least one or more previously reconstructed blocks, e.g., reconstructed blocks of one or more other / different previously decoded pictures 231, for motion estimation. For example, a video sequence may have the current picture and the previously decoded picture 231; that is, in other words, the current picture and the previously decoded picture 231 may be part of or form a sequence of pictures that form a video sequence.
[0101] The encoder 20 may be configured, for example, to select a reference block from multiple reference blocks of the same or different pictures among multiple other pictures, and to provide the reference picture (or reference picture index) and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block to the motion estimation unit as inter-prediction parameters. This offset is also called a motion vector (MV).
[0102] The motion compensation unit is configured to obtain, e.g., receive, inter prediction parameters and perform inter prediction based on or using the inter prediction parameters to obtain inter prediction block 265. The motion compensation performed by the motion compensation unit may include fetching or generating a prediction block based on motion / block vectors determined by motion estimation, possibly performing interpolation to sub-pixel accuracy. Interpolation filtering may generate additional pixel samples from known pixel samples, thus possibly increasing the number of candidate prediction blocks that can be used to code the picture block. Upon receiving a motion vector for the PU of the current picture block, the motion compensation unit may find the prediction block to which the motion vector points in one of the reference picture lists.
[0103] The motion compensation unit may also generate syntax elements associated with the blocks and video slices that are used by video decoder 30 in decoding picture blocks of the video slices.
[0104] Entropy Coding Entropy encoding unit 270 may be configured to apply, for example, an entropy encoding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context-adaptive VLC scheme (CAVLC), an arithmetic coding scheme, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy encoding methodology or technique) to quantized coefficients 209, inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements to obtain encoded picture data 21, which may be output via output 272, for example, in the form of encoded bitstream 21, so that, for example, video decoder 30 may receive and use the parameters for decoding. Encoded bitstream 21 may be sent to video decoder 30 or may be stored in memory for later transmission or retrieval by video decoder 30.
[0105] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform-based encoder 20 may quantize the residual signal directly for a particular block or frame without relying on the transform processing unit 206. In other implementations, the encoder 20 may have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.
[0106] Decoder and decoding method 3 shows an example of a video decoder 30 configured to implement the techniques of the present application. The video decoder 30 is configured to receive encoded picture data 21 (e.g., encoded bitstream 21), for example, encoded by encoder 20, to obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, for example, data representing picture blocks of an encoded video slice and associated syntax elements.
[0107] 3, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., summer 314), a loop filter 320, a decoding picture buffer (DPB) 330, an inter prediction unit 344, and an intra prediction unit 354. Inter prediction unit 344 may be or include a motion compensation unit. Video decoder 30, in some examples, may perform a decoding path that is generally the reverse of the encoding path described with respect to video encoder 20 of FIG. 2.
[0108] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoding picture buffer (DPB) 230, inter prediction unit 244, and intra prediction unit 254 are also said to form a “built-in decoder” of video encoder 20. Accordingly, inverse quantization unit 310 may be the same in function as inverse quantization unit 210, inverse transform processing unit 312 may be the same in function as inverse transform processing unit 212, reconstruction unit 314 may be the same in function as reconstruction unit 214, loop filter 320 may be the same in function as loop filter 220, and decoding picture buffer 330 may be the same in function as decoding picture buffer 230. Accordingly, the descriptions provided for each unit and function of video encoder 20 also apply correspondingly to each unit and function of video decoder 30.
[0109] Entropy Decoding The entropy decoding unit 304 is configured to parse the bitstream 21 (or generally, the encoded picture data 21) and perform, e.g., entropy decoding, on the encoded picture data 21 to obtain, e.g., quantized coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), such as any or all of inter-prediction parameters (e.g., reference picture indices and motion vectors), intra-prediction parameters (e.g., intra-prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme described with respect to the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may further be configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode select unit 360, and other parameters to other units of the decoder 30. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0110] inverse quantization Inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally, information regarding inverse quantization) and quantized coefficients from encoded picture data 21 (e.g., by parsing and / or decoding by entropy decoding unit 304), and apply inverse quantization to decoded quantized coefficients 309 based on the quantization parameter to obtain inverse quantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may involve using the quantization parameter determined by video encoder 20 for each video block within a video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0111] Inverse transformation The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain reconstructed residual blocks 313 in the sample domain. The reconstructed residual blocks 313 may also be referred to as transform blocks 313. The transform may be an inverse transform, such as an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may further be configured to receive transform parameters or corresponding information from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304) to determine the transform to be applied to the dequantized coefficients 311.
[0112] Reconstruction The reconstruction unit 314 (e.g., an adder or summator 314) may be configured to add the reconstructed residual block 313 to the prediction block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365 to obtain the reconstructed block 315 in the sample domain.
[0113] filtering Loop filter unit 320 (either in the coding loop or after the coding loop) is configured to filter reconstructed block 315, e.g., to smooth pixel transitions or otherwise improve video quality, to obtain filtered block 321. Loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although loop filter unit 320 is shown in FIG. 3 as being an in-loop filter, in other configurations, loop filter unit 320 may be implemented as a post-loop filter.
[0114] Decoding Picture Buffer The decoded video blocks 321 of the picture are then stored in a decoding picture buffer 330. The decoding picture buffer 330 stores the decoded picture 331 as a reference picture for subsequent motion compensation of other pictures and / or for output display, respectively.
[0115] The decoder 30 is arranged to output the decoded pictures 331, for example via an output 332, to a user for presentation or viewing.
[0116] prediction The inter prediction unit 344 may be the same as the inter prediction unit 244 (in particular the motion compensation unit), and the intra prediction unit 354 may be the same in function as the intra prediction unit 254, performing the division or partitioning decision and prediction based on the partitioning and / or prediction parameters, or each information, received from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304). The mode selection unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the reconstructed picture, block, or each sample (filtered or unfiltered) to obtain a prediction block 365.
[0117] If the video slice is coded as an intra-coded (I) slice, intra prediction unit 354 of mode select unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current picture. If the video picture is coded as an inter-coded (i.e., B or P) slice, inter prediction unit 344 (e.g., a motion compensation unit) of mode select unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from entropy decoding unit 304. In the case of inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct reference frame lists List 0 and List 1 using a default construction technique based on the reference pictures stored in DPB 330.
[0118] Mode select unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing the motion vectors and other syntax elements, and to generate a prediction block for the current video block being decoded using the prediction information. For example, mode select unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra- or inter-prediction) used to code the video blocks of the video slice, the inter-prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the slice's reference picture lists, the motion vectors for each inter-encoded video block of the slice, the inter-prediction status for each inter-coded video block of the slice, and other information to decode video blocks in the current video slice.
[0119] Other variations of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 may generate an output video stream without the loop filtering unit 320. For example, a non-transform-based decoder 30 may inverse quantize the residual signal directly for a particular block or frame without the inverse transform processing unit 312. In other implementations, the video decoder 30 may have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.
[0120] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of the interpolation filtering, motion vector derivation, or loop filtering.
[0121] It should be noted that further operations may be applied to the derived motion vector of the current block (including, but not limited to, control point motion vectors in affine mode, sub-block motion vectors in affine, planar, and advanced temporal motion vector prediction (ATMVP) modes, temporal motion vectors, etc.). For example, the value of a motion vector is constrained to a predefined range according to its representation bit. When the representation bit of a motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" means exponent calculation. For example, when bitDepth is set equal to 16, the range is -32768 to 32767, and when bitDepth is set equal to 18, the range is -131072 to 131071. Two methods for constraining a motion vector are provided here.
[0122] Method 1: Remove the overflow MSB (most significant bit) by the following operation: ux=(mvx+2 bitDepth )%2 bitDepth (1) mvx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux (2) uy=(mvy+2 bitDepth )%2 bitDepth (3) mvy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy (4) For example, if the value of mvx is -32769, after applying equations (1) and (2), the resulting value is 32767. In computer systems, negative decimal numbers are stored as two's complement numbers. The two's complement of -32769 is 1, 0111, 1111, 1111, 1111 (17 bits), in which case the MSB is discarded, so the resulting two's complement is 0111, 1111, 1111, 1111 (decimal number is 32767), which is the same as the output by applying equations (1) and (2). ux=(mvpx+mvdx+2 bitDepth )%2 bitDepth (5) mvx=(ux>=2 bitDepth-1 )?(ux-2 bitDepth ):ux (6) uy=(mvpy+mvdy+2 bitDepth )%2 bitDepth (7) mvy=(uy>=2 bitDepth-1 )?(uy-2 bitDepth ):uy (8) As shown in equations (5) to (8), the operations may be applied during the summation of mvp (motion vector predictor) and mvd (motion vector difference).
[0123] Method 2: Eliminate overflow MSB by clipping the value vx=Clip3(-2 bitDepth-1 ,2 bitDepth-1 -1,vx) vy=Clip3(-2 bitDepth-1 ,2 bitDepth-1 -1,vy) Here, the function Clip3 is defined as follows:
number
[0124] 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments described herein. In an embodiment, the video coding device 400 may be a decoder, such as the video decoder 30 of FIG. 1A, or an encoder, such as the video encoder 20 of FIG. 1A.
[0125] Video coding device 400 includes an ingress port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an egress port 450 (or output port 450) for transmitting data, and a memory 460 for storing data. Video coding device 400 may also include optical-electrical (OE) and electro-optical (EO) components coupled to ingress port 410, receiver unit 420, transmitter unit 440, and egress port 450 for the egress or ingress of optical or electrical signals.
[0126] The processor 430 is implemented in hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGA, ASIC, and DSP. The processor 430 communicates with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. The inclusion of the coding module 470 thus provides a substantial improvement in the functionality of the video coding device 400 and achieves transformation of the video coding device 400 into different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0127] Memory 460 may include one or more disks, tape drives, and solid-state drives, and may be used as overflow data storage devices to store programs when such programs are selected for execution and to store instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).
[0128] FIG. 5 is a schematic block diagram of an apparatus 500 that may be used as either or both of source device 12 and destination device 14 of FIG. 1 according to an example embodiment.
[0129] Processor 502 in device 500 can be a central processing unit. Alternatively, processor 502 can be any other type of device or devices, now existing or later developed, capable of manipulating or processing information. While the disclosed implementations can be performed with a single processor, e.g., processor 502, as shown, advantages of speed and efficiency can be achieved using more than one processor.
[0130] The memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device in implementation. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that is accessed by the processor 502 using a bus 512. The memory 504 may further include an operating system 508 and application programs 510, which include at least one program that enables the processor 502 to perform the methods described herein. For example, the application programs 510 may include applications 1 through N, which may further include a video coding application that performs the methods described herein.
[0131] The apparatus 500 may also include one or more output devices, such as a display 518. The display 518, in one example, may be a touch-sensitive display that combines a display with touch-sensitive elements operable to detect touch input. The display 518 may be coupled to the processor 502 via the bus 512.
[0132] Although depicted here as a single bus, bus 512 of device 500 may be comprised of multiple buses. Additionally, secondary storage 514 may be directly coupled to other components of device 500 or may be accessible over a network and may comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. Device 500 may thus be implemented in a wide variety of configurations.
[0133] Recent developments in video coding have seen the emergence of more sophisticated techniques and schemes for prediction.
[0134] One such technique is multi-hypothesis prediction. Initially, the term "multi-hypothesis prediction" was introduced to extend motion compensation with one prediction signal to the linear superposition of several motion-compensated prediction signals. More recently, this approach has been generalized to combining existing prediction modes with extra merge-indexing prediction. This includes multi-hypothesis prediction, particularly for intra and inter modes (i.e., combined intra and inter modes; see, e.g., Joint Video Experts Team (JVET), document JVET-L0100-v3, titled "CE10.1.1: Multi-hypothesis prediction for improving AMVP mode, skip or merge mode, and intra mode," 12th Meeting, Macau, China, October 3-12, 2018). This approach applies multi-hypothesis prediction to improve intra modes by combining one intra prediction and one merge-indexing prediction. That is, a linear combination of both intra and inter prediction blocks is used.
[0135] Another technique is to introduce and use a separate merge list for sub-block merge candidates at the block level, i.e., separate from the merge candidate list for the regular merge mode (see, e.g., Joint Video Experts Team (JVET), document JVET-L0369-v2, titled “CE4: Separate list for sub-block merge candidates (Test 4.2.8)”), 12th Meeting, Macau, China, October 3-12, 2018. This technique is particularly applicable to the ATMVP and affine modes mentioned above.
[0136] Multiple hypothesis prediction for intra and inter modes is controlled by mh_intra_flag, which specifies whether multiple hypothesis prediction for intra and inter modes is enabled for the current block. In the prior art, the flag mh_intra_flag is conditionally signaled depending on merge_affine_flag according to the following syntax table: [Table 1] Therefore, mh_intra_flag is conditionally signaled under the condition that merge_affine_flag is zero. This means that the joint use of affine merging and mh_intra_flag, i.e., multi-hypothesis prediction for intra and inter modes, is not possible. However, it is possible to use ATMVP with mh_intra_flag.
[0137] After the recent adoption of separate merge lists for sub-block merge candidates, affine merge candidates are combined with atmvp merge candidates, and the utilization of these sub-block candidates is controlled by a newly introduced parameter called merge_subblock_flag. The separate merge_affine_flag is no longer used. Thus, the syntax no longer allows to distinguish the cases when merge affine mode is not used but atmvp is used. A new method for signaling multi-hypothesis prediction that combines intra and inter modes in the presence of separate merge lists for sub-block merge candidates has been developed within the framework of the present invention.
[0138] The present invention proposes multiple methods of utilizing and signaling multiple hypothesis prediction for intra and inter modes, assuming the existence of separate merge lists for sub-block merging candidates within the codec. Multiple harmonization methods are possible and are included in the current disclosure. The utilization of multiple hypothesis prediction for intra and inter modes in the case of skip mode is also disclosed.
[0139] According to one general aspect of the present disclosure, a method for video encoding of video data into a bitstream and a method for video decoding of video data received in the bitstream are provided.
[0140] A method of video encoding includes applying a first technique and / or a second technique. The first technique includes using separate merge lists for sub-block merging candidates. The second technique includes multiple hypothesis prediction for intra and inter modes. The method includes transmitting a first control flag in the bitstream, the first control flag indicating whether the first technique should be used, and transmitting a second control flag in the bitstream, the second control flag indicating whether the second technique should be used.
[0141] A method of video decoding includes applying a first technique and / or a second technique. The first technique includes using separate merge lists for sub-block merging candidates. The second technique includes multiple hypothesis prediction for intra and inter modes. The method includes receiving a first control flag from a bitstream, the first control flag indicating whether the first technique should be used, and receiving a second control flag from the bitstream, the second control flag indicating whether the second technique should be used.
[0142] According to an embodiment, the use of the multiple hypothesis prediction technique for intra and inter modes is controlled independently of the use of the separate merge list technique for subblock merging candidates, i.e., in such an embodiment, the signaling of the first control flag (merge_subblock_flag) is performed independently of the signaling of the second control flag (mh_intra_flag).
[0143] In one embodiment, multi-hypothesis prediction for intra and inter modes is controlled by mh_intra_flag, which is signaled independently of merge_subblock_flag. In this case, multi-hypothesis prediction for intra and inter modes is possible for both sub-block modes, i.e., affine and atmvp, and for normal merge. The following syntax table shows the possible signaling methods for mh_intra_flag in this embodiment: [Table 2] In the above table, the use of multiple hypothesis prediction combining intra and inter modes is restricted to coding blocks in merge mode.
[0144] [Table 3] In the above table, the use of multiple hypothesis prediction combining intra and inter modes is applicable to both coding blocks in merge mode and coding blocks in skip mode.
[0145] In an embodiment of the present invention, multiple hypothesis prediction for intra and inter modes is controlled by mh_intra_flag, which is signaled based on merge_subblock_flag. In this case, multiple hypothesis prediction for intra and inter modes is only possible when both subblock modes, i.e., affine and atmvp, are disabled. That is, multiple hypothesis prediction for intra and inter modes is possible in normal merge mode, but the combination of multiple hypothesis prediction for intra and inter modes and separate merge lists for subblock merge candidates is disabled. The following syntax table shows possible signaling methods for mh_intra_flag in this embodiment. [Table 4] In the above table, the use of multiple hypothesis prediction combining intra and inter modes is restricted to coding blocks in merge mode.
[0146] [Table 5] In the above table, the use of multiple hypothesis prediction combining intra and inter modes is applicable to both coding blocks in merge mode and coding blocks in skip mode.
[0147] Within the framework of the above embodiment, the encoder signals mh_intra_flag (the "second flag", also called ciip_flag) only if separate merge lists for sub-block candidates are disabled, i.e., if merge_subblock_flag (the "first flag") is zero. On the other hand, a decoder receiving a bitstream containing mh_intra_flag determines by default that separate merge lists for sub-block candidates are disabled.
[0148] In the following, the process according to the above embodiment will be described with reference to the flowcharts of FIGS.
[0149] FIG. 6 is a flow chart illustrating an encoder-side process according to an embodiment of the present invention.
[0150] The process begins in step S101 with a determination of whether separate merge lists for sub-block candidates should be used. If this is not the case (S101: No), processing proceeds to step S103. In step S103, a determination is made of whether multiple hypothesis prediction combining intra and inter modes will be used. Regardless of the outcome of that determination, processing then proceeds to step S105. In step S105, mh_intra_flag is generated as a parameter to be signaled in the bitstream. More specifically, if multiple hypothesis prediction for intra and inter modes is used (S103: Yes), a value of "1" is set to mh_intra_flag, and if multiple hypothesis prediction for intra and inter modes is not used (S103: No), a value of "0" is set to mh_intra_flag (not shown).
[0151] Thereafter, the process proceeds to step S 107. If it is determined in step S101 to use a separate merge list for the sub-block candidate (S101: YES), the process proceeds directly from step S101 to step S107.
[0152] In step S107, merge_subblock_flag is generated as a parameter to be notified in the bitstream. More specifically, in the case of use of an individual merge list, i.e., if the flow proceeds from S101: Yes, the value "1" is set to merge_subblock_flag, and in the case of flow proceeding from S101: No through S103 and S105, i.e., if an individual merge list is not used, merge_subblock_flag is set to the value "0" (not shown).
[0153] Therefore, multiple hypothesis prediction for intra and inter modes is only used conditionally under the condition that separate merge lists for sub-block candidates are not used. Thus, merge_subblock_flag is always generated, but mh_intra_flag is generated only if separate merge lists for sub-block candidates are not used. Then, a bitstream including merge_subblock_flag and conditionally including mh_intra_flag (if merge_subblock_flag is 0) is generated in step S109, and the process ends.
[0154] FIG. 7 is a flowchart illustrating decoder-side processing according to an embodiment of the present invention.
[0155] The process begins at step S201, in which the received bitstream is parsed. In the following step S203, it is checked whether the bitstream has an mh_intra_flag. If this is the case (S203: Yes), the process proceeds to step S207. In step S207, the mh_intra_flag is evaluated to determine whether multiple hypothesis prediction for intra and inter modes should be used in decoding, depending on the value of the parsed mh_intra_flag. If the mh_intra_flag has a value of 1, decoding is performed using multiple hypothesis prediction for intra and inter modes in step (S207: Yes → S211), and the process ends. Otherwise, i.e., if the mh_intra_flag has a value of 0 (S207: No), multiple hypothesis prediction is not used (and separate merge lists for sub-block candidates are not used).
[0156] On the other hand, if it is determined in step S203 that mh_intra_flag is not present (S203: No), processing proceeds to step S204. In step S204, the merge_subblock_flag received from the bitstream is evaluated to determine in the next step (S205) whether to use a separate merge list for the subblock candidate.
[0157] In step S205, if merge_subblock_flag has a value of 1, it is determined that an individual merge list should be used (S205: YES), and the process proceeds to step S209 where decoding is performed using the individual merge list. Otherwise, that is, if merge_subblock_flag has a value of 0 (S205: NO), it is determined that decoding is performed without using an individual merge list (and without using multiple hypothesis prediction for intra and inter modes), and the process flow ends.
[0158] Therefore, according to an embodiment, it is directly inferred from the presence of mh_intra_flag in the parsed bitstream (S203: yes) that separate merge lists for sub-block candidates are not used.
[0159] That is, merge_subblock_flag, which is always received in the bitstream, is parsed only if mh_intra_flag is not present in the bitstream, so that only a single flag needs to be evaluated at the decoder in any given case.
[0160] More generally, in accordance with the present invention, the use of the multiple hypothesis prediction technique for intra and inter modes is controlled based on the use of the separate merge list technique for sub-block merging candidates. More specifically, in accordance with the embodiment of Figure 7, the multiple hypothesis prediction technique for intra and inter modes may be used if and only if the separate merge list technique for sub-block merging candidates is unavailable.
[0161] The above embodiments allow for the coordination of multiple hypothesis prediction for intra and inter modes with separate merge lists for sub-block merging candidates. An embodiment that also allows for the use of multiple hypothesis prediction for intra and inter modes in block skip mode further allows for achieving coding gains with respect to the basic design. An embodiment that includes conditional signaling of mh_intra_flag based on the use of separate merge lists for sub-block merging candidates has the additional advantage of reducing signaling overhead, since mh_intra_flag does not need to be signaled in each case.
[0162] Mathematical Operators The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more strictly defined, and additional operations such as exponentiation and real division are defined. Numbering and counting conventions generally start from 0, e.g., "first" corresponds to 0, "second" corresponds to 1, etc.
[0163] Arithmetic operators The following arithmetic operations are defined as follows: [Table 6]
[0164] Logical operators The following arithmetic operators are defined as follows: x&&y Boolean logic "AND" of x and y x||y Boolean logic "OR" of x and y ! Boolean logic “NOT” x?y:zIf x is true or not equal to 0, evaluates to the value of y, otherwise evaluates to the value of z.
[0165] Relational operators The following relational operators are defined as follows: > greater than >= greater than or equal to < less than <= less than or equal to == equal to != not equal to
[0166] When a relational operator is applied to a syntactic element or variable to which the value "NA" (not applicable) is assigned, the value "NA" is treated as a distinct value of that syntactic element or variable. The value "NA" is considered not equal to any other value.
[0167] Bitwise operators The following bitwise operators are defined as follows: & Bitwise "AND". When acting on integer arguments, it acts on the two's complement representation of the integer value. When acting on a binary argument that contains fewer bits than the other arguments, the shorter argument is extended by adding additional significant bits equal to 0. | Bitwise "OR". When acting on integer arguments, it acts on the two's complement representation of the integer value. When acting on a binary argument that contains fewer bits than the other arguments, the shorter argument is extended by adding additional significant bits equal to 0. ^ Bitwise "XOR". When acting on integer arguments, it acts on the two's complement representation of the integer value. When acting on a binary argument that contains fewer bits than the other arguments, the shorter argument is extended by adding additional significant bits equal to 0. x>>y Arithmetic right shift of the two's complement integer representation of x by y binary digits. This function is defined only for non - negative integer values of y. The bits shifted into the most significant bit (MSB) as a result of the right shift have the same value as the MSB of x before the shift operation. x<<y Arithmetic left shift of the two's complement integer representation of x by y binary digits. This function is defined only for non - negative integer values of y. The bits shifted into the least significant bit (LSB) as a result of the left shift have a value equal to 0.
[0168] assignment operator The following assignment operators are defined as follows: = assignment operator ++ increment, i.e., x++, is equivalent to x=x+1 and, when used in an array index, evaluates to the value of the variable before the increment operation. -- Decrement, i.e., x--, is equivalent to x=x-1 and, when used in an array index, evaluates to the value of the variable before the decrement operation. += Increment by the specified amount, i.e. x+=3 is equivalent to x=x+3, and x+=(-3) is equivalent to x=x+(-3). -= Decrement by the specified amount, i.e. x-=3 is equivalent to x=x-3, and x-=(-3) is equivalent to x=x-(-3).
[0169] Range Notation The following notation is used to specify a range of values: x=y..zx, y, and z are integers, and z is greater than y, so that x takes on an integer value greater than or equal to y and less than or equal to z.
[0170] Mathematical Functions The following mathematical functions are defined:
number
number
number
number
number
[0171] Operation precedence When precedence in an expression is not explicitly indicated by the use of parameters, the following rules apply: Operations with higher precedence are evaluated before operations with lower precedence. Operations of equal precedence are evaluated in order from left to right.
[0172] The following table defines the priority of operations from highest to lowest, with higher positions in the table indicating higher priority.
[0173] For operations that are also used in the C programming language, the precedence used herein is the same as that used in the C programming language.
[0174] Table: Priority of operations from highest (top of table) to lowest (bottom of table) [Table 7]
[0175] Text description of logical operations A statement of logical operations that will be written mathematically in the text in the following form: if(condition 0) Statement 0 else if(condition 1) Statement 1 ... else / * Explanatory findings regarding the remaining conditions * / Statement n may be written in the following manner: ···As follows / ···The following applies: - If condition 0, then statement 0 - Otherwise, if condition 1, then statement 1 - - Otherwise (descriptive observations regarding the remaining conditions), statement n
[0176] Each "if... otherwise, if... otherwise then..." statement in the text is introduced by "...as follows" or "...the following applies" immediately followed by "if...." The final condition of "if... otherwise, if... otherwise then..." is always "otherwise...." Interleaved "if... otherwise, if... otherwise then..." statements can be identified by matching "...as follows" or "...the following applies" with the closing "otherwise then...."
[0177] A statement of logical operations that will be written mathematically in the text in the following form: if(condition0a&&condition0b) Statement 0 else if(condition 1a||condition 1b) Statement 1 ... else Statement n may be written in the following manner: ···As follows / ···The following applies: - Statement 0 if all of the following conditions are true: - Condition 0a - Condition 0b - Otherwise, if one or more of the following conditions are true, then statement 1: - Condition 1a - Condition 1b - - Otherwise, statement n
[0178] A statement of logical operations that will be written mathematically in the text in the following form: if(condition 0) Statement 0 if(condition1) Statement 1 may be written in the following manner: If condition 0, then statement 0 If condition 1, then statement 1
[0179] For example, embodiments of the encoder 20 and decoder 30, and functions described herein with reference to the encoder 20 and decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on a computer-readable medium or transmitted over a communication medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) communication media, such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0180] By way of example, and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio waves, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio waves, and microwaves are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory, tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where a disk typically reproduces data magnetically, while a disc reproduces data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0181] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the above structures, or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a hybrid codec. Alternatively, the techniques may be implemented entirely in one or more circuits or logic elements.
[0182] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, the various units may be combined into a codec hardware unit, as described above, or may be provided by a collection of interoperating hardware units including one or more processors, as described above, along with appropriate software and / or firmware.
Claims
1. calculating a residual block of the coding block; applying a transform to sample values of the residual block to obtain transform coefficients of the coding block; obtaining quantized coefficients of the coding block based on the transform coefficients; obtaining a value of a first control flag, the value of the first control flag indicating whether a first technique should be used for the coding block, the first technique comprising using a separate merge list for sub-block merging candidates; obtaining a value of a second control flag for the coding block in response to determining that the first technique is not used for the coding block, the value of the second control flag being equal to 1 indicating that a second technique should be used, the second technique comprising multiple hypothesis prediction for intra and inter modes, and the value of the second control flag being transmitted only if the first technique is not used for the coding block; entropy encoding the quantized coefficients and the second control flag by applying an entropy encoding algorithm to obtain an encoded bitstream; A method having the following.
2. transmitting the second control flag if the coding block is coded in merge mode. The method of claim 1.
3. and determining that the second control flag should not be transmitted if the coding block is not coded in merge mode. The method of claim 2.
4. transmitting the second control flag if the coding block is coded in skip mode. The method of claim 1.
5. determining that the second control flag should not be transmitted if the coding block is not coded in skip mode or merge mode. The method of claim 4.
6. receiving a bitstream; entropy decoding the bitstream by applying an entropy decoding algorithm to obtain quantized coefficients of coding blocks; inferring a value of a first control flag of the coding block to be equal to 0, the value of the first control flag indicating whether a first technique should be used, the first technique comprising using a separate merge list for sub-block merging candidates; and obtaining a value of a second control flag of the coding block based on the value of the first control flag, the value of the second control flag indicating whether a second technique should be used, the second technique comprising multiple hypothesis prediction for intra and inter modes, and the value of the second control flag being transmitted only if the first control flag is equal to 0; obtaining transform coefficients based on the quantized coefficients; obtaining a residual block of the coding block based on the transform coefficients; reconstructing the coding block based at least on the residual block of the coding block, the value of the first control flag, and the value of the second control flag; A method having the following.
7. receiving the second control flag if the coding block is coded in merge mode. The method of claim 6.
8. determining that the second control flag should not be received if the coding block is not coded in merge mode. The method of claim 7.
9. receiving the second control flag if the coding block is coded in skip mode. The method of claim 6.
10. determining that the second control flag should not be received if the coding block is not coded in skip mode or merge mode.
10. The method of claim 9.
11. A decoder comprising: at least one processor; one or more memories coupled to the at least one processor; The memory includes: storing programming instructions which, when executed by at least one processor, cause the decoder to perform the method of any one of claims 6 to 10; decoder.
12. 1. An encoder comprising: at least one processor; one or more memories coupled to the at least one processor; The memory includes: storing programming instructions that, when executed by at least one processor, cause the encoder to perform the method of any one of claims 1 to 5; Encoder.
13. A non-transitory storage medium for storing an encoded bitstream of a video signal, comprising: the encoded bitstream includes a plurality of syntax elements and quantized coefficients of a coding block, the quantized coefficients being obtained based on transform coefficients of the coding block, the transform coefficients being obtained by applying a transform to sample values of a residual block of the coding block; When the value of a first control flag is equal to 0, the plurality of syntax elements include a value of a second control flag, the value of the first control flag indicating whether a first technique should be used for the coding block, the first technique comprising using a separate merge list for sub-block merging candidates, the value of the second control flag indicating whether a second technique should be used for the coding block, the second technique comprising multiple hypothesis prediction for intra and inter modes, and the value of the second control flag being included in the plurality of syntax elements only when the first control flag is equal to 0. Non-transitory storage media.
14. 1. A method for storing a bitstream, comprising: receiving one or more bitstreams by at least one receiver; storing the one or more bitstreams in one or more memories; the bitstream includes a plurality of syntax elements and quantized coefficients of a coding block, the quantized coefficients being obtained based on transform coefficients of the coding block, the transform coefficients being obtained by applying a transform to sample values of a residual block of the coding block; When the value of a first control flag is equal to 0, the plurality of syntax elements include a value of a second control flag, the value of the first control flag indicating whether a first technique should be used for the coding block, the first technique comprising using a separate merge list for sub-block merging candidates, the value of the second control flag indicating whether a second technique should be used for the coding block, the second technique comprising multiple hypothesis prediction for intra and inter modes, and the value of the second control flag being included in the plurality of syntax elements only when the first control flag is equal to 0. method.
Citation Information
Patent Citations
Multi-hypothesis motion compensation for scalable video coding and 3D video coding
US20140044179A1