Video coding method and apparatus based on motion prediction
By optimizing inter prediction through a CIIP availability flag and sequence parameter set, the method addresses inefficiencies in video coding for high-resolution and immersive media, improving compression efficiency and reducing signaling overhead.
Patent Information
- Application Number
- JP2021575502
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-19
- Filing Date
- 2020-06-19
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2040-06-19
AI Technical Summary
The increasing demand for high-resolution and immersive media has led to a need for highly efficient video coding methods that can compress and transmit high-quality image/video information while minimizing transmission and storage costs, and existing technologies face inefficiencies in inter prediction and unnecessary signaling during the process.
The method involves deriving a prediction mode for a current block based on a bitstream that includes a CIIP availability flag, parsing a regular merge flag under specific conditions, and encoding video information with a sequence parameter set to optimize inter prediction and reduce unnecessary signaling.
This approach improves overall image/video compression efficiency, enables efficient inter prediction, and reduces unnecessary syntax signaling, thereby enhancing the coding process.
Smart Images

Figure 0007807239000044 
Figure 0007807239000045 
Figure 0007807239000046
Abstract
Description
[Technical Field]
[0001] The present technology relates to a method and apparatus for coding video based on motion prediction. [Background technology]
[0002] In recent years, the demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has been increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits to be transmitted increases relatively compared to existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line or storing image / video data using an existing storage medium, transmission costs and storage costs increase.
[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, has been increasing.
[0004] Therefore, there is a need for highly efficient image / video compression technology to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information having the above-mentioned various characteristics. Summary of the Invention [Problem to be solved by the invention]
[0005] The technical problem of this document is to provide a method and apparatus for improving video coding efficiency.
[0006] Another technical problem of this document is to provide a method and apparatus for efficiently performing inter prediction.
[0007] Another technical problem of this document is to provide a method and apparatus for preventing unnecessary signaling during inter prediction. [Means for solving the problem]
[0008] According to one embodiment of this document, a decoding method performed by a decoding device includes steps of obtaining information regarding a prediction mode of a current block from a bitstream, deriving a prediction mode of the current block based on the information about the prediction mode, generating a predicted sample of the current block based on the prediction mode, and generating a reconstructed sample based on the predicted sample, wherein the bitstream includes a sequence parameter set, the sequence parameter set includes a CIIP (combined inter-picture merge and intra-picture prediction) availability flag, and the deriving step may include a step of parsing a regular merge flag from the bitstream based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied.
[0009] According to another embodiment of this document, an encoding method performed by an encoding device includes a step of determining a prediction mode of a current block, a step of generating information about the prediction mode based on the prediction mode, and a step of encoding video information including information about the prediction mode, wherein the video information includes a sequence parameter set, the sequence parameter set includes a CIIP availability flag, and the video information includes a grammar merge flag based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied.
[0010] According to another embodiment of the present document, there is provided a computer-readable digital storage medium, the digital storage medium including information that enables a decoding device to perform a decoding method, the decoding method including the steps of obtaining information regarding a prediction mode of a current block from a bitstream, deriving a prediction mode of the current block based on the information regarding the prediction mode, generating a predicted sample of the current block based on the prediction mode, and generating a reconstructed sample based on the predicted sample, wherein the bitstream includes a sequence parameter set, the sequence parameter set includes a CIIP availability flag, and the deriving step includes the step of parsing a regular merge flag from the bitstream based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied. [Effects of the Invention]
[0011] According to one embodiment of this document, the overall image / video compression efficiency can be improved.
[0012] According to one embodiment of this document, inter prediction can be performed efficiently.
[0013] According to one embodiment of this document, unnecessary syntax signaling can be efficiently removed during inter prediction. [Brief explanation of the drawings]
[0014] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document may be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which an embodiment of this document can be applied. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which an embodiment of the present document can be applied. [Figure 4]1 illustrates an example of an inter-prediction based video / image encoding method. [Figure 5] 1 illustrates an example of an inter-prediction based video / image decoding method. [Figure 6] 1 illustrates an exemplary inter-prediction procedure. [Figure 7] FIG. 1 illustrates spatial candidates that can be used for inter prediction. [Figure 8] FIG. 10 is a diagram illustrating a merge mode in which there is a motion vector difference that can be used during inter prediction. [Figure 9] FIG. 1 illustrates a sub-block-based temporal motion vector prediction process that can be used during inter prediction. [Figure 10] FIG. 1 illustrates a sub-block-based temporal motion vector prediction process that can be used during inter prediction. [Figure 11] FIG. 1 is a diagram illustrating partitioning modes applicable to inter prediction. [Figure 12] FIG. 10 is a diagram illustrating a CIIP mode that can be applied to inter prediction. [Figure 13] 1 illustrates an example of a video / picture encoding method and related components including an inter-prediction method according to an embodiment of the present document. [Figure 14] 1 illustrates an example of a video / picture encoding method and related components including an inter-prediction method according to an embodiment of the present document. [Figure 15] 1 illustrates an example of a video / picture decoding method and related components including an inter-prediction method according to an embodiment of the present document. [Figure 16] 1 illustrates an example of a video / picture decoding method and related components including an inter-prediction method according to an embodiment of the present document. [Figure 17] 1 illustrates an example of a content streaming system to which the embodiments disclosed herein may be applied. DETAILED DESCRIPTION OF THE INVENTION
[0015] Because the disclosure of this document can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. The terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. The singular expressions "a," "an," "an," "the," and the like include the expression "at least one" unless the context clearly dictates otherwise. In this document, the terms "comprise," "have," and the like are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0016] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of describing different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of this document as long as they do not deviate from the essence of the method disclosed herein.
[0017] Hereinafter, the embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and redundant description of the same components will be omitted.
[0018] FIG. 1 illustrates a schematic diagram of an example video / image coding system to which embodiments of this document may be applied.
[0019] As shown in Figure 1, a video / image coding system includes a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.
[0020] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0021] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.
[0022] An encoding device can encode input video / images. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0023] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.
[0024] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.
[0025] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.
[0026] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation audio video coding standard), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0027] Various embodiments relating to video / image coding are presented in this document, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0028] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may be referred to as transform coefficients for uniformity of expression.
[0029] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information includes information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s)), and scaled transform coefficients may be derived by inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on inverse transform (transform) of the scaled transform coefficients. This can be similarly applied / expressed in other parts of this document.
[0030] In this document, video can refer to a collection of a series of images over time. A picture generally refers to a unit that shows an image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile contains one or more coding tree units (CTUs). A picture consists of one or more slices / tiles. A picture consists of one or more tile groups. A tile group contains one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan refers to a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan within a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the height of the picture. A tile scan indicates a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably.For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0031] A pixel or a pel may refer to the smallest unit constituting one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or may refer to a transform coefficient in the frequency domain when such a pixel value is transformed into the frequency domain.
[0032] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.
[0033] In this document, " / " and "," should be interpreted to indicate "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." (In this document, the term " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A, B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")
[0034] Additionally, in this document, "or" should be interpreted as "and / or." For example, "A or B" may mean 1) only "A," 2) only "B," or 3) "A and B." Further, in the document, the term "or" should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively."
[0035] 2 is a diagram illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the term "video encoding device" includes the image encoding device.
[0036] As shown in FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0037] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0038] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or pixel value, and can refer to only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).
[0039] The encoding apparatus 200 subtracts a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input video signal (original block, original sample array) is called a subtraction unit 231. The prediction unit 220 may perform prediction on a current block (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 determines whether intra prediction or inter prediction is to be applied for the current block or CU. The prediction unit 220 may generate various information related to prediction, such as prediction mode information, as will be described later in the description of each prediction mode, and transmit the information to the entropy encoding unit 240. The prediction information can be encoded in the entropy encoding unit 240 and output in the form of a bitstream.
[0040] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0041] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (col CU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the neighboring block as a motion vector predictor and signaling the motion vector difference.
[0042] The prediction unit 220 generates a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit 200 can apply intra prediction or inter prediction for predicting a block, or can simultaneously apply intra prediction and inter prediction. This is called combined inter and intra prediction (CIIP). The prediction unit can also use intra block copy (IBC) prediction mode or palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content image / video coding, such as games, as in screen content coding (SCC). IBC basically performs prediction within a current picture, but is similar to inter prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. Palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.
[0043] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a reconstructed signal or a residual signal.
[0044] The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0045] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0046] The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 240 can encode information required for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / video information can further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information can also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device are included in video / image information. The video / image information is encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoding unit 240.
[0047] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) is reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal is used for intra prediction of the next block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as described below.
[0048] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in the picture encoding and / or reconstruction process.
[0049] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, and a bilateral filter. The filtering unit 260 generates various information related to filtering and transmits it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The entropy encoding unit 240 encodes the filtering information and outputs it in the form of a bitstream.
[0050] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter prediction unit 221. This allows the encoding apparatus to avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus when inter prediction is applied, and also improves coding efficiency.
[0051] The DPB of the memory 270 may store a modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0052] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0053] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be implemented as a single hardware component (e.g., a decoder chipset or processor). The memory 360 may include a decoded picture buffer (DPB) or may be implemented as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0054] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process in which the video / image information was processed by the encoding apparatus of FIG. 3. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processing unit applied by the encoding apparatus. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced via a reproduction device.
[0055] The decoding apparatus 300 receives a signal output from the encoding apparatus of FIG. 2 in the form of a bitstream, and the received signal is decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / video information) necessary for image restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding apparatus may decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 decodes information in a bitstream based on a coding method such as exponential-Golomb coding, context-adaptive variable length coding (CAVLC), or context-adaptive arithmetic coding (CABAC), and outputs values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoding information on neighboring and current blocks or information on symbols / bins decoded in previous steps, predicts the occurrence probability of bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element.In this case, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Prediction information from the information decoded by the entropy decoding unit 310 is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320.
[0056] The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). Information related to filtering among the information decoded by the entropy decoding unit 310 is provided to the filtering unit 350. A receiving unit (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, and the receiving unit may be a component of the entropy decoding unit 310. The decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder includes the entropy decoding unit 310, and the sample decoder includes at least one of the inverse quantization unit 321, the inverse transform unit 322, the adder 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0057] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0058] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0059] The prediction unit 330 performs prediction on a current block and generates a predicted block including prediction samples for the current block. The prediction unit 330 may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0060] The prediction unit 330 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content video / movie coding, such as games, as in screen content coding (SCC). IBC basically performs prediction within a current picture, but can be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index is included in the video / picture information and signaled.
[0061] The intra prediction unit 331 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. Prediction modes in intra prediction include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0062] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0063] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as a reconstructed block.
[0064] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.
[0065] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during picture decoding.
[0066] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 60, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0067] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information is transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0068] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0069] Meanwhile, as described above, prediction is performed to improve compression efficiency during video coding. Accordingly, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device can improve image coding efficiency by signaling to a decoding device information (residual information) regarding the residual between the original block and the predicted block, rather than the original sample values of the original block themselves. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.
[0070] The residual information may be generated through a transform and quantization procedure. For example, an encoding apparatus may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and then signal the related residual information (via a bitstream) to a decoding apparatus. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding apparatus may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding apparatus may generate a reconstructed picture based on the predicted block and the residual block. The encoding apparatus may also derive a residual block by inverse quantizing / inverse transforming the quantized transform coefficients for reference for inter-prediction of a future picture, and generate a reconstructed picture based on the residual block.
[0071] When inter prediction is applied to the current block, a prediction unit of an encoding / decoding apparatus may perform inter prediction on a block-by-block basis to derive predicted samples. Inter prediction refers to a prediction derived in a manner that is dependent on data elements (e.g., sample values or motion information) of picture(s) other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector on a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block, sub-block, or sample-by-block basis based on the correlation of motion information between neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include information on an inter prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, the neighboring blocks include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including the temporal neighboring blocks may be called collocated pictures (colPics).For example, a motion information candidate list may be constructed based on neighboring blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled. Inter prediction may be performed based on various prediction modes. For example, in skip mode and (normal) merge mode, the motion information of the current block may be the same as that of the selected neighboring block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference.
[0072] A video / image encoding procedure based on inter prediction generally includes, for example, the following.
[0073] FIG. 4 illustrates an example of an inter-prediction based video / image encoding method.
[0074] An encoding apparatus performs inter prediction on a current block (S400). The encoding apparatus derives an inter prediction mode and motion information for the current block and generates a predicted sample for the current block. Here, the steps of determining the inter prediction mode, deriving the motion information, and generating the predicted sample may be performed simultaneously, or one step may be performed before the other steps. For example, the inter prediction unit of the encoding apparatus includes a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit, in which the prediction mode determination unit determines a prediction mode for the current block, the motion information derivation unit derives motion information for the current block, and the predicted sample derivation unit derives a predicted sample for the current block. For example, the inter prediction unit of the encoding apparatus may search for a block similar to the current block within a certain region (search area) of a reference picture using motion estimation, and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located is derived, and a motion vector is derived based on a position difference between the reference block and the current block. The encoding apparatus may determine a mode to be applied to the current block from various prediction modes. The encoding apparatus may compare rate-distortion (RD) costs for the various prediction modes to determine an optimal prediction mode for the current block.
[0075] For example, when a skip mode or a merge mode is applied to the current block, the encoding apparatus may construct a merge candidate list (described below) and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate is generated and signaled to the decoding apparatus. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0076] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list (described below) and use the motion vector of a selected MVP (motion vector predictor) candidate from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block may be the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information about the MVD may be signaled to the decoding apparatus. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.
[0077] The encoding apparatus derives residual samples based on the predicted samples (S410) by comparing the original samples of the current block with the predicted samples.
[0078] The encoding apparatus encodes video information including prediction information and residual information (S420). The encoding apparatus may output the encoded video information in the form of a bitstream. The prediction information is information related to the prediction procedure and includes prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information on motion information. The information on the motion information includes candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) for deriving a motion vector. The information on the motion information also includes information on the MVD and / or reference picture index information. The information on the motion information also includes information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information on the residual sample. The residual information includes information on quantized transform coefficients for the residual sample.
[0079] The output bitstream can be stored in a (digital) storage medium and transmitted to a decoding device, or can be transmitted to a decoding device via a network.
[0080] Meanwhile, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the encoding apparatus derives the same prediction result as that performed in the decoding apparatus, thereby improving coding efficiency. Therefore, the encoding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0081] A video / picture decoding procedure based on inter prediction generally includes, for example, the following.
[0082] FIG. 5 illustrates an example of an inter-prediction based video / picture decoding method.
[0083] As shown in Figure 5, the decoding apparatus may perform operations corresponding to those performed by the encoding apparatus, such as performing prediction on the current block based on received prediction information and deriving predicted samples.
[0084] Specifically, the decoding apparatus determines a prediction mode for the current block based on received prediction information (S500). The decoding apparatus determines which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0085] For example, it may determine whether the merge mode or the (A)MVP mode is applied to the current block based on the merge flag, or may select one of various inter prediction mode candidates based on the mode index. The inter prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or various inter prediction modes described below.
[0086] The decoding apparatus derives motion information of the current block based on the determined inter prediction mode (S510). For example, when a skip mode or a merge mode is applied to the current block, the decoding apparatus may construct a merge candidate list (described below) and select one of the merge candidates included in the merge candidate list. The selection is performed based on the selection information (merge index) described above. Motion information of the selected merge candidate may be used to derive motion information of the current block. The motion information of the selected merge candidate may be used as motion information of the current block.
[0087] As another example, when the (A)MVP mode is applied to the current block, the decoding apparatus may construct an (A)MVP candidate list (described below) and use a motion vector of a motion vector predictor (MVP) candidate selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Furthermore, the decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. A picture pointed to by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.
[0088] Alternatively, the motion information of the current block may be derived without constructing a candidate list, as described below. In this case, the motion information of the current block may be derived according to a procedure disclosed in a prediction mode, as described below. In this case, the candidate list construction as described above may be omitted.
[0089] The decoding apparatus generates prediction samples for the current block based on the motion information of the current block (S520). In this case, the reference picture may be derived based on a reference picture index of the current block, and the prediction samples of the current block may be derived using samples of a reference block pointed to in the reference picture by the motion vector of the current block. In this case, as described below, a prediction sample filtering procedure may be further performed on all or some of the prediction samples of the current block, depending on the case.
[0090] For example, the inter-prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for the current block based on prediction mode information received in the prediction mode determination unit, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information regarding the motion information received in the motion information derivation unit, and derive a prediction sample of the current block in the prediction sample derivation unit.
[0091] The decoding apparatus generates residual samples for the current block based on the received residual information (S530). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and generates a reconstructed picture based on the reconstructed samples (S540). Thereafter, an in-loop filtering procedure, etc., can be further applied to the reconstructed picture, as described above.
[0092] FIG. 6 exemplarily illustrates an inter prediction procedure.
[0093] As shown in Figure 6, as described above, the inter prediction procedure includes an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction (prediction sample generation) step based on the derived motion information. As described above, the inter prediction procedure is performed in an encoding device and a decoding device. In this document, a coding device includes an encoding device and / or a decoding device.
[0094] As shown in FIG. 6, the coding apparatus determines an inter prediction mode for a current block (S600). Various inter prediction modes are used for predicting the current block in a picture. For example, various modes such as merge mode, skip mode, motion vector prediction (MVP) mode, affine mode, sub-block merge mode, merge with MVD (MMVD) mode, and historical motion vector prediction (HMVP) mode may be used. Decoder side motion vector refinement (DMVR) mode, adaptive motion vector resolution (AMVR) mode, bi-prediction with CU-level weight (BCW), bi-directional optical flow (BDOF), etc. may be used in addition to or instead of the accompanying modes. The affine mode may also be referred to as an "affine motion prediction mode." The MVP mode may also be referred to as "AMVP (advanced motion vector prediction) mode." In this document, some modes and / or motion information candidates derived by some modes may be included as one of the motion information-related candidates of other modes. For example, an HMVP candidate may be added as a merge candidate of the merge / skip mode, or may be added as an MVP candidate of the MVP mode.
[0095] Prediction mode information indicating the inter prediction mode of the current block may be signaled from the encoding apparatus to the decoding apparatus. The prediction mode information is included in a bitstream and received by the decoding apparatus. The prediction mode information includes index information indicating one of multiple candidate modes. Alternatively, the inter prediction mode may be indicated through hierarchical signaling of flag information. In this case, the prediction mode information includes one or more flags. For example, a skip flag may be signaled to indicate whether to apply the skip mode, and if the skip mode is not applied, a merge flag may be signaled to indicate whether to apply the merge mode, and if the merge mode is not applied, an MVP mode may be applied. Alternatively, a flag for additional classification may be further signaled. The affine mode may be signaled as an independent mode or as a mode dependent on the merge mode or MVP mode. For example, the affine mode may include an affine merge mode and an affine MVP mode.
[0096] Meanwhile, information indicating whether the list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit) is signaled. This information may be referred to as motion prediction direction information, inter-prediction direction information, or inter-prediction indication information, and may be configured / encoded / signaled in the form of, for example, an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element may indicate whether the list0 (L0) prediction, list1 (L1) prediction, or bi-prediction is used for the current block (current coding unit). For convenience of explanation, in this document, the inter-prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element may be referred to as a motion prediction direction. L0 prediction may be denoted as pred_L0, L1 prediction as pred_L1, and bi-prediction as pred_BI. For example, the following prediction types can be indicated depending on the value of the inter_pred_idc syntax element:
[0097] As described above, one picture includes one or more slices. A slice has one of the slice types, including an I (intra) slice, a P (predictive) slice, and a B (bi-predictive) slice. The slice type is indicated based on slice type information. For blocks in an I slice, inter prediction is not used for prediction, and only intra prediction is used. Of course, even in this case, original sample values can be coded and signaled without prediction. For blocks in a P slice, intra prediction or inter prediction is used, and when inter prediction is used, only uni prediction is used. On the other hand, for blocks in a B slice, intra prediction or inter prediction is used, and when inter prediction is used, up to bi prediction is used.
[0098] L0 and L1 include reference pictures encoded / decoded before the current picture. For example, L0 includes reference pictures before and / or after the current picture in POC order, and L1 includes reference pictures after and / or before the current picture in POC order. In this case, L0 is assigned a reference picture index lower than the reference picture before the current picture in POC order, and L1 is assigned a reference picture index lower than the reference picture after the current picture in POC order. In the case of a B slice, bi-prediction is applied, and either unidirectional bi-prediction or bi-directional bi-prediction may be applied in this case. Bi-directional bi-prediction may be referred to as true bi-prediction.
[0099] Specifically, for example, information regarding the inter prediction mode of the current block may be coded and signaled at a level such as a CU (CU syntax), or may be implicitly determined according to conditions. In this case, some modes may be explicitly signaled, and the remaining modes may be implicitly derived.
[0100] For example, the CU syntax can carry information about the (inter) prediction mode as shown in Table 1 below.
[0101] [Table 1-1]
[0102] [Table 1-2]
[0103] [Table 1-3]
[0104] [Table 1-4]
[0105] [Table 1-5]
[0106] [Table 1-6]
[0107] [Table 1-7]
[0108] [Table 1-8]
[0109] [Table 1-9]
[0110] [Table 1-10]
[0111] [Table 1-11]
[0112] [Table 1-12]
[0113] [Table 1-13]
[0114] Here, cu_skip_flag indicates whether the skip mode is applied to the current block (CU).
[0115] A value of 0 in pred_mode_flag indicates that the current coding unit is coded in inter prediction mode. A value of 1 in pred_mode_flag indicates that the current coding unit is coded in intra prediction mode. (pred_mode_flag equal to 0 specifies that the current coding unit is coded in inter prediction mode. pred_mode_flag equal to 1 specifies that the current coding unit is coded in intra prediction mode.)
[0116] A value of 1 in pred_mode_ibc_flag indicates that the current coding unit is coded in IBC prediction mode. A value of 0 in pred_mode_ibc_flag indicates that the current coding unit is not coded in IBC prediction mode (pred_mode_ibc_flag equal to 1 specifies that the current coding unit is coded in IBC prediction mode. pred_mode_ibc_flag equal to 0 specifies that the current coding unit is not coded in IBC prediction mode).
[0117] A value of 1 in pcm_flag[x0][y0] indicates that the pcm_sample() syntax structure is present and the transform_tree() syntax structure is not present in the coding unit including the luma coding block at the location (x0, y0). A value of 0 in pcm_flag[x0][y0] indicates that the pcm_sample() syntax structure is not present. (pcm_flag[x0][y0] equal to 1 specifies that the pcm_sample() syntax structure is present and the transform_tree() syntax structure is not present in the coding unit including the luma coding block at the location (x0, y0). pcm_flag[x0][y0] equal to 0 specifies that the pcm_sample() syntax structure is not present.) In other words, pcm_flag indicates whether pulse coding modulation (PCM) mode is applied to the current block. When PCM mode is applied to the current block, prediction, transformation, quantization, etc. are not applied, and the original sample values in the current block are coded and signaled.
[0118] A value of intra_mip_flag[x0][y0] of 1 indicates that the intra prediction type for the luma sample is matrix-based intra prediction (MIP). A value of intra_mip_flag[x0][y0] of 0 indicates that the intra prediction type for the luma sample is not matrix-based intra prediction. (intra_mip_flag[x0][y0] equal to 1 specifies that the intra prediction type for luma samples is matrix-based intra prediction (MIP). intra_mip_flag[x0][y0] equal to 0 specifies that the intra prediction type for luma samples is not matrix-based intra prediction.) In other words, intra_mip_flag indicates whether the MIP prediction mode (type) is applied to the current block (luma sample).
[0119] intra_chroma_pred_mode[x0][y0] specifies the intra prediction mode for chroma samples in the current block.
[0120] general_merge_flag[x0][y0] indicates whether the inter prediction parameters for the current coding unit are inferred from a neighboring inter-predicted partition. That is, general_merge_flag indicates that general merging is available, and when the value of general_merge_flag is 1, regular merge mode, mmvd mode, and merge subblock mode (subblock merge mode) are available. For example, when the value of general_merge_flag is 1, merge data syntax is parsed from the encoded video / image information (or bitstream), and the merge data syntax is configured / coded to include information such as that shown in Table 2.
[0121] [Table 2-1]
[0122] [Table 2-2]
[0123] [Table 2-3]
[0124] Here, if the value of regular_merge_flag[x0][y0] is 1, it indicates that regular merge mode is used to generate the inter prediction parameters of the current coding unit. That is, regular_merge_flag indicates whether the merge mode (regular merge mode) is applied to the current block.
[0125] When the value of mmvd_merge_flag[x0][y0] is 1, it indicates that merge mode with motion vector difference is used to generate the inter prediction parameters of the current coding unit. In other words, mmvd_merge_flag indicates whether MMVD is applied to the current block.
[0126] mmvd_cand_flag[x0][y0] specifies whether the first (0) or the second (1) candidate in the merging candidate list is used with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0].
[0127] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0].
[0128] mmvd_direction_idx[x0][y0] specifies the index used to derive MmvdSign[x0][y0].
[0129] merge_subblock_flag[x0][y0] indicates the subblock-based inter prediction parameters for the current coding. That is, merge_subblock_flag[x0][y0] indicates whether the subblock merge mode (or affine merge mode) is applied to the current block.
[0130] merge_subblock_idx[x0][y0] specifies the merging candidate index of the subblock-based merging candidate list.
[0131] ciip_flag[x0][y0] specifies whether the combined inter-picture merge and intra-picture prediction is applied for the current coding unit.
[0132] merge_triangle_idx0[x0][y0] specifies the first merging candidate index of the triangular shape based motion compensation candidate list.
[0133] merge_triangle_idx1[x0][y0] specifies the second merging candidate index of the triangular shape based motion compensation candidate list.
[0134] merge_idx[x0][y0] specifies the merging candidate index of the merging candidate list.
[0135] Meanwhile, referring again to the CU syntax in Table 1, mvp_l0_flag[x0][y0] indicates the index of the motion vector predictor of list 0. That is, when the MVP mode is applied, mvp_l0_flag indicates the candidate selected for MVP derivation of the current block in MVP candidate list 0.
[0136] mvp_l1_flag[x0][y0] has the same semantics as mvp_l0_flag, with l0, L0 and list 0 replaced by l1, L1 and list 1, respectively.
[0137] inter_pred_idc[x0][y0] specifies whether list0, list1, or bi-prediction is used for the current coding unit.
[0138] A value of 1 in sym_mvd_flag[x0][y0] indicates the syntax elements ref_idx_l0[x0][y0] and ref_idx_l1[x0][y0], and indicates that the mvd_coding(x0, y0, refList, cpIdx) syntax structure for refList equal to 1 does not exist (sym_mvd_flag[x0][y0] equal to 1 specifies that the syntax elements ref_idx_l0[x0][y0] and ref_idx_l1[x0][y0], and the mvd_coding(x0, y0, refList, cpIdx) syntax structure for refList equal to 1 are not present. In other words, sym_mvd_flag indicates whether symmetric MVD is used in mvd coding.
[0139] ref_idx_l0[x0][y0] specifies the list 0 reference picture index for the current coding unit.
[0140] ref_idx_l1[x0][y0] has the same semantics as ref_idx_l0, with l0, L0 and list 0 replaced by l1, L1 and list 1, respectively.
[0141] A value of 1 in inter_affine_flag[x0][y0] specifies that affine model based motion compensation is used to generate the prediction samples of the current coding unit when decoding a P or B slice for the current coding unit.
[0142] A value of 1 in cu_affine_type_flag[x0][y0] indicates that 6-parameter affine model-based motion compensation is used to generate the prediction samples of the current coding unit when decoding a P or B slice for the current coding unit. A value of 0 in cu_affine_type_flag[x0][y0] indicates that 4-parameter affine model-based motion compensation is used to generate the prediction samples of the current coding unit. (cu_affine_type_flag[x0][y0] equal to 1 specifies that for the current coding unit, when decoding a P or B slice, 6-parameter affine model-based motion compensation is used to generate the prediction samples of the current coding unit. cu_affine_type_flag[x0][y0] equal to 0 specifies that 4-parameter affine model-based motion compensation is used to generate the prediction samples of the current coding unit.)
[0143] amvr_flag[x0][y0] indicates the resolution of the motion vector differential. The array indexes x0, y0 indicate the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. A value of 0 in amvr_flag[x0][y0] indicates that the resolution of the motion vector differential is 1 / 4 of the luma sample. A value of 1 in amvr_flag[x0][y0] indicates that the resolution of the motion vector difference is additionally specified by amvr_precision_flag[x0][y0]. (amvr_flag[x0][y0] specifies the resolution of motion vector difference. The array indices x0, y0 specify the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. amvr_flag[x0][y0] equal to 0 specifies that the resolution of the motion vector difference is 1 / 4 of a luma sample. amvr_flag[x0][y0] equal to 1 specifies that the resolution of the motion vector difference is further specified by amvr_precision_flag[x0][y0].)
[0144] A value of 0 in amvr_precision_flag[x0][y0] indicates that the resolution of the motion vector differential is 1 integer luma sample if the value of inter_affine_flag[x0][y0] is 0, or 1 / 16 of a luma sample if not. A value of 1 in amvr_precision_flag[x0] indicates that the resolution of the motion vector differential is 4 luma samples if the value of inter_affine_flag[x0][y0] is 0, or 1 integer luma sample if not. The array indices x0, y0 indicate the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. (amvr_precision_flag[x0][y0] equal to 0 specifies that the resolution of the motion vector difference is one integer luma sample if inter_affine_flag[x0][y0] is equal to 0, and 1 / 16 of a luma sample otherwise. amvr_precision_flag[x0][y0] equal to 1 specifies that the resolution of the motion vector difference is four luma samples if inter_affine_flag[x0][y0] is equal to 0, and one integer luma sample otherwise. The array indices x0, y0 specify the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.)
[0145] bcw_idx[x0][y0] specifies the weight index of bi-prediction with CU weights.
[0146] When the (inter) prediction mode for the current block is determined, the coding apparatus derives motion information for the current block based on the prediction mode (S610).
[0147] A coding apparatus performs inter prediction using motion information of a current block. An encoding apparatus can derive optimal motion information for a current block through a motion estimation procedure. For example, the encoding apparatus can use an original block in an original picture for the current block to search for a similar reference block with high correlation in fractional pixel units within a predetermined search range in the reference picture, thereby deriving motion information. Block similarity is derived based on a phase-based sample value difference. For example, block similarity is calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, motion information can be derived based on the reference block with the smallest SAD within the search range. The derived motion information is signaled to a decoding apparatus in various ways depending on the inter prediction mode.
[0148] After deriving motion information for the current block, the coding apparatus performs inter prediction based on the motion information for the current block (S620). The coding apparatus may derive prediction sample(s) for the current block based on the motion information. The current block including the prediction sample(s) may be referred to as a predicted block.
[0149] Reconstructed samples and pictures are generated based on the derived predicted samples, after which procedures such as in-loop filtering can be performed.
[0150] FIG. 7 is a diagram illustrating merge mode and skip mode that can be used in inter prediction.
[0151] When a merge mode is applied during inter prediction, motion information of a current block is not directly transmitted, but is derived using motion information of neighboring predicted blocks. Therefore, the encoding apparatus can indicate motion information of the current block by transmitting flag information indicating that the merge mode is used and a merge index indicating which neighboring predicted block is used. The merge mode may also be referred to as a regular merge mode.
[0152] The coding device searches for merge candidate blocks to be used to derive motion information of the current block to perform the merge mode. For example, up to five merge candidate blocks may be used, but this embodiment is not limited to this. Also, information regarding the maximum number of merge candidate blocks may be transmitted in a slice header or a tile group header, but this embodiment is not limited to this. After finding the merge candidate blocks, the coding device may generate a merge candidate list and select the merge candidate block with the smallest cost as the final merge candidate block.
[0153] This document provides various examples for the merge candidate blocks that make up the merge candidate list.
[0154] The merge candidate list includes, for example, five merge candidate blocks. For example, four spatial merge candidates and one temporal merge candidate can be used. As a specific example, in the case of spatial merge candidates, the blocks (A0, A1, B0, B1, B2) shown in FIG. 7 can also be used as spatial merge candidates. Hereinafter, the spatial merge candidates or spatial MVP candidates described later may be referred to as SMVPs, and the temporal merge candidates or temporal MVP candidates described later may be referred to as TMVPs.
[0155] The merge candidate list for the current block is constructed, for example, according to the following procedure.
[0156] First, a coding apparatus (encoding apparatus / decoding apparatus) may search spatial neighboring blocks of a current block and insert derived spatial merge candidates into a merge candidate list. For example, the spatial neighboring blocks may include the lower left corner neighboring block (A0), the left side neighboring block (A1), the upper right corner neighboring block (B0), the upper side neighboring block (B1), and the upper left corner neighboring block (B2) of the current block. However, this is merely an example, and additional neighboring blocks such as a right side neighboring block, a lower side neighboring block, and a lower right side neighboring block may also be used as the spatial neighboring blocks. The coding apparatus may search the spatial neighboring blocks based on priority to detect available blocks and derive motion information of the detected blocks as the spatial merge candidates. For example, the encoding apparatus and / or decoding apparatus may search the five blocks shown in FIG. 7 in the order of A1, B1, B0, A0, and B2, and sequentially index available candidates to form a merge candidate list.
[0157] Furthermore, the coding apparatus may search for temporal neighboring blocks of the current block and insert derived temporal merge candidates into the merge candidate list. The temporal neighboring blocks may be located on a reference picture that is a different picture from the current picture in which the current block is located. The reference picture in which the temporal neighboring blocks are located may be called a collocated picture or col picture. The temporal neighboring blocks may be searched for in the order of a lower right corner neighboring block and a lower right center block of a co-located block with respect to the current block on the col picture.
[0158] Meanwhile, the coding apparatus checks whether the number of current merging candidates is smaller than the maximum number of merging candidates. The maximum number of merging candidates may be predefined or signaled from the encoding apparatus to the decoding apparatus. For example, the encoding apparatus generates information regarding the maximum number of merging candidates, encodes it, and transmits it to the decoding apparatus in the form of a bitstream. Once the maximum number of merging candidates is filled, no further candidate addition processes may be performed.
[0159] If the check result indicates that the number of current merge candidates is less than the maximum number of merge candidates, the coding device inserts additional merge candidates into the merge candidate list, including, for example, at least one of history-based merge candidate(s), pair-wise average merge candidate(s), ATMP, combined bi-predictive merge candidate (if the slice / tile group type of the current slice / tile group is type B), and / or zero vector merge candidate.
[0160] If the check result indicates that the number of current merge candidates is not less than the maximum number of merge candidates, the coding device terminates construction of the merge candidate list. In this case, the encoding device may select an optimal candidate from among the merge candidates constituting the merge candidate list based on a rate-distortion (RD) cost and signal selection information (e.g., a merge index) indicating the selected merge candidate to the decoding device. The decoding device may select the optimal merge candidate based on the merge candidate list and the selection information.
[0161] As described above, the motion information of the selected merging candidate may be used as the motion information of the current block, and a predicted sample of the current block may be derived based on the motion information of the current block. The encoding apparatus may derive a residual sample of the current block based on the predicted sample and signal residual information regarding the residual sample to a decoding apparatus. As described above, the decoding apparatus may generate reconstructed samples based on the residual samples derived based on the residual information and the predicted sample, and generate a reconstructed picture based on the reconstructed samples.
[0162] When a skip mode is applied during inter prediction, the motion information of the current block can be derived in the same manner as when the merge mode is applied, except that when the skip mode is applied, the residual signal for the corresponding block is omitted, and therefore, the predicted samples can be used directly as reconstructed samples.
[0163] FIG. 8 is a diagram illustrating a merge mode in which there is a motion vector difference that can be used during inter prediction.
[0164] In addition to the merge mode, where implicitly derived motion information is directly used for generating prediction samples of the current CU, the merge mode with motion vector differences (MMVD) is introduced in VVC. Because similar motion information derivation methods are used for the skip mode and the merge mode, MMVD may be applied to the skip mode. An MMVD flag (e.g., mmvd_flag) may be signaled right after sending a skip flag and a merge flag to specify whether the MMVD mode is used for a CU.
[0165] In MMVD, after a merge candidate is selected, it is further refined by the signaled MVD information. When MMVD is applied to the current block (i.e., when the mmvd_flag is equal to 1), further information for the MMVD may be signaled.
[0166] The additional information includes a merge candidate flag (e.g., mmvd_merge_flag) indicating whether the first (0) or second (1) candidate in the merging candidate list is used with the motion vector difference, an index to specify the motion magnitude (e.g., mmvd_distance_idx), and an index for indication of the motion direction (e.g., mmvd_direction_idx). In MMVD mode, one of the first two candidates in the merge list is selected to be used as the basis for MV. The merge candidate flag is signaled to specify which one is used.
[0167] The distance index specifies motion magnitude information and indicates the pre-defined offset from the starting point.
[0168] As shown in Figure 8, an offset can be added to either the horizontal component or the vertical component of the starting MV. The relationship between the distance index and the pre-defined offset is shown in Table 3 below.
[0169] [Table 3]
[0170] Here, slice_fpel_mmvd_enabled_flag equal to 1 specifies that merge mode with motion vector difference uses integer sample precision in the current slice. When slice_fpel_mmvd_enabled_flag is 0, it specifies that merge mode with motion vector difference can use fractional sample precision in the current slice. When not present, the value of slice_fpel_mmvd_enabled_flag is inferred to be 0. The slice_fpel_mmvd_enabled_flag syntax element may be signaled through (may be comprised in) a slice header.
[0171] The direction index indicates the direction of the MVD relative to the starting point. The direction index represents the direction of the MVD relative to the starting point, as shown in Table 4. The meaning of the MVD sign can be varied depending on the starting MVs. When the starting MV is a non-prediction MV or a bi-prediction MV with two lists point to the same side of the current picture (for example, when the POCs of the two references are both larger than the POC of the current picture, or both smaller than the POC of the current picture), the sign in Table 4 specifies the sign of the MV offset added to the starting MV.When the starting MV is a bi-prediction MV with two MVs pointing to other sides of the current picture (i.e., the POC of one reference is larger than the POC of the current picture, and the POC of the other reference is smaller than the POC of the current picture), the sign in Table 4 specifies the sign of the MV offset added to the list0 MV component of the starting MV, and the sign for the list1 MV has the opposite value.
[0172] [Table 4]
[0173] Both components of the merge plus MVD offset MmvdOffset[x0][y0] are derived as follows.
[0174]
number
[0175] 9 and 10 are diagrams illustrating a sub-block-based temporal motion vector prediction process that can be used during inter prediction.
[0176] The subblock-based temporal motion vector prediction (SbTMVP) method can be used for inter prediction. Similar to TMVP (temporal motion vector prediction), SbTMVP uses the motion field in the collocated picture to improve motion vector prediction and merge mode for CUs in the current picture. The same collocated picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects:
[0177] 1. TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level.
[0178] 2. Whereas TMVP fetches the temporal motion vectors from the collocated block in the collocated picture (the collocated block is the bottom-right or center (below-right center) block relative to the current CU), SbTMVP applies a motion shift before fetching the temporal motion information from the collocated picture, where the motion shift is obtained from the motion vector from one of the spatial neighboring blocks of the current CU.
[0179] The SbTVMP process is shown in Figures 9 and 10. SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 9 is examined. When a reference picture is identified, if A1 has a motion vector to use as a collocated picture, this motion vector (which may be referred to as a temporal MV (tempMV)) is selected as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0,0).
[0180] In the second step, the motion shift identified in Step 1 is applied (i.e., added as a candidate for the current block) to obtain sub-CU-level motion information (motion vectors and reference indices) from the collocated picture as shown in Figure 10. The example in Figure 10 assumes the motion shift is set to block A1's motion. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid that covers the center sample) in the collocated picture is used to derive the motion information for the sub-CU.The center block (below right center sample) may correspond to a below-right sample among four central samples in a sub-CU when the sub-block has even length, width, and height.
[0181] After the motion information of the collocated sub-CU is identified, it is converted to the motion vectors and reference indices of the current sub-CU in a similar way as the TMVP process, where temporal motion scaling may be applied to align the reference pictures of the temporal motion vectors to those of the current CU.
[0182] A combined sub-block based merge list containing both SbTVMP candidates and affine merge candidates may be used for signaling affine merge mode (may be referred to as sub-block (based) merge mode). The SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block merge candidates, followed by the affine merge candidates. The maximum allowed size of the affine merge candidate list may be 5.
[0183] The sub-CU size used in SbTMVP may be fixed to be 8x8, and as done for affine merge mode, SbTMVP mode may be only applicable to the CU with both width and height larger than or equal to 8.
[0184] The encoding logic of the additional SbTMVP merge candidate is the same as for the other merge candidates, that is, for each CU in a P or B slice, an additional RD check may be performed to decide whether to use the SbTMVP candidate.
[0185] FIG. 11 is a diagram illustrating partitioning modes that can be applied to inter prediction.
[0186] A triangle partition mode may be used for inter prediction. The triangle partition mode may only be applied to CUs that are 8x8 or larger. The triangle partition mode is signaled using a CU-level flag as one kind of merge mode, with other merge modes including the regular merge mode, the MMVD mode, the CIIP mode, and the subblock merge mode.
[0187] When this mode is used, a CU may be split evenly into two triangle-shaped partitions using a diagonal or semi-diagonal split, as shown in Figure 11. Each triangle partition in a CU is inter-predicted using its own motion; only uni-prediction is allowed for each partition. That is, each partition has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure that only two motion-compensated predictions are needed for each CU, just like in conventional bi-prediction.
[0188] If triangle partition mode is used for the current CU, then a flag indicating the direction of the triangle partition (diagonal or semi-diagonal) and two merge indices (one for each partition) are further signaled. The number of maximum TPM (triangle partition mode) candidate sizes is signaled explicitly at the slice level and specifies syntax binarization for TPM merge indices. After predicting each triangle partition, the sample values along the diagonal or semi-diagonal edge are adjusted using a blending process with adaptive weights.This is the prediction signal for the whole CU, and the transform and quantization process will be applied to the whole CU as in other prediction modes. Finally, the motion field of a CU predicted using the triangle partition mode is stored in 4x4 units. The triangle partition mode is not used in combination with SBT (subblock transform). That is, when the signaled triangle mode value is 1, the cu_sbt_flag is inferred to be 0 without signaling.
[0189] The uni-prediction candidate list is derived directly from the merge candidate list constructed as described above.
[0190] After predicting each triangle partition using its own motion, blending is applied to the two prediction signals to derive samples around the diagonal or anti-diagonal edge.
[0191] FIG. 12 is a diagram illustrating a CIIP mode that can be applied to inter prediction.
[0192] Combined inter and intra prediction can be applied to a current block. An additional flag (e.g., ciip_flag) may be signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. For example, when a CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., the product of the CU width and CU height is greater than or equal to 64 luma samples) and the CU width and CU height are all less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As its name indicates, the CIIP prediction combines an inter prediction signal with an intra prediction signal.The inter prediction signal in the CIIP mode P_inter is derived using the same inter prediction process applied to regular merge mode, and the intra prediction signal P_intra is derived following the regular intra prediction process with the planar mode. Then, the intra and inter prediction signals are combined using weighted averaging, where the weight value is calculated depending on the coding modes of the top and left neighboring blocks (shown in Figure 12) as follows:
[0193] If the top neighbor is available and intra-coded, then set isIntraTop to 1; otherwise, set isIntraTop to 0.
[0194] If the left neighbor is available and intra-coded, then set isIntraLeft to 1, otherwise set isIntraLeft to 0.
[0195] If (isIntraLeft + isIntraLeft) is equal to 2, then wt is set to 3.
[0196] Otherwise, if (isIntraLeft + isIntraLeft) is equal to 1, then wt is set to 2.
[0197] Otherwise, set wt to 1.
[0198] The CIIP prediction is formed as follows:
[0199]
number
[0200] Meanwhile, to generate a prediction block in a coding device, motion information can be derived based on the regular merge mode, skip mode, SbTMVP mode, MMVD mode, triangle partition mode (partitioning mode), and / or CIIP mode. Each mode is activated / disabled via an on / off flag for each mode included in a sequence parameter set (SPS). If the on / off flag for a specific mode is deactivated, the encoding device does not signal an explicit syntax for the corresponding prediction mode on a CU or PU basis.
[0201] Therefore, if all or some of the specific modes for merge / skip modes are disabled in the existing operation process, a problem occurs in which the on / off flag is signaled redundantly. Therefore, in this document, one of the following methods is used to prevent the same information (flag) from being signaled redundantly in the process of selecting the merge mode to be applied to the current block based on the merge data syntax in Table 2.
[0202] In order to select a prediction mode that can be used in the process of deriving motion information, the encoding apparatus can signal a flag based on a sequence parameter set such as that shown in Table 5. Each prediction mode is turned on / off based on the sequence parameter set shown in Table 5, and each syntax element of the merge data syntax shown in Table 2 is parsed or induced according to the flags shown in Table 5 and the conditions under which each mode can be used.
[0203] [Table 5-1]
[0204] [Table 5-2]
[0205] [Table 5-3]
[0206] [Table 5-4]
[0207] [Table 5-5]
[0208] [Table 5-6]
[0209] [Table 5-7]
[0210] [Table 5-8]
[0211] [Table 5-9]
[0212] [Table 5-10]
[0213] The following drawings are created to explain a specific example of the present document. The names of specific devices and names of specific signals / information shown in the drawings are presented for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.
[0214] 13 and 14 illustrate an example of a video / picture encoding method and associated components including an inter-prediction method according to an embodiment of the present document.
[0215] The encoding method disclosed in Fig. 13 may be performed by the encoding apparatus 200 disclosed in Fig. 2. Specifically, for example, steps S1300 and S1310 in Fig. 13 are performed by the prediction unit 220 of the encoding apparatus 200, and step S1320 is performed by the entropy encoding unit 240 of the encoding apparatus 200. The encoding method disclosed in Fig. 13 includes the embodiments described above in this document.
[0216] 13 and 14, a prediction unit of an encoding apparatus determines a prediction mode of a current block (S1300). For example, if inter prediction is applied to the current block, the prediction unit of the encoding apparatus determines one of a regular merge mode, a skip mode, an MMVD mode, a sub-block merge mode, a partitioning mode, and a CIIP mode as the prediction mode of the current block.
[0217] Here, the regular merge mode is defined as a mode in which motion information of a current block is derived using motion information of neighboring blocks. The skip mode is defined as a mode in which a predicted block is used as a reconstruction block. The MMVD mode is applied to the merge mode or the skip mode and is defined as a merge (or skip) mode using a motion vector difference. The sub-block merge mode is defined as a merge mode based on a sub-block. The partitioning mode is defined as a mode in which prediction is performed by dividing the current block into two partitions (diagonal or semi-diagonal). The CIIP mode is defined as a mode in which inter-picture merge and intra-picture prediction are combined.
[0218] Meanwhile, the prediction unit of the encoding apparatus may search for blocks similar to the current block within a certain region (search region) of the reference picture through motion estimation, derive a reference block whose difference with the current block is minimum or equal to or less than a certain criterion, and based on the reference block, derive a reference picture index indicating the reference picture in which the reference block is located. Also, based on the position difference between the reference block and the current block, derive a motion vector.
[0219] A prediction unit of an encoding apparatus may generate a prediction sample (prediction block) of a current block based on a prediction mode of the current block and a motion vector of the current block, and may generate information about the prediction mode based on the prediction mode (S1310). Here, the information about the prediction mode may include inter / intra prediction classification information, inter prediction mode information, etc., and various syntax elements related thereto.
[0220] The residual processing unit of the encoding device generates residual samples based on original samples (original block) for the current block and predicted samples (predicted block) for the current block, and can derive information about the residual samples based on the residual samples.
[0221] An encoding unit of the encoding device encodes video information including information on the residual samples and information on the prediction mode (S1320). The video information includes partitioning-related information, prediction mode-related information, residual information, in-loop filtering-related information, and various syntax elements related thereto. The information encoded by the encoding unit of the encoding device is output in the form of a bitstream. The bitstream is transmitted to a decoding device via a network or a storage medium.
[0222] For example, the video information includes information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video information also includes information on a prediction mode of a current block such as a coding unit syntax, a merge data syntax, etc. Here, the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, an enable flag for a partitioning mode, etc. The coding unit syntax includes a CU skip flag indicating whether a skip mode is applied to the current block.
[0223] According to an embodiment, for example, the encoding device may include a regular merge flag in the video information when a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied to prevent duplicate transmission of the same syntax. Here, the condition based on the size of the current block may be when the product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag may be when the value of the CIIP availability flag is 1. That is, the encoding device may signal a regular merge flag when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1.
[0224] As another example, the encoding device may include the regular merge flag in the video information when a condition based on the CU skip flag and the size of the current block is satisfied. Here, the condition based on the CU skip flag may be when the value of the CU skip flag is 0. In other words, the encoding device may signal the regular merge flag when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, and the value of the CU skip flag is 0.
[0225] As another example, the encoding device may include the regular merge flag in the video information when a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP availability flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag may be when the value of the CU skip flag is 0. In other words, the encoding device may signal a regular merge flag when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, the value of the CIIP availability flag is 1, and the value of the CU skip flag is 0.
[0226] As another example, the encoding device may include a regular merge flag in the video information when a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes when the product of the width and height of the current block is 64 or more and / or when the type of the slice including the current block is a B slice. The condition based on the partitioning mode available flag may be when the value of the partitioning mode available flag is 1. That is, the encoding device may signal a regular merge flag when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0227] The encoding device determines whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the CIIP availability flag and the size of the current block are not satisfied, or determines whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the information about the current block and the partitioning mode available flag is not satisfied,
[0228] On the other hand, the encoding device may signal a regular merge flag if the product of the width and height of the current block is not 32 and the value of the MMVD available flag is 1, or if the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are each 8 or greater.
[0229] For this purpose, as an example, the merge data syntax is configured as shown in Table 6 below.
[0230] [Table 6-1]
[0231] [Table 6-2]
[0232] [Table 6-3]
[0233] In Table 6, a value of 1 for the regular_merge_flag indicates that regular merge mode is used to generate the inter prediction parameters of the current coding unit (current block). The array indices x0, y0 specify the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.
[0234] When regular_merge_flag[x0][y0] is not present, it is inferred as follows:
[0235] If all of the following conditions are true, the value of regular_merge_flag[x0][y0] is inferred to be equal to 1.
[0236] -The value of the general merge flag is 1 (general_merge_flag[x0][y0] is equal to 1)
[0237] -The value of the SPS MMVD enable flag is 0 or the product of the width and height of the current block is 32 (sps_mmvd_enable_flag is equal to 0 or cbWidth*cbHeight==32)
[0238] -The maximum number of subblock merge candidates is 0 or less, or the width of the current block is less than 8, or the height of the current block is less than 8 (MaxNumSubblockMergeCand<=0 or cbWidth<8 or cbHeight<8)
[0239] -The value of the SPS CIIP enabled flag is 0, or the product of the width and height of the current block is less than 64, or the width of the current block is 128 or more, or the value of the CU skip flag is 1 (sps_ciip_enabled_flag is equal to 0 or cbWidth*cbHeight<64 or cbWidth>=128 or cbHeight>=128 or cu_skip_flag[x0][y0] is equal to 1)
[0240] -The value of the SPS partitioning enabled flag is 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not B slice (sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand<2 or slice_type is not equal to B_SLICE)
[0241] Otherwise, the value of regular_merge_flag[x0][y0] is inferred to be equal to 0.
[0242] Meanwhile, according to another embodiment, for example, the encoding device may include an MMVD merge flag in the video information when a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied, so that the same syntax is not transmitted redundantly. Here, the condition based on the size of the current block may be when the product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. In other words, the encoding device may signal an MMVD merge flag when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1.
[0243] As another example, the encoding device may include the MMVD merge flag in the video information if a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP availability flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag is when the value of the CU skip flag is 0. That is, the encoding device may signal the MMVD merge flag when the product of the height and width of the current block is 64 or more, the height and width of the current block are each less than 128, the value of the CIIP availability flag is 1, and the value of the CU skip flag is 0.
[0244] As another example, the encoding device may include an MMVD merge flag in the video information when a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes when the product of the width and height of the current block is 64 or more and / or when the type of the slice including the current block is a b-slice. The condition based on the partitioning mode available flag is when the value of the partitioning mode available flag is 1. That is, the encoding device may signal an MMVD merge flag when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0245] The encoding device may determine whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the CIIP available flag and the size of the current block are not satisfied, or may determine whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the information about the current block and the partitioning mode available flag is not satisfied.
[0246] On the other hand, the encoding device may also signal the MMVD merge flag if the product of the width and height of the current block is not 32 and the value of the MMVD available flag is 1, or if the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are each 8 or greater.
[0247] For this purpose, as an example, the merge data syntax is configured as shown in Table 7 below.
[0248] [Table 7-1]
[0249] [Table 7-2]
[0250] [Table 7-3]
[0251] [Table 7-4]
[0252] A value of 1 in the MMVD merge flag indicates that merge mode with motion vector difference is used to generate the inter prediction parameters of the current coding unit (current block). The array indices x0, y0 specify the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.
[0253] When the MMVD merge flag is not present, it is derived as follows:
[0254] If all of the following conditions are true, the value of the MMVD merge flag is inferred to be equal to 1.
[0255] -The value of the general merge flag is 1 (general_merge_flag[x0][y0] is equal to 1)
[0256] -The value of the regular merge flag is 0 (regular_merge_flag[x0][y0] is equal to 0)
[0257] -The value of the SPS MMVD enable flag is 1 (sps_mmvd_enable_flag is equal to 1)
[0258] -The product of the current block's width and height is not 32 (cbWidth*cbHeight!=32)
[0259] -The maximum number of subblock merge candidates is 0 or less, or the width of the current block is less than 8, or the height of the current block is less than 8 (MaxNumSubblockMergeCand<=0 or cbWidth<8 or cbHeight<8)
[0260] -The value of the SPS CIIP enabled flag is 0 or the width of the current block is 128 or more or the height of the current block is 128 or more or the value of the CU skip flag is 1 (sps_ciip_enabled_flag is equal to 0 or cbWidth>=128 or cbHeight>=128 or cu_skip_flag[x0][y0] is equal to 1)
[0261] -The value of the SPS partitioning enabled flag is 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not B slice (sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand<2 or slice_type is not equal to B_SLICE)
[0262] Otherwise, the value of the MMVD merge flag is inferred to be equal to 0.
[0263] Meanwhile, according to another embodiment, for example, the encoding device may include a merge sub-block flag in the video information when a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied so that the same syntax is not transmitted redundantly. Here, the condition based on the size of the current block is when the product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. In other words, the encoding device may signal a merge sub-block flag when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1.
[0264] As another example, the encoding device may include the remaining sub-block flags in the video information if a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP available flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag is when the value of the CU skip flag is 0. In other words, the encoding device may signal a merge sub-block flag when the product of the height and width of the current block is 64 or more, the height and width of the current block are each less than 128, the value of the CIIP available flag is 1, and the value of the CU skip flag is 0.
[0265] As another example, the encoding device may include a merge sub-block flag in the video information when a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes when a product of the width and height of the current block is 64 or more and / or when a type of slice including the current block is a B slice. The condition based on the partitioning mode available flag is when the value of the partitioning mode available flag is 1. That is, the encoding device may signal a merge sub-block flag when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0266] The encoding device may determine whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the CIIP available flag and the size of the current block are not satisfied, or may determine whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the information about the current block and the partitioning mode available flag is not satisfied.
[0267] Meanwhile, the encoding apparatus may signal a merge sub-block flag if the maximum number of sub-block merging candidates is greater than 0 and the width and height of the current block are equal to or greater than 8.
[0268] For this purpose, as an example, the merge data syntax is configured as shown in Table 8 below.
[0269] [Table 8-1]
[0270] [Table 8-2]
[0271] [Table 8-3]
[0272] The merge_subblock_flag[x0][y0] specifies whether the subblock-based inter prediction parameters for the current coding unit are inferred from neighboring blocks. The array indices x0, y0 specify the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.
[0273] If the merge_subblock_flag[x0][y0] is not present, it is inferred as follows:
[0274] If all of the following conditions are true, the value of the merge subblock flag is inferred to be equal to 1.
[0275] -The value of the general merge flag is 1 (general_merge_flag[x0][y0] is equal to 1)
[0276] -The value of the regular merge flag is 0 (regular_merge_flag[x0][y0] is equal to 0)
[0277] -The value of the merge subblock flag is 0 (merge_subblock_flag[x0][y0] is equal to 0)
[0278] -MMVD merge flag value is 0 (mmvd_merge_flag[x0][y0] is equal to 0)
[0279] -The maximum number of subblock merge candidates is greater than 0 (MaxNumSubblockMergeCand>0).
[0280] -The width and height of the current block are greater than or equal to 8 (cbWidth>=8 and cbHeight>=8)
[0281] -The value of the SPS CIIP enabled flag is 0 or the width of the current block is 128 or more or the height of the current block is 128 or more or the value of the CU skip flag is 1 (sps_ciip_enabled_flag is equal to 0 or cbWidth>=128 or cbHeight>=128 or cu_skip_flag[x0][y0] is equal to 1)
[0282] -The value of the SPS partitioning enabled flag is 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not B slice (sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand<2 or slice_type is not equal to B_SLICE)
[0283] Otherwise, the value of the merge subblock flag is inferred as 0 (Otherwise, merge_subblock_flag[x0][y0] is inferred to be equal to 0.).
[0284] Meanwhile, according to another embodiment, for example, the encoding device may include a CIIP flag in the video information when a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied so that the same syntax is not transmitted redundantly. Here, the condition based on the size of the current block may be when the product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. That is, the encoding device may signal a CIIP flag when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1.
[0285] As another example, the encoding device may include the CIIP flag in the video information when a condition based on a CU Skip flag is further satisfied in addition to the condition based on the CIIP available flag and the condition based on the size of the current block. Here, the condition based on the CU Skip flag is when the value of the CU Skip flag is 0. In other words, the encoding device may signal the CIIP flag when the product of the height and width of the current block is 64 or more, the height and width of the current block are each less than 128, the value of the CIIP available flag is 1, and the value of the CU Skip flag is 0.
[0286] As another example, the encoding device may include a CIIP flag in the video information based on whether a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes whether the product of the width and height of the current block is 64 or greater and / or whether the type of slice including the current block is a B slice. The condition based on the partitioning mode available flag is whether the value of the partitioning mode available flag is 1. In other words, the encoding device may signal a CIIP flag when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0287] The encoding device determines whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the CIIP availability flag and the size of the current block are not satisfied, or determines whether a condition based on the information about the current block and the partitioning mode available flag is satisfied when the condition based on the information about the current block and the partitioning mode available flag is not satisfied,
[0288] For this purpose, as an example, the merge data syntax is configured as shown in Table 9 below.
[0289] [Table 9-1]
[0290] [Table 9-2]
[0291] [Table 9-3]
[0292] The CIIP flag indicates whether the combined inter-picture merge and intra-picture prediction is applied for the current coding unit (ciip_flag[x0][y0] specifies whether the combined inter-picture merge and intra-picture prediction is applied for the current coding unit. The array indices x0, y0 specify the location (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.
[0293] When the CIIP flag is not present, it is derived as follows:
[0294] If all of the following conditions are true, the value of the CIIP flag is inferred to be equal to 1.
[0295] -The value of the general merge flag is 1 (general_merge_flag[x0][y0] is equal to 1)
[0296] -Regular merge flag value is 0 (regular_merge_flag[x0][y0] is euqal to 0)
[0297] -The value of the merge subblock flag is 0 (merge_subblock_flag[x0][y0] is equal to 0)
[0298] -MMVD merge flag value is 0 (mmvd_merge_flag[x0][y0] is equal to 0)
[0299] -The value of the SPS CIIP enabled flag is 1 (sps_ciip_enabled_flag is equal to 1)
[0300] -CU skip flag value is 0 (cu_skip_flag[x0][y0] is equal to 0)
[0301] - The product of the width and height of the current block is 64 or more, and the width and height of the current block are less than 128 (cbWidth*cbHeight>=64 and cbWidth<128 and cbHeight<128).
[0302] -The value of the SPS partitioning enabled flag is 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not B slice (sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand<2 or slice_type is not equal to B_SLICE)
[0303] Otherwise, the value of the CIIP flag is inferred to be equal to 0.
[0304] 15 and 16 illustrate an example of a video / picture decoding method and associated components including an inter-prediction method according to an embodiment of the present document.
[0305] The decoding method disclosed in Figure 15 may be performed by the decoding apparatus 300 disclosed in Figures 3 and 16. Specifically, for example, steps S1500 to S1520 in Figure 15 are performed by the prediction unit 330 of the decoding apparatus 300, and step S1530 is performed by the addition unit 340 of the decoding apparatus 300. The decoding method disclosed in Figure 15 includes the embodiments described above in this document.
[0306] As shown in Figures 15 and 16, a decoding apparatus obtains information about a prediction mode of a current block from a bitstream (S1500). Specifically, the entropy decoding unit 310 of the decoding apparatus derives residual information and information about the prediction mode from a signal received in the form of a bitstream from the encoding apparatus of Figure 2. Here, the information about the prediction mode may be referred to as prediction-related information. The information about the prediction mode includes inter / intra prediction classification information, inter prediction mode information, etc., and includes various syntax elements related thereto.
[0307] The bitstream includes video information including information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video information further includes information on a prediction mode of a current block such as a coding unit syntax and a merge data syntax. The sequence parameter set includes a CIIP available flag, an available flag for a partitioning mode, etc. The coding unit syntax includes a CU skip flag indicating whether a skip mode is applied to the current block.
[0308] Meanwhile, the residual processing unit 320 of the decoding device generates residual samples based on the residual information. The prediction unit 330 of the decoding device derives a prediction mode for the current block based on the prediction mode information (S1510). The prediction unit 330 can also derive motion information for the current block based on the derived prediction mode. The prediction unit of the decoding device can generate a motion information candidate list based on neighboring blocks of the current block and derive a motion vector and / or a reference picture index for the current block based on candidate selection information received from the encoding device. After the motion information for the current block is derived, the prediction unit of the decoding device generates predicted samples for the current block based on the motion information for the current block (S1520). The adder 340 of the decoding device then generates reconstructed samples based on the predicted samples generated by the prediction unit 330 and the residual samples generated by the residual processing unit 320 (S1530). A reconstructed picture can be generated based on the reconstructed samples. Afterwards, in-loop filtering procedures such as deblocking filtering, SAO and / or ALF procedures can be applied to the reconstructed pictures to improve the subjective / objective image quality if necessary.
[0309] In one embodiment, the decoding device obtains a regular merge flag from the bitstream when deriving the prediction mode of the current block, based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied. Here, the condition based on the size of the current block may be when a product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. In other words, when a product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1, the decoding device may parse the regular image flag from merge data syntax included in the bitstream.
[0310] As another example, a regular merge flag may be obtained from the bitstream based on whether a condition based on a CU skip flag and a size of a current block is satisfied. Here, the condition based on the size of the current block may be that the product of the height and width of the current block is greater than or equal to 64, and that the height and width of the current block are each less than 128. The condition based on the CU skip flag is that the value of the CU skip flag is 0. In other words, if the product of the height and width of the current block is greater than or equal to 64, that the height and width of the current block are each less than 128, and the value of the CU skip flag is 0, the decoding device may parse the regular image flag from merge data syntax included in the bitstream.
[0311] As another example, the decoding device may acquire the regular merge flag from the bitstream based on whether a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP available flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag is when the value of the CU skip flag is 0. In other words, the decoding device may parse the regular merge flag from the merge data syntax when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, the value of the CIIP available flag is 1, and the value of the CU skip flag is 0.
[0312] As another example, the decoding device may acquire the regular merge flag from the bitstream based on whether a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes whether the product of the width and height of the current block is 64 or greater and / or whether the type of slice including the current block is a B slice. The condition based on the partitioning mode available flag is whether the value of the partitioning mode available flag is 1. In other words, the decoding device may parse the regular merge flag from the merge data syntax when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0313] If the condition based on the CIIP availability flag and the condition based on the size of the current block are not satisfied, the decoding device determines whether the condition based on information about the current block and the partitioning mode availability flag is satisfied. Alternatively, if the condition based on information about the current block and the partitioning mode availability flag is not satisfied, the decoding device may determine whether the condition based on the CIIP availability flag and the condition based on the size of the current block are satisfied.
[0314] Meanwhile, the decoding device may parse the regular image flag from the bitstream if the product of the width and height of the current block is not 32 and the value of the MMVD available flag is 1, or if the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are each equal to or greater than 8. For this purpose, the merge data syntax is configured as shown in Table 6 above.
[0315] If the bitstream does not contain a regular merge flag, the decoding device may derive the regular merge plug as a value of 1 if the general merge flag has a value of 1, the MMVD available flag of the SPS has a value of 0, or the product of the width and height of the current block is 32, and the maximum number of sub-block merge candidates is 0 or less, or the width of the current block is less than 8, or the height of the current block is less than 8, the CIIP available flag of the SPS has a value of 0, or the product of the width and height of the current block is less than 64, or the width of the current block is 128 or more, or the CU skip flag has a value of 1, and the partitioning available flag of the SPS has a value of 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not a B slice. Otherwise, the regular merge plug is derived as a value of 0.
[0316] In another embodiment, the decoding device may acquire an MMVD merge flag from the bitstream when deriving the prediction mode of the current block, based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied. Here, the condition based on the size of the current block may be when the product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. In other words, when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1, the decoding device may parse an MMVD merge flag from merge data syntax included in the bitstream.
[0317] As another example, the MMVD merge flag may be obtained from the bitstream based on whether a condition based on a CU skip flag and a size of the current block is satisfied. Here, the condition based on the size of the current block may be that the product of the height and width of the current block is greater than or equal to 64, and that the height and width of the current block are each less than 128. The condition based on the CU skip flag is that the value of the CU skip flag is 0. In other words, if the product of the height and width of the current block is greater than or equal to 64, that the height and width of the current block are each less than 128, and the value of the CU skip flag is 0, the decoding device may parse the MMVD merge flag from the merge data syntax included in the bitstream.
[0318] As another example, the decoding device may acquire the MMVD merge flag from the bitstream based on whether a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP availability flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag is when the value of the CU skip flag is 0. In other words, the decoding device may parse the MMVD merge flag from the merge data syntax when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, the value of the CIIP availability flag is 1, and the value of the CU skip flag is 0.
[0319] As another example, the decoding device may obtain an MMVD merge flag from the bitstream based on whether a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes whether the product of the width and height of the current block is 64 or greater and / or whether the type of slice including the current block is a B slice. The condition based on the partitioning mode available flag is whether the value of the partitioning mode available flag is 1. In other words, the decoding device may parse an MMVD merge flag from the merge data syntax when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0320] If the condition based on the CIIP availability flag and the condition based on the size of the current block are not satisfied, the decoding device determines whether the condition based on information about the current block and the partitioning mode availability flag is satisfied. Alternatively, if the condition based on information about the current block and the partitioning mode availability flag is not satisfied, the decoding device may determine whether the condition based on the CIIP availability flag and the condition based on the size of the current block are satisfied.
[0321] Meanwhile, the decoding device can also parse the MMVD merge flag from the bitstream if the product of the width and height of the current block is not 32 and the value of the MMVD available flag is 1, or if the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are each equal to or greater than 8. For this purpose, the merge data syntax is configured as shown in Table 7 above.
[0322] The decoding device may derive the value of the MMVD merge flag as 1 if the bitstream does not contain an MMVD merge flag, the value of the general merge flag is 1, the value of the regular merge flag is 0, the value of the MMVD available flag in the SPS is 1, the product of the width and height of the current block is not 32, the maximum number of sub-block merge candidates is 0 or less, or the width of the current block is less than 8, or the height of the current block is less than 8, the value of the CIIP available flag in the SPS is 0, or the width of the current block is 128 or more, or the height of the current block is 128 or more, or the value of the CU skip flag is 1, the value of the partitioning available flag in the SPS is 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not a B slice. Otherwise, the value of the MMVD merge flag may be derive the value of 0.
[0323] In another embodiment, the decoding device may acquire merge sub-block flags from the bitstream when deriving the prediction mode of the current block, based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied. Here, the condition based on the size of the current block may be when a product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. In other words, when a product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1, the decoding device may parse the merge sub-block flags from merge data syntax included in the bitstream.
[0324] As another example, a merge sub-block flag may be acquired from the bitstream based on whether a condition based on a CU skip flag and a size of a current block is satisfied. Here, the condition based on the size of the current block may be that the product of the height and width of the current block is greater than or equal to 64, and that the height and width of the current block are each less than 128. The condition based on the CU skip flag is that the value of the CU skip flag is 0. In other words, if the product of the height and width of the current block is greater than or equal to 64, that the height and width of the current block are each less than 128, and the value of the CU skip flag is 0, the decoding device may parse the merge sub-block flag from merge data syntax included in the bitstream.
[0325] As another example, the decoding device may acquire merge sub-block flags from the bitstream based on whether a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP availability flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag is when the value of the CU skip flag is 0. In other words, the decoding device may parse merge sub-block flags from the merge data syntax when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, the value of the CIIP availability flag is 1, and the value of the CU skip flag is 0.
[0326] As another example, the decoding device may acquire a merge sub-block flag from the bitstream based on whether a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes whether the product of the width and height of the current block is 64 or greater and / or whether the type of slice including the current block is a B slice. The condition based on the partitioning mode available flag is whether the value of the partitioning mode available flag is 1. In other words, the decoding device may parse a merge sub-block flag from the merge data syntax when both the condition based on the height of the current block and information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0327] If the condition based on the CIIP availability flag and the condition based on the size of the current block are not satisfied, the decoding device may determine whether the condition based on information about the current block and the partitioning mode availability flag is satisfied. Alternatively, if the condition based on information about the current block and the partitioning mode availability flag is not satisfied, the decoding device may determine whether the condition based on the CIIP availability flag and the condition based on the size of the current block are satisfied.
[0328] Meanwhile, the decoding apparatus may parse the merge sub-block flag from the bitstream if the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are each greater than or equal to 8. For this purpose, the merge data syntax is configured as shown in Table 8 above.
[0329] If the bitstream does not contain a merge sub-block flag, the decoding device may derive the value of the merge sub-block flag as 1 if the general merge flag has a value of 1, the regular merge flag has a value of 0, the merge sub-block flag has a value of 0, the MMVD merge flag has a value of 0, the maximum number of sub-block merge candidates is greater than 0, the width and height of the current block are each 8 or greater, the CIIP available flag of the SPS has a value of 0, or the width of the current block is 128 or greater, or the height of the current block is 128 or greater, or the CU skip flag has a value of 1, or the partitioning available flag of the SPS has a value of 9, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not a B slice. Otherwise, the decoding device may derive the value of the merge sub-block flag as 0.
[0330] In another embodiment, the decoding device may acquire the CIIP flag from the bitstream when deriving the prediction mode of the current block, based on whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied. Here, the condition based on the size of the current block may be when a product of the height and width of the current block is greater than or equal to 64, and the height and width of the current block are each less than 128. The condition based on the CIIP availability flag is when the value of the CIIP availability flag is 1. In other words, the decoding device may parse the CIIP flag from the merge data syntax when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, and the value of the CIIP availability flag is 1.
[0331] As another example, the decoding device may acquire the CIIP flag from the bitstream based on whether a condition based on a CU skip flag is further satisfied in addition to the condition based on the CIIP available flag and the condition based on the size of the current block. Here, the condition based on the CU skip flag is when the value of the CU skip flag is 0. In other words, the decoding device may parse the CIIP flag from the merge data syntax when the product of the height and width of the current block is greater than or equal to 64, the height and width of the current block are each less than 128, the value of the CIIP available flag is 1, and the value of the CU skip flag is 0.
[0332] As another example, the decoding device may acquire a CIIP flag from a bitstream based on whether a condition based on information about the current block and the partitioning mode available flag is satisfied. Here, the condition based on the information about the current block includes whether the product of the width and height of the current block is 64 or greater and / or whether the type of slice including the current block is a B slice. The condition based on the partitioning mode available flag is whether the value of the partitioning mode available flag is 1. In other words, the decoding device may parse a CIIP flag from a merge data syntax when both the condition based on the height of the current block and the information about the current block and the condition based on the partitioning mode available flag are satisfied.
[0333] The decoding device determines whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied when the condition based on the CIIP availability flag and a condition based on the size of the current block are not satisfied. Alternatively, the decoding device may determine whether a condition based on the CIIP availability flag and a condition based on the size of the current block are satisfied when the condition based on the information on the current block and a condition based on the partitioning mode availability flag are not satisfied. For this purpose, the merge data syntax is configured as shown in Table 9 above.
[0334] If the CIIP flag is not present in the bitstream, the decoding device may derive the CIIP flag value as 1 if the general merge flag value is 1, the regular merge flag value is 0, the merge sub-block flag value is 0, the MMVD merge flag value is 0, the CIIP available flag value in the SPS is 1, the CU skip flag value is 0, the product of the width and height of the current block is 64 or greater, the width and height of the current block are each less than 128, the partitioning available flag value in the SPS is 0, or the maximum number of partitioning merge candidates is less than 2, or the slice type is not a B slice. Otherwise, the CIIP flag value is derive as 0.
[0335] In the above-described embodiments, the methods are described based on flow charts as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and certain steps may occur in a different order or simultaneously with other steps than those described. Furthermore, those skilled in the art will understand that the steps shown in the flow charts are not exclusive, and other steps may be included, or one or more steps in the flow charts may be deleted without affecting the scope of the embodiments herein.
[0336] The methods according to the embodiments of the present document described above can be implemented in software form, and the encoding device and / or decoding device according to the present document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.
[0337] In this document, when an embodiment is implemented in software, the method described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0338] In addition, the decoding device and encoding device to which the embodiment(s) of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process a video signal or a data signal. For example, over-the-top (OTT) video (over-the-top) device may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0339] In addition, a processing method to which the embodiment(s) of this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. Examples of the computer-readable recording medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission via the Internet). A bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0340] Furthermore, the embodiment(s) of this document may be embodied in a computer program product by program code, which may be executed by a computer in accordance with the embodiment(s) of this document, and which may be stored on a computer-readable carrier.
[0341] FIG. 17 illustrates an example of a content streaming system to which the embodiments disclosed herein can be applied.
[0342] As shown in FIG. 17, a content streaming system to which the embodiments of this document are applied mainly includes an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0343] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0344] The bitstream may be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0345] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0346] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0347] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head mounted displays (HMDs)), digital TVs, desktop computers, and digital signs.
[0348] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.
Claims
1. A decoding method performed by a decoding device, obtaining information about a prediction mode of a current block from a bitstream; deriving the prediction mode of the current block based on the information regarding the prediction mode; generating a predicted sample of the current block based on the prediction mode; generating reconstructed samples based on the predicted samples; the bitstream includes a sequence parameter set; the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) availability flag and a partitioning mode availability flag; the deriving step includes parsing a regular merge flag from the bitstream based on a first condition and a second condition being satisfied; the first condition is met only based on a condition related to the height of the current block and the width of the current block; The second condition is: (i) the value of the CIIP available flag is equal to 1, and the product of the height of the current block and the width of the current block is greater than or equal to 64; or (ii) a value of the partitioning mode available flag is equal to 1.
2. In an encoding method performed by an encoding device, determining a prediction mode for a current block; generating information about the prediction mode based on the prediction mode; encoding video information including the information regarding the prediction mode; the video information includes a sequence parameter set; the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) availability flag and a partitioning mode availability flag; the video information includes a regular merge flag based on whether a first condition and a second condition are satisfied; the first condition is met only based on a condition related to the height of the current block and the width of the current block; The second condition is: (i) the value of the CIIP available flag is equal to 1, and the product of the height of the current block and the width of the current block is greater than or equal to 64; or (ii) an encoding method that is satisfied based on the value of the partitioning mode available flag being equal to 1.
3. In a method for transmitting video data, obtaining a bitstream relating to the video, The bitstream comprises: determining a prediction mode for a current block; generating information about the prediction mode based on the prediction mode; encoding video information including the information about the prediction mode; transmitting the data including the bitstream; the video information includes a sequence parameter set; the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) availability flag and a partitioning mode availability flag; the video information includes a regular merge flag based on whether a first condition and a second condition are satisfied; the first condition is met only based on a condition related to the height of the current block and the width of the current block; The second condition is: (i) the value of the CIIP available flag is equal to 1, and the product of the height of the current block and the width of the current block is greater than or equal to 64; or (ii) the value of the partitioning mode available flag is equal to 1.
Citation Information
Patent Citations
JPP7279208B
Video signal processing method and device using motion compensation
WO2020149725A1