Image compiling method and device based on motion prediction
By introducing CIIP enable flags and bitstream parsing conventional merge flags during image encoding, the problems of low high-resolution image/video transmission efficiency and inter-prediction signaling redundancy are solved, and more efficient image compression and encoding are achieved.
Patent Information
- Application Number
- CN202510529252.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-19
- Filing Date
- 2020-06-19
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art is inefficient in transmitting and storing high-resolution, high-quality images/videos, and there are unnecessary signaling problems in the inter-frame prediction process.
Using combined inter-image merging and in-image prediction (CIIP) enable flags, parsing conventional merge flags through bitstreams, optimizing the image encoding process and reducing unnecessary signaling.
Improve image/video compression efficiency, effectively perform inter-frame prediction, reduce unnecessary signaling, and improve encoding efficiency.
Smart Images

Figure CN120416461A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the application number 202080058624.2 (PCT / KR2020 / 008000), the filing date of June 19, 2020, and the title of "Image Compilation Method and Apparatus Based on Motion Prediction", which was filed on February 18, 2022. Technical Field
[0002] The technology relates to a method and apparatus for compiling an image based on motion prediction. Background Art
[0003] Recently, the demand for high-resolution and high-quality images / videos such as 4K or 8K ultra-high definition (UHD) images / videos has been increasing in various fields. As the resolution or quality of images / videos becomes higher, relatively more information or bits are sent compared to traditional image / video data. Therefore, if image / video data is transmitted via a medium such as an existing wired / wireless broadband line or stored in a traditional storage medium, the costs of transmission and storage are likely to increase.
[0004] In addition, the interest and demand for virtual reality (VR) and artificial reality (AR) content, as well as immersive media such as holograms, are growing; and the broadcasting of images / videos (e.g., game images / videos) that exhibit different image / video characteristics from actual images / videos is also increasing.
[0005] Therefore, highly efficient image / video compression technologies are needed to effectively compress, transmit, store, or play high-resolution and high-quality images / videos that exhibit various characteristics as described above. Summary of the Invention
[0006] Technical Problem
[0007] The present disclosure provides a method and apparatus for improving image compilation efficiency.
[0008] The present disclosure also provides a method and apparatus for effectively performing inter-frame prediction.
[0009] The present disclosure also provides a method and apparatus for preventing unnecessary signaling during inter-frame prediction.
[0010] Technical Solution
[0011] In one aspect, a decoding method performed by a decoding device includes: obtaining information about a prediction mode of a current block from a bitstream; deriving the prediction mode of the current block based on the information about the prediction mode; generating prediction samples of the current block based on the prediction mode; and generating reconstructed samples based on the prediction samples, wherein the bitstream includes a sequence parameter set, the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, and the derivation includes parsing a regular merge flag from the bitstream based on conditions that satisfy a condition based on the CIIP enable flag and a condition based on the size of the current block.
[0012] In another aspect, an encoding method performed by an encoding device includes: determining a prediction mode of a current block; generating information about the prediction mode based on the prediction mode; and encoding image information including the information about the prediction mode, wherein the image information includes a sequence parameter set, the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, and based on conditions that satisfy a condition based on the CIIP enable flag and a condition based on the size of the current block, the image information includes a regular merge flag.
[0013] In another aspect, a computer-readable digital storage medium includes information that causes a decoding device to perform a decoding method, wherein the decoding method includes: obtaining information about a prediction mode of a current block from a bitstream; deriving the prediction mode of the current block based on the information about the prediction mode; generating prediction samples of the current block based on the prediction mode; and generating reconstructed samples based on the prediction samples, wherein the bitstream includes a sequence parameter set, the sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, and the derivation includes parsing a regular merge flag from the bitstream based on conditions that satisfy a condition based on the CIIP enable flag and a condition based on the size of the current block.
[0014] Advantageous Effects
[0015] According to embodiments of the present disclosure, the overall image / video compression efficiency can be improved.
[0016] According to embodiments of the present disclosure, inter-frame prediction can be effectively performed.
[0017] According to embodiments of the present disclosure, signaling of unnecessary syntax can be effectively removed during inter-frame prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 An example of a video / image compilation system to which embodiments of the present disclosure can be applied is schematically illustrated.
[0019] Figure 2 FIG. schematically illustrates a configuration of a video / image encoding device to which embodiments of the present disclosure can be applied.
[0020] Figure 3 FIG. is a diagram schematically illustrating a configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0021] Figure 4 FIG. shows an example of a video / image encoding method based on inter-frame prediction.
[0022] Figure 5 FIG. shows an example of a video / image decoding method based on inter-frame prediction.
[0023] Figure 6 FIG. exemplarily shows an inter-frame prediction process.
[0024] Figure 7 FIG. is a diagram illustrating spatial candidates that can be used for inter-frame prediction.
[0025] Figure 8 FIG. is a diagram illustrating a merge mode having a motion vector difference that can be used in inter-frame prediction.
[0026] Figure 9 and Figure 10 FIG. is a diagram illustrating a sub-block based temporal motion vector prediction process that can be used in inter-frame prediction.
[0027] Figure 11 FIG. is a diagram illustrating a segmentation mode applicable to inter-frame prediction.
[0028] Figure 12 FIG. is a diagram illustrating a CIIP mode applicable to inter-frame prediction.
[0029] Figure 13 and 14 FIG. schematically shows an example of a video / image encoding method including an inter-frame prediction method according to an embodiment of the present disclosure and associated components.
[0030] Figure 15 and Figure 16 FIG. schematically shows an example of a video / image decoding method including an inter-frame prediction method according to an embodiment of the present disclosure and associated components.
[0031] Figure 17 FIG. shows an example of a content streaming system to which embodiments disclosed in this document can be applied. DETAILED DESCRIPTION
[0032] The disclosure of the present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. The terms used in the present disclosure are only for describing specific embodiments and are not intended to limit the methods disclosed in the present disclosure. Expressions in the singular include expressions of "at least one" as long as they are clearly interpreted differently. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the document, and thus it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0033] In addition, each configuration of the accompanying drawings described in this document is an independent illustration for explaining the functions of features that are different from each other, and does not mean that each configuration is implemented by mutually different hardware or different software. For example, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments of combining and / or separating configurations are included within the scope of the disclosure of the present disclosure without departing from the gist of the disclosed method of the present disclosure.
[0034] Hereinafter, embodiments of this document will be described in detail with reference to the accompanying drawings. In addition, in all the drawings, the same reference numerals may be used to indicate the same elements, and the same description of the same elements will be omitted.
[0035] Figure 1 An example of a video / image compilation system to which embodiments of the present disclosure can be applied is illustrated.
[0036] Referring to Figure 1 , the video / image compilation system may include a first device (source device) and a second device (receiving device). The source device can send encoded video / image information or data in the form of a file or a stream to the receiving device through a digital storage medium or a network.
[0037] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0038] The video source can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source can include a video / image capture device and / or a video / image generation device. For example, the video / image capture device can include one or more cameras, a video / image archive including previously captured video / images, etc. For example, the video / image generation device can include a computer, a tablet computer, and a smart phone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating relevant data.
[0039] The encoding device can encode the input video / image. For compression and compilation efficiency, the encoding device can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0040] The transmitter can send the encoded image / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for transmitting through a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to the decoding device.
[0041] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0042] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0043] This document relates to video / image compilation. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the General Video Compilation (VVC) standard. In addition, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the Essential Video Compilation (EVC) standard, the AOMedia Video 1 (AV1) standard, the Second Generation Audio Video Compilation Standard (AVS2), or the next-generation video / image compilation standards (e.g., H.267, H.268, etc.).
[0044] Various embodiments related to video / image compilation are presented in this document, and the above embodiments can also be executed in combination with each other unless otherwise specified.
[0045] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients in a unified expression.
[0046] In this document, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled by the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived by the inverse transform (scaling) of the transform coefficients. The residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0047] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image at a specific time frame, and a slice / tile refers to a unit that forms part of a picture in terms of compilation. A slice / tile may include one or more Compilation Tree Units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of CTU rows within a tile in a picture. A tile may be divided into multiple bricks, and each brick may consist of one or more CTU rows within the tile. A tile that is not divided into multiple bricks may also be referred to as a brick. A tile scan may represent a specific sequential ordering of CTUs that divide a picture, where the CTUs are sequentially ordered in raster scan within a tile, the bricks within a tile are sequentially ordered in raster scan of the tile of the tile, and the tiles in a picture are sequentially ordered in raster scan of the tiles of the picture. A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential ordering of CTUs that divide a picture, where the CTUs are sequentially ordered in raster scan within a tile while the tiles in a picture are sequentially ordered in raster scan of the tiles of the picture. A slice includes an integer number of tiles of a picture that can be contained in only a single NAL unit. A slice may consist of multiple complete tiles, or may consist of only a sequential sequence of complete tiles of one tile. In this document, tile groups and slices may be used in place of each other. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0048] A pixel or pel may mean the smallest unit that constitutes a picture (or image). Additionally, the term "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0049] A unit may represent the basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or a region. In general, an M×N block may include a set (or array) of samples (or a sample array) of M columns and N rows or a set (or array) of transform coefficients. Alternatively, a sample may mean a pixel value in the spatial domain, and when such a pixel value is transformed into the frequency domain, it may mean a transform coefficient in the frequency domain.
[0050] In this document, the terms " / " and "," shall be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Additionally, "A, B" may mean "A and / or B". Additionally, "A / B / C" may mean at least one of "A, B, and / or C". Additionally, "A / B / C" may mean at least one of "A, B, and / or C".
[0051] Additionally, in the document, the term "or" shall be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document shall be interpreted as indicating "additionally or alternatively".
[0052] Figure 2 FIG. is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied. Hereinafter, a device referred to as a video encoding device may include an image encoding device.
[0053] Refer to Figure 2, the encoding device 200 includes an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB), or may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0054] The image partitioner 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and later the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding process according to the present disclosure may be performed based on the final coding unit that is no longer divisible. In this case, depending on the image characteristics and coding efficiency, etc., the largest coding unit may be used as the final coding unit, or if necessary, the coding unit may be recursively divided into coding units with a deeper depth and the coding unit with an optimal size may be used as the final coding unit. Here, the coding process may include processes of prediction, transformation, and reconstruction (which will be described later). As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be divided or partitioned from the above-mentioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0055] In some cases, a unit may be used interchangeably with terms such as a block or a region. Generally, an M×N block may represent a set of samples or transformation coefficients composed of M columns and N rows. Samples may generally represent pixels or pixel values, and may represent only the pixels / pixel values of the luminance component or only the pixels / pixel values of the chrominance component. Samples may be used as a term corresponding to pixels or pels of a picture (or image).
[0056] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoder 200 can be referred to as the subtractor 231. The predictor can perform prediction on the processing target block (hereinafter, referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction in units of the current block or CU. As will be described later in the description of each prediction mode, the predictor can generate various types of information about the prediction (e.g., prediction mode information) and send the generated information to the entropy encoder 240. The information about the prediction can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0057] The intra-frame predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the samples referred to can be located near the current block or can be separated. In intra-frame prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. For example, the non-directional modes can include the DC mode and the planar mode. For example, depending on the level of detail of the prediction direction, the directional modes can include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used according to the settings. The intra-frame predictor 222 can use the prediction mode applied to the neighboring block to determine the prediction mode applied to the current block.
[0058] The inter - frame predictor 221 can derive a prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling a motion vector difference.
[0059] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor 220 can apply intra - frame prediction or inter - frame prediction to predict a block, and can apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). In addition, the predictor can be based on the intra - block copy (IBC) prediction mode or the palette mode for predicting the block. The IBC prediction mode or the palette mode can be used for image / video compilation of content such as games, for example, screen content compilation (SCC). IBC basically performs prediction in the current picture, but it can perform similarly to inter - frame prediction in that it derives a reference block in the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in this document. The palette mode can be regarded as an example of intra - frame compilation or intra - frame prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information about the palette table and the palette index.
[0060] The prediction signal generated by the predictor (including the inter - frame predictor 221 and / or the intra - frame predictor 222) can be used to generate a reconstructed signal or to generate a residual signal.
[0061] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of the following: discrete cosine transform (DCT), discrete sine transform (DST), graph-based transform (GBT), or conditional non-linear transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT means a transform obtained from the graph. CNT means a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. Additionally, the transform processing can also be applied to a pixel block of a square with the same size, or can also be applied to a variable-size block that is not square.
[0062] The quantizer 233 quantizes the transform coefficients and sends the quantized transform coefficients to the entropy encoder 240, and the entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the quantized transform coefficients in a block form into a one-dimensional vector based on the coefficient scan order, and also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0063] The entropy encoder 240 can perform various encoding methods such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can also encode, together or separately, information necessary for video / image reconstruction other than the quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be sent or stored in the form of a bitstream in units of network abstraction layer (NAL). The video / image information can also include information about various parameter sets, such as adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). Additionally, the video / image information can also include general constraint information. In this document, the information signaled / sent from the encoding device to the decoding device and / or syntax elements can be included in the video / image information. The video / image information can be encoded through the aforementioned encoding process and thus be included in the bitstream. The bitstream can be sent through a network or can be stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A sending unit (not shown) for sending the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal can be configured as internal / external elements of the encoding device 200, or the sending unit can also be included in the entropy encoder 240.
[0064] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, inverse quantization and inverse transformation can be applied to the quantized transform coefficients by the inverse quantizer 234 and the inverse transform unit 235 to reconstruct a residual signal (residual block or residual samples). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-prediction unit 221 or the intra-prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). For example, when applying the skip mode, when there is no residual for the target block to be processed, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a restorer or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of the next target block within the current picture, and can also be used for inter-prediction of the next picture after filtering, as described below.
[0065] Meanwhile, luminance mapping and chrominance scaling (LMCS) can also be applied during picture encoding and / or reconstruction processing.
[0066] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). For example, various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering and transmit the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0067] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-prediction unit 221. When inter-prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided and the coding efficiency can be improved.
[0068] The DPB of the memory 270 can store the corrected reconstructed picture to be used as a reference picture in the inter-prediction unit 221. The memory 270 can store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be transmitted to the inter-prediction unit 221 to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit the reconstructed samples to the intra-prediction unit 222.
[0069] Figure 3This is a diagram for schematically explaining the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0070] Referring to Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0071] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the processing of the video / image information in the Figure 2 encoding device. For example, the decoding device 300 may derive units / blocks based on block partition-related information obtained from the bitstream. The decoding device 300 may perform decoding using the processing units applied in the encoding device. Thus, for example, the decoding processing unit may be a coding unit, and the coding unit may be split from a coding tree unit or a largest coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.
[0072] The decoding device 300 may receive, in the form of a bitstream, from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as the adaptation parameter set (APS), the picture parameter set (PPS), the sequence parameter set (SPS), or the video parameter set (VPS). Additionally, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described later in this document may be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information within the bitstream based on a coding method such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC), and output the syntax elements required for image reconstruction and the quantization values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive the bin corresponding to each syntax element in the bitstream, determine the context model by using the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage, and perform arithmetic decoding on the bin by predicting the probability of the occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. The information related to prediction among the information decoded by the entropy decoder 310 may be provided to the predictors (the inter-frame predictor 332 and the intra-frame predictor 331), and the residual values (i.e., the quantized transform coefficients and the related parameter information) for which the entropy decoder 310 has performed entropy decoding may be input to the residual processor 320.
[0073] The residual processor 320 may derive a residual signal (residual block, residual sample, residual sample array). Additionally, information about filtering among the information decoded by the entropy decoder 310 may be provided to the filter 350. Meanwhile, a receiver (not shown) for receiving a signal output from an encoding device may also be configured as an internal / external component of the decoding device 300, or the receiver may be a component of the entropy decoder 310. Meanwhile, the decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the following: a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0074] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 may perform dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step information) and obtain the transform coefficients.
[0075] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0076] The predictor 330 may perform prediction on a current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction to the current block and determine a specific intra-frame / inter-frame prediction mode based on the information about prediction output from the entropy decoder 310.
[0077] The predictor 330 may generate a prediction signal based on various prediction methods described below. For example, the predictor may apply intra-frame prediction or inter-frame prediction for predicting a block, and may apply intra-frame prediction and inter-frame prediction simultaneously. This may be referred to as combined inter-frame and intra-frame prediction (CIIP). Additionally, the predictor may predict a block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or the palette mode may be used for image / video compilation of content such as games, e.g., screen content compilation (SCC). IBC may basically perform prediction in the current picture, but may be performed similarly to inter-frame prediction such that a reference block is derived within the current picture. That is, IBC may use at least one of the inter-frame prediction techniques described in this document. The palette mode may be regarded as an example of intra-frame compilation or intra-frame prediction. When the palette mode is applied, information about the palette table and the palette index may be included in the video / image information and signaled.
[0078] The intra predictor 331 can predict the current block by referring to samples in the current picture. Depending on the prediction mode, the samples referred to may be located near the current block or may be separated from the current block. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to adjacent blocks.
[0079] The inter predictor 332 can derive the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the adjacent blocks can include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. For example, the inter predictor 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the inter prediction mode for the current block.
[0080] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter predictor 332 and / or the intra predictor 331). If there is no residual for the target block to be processed, for example, in the case of applying the skip mode, the prediction block can be used as the reconstructed block.
[0081] The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of the next block to be processed in the current picture, and as described later, can also be output through filtering or can be used for inter prediction of the next picture.
[0082] In addition, a luminance mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.
[0083] Filter 350 may improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 may generate a corrected reconstructed picture by applying various filtering methods to the reconstructed picture, and store the corrected reconstructed picture in memory 360, specifically, in the DPB of memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0084] The (modified) reconstructed picture stored in the DPB of memory 360 may be used as a reference picture in inter predictor 332. Memory 360 may store motion information of blocks from which motion information within the current picture is derived (decoded) and / or motion information of blocks within the already reconstructed pictures. The stored motion information may be transmitted to inter predictor 260 to be used as motion information of spatially adjacent blocks or temporally adjacent blocks. Memory 360 may store reconstructed samples of reconstructed blocks within the current picture, and transmit the reconstructed samples to intra predictor 331.
[0085] In this specification, the embodiments described in filter 260, inter predictor 221, and intra predictor 222 of encoding device 200 may be equally applied to or respectively correspond to filter 350, inter predictor 332, and intra predictor 331.
[0086] In addition, as described above, when performing video encoding, prediction is performed to enhance compression efficiency. By doing so, a prediction block including prediction samples of a current block to be encoded (i.e., an encoding target block) may be generated. In this case, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically in the encoding device and the decoding device, and the encoding device may enhance image encoding efficiency by signaling to the decoding device information (residual information) about the residual between the original block (rather than the original sample values of the original block) and the prediction block. The decoding device may derive a residual block including residual samples based on the residual information, may generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and may generate a reconstructed picture including the reconstructed block.
[0087] Residual information can be generated through a transformation process and a quantization process. For example, an encoding device can derive a residual block between an original block and a prediction block, can derive transform coefficients by performing a transformation process on residual samples (an array of residual samples) included in the residual block, can derive quantized transform coefficients by performing a quantization process on the transform coefficients, and can signal relevant residual information (through a bitstream) to a decoding device. In this case, the residual information can include value information such as quantized transform coefficients, position information, a transformation scheme, a transformation kernel, and quantization parameters. The decoding device can perform an inverse quantization / inverse transformation process based on the residual information and can derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. In addition, for inter-prediction reference of subsequent pictures, the encoding device can derive a residual block by performing inverse quantization / inverse transformation on the quantized transform coefficients and can generate a reconstructed picture based on this.
[0088] In the case where inter - frame prediction is applied to a current block, the predictor of an encoding device / decoding device can perform inter - frame prediction on a block - by - block basis and derive predicted samples. Inter - frame prediction can be prediction derived in a manner that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. In the case where inter - frame prediction is applied to the current block, a predicted block (predicted sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information of the current block can be predicted on a block, sub - block, or sample basis based on the correlation of motion information between adjacent blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include information on an inter - frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). In the case where inter - frame prediction is applied, adjacent blocks can include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent block can be equal to or different from each other. Temporally adjacent blocks can be referred to as co - located reference blocks or co - located CUs (colCUs), and the reference picture including the temporally adjacent block can be referred to as a co - located picture (colPic). For example, a motion information candidate list can be configured based on adjacent blocks of the current block, and a flag or index information indicating which candidate is selected (used) can be signaled to facilitate the derivation of the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the case of the skip mode and (normal) merge mode, the motion information of the current block can be the same as the motion information of the selected adjacent block. In the case of the skip mode, different from the merge mode, a residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and a motion vector difference can be signaled. In this case, the sum of the motion vector prediction value and the motion vector difference can be used to derive the motion vector of the current block.
[0089] The video / image encoding process based on inter - frame prediction can schematically include, for example, the following processes.
[0090] Figure 4 An example of a video / image encoding method based on inter - frame prediction is shown.
[0091] The encoding device performs inter prediction on the current block (S400). The encoding device can derive the inter prediction mode and motion information of the current block, and generate prediction samples of the current block. Here, the processes of determining the inter prediction mode, deriving the motion information, and generating the prediction samples can be executed simultaneously, or one process after another. For example, the inter predictor 221 of the encoding device may include a prediction mode determiner, a motion information deriver, and a prediction sample deriver. The prediction mode determiner can determine the prediction mode for the current block, and the motion information deriver can derive the motion information of the current block, and the prediction sample deriver can derive the prediction samples of the current block. For example, the inter predictor of the encoding device can search for a block similar to the current block in a specific region (search region) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is the smallest or less than or equal to a predetermined reference. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the position difference between the reference block and the current block. The encoding device can determine the mode applied to the current block from among various prediction modes. The encoding device can compare the rate distortion (RD) costs of various prediction modes, and determine the best prediction mode for the current block.
[0092] For example, in the case where the skip mode or the merge mode is applied to the current block, the encoding device can construct a merge candidate list described later, and derive a reference block whose difference from the current block is the smallest or less than or equal to a predetermined reference among the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, the merge candidate associated with the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.
[0093] As another example, in the case where the (A)MVP mode is applied to the current block, the encoding device can construct an (A)MVP candidate list described later, and use the motion vector of the MVP candidate selected from among the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, the motion vector indicating the reference block derived through the above motion estimation can be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates can become the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, can be derived. In this case, information about the MVD can be signaled to the decoding device. In addition, in the case where the (A)MVP mode is applied, the reference image index value can be configured as reference image index information, and can be signaled to the decoding device separately.
[0094] The encoding device may derive a residual sample based on a prediction sample (S410). The encoding device may derive a residual sample by comparing the original sample of the current block with the prediction sample.
[0095] The encoding device encodes image information including prediction information and residual information. The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include information about prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information as information related to the prediction process. The information about the motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) as information for deriving a motion vector. In addition, the information about the motion information may include the above-mentioned MVD information and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual sample. The residual information may include information about transform coefficients for quantization of the residual sample.
[0096] The output bitstream may be stored in a (digital) storage medium and sent to the decoding device or may be sent to the decoding device via a network.
[0097] Meanwhile, as described above, the encoding device may generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference sample and the residual sample. This is because the encoding device can derive the same prediction result as the prediction result executed by the decoding device, and by doing so, the encoding efficiency can be improved. Therefore, the encoding device may store the reconstructed picture (or reconstructed samples or reconstructed blocks) in the memory and use the reconstructed picture as a reference picture for inter-frame prediction. As described above, the in-loop filtering process may also be applied to the reconstructed picture.
[0098] The video / image decoding process based on inter-frame prediction may schematically include, for example, the following processes.
[0099] Figure 5 An example of a video / image decoding method based on inter-frame prediction is shown.
[0100] Referring to Figure 5 , the decoding device may perform operations corresponding to the operations performed by the encoding device. The decoding device may perform prediction on the current block based on the received prediction information and derive a prediction sample.
[0101] Specifically, the decoding device may derive a prediction mode for the current block based on the received prediction information (S500). The decoding device may determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.
[0102] For example, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the (A)MVP mode is determined. Alternatively, one of various inter prediction mode candidates can be selected based on a mode index. The inter prediction mode candidates can include the skip mode, the merge mode, and / or the (A)MVP mode, or can include various inter prediction modes described later.
[0103] The decoding device derives the motion information of the current block based on the determined inter prediction mode (S510). For example, in a case where the skip mode or the merge mode is applied to the current block, the decoding device can construct a merge candidate list described later and select one merge candidate from the merge candidates included in the merge candidate list. This selection can be performed based on the above-described selection information (merge index). The motion information of the selected merge candidate can be used to derive the motion information of the current block. The motion information of the selected merge candidate can be used as the motion information of the current block.
[0104] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list described later and use the motion vector of the MVP candidate selected from the motion vector predictor (MVP) candidates included in the (A)MVP candidate list as the MVP of the current block. This selection can be performed based on the above-described selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on the information about the MVD, and the motion vector of the current block can be derived based on the MVP and the MVD of the current block. In addition, the reference picture index of the current block can be derived based on the reference picture index information. The picture indicated by the reference picture index in the reference picture list of the current block can be derived as the reference picture that is referred to for the inter prediction of the current block.
[0105] Meanwhile, as described later, the motion information of the current block can be derived without configuring a candidate list. In this case, the motion information of the current block can be derived according to the process disclosed in the prediction mode described later. In this case, the configuration of the candidate list as described above can be omitted.
[0106] The decoding device can generate prediction samples for the current block based on the motion information of the current block (S520). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction samples of the current block can be derived using the samples of the reference block indicated by the motion vector of the current block on the reference picture. In this case, as described later, the prediction sample filtering process can be further performed on all or part of the prediction samples of the current block in some cases.
[0107] For example, an inter-prediction unit of a decoding device may include a prediction mode determiner, a motion information derivator, and a predicted sample derivator. The prediction mode determiner may determine a prediction mode for a current block based on received prediction mode information. The motion information derivator may derive motion information (such as a motion vector and / or a reference picture index, etc.) of the current block based on information about the received motion information. And the predicted sample derivator may derive predicted samples of the current block.
[0108] The decoding device generates residual samples for a current block based on received residual information (S530). The decoding device may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based thereon (S540). Thereafter, as described above, an in-loop filtering process may be further applied to the reconstructed picture.
[0109] Figure 6 An inter-prediction process is exemplarily illustrated
[0110] Referring to Figure 6 , as described above, the inter-prediction process may include determining an inter-prediction mode, deriving motion information according to the determined prediction mode, and performing prediction (generating predicted samples) based on the derived motion information. As described above, the inter-prediction process may be performed by an encoding device and a decoding device. In this document, a compiling device may include an encoding device and / or a decoding device.
[0111] Referring to Figure 6 , the compiling device determines an inter-prediction mode for a current block (S600). Various inter-prediction modes may be used to predict a current block within a picture. For example, various modes such as a merge mode, a skip mode, a motion vector prediction (MVP) mode, an affine mode, a sub-block merge mode, a merge with MVD (MMVD) mode, and a history motion vector prediction (HMVP) mode, etc. may be used. A decoder-side motion vector refinement (DMVR) mode, an adaptive motion vector resolution (AMVR) mode, a bi-prediction with CU-level weights (BCW), and a bi-directional optical flow (BDOF) may be further used or alternatively used as incidental modes. The affine mode may be referred to as an affine motion prediction mode. The MVP mode may be referred to as an advanced motion vector prediction (AMVP) mode. In this document, some modes and / or candidates of motion information derived by some modes may be included as one of candidates related to motion information of other modes. For example, an HMVP candidate may be added as a merge candidate for the merge / skip mode, or may be added as an MVP candidate for the MVP mode.
[0112] Prediction mode information indicating an inter - frame prediction mode of a current block can be signaled from an encoding device to a decoding device. The prediction mode information can be included in a bitstream and received at the decoding device. The prediction mode information can include index information indicating one of a plurality of candidate modes. Additionally, the inter - frame prediction mode can be indicated by hierarchical signaling of flag information. In this case, the prediction mode information can include one or more flags. For example, a skip flag can be signaled to indicate whether a skip mode is applied; for a non - applied skip mode, a merge flag can be signaled to indicate whether a merge mode is applied; and when the merge mode is not applied, an MVP mode can be indicated or a flag for further splitting can also be signaled. The affine mode can be signaled as an independent mode or as a mode depending on the merge mode, MVP mode, etc. For example, the affine mode can include an affine merge mode and an affine MVP mode.
[0113] Meanwhile, information indicating whether list 0 (L0) prediction, list 1 (L1) prediction, or bi - prediction is used for the current block (current coding unit) is signaled to the current block. This information can be referred to as motion prediction direction information, inter - frame prediction direction information, or inter - frame prediction indication information, and can be constructed / encoded / signaled in the form of, for example, an inter_pred_idc syntax element. That is, the inter_pred_idc syntax element can indicate whether list 0 (L0) prediction, list 1 (L1) prediction, or bi - prediction is used for the current block (current coding unit). In this document, for ease of explanation, the inter - frame prediction type (L0 prediction, L1 prediction, or BI prediction) indicated by the inter_pred_idc syntax element can be represented as a motion prediction direction. L0 prediction can be represented by pred_L0; L1 prediction can be represented by pred_L1; and bi - prediction can be represented by pred_BI. For example, the following prediction types can be indicated according to the value of the inter_pred_idc syntax element.
[0114] As described above, a picture can include one or more slices. A slice can have one of slice types including an intra (I) slice, a predictive (P) slice, and a bi - predictive (B) slice. The slice type can be indicated based on slice type information. For blocks in an I slice, inter - frame prediction is not used for prediction, and only intra - frame prediction can be used. Of course, even in this case, the original sample values can be coded and signaled without prediction. For blocks in a P slice, intra - frame prediction or inter - frame prediction can be used, and in the case of using inter - frame prediction, only uni - directional prediction can be used. Meanwhile, for blocks in a B slice, intra - frame prediction or inter - frame prediction can be used, and in the case of using inter - frame prediction, up to maximum bi - prediction can be used.
[0115] L0 and L1 may include reference pictures that have been encoded / decoded before the current picture. For example, L0 may include reference pictures before and / or after the current picture in POC order, and L1 may include reference pictures after and / or before the current picture in POC order. In this case, a lower reference picture index relative to the reference pictures earlier than the current picture in POC order may be assigned to L0, and a lower reference picture index relative to the reference pictures later than the current picture in POC order may be assigned to L1. In the case of B slices, bidirectional prediction may be applied, and in this case, unidirectional bidirectional prediction or bidirectional bidirectional prediction may be applied. Bidirectional bidirectional prediction may be referred to as true bidirectional prediction.
[0116] Specifically, for example, information about the inter prediction mode of the current block can be compiled and signaled at the CU (CU syntax) level or the like, or can be implicitly determined according to conditions. In this case, some modes can be signaled explicitly while other modes can be implicitly derived.
[0117] For example, the CU syntax may carry information about the (inter) prediction modes in Table 1 below.
[0118] [Table 1 ]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] Here, cu_skip_flag can indicate whether the skip mode is applied to the current block (CU).
[0133] pred_mode_flag being equal to 0 specifies that the current coding unit is coded in an inter prediction mode. pred_mode_flag being equal to 1 specifies that the current coding unit is coded in an intra prediction mode.
[0134] pred_mode_ibc_flag being equal to 1 specifies that the current coding unit is coded in an IBC prediction mode. pred_mode_ibc_flag being equal to 0 specifies that the current coding unit is not coded in an IBC prediction mode.
[0135] pcm_flag [x0][y0] being equal to 1 specifies that there is a pcm_sample() syntax structure and no transform_tree() syntax structure in the coding unit including the luma coding block at position (x0, y0). pcm_flag [x0][y0] being equal to 0 specifies that there is no pcm_sample() syntax structure. That is, pcm_flag can indicate whether the Pulse Code Modulation (PCM) mode is applied to the current block. If the PCM mode is applied to the current block, prediction, transformation, quantization, etc. may not be applied, and the original sample values in the current block can be coded and signaled.
[0136] intra_mip_flag[x0][y0] being equal to 1 specifies that the intra prediction type of the luma samples is matrix-based intra prediction (MIP). intra_mip_flag[x0][y0] being equal to 0 specifies that the intra prediction type of the luma samples is not matrix-based intra prediction. That is, intra_mip_flag can indicate whether the MIP prediction mode (type) is applied to the current block (for the luma samples).
[0137] intra_chroma_pred_mode[x0][y0] specifies the intra prediction mode of the chroma samples in the current block.
[0138] The general_merge_flag[x0][y0] specifies whether to infer the inter-prediction parameters of the current compilation unit from the inter-frame prediction partitions of adjacent frames. That is, the general_merge_flag can indicate that general merging is available, and when the general_merge_flag value is 1, the regular merge mode, the MMVD mode, and the merged sub-block mode (sub-block merge mode) can be available. For example, when the general_merge_flag value is 1, the merge data syntax can be parsed from the encoded video / image information (or bitstream), and the merge data syntax is configured / compiled to include the information shown in Table 2 below.
[0139] [Table 2]
[0140]
[0141]
[0142]
[0143] Here, the regular_merge_flag[x0][y0] being equal to 1 specifies that the regular merge mode is used to generate the inter-prediction parameters of the current compilation unit. That is, the regular_merge_flag indicates whether the merge mode (regular merge mode) is applied to the current block.
[0144] The mmvd_merge_flag[x0][y0] being equal to 1 specifies that the merge mode with motion vector difference is used to generate the inter-prediction parameters of the current compilation unit. That is, the mmvd_merge_flag indicates whether MMVD is applied to the current block.
[0145] The mmvd_cand_flag [x0][y0] specifies whether the first (0) or second (1) candidate in the merge candidate list is used together with the motion vector difference derived from mmvd_distance_idx [x0][y0] and mmvd_direction_idx[x0][y0].
[0146] The mmvd_distance_idx [x0][y0] specifies the index used to derive MmvdDistance [x0][y0].
[0147] The mmvd_direction_idx [x0][y0] specifies the index used to derive MmvdSign[x0][y0].
[0148] merge_sub_flag[x0][y0] specifies the sub-block based inter prediction parameters for the current compilation. That is, merge_sub_flag can indicate whether the sub-block merge mode (or affine merge mode) is applied to the current block.
[0149] merge_subblock_idx[x0][y0] specifies the merge candidate index of the sub-block based merge candidate list.
[0150] ciip_flag[x0][y0] specifies whether combined inter picture merge and intra prediction are applied to the current compilation unit.
[0151] merge_triagle_idx0[x0][y0] specifies the first merge candidate index of the triangle shape based motion compensation candidate list.
[0152] merge _ triagle _ idx1[x0][y0] specifies the second merge candidate index of the triangle shape based motion compensation candidate list.
[0153] merge_idx [x0][y0] specifies the merge candidate index of the merge candidate list.
[0154] Meanwhile, returning to the CU syntax in Table 1 for reference, mvp_l0_flag [x0][y0] specifies the motion vector prediction index for list 0. That is, when the MVP mode is applied, mvp_l0_flag can indicate the candidate selected from the MVP candidate list 0 for the MVP derivation for the current block.
[0155] ref_idx_l1[x0][y0] has the same semantics as ref_idx_l0, where l0 and list 0 can be replaced by l1 and list 1 respectively. (ref_idx_l1[x0][y0] has the same semantics as ref_idx_L0, where l0, L0 and list 0 are replaced by l1, L1 and list 1 respectively).
[0156] inter_pred_idc [x0][y0] specifies whether list 0, list 1 or bi-prediction is used for the current compilation unit.
[0157] sym_mvd_flag [x0][y0] being equal to 1 specifies that the syntax elements ref_idx_l0[x0][y0] and ref_idx_l1[x0][y0] and the mvd_coding (x0, y0, refList, cpdix) syntax structure for refList being equal to 1 do not exist. That is, sym_mvd_flag indicates whether symmetric MVD is used for mvd compilation.
[0158] ref_idx_l0[x0][y0] specifies the list 0 reference picture index for the current coding unit.
[0159] ref_idx_l1[x0][y0] has the same semantics as ref_idx_l0, where l0, L0, and list 0 are replaced by l1, L1, and list 1, respectively.
[0160] inter_affif_flag[x0][y0] being equal to 1 specifies that, for the current coding unit, when decoding a P or B slice, motion compensation based on an affine model is used to generate the prediction samples for the current coding unit.
[0161] cu_affif_type_flag[x0][y0] being equal to 1 specifies that, for the current coding unit, when decoding a P or B slice, motion compensation based on a 6-parameter affine model is used to generate the prediction samples for the current coding unit. cu_affif_type_flag[x0][y0] being equal to 0 specifies that motion compensation based on a 4-parameter affine model is used to generate the prediction samples for the current coding unit.
[0162] amvr_flag [x0][y0] specifies the resolution of the motion vector difference. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. amvr_ flag [x0][y0] being equal to 0 specifies that the resolution of the motion vector difference is 1 / 4 of a luma sample. amvr_ flag [x0][y0] being equal to 1 specifies that the resolution of the motion vector difference is further specified by amvr_ precision _ flag [x0][y0].
[0163] If inter_affice_flag [x0][y0] is equal to 0, amvr_precision_flag[x0][y0] being equal to 0 specifies that the resolution of the motion vector difference is an integer luma sample, otherwise 1 / 16 of a luma sample. If inter_affice_flag[x0][y0] is equal to 0, amvr_precision_flag [x0][y0] being equal to 1 specifies that the resolution of the motion vector difference is four luma samples, otherwise an integer luma sample. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0164] bcw_idx [x0][y0] specifies the weight index for bi-prediction with CU weights.
[0165] When determining a (inter - frame) prediction mode for a current block, an encoding device derives motion information for the current block based on the prediction mode (S610).
[0166] The encoding device can perform inter - frame prediction using the motion information of the current block. The encoding device can derive the optimal motion information of the current block through a motion estimation process. For example, the encoding device can search for a similar reference block with high correlation within a search range determined in terms of fractional pixels for an original block in the original image of the current block in the reference image, and derive the motion information therefrom. The block similarity can be derived based on the difference between the sample values based on the phase. For example, the block similarity can be calculated based on the SAD between the current block (or a template of the current block) and the reference block (or a template of the reference block). In this case, the motion information can be derived based on the reference block with the minimum SAD in the search area. The derived motion information can be signaled to the decoding device according to various methods based on the inter - frame prediction mode.
[0167] When deriving the motion information about the current block, the encoding device performs inter - frame prediction based on the motion information about the current block (S620). The encoding device can derive the predicted samples of the current block based on the motion information. The current block including the predicted samples can be referred to as a prediction block.
[0168] The reconstructed samples and the reconstructed picture can be generated based on the derived predicted samples, and thereafter, processes such as in - loop filtering can be performed.
[0169] Figure 7 FIG. is a diagram illustrating the merge mode and the skip mode that can be used for inter - frame prediction.
[0170] In the case where the merge mode is applied during inter - frame prediction, the motion information of the current block is not directly sent, and the motion information of the current block is derived using the motion information of adjacent prediction blocks. Therefore, the encoding device can indicate the motion information of the current block by sending flag information indicating the use of the merge mode and a merge index indicating which adjacent prediction block to use. The merge mode can be referred to as a regular merge mode.
[0171] To perform the merge mode, the encoding device searches for merge candidate blocks for deriving the motion information of the current block. For example, up to 5 merge candidate blocks can be used, but this embodiment is not limited thereto. In addition, information about the maximum number of merge candidate blocks can be sent in the slice header or the tile group header, but this embodiment is not limited thereto. After finding the merge candidate blocks, the encoding device can generate a merge candidate list, and can select the merge candidate block with the minimum cost among the merge candidate blocks as the final merge candidate block.
[0172] This document provides various embodiments of merge candidate blocks for constructing a merge candidate list.
[0173] The merge candidate list may include, for example, 5 merge candidate blocks. For example, 4 spatial merge candidates and one temporal merge candidate may be used. As a specific example, in the case of spatial merge candidates, Figure 7 the blocks A0, A1, B0, B1, and B2 shown in may be used as spatial merge candidates. Hereinafter, the spatial merge candidates or spatial MVP candidates described later may be referred to as SMVPs, and the temporal merge candidates or temporal MVP candidates described later may be referred to as TMVPs
[0174] For example, the merge candidate list for the current block may be constructed based on the following process.
[0175] First, the compiling device (encoding device / decoding device) may insert the spatial merge candidates derived by searching the spatial neighboring blocks of the current block into the merge candidate list. For example, the spatial neighboring blocks may include the bottom-left neighboring block A0, left neighboring block A1, top-right neighboring block B0, top neighboring block B1, and top-left neighboring block B2 of the current block. However, this is an example, and in addition to the above spatial neighboring blocks, additional neighboring blocks such as right neighboring blocks, bottom neighboring blocks, and bottom-right neighboring blocks may further be used as spatial neighboring blocks. The compiling device may detect available blocks by searching the spatial neighboring blocks based on the priority, and may derive the motion information of the detected blocks as spatial merge candidates. For example, the encoding device and / or the decoding device may search the 5 blocks shown in in the order of A1, B1, B0, A0, and B2, Figure 7 and sequentially index the available candidates to form the merge candidate list.
[0176] In addition, the compiling device may insert the temporal merge candidates derived by searching the temporal neighboring blocks of the current block into the merge candidate list. The temporal neighboring blocks may be placed on a reference picture, which is a picture different from the current picture on which the current block is placed. The reference picture on which the temporal neighboring blocks are placed may be referred to as the collocated picture or col picture. The temporal neighboring blocks may be searched in the order of the bottom-right neighboring block and the bottom-right center block of the collocated block of the current block on the col picture.
[0177] Meanwhile, the compiling device may check whether the number of current merge candidates is less than the number of maximum merge candidates. The number of maximum merge candidates may be predefined or may be signaled from the encoding device to the decoding device. For example, the encoding device may generate information about the number of maximum merge candidates, encode the information, and send the encoded information to the decoding device in the form of a bitstream. When the number of maximum merge candidates is filled, the subsequent candidate addition process may not be performed.
[0178] When the number of current merge candidates is less than the maximum number of merge candidates as a result of the check, the compiling device may insert additional merge candidates into the merge candidate list. For example, the additional merge candidates may include history-based merge candidates, pairwise average merge candidates, ATMVP, combined dual-prediction merge candidates (when the slice / tile group type of the current slice / tile group is of type B), and / or zero vector merge candidates.
[0179] If, as a result of the check, the number of current merge candidates is not less than the maximum number of merge candidates, the compiling device may end the construction of the merge candidate list. In such a case, the encoding device may select the best merge candidate from the merge candidates constituting the merge candidate list based on rate-distortion (RD) cost, and may signal the selection information (e.g., merge index) indicating the selected merge candidate to the decoding device. The decoding device may select the best merge candidate based on the merge candidate list and the selection information.
[0180] As described above, the motion information of the selected merge candidate may be used as the motion information of the current block, and the predicted samples of the current block may be derived based on the motion information of the current block. The encoding device may derive the residual samples of the current block based on the predicted samples, and may signal the residual information regarding the residual samples to the decoding device. As described above, the decoding device may generate reconstructed samples based on the residual samples and the predicted samples derived from the residual information, and generate a reconstructed picture based thereon.
[0181] When the skip mode is applied during inter prediction, the motion information of the current block can be derived in the same manner as when the merge mode is applied as described above. However, when the skip mode is applied, the residual signal of the corresponding block is omitted, so the predicted samples can be directly used as the reconstructed samples.
[0182] Figure 8 FIG. is a diagram illustrating a merge mode having a motion vector difference that can be used in inter prediction.
[0183] In addition to the merge mode in which implicitly derived motion information is directly used for the generation of predicted samples of the current CU, a merge mode with motion vector difference (MMVD) is introduced in VVC. Since a similar motion information derivation method is used for the skip mode and the merge mode, MMVD can be applied to the skip mode. The MMVD flag (e.g., MMVD_flag) may be signaled immediately after the skip flag and the merge flag are sent to specify whether the MMVD mode is used for the CU.
[0184] In MMVD, after selecting merge candidates, the merge candidates are further refined with signaled MVD information. When MMVD is applied to the current block (i.e., when the MMVD_flag is equal to 1), further information of MMVD can be signaled.
[0185] The further information includes a merge candidate flag (e.g., mmvd_merge_flag), an index specifying the motion amplitude (e.g., mmvd_distance_idx), and an index for indicating the motion direction (e.g., mmvd_direction_idx). The merge candidate flag indicates whether the first candidate (0) or the second candidate (1) in the merge candidate list is used together with the motion vector difference. In the MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. The merge candidate flag is signaled to specify which one is used.
[0186] The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point.
[0187] As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 3.
[0188] [Table 3 ]
[0189]
[0190] Here, slice_fpel_mmvd_enabled_flag being equal to 1 specifies that the merge mode with motion vector difference uses the integer sample precision in the current slice. slice_fpel_mmvd_enabled_flag being equal to 0 specifies that the merge mode with motion vector difference can use the fractional sample precision in the current slice. When absent, the value of slice_fpel_mmvd_enabled_flag is inferred as 0. The slice_fpel_mmvd_enabled_flag syntax element can be signaled (can be included) by the slice header.
[0191] The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown in Table 4. Note that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a non-predictive MV or a bi-predictive MV where both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 4 specify the sign of the MV offset added to the starting MV. When the starting MV is a non-predictive MV or a bi-predictive MV where both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 4 specify the sign of the MV offset added to the starting MV.
[0192] [Table 4 ]
[0193]
[0194] The two merged components plus the MVD offset MmvdOffset[x0][y0] are derived as follows.
[0195] [Equation 1]
[0196]
[0197] Figure 9 and Figure 10 is a diagram explaining the sub-block based temporal motion vector prediction process that can be used during inter prediction.
[0198] The sub-block based temporal motion vector prediction (SbTMVP) method can be used for inter prediction. Similar to the temporal motion vector prediction (TMVP), SbTMVP uses the motion field in the collocated picture to improve the motion vector prediction and merge mode of the CUs in the current picture. The same collocated image used by TMVP is used for SbTVMP. The differences between SbTMVP and TMVP are in the following two main aspects.
[0199] 1. TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level.
[0200] 2. TMVP extracts the temporal motion vector from the collocated block in the collocated picture (the collocated block is the bottom-right or center (bottom-right center) block relative to the current CU), while SbTMVP applies a motion shift before extracting the temporal motion information from the collocated picture, where the motion shift is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0201] Figure 9 and Figure 10Shows the SbTVMP process. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 9 is checked. If A1 has a motion vector using a collocated picture as its reference picture identified, then that motion vector (which can be called the temporal MV (tempVM)) is selected as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0, 0).
[0202] In the second step, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vectors and reference indices) from the collocated picture as shown in Figure 10 . Figure 10 The example in assumes that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample) in the collocated picture is used to derive the motion information for the sub-CU. When the sub-blocks have uniform length, width, and height, the central sample (the lower-right central sample) can correspond to the lower-right sample among the 4 central samples in the sub-CU.
[0203] After identifying the motion information of the collocated sub-CU, it is converted to the motion vectors and reference indices of the current sub-CU in a similar manner to the TMVP process, where temporal motion scaling can be applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0204] A sub-block based merge list containing both SbTVMP candidates and affine merge candidates can be used for signaling in the affine merge mode (which can be called the (sub-block based) merge mode). The SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the sub-block merge candidate list, followed by the affine merge candidates. The maximum allowed size of the affine merge candidate list can be 5.
[0205] The sub-CU size used in SbTMVP can be fixed to 8×8, and similar to what is done for the affine merge mode, the SbTMVP mode can only be applied to CUs where both the width and height are greater than or equal to 8.
[0206] The encoding logic for the additional SbTMVP merge candidates is the same as that for the other merge candidates, that is, for each CU in a P or B slice, an additional RD check can be performed to decide whether to use the SbTMVP candidates.
[0207] Figure 11 is a diagram illustrating the partition patterns that can be applied to inter prediction.
[0208] The triangular partitioning mode can be used for inter prediction. The triangular partitioning mode can be applied only to CUs of 8x8 or larger. In cases where other merge modes include the regular merge mode, the MMVD mode, the CIIP mode, and the sub-block merge mode, a CU-level flag is used as a merge mode to signal the triangular partitioning mode.
[0209] When this mode is used, a CU can be evenly divided into two triangular partitions using diagonal partitioning or anti-diagonal partitioning, as Figure 11 shown. Each triangular partition in the CU is inter predicted using its own motion; only single prediction is allowed for each partition, i.e., each partition has one motion vector and one reference index. A single prediction motion constraint is applied to ensure that, as with regular dual prediction, only two motion compensated predictions are required for each CU.
[0210] If the triangular partitioning mode is used for the current CU, a flag indicating the direction (diagonal or anti-diagonal) of the triangular partition and two merge indices (one merge index for each partition) are further signaled. The number of maximum TPM candidate sizes is explicitly signaled at the slice level, and the syntax for specifying the TMP merge index is binaryized. After each of the triangular partitions is predicted, a hybrid process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edge. This is the prediction signal for the entire CU, and the transform and quantization processes will be applied to the entire CU, as in other prediction modes. Finally, the motion field of the CU predicted using the triangular partitioning mode is stored in 4x4 units. The triangular partitioning mode is not combined with SBT, that is, when the signaled triangular mode is equal to 1, cu_SBT_flag is inferred to be 0 without signaling.
[0211] The uni-prediction candidate list is directly derived from the merge candidate list constructed as described above.
[0212] After each triangular partition is predicted using its own motion, a hybrid is applied to the two prediction signals to derive the samples around the diagonal or anti-diagonal edge.
[0213] Figure 12 is a diagram illustrating the CIIP mode applicable to inter prediction.
[0214] Combined inter - and intra - prediction can be applied to the current block. An additional flag (e.g., CIIP_flag) can be signaled to indicate whether the combined inter / intra - prediction (CIIP) mode is applied to the current CU. For example, when compiling a CU in merge mode, if the CU includes at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, then the additional flag is signaled to indicate whether the combined inter / intra - prediction (CIIP) mode is applied to the current CU. As its name indicates, CIIP prediction combines the inter - prediction signal with the intra - prediction signal. The same inter - prediction process applied to the regular merge mode is used to derive the inter - prediction signal in the CIIP mode P_inter; and the intra - prediction signal P_intra is derived using the planar mode after the regular intra - prediction process. Then, a weighted average is used to combine the intra - and inter - prediction signals, where the weight values are calculated as follows according to the Figure 12 depicted in the) top - adjacent block and left - adjacent block's coding modes.
[0215] If the top neighbor is available and is intra - coded, set isIntrantop to 1, otherwise set isIntrantop to 0.
[0216] If the left neighbor is available and is intra - coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0.
[0217] If (isIntraLeft + isIntraLeft) equals 2, set wt to 3.
[0218] Otherwise, if (isIntraLeft + isIntraLeft) equals 1, set wt to 2.
[0219] Otherwise, if (isIntraLeft + isIntraLeft) equals 1, set wt to 2.
[0220] The CIIP prediction is formed as follows.
[0221] [Equation 2]
[0222]
[0223] Meanwhile, to generate prediction blocks, the encoding device may derive motion information based on the conventional merge mode, skip mode, SbTMVP mode, MMVD mode, triangle partitioning mode (partitioning mode), and / or CIIP mode as described above. Each mode may be enabled / disabled by an on / off flag for each mode included in the sequence parameter set (SPS). If the on / off flag for a specific mode is disabled, the encoding device does not signal the syntax explicitly sent for the corresponding prediction mode in units of CU or PU.
[0224] Therefore, when all specific modes of the merge / skip mode are disabled or partially disabled in the existing operation processing, there is a problem that the on / off flag is redundantly signaled. Therefore, in this document, to prevent the same information (flag) from being redundantly signaled in the process of selecting the merge mode applied to the current block based on the merge data syntax in Table 2, any of the following methods may be used.
[0225] The encoding device may signal the flag based on the sequence parameter set shown in Table 5 below, so as to select the prediction mode that can be used in the process of deriving motion information. Each prediction mode may be turned on / off based on the sequence parameter set in Table 5, and each syntax element of the merge data syntax in Table 2 may be parsed or induced according to the flag in Table 5 and the conditions of each mode.
[0226] [Table 5]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237] The following drawings are prepared to explain specific examples of the present document. Since the names of specific devices and signals / information described in the drawings are presented in an example manner, the technical features of the present disclosure are not limited to the specific names used in the following drawings.
[0238] Figure 13 and Figure 14 Schematically shows an example of a video / image encoding method including an inter-frame prediction method according to an embodiment of the present disclosure and associated components.
[0239] Figure 13 The encoding method disclosed in Figure 2 can be executed by the encoding device 200 disclosed in Figure 13 Specifically, for example, Figure 13 S1300 to S1310 of
[0240] can be executed by the predictor 220 of the encoding device 200, and S1320 can be executed by the entropy encoder 240 of the encoding device 200. Figure 13 and Figure 14 Specifically, referring to
[0241] the predictor of the encoding device can determine the prediction mode of the current block (S1300). As an example, in the case where inter-frame prediction is applied to the current block, the predictor of the encoding device can determine any one of the regular merge mode, skip mode, MMVD mode, sub-block merge mode, partition mode, and CIIP mode as the prediction mode of the current block.
[0242] Here, the regular merge mode can be defined as a mode that uses the motion information of adjacent blocks to derive the motion information of the current block. The skip mode can be defined as a mode in which the predicted block is used as the reconstructed block. The MMVD mode can be applied to the merge mode or the skip mode, and can be defined as a merge (or skip) mode that uses the motion vector difference. The sub-block merge mode can be defined as a merge mode based on sub-blocks. The partition mode can be defined as a mode that performs prediction by dividing the current block into two partitions (diagonal or anti-diagonal). The CIIP mode can be defined as a mode in which inter-picture merge and intra-picture prediction are combined with each other.
[0243] The predictor of the encoding device may generate a prediction sample (prediction block) of the current block based on the prediction mode of the current block and the motion vector of the current block. In addition, the predictor of the encoding device may generate information about the prediction mode based on the prediction mode (S1310). Here, the information about the prediction mode may include inter / intra prediction classification information, inter prediction mode information, etc., and may include various syntax elements related thereto.
[0244] The residual processor of the encoding device may generate a residual sample based on the original sample (original block) of the current block and the prediction sample (prediction block) of the current block. Additionally, information about the residual sample may be derived based on the residual sample.
[0245] The encoder of the encoding device may encode image information, which includes information about the residual sample, information about the prediction mode, etc. (S1320). The image information may include partition-related information, information about the prediction mode, residual information, in-loop filtering-related information, etc., and may include various syntax elements related thereto. The information encoded by the encoder of the encoding device can be output in the form of a bitstream. The bitstream can be sent to the decoding device via a network or a storage medium.
[0246] For example, the image information may include information about various parameter sets, such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). In addition, the image information may include information about the prediction mode of the current block, such as compile unit syntax and merge data syntax. Here, the Sequence Parameter Set may include a Combined Inter and Intra Prediction (CIIP) enable flag, an enable flag for the partition mode, etc. The compile unit syntax may include a CU skip flag indicating whether the skip mode is applied to the current block.
[0247] According to an embodiment, as an example, the encoding device may include a regular merge flag in the image information based on conditions that satisfy a condition based on the CIIP enable flag and a condition based on the size of the current block, so as not to repeatedly transmit the same syntax. Here, the condition based on the size of the current block may be a case where the product of the height and the width of the current block is 64 or greater and the height and the width of the current block are each less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and the width of the current block is 64 or greater, the height and the width of the current block are each less than 128, and the value of the CIIP enable flag is 1, the encoding device may subsequently signal the regular merge flag.
[0248] As another example, the encoding device may include a regular merge flag in the image information based on a condition that satisfies a condition based on the CU skip flag and the size of the current block. Here, the condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the product of the height and width of the current block is less than 128, and the value of the CU skip flag is 0, the encoding device may subsequently signal the regular merge flag.
[0249] As another example, in addition to the condition based on the CIIP enable flag and the condition based on the size of the current block, based on further satisfying the condition based on the CU skip flag, the encoding device may include a regular merge flag in the image information. Here, the condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the product of the height and width of the current block is less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the encoding device may subsequently signal the regular merge flag.
[0250] As another example, the encoding device may include a regular merge flag in the image information based on a condition that satisfies a condition based on information about the current block and the partition mode enable flag. Here, the condition based on information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or include a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block and information about the current block and the condition based on the partition mode enable flag are satisfied, the encoding device may signal the regular merge flag.
[0251] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the encoding device may determine whether the conditions based on information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on information about the current block and the partition mode enable flag are not satisfied, the encoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied.
[0252] Meanwhile, when the product of the width and height of the current block is not 32 and the value of the MMVD enable flag is 1, or when the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are equal to or greater than 8, the encoding device may signal the regular merge flag.
[0253] To this end, as an example, the merge data syntax may be configured as shown in Table 6 below.
[0254] [Table 6]
[0255]
[0256]
[0257]
[0258] In Table 6, regular_merge_flag[x0][y0] being equal to 1 specifies that the regular merge mode is used to generate the inter-prediction parameters of the current coding unit. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0259] When regular_merge_flag[x0][y0] does not exist, the following inference is made.
[0260] If all of the following conditions are true, regular_merge_flag[x0][y0] is inferred to be equal to 1.
[0261] - regular_merge_flag[x0][y0] is equal to 1
[0262] - sps_mmvd_enable_flag is equal to 0 or (cbWidth * cbHeight == 32)
[0263] - MaxNumSubblockMergeCand <= 0 or cbWidth < 8 or cbHeight < 8
[0264] - sps_ciip_enabled_flag is equal to 0 or cbWidth * cbHeight < 64 or cbWidth >= 128 or cu_skip_flag[x0][y0] is equal to 1
[0265] - sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand < 2 or slice_type is not equal to B_SLICE
[0266] Otherwise, regular_merge_flag[x0][y0] is inferred to be equal to 0.
[0267] Meanwhile, according to another embodiment, as an example, the encoding device may include an MMVD merge flag in the image information based on satisfying a condition based on the CIIP enable flag and a condition based on the size of the current block, so as not to repeatedly transmit the same syntax. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are less than 128, and the value of the CIIP enable flag is 1, the encoding device may signal the MMVD merge flag.
[0268] As another example, in addition to the condition based on the CIIP enable flag and the condition based on the size of the current block, based on further satisfying a condition based on the CU skip flag, the encoding device may include an MMVD merge flag in the image information. Here, the condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the encoding device may subsequently signal the MMVD merge flag.
[0269] As another example, the encoding device may include an MMVD merge flag in the image information based on satisfying a condition based on information about the current block and the partition mode enable flag. Here, the condition based on information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or include a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block, information about the current block, and the partition mode enable flag are satisfied, the encoding device may signal the MMVD merge flag.
[0270] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the encoding device may determine whether the conditions based on information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on information about the current block and the partition mode enable flag are not satisfied, the encoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied.
[0271] Meanwhile, when the product of the width and height of the current block is not 32 and the value of the MMVD enable flag is 1, or when the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are 8 or greater, the encoding device may signal the MMVD merge flag.
[0272] To this end, as an example, the merge data syntax may be configured as shown in Table 7 below.
[0273] [Table 7]
[0274]
[0275]
[0276]
[0277]
[0278] mmvd_merge_flag[x0][y0] equal to 1 specifies that merge mode with motion vector difference is used to generate inter prediction parameters for the current coding unit. The array index x0, y0 specifies the position (x0, y0) of the top left luma sample of the considered coding block relative to the top left luma sample of the picture.
[0279] When mmvd_merge_flag[x0][y0] does not exist, the inference is as follows.
[0280] mmvd_merge_flag[x0][y0] is inferred to be equal to 1 if all of the following conditions are true.
[0281] -general_merge_flag[x0][y0] is equal to 1
[0282] -regular_merge_flag[x0][y0] is equal to 0
[0283] -sps_mmvd_enable_flag equal to 1
[0284] - cbWidth * cbHeight != 32
[0285] -MaxNumSubblockMergeCand<=0 or cbWidth<8 or cbHeight<8
[0286] -sps_ciip_enabled_flag is equal to 0 or cbWidth >= 128 or cbHeight >= 128 or cu_skip_flag[x0][y0] is equal to 1
[0287] - The -sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand < 2 or the slice_type is not equal to B_SLICE
[0288] Otherwise, mmvd_merge_flag[x0][y0] is inferred to be equal to 0.
[0289] Meanwhile, according to another embodiment, as an example, an encoding device may include a merge sub-block flag in image information based on satisfying a condition based on a CIIP enable flag and a condition based on the size of a current block, so as not to repeatedly transmit the same syntax. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, and the value of the CIIP enable flag is 1, the encoding device may signal the merge sub-block flag.
[0290] As another example, in addition to the condition based on the CIIP enable flag and the condition based on the size of the current block, based on further satisfying a condition based on a CU skip flag, the encoding device may include a merge sub-block flag in image information. Here, the condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the encoding device may subsequently signal the merge sub-block flag.
[0291] As another example, the encoding device may include a merge sub-block flag in image information based on satisfying a condition based on information about the current block and a partition mode enable flag. Here, the condition based on information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or include a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block and the information about the current block and the condition based on the partition mode enable flag are satisfied, the encoding device may signal the merge sub-block flag.
[0292] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the encoding device may determine whether the conditions based on the information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on the information about the current block and the partition mode enable flag are not satisfied, the encoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied.
[0293] Meanwhile, when the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are each 8 or greater, the encoding device may signal a merge sub-block flag.
[0294] For this purpose, as an example, the merge data syntax may be configured as shown in Table 8 below.
[0295] [Table 8]
[0296]
[0297]
[0298]
[0299] merge_subblock_flag[x0][y0] specifies whether to infer sub-block based inter prediction parameters for the current coding unit from adjacent blocks. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block being considered relative to the top-left luma sample of the picture.
[0300] When merge_subblock_flag[x0][y0] does not exist, the inference is as follows.
[0301] If all of the following conditions are true, then merge_subblock_flag[x0][y0] is inferred to be equal to 1.
[0302] - general_merge_flag[x0][y0] is equal to 1
[0303] - regular_merge_flag[x0][y0] is equal to 0
[0304] - merge_subblock_flag[x0][y0] is equal to 0
[0305] - mmvd_merge_flag[x0][y0] is equal to 0
[0306] - MaxNumSubblockMergeCand > 0
[0307] - cbWidth >= 8 and cbHeight >= 8
[0308] - The sps_ciip_enabled_flag is equal to 0 or cbWidth >= 128 or cbHeight >= 128 or cu_skip_flag[x0][y0] is equal to 1
[0309] - The sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand < 2 or the slice_type is not equal to B_SLICE
[0310] Otherwise, merge_subblock_flag[x0][y0] is inferred to be equal to 0. [[ID=II]]
[0311] Meanwhile, according to another embodiment, by way of example, the encoding device may include a CIIP flag in the image information based on satisfying a condition based on the CIIP enable flag and a condition based on the size of the current block, so as not to repeatedly transmit the same syntax. Here, the condition based on the size of the current block may be the case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CIIP enable flag may be the case where the value of the CIIP enable flag is 1. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are less than 128, and the value of the CIIP enable flag is 1, the encoding device may signal the CIIP flag.
[0312] As another example, in addition to the condition based on the CIIP enable flag and the condition based on the size of the current block, based on further satisfying a condition based on the CU skip flag, the encoding device may include a CIIP flag in the image information. Here, the condition based on the CU skip flag may be the case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the encoding device may subsequently signal the CIIP flag.
[0313] As another example, the encoding device may include a CIIP flag in the image information based on a condition that is satisfied based on information about the current block and a partition mode enable flag. Here, the condition based on information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block and information about the current block and the condition based on the partition mode enable flag are satisfied, the encoding device may signal the CIIP flag.
[0314] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the encoding device may determine whether the conditions based on information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on information about the current block and the partition mode enable flag are not satisfied, the encoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied.
[0315] For this purpose, as an example, the merge data syntax may be configured as shown in Table 9 below.
[0316] [Table 9]
[0317]
[0318]
[0319]
[0320] ciip_flag[x0][y0] specifies whether combined inter-picture merge and intra-picture prediction are applied to the current coding unit. Ciip_flag (x0) (y0) specifies whether combined inter-picture merge and intra-picture prediction are applied to the current coding unit.
[0321] When ciip_flag[x0][y0] does not exist, the following inference is made.
[0322] If all of the following conditions are true, then ciip_flag[x0][y0] is inferred to be equal to 1.
[0323] - general_merge_flag[x0][y0] is equal to 1
[0324] - regular_merge_flag[x0][y0] is equal to 0
[0325] - merge_subblock_flag[x0][y0] is equal to 0
[0326] -mmvd_merge_flag[x0][y0] is equal to 0
[0327] -sps_ciip_enabled_flag is equal to 1
[0328] -cu_skip_flag[x0][y0] is equal to 0
[0329] -cbWidth * cbHeight >= 64 and cbWidth < 128 and cbHeight < 128
[0330] -sps_triangle_enabled_flag is equal to 0 or MaxNumTriangleMergeCand < 2 or slice_type is not equal to B_SLICE
[0331] Otherwise, infer that ciip_flag[x0][y0] is equal to 0.
[0332] Figure 15 and Figure 16 Schematically shows an example of a video / image decoding method including an inter-frame prediction method according to an embodiment of the present disclosure and related components.
[0333] Figure 15 The decoding method disclosed in Figure 3 and Figure 16 can be executed by the decoding device 300 illustrated in Figure 15 Specifically, for example, steps S1500 to S1520 can be executed by the predictor 330 of the decoding device 300, and step S1530 can be executed by the adder 340 of the decoding device 300. Figure 15 The decoding method disclosed in
[0334] Referring to Figure 15 and Figure 16 , the decoding device can obtain information about the prediction mode of the current block from the bitstream (S150). Specifically, the entropy decoder 310 of the decoding device can derive residual information and information about the prediction mode from the signal received from the Figure 2 encoding device in the form of a bitstream. Here, the information about the prediction mode can be referred to as prediction-related information. The information about the prediction mode can include inter-frame / intra-frame prediction classification information, inter-frame prediction mode information, etc., and can include various syntax elements related thereto.
[0335] The bitstream may include picture information, which includes information about various parameter sets, such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). The picture information may also include information about the prediction mode of the current block, such as the Compilation Unit syntax and the Merge Data syntax. The Sequence Parameter Set may include a CIIP enable flag, an enable flag for a partitioning mode, etc. The Compilation Unit syntax may include a CU skip flag indicating whether the skip mode is applied to the current block.
[0336] Meanwhile, the residual processor 320 of the decoding device may generate residual samples based on the residual information. In addition, the predictor 330 of the decoding device may derive the prediction mode of the current block based on the information about the prediction mode (S1510). Additionally, the motion information of the current block may be derived based on the derived prediction mode. In this case, the predictor of the decoding device may construct a motion information candidate list based on adjacent blocks of the current block and derive the motion vector and / or reference picture index of the current block based on the candidate selection information received from the encoding device. When deriving the motion information of the current block, the predictor of the decoding device may generate prediction samples of the current block based on the motion information of the current block (S1520). Thereafter, the adder 340 of the decoding device may generate reconstructed samples based on the prediction samples generated by the predictor 330 and the residual samples generated by the residual processor 320 (S1530). Reconstructed pictures may be generated based on the reconstructed samples. Thereafter, in-loop filtering processes such as deblocking filtering, SAO, and / or ALF processes may be applied to the reconstructed pictures to improve subjective / objective picture quality as needed.
[0337] As an example, when deriving the prediction mode of the current block, the decoding device may obtain a regular merge flag from the bitstream based on a condition based on the CIIP enable flag and a condition based on the size of the current block. Here, the condition based on the size of the current block may be a case where the product of the height and the width of the current block is 64 or greater and the height and the width of the current block are each less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and the width of the current block is 64 or greater, the height and the width of the current block are each less than 128, and the value of the CIIP enable flag is 1, the decoding device may then parse the regular merge flag from the merge data syntax included in the bitstream.
[0338] As another example, a conventional merge flag may be obtained from a bitstream based on conditions that satisfy a CU skip flag and the size of a current block. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, and the value of the CU skip flag is 0, the decoding device may subsequently parse the conventional merge flag from the merge data syntax included in the bitstream.
[0339] As another example, in addition to the conditions based on the CIIP enable flag and the size of the current block, based on further satisfying the condition based on the CU skip flag, the decoding device may obtain the conventional merge flag from the bitstream. Here, the condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the decoding device may subsequently parse the conventional merge flag from the merge data syntax.
[0340] As another example, the decoding device may obtain the conventional merge flag from the bitstream based on conditions that satisfy information about the current block and a partition mode enable flag. Here, the condition based on the information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block, the information about the current block, and the partition mode enable flag are satisfied, the decoding device may parse the conventional merge flag from the merge data syntax.
[0341] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the decoding device may determine whether the conditions based on the information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on the information about the current block and the partition mode enable flag are not satisfied, the decoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied. [[ID=z]]
[0342] Meanwhile, when the product of the width and height of the current block is not 32 and the value of the MMVD enable flag is 1, or when the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are 8 or greater, the decoding device may parse the conventional merge flag from the bitstream. To this end, the merge data syntax may be configured as shown in Table 6 above.
[0343] When there is no regular merge flag in the bitstream, if the value of the general merge flag is 1, the value of the MMVD enable flag in the SPS is 0, the product of the width and height of the current block is 32, the maximum number of sub-block merge candidates is 0 or less, the width of the current block is less than 8, the height of the current block is less than 8, the value of the CIIP enable flag in the SPS is 0, the product of the width and height of the current block is less than 64, the width of the current block is 128 or greater, the value of the CU skip flag is 1, the value of the partition enable flag in the SPS is 0, the maximum number of partition merge candidates is less than 2, or the slice type is not a B slice, the decoding device may deduce that the value of the regular merge flag is 1. Otherwise, the value of the regular merge flag may be deduced as 0.
[0344] As another example, when deriving the prediction mode of the current block, the decoding device may obtain the MMVD merge flag from the bitstream based on conditions that satisfy a condition based on the CIIP enable flag and a condition based on the size of the current block. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and each of the height and width of the current block is less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and width of the current block is 64 or greater, each of the height and width of the current block is less than 128, and the value of the CIIP enable flag is 1, the decoding device may then parse the MMVD merge flag from the merge data syntax included in the bitstream.
[0345] As another example, the decoding device may obtain the MMVD merge flag from the bitstream based on conditions that satisfy a condition based on the CU skip flag and the size of the current block. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and each of the height and width of the current block is less than 128. The condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, each of the height and width of the current block is less than 128, and the value of the CU skip flag is 0, the decoding device may then parse the MMVD merge flag from the merge data syntax included in the bitstream.
[0346] As another example, in addition to the conditions based on the CIIP enable flag and the size of the current block, based on further satisfying the condition based on the CU skip flag, the decoding device may obtain the MMVD merge flag from the bitstream. Here, the condition based on the CU skip flag may be the case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the decoding device may subsequently parse the MMVD merge flag from the merge data syntax.
[0347] As another example, the decoding device may obtain the MMVD merge flag from the bitstream based on satisfying the condition based on the information about the current block and the partition mode enable flag. Here, the condition based on the information about the current block may include the case where the product of the width and height of the current block is 64 or greater and / or include the case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be the case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block, the information about the current block, and the partition mode enable flag are satisfied, the decoding device may parse the MMVD merge flag from the merge data syntax.
[0348] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the decoding device may determine whether the conditions based on the information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on the information about the current block and the partition mode enable flag are not satisfied, the decoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied.
[0349] Meanwhile, when the product of the width and height of the current block is not 32 and the value of the MMVD enable flag is 1, or when the maximum number of sub-block merge candidates is greater than 0 and the width and height of the current block are 8 or greater, the decoding device may parse the MMVD merge flag from the bitstream. To this end, the merge data syntax may be configured as shown in Table 7 above.
[0350] When the MMVD merge flag does not exist in the bitstream, if the value of the general merge flag is 1, the value of the general merge flag is 0, the value of the MMVD enable flag of the SPS is 1, the product of the width and height of the current block is not 32, the maximum number of sub-block merge candidates is 0 or less, the width of the current block is less than 8, the height of the current block is less than 8, the value of the CIIP enable flag of the SPS is 0, the width of the current block is 128 or greater, the height of the current block is 128 or greater, the value of the CU skip flag is 1, the value of the partition enable flag of the SPS is 0, the maximum number of partition merge candidates is less than 2, or the slice type is not a B slice, the decoding device can deduce that the value of the MMVD merge flag is 1. Otherwise, the value of the MMVD merge flag can be deduced as 0.
[0351] As another embodiment, when deriving the prediction mode of the current block, the decoding device may obtain a merged sub-block flag from the bitstream based on conditions that satisfy the condition based on the CIIP enable flag and the condition based on the size of the current block. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, and the value of the CIIP enable flag is 1, the decoding device may subsequently parse the merged sub-block flag from the merge data syntax included in the bitstream.
[0352] As another example, the merged sub-block flag may be obtained from the bitstream based on conditions that satisfy the condition based on the CU skip flag and the size of the current block. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, and the value of the CU skip flag is 0, the decoding device may subsequently parse the merged sub-block flag from the merge data syntax included in the bitstream.
[0353] As another example, in addition to the conditions based on the CIIP enable flag and the current block size, the decoding device may also obtain the merge sub-block flag from the bitstream based on the further satisfaction of the conditions based on the CU skip flag. Here, the condition based on the CU skip flag may be when the CU skip flag value is 0. In other words, when the product of the current block height and the current block width is 64 or greater, the current block height and the current block width are less than 128, the CIIP enable flag value is 1, and the CU skip flag value is 0, the decoding device may then parse the merge sub-block flag from the merge data syntax.
[0354] As another example, the decoding device may obtain a merge sub-block flag from the bitstream based on a condition based on information about the current block and a partition mode enable flag. Here, the condition based on the information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the condition based on the height of the current block and the condition based on the information about the current block and the partition mode enable flag are satisfied, the decoding device may parse the merge sub-block flag from the merge data syntax.
[0355] When the condition based on the CIIP enable flag and the condition based on the size of the current block are not satisfied, the decoding device may determine whether the condition based on the information about the current block and the partition mode enable flag is satisfied. Alternatively, when the condition based on the information about the current block and the partition mode enable flag is not satisfied, the decoding device may determine whether the condition based on the CIIP enable flag and the condition based on the size of the current block are satisfied.
[0356] At the same time, when the maximum number of sub-block merging candidates is greater than 0 and the width and height of the current block are 8 or more, the decoding device can parse the merge sub-block flag from the bitstream. To this end, the merge data syntax can be configured as shown in Table 8 above.
[0357] When the merge sub-block flag does not exist in the bitstream, if the value of the general merge flag is 1, the value of the regular merge flag is 0, the value of the merge sub-block flag is 0, the value of the MMVD merge flag is 0, the maximum number of sub-block merge candidates is greater than 0, the width and height of the current block are 8 or greater, the value of the CIIP enable flag of the SPS is 0, the width of the current block is 128 or greater, the height of the current block is 128 or greater, the value of the CU skip flag is 1, the value of the partition enable flag of the SPS is 0, the maximum number of partition merge candidates is less than 2, or the slice type is not a B slice, then the decoding device may derive the value of the merge sub-block flag as 1. Otherwise, the value of the merge sub-block flag may be derived as 1.
[0358] As another example, when deriving the prediction mode of the current block, the decoding device may obtain the MMVD merge flag from the bitstream based on a condition based on the CIIP enable flag and a condition based on the size of the current block. Here, the condition based on the size of the current block may be a case where the product of the height and width of the current block is 64 or greater and the height and width of the current block are each less than 128. The condition based on the CIIP enable flag may be a case where the value of the CIIP enable flag is 1. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are each less than 128, and the value of the CIIP enable flag is 1, the decoding device may then parse the CIIP flag from the merge data syntax.
[0359] As another example, in addition to the condition based on the CIIP enable flag and the condition based on the size of the current block, based on further satisfying a condition based on the CU skip flag, the decoding device may obtain the CIIP flag from the bitstream. Here, the condition based on the CU skip flag may be a case where the value of the CU skip flag is 0. In other words, when the product of the height and width of the current block is 64 or greater, the height and width of the current block are less than 128, the value of the CIIP enable flag is 1, and the value of the CU skip flag is 0, the decoding device may then parse the CIIP flag from the merge data syntax.
[0360] As another example, the decoding device may obtain the CIIP flag from the bitstream based on satisfying a condition based on information about the current block and the partition mode enable flag. Here, the condition based on information about the current block may include a case where the product of the width and height of the current block is 64 or greater and / or include a case where the slice type of the current block is a B slice. The condition based on the partition mode enable flag may be a case where the value of the partition mode enable flag is 1. In other words, when the conditions based on the height of the current block, information about the current block, and the partition mode enable flag are satisfied, the decoding device may parse the CIIP flag from the merge data syntax.
[0361] When the conditions based on the CIIP enable flag and the size of the current block are not satisfied, the decoding device may determine whether the conditions based on information about the current block and the partition mode enable flag are satisfied. Alternatively, when the conditions based on information about the current block and the partition mode enable flag are not satisfied, the decoding device may determine whether the conditions based on the CIIP enable flag and the size of the current block are satisfied. To this end, the merge data syntax may be configured as shown in Table 9 above.
[0362] When the CIIP flag is not present in the bitstream, if the value of the general merge flag is 1, the value of the regular merge flag is 0, the value of the merge sub-block flag is 0, the value of the MMVD merge flag is 0, the value of the CIIP enable flag in the SPS is 1, the value of the CU skip flag is 0, the product of the width and height of the current block is 64 or greater, the width and height of the current block are less than 128, the value of the partition enable flag in the SPS is 0, the maximum number of partition merge candidates is less than 2, or the slice type is not a B slice, the decoding device may deduce that the value of the CIIP flag is 1. Otherwise, the value of the CIIP flag may be deduced as 0.
[0363] Although the method has been described based on the flowchart that sequentially lists steps or blocks in the above embodiments, the steps of the present disclosure are not limited to a specific order, and specific steps may be performed in different steps or in a different order or simultaneously with respect to the above steps. Additionally, those of ordinary skill in the art will understand that the steps in the flowchart are not exclusive, and another step may be included therein, or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0364] The above-mentioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in a device for performing image processing (e.g., TV, computer, smart phone, set-top box, display device, etc.).
[0365] When implementing the embodiments of the present disclosure in software, the above-mentioned method may be implemented by modules (processes or functions) that execute the above-mentioned functions. The modules may be stored in a memory and executed by a processor. The memory may be installed inside or outside the processor and may be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. In other words, the embodiments according to the present disclosure may be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in the corresponding figures may be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information about the implementation (e.g., information about instructions) or algorithms may be stored in a digital storage medium.
[0366] In addition, the decoding device and the encoding device according to the embodiments of the present disclosure may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital cinema video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video-on-demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, a vehicle-mounted terminal (e.g., a vehicle (including an autonomous vehicle)-mounted terminal, an aircraft terminal, or a ship terminal), and a medical video device; and may be used to process image signals or data. For example, the OTT video device may include a game console, a Blueray player, an Internet-connected TV, a home theater system, a smart phone, a tablet PC, and a digital video recorder (DVR).
[0367] In addition, the processing method according to the embodiments of the present disclosure can be generated in the form of a program executable by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of the present disclosure can also be stored in the computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices storing computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (e.g., transmission on the Internet). In addition, the bitstream generated by the encoding method can be stored in the computer-readable recording medium or can be transmitted through a wired or wireless communication network.
[0368] In addition, the embodiments of the present disclosure can be implemented as a computer program product based on program code, and the program code can be executed on a computer according to the embodiments of this document. The program code can be stored on a computer-readable carrier.
[0369] Figure 17 An example of a content stream system to which the embodiments of the present disclosure can be applied is shown.
[0370] Referring to Figure 17 , the content stream system to which the embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a web server, a media warehouse, a user device, and a multimedia input device.
[0371] The encoding server is used to compress the content input from multimedia input devices such as smart phones, cameras, and portable video cameras into digital data, generate a bitstream, and transmit it to the streaming server. As another example, in the case where a multimedia input device such as a smart phone, a camera, or a portable video camera directly generates a bitstream, the encoding server can be omitted.
[0372] The bitstream can be generated by an encoding method or a bitstream generation method applicable to the embodiments of the present disclosure. And the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0373] Based on the user's request, the streaming server transmits multimedia data to the user device through the network server, which serves as a tool to notify the user of what services are available. When the user requests a service that the user wants, the network server transfers the request to the streaming server, and the streaming server transmits the multimedia data to the user. In this regard, the content streaming system can include a separate control server, and in this case, the control server is used to control the commands / responses between the various devices in the content streaming system.
[0374] The streaming server can receive content from the media storage device and / or the encoding server. For example, in the case of receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide the streaming service.
[0375] For example, the user device can include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a watch-type terminal (smart watch), a glasses-type terminal (smart glasses), a head-mounted display (HMD)), a digital TV, a desktop computer, a digital signage, etc.
[0376] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
Claims
1. A decoding device for image decoding, the decoding device comprising: A memory; And At least one processor, the at least one processor being connected to the memory, the at least one processor being configured to: Obtain information about a prediction mode of a current block from a bitstream; Derive the prediction mode of the current block based on the information about the prediction mode; Generate a prediction sample of the current block based on the prediction mode; And Generate a reconstructed sample based on the prediction sample; Wherein, the bitstream includes a sequence parameter set, The sequence parameter set includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, and The at least one processor is configured to derive the prediction mode of the current block by parsing a regular merge flag from the bitstream based on satisfying a condition based on the CIIP enable flag and a condition based on the size of the current block, Wherein, the condition based on the size of the current block is (i) the product of the height of the current block and the width of the current block is 64 or greater and (ii) the height and the width of the current block are each less than 128, Wherein, the condition based on the CIIP enable flag is that the value of the CIIP enable flag is 1.
2. An encoding device for image decoding, the encoding device comprising: A memory; And At least one processor, the at least one processor being connected to the memory, the at least one processor being configured to: Determine the prediction mode of a current block; Generate information about the prediction mode based on the prediction mode; and Encode image information including the information about the prediction mode; Wherein, the image information includes a sequence parameter set, The sequence parameter includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, and Based on satisfying a condition based on the CIIP enable flag and a condition based on the size of the current block, the image information includes a regular merge flag, Wherein, the condition based on the size of the current block is (i) the product of the height of the current block and the width of the current block is 64 or greater and (ii) the height and the width of the current block are each less than 128, Wherein, the condition based on the CIIP enable flag is that the value of the CIIP enable flag is 1.
3. A device for sending data for an image, the device comprising: At least one processor, the at least one processor being configured to obtain a bitstream for the image, wherein the bitstream is generated based on: determining the prediction mode of a current block, generating information about the prediction mode based on the prediction mode, and encoding image information including the information about the prediction mode; And A transmitter, the transmitter being configured to transmit the data including the bitstream, Wherein, the image information includes a sequence parameter set, The sequence parameter includes a combined inter-picture merge and intra-picture prediction (CIIP) enable flag, and Based on satisfying conditions based on the CIIP enable flag and conditions based on the size of the current block, the image information includes a regular merge flag, wherein the condition based on the size of the current block is that (i) the product of the height of the current block and the width of the current block is 64 or greater and (ii) the height and the width of the current block are each less than 128, wherein the condition based on the CIIP enable flag is that the value of the CIIP enable flag is 1.