Image encoding / decoding method and apparatus, and recording medium storing bit stream

Through the neural network-based image generation method, multiple image inputs are used to perform prediction operations, the problem of low image coding efficiency in the prior art is solved, and higher quality image compression and less encoded data volume are achieved.

CN120359748APending Publication Date: 2025-07-22ELECTRONICS & TELECOMM RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088132.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing image encoding technology is inefficient in high resolution and high definition image compression, making it difficult to meet users' needs for higher quality images.

Method used

Using a neural network-based image generation method, the encoding efficiency and image quality are improved by performing prediction operations and generating recovery images using two or more images as inputs.

Benefits of technology

The predicted image quality generated by neural networks is better, reducing the amount of encoded data, and improving image compression performance and coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359748A_ABST
    Figure CN120359748A_ABST
Patent Text Reader

Abstract

The invention provides an image encoding / decoding method and apparatus, and a recording medium storing a bit stream. An image decoding method according to an embodiment may comprise: performing a prediction operation; generating a restored image based on an image generated by the prediction operation; and generating an image to be used for prediction based on the restored image, in which at least one or more of performing the prediction operation and generating the image includes generating an image through a neural network using two or more images as an input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods, apparatuses, and recording media for image encoding / decoding. More specifically, the present disclosure relates to neural network-based image generation techniques. Background Art

[0002] With the continuous development of the information and communication industry, broadcast services supporting high-definition (HD) resolution have become widespread worldwide. Through this widespread adoption, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.

[0003] To meet the users' demand for high definition, a large number of institutions have accelerated the development of next-generation imaging devices. In addition to high-definition TVs (HDTVs) and full high-definition (FHD) TVs, users' interest in ultra-high-definition (UHD) TVs has also increased, where the resolution of a UHD TV is more than four times that of an FHD TV. With this increasing interest, there is now a need for image encoding / decoding techniques for images with higher resolution and higher clarity.

[0004] As image compression techniques, there are various techniques (such as inter-frame prediction techniques, intra-frame prediction techniques, transformation, quantization techniques, filtering techniques, and entropy coding techniques).

[0005] The inter-frame prediction technique is a technique for predicting the values of pixels included in a current picture using a picture before the current picture and / or a picture after the current picture. The intra-frame prediction technique is a technique for predicting the values of pixels included in a current picture using information about pixels in the current picture. The transformation and quantization techniques can be techniques for compressing the energy of a residual signal. The entropy coding technique is a technique for assigning short codewords to frequently occurring values and long codewords to less frequently occurring values.

[0006] By utilizing these image compression techniques, data regarding images can be effectively compressed, transmitted, and stored. Summary of the Invention

[0007] Technical Problem An object of the present disclosure is to improve the encoding efficiency in image encoding.

[0008] Another object of the present disclosure is to provide a neural network-based method for generating an image signal, which improves the encoding efficiency.

[0009] Still another object of the present disclosure is to provide a neural network-based method for generating an image signal, which improves the quality of the output image signal.

[0010] Technical Solution An image decoding method according to an embodiment of the present disclosure may include: performing a prediction operation; generating a restored image based on the image generated by the prediction operation; and generating an image to be used for prediction based on the restored image, wherein at least one or more steps of performing the prediction operation and generating the image include generating an image by using a neural network with two or more images as inputs.

[0011] An image encoding method according to another embodiment of the present disclosure may include: downsampling an input image based on a scaling factor; performing a prediction operation on the downsampled image; generating a residual image based on the image generated by the prediction operation and encoding the residual image; decoding the encoded residual image to generate a restored image; and generating an image to be used for prediction based on the restored image, wherein at least one or more steps of performing the prediction operation and generating the image include generating an image by using a neural network with two or more images as inputs.

[0012] A computer-readable storage medium according to still another embodiment of the present disclosure may store a bitstream of image information, the image information may be generated by performing an image encoding method, and the image encoding method may include: downsampling an input image based on a scaling factor; performing a prediction operation on the downsampled image; generating a residual image based on the image generated by the prediction operation and encoding the residual image; decoding the encoded residual image to generate a restored image; and generating an image to be used for prediction based on the restored image, wherein at least one or more steps of performing the prediction operation and generating the image include generating an image by using a neural network with two or more images as inputs.

[0013] Advantageous Effects According to an embodiment of the present disclosure, a prediction image with better quality can be generated through a neural network, thereby improving the compression performance of the image.

[0014] Moreover, the quality of the output image of the neural network can be improved by diversifying the inputs of the neural network for generating the image.

[0015] In addition, the image can be upsampled through a neural network to have better quality, and the input image can be downsampled for encoding, thereby reducing the amount of encoded data.

[0016] In addition, the quality of the reference image for prediction can be improved through a neural network, and a residual image with less energy can be obtained, thereby achieving higher encoding efficiency. Brief Description of the Drawings

[0017] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding apparatus to which the present disclosure is applied; Figure 2is a block diagram showing a configuration of an embodiment of a decoding apparatus to which the present disclosure is applied; Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded; Figure 4 is a diagram showing forms of prediction units (PUs) that a coding unit (CU) can include; Figure 5 is a diagram showing forms of transform units (TUs) that can be included in a CU; Figure 6 shows a division of a block according to an example; Figure 7 is a diagram for explaining an embodiment of an intra prediction process; Figure 8 is a diagram showing reference sample points used in the intra prediction process; Figure 9 is a diagram for explaining an embodiment of an inter prediction process; Figure 10 shows spatial candidates according to an embodiment; Figure 11 shows an order of adding motion information of spatial candidates to a merge list according to an embodiment; Figure 12 shows transform and quantization processing according to an example; Figure 13 shows diagonal scanning according to an example; Figure 14 shows horizontal scanning according to an example; Figure 15 shows vertical scanning according to an example; Figure 16 is a configuration diagram of an encoding apparatus according to an embodiment; Figure 17 is a configuration diagram of a decoding apparatus according to an embodiment; Figure 18 is a flowchart showing a process of processing an image using a neural network according to an embodiment; Figure 19 is a flowchart showing an encoding method for generating a bitstream by prediction of a target block according to an embodiment; Figure 20 is a flowchart showing a method for predicting a target block using a bitstream according to an embodiment; Figure 21 is a flowchart showing a prediction method according to an embodiment; Figure 22 shows an operation flowchart for applying neural network technology according to an embodiment; Figure 23is a flowchart showing a unidirectional prediction method according to an embodiment; Figure 24 is a flowchart showing a first bidirectional prediction method according to an embodiment; Figure 25 is a flowchart showing a second bidirectional prediction method according to an embodiment; Figure 26 is a flowchart showing a third bidirectional prediction method according to an embodiment; Figure 27 is a flowchart showing a fourth bidirectional prediction method according to an embodiment; Figure 28 shows the structure of an attention residual convolutional neural network model according to an embodiment; Figure 29 is a flowchart showing a prediction method using bidirectional prediction according to an embodiment; Figure 30 is a flowchart briefly showing an image encoding / decoding method using a neural network-based super-resolution technique according to an embodiment; Figure 31 is an exemplary block diagram showing the structure of an encoding apparatus for implementing a neural network-based super-resolution technique; Figure 32 shows a 12-tap interpolation filter defined in a reference picture resampling technique; Figure 33 is to show Figure 30 the flowchart of steps of a selected upsampling method; Figure 34 shows an 8-tap interpolation filter defined in a reference picture resampling technique; Figure 35 shows a 4-tap interpolation filter defined in a reference picture resampling technique; Figure 36 is to show Figure 30 the flowchart of steps of a selected neural network-based upsampling model combination; Figures 37 to 42 shows the structure of an upsampling replacement residual convolutional neural network model for reference picture resampling based on an interpolation filter; Figures 43a to 43g describes an image input to a neural network model; Figure 44 is an exemplary block diagram showing the structure of a decoding apparatus for implementing a neural network-based super-resolution technique; Figure 45 is a flowchart of operations of an image encoding method according to an embodiment; and Figure 46 is a flowchart of operations of an image decoding method according to an embodiment. Detailed implementation mode

[0018] Best mode The image decoding method according to an embodiment of the present disclosure may include: performing a prediction operation; generating a restored image based on the image generated by the prediction operation; and generating an image to be used for prediction based on the restored image, wherein at least one or more of performing the prediction operation and generating the image include generating an image by using a neural network with two or more images as inputs.

[0019] In one embodiment, the execution of the prediction operation may include generating a prediction block by using two intermediate prediction blocks obtained from one or more reference pictures through motion compensation as inputs to the neural network.

[0020] In one embodiment, the execution of the prediction operation further includes: when only one of the intermediate prediction blocks is available, generating a second intermediate prediction block based on the available first intermediate prediction block.

[0021] In one embodiment, generating an image to be used for prediction may include generating a reference picture by using two restored images as inputs to the neural network.

[0022] In one embodiment, generating an image to be used for prediction may further include adding the generated reference picture to reference picture list 0 and reference picture list 1 with the same POC as the current picture for which the prediction operation is performed.

[0023] In one embodiment, the method may further include obtaining an upsampled image through a neural network that uses the restored image and an image obtained by upsampling the restored image through an interpolation filter as inputs.

[0024] In one embodiment, obtaining the upsampled image may additionally use a predicted image generated by performing the prediction operation as an input to the neural network.

[0025] In one embodiment, the neural network for obtaining the upsampled image may use one or more of quantization parameters, slice information, and block partition information as additional inputs.

[0026] In one embodiment, the neural network for obtaining the upsampled image may perform pixel de-shuffling on an image upsampled to the resolution of the restored image through an interpolation filter for use as an input or merge the upsampled image into the output of the neural network.

[0027] In one embodiment, an image generated by decomposing an image upsampled through an interpolation filter into high-frequency components and low-frequency components in the horizontal and vertical directions may be used as an input to the neural network for obtaining the upsampled image.

[0028] In one embodiment, generating an image by a neural network may include: when generating an image of a chrominance component, using an image of a luminance component corresponding to the chrominance component as an additional input.

[0029] An image encoding method according to another embodiment of the present disclosure may include: downsampling an input image based on a scaling factor; performing a prediction operation on the downsampled image; generating a residual image based on the image generated by the prediction operation and encoding the residual image; decoding the encoded residual image to generate a restored image; and generating an image to be used for prediction based on the restored image, where at least one or more of performing the prediction operation and generating the image include generating an image by a neural network using two or more images as inputs.

[0030] A computer-readable storage medium according to still another embodiment of the present disclosure may store a bitstream of image information, where the image information may be generated by an image encoding method including: downsampling an input image based on a scaling factor; performing a prediction operation on the downsampled image; generating a residual image based on the image generated by the prediction operation and encoding the residual image; decoding the encoded residual image to generate a restored image; and generating an image to be used for prediction based on the restored image, where at least one or more of performing the prediction operation and generating the image include generating an image by a neural network using two or more images as inputs.

[0031] Modes for the Invention The present invention can be variously changed and can have various embodiments. Specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the present invention to a specific disclosed form, and they include all changes, equivalents, or modifications included within the spirit and scope of the present invention.

[0032] The following exemplary embodiments will be described in detail with reference to the accompanying drawings showing specific embodiments. These embodiments are described such that those of ordinary skill in the art to which the present disclosure pertains can easily implement these embodiments. It should be noted that the various embodiments are different from each other but do not need to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be implemented as other embodiments without departing from the spirit and scope of other embodiments related to one embodiment. In addition, it should be understood that the positions or arrangements of the respective components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Therefore, the following detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is only defined by the appended claims and their equivalents (as long as they are properly described).

[0033] In the drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, dimensions, etc. of the components in the drawings may be exaggerated to make the description clear.

[0034] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by such terms. The terms are only used to distinguish one component from another. For example, without departing from the scope of this specification, the first component may be referred to as the second component. Similarly, the second component may be referred to as the first component. The term "and / or" may include combinations of multiple related description items or any one of multiple related description items.

[0035] It will be understood that when a component is referred to as "connected" or "coupled" to another component, the two components may be directly connected or coupled to each other, or there may be an intermediate component between the two components. On the other hand, it will be understood that when a component is referred to as "directly connected or coupled", there is no intermediate component between the two components.

[0036] In addition, the components described in the embodiments are independently shown to indicate different characteristic functions, but this does not mean that each component is formed by a single piece of hardware or software. That is, for convenience of description, multiple components are separately arranged and included. For example, at least two of the multiple components may be integrated into a single component. Conversely, one component may be divided into multiple components. As long as the essence of this specification is not departed from, embodiments in which multiple components are integrated or embodiments in which some components are separated are included in the scope of this specification.

[0037] The terms used in the embodiments are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless specifically stated to the contrary in the context. In the embodiments, it should be understood that terms such as "comprising" or "having" are only intended to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. That is, in the embodiments, the expression describing a component "comprising" a specific component means that other components may be included within the scope of the practice of the present invention or the technical spirit of the present invention, but does not exclude the presence of components other than the specific component.

[0038] In the embodiments, the term "at least one" may mean one of one or more quantities (such as 1, 2, 3, and 4). In the embodiments, the term "a plurality" may mean one of two or more quantities (such as 2, 3, and 4).

[0039] Some components of the embodiments are not essential components for performing necessary functions, but can be optional components only for improving performance. The embodiments can be implemented using only the essential components for realizing the essence of the embodiments. For example, a structure that includes only the essential components (excluding the optional components only for improving performance) is also within the scope of the embodiments.

[0040] Embodiments will be described in detail below with reference to the accompanying drawings so that those of ordinary skill in the art to which the embodiments belong can easily implement the embodiments. In the following description of the embodiments, a detailed description of well-known functions or configurations that are considered to obscure the gist of the present specification will be omitted. In addition, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.

[0041] Hereinafter, an "image" may represent a single frame constituting a video, or may represent the video itself. For example, "encoding and / or decoding of an image" may represent "encoding and / or decoding of a video", and may also represent "encoding and / or decoding of any one of a plurality of images constituting a video".

[0042] Hereinafter, the terms "video" and "moving picture" may be used with the same meaning and may be used interchangeably with each other.

[0043] Hereinafter, a target image may be an encoding target image that is a target to be encoded and / or a decoding target image that is a target to be decoded. In addition, the target image may be an input image input to an encoding device or an input image input to a decoding device. And, the target image may be a current image, that is, a target currently to be encoded and / or decoded. For example, the terms "target image" and "current image" may be used with the same meaning and may be used interchangeably with each other.

[0044] Hereinafter, the terms "image", "frame", "picture", and "screen" may be used with the same meaning and may be used interchangeably with each other.

[0045] Hereinafter, a target block may be an encoding target block (i.e., a target to be encoded) and / or a decoding target block (i.e., a target to be decoded). In addition, the target block may be a current block, that is, a target currently to be encoded and / or decoded. Here, the terms "target block" and "current block" may be used with the same meaning and may be used interchangeably with each other. The current block may represent an encoding target block that is an encoding target during encoding and / or a decoding target block that is a decoding target during decoding. In addition, the current block may be at least one of an encoding block, a prediction block, a residual block, and a transform block.

[0046] Hereinafter, the terms "block" and "unit" may be used with the same meaning and may be used interchangeably with each other. Optionally, a "block" may represent a specific unit.

[0047] Hereinafter, the terms "region" and "segment" may be used interchangeably with each other.

[0048] In the following embodiments, specific information, data, flags, indexes, elements, and attributes may have their respective values. The value "0" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate false, logical false, or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value may be used interchangeably with each other. The value "1" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate true, logical true, or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value may be used interchangeably with each other.

[0049] When variables such as i or j are used to indicate a row, column, or index, the value of i may be an integer 0 or an integer greater than 0, or may be an integer 1 or an integer greater than 1. In other words, in an embodiment, each of the row, column, and index may be counted starting from 0, or may be counted starting from 1.

[0050] In an embodiment, the term "one or more" or the term "at least one" may represent the term "multiple". The term "one or more" or the term "at least one" may be used interchangeably with "multiple".

[0051] Below, terms to be used in the embodiments will be described.

[0052] Encoder: An encoder represents a device for performing encoding. That is, an encoder may represent an encoding device.

[0053] Decoder: A decoder represents a device for performing decoding. That is, a decoder may represent a decoding device.

[0054] Unit: A unit may represent a unit of image encoding and decoding. The terms "unit" and "block" may be used with the same meaning and may be used interchangeably with each other.

[0055] – A unit may be an array of samples of M×N. Each of M and N may be a positive integer. A unit generally may represent an array of samples in a two-dimensional form.

[0056] – During the encoding and decoding processes of an image, a "unit" can be a region generated by partitioning an image. In other words, a "unit" can be a region specified within an image. A single image can be partitioned into multiple units. Optionally, an image can be partitioned into sub-parts, and a unit can represent each of the partitioned sub-parts when encoding or decoding is performed on the partitioned sub-parts.

[0057] – During the encoding and decoding processes of an image, predefined processing can be performed on each unit according to the type of the unit.

[0058] – According to functions, unit types can be classified as macro units, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, according to functions, a unit can represent blocks, macro blocks, coding tree units, coding tree blocks, coding units, coding blocks, prediction units, prediction blocks, residual units, residual blocks, transform units, transform blocks, etc. For example, a target unit that is the target of encoding and / or decoding can be at least one of a CU, a PU, a residual unit, and a TU.

[0059] – The term "unit" can represent information including a luma component block, a chroma component block corresponding to the luma component block, and syntax elements for each block, such that the unit is specified as being distinct from a block.

[0060] – The size and shape of a unit can be implemented differently. Moreover, a unit can have any one of various sizes and shapes. Specifically, the shape of a unit can include not only a square but also geometric shapes that can be represented in two dimensions (2D) (such as a rectangle, a trapezoid, a triangle, and a pentagon).

[0061] In addition, unit information can include one or more of the type of the unit, the size of the unit, the depth of the unit, the encoding order of the unit, and the decoding order of the unit, etc. For example, the type of a unit can indicate one of a CU, a PU, a residual unit, and a TU.

[0062] – A unit can be partitioned into sub-units, each sub-unit having a size smaller than that of the relevant unit.

[0063] Depth: Depth can represent the degree to which a unit is partitioned. Moreover, the depth of a unit can indicate the level at which the corresponding unit exists when the unit is represented by a tree structure.

[0064] – Unit partition information can include a depth indicating the depth of the unit. The depth can indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned.

[0065] – In a tree structure, the root node can be considered to have the minimum depth and the leaf node can be considered to have the maximum depth. The root node can be the highest (top) node. The leaf node can be the lowest node.

[0066] – A single unit can be hierarchically partitioned into multiple sub-units, while the single unit has depth information based on a tree structure. In other words, the unit and the sub-units generated by partitioning the unit can respectively correspond to a node and the children nodes of the node. Each partitioned sub-unit can have a unit depth. Since the depth indicates the number of times the unit is partitioned and / or the degree to which the unit is partitioned, the partition information of the sub-unit can include information about the size of the sub-unit.

[0067] In a tree structure, the top node can correspond to the initial node before partitioning. The top node can be referred to as the "root node". In addition, the root node can have the minimum depth value. Here, the depth of the top node can be level "0".

[0068] – A node with a depth of level "1" can represent the unit generated when the initial unit is partitioned once. A node with a depth of level "2" can represent the unit generated when the initial unit is partitioned twice.

[0069] – A leaf node with a depth of level "n" can represent the unit generated when the initial unit is partitioned n times.

[0070] – A leaf node can be the bottom node that cannot be further partitioned. The depth of the leaf node can be the maximum level. For example, the predefined value for the maximum level can be 3.

[0071] – QT depth can represent the depth for four-way partitioning. BT depth can represent the depth for two-way partitioning. TT depth can represent the depth for three-way partitioning.

[0072] – Sample: A sample can be the basic unit that constitutes a block. Samples can be represented by values from 0 to 2 Bd -1 according to the bit depth (Bd).

[0073] – A sample can be a pixel or a pixel value.

[0074] – Hereinafter, the terms "pixel" and "sample" can be used with the same meaning and can be used interchangeably with each other.

[0075] Coding tree unit (CTU): A CTU can be composed of a single luma component (Y) coding tree block and two chroma component (i.e., Cb, Cr) coding tree blocks related to the luma component coding tree block. In addition, a CTU can represent the information including the above-mentioned blocks and the syntax elements for each block.

[0076] – Each coding tree unit (CTU) may be partitioned using one or more partitioning methods such as quadtree (QT), binary tree (BT), and ternary tree (TT) to configure subunits such as coding units, prediction units, and transform units. A quadtree may represent a quaternary tree. Additionally, each coding tree unit may be partitioned using one or more partitioning methods with multi-type tree (MTT).

[0077] – “CTU” may be used as a term to specify a pixel block that is a processing unit in image decoding and encoding processes, such as in the case of partitioning an input image.

[0078] Coding tree block (CTB): “CTB” may be used as a term to specify any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.

[0079] Neighboring block: A neighboring block (or adjacent block) may represent a block adjacent to a target block. A neighboring block may represent a reconstructed neighboring block.

[0080] Hereinafter, the terms “neighboring block” and “adjacent block” may be used to have the same meaning and may be used interchangeably with each other.

[0081] A neighboring block may represent a reconstructed neighboring block.

[0082] Spatial neighboring block: A spatial neighboring block may be a block that is spatially adjacent to a target block. A neighboring block may include a spatial neighboring block.

[0083] – A target block and a spatial neighboring block may be included in a target picture.

[0084] – A spatial neighboring block may represent a block whose boundary touches the target block or a block located within a predetermined distance from the target block.

[0085] – A spatial neighboring block may represent a block adjacent to a vertex of the target block. Here, a block adjacent to a vertex of the target block may represent a block that is vertically adjacent to a horizontally adjacent neighboring block of the target block or a block that is horizontally adjacent to a vertically adjacent neighboring block of the target block.

[0086] Temporal neighboring block: A temporal neighboring block may be a block that is temporally adjacent to a target block. A neighboring block may include a temporal neighboring block.

[0087] – A temporal neighboring block may include a collocated block (col block).

[0088] – A col block may be a block in a previously reconstructed collocated picture (col picture). The position of the col block in the col picture may correspond to the position of the target block in the target picture. Optionally, the position of the col block in the col picture may be equal to the position of the target block in the target picture. The col picture may be a picture included in a reference picture list.

[0089] – A temporally adjacent block may be a block that is temporally adjacent to a spatially adjacent block of the target block.

[0090] Prediction mode: The prediction mode may be information indicating a mode for intra prediction or a mode for inter prediction.

[0091] Prediction unit: The prediction unit may be a basic unit for prediction (such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation).

[0092] – A single prediction unit may be divided into a plurality of partitions or sub-prediction units having smaller sizes. The plurality of partitions may also be basic units during the execution of prediction or compensation. The partitions generated by dividing the prediction unit may also be prediction units.

[0093] Prediction unit partition: The prediction unit partition may be the shape into which the prediction unit is divided.

[0094] Reconstructed neighboring unit: The reconstructed neighboring unit may be a unit that has been decoded and reconstructed and is neighboring to the target unit.

[0095] – The reconstructed neighboring unit may be a unit that is spatially adjacent to the target unit or temporally adjacent to the target unit.

[0096] – The reconstructed spatially adjacent unit may be a unit that has been reconstructed through encoding and / or decoding and is included in the target picture.

[0097] – The reconstructed temporally adjacent unit may be a unit that has been reconstructed through encoding and / or decoding and is included in the reference picture. The position of the reconstructed temporally adjacent unit in the reference picture may be the same as the position of the target unit in the target picture, or may correspond to the position of the target unit in the target picture. In addition, the reconstructed temporally adjacent unit may be a block that is neighboring to the corresponding block in the reference picture. Here, the position of the corresponding block in the reference picture may correspond to the position of the target block in the target picture. Here, the fact that the positions of the blocks correspond to each other may mean that the positions of the blocks are the same as each other, may mean that one block is included in another block, or may mean that one block occupies a specific position in another block.

[0098] Sub-picture: A picture may be divided into one or more sub-pictures. A sub-picture may be composed of one or more parallel block rows and one or more parallel block columns.

[0099] – The sub-picture may be an area in the picture having a square shape or a rectangular (i.e., non-square rectangle) shape. In addition, the sub-picture may include one or more CTUs.

[0100] – The sub-picture may be a rectangular area of one or more stripes in the picture.

[0101] – A sub - picture may include one or more parallel blocks, one or more bricks, and / or one or more stripes.

[0102] Parallel block: A parallel block may be an area in the picture having a square shape or a rectangular (i.e., non - square rectangle) shape.

[0103] – A parallel block may include one or more CTUs.

[0104] – A parallel block may be partitioned into one or more bricks.

[0105] Brick: A brick may represent one or more CTU rows in a parallel block.

[0106] – A parallel block may be partitioned into one or more bricks. Each brick may include one or more CTU rows.

[0107] – A parallel block that is not partitioned into two parts may also represent a brick.

[0108] Stripe: A stripe may include one or more parallel blocks in the picture. Optionally, a stripe may include one or more bricks in a parallel block.

[0109] - A sub - picture may contain one or more stripes that jointly cover a rectangular area of the picture. Thus, each sub - picture boundary is always a stripe boundary, and each vertical sub - picture boundary is always a vertical parallel - block boundary.

[0110] Parameter set: A parameter set may correspond to the header information in the internal structure of a bitstream.

[0111] A parameter set may include at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a decoding parameter set (DPS), etc.

[0112] - The information signaled by each parameter set may be applied to the picture that references the corresponding parameter set. For example, the information in the VPS may be applied to the picture that references the VPS. The information in the SPS may be applied to the picture that references the SPS. The information in the PPS may be applied to the picture that references the PPS.

[0113] - Each parameter set may reference a higher - level parameter set. For example, the PPS may reference the SPS. The SPS may reference the VPS.

[0114] - In addition, a parameter set may include a parallel - block group, stripe - header information, and parallel - block header information. A parallel - block group may be a group that includes multiple parallel blocks. In addition, the meaning of “parallel - block group” may be the same as the meaning of “stripe”.

[0115] Rate - distortion optimization: The encoding device may use rate - distortion optimization to provide high encoding efficiency by utilizing a combination of the following: the size of the coding unit (CU), the prediction mode, the size of the prediction unit (PU), the motion information, and the size of the transform unit (TU).

[0116] – The rate - distortion optimization scheme can calculate the rate - distortion cost for each combination to select the optimal combination from these combinations. The equation “ ” can be used to calculate the rate - distortion cost. Generally, the combination that minimizes the rate - distortion cost can be selected as the optimal combination under the rate - distortion optimization scheme.

[0117] – D can represent the distortion. D can be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., the mean - square error).

[0118] – R can represent the rate, which can use relevant context information to represent the bit - rate.

[0119] – represents the Lagrange multiplier. R can include not only coding parameter information (such as prediction mode, motion information, and coding block flag), but also the bits generated due to encoding the transform coefficients.

[0120] – The encoding device can perform processes such as inter - frame prediction and / or intra - frame prediction, transformation, quantization, entropy coding, inverse quantization (de - quantization), and / or inverse transformation to calculate accurate D and R. These processes will greatly increase the complexity of the encoding device.

[0121] – Bitstream: The bitstream can represent a stream of bits including encoded image information.

[0122] Parsing: Parsing can be the determination of the value of a syntax element by performing entropy decoding on the bitstream. Optionally, the term “parsing” can represent this entropy decoding itself.

[0123] Symbol: A symbol can be at least one of a syntax element, a coding parameter, and a transform coefficient of an encoding target unit and / or a decoding target unit. Additionally, a symbol can be the target of entropy coding or the result of entropy decoding.

[0124] Reference picture: A reference picture can be an image that is referenced by a unit to perform inter - frame prediction or motion compensation. Optionally, a reference picture can be an image including reference units that are referenced by the target unit to perform inter - frame prediction or motion compensation.

[0125] Hereinafter, the terms “reference picture” and “reference image” can be used with the same meaning and can be used interchangeably with each other.

[0126] Reference picture list: A reference picture list can be a list including one or more reference images used for inter - frame prediction or motion compensation.

[0127] – The types of reference picture lists can include a combined list (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.

[0128] – For inter - frame prediction, one or more reference picture lists can be used.

[0129] Inter - frame prediction indicator: The inter - frame prediction indicator can indicate the inter - frame prediction direction for a target unit. Inter - frame prediction can be one of unidirectional prediction and bidirectional prediction. Optionally, the inter - frame prediction indicator can represent the number of reference pictures used to generate the prediction unit for the target unit. Optionally, the inter - frame prediction indicator can represent the number of prediction blocks used for inter - frame prediction or motion compensation of the target unit.

[0130] Prediction list utilization flag: The prediction list utilization flag can indicate whether at least one reference picture in a specific reference picture list is used to generate a prediction unit.

[0131] – The prediction list utilization flag can be used to derive the inter - frame prediction indicator. Conversely, the inter - frame prediction indicator can be used to derive the prediction list utilization flag. For example, the case where the prediction list utilization flag indicates "0" (as the first value) can indicate that for the target unit, the reference pictures in the reference picture list are not used to generate the prediction block. The case where the prediction list utilization flag indicates "1" (as the second value) can indicate that for the target unit, the reference picture list is used to generate the prediction unit.

[0132] Reference picture index: The reference picture index can be an index indicating a specific reference picture in the reference picture list.

[0133] Picture Order Count (POC): The POC value of a picture can represent the order of displaying the corresponding picture.

[0134] Motion Vector (MV): The motion vector can be a 2D vector used for inter - frame prediction or motion compensation. The motion vector can represent the offset between the target image and the reference image.

[0135] – For example, the MV can be represented in the form of (mvx, mvy). mvx can indicate the horizontal component, and mvy can indicate the vertical component.

[0136] – Search range: The search range can be a 2D area where the search for the MV is performed during inter - frame prediction. For example, the size of the search range can be M×N. M and N can be positive integers respectively.

[0137] Motion vector candidate: A motion vector candidate can be a block that is a prediction candidate when a motion vector is predicted or a motion vector of a block that is a prediction candidate.

[0138] – A motion vector candidate can be included in a motion vector candidate list.

[0139] Motion vector candidate list: A motion vector candidate list can be a list configured using one or more motion vector candidates.

[0140] Motion vector candidate index: A motion vector candidate index can be an indicator for indicating a motion vector candidate in a motion vector candidate list. Optionally, a motion vector candidate index can be an index of a motion vector predictor.

[0141] Motion information: Motion information can be information including at least one of a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index, and a motion vector, a reference picture index, and an inter prediction indicator.

[0142] Merge candidate list: A merge candidate list can be a list configured using one or more merge candidates.

[0143] Merge candidate: A merge candidate can be a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-prediction merge candidate, a history-based candidate, a candidate based on the average of two candidates, a zero merge candidate, etc. A merge candidate can include an inter prediction indicator and can include motion information such as prediction type information, a reference picture index for each list, a motion vector, a prediction list utilization flag, and an inter prediction indicator.

[0144] Merge index: A merge index can be an indicator for indicating a merge candidate in a merge candidate list.

[0145] – A merge index can indicate a reconstructed unit for deriving a merge candidate among reconstructed units that are spatially adjacent to a target unit and reconstructed units that are temporally adjacent to the target unit.

[0146] – A merge index can indicate at least one of multiple pieces of motion information of a merge candidate.

[0147] Transform unit: A transform unit can be a basic unit for encoding and / or decoding a residual signal (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit can be partitioned into multiple sub-transform units having smaller sizes. Here, the transform can include one or more of a primary transform and a secondary transform, and the inverse transform can include one or more of a primary inverse transform and a secondary inverse transform.

[0148] Scaling: Scaling may represent the process of multiplying a factor by a level of a transform coefficient.

[0149] – As a result of scaling the level of the transform coefficient, a transform coefficient may be generated. Scaling may also be referred to as "inverse quantization".

[0150] Quantization Parameter (QP): The quantization parameter may be a value used to generate a level of a transform coefficient for a transform coefficient in quantization. Optionally, the quantization parameter may also be a value used to generate a transform coefficient by scaling the level of the transform coefficient in inverse quantization. Optionally, the quantization parameter may be a value mapped to a quantization step size.

[0151] Delta Quantization Parameter: The delta quantization parameter may represent the difference between the quantization parameter of a target unit and the predicted quantization parameter.

[0152] Scanning: Scanning may represent a method of arranging the order of coefficients in a unit, block, or matrix. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array may be referred to as "scanning". Optionally, a method for arranging a 1D array in the form of a 2D array may also be referred to as "scanning" or "inverse scanning".

[0153] Transform Coefficient: The transform coefficient may be a coefficient value generated when a coding device performs a transform. Optionally, the transform coefficient may be a coefficient value generated when a decoding device performs at least one of entropy decoding and inverse quantization.

[0154] – A quantized level or a quantized transform coefficient level generated by applying quantization to a transform coefficient or a residual signal may also be included in the meaning of the term "transform coefficient".

[0155] Quantized Level: The quantized level may be a value generated when a coding device performs quantization on a transform coefficient or a residual signal. Optionally, the quantized level may be a value targeted for inverse quantization when a decoding device performs inverse quantization.

[0156] – A quantized transform coefficient level as a result of transform and quantization may also be included in the meaning of the quantized level.

[0157] Non-Zero Transform Coefficient: The non-zero transform coefficient may be a transform coefficient having a value other than 0, or may be a level of a transform coefficient having a value other than 0. Optionally, the non-zero transform coefficient may be a transform coefficient whose value magnitude is not 0, or may be a level of a transform coefficient whose value magnitude is not 0.

[0158] Quantization Matrix: The quantization matrix may be a matrix used in the quantization process or the inverse quantization process to improve the subjective image quality or objective image quality of an image. The quantization matrix may also be referred to as a "scaling list".

[0159] Quantization matrix coefficient: The quantization matrix coefficient can be each element in the quantization matrix. The quantization matrix coefficient can also be referred to as the "matrix coefficient".

[0160] Default matrix: The default matrix can be a quantization matrix predefined by the encoding device and the decoding device.

[0161] Non-default matrix: The non-default matrix can be a quantization matrix not predefined by the encoding device and the decoding device. The non-default matrix can represent a quantization matrix signaled by the user from the encoding device to the decoding device.

[0162] Most Probable Mode (MPM): The MPM can represent an intra prediction mode that is highly likely to be used for intra prediction of a target block.

[0163] The encoding device and the decoding device can determine one or more MPMs based on the encoding parameters related to the target block and the attributes of the entity related to the target block.

[0164] The encoding device and the decoding device can determine one or more MPMs based on the intra prediction mode of a reference block. The reference block can include multiple reference blocks. The multiple reference blocks can include a spatially adjacent block adjacent to the left of the target block and a spatially adjacent block adjacent to the upper side of the target block. In other words, one or more different MPMs can be determined according to which intra prediction modes have been used for the reference blocks.

[0165] – One or more MPMs can be determined in the same way in both the encoding device and the decoding device. That is, the encoding device and the decoding device can share the same MPM list including one or more MPMs.

[0166] MPM list: The MPM list can be a list including one or more MPMs. The number of one or more MPMs in the MPM list can be predefined.

[0167] MPM indicator: The MPM indicator can indicate the MPM among one or more MPMs in the MPM list that will be used for intra prediction of the target block. For example, the MPM indicator can be an index for the MPM list.

[0168] – Since the MPM list is determined in the same way in both the encoding device and the decoding device, it may not be necessary to send the MPM list itself from the encoding device to the decoding device.

[0169] – The MPM indicator can be signaled from the encoding device to the decoding device. Since the MPM indicator is signaled, the decoding device can determine the MPM among the MPMs in the MPM list that will be used for intra prediction of the target block.

[0170] MPM Usage Indicator: The MPM usage indicator may indicate whether the MPM usage mode will be used for prediction of a target block. The MPM usage mode may be a mode of using an MPM list to determine the MPM to be used for intra prediction of a target block.

[0171] – The MPM usage indicator may be signaled from an encoding device to a decoding device.

[0172] Signaling: "Signaling" may mean that information is sent from an encoding device to a decoding device. Optionally, "signaling" may mean that the encoding device includes the information in a bitstream or a recording medium. The information signaled by the encoding device may be used by the decoding device.

[0173] – The encoding device may generate encoded information by encoding the information to be signaled. The encoded information may be sent from the encoding device to the decoding device. The decoding device may obtain the information by decoding the sent encoded information. Here, encoding may be entropy encoding, and decoding may be entropy decoding.

[0174] Selective Signaling: Information may be selectively signaled. Selective signaling for information may mean that the encoding device selectively includes the information in the bitstream or the recording medium (according to specific conditions). Selective signaling for information may mean that the decoding device selectively extracts the information from the bitstream (according to specific conditions).

[0175] Omission of Signaling: Signaling for information may be omitted. Omission of signaling for information may mean that the encoding device does not include the information in the bitstream or the recording medium (according to specific conditions). Omission of signaling for information may mean that the decoding device does not extract the information from the bitstream (according to specific conditions).

[0176] Statistical Value: A variable, an encoding parameter, a constant, etc. may have a computable value. A statistical value may be a value generated by performing a calculation (operation) on a value of a specified target. For example, the statistical value may indicate one or more of an average value, a weighted average value, a weighted sum, a minimum value, a maximum value, a mode, a median, and an interpolation of values of a specific variable, a specific encoding parameter, a specific constant, etc.

[0177] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied.

[0178] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images (frames). The encoding device 100 may sequentially encode one or more images of the video.

[0179] Refer to Figure 1, the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transformation unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transformation unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0180] The encoding device 100 can perform encoding on a target image using the intra-frame mode and / or the inter-frame mode. In other words, the prediction mode of the target block can be one of the intra-frame mode and the inter-frame mode.

[0181] Hereinafter, the terms "intra-frame mode", "intra-frame prediction mode", "intra-picture mode", and "intra-picture prediction mode" can be used with the same meaning and can be used interchangeably with each other.

[0182] Hereinafter, the terms "inter-frame mode", "inter-frame prediction mode", "inter-picture mode", and "inter-picture prediction mode" can be used with the same meaning and can be used interchangeably with each other.

[0183] Hereinafter, the term "image" can only indicate a partial image or can indicate a block. In addition, the processing of "image" can indicate the sequential processing of multiple blocks.

[0184] In addition, the encoding device 100 can generate a bitstream including encoded information by encoding the target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed through wired and / or wireless transmission media.

[0185] When the intra-frame mode is used as the prediction mode, the switch 115 can switch to the intra-frame mode. When the inter-frame mode is used as the prediction mode, the switch 115 can switch to the inter-frame mode.

[0186] The encoding device 100 can generate a prediction block for the target block. In addition, after the prediction block has been generated, the encoding device 100 can encode the residual block for the target block using the residual between the target block and the prediction block.

[0187] When the prediction mode is the intra-frame mode, the intra-frame prediction unit 120 can use the pixels of the previously encoded / decoded neighboring blocks adjacent to the target block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the target block using the reference samples and can generate prediction samples for the target block via spatial prediction. The prediction samples can represent the samples in the prediction block.

[0188] The inter-frame prediction unit 110 can include a motion prediction unit and a motion compensation unit.

[0189] When the prediction mode is an inter-frame mode, the motion prediction unit may search for the region that best matches the target block in the reference image during the motion prediction process, and may derive a motion vector for the target block and the found region based on the found region. Here, the motion prediction unit may use the search range as the target region for the search.

[0190] The reference image may be stored in the reference picture buffer 190. More specifically, when the encoding and / or decoding of the reference image has been processed, the encoded and / or decoded reference image may be stored in the reference picture buffer 190.

[0191] Since the decoded pictures are stored, the reference picture buffer 190 may be a decoded picture buffer (DPB).

[0192] The motion compensation unit may generate a prediction block for the target block by performing motion compensation using the motion vector. Here, the motion vector may be a two-dimensional (2D) vector for inter-frame prediction. In addition, the motion vector may indicate the offset between the target image and the reference image.

[0193] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit may generate a prediction block by applying an interpolation filter to a partial region of the reference image. To perform inter-frame prediction or motion compensation, it may be determined which one of the skip mode, the merge mode, the advanced motion vector prediction (AMVP) mode, and the current picture reference mode corresponds to the method for predicting and compensating for the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation may be performed according to the mode.

[0194] The subtractor 125 may generate a residual block, where the residual block is the difference between the target block and the prediction block. The residual block may also be referred to as a "residual signal".

[0195] The residual signal may be the difference between the original signal and the prediction signal. Optionally, the residual signal may be a signal generated by transforming or quantifying the difference between the original signal and the prediction signal or a signal generated by transforming and quantifying the difference. The residual block may be the residual signal for a block unit.

[0196] The transform unit 130 may generate transform coefficients by transforming the residual block and may output the generated transform coefficients. Here, the transform coefficients may be the coefficient values generated by transforming the residual block.

[0197] The transform unit 130 may use one of a plurality of predefined transform methods when performing the transform.

[0198] The multiple predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), etc.

[0199] The transformation method for transforming the residual block may be determined according to at least one of the coding parameters for the target block and / or neighboring blocks. For example, the transformation method may be determined based on at least one of the inter prediction mode for the PU, the intra prediction mode for the PU, the size of the TU, and the shape of the TU. Optionally, the transformation information indicating the transformation method may be signaled from the coding device 100 to the decoding device 200.

[0200] When using the transform skip mode, the transform unit 130 may omit the operation of transforming the residual block.

[0201] By performing quantization on the transform coefficients, quantized transform coefficient levels or quantized levels may be generated. Hereinafter, in the embodiments, each of the quantized transform coefficient levels and the quantized levels may also be referred to as "transform coefficients".

[0202] The quantization unit 140 may generate quantized transform coefficient levels (i.e., quantized levels or quantized coefficients) by quantizing the transform coefficients according to quantization parameters. The quantization unit 140 may output the generated quantized transform coefficient levels. In this case, the quantization unit 140 may use a quantization matrix to quantize the transform coefficients.

[0203] The entropy coding unit 150 may generate a bitstream by performing entropy coding based on probability distribution on the values calculated by the quantization unit 140 and / or the coding parameter values calculated during the coding process. The entropy coding unit 150 may output the generated bitstream.

[0204] The entropy coding unit 150 may perform entropy coding on the information about the pixels of the image and the information required for decoding the image. For example, the information required for decoding the image may include syntax elements, etc.

[0205] When applying entropy coding, fewer bits may be allocated to more frequently occurring symbols, and more bits may be allocated to less frequently occurring symbols. Since the symbols are represented by this allocation, the size of the bitstring for the target symbols to be encoded may be reduced. Therefore, the compression performance of video coding may be improved by entropy coding.

[0206] In addition, for entropy encoding, the entropy encoding unit 150 may use an encoding method such as exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC). For example, the entropy encoding unit 150 may use a variable length coding / table (VLC) to perform entropy encoding. For example, the entropy encoding unit 150 may derive a binarization method for a target symbol. In addition, the entropy encoding unit 150 may derive a probability model for the target symbol / binary bit. The entropy encoding unit 150 may perform arithmetic encoding using the derived binarization method, probability model, and context model.

[0207] The entropy encoding unit 150 may transform the coefficients in 2D block form into 1D vector form through a transform coefficient scanning method in order to encode the quantized transform coefficient levels.

[0208] Encoding parameters may be information required for encoding and / or decoding. Encoding parameters may include information encoded by the encoding device 100 and sent from the encoding device 100 to the decoding device, and may also include information that can be derived during the encoding or decoding process. For example, the information sent to the decoding device may include syntax elements.

[0209] Coding parameters may include not only information such as syntax elements (or flags or indices) encoded by an encoding device and signaled by the encoding device to a decoding device, but also information derived during the encoding or decoding process. Additionally, the coding parameters may include information required for encoding or decoding an image. For example, the coding parameters may include at least one value, a combination of the following items, or statistics: the size of a unit / block, the shape / form of a unit / block, the depth of a unit / block, the partitioning information of a unit / block, the partitioning structure of a unit / block, information indicating whether a unit / block is partitioned in a quadtree structure, information indicating whether a unit / block is partitioned in a binary tree structure, the partitioning direction (horizontal or vertical direction) of a binary tree structure, the partitioning form (symmetric partitioning or asymmetric partitioning) of a binary tree structure, information indicating whether a unit / block is partitioned in a ternary tree structure, the partitioning direction (horizontal or vertical direction) of a ternary tree structure, the partitioning form (symmetric partitioning or asymmetric partitioning, etc.) of a ternary tree structure, information indicating whether a unit / block is partitioned in a multi-type tree structure, the combination and direction (horizontal or vertical direction, etc.) of partitioning in a multi-type tree structure, the partitioning form (symmetric partitioning or asymmetric partitioning, etc.) of a multi-type tree structure, the partitioning tree (binary tree or ternary tree) of a multi-type tree form, the prediction type (intra prediction or inter prediction), the intra prediction mode / direction, the intra luminance prediction mode / direction, the intra chrominance prediction mode / direction, the intra partitioning information, the inter partitioning information, the coding block partitioning flag, the prediction block partitioning flag, the transform block partitioning flag, the reference sample filtering method, the reference sample filter taps, the reference sample filter coefficients, the prediction block filtering method, the prediction block filter taps, the prediction block filter coefficients, the prediction block boundary filtering method, the prediction block boundary filter taps, the prediction block boundary filter coefficients, the inter prediction mode, the motion information, the motion vector, the motion vector difference, the reference picture index, the inter prediction direction, the inter prediction indicator, the prediction list utilization flag, the reference picture list, the reference image, the POC, the motion vector prediction factor, the motion vector prediction index, the motion vector prediction candidate, the motion vector candidate list, information indicating whether the merge mode is used, the merge index, the merge candidate, the merge candidate list, information indicating whether the skip mode is used, the type of interpolation filter, the taps of the interpolation filter, the filter coefficients of the interpolation filter, the magnitude of the motion vector, the precision of the motion vector representation, the transform type, the transform size, information indicating whether the first transform is used, information indicating whether an additional (second) transform is used, the first transform selection information (or the first transform index), the second transform selection information (or the second transform index), information indicating the presence or absence of a residual signal, the coding block style, the coding block flag, the quantization parameter, the residual quantization parameter, the quantization matrix, information about the loop filter, information indicating whether the loop filter is applied, the coefficients of the loop filter, the taps of the loop filter, the shape / form of the loop filter, information indicating whether the deblocking filter is applied,Coefficients of the deblocking filter, taps of the deblocking filter, deblocking filter strength, shape / form of the deblocking filter, information indicating whether adaptive sample offset is applied, value of the adaptive sample offset, category of the adaptive sample offset, type of the adaptive sample offset, information indicating whether the adaptive loop filter is applied, coefficients of the adaptive loop filter, taps of the adaptive loop filter, shape / form of the adaptive loop filter, binarization / de-binarization method, context model, context model determination method, context model update method, information indicating whether the normal mode is executed, information indicating whether the bypass mode is executed, valid coefficient flag, last valid coefficient flag, coding flag of the coefficient group, position of the last valid coefficient, information indicating whether the value of the coefficient is greater than 1, information indicating whether the value of the coefficient is greater than 2, information indicating whether the value of the coefficient is greater than 3, remaining coefficient value information, positive / negative sign information, reconstructed luminance sample, reconstructed chrominance sample, context binary bit, bypass binary bit, residual luminance sample, residual chrominance sample, transform coefficient, luminance transform coefficient, chrominance transform coefficient, quantization level, luminance quantization level, chrominance quantization level, transform coefficient level, transform coefficient level scanning method, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of times of motion vector search on the decoding device side, size of the CTU, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, slice identification information, slice type, slice partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luminance signal, information about the chrominance signal, color space of the target block and color space of the residual block. In addition, the above information related to the coding parameters may also be included in the coding parameters. Information used to calculate and / or derive the above coding parameters may also be included in the coding parameters. Information calculated or derived using the above coding parameters may also be included in the coding parameters.

[0210] The first transform selection information may indicate the first transform applied to the target block.

[0211] The second transform selection information may indicate the second transform applied to the target block.

[0212] The residual signal may represent the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be the residual signal for the block.

[0213] Here, transmitting information by a signal may indicate that the encoding device 100 includes the entropy-encoded information generated by performing entropy encoding on a flag or an index in a bitstream, and may indicate that the decoding device 200 obtains information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. Here, the information may include a flag, an index, and the like.

[0214] A signal may mean the information to be transmitted by the signal. Hereinafter, information for an image and a block may be referred to as a "signal". In addition, hereinafter, the terms "information" and "signal" may be used to have the same meaning and may be used interchangeably with each other. For example, a specific signal may be a signal representing a specific block. The original signal may be a signal representing a target block. The prediction signal may be a signal representing a prediction block. The residual signal may be a signal representing a residual block.

[0215] The bitstream may include information based on a specific syntax. The encoding device 100 may generate a bitstream including information according to the specific syntax. The decoding device 200 may obtain information from the bitstream according to the specific syntax.

[0216] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image may be used as a reference image for another image to be subsequently processed. Therefore, the encoding device 100 may reconstruct or decode the encoded target image and store the reconstructed or decoded image in the reference picture buffer 190 as a reference image. For decoding, inverse quantization and inverse transformation of the encoded target image may be performed.

[0217] The quantization level may be inverse-quantized by the inverse quantization unit 160 and may be inverse-transformed by the inverse transformation unit 170. The inverse quantization unit 160 may generate inverse-quantized coefficients by performing an inverse transformation on the quantization level. The inverse transformation unit 170 may generate coefficients that have been inverse-quantized and inverse-transformed by performing an inverse transformation on the inverse-quantized coefficients.

[0218] The coefficients that have been inverse-quantized and inverse-transformed may be added to the prediction block by the adder 175. Adding the coefficients that have been inverse-quantized and inverse-transformed and the prediction block may then generate a reconstructed block. Here, the coefficients that have been inverse-quantized and / or inverse-transformed may represent coefficients on which one or more of inverse quantization and inverse transformation have been performed, and may also represent a reconstructed residual block. Here, the reconstructed block may represent a restored block or a decoded block.

[0219] The reconstructed block can be filtered by the filter unit 180. The filter unit 180 can apply one or more of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), and a non-local filter (NLF) to the reconstructed samples, the reconstructed block, or the reconstructed picture. The filter unit 180 can also be referred to as a "loop filter".

[0220] The deblocking filter can remove block distortion that appears at the boundary between blocks in the reconstructed picture. To determine whether to apply the deblocking filter, the number of columns or rows that are included in the block and that include the pixels based on which it is determined whether to apply the deblocking filter to the target block can be determined.

[0221] When the deblocking filter is applied to the target block, the filter applied can be different according to the intensity of the required deblocking filtering. In other words, among different filters, the filter determined in consideration of the intensity of the deblocking filtering can be applied to the target block. When the deblocking filter is applied to the target block, one or more of a long tap filter, a strong filter, a weak filter, and a Gaussian filter can be applied to the target block according to the intensity of the required deblocking filtering.

[0222] In addition, when performing vertical filtering and horizontal filtering on the target block, the horizontal filtering and the vertical filtering can be performed in parallel.

[0223] SAO can add an appropriate offset to the pixel value to compensate for the coding error. SAO can perform correction on the image to which deblocking has been applied based on pixels, where the correction uses the offset of the difference between the original image and the image to which deblocking has been applied. To perform offset correction on the image, a method of dividing the pixels included in the image into a specific number of regions, determining the regions to which the offset will be applied among the divided regions, and applying the offset to the determined regions can be used, and a method of applying the offset in consideration of the edge information of each pixel can also be used.

[0224] ALF can perform filtering based on the value obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether to apply the adaptive loop filter can be signaled for each CU. Such information can be signaled for the luminance signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Optionally, an ALF with a fixed form can be applied to the block regardless of the characteristics of the block.

[0225] The non-local filter can perform filtering based on a reconstructed block similar to a target block. An area similar to the target block can be selected from the reconstructed picture, and the statistical attributes of the selected similar area can be used to perform filtering of the target block. Information on whether to apply the non-local filter can be signaled for a coding unit (CU). In addition, the shape and filter coefficients of the non-local filter applied to a block can vary according to the block.

[0226] The reconstructed block or reconstructed image filtered by the filter unit 180 can be stored in the reference picture buffer 190 as a reference picture. The reconstructed block filtered by the filter unit 180 can be part of a reference picture. In other words, the reference picture can be a reconstructed picture composed of the reconstructed blocks filtered by the filter unit 180. The stored reference picture can then be used for inter-frame prediction or motion compensation.

[0227] Figure 2 is a block diagram showing the configuration of an embodiment of a decoding apparatus to which the present disclosure is applied.

[0228] The decoding apparatus 200 can be a decoder, a video decoding apparatus, or an image decoding apparatus.

[0229] Referring to Figure 2 , the decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, a switch 245, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0230] The decoding apparatus 200 can receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 can receive a bitstream stored in a computer-readable storage medium, and can receive a bitstream streamed through a wired / wireless transmission medium.

[0231] The decoding apparatus 200 can perform decoding of the bitstream in an intra mode and / or an inter mode. In addition, the decoding apparatus 200 can generate a reconstructed image or a decoded image via decoding, and can output the reconstructed image or the decoded image.

[0232] For example, the operation of switching to the intra mode or the inter mode based on the prediction mode for decoding can be performed by the switch 245. When the prediction mode for decoding is the intra mode, the switch 245 can be operated to switch to the intra mode. When the prediction mode for decoding is the inter mode, the switch 245 can be operated to switch to the inter mode.

[0233] The decoding device 200 can obtain the reconstructed residual block by decoding the input bitstream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as the target to be decoded by adding the reconstructed residual block and the prediction block.

[0234] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols can include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method can be similar to the entropy encoding method described above. That is, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0235] The entropy decoding unit 210 can change the coefficients in the form of a one-dimensional (1D) vector to a 2D block shape by a transform coefficient scanning method to decode the quantized transform coefficient levels.

[0236] For example, the coefficients of the block can be scanned by using a right upper diagonal scan to change the coefficients of the block to a 2D block shape. Optionally, which one of the right upper diagonal scan, vertical scan, and horizontal scan will be used can be determined according to the size of the corresponding block and / or the intra prediction mode.

[0237] The quantized coefficients can be inverse quantized by the inverse quantization unit 220. The inverse quantization unit 220 can generate inverse quantized coefficients by performing inverse quantization on the quantized coefficients. In addition, the inverse quantized coefficients can be inverse transformed by the inverse transform unit 230. The inverse transform unit 230 can generate a reconstructed residual block by performing inverse transform on the inverse quantized coefficients. As a result of performing inverse quantization and inverse transform on the quantized coefficients, a reconstructed residual block can be generated. Here, when generating the reconstructed residual block, the inverse quantization unit 220 can apply a quantization matrix to the quantized coefficients.

[0238] When using the intra mode, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the target block, where the spatial prediction uses the pixel values of previously decoded neighboring blocks adjacent to the target block.

[0239] The inter prediction unit 250 can include a motion compensation unit. Optionally, the inter prediction unit 250 can be designated as a "motion compensation unit".

[0240] When using the inter mode, the motion compensation unit can generate a prediction block by performing motion compensation on the target block, where the motion compensation uses a motion vector and a reference image stored in the reference picture buffer 270.

[0241] The motion compensation unit may apply an interpolation filter to a partial region of a reference image when the motion vector has a value other than an integer, and may generate a prediction block using the reference image to which the interpolation filter has been applied. To perform motion compensation, the motion compensation unit may determine, based on the CU, which one of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the motion compensation method for the PUs included in the CU, and may perform motion compensation according to the determined mode.

[0242] The reconstructed residual block and the prediction block may be added to each other by an adder 255. The adder 255 may generate a reconstructed block by adding the reconstructed residual block and the prediction block.

[0243] The reconstructed block may be filtered by a filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, SAO filter, ALF, and NLF to the reconstructed block or the reconstructed image. The reconstructed image may be a picture including the reconstructed block.

[0244] The filter unit may output the reconstructed image.

[0245] The reconstructed image and / or the reconstructed block filtered by the filter unit 260 may be stored as a reference picture in a reference picture buffer 270. The reconstructed block filtered by the filter unit 260 may be a part of the reference picture. In other words, the reference picture may be an image composed of the reconstructed blocks filtered by the filter unit 260. The stored reference picture may then be used for inter prediction or motion compensation.

[0246] Figure 3 is a diagram schematically showing a partitioning structure of an image when the image is encoded and decoded.

[0247] Figure 3 An example in which a single unit is partitioned into a plurality of sub-units may be schematically shown.

[0248] To partition an image effectively, a coding unit (CU) may be used in encoding and decoding. The term "unit" may be used to commonly specify 1) a block including image samples and 2) syntax elements. For example, "partitioning of a unit" may mean "partitioning of a block corresponding to the unit".

[0249] The CU may be used as a basic unit for image encoding / decoding. The CU may be used as a unit to which one mode selected from an intra mode and an inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it may be determined which one of the intra mode and the inter mode will be applied to each CU.

[0250] In addition, a CU may be a basic unit that performs prediction, transformation, quantization, inverse transformation, dequantization, and encoding / decoding on transform coefficients.

[0251] Referring Figure 3 , image 300 may be sequentially partitioned into units corresponding to largest coding units (LCUs), and the partitioning structure may be determined for each LCU. Here, an LCU may be used to have the same meaning as a coding tree unit (CTU).

[0252] Partitioning a unit may mean partitioning a block corresponding to the unit. Block partitioning information may include depth information regarding the depth of the unit. The depth information may indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned. A single unit may be hierarchically partitioned into multiple sub-units while the single unit has depth information based on a tree structure.

[0253] Each partitioned sub-unit may have depth information. The depth information may be information indicating the size of a CU. The depth information may be stored for each CU.

[0254] Each CU may have depth information. When a CU is partitioned, the depth of the CU generated from the partition may be increased by 1 from the depth of the partitioned CU.

[0255] The partitioning structure may represent the distribution of coding units (CUs) in LCU 310 for efficiently encoding an image. Such a distribution may be determined based on whether a single CU will be partitioned into multiple CUs. The number of CUs generated by partitioning may be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, etc.

[0256] According to the number of CUs generated by partitioning, the horizontal size and vertical size of each CU generated by partitioning may be smaller than the horizontal size and vertical size of the CU before partitioning. For example, the horizontal size and vertical size of each CU generated by partitioning may be half of the horizontal size and vertical size of the CU before partitioning.

[0257] Each partitioned CU may be recursively partitioned into four CUs in the same manner. Compared with at least one of the horizontal size and vertical size of the CU before partitioning, at least one of the horizontal size and vertical size of each partitioned CU may be reduced via recursive partitioning.

[0258] The partitioning of a CU may be recursively performed until a predefined depth or a predefined size.

[0259] For example, the depth of a CU may have a value ranging from 0 to 3. The size range of a CU may be from a size of 64×64 to a size of 8×8 depending on the depth of the CU.

[0260] For example, the depth of the LCU 310 can be 0, and the depth of the smallest coding unit (SCU) can be a predefined maximum depth. Here, as described above, the LCU can be a CU having the maximum coding unit size, and the SCU can be a CU having the smallest coding unit size.

[0261] Partitioning can start at the LCU 310, and whenever the horizontal size and / or vertical size of a CU is reduced by partitioning, the depth of the CU can be incremented by 1.

[0262] For example, for each depth, the non-partitioned CU can have a size of 2N×2N. In addition, in the case where the CU is partitioned, the CU with a size of 2N×2N can be partitioned into four CUs each with a size of N×N. Whenever the depth is incremented by 1, the value of N can be halved.

[0263] Refer to Figure 3 , the LCU with a depth of 0 can have 64×64 pixels or a block of 64×64. 0 can be the minimum depth. The SCU with a depth of 3 can have 8×8 pixels or a block of 8×8. 3 can be the maximum depth. Here, the CU with a 64×64 block as the LCU can be represented by a depth of 0. The CU with a 32×32 block can be represented by a depth of 1. The CU with a 16×16 block can be represented by a depth of 2. The CU with an 8×8 block as the SCU can be represented by a depth of 3.

[0264] Information about whether the corresponding CU is partitioned can be represented by the partitioning information of the CU. The partitioning information can be 1-bit information. All CUs except the SCU can include the partitioning information. For example, the value of the partitioning information of the non-partitioned CU can be a first value. The value of the partitioning information of the partitioned CU can be a second value. When the partitioning information indicates whether the CU is partitioned, the first value can be "0" and the second value can be "1".

[0265] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by partitioning can be half of the horizontal size and vertical size of the CU before partitioning. When the CU with a size of 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned in a quadtree structure. In other words, it can be considered that quadtree partitioning has been applied to the CU.

[0266] For example, when a single CU is partitioned into two CUs, the horizontal or vertical size of each of the two CUs generated by the partitioning can be half of the horizontal or vertical size of the CU before partitioning. When a CU with a size of 32×32 is vertically partitioned into two CUs, the size of each of the two partitioned CUs can be 16×32. When a CU with a size of 32×32 is horizontally partitioned into two CUs, the size of each of the two partitioned CUs can be 32×16. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned in a binary tree structure. In other words, it can be considered that binary tree partitioning has been applied to the CU.

[0267] For example, when a single CU is partitioned (or divided) into three CUs, the original CU before partitioning is partitioned such that its horizontal or vertical size is divided in a ratio of 1:2:1, thus enabling the generation of three sub-CUs. For example, when a CU with a size of 16×32 is horizontally partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have sizes of 16×8, 16×16, and 16×8 respectively in the top-to-bottom direction. For example, when a CU with a size of 32×32 is vertically partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have sizes of 8×32, 16×32, and 8×32 respectively in the left-to-right direction. When a single CU is partitioned into three CUs, it can be considered that the CU is partitioned in a ternary tree form. In other words, it can be considered that ternary tree partitioning has been applied to the CU.

[0268] Both quadtree partitioning and binary tree partitioning are applied to Figure 3 the LCU 310.

[0269] In the encoding device 100, a coding tree unit (CTU) with a size of 64×64 can be partitioned into multiple smaller CUs through a recursive quadtree structure. A single CU can be partitioned into four CUs with the same size. Each CU can be recursively partitioned and can have a quadtree structure.

[0270] Through the recursive partitioning of the CU, an optimal partitioning method that incurs the minimum rate-distortion cost can be selected.

[0271] Figure 3 The coding tree unit (CTU) 320 in

[0272] is an example of a CTU to which all of quadtree partitioning, binary tree partitioning, and ternary tree partitioning are applied. As described above, in order to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning can be applied to the CTU. The partitioning can be applied based on a specific priority.

[0273] For example, quadtree partitioning can be preferentially applied to CTUs. A CU that cannot be further partitioned in the form of a quadtree can correspond to a leaf node of the quadtree. A CU corresponding to a leaf node of the quadtree can be the root node of a binary tree and / or a ternary tree. That is, a CU corresponding to a leaf node of the quadtree can be partitioned in the form of a binary tree or a ternary tree, or may not be further partitioned. In this case, it is prevented that each CU generated by applying binary tree partitioning or ternary tree partitioning to a CU corresponding to a leaf node of the quadtree is again quadtree partitioned, thereby effectively performing the operation of partitioning the block and / or signaling block partition information.

[0274] Quad-partition information can be used to signal the partitioning of a CU corresponding to each node of the quadtree. Quad-partition information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a quadtree. Quad-partition information having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in the form of a quadtree. The quad-partition information can be a flag having a specific length (e.g., 1 bit).

[0275] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, a CU corresponding to a leaf node of the quadtree can be partitioned in the form of a binary tree or a ternary tree. In addition, a CU generated by binary tree partitioning or ternary tree partitioning can be further partitioned in the form of a binary tree or a ternary tree, or may not be further partitioned.

[0276] The partitioning performed when there is no priority between binary tree partitioning and ternary tree partitioning can be referred to as "multi-type tree partitioning". That is, a CU corresponding to a leaf node of the quadtree can be the root node of a multi-type tree. At least one of information indicating whether a CU is partitioned according to a multi-type tree, partitioning direction information, and partitioning tree information can be used to signal the partitioning of a CU corresponding to each node of the multi-type tree. For the partitioning of a CU corresponding to each node of the multi-type tree, information indicating whether the partitioning according to the multi-type tree is performed, partitioning direction information, and partitioning tree information can be signaled sequentially.

[0277] For example, information indicating whether a CU is partitioned according to a multi-type tree and having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a multi-type tree. Information indicating whether a CU is partitioned according to a multi-type tree and having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in the form of a multi-type tree.

[0278] When a CU corresponding to each node of the multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU can further include partitioning direction information.

[0279] The partition direction information can indicate the partition direction of multiple types of tree partitions. The partition direction information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the vertical direction. The partition direction information having a second value (e.g., "0") can indicate that the corresponding CU is partitioned in the horizontal direction.

[0280] When the CU corresponding to each node of the multiple-type tree is partitioned in the form of the multiple-type tree, the corresponding CU can further include partition tree information. The partition tree information can indicate the tree used for the multiple-type tree partition.

[0281] For example, the partition tree information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a binary tree. The partition tree information having a second value (e.g., "0") can indicate that the corresponding CU is partitioned in the form of a ternary tree.

[0282] Here, each of the above information indicating whether the partition of the multiple-type tree is performed, the partition tree information, and the partition direction information can be a flag having a specific length (e.g., 1 bit).

[0283] Entropy encoding and / or entropy decoding can be performed on at least one of the above four-partition information, the information indicating whether the partition of the multiple-type tree is performed, the partition direction information, and the partition tree information. To perform the entropy encoding / decoding of this information, the information of neighboring CUs adjacent to the target CU can be used.

[0284] For example, it can be considered that the probability that the partition form (i.e., partition / non-partition, partition tree, and / or partition direction) of the left CU and / or the upper CU is similar to that of the target CU is very high. Therefore, based on the information of the neighboring CUs, the context information for the entropy encoding and / or entropy decoding of the information for the target CU can be derived. Here, the information of the neighboring CUs can include at least one of the following: 1) the four-partition information of the neighboring CU, 2) the information indicating whether the neighboring CU is partitioned in the form of the multiple-type tree, 3) the partition direction information of the neighboring CU, and 4) the partition tree information of the neighboring CU.

[0285] In another embodiment of the binary tree partition and the ternary tree partition, the binary tree partition can be preferentially performed. That is to say, the binary tree partition can be applied first, and then the CU corresponding to the leaf node of the binary tree can be set as the root node of the ternary tree. In this case, the four-tree partition or the binary tree partition is not performed on the CU corresponding to the node of the ternary tree.

[0286] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning can be a unit for encoding, prediction, and / or transformation. That is, the CU may not be further partitioned for prediction and / or transformation. Therefore, a partitioning structure for partitioning the CU into prediction units (PUs) / or transformation units (TUs), its partitioning information, etc. may not exist in the bitstream.

[0287] However, when the size of the CU as a partitioning unit is larger than the size of the maximum transform block, the CU can be recursively partitioned until the size of the CU becomes less than or equal to the size of the maximum transform block. For example, when the size of the CU is 64×64 and the size of the maximum transform block is 32×32, the CU can be partitioned into four 32×32 blocks to perform transformation. For example, when the size of the CU is 32×64 and the size of the maximum transform block is 32×32, the CU can be partitioned into two 32×32 blocks.

[0288] In this case, information indicating whether the CU is partitioned for transformation may not be signaled separately. Without signaling, it can be determined whether the CU is partitioned via a comparison between the horizontal size (and / or vertical size) of the CU and the horizontal size (and / or vertical size) of the maximum transform block. For example, when the horizontal size of the CU is larger than the horizontal size of the maximum transform block, the CU can be bisected vertically. In addition, when the vertical size of the CU is larger than the vertical size of the maximum transform block, the CU can be bisected horizontally.

[0289] Information about the maximum size and / or minimum size of the CU and information about the maximum size and / or minimum size of the transform block can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be a sequence level, a picture level, a parallel block level, a parallel block group level, or a slice level. For example, the minimum size of the CU can be set to 4×4. For example, the maximum size of the transform block can be set to 64×64. For example, the maximum size of the transform block can be set to 4×4.

[0290] Information about the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or information about the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be a sequence level, a picture level, a slice level, a parallel block group level, or a parallel block level. Information about the minimum size of the quadtree and / or information about the maximum depth of the multi-type tree can be signaled or determined separately at each of the intra-slice level and the inter-slice level.

[0291] Information about the difference between the size of a CTU and the maximum size of a transform block may be signaled or determined at a level higher than the level of the CU. For example, the higher level may be a sequence level, a picture level, a slice level, a parallel block group level, or a parallel block level. Information about the maximum size (i.e., the maximum size of the binary tree) of a CU corresponding to each node of the binary tree may be determined based on the size of the CTU and the information about the difference. The maximum size (i.e., the maximum size of the ternary tree) of a CU corresponding to each node of the ternary tree may have different values according to the type of the slice. For example, the maximum size of the ternary tree at the intra-slice level may be 32×32. For example, the maximum size of the ternary tree at the inter-slice level may be 128×128. For example, the minimum size (i.e., the minimum size of the binary tree) of a CU corresponding to each node of the binary tree and / or the minimum size (i.e., the minimum size of the ternary tree) of a CU corresponding to each node of the ternary tree may be set to the minimum size of the CU.

[0292] In another example, the maximum size of the binary tree and / or the maximum size of the ternary tree may be signaled or determined at the slice level. In addition, the minimum size of the binary tree and / or the minimum size of the ternary tree may be signaled or determined at the slice level.

[0293] Based on the various block sizes and depths described above, the quad-partition information, the information indicating whether the partitioning according to the multi-type tree is performed, the partition tree information, and / or the partition direction information may or may not be present in the bitstream.

[0294] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include the quad-partition information, and the quad-partition information of the CU may be inferred as a second value.

[0295] For example, when the size (horizontal size and vertical size) of a CU corresponding to each node of the multi-type tree is greater than the maximum size (horizontal size and vertical size) of the binary tree and / or the maximum size (horizontal size and vertical size) of the ternary tree, the CU may not be partitioned in the form of the binary tree and / or the ternary tree. By this determination method, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.

[0296] Optionally, when the size (horizontal size and vertical size) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal size and vertical size) of the binary tree, or when the size (horizontal size and vertical size) of the CU is equal to twice the minimum size (horizontal size and vertical size) of the ternary tree, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. In this determination method, the information indicating whether the partition according to the multi-type tree is performed may not be signaled, but may be inferred as a second value. The reason is that when the CU is partitioned in the form of a binary tree and / or a ternary tree, a CU smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.

[0297] Optionally, the binary tree partition or the ternary tree partition may be restricted based on the size of the virtual pipeline data unit (i.e., the size of the pipeline buffer). For example, when the CU is partitioned into sub-CUs that do not fit the size of the pipeline buffer through binary tree partition or ternary tree partition, the binary tree partition or the ternary tree partition may be restricted. The size of the pipeline buffer may be equal to the maximum size of the transform block (e.g., 64×64).

[0298] For example, when the size of the pipeline buffer is 64×64, the following partitions may be restricted.

[0299] - Ternary tree partition for an N×M CU (where N and / or M is 128) - Horizontal binary tree partition for a 128×N CU (where N <= 64) - Vertical binary tree partition for an N×128 CU (where N <= 64) Optionally, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. In this determination method, the information indicating whether the partition according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.

[0300] Optionally, the information indicating whether the partition according to the multi-type tree is performed may be signaled only when at least one of the vertical binary tree partition, the horizontal binary tree partition, the vertical ternary tree partition, and the horizontal ternary tree partition is possible for the CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. In this determination method, the information indicating whether the partition according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.

[0301] Optionally, for a CU corresponding to each node of the multi-type tree, the partitioning direction information may be signaled only when both the vertical binary tree partitioning and the horizontal binary tree partitioning are feasible or only when both the vertical ternary tree partitioning and the horizontal ternary tree partitioning are feasible. Otherwise, the partitioning direction information may not be signaled, but may be inferred as a value indicating the direction in which the CU may be partitioned.

[0302] Optionally, for a CU corresponding to each node of the multi-type tree, the partitioning tree information may be signaled only when both the vertical binary tree partitioning and the vertical ternary tree partitioning are feasible or only when both the horizontal binary tree partitioning and the horizontal ternary tree partitioning are feasible. Otherwise, the partitioning tree information may not be signaled, but may be inferred as a value indicating the tree of the partitioning applicable to the CU.

[0303] Figure 4 is a diagram showing the forms of prediction units that a coding unit can include.

[0304] Among the CUs partitioned from an LCU, the CUs that are no longer partitioned may be divided into one or more prediction units (PUs). This division is also referred to as "partitioning".

[0305] A PU can be a basic unit for prediction. A PU can be encoded and decoded in any one of the skip mode, the inter-frame mode, and the intra-frame mode. A PU can be partitioned into various shapes according to each mode. For example, the target blocks described above with reference to Figure 1 and the target blocks described above with reference to Figure 2 can both be PUs.

[0306] A CU may not be divided into PUs. When a CU is not divided into PUs, the size of the CU and the size of the PU may be equal to each other.

[0307] In the skip mode, there may be no partitioning in the CU. In the skip mode, the 2N×2N mode 410 may be supported without partitioning, where in the 2N×2N mode 410, the size of the PU and the size of the CU are the same.

[0308] In the inter-frame mode, there may be 8 types of partitioning shapes in the CU. For example, in the inter-frame mode, the 2N×2N mode 410, the 2N×N mode 415, the N×2N mode 420, the N×N mode 425, the 2N×nU mode 430, the 2N×nD mode 435, the nL×2N mode 440, and the nR×2N mode 445 may be supported.

[0309] In the intra-frame mode, the 2N×2N mode 410 and the N×N mode 425 may be supported.

[0310] In the 2N×2N mode 410, a PU with a size of 2N×2N can be encoded. The PU with a size of 2N×2N can represent a PU having the same size as the CU. For example, the PU with a size of 2N×2N can have a size of 64×64, 32×32, 16×16, or 8×8.

[0311] In the N×N mode 425, a PU with a size of N×N can be encoded.

[0312] For example, in intra prediction, when the size of the PU is 8×8, the PUs divided into four partitions can be encoded. The size of each partitioned PU can be 4×4.

[0313] When encoding a PU in the intra mode, any one of multiple intra prediction modes can be used to encode the PU. For example, the HEVC technology can provide 35 intra prediction modes, and the PU can be encoded under any one of the 35 intra prediction modes.

[0314] It is possible to determine which one of the 2N×2N mode 410 and the N×N mode 425 will be used to encode the PU based on the rate-distortion cost.

[0315] The encoding device 100 can perform an encoding operation on a PU with a size of 2N×2N. Here, the encoding operation can be an operation of encoding the PU under each of multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the best intra prediction mode for the PU with a size of 2N×2N can be derived. The best intra prediction mode can be the intra prediction mode that incurs the minimum rate-distortion cost when encoding the PU with a size of 2N×2N among the multiple intra prediction modes that can be used by the encoding device 100.

[0316] In addition, the encoding device 100 can sequentially perform an encoding operation on each PU obtained by performing N×N partitioning. Here, the encoding operation can be an operation of encoding the PU under each of multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the best intra prediction mode for the PU with a size of N×N can be derived. The best intra prediction mode can be the intra prediction mode that incurs the minimum rate-distortion cost when encoding the PU with a size of N×N among the multiple intra prediction modes that can be used by the encoding device 100.

[0317] The encoding device 100 can determine which one of the PU with a size of 2N×2N and the PU with a size of N×N will be encoded based on a comparison between the rate-distortion cost of the PU with a size of 2N×2N and the rate-distortion cost of the PU with a size of N×N.

[0318] A single CU can be partitioned into one or more PUs, and a PU can be partitioned into multiple PUs.

[0319] For example, when a single PU is partitioned into four PUs, the horizontal dimension and vertical dimension of each of the four PUs generated by the partitioning can be half of the horizontal dimension and vertical dimension of the PU before partitioning. When a PU with a size of 32×32 is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, it can be considered that the PU has been partitioned in a quadtree structure.

[0320] For example, when a single PU is partitioned into two PUs, the horizontal dimension or vertical dimension of each of the two PUs generated by the partitioning can be half of the horizontal dimension or vertical dimension of the PU before partitioning. When a PU with a size of 32×32 is vertically partitioned into two PUs, the size of each of the two partitioned PUs can be 16×32. When a PU with a size of 32×32 is horizontally partitioned into two PUs, the size of each of the two partitioned PUs can be 32×16. When a single PU is partitioned into two PUs, it can be considered that the PU has been partitioned in a binary tree structure.

[0321] Figure 5 is a diagram showing the form of transform units that can be included in a coding unit.

[0322] A transform unit (TU) can be the basic unit in a CU that is used for processes such as transformation, quantization, inverse transformation, dequantization, entropy coding, and entropy decoding.

[0323] A TU can have a square shape or a rectangular shape. The shape of the TU can be determined based on the size and / or shape of the CU.

[0324] In a CU partitioned from an LCU, a CU that is no longer partitioned into CUs can be partitioned into one or more TUs. Here, the partitioning structure of the TUs can be a quadtree structure. For example, as Figure 5 shown, a single CU 510 can be partitioned one or more times according to the quadtree structure. Through this partitioning, a single CU 510 can be composed of TUs with various sizes.

[0325] It can be considered that when a single CU is divided two or more times, the CU is recursively divided. Through the division, a single CU can be composed of transform units (TUs) with various sizes.

[0326] Optionally, a single CU can be divided into one or more TUs based on the number of vertical lines and / or horizontal lines for partitioning the CU.

[0327] A CU can be divided into a symmetric TU or an asymmetric TU. To divide into an asymmetric TU, information about the size and / or shape of each TU can be signaled from the encoding device 100 to the decoding device 200. Optionally, the size and / or shape of each TU can be derived from the information about the size and / or shape of the CU.

[0328] A CU may not be divided into TUs. When a CU is not divided into TUs, the size of the CU and the size of the TU may be equal to each other.

[0329] A single CU can be partitioned into one or more TUs, and a TU can be partitioned into multiple TUs.

[0330] For example, when a single TU is partitioned into four TUs, the horizontal size and the vertical size of each of the four TUs generated by the partitioning can be half of the horizontal size and the vertical size of the TU before partitioning. When a TU with a size of 32×32 is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, it can be considered that the TU has been partitioned in a quadtree structure.

[0331] For example, when a single TU is partitioned into two TUs, the horizontal size or the vertical size of each of the two TUs generated by the partitioning can be half of the horizontal size or the vertical size of the TU before partitioning. When a TU with a size of 32×32 is vertically partitioned into two TUs, the size of each of the two partitioned TUs can be 16×32. When a TU with a size of 32×32 is horizontally partitioned into two TUs, the size of each of the two partitioned TUs can be 32×16. When a single TU is partitioned into two TUs, it can be considered that the TU has been partitioned in a binary tree structure.

[0332] The CU can be divided in a way different from Figure 5 the way shown in

[0333] For example, a single CU can be divided into three CUs. The horizontal size or the vertical size of the three CUs generated by the division can be 1 / 4, 1 / 2, and 1 / 4 of the horizontal size or the vertical size of the original CU before division, respectively.

[0334] For example, when a CU with a size of 32×32 is vertically divided into three CUs, the sizes of the three CUs generated by the division can be 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is divided into three CUs, it can be considered that the CU is divided in the form of a ternary tree.

[0335] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) can be applied to the partitioning of the CU, and multiple partitioning schemes can be combined and used together for the partitioning of the CU. Here, the case of combining and using multiple partitioning schemes together can be referred to as "compound tree form partitioning".

[0336] Figure 6 Shows the partitioning of a block according to an example.

[0337] In video encoding and / or decoding processing, as Figure 6 shown in, the target block can be partitioned. For example, the target block can be a CU.

[0338] For the partitioning of the target block, an indicator indicating the partitioning information can be signaled from the encoding device 100 to the decoding device 200. The partitioning information can be information indicating how the target block is partitioned.

[0339] The partitioning information can be one or more of a partitioning flag (hereinafter referred to as "split_flag"), a quadtree-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").

[0340] "split_flag" can be a flag indicating whether the block is partitioned. For example, a split_flag value of 1 can indicate that the corresponding block is partitioned. A split_flag value of 0 can indicate that the corresponding block is not partitioned.

[0341] "QB_flag" can be a flag indicating which of the quadtree form and the binary tree form corresponds to the shape of the partitioned block. For example, a QB_flag value of 0 can indicate that the block is partitioned in the quadtree form. A QB_flag value of 1 can indicate that the block is partitioned in the binary tree form. Optionally, a QB_flag value of 0 can indicate that the block is partitioned in the binary tree form. A QB_flag value of 1 can indicate that the block is partitioned in the quadtree form.

[0342] "quadtree_flag" can be a flag indicating whether the block is partitioned in the quadtree form. For example, a quadtree_flag value of 1 can indicate that the block is partitioned in the quadtree form. A quadtree_flag value of 0 can indicate that the block is not partitioned in the quadtree form.

[0343] "binarytree_flag" can be a flag indicating whether a block is divided in a binary tree form. For example, a binarytree_flag value of 1 can indicate that the block is divided in a binary tree form. A binarytree_flag value of 0 can indicate that the block is not divided in a binary tree form.

[0344] "Btype_flag" can be a flag indicating which of the vertical division and the horizontal division corresponds to the division direction when the block is divided in a binary tree form. For example, a Btype_flag value of 0 can indicate that the block is divided in the horizontal direction. A Btype_flag value of 1 can indicate that the block is divided in the vertical direction. Optionally, a Btype_flag value of 0 can indicate that the block is divided in the vertical direction. A Btype_flag value of 1 can indicate that the block is divided in the horizontal direction.

[0345] For example, the division information of the block in Figure 6 can be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag, as shown in Table 1 below.

[0346] Table 1

[0347] For example, the division information of the block in Figure 6 can be derived by signaling at least one of split_flag, QB_flag, and Btype_flag, as shown in Table 2 below.

[0348] Table 2

[0349] The division method can be limited to a quadtree or a binary tree according to the size and / or shape of the block. When this limitation is applied, split_flag can be a flag indicating whether the block is divided in a quadtree form or a flag indicating whether the block is divided in a binary tree form. The size and shape of the block can be derived from the depth information of the block, and the depth information can be signaled from the encoding device 100 to the decoding device 200.

[0350] When the size of the block falls within a specific range, division can only be performed in a quadtree form. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can only be divided in a quadtree form.

[0351] Information indicating the maximum block size and the minimum block size that can be partitioned only in the form of a quadtree can be signaled from the encoding device 100 to the decoding device 200 via a bitstream. In addition, this information can be signaled for at least one of units such as video, sequence, picture, parameter, parallel block group, and slice (or segment).

[0352] Optionally, the maximum block size and / or the minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the size of a block is greater than 64×64 and less than 256×256, partitioning only in the form of a quadtree is possible. In this case, split_flag can be a flag indicating whether to perform partitioning in the form of a quadtree.

[0353] When the size of a block is greater than the maximum size of a transform block, partitioning only in the form of a quadtree is possible. Here, the sub-blocks generated by partitioning can be at least one of a CU and a TU.

[0354] In this case, split_flag can be a flag indicating whether the CU is partitioned in the form of a quadtree.

[0355] When the size of a block falls within a specific range, partitioning only in the form of a binary tree or a ternary tree is possible. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can be partitioned only in the form of a binary tree or a ternary tree.

[0356] Information indicating the maximum block size and / or the minimum block size that can be partitioned only in the form of a binary tree or in the form of a ternary tree can be signaled from the encoding device 100 to the decoding device 200 via a bitstream. In addition, this information can be signaled for at least one of units such as sequence, picture, and slice (or segment).

[0357] Optionally, the maximum block size and / or the minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the size of a block is greater than 8×8 and less than 16×16, partitioning only in the form of a binary tree is possible. In this case, split_flag can be a flag indicating whether to perform partitioning in the form of a binary tree or a ternary tree.

[0358] The above description of partitioning in the form of a quadtree can be equally applied to the form of a binary tree and / or the form of a ternary tree.

[0359] The partitioning of a block may be restricted by a previous partitioning. For example, when a block is partitioned in a specific binary tree form and a plurality of sub-blocks are generated from the partitioning, each sub-block may be further partitioned only in a specific tree form. Here, the specific tree form may be at least one of a binary tree form, a ternary tree form, and a quaternary tree form.

[0360] When the horizontal size or the vertical size of a partitioned block is a size that cannot be further divided, the above indicator may not be signaled.

[0361] Figure 7 is a diagram for explaining an embodiment of intra prediction processing.

[0362] From Figure 7 The arrow extending radially from the center of the diagram in indicates the prediction direction of the intra prediction mode. In addition, the numbers appearing near the arrow indicate examples of mode values assigned to the intra prediction mode or the prediction direction of the intra prediction mode.

[0363] In Figure 7 the number 0 may represent a planar mode as a non-directional intra prediction mode. The number 1 may represent a DC mode as a non-directional intra prediction mode.

[0364] Intra coding and / or decoding may be performed using reference samples of neighboring blocks of a target block. The neighboring blocks may be reconstructed neighboring blocks. The reference samples may represent neighboring samples.

[0365] For example, intra coding and / or decoding may be performed using the values of the reference samples included in the reconstructed neighboring blocks or the coding parameters of the reconstructed neighboring blocks.

[0366] The encoding device 100 and / or the decoding device 200 may generate a prediction block by performing intra prediction on a target block based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may generate a prediction block for the target block by performing intra prediction based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.

[0367] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.

[0368] The unit of the prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The size N×N may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.

[0369] Optionally, the prediction block may be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc. or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.

[0370] Intra prediction may be performed considering the intra prediction mode for the target block. The number of intra prediction modes that the target block may have may be a predefined fixed value and may be a value determined differently according to the attributes of the prediction block. For example, the attributes of the prediction block may include the size of the prediction block, the type of the prediction block, etc. In addition, the attributes of the prediction block may indicate the coding parameters for the prediction block.

[0371] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to N. Optionally, the number of intra prediction modes may be, for example, 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95.

[0372] The intra prediction mode may be a non-directional mode or a directional mode.

[0373] For example, the intra prediction mode may include two non-directional modes corresponding to the numbers 0 to 66 shown in Figure 7 and 65 directional modes.

[0374] For example, in the case of using a specific intra prediction method, the intra prediction mode may include two non-directional modes corresponding to the numbers -14 to 80 shown in Figure 7 and 93 directional modes.

[0375] The two non-directional modes may include a DC mode and a planar mode.

[0376] The directional mode may be a prediction mode having a specific direction or a specific angle. The directional mode may also be referred to as an "angle mode".

[0377] The intra prediction mode may be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "(mode) number of the intra prediction mode", "(mode) value of the intra prediction mode", "(mode) angle of the intra prediction mode", and "(mode) direction of the intra prediction mode" may be used with the same meaning and may be used interchangeably with each other.

[0378] The number of intra prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.

[0379] The number of intra prediction modes can be fixed to M, regardless of the size of the block and / or color component. For example, the number of intra prediction modes can be fixed to either 35 or 67, regardless of the size of the block.

[0380] Optionally, the number of intra prediction modes can vary according to the shape, size, and / or type of color component of the block.

[0381] For example, in Figure 7 the directional prediction mode shown by the dashed line can be applied only to the prediction for non-square blocks.

[0382] For example, the larger the size of the block, the more intra prediction modes there are. Optionally, the larger the size of the block, the fewer intra prediction modes there are. When the size of the block is 4×4 or 8×8, the number of intra prediction modes can be 67. When the size of the block is 16×16, the number of intra prediction modes can be 35. When the size of the block is 32×32, the number of intra prediction modes can be 19. When the size of the block is 64×64, the number of intra prediction modes can be 7.

[0383] For example, the number of intra prediction modes can vary according to whether the color component is a luminance signal or a chrominance signal. Optionally, the number of intra prediction modes corresponding to the luminance component block can be greater than the number of intra prediction modes corresponding to the chrominance component block.

[0384] For example, in the vertical mode with a mode value of 50, prediction can be performed in the vertical direction based on the pixel values of the reference samples. For example, in the horizontal mode with a mode value of 18, prediction can be performed in the horizontal direction based on the pixel values of the reference samples.

[0385] Even in a directional mode other than the above modes, the encoding device 100 and the decoding device 200 can still perform intra prediction on the target unit using the reference samples according to the angle corresponding to the directional mode.

[0386] The intra prediction mode located to the right of the vertical mode can be referred to as the "vertical - right mode". The intra prediction mode located below the horizontal mode can be referred to as the "horizontal - below mode". For example, in Figure 7 the intra prediction mode with a mode value being one of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 can be the vertical - right mode. The intra prediction mode with a mode value being one of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be the horizontal - below mode.

[0387] The non-directional mode may include a DC mode and a planar mode. For example, the value of the DC mode may be 1. The value of the planar mode may be 0.

[0388] The directional mode may include an angular mode. Among the intra-frame prediction modes, the remaining modes other than the DC mode and the planar mode may be the directional mode.

[0389] When the intra-frame prediction mode is the DC mode, a prediction block may be generated based on the average value of the pixel values of a plurality of reference pixels. For example, the pixel value of the prediction block may be determined based on the average value of the pixel values of a plurality of reference pixels.

[0390] The number of the above-described intra-frame prediction modes and the mode values of each intra-frame prediction mode are merely exemplary. The number of the above-described intra-frame prediction modes and the mode values of each intra-frame prediction mode may be defined differently according to embodiments, implementations, and / or requirements.

[0391] In order to perform intra-frame prediction on a target block, a step of checking whether the samples included in the reconstructed neighboring blocks can be used as reference samples for the target block may be performed. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, the value generated by interpolation and / or copying of at least one of the sample values included in the reconstructed neighboring blocks may replace the sample value of the sample that cannot be used as a reference sample. When the value generated by copying and / or interpolation replaces the sample value of the existing sample, the sample may be used as a reference sample for the target block.

[0392] When intra-frame prediction is used, a filter may be applied to at least one of the reference samples and the prediction samples based on at least one of the size of the target block and the intra-frame prediction mode.

[0393] The type of the filter to be applied to at least one of the reference samples and the prediction samples may be different according to at least one of the intra-frame prediction mode of the target block, the size of the target block, and the shape of the target block. The type of the filter may be classified according to one or more of the length of the filter taps, the value of the filter coefficients, and the filter strength. The length of the filter taps may represent the number of the filter taps. In addition, the number of the filter taps may represent the length of the filter.

[0394] When the intra-frame prediction mode is the planar mode, when generating the prediction block of the target block, the sample value of the prediction target block may be generated using the weighted sum of the upper reference sample of the target block, the left reference sample of the target block, the upper-right reference sample of the target block, and the lower-left reference sample of the target block according to the position of the prediction target sample in the prediction block.

[0395] When the intra prediction mode is the DC mode, the average value of the reference samples above the target block and the reference samples to the left of the target block may be used when generating the prediction block of the target block. In addition, filtering using the values of the reference samples may be performed on a specific row or a specific column in the target block. The specific row may be one or more upper rows adjacent to the reference samples. The specific column may be one or more left columns adjacent to the reference samples.

[0396] When the intra prediction mode is the direction mode, the upper reference sample, the left reference sample, the upper-right reference sample, and / or the lower-left reference sample of the target block may be used to generate the prediction block.

[0397] To generate the above prediction samples, interpolation based on real numbers may be performed.

[0398] The intra prediction mode of the target block may be predicted from the intra prediction modes of neighboring blocks adjacent to the target block, and the information used for prediction may be entropy-coded / entropy-decoded.

[0399] For example, when the intra prediction modes of the target block and the neighboring block are the same, a predefined flag may be used to signal that the intra prediction modes of the target block and the neighboring block are the same.

[0400] For example, an indicator indicating the intra prediction mode that is the same as the intra prediction mode of the target block among the intra prediction modes of a plurality of neighboring blocks may be signaled.

[0401] When the intra prediction modes of the target block and the neighboring block are different from each other, the information about the intra prediction mode of the target block may be encoded and / or decoded using entropy coding and / or entropy decoding.

[0402] Figure 8 is a diagram showing the reference samples used in the intra prediction process.

[0403] The reconstructed reference samples for intra prediction of the target block may include a lower-left reference sample, a left reference sample, an upper-left reference sample, an upper reference sample, and an upper-right reference sample.

[0404] For example, the left reference sample may represent a reconstructed reference pixel adjacent to the left side of the target block. The upper reference sample may represent a reconstructed reference pixel adjacent to the top of the target block. The upper-left reference sample may represent a reconstructed reference pixel located at the upper left corner of the target block. The lower-left reference sample may represent a reference sample located below the left sample line among the samples on the same line as the left sample line composed of the left reference samples. The upper-right reference sample may represent a reference sample located to the right of the upper sample line among the samples on the same line as the upper sample line composed of the upper reference samples.

[0405] When the size of the target block is N×N, the number of the lower-left reference sample points, left reference sample points, upper reference sample points, and upper-right reference sample points can all be N.

[0406] By performing intra prediction on the target block, a prediction block can be generated. The process of generating the prediction block can include determining the values of the pixels in the prediction block. The sizes of the target block and the prediction block can be the same.

[0407] The reference sample points used for intra prediction of the target block can be changed according to the intra prediction mode of the target block. The direction of the intra prediction mode can represent the dependency relationship between the reference sample points and the pixels of the prediction block. For example, the value of the specified reference sample point can be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample point and the one or more specified pixels in the prediction block can be the sample points and pixels located on a straight line along the direction of the intra prediction mode. In other words, the value of the specified reference sample point can be copied as the value of the pixels located in the direction opposite to the direction of the intra prediction mode. Optionally, the value of the pixels in the prediction block can be the value of the reference sample point located in the direction of the intra prediction mode relative to the position of the pixel.

[0408] In an example, when the intra prediction mode of the target block is the vertical mode, the upper reference sample points can be used for intra prediction. When the intra prediction mode is the vertical mode, the value of the pixels in the prediction block can be the value of the reference sample points vertically located above the position of the pixel. Therefore, the upper reference sample points adjacent to the top of the target block can be used for intra prediction. In addition, the values of the pixels in a row of the prediction block can be the same as the values of the pixels of the upper reference sample points.

[0409] In an example, when the intra prediction mode of the target block is the horizontal mode, the left reference sample points can be used for intra prediction. When the intra prediction mode is the horizontal mode, the value of the pixels in the prediction block can be the value of the reference sample points horizontally located to the left of the position of the pixel. Therefore, the left reference sample points adjacent to the left side of the target block can be used for intra prediction. In addition, the values of the pixels in a column of the prediction block can be the same as the values of the pixels of the left reference sample points.

[0410] In an example, when the mode value of the intra prediction mode of the current block is 34, at least some of the left reference sample points, the upper-left reference sample points, and at least some of the upper reference sample points can be used for intra prediction. When the mode value of the intra prediction mode is 34, the value of the pixels in the prediction block can be the value of the reference sample points diagonally located at the upper left corner of the pixel.

[0411] In addition, in the case of the intra prediction mode with the mode value in the range from 52 to 66, at least a part of the upper-right reference sample points can be used for intra prediction.

[0412] In addition, in the case of an intra prediction mode with a mode value in the range from 2 to 17, at least a part of the lower left reference samples can be used for intra prediction.

[0413] In addition, in the case of an intra prediction mode with a mode value in the range from 19 to 49, the upper left reference sample can be used for intra prediction.

[0414] The number of reference samples used to determine the pixel value of a pixel in the prediction block can be 1 or 2 or more.

[0415] As described above, the pixel value of a pixel in the prediction block can be determined according to the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of a reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.

[0416] When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are not integer positions, an interpolated reference sample based on the two reference samples closest to the position of the reference sample can be generated. The value of the interpolated reference sample can be used to determine the pixel value of the pixel in the prediction block. In other words, when the position of the pixel in the prediction block and the position of the reference sample indicated by the direction of the intra prediction mode indicate a position between two reference samples, an interpolation based on the values of these two samples can be generated.

[0417] The prediction block generated through prediction can be different from the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.

[0418] Hereinafter, the terms "difference", "error" and "residual" can be used with the same meaning and can be used interchangeably with each other.

[0419] For example, in the case of directional intra prediction, the longer the distance between the pixels of the prediction block and the reference sample, the greater the possible prediction error. Such a prediction error can lead to discontinuity between the generated prediction block and neighboring blocks.

[0420] To reduce the prediction error, a filtering operation for the prediction block can be used. The filtering operation can be configured to adaptively apply a filter to the area in the prediction block that is considered to have a larger prediction error. For example, the area considered to have a larger prediction error can be the boundary of the prediction block. In addition, the area in the prediction block considered to have a larger prediction error can be different according to the intra prediction mode, and the characteristics of the filter can also be different according to the intra prediction mode.

[0421] AsFigure 8 As shown, for intra prediction of a target block, at least one of reference lines 0 to 3 can be used.

[0422] In Figure 8 each of the reference lines can indicate a reference sample line including one or more reference samples. When the number of the reference line is small, it can indicate a reference sample line closer to the target block.

[0423] The samples in segment A and segment F can be obtained by padding instead of from the reconstructed neighboring blocks, where the padding uses the samples in segment B and segment E that are closest to the target block.

[0424] Index information indicating the reference sample line to be used for intra prediction of the target block can be signaled. The index information can indicate the reference sample line among the multiple reference sample lines to be used for intra prediction of the target block. For example, the index information can have a value corresponding to any one of 0 to 3.

[0425] When the upper boundary of the target block is the boundary of the CTU, only reference line 0 can be available. Therefore, in this case, the index information may not be signaled. When additional reference sample lines other than reference line 0 are used, filtering of the prediction block described later may not be performed.

[0426] In the case of inter-color intra prediction, a prediction block of the target block of the second color component can be generated based on the corresponding reconstructed block of the first color component.

[0427] For example, the first color component can be a luminance component and the second color component can be a chrominance component.

[0428] To perform inter-color intra prediction, parameters of a linear model between the first color component and the second color component can be derived based on a template.

[0429] The template can include reference samples (upper reference samples) above the target block and / or reference samples (left reference samples) to the left of the target block, and can include upper reference samples and / or left reference samples of the reconstructed blocks of the first color component corresponding to the reference samples.

[0430] For example, the following values can be used to derive the parameters of the linear model: 1) the value of the sample of the first color component having the maximum value among the samples in the template, 2) the value of the sample of the second color component corresponding to the sample of the first color component, 3) the value of the sample of the first color component having the minimum value among the samples in the template, and 4) the value of the sample of the second color component corresponding to the sample of the first color component.

[0431] When deriving the parameters of the linear model, a predicted block of the target block can be generated by applying the corresponding reconstructed block to the linear model.

[0432] According to the image format, subsampling can be performed on the samples adjacent to the reconstructed block of the first color component and the corresponding reconstructed block of the first color component. For example, when one sample of the second color component corresponds to four samples of the first color component, one corresponding sample can be calculated by performing subsampling on the four samples of the first color component. When performing subsampling, the derivation of the parameters of the linear model and the inter-color intra prediction can be performed based on the corresponding samples to be subsampled.

[0433] Information about whether to perform inter-color intra prediction and / or the range of the template can be signaled in the intra prediction mode.

[0434] The target block can be partitioned into two or four sub-blocks in the horizontal direction and / or the vertical direction.

[0435] The sub-blocks generated by the partitioning can be reconstructed sequentially. That is, when performing intra prediction on each sub-block, a sub-predicted block of the sub-block can be generated. In addition, when performing inverse quantization and / or inverse transformation on each sub-block, a sub-residual block for the corresponding sub-block can be generated. The reconstructed sub-block can be generated by adding the sub-predicted block and the sub-residual block. The reconstructed sub-block can be used as a reference sample for the intra prediction of the sub-block with the next priority.

[0436] The sub-block can be a block including a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block can be partitioned into two sub-blocks. In addition, when the target block is a 4×4 block, the target block cannot be partitioned into sub-blocks. When the target block has another size, the target block can be partitioned into four sub-blocks.

[0437] Information about whether to perform intra prediction based on these sub-blocks and / or information about the partitioning direction (horizontal direction or vertical direction) can be signaled.

[0438] This intra prediction based on sub-blocks can be restricted so that it is only performed when the reference sample line 0 is used. When performing intra prediction based on sub-blocks, the filtering of the predicted block described below may not be performed.

[0439] The final predicted block can be generated by performing filtering on the predicted block generated by intra prediction.

[0440] Filtering can be performed by applying specific weights to the filtering target samples, the left reference sample, the upper reference sample, and / or the upper left reference sample that are the targets to be filtered.

[0441] The weights for filtering and / or the reference samples (e.g., the range of reference samples, the positions of reference samples, etc.) can be determined based on at least one of the block size, the intra prediction mode, and the positions of the filtered target samples in the prediction block.

[0442] For example, filtering can be performed only in specific intra prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).

[0443] The adjacent diagonal mode can be a mode having a number obtained by adding k to the number of the diagonal mode, and can be a mode having a number obtained by subtracting k from the number of the diagonal mode. In other words, the number of the adjacent diagonal mode can be the sum of the number of the diagonal mode and k, or can be the difference between the number of the diagonal mode and k. For example, k can be a positive integer of 8 or less.

[0444] The intra prediction mode of the target block can be derived using the intra prediction modes of neighboring blocks existing near the target block, and such derived intra prediction mode can be entropy-coded and / or entropy-decoded.

[0445] For example, when the intra prediction mode of the target block is the same as the intra prediction mode of a neighboring block, specific flag information can be used to signal information indicating that the intra prediction mode of the target block is the same as the intra prediction mode of the neighboring block.

[0446] In addition, for example, indicator information of neighboring blocks in which the intra prediction mode among the intra prediction modes of multiple neighboring blocks is the same as the intra prediction mode of the target block can be signaled.

[0447] For example, when the intra prediction mode of the target block is different from the intra prediction mode of a neighboring block, entropy coding and / or entropy decoding can be performed on the information regarding the intra prediction mode of the target block by performing entropy coding and / or entropy decoding based on the intra prediction mode of the neighboring block.

[0448] Figure 9 is a diagram for explaining an embodiment of the inter prediction process.

[0449] Figure 9 The rectangle shown in can represent an image (or picture). In addition, in Figure 9 the arrow can represent the prediction direction. The arrow pointing from the first picture to the second picture indicates that the second picture references the first picture. That is, each image can be encoded and / or decoded according to the prediction direction.

[0450] Images can be classified into intra pictures (I pictures), forward predicted pictures or predictive coded pictures (P pictures), and bi - directionally predicted pictures or bi - directionally predictive coded pictures (B pictures) according to the coding type. Each picture can be coded and / or decoded according to the coding type of each picture.

[0451] When the target image to be coded is an I picture, the target image can be coded using the data contained in the image itself without performing inter - frame prediction with reference to other images. For example, an I picture can be coded only via intra - frame prediction.

[0452] When the target image is a P picture, the target image can be coded via inter - frame prediction using a reference picture existing in one direction. Here, the one direction can be the forward direction or the backward direction.

[0453] When the target image is a B picture, the image can be coded via inter - frame prediction using reference pictures existing in two directions, or can be coded via inter - frame prediction using a reference picture existing in one of the forward direction and the backward direction. Here, the two directions can be the forward direction and the backward direction.

[0454] P pictures and B pictures coded and / or decoded using reference pictures can be regarded as images using inter - frame prediction.

[0455] Hereinafter, inter - frame prediction in the inter - frame mode according to an embodiment will be described in detail.

[0456] Inter - frame prediction or motion compensation can be performed using a reference image and motion information.

[0457] In the inter - frame mode, the encoding device 100 can perform inter - frame prediction and / or motion compensation on a target block. The decoding device 200 can perform inter - frame prediction and / or motion compensation corresponding to the inter - frame prediction and / or motion compensation performed by the encoding device 100 on the target block.

[0458] The motion information of the target block can be separately derived by the encoding device 100 and the decoding device 200 during inter - frame prediction. The motion information can be derived using the motion information of the reconstructed neighboring blocks, the motion information of the col blocks, and / or the motion information of the blocks adjacent to the col blocks.

[0459] For example, the encoding device 100 or the decoding device 200 can perform prediction and / or motion compensation by using the motion information of spatial candidates and / or temporal candidates as the motion information of the target block. The target block can represent a PU and / or a PU partition.

[0460] Spatial candidates can be reconstructed blocks that are spatially adjacent to the target block.

[0461] The temporal candidate may be a reconstructed block corresponding to the target block in a previously reconstructed co-located picture (col picture).

[0462] In inter prediction, the encoding device 100 and the decoding device 200 may improve the encoding efficiency and the decoding efficiency by using the motion information of the spatial candidate and / or the temporal candidate. The motion information of the spatial candidate may be referred to as "spatial motion information". The motion information of the temporal candidate may be referred to as "temporal motion information".

[0463] Hereinafter, the motion information of the spatial candidate may be the motion information of the PU including the spatial candidate. The motion information of the temporal candidate may be the motion information of the PU including the temporal candidate. The motion information of the candidate block may be the motion information of the PU including the candidate block.

[0464] Inter prediction may be performed using a reference picture.

[0465] The reference picture may be at least one of a picture before the target picture and a picture after the target picture. The reference picture may be an image for prediction of the target block.

[0466] In inter prediction, a reference picture index (or refIdx) for indicating the reference picture, a motion vector to be described later, etc. may be used to specify a region in the reference picture. Here, the region specified in the reference picture may indicate a reference block.

[0467] Inter prediction may select a reference picture, and may also select a reference block corresponding to the target block from the reference picture. In addition, inter prediction may use the selected reference block to generate a prediction block for the target block.

[0468] Each of the encoding device 100 and the decoding device 200 may derive motion information during inter prediction.

[0469] The spatial candidate may be a block that 1) exists in the target picture, 2) has been previously reconstructed via encoding and / or decoding, and 3) is adjacent to the target block or located at a corner of the target block. Here, "a block located at a corner of the target block" may be a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block. In addition, "a block located at a corner of the target block" may have the same meaning as "a block adjacent to a corner of the target block". The meaning of "a block located at a corner of the target block" may be included in the meaning of "a block adjacent to the target block".

[0470] For example, the spatial candidate may be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower left corner of the target block, a reconstructed block located at the upper right corner of the target block, or a reconstructed block located at the upper left corner of the target block.

[0471] Each of the encoding device 100 and the decoding device 200 can identify a block at a position in the col picture that spatially corresponds to the target block. The position of the target block in the target picture and the position of the identified block in the col picture can correspond to each other.

[0472] Each of the encoding device 100 and the decoding device 200 can determine a col block at a predefined relative position for the identified block as a temporal candidate. The predefined relative position can be a position inside and / or outside the identified block.

[0473] For example, the col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at the coordinates (xP + nPSW, yP + nPSH). The second col block can be the block located at the coordinates (xP + (nPSW >> 1), yP + (nPSH >> 1)). When the first col block is not available, the second col block can be selectively used.

[0474] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. In addition, the motion vector of the motion information of the temporal candidate stored in the list can be a scaled motion vector.

[0475] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first temporal distance to the second temporal distance. The first temporal distance can be the distance between the reference picture and the target picture of the target block. The second temporal distance can be the distance between the reference picture and the col picture of the col block.

[0476] The scheme for deriving motion information can be changed according to the inter-frame prediction mode of the target block. For example, as the inter-frame prediction mode applied to inter-frame prediction, there can be an advanced motion vector predictor (AMVP) mode, a merge mode, a skip mode, a merge mode with a motion vector difference, a sub-block merge mode, a triangular partitioning mode, an inter-frame / intra-frame combined prediction mode, an affine inter-frame mode, a current picture reference mode, etc. The merge mode can also be referred to as the "motion merge mode". Each mode will be described in detail below.

[0477] 1) AMVP mode When using the AMVP mode, the encoding device 100 may search for similar blocks in the neighboring area of the target block. The encoding device 100 may obtain a predicted block by performing prediction on the target block by using the motion information of the found similar blocks. The encoding device 100 may encode a residual block that is the difference between the target block and the predicted block.

[0478] 1-1) Create a list of predicted motion vector candidates When the AMVP mode is used as a prediction mode, each of the encoding device 100 and the decoding device 200 may use a spatial candidate motion vector, a temporal candidate motion vector, and a zero vector to create a list of predicted motion vector candidates. The list of predicted motion vector candidates may include one or more predicted motion vector candidates. At least one of the spatial candidate motion vector, the temporal candidate motion vector, and the zero vector may be determined and used as a predicted motion vector candidate.

[0479] Hereinafter, the terms "predicted motion vector (candidate)" and "motion vector (candidate)" may be used to have the same meaning and may be used interchangeably with each other.

[0480] Hereinafter, the terms "predicted motion vector candidate" and "AMVP candidate" may be used to have the same meaning and may be used interchangeably with each other.

[0481] Hereinafter, the terms "list of predicted motion vector candidates" and "list of AMVP candidates" may be used to have the same meaning and may be used interchangeably with each other.

[0482] Spatial candidates may include reconstructed spatially neighboring blocks. In other words, the motion vector of the reconstructed neighboring block may be referred to as a "spatial predicted motion vector candidate".

[0483] Temporal candidates may include col blocks and blocks adjacent to the col blocks. In other words, the motion vector of the col block or the motion vector of the block adjacent to the col block may be referred to as a "temporal predicted motion vector candidate".

[0484] The zero vector may be a (0,0) motion vector.

[0485] A predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. In addition, in the encoding device 100, each predicted motion vector candidate may be an initial search position for the motion vector.

[0486] 1-2) Search for motion vectors using the list of predicted motion vector candidates The encoding device 100 can determine a motion vector to be used for encoding a target block within a search range by using a list of predicted motion vector candidates. In addition, the encoding device 100 can determine, among the predicted motion vector candidates present in the predicted motion vector candidate list, a predicted motion vector candidate to be used as the predicted motion vector of the target block.

[0487] The motion vector to be used for encoding the target block can be a motion vector that can be encoded at the minimum cost.

[0488] In addition, the encoding device 100 can determine whether to use the AMVP mode to encode the target block.

[0489] 1-3) Transmission of inter-frame prediction information The encoding device 100 can generate a bitstream including the inter-frame prediction information required for inter-frame prediction. The decoding device 200 can perform inter-frame prediction on the target block by using the inter-frame prediction information of the bitstream.

[0490] The inter-frame prediction information can include 1) mode information indicating whether the AMVP mode is used, 2) a predicted motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.

[0491] Hereinafter, the terms "predicted motion vector index" and "AMVP index" can be used with the same meaning and can be used interchangeably with each other.

[0492] In addition, the inter-frame prediction information can include a residual signal.

[0493] When the mode information indicates that the AMVP mode is used, the decoding device 200 can obtain the predicted motion vector index, MVD, reference direction, and reference picture index from the bitstream by entropy decoding.

[0494] The predicted motion vector index can indicate a predicted motion vector candidate to be used for predicting the target block among the predicted motion vector candidates included in the predicted motion vector candidate list.

[0495] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information The decoding device 200 can use the predicted motion vector candidate list to derive predicted motion vector candidates, and can determine the motion information of the target block based on the derived predicted motion vector candidates.

[0496] The decoding device 200 can use the predicted motion vector index to determine a motion vector candidate for the target block among the predicted motion vector candidates included in the predicted motion vector candidate list. The decoding device 200 can select, as the predicted motion vector of the target block, the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list.

[0497] The encoding device 100 can generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and can generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index can be signaled from the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 can extract the entropy-coded predicted motion vector index from the bitstream, and can obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.

[0498] The motion vector actually used for inter prediction of the target block may not match the predicted motion vector. To indicate the difference between the motion vector actually used for inter prediction of the target block and the predicted motion vector, an MVD can be used. The encoding device 100 can derive a predicted motion vector similar to the motion vector actually used for inter prediction of the target block so as to use the smallest possible MVD.

[0499] The motion vector difference (MVD) can be the difference between the motion vector of the target block and the predicted motion vector. The encoding device 100 can calculate the MVD, and can generate an entropy-coded MVD by applying entropy coding to the MVD. The encoding device 100 can generate a bitstream including the entropy-coded MVD.

[0500] The MVD can be sent from the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 can extract the entropy-coded MVD from the bitstream, and can obtain the MVD by applying entropy decoding to the entropy-coded MVD.

[0501] The decoding device 200 can derive the motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 can be the sum of the MVD and the motion vector candidate.

[0502] In addition, the encoding device 100 can generate an entropy-coded MVD resolution information by applying entropy coding to the calculated MVD resolution information, and can generate a bitstream including the entropy-coded MVD resolution information. The decoding device 200 can extract the entropy-coded MVD resolution information from the bitstream, and can obtain the MVD resolution information by applying entropy decoding to the entropy-coded MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the resolution of the MVD.

[0503] Additionally, the encoding device 100 can calculate the MVD based on an affine model. The decoding device 200 can derive the affine control motion vector of the target block by the sum of the MVD and the affine control motion vector candidate, and can use the affine control motion vector to derive the motion vector of the sub-block.

[0504] The reference direction may indicate a list of reference pictures to be used for predicting a target block. For example, the reference direction may indicate one of reference picture list L0 and reference picture list L1.

[0505] The reference direction only indicates the list of reference pictures to be used for predicting a target block, and does not necessarily mean that the direction of the reference picture is limited to the forward direction or the backward direction. In other words, each of reference picture list L0 and reference picture list L1 may include pictures in the forward direction and / or the backward direction.

[0506] That the reference direction is unidirectional may mean using a single reference picture list. That the reference direction is bidirectional may mean using two reference picture lists. In other words, the reference direction may indicate one of the following cases: the case of using only reference picture list L0, the case of using only reference picture list L1, and the case of using two reference picture lists.

[0507] The reference picture index may indicate the reference picture among the reference pictures existing in the reference picture list for predicting the target block. The encoding device 100 may generate an entropy-encoded reference picture index by applying entropy coding to the reference picture index, and may generate a bitstream including the entropy-encoded reference picture index. The entropy-encoded reference picture index may be signaled from the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-encoded reference picture index from the bitstream, and may obtain the reference picture index by applying entropy decoding to the entropy-encoded reference picture index.

[0508] When two reference picture lists are used for predicting a target block, a single reference picture index and a single motion vector may be used for each of the reference picture lists. In addition, when two reference picture lists are used for predicting a target block, two prediction blocks may be specified for the target block. For example, the average value or weighted sum of the two prediction blocks for the target block may be used to generate the (final) prediction block of the target block.

[0509] The motion vector of the target block may be derived from the predicted motion vector index, MVD, reference direction, and reference picture index.

[0510] The decoding device 200 may generate a prediction block for the target block based on the derived motion vector and reference picture index. For example, the prediction block may be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.

[0511] Since the predicted motion vector index and MVD are encoded, while the motion vector of the target block itself is not encoded, the number of bits sent from the encoding device 100 to the decoding device 200 may be reduced, and the encoding efficiency may be improved.

[0512] For a target block, the motion information of reconstructed neighboring blocks can be used. In a specific inter prediction mode, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the motion information of the target block, additional information may be encoded, where the additional information enables the motion information of the target block to be derived using the motion information of the reconstructed neighboring blocks. Since the additional information is encoded, the number of bits sent to the decoding device 200 can be reduced, and the encoding efficiency can be improved.

[0513] For example, as inter prediction modes in which the motion information of the target block is not directly encoded, there may be a skip mode and / or a merge mode. Here, each of the encoding device 100 and the decoding device 200 may use an identifier and / or an index of a unit indicating that the motion information of the unit among the reconstructed neighboring units will be used as the motion information of the target unit.

[0514] 2) Merge mode As a scheme for deriving the motion information of the target block, there is merge. The term "merge" may mean merging the motions of multiple blocks. "Merge" may also mean that the motion information of one block is also applied to other blocks. In other words, the merge mode may be a mode of deriving the motion information of the target block from the motion information of neighboring blocks.

[0515] When using the merge mode, the encoding device 100 may use the motion information of spatial candidates and / or the motion information of temporal candidates to predict the motion information of the target block. Spatial candidates may include reconstructed spatial neighboring blocks that are spatially adjacent to the target block. Spatial neighboring blocks may include a left neighboring block and an upper neighboring block. Temporal candidates may include col blocks. The terms "spatial candidate" and "spatial merge candidate" may be used with the same meaning and may be used interchangeably with each other. The terms "temporal candidate" and "temporal merge candidate" may be used with the same meaning and may be used interchangeably with each other.

[0516] The encoding device 100 may obtain a prediction block through prediction. The encoding device 100 may encode a residual block that is the difference between the target block and the prediction block.

[0517] 2-1) Create a merge candidate list When using the merge mode, each of the encoding device 100 and the decoding device 200 may use the motion information of spatial candidates and / or the motion information of temporal candidates to create a merge candidate list. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may be unidirectional or bidirectional. The reference direction may represent an inter prediction indicator.

[0518] The merge candidate list may include merge candidates. A merge candidate may be motion information. In other words, the merge candidate list may be a list storing multiple pieces of motion information.

[0519] A merge candidate may be motion information of multiple temporal candidates and / or spatial candidates. In other words, the merge candidate list may include motion information of temporal candidates and / or spatial candidates, etc.

[0520] In addition, the merge candidate list may include new merge candidates generated by combining the merge candidates already existing in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list.

[0521] In addition, the merge candidate list may include history-based merge candidates. A history-based merge candidate may be the motion information of a block that has been encoded and / or decoded before the target block.

[0522] In addition, the merge candidate list may include merge candidates based on the average value of two merge candidates.

[0523] A merge candidate may be a specific mode for deriving inter-frame prediction information. A merge candidate may be information indicating a specific mode for deriving inter-frame prediction information. The inter-frame prediction information of the target block may be derived according to the specific mode indicated by the merge candidate. In addition, the specific mode may include a process for deriving a series of inter-frame prediction information. Such a specific mode may be an inter-frame prediction information derivation mode or a motion information derivation mode.

[0524] The inter-frame prediction information of the target block may be derived according to the mode indicated by the merge candidate selected from the merge candidates in the merge candidate list through a merge index.

[0525] For example, the motion information derivation mode in the merge candidate list may be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit and 2) an affine motion information derivation mode.

[0526] In addition, the merge candidate list may include motion information of a zero vector. The zero vector may also be referred to as a "zero merge candidate".

[0527] In other words, multiple pieces of motion information in the merge candidate list may be at least one of the following information: 1) motion information of a spatial candidate, 2) motion information of a temporal candidate, 3) motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list, and 4) a zero vector.

[0528] The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter-frame prediction indicator". The reference direction may be unidirectional or bidirectional. The unidirectional reference direction may indicate L0 prediction or L1 prediction.

[0529] A merge candidate list may be created before performing prediction in the merge mode.

[0530] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, such that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made the same as each other using the predefined scheme and the predefined priority.

[0531] Merge may be applied based on a CU or a PU. When performing merge based on a CU or a PU, the encoding device 100 may send a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merge is performed for each block partition, and 2) information about the block to be merged among the blocks that are spatial candidates and / or temporal candidates for the target block.

[0532] 2-2) Search for motion vectors using the merge candidate list The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using the merge candidates in the merge candidate list and may generate a residual block for the merge candidate. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in the prediction and the encoding of the residual block.

[0533] In addition, the encoding device 100 may determine whether to encode the target block using the merge mode.

[0534] 2-3) Transmission of inter-frame prediction information The encoding device 100 may generate a bitstream including the inter-frame prediction information required for inter-frame prediction. The encoding device 100 may generate entropy-encoded inter-frame prediction information by performing entropy encoding on the inter-frame prediction information, and may send the bitstream including the entropy-encoded inter-frame prediction information to the decoding device 200. The entropy-encoded inter-frame prediction information may be signaled by the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-encoded inter-frame prediction information from the bitstream and may obtain the inter-frame prediction information by applying entropy decoding to the entropy-encoded inter-frame prediction information.

[0535] The decoding device 200 may perform inter-frame prediction on the target block using the inter-frame prediction information of the bitstream.

[0536] The inter-frame prediction information may include 1) mode information indicating whether the merge mode is used, 2) a merge index, and 3) correction information.

[0537] In addition, the inter-frame prediction information may include a residual signal.

[0538] The decoding device 200 may obtain the merge index from the bitstream only when the mode information indicates that the merge mode is used.

[0539] The mode information may be a merge flag. The unit of the mode information may be a block. Information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.

[0540] The merge index may indicate a merge candidate among the merge candidates included in the merge candidate list that will be used to predict the target block. Optionally, the merge index may indicate a block among the neighboring blocks that are spatially or temporally adjacent to the target block and that will be merged with the target block.

[0541] The encoding device 100 may select a merge candidate having the highest encoding performance among the merge candidates included in the merge candidate list, and may set the value of the merge index to indicate the selected merge candidate.

[0542] The correction information may be information for correcting a motion vector. The encoding device 100 may generate the correction information. The decoding device 200 may correct the motion vector of the merge candidate selected by the merge index based on the correction information.

[0543] The correction information may include at least one of information indicating whether correction will be performed, correction direction information, and correction size information. The prediction mode in which the motion vector is corrected based on the signaled correction information may be referred to as "merge mode with motion vector difference".

[0544] 2-4) Inter-frame prediction in merge mode using inter-frame prediction information The decoding device 200 may perform prediction on the target block using the merge candidate indicated by the merge index among the merge candidates included in the merge candidate list.

[0545] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the merge index.

[0546] 3) Skip mode The skip mode may be a mode in which the motion information of a spatial candidate or the motion information of a temporal candidate is applied to the target block without change. In addition, the skip mode may be a mode in which the residual signal is not used. In other words, when the skip mode is used, the reconstructed block may be the same as the predicted block.

[0547] The difference between the merge mode and the skip mode lies in whether to send or use the residual signal. That is, the skip mode can be similar to the merge mode except that the residual signal is not sent or used.

[0548] When the skip mode is used, the encoding device 100 can send information related to the block whose motion information among the spatial candidates or temporal candidates will be used as the motion information of the target block to the decoding device 200 through the bitstream. The encoding device 100 can generate entropy-encoded information by performing entropy encoding on this information, and can signal the entropy-encoded information to the decoding device 200 through the bitstream. The decoding device 200 can extract the entropy-encoded information from the bitstream, and can obtain the information by applying entropy decoding to the entropy-encoded information.

[0549] In addition, when the skip mode is used, the encoding device 100 may not send other syntax information (such as MVD) to the decoding device 200. For example, when the skip mode is used, the encoding device 100 may not signal syntax elements related to at least one of MVD, coding block flag, and transform coefficient level to the decoding device 200.

[0550] 3-1) Create a merge candidate list The skip mode can also use the merge candidate list. In other words, the merge candidate list can be used in both the merge mode and the skip mode. In this regard, the merge candidate list can also be referred to as the "skip candidate list" or the "merge / skip candidate list".

[0551] Optionally, the skip mode can use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate can be replaced by the skip candidate list and the skip candidate, respectively.

[0552] The merge candidate list can be created before performing prediction in the skip mode.

[0553] 3-2) Search for motion vectors using the merge candidate list The encoding device 100 can determine the merge candidate to be used for encoding the target block. For example, the encoding device 100 can perform prediction on the target block using the merge candidate in the merge candidate list. The encoding device 100 can encode the target block using the merge candidate that generates the minimum cost in the prediction.

[0554] In addition, the encoding device 100 can determine whether to use the skip mode to encode the target block.

[0555] 3-3) Transmission of inter-frame prediction information The encoding device 100 can generate a bitstream including inter-prediction information required for inter-frame prediction. The decoding device 200 can perform inter-frame prediction on a target block using the inter-prediction information of the bitstream.

[0556] The inter-prediction information can include 1) mode information indicating whether the skip mode is used and 2) a skip index.

[0557] The skip index can be the same as the merge index described above.

[0558] When the skip mode is used, the target block can be encoded without using a residual signal. The inter-prediction information may not include a residual signal. Optionally, the bitstream may not include a residual signal.

[0559] The decoding device 200 can obtain the skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index can be the same as each other. The decoding device 200 can obtain the skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.

[0560] The skip index can indicate a merge candidate among the merge candidates included in the merge candidate list that will be used to perform prediction on the target block.

[0561] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information The decoding device 200 can perform prediction on the target block using the merge candidate indicated by the skip index among the merge candidates included in the merge candidate list.

[0562] The motion vector of the target block can be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the skip index.

[0563] 4) Current picture reference mode The current picture reference mode can represent such a prediction mode: this prediction mode uses a previously reconstructed region in the target picture to which the target block belongs.

[0564] A motion vector used to specify the previously reconstructed region can be used. The reference picture index of the target block can be used to determine whether the target block has been encoded in the current picture reference mode.

[0565] A flag or index indicating whether the target block is a block encoded in the current picture reference mode can be signaled by the encoding device 100 to the decoding device 200. Optionally, it can be inferred whether the target block is a block encoded in the current picture reference mode through the reference picture index of the target block.

[0566] When the target block is encoded in the current picture reference mode, the current picture can be present at a fixed position or any position in the reference picture list for the target block.

[0567] For example, the fixed position may be a position where the value of the reference picture index is 0 or the last position.

[0568] When the target picture exists at any position in the reference picture list, an additional reference picture index indicating such an arbitrary position may be signaled by the encoding device 100 to the decoding device 200.

[0569] 5) Sub-block merge mode The sub-block merge mode may be a mode of deriving motion information from sub-blocks of a CU.

[0570] When the sub-block merge mode is applied, motion information of a co-located sub-block (col-sub-block) of a target sub-block (i.e., a temporal merge candidate based on the sub-block) in a reference image and / or an affine control point motion vector merge candidate may be used to generate a sub-block merge candidate list.

[0571] 6) Triangular partition mode In the triangular partitioning mode, a target block may be partitioned in a diagonal direction, and sub-target blocks generated by the partitioning may be generated. For each sub-target block, motion information corresponding to the sub-target block may be derived, and the derived motion information may be used to derive prediction samples of each sub-target block. The prediction samples of the target block may be derived by a weighted sum of the prediction samples of the sub-target blocks generated via the partitioning.

[0572] 7) Combined inter-frame - intra-frame prediction mode The combined inter-intra prediction mode may be a mode of deriving prediction samples of a target block by using a weighted sum of prediction samples generated via inter prediction and prediction samples generated via intra prediction.

[0573] In the above modes, the decoding device 200 may autonomously correct the derived motion information. For example, the decoding device 200 may search for motion information having a minimum sum of absolute differences (SAD) in a specific region based on a reference block indicated by the derived motion information, and may derive the found motion information as the corrected motion information.

[0574] In the above modes, the decoding device 200 may use optical flow to compensate prediction samples derived via inter prediction.

[0575] In the AMVP mode, merge mode, skip mode, etc. described above, index information of a list may be used to specify motion information among multiple pieces of motion information in the list that will be used for predicting a target block.

[0576] To improve the coding efficiency, the coding device 100 may use only the index of the element among the elements in the signal transmission list that generates the minimum cost in the inter prediction of the target block. The coding device 100 may encode the index and signal the encoded index.

[0577] Therefore, the coding device 100 and the decoding device 200 must be able to derive the above-described lists (i.e., the predicted motion vector candidate list and the merge candidate list) based on the same data using the same scheme. Here, the same data may include the reconstructed picture and the reconstructed block. In addition, to specify an element using an index, the order of the elements in the list must be fixed.

[0578] Figure 10 Shows spatial candidates according to an embodiment.

[0579] In Figure 10 the positions of the spatial candidates are shown.

[0580] The large block at the center of the figure may represent the target block. Five small blocks may represent spatial candidates.

[0581] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).

[0582] Spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying the pixels at coordinates (xP - 1, yP + nPSH).

[0583] Spatial candidate A1 may be a block adjacent to the left side of the target block. A1 may be the lowermost block among the blocks adjacent to the left side of the target block. Optionally, A1 may be a block adjacent to the top of A0. A1 may be a block occupying the pixels at coordinates (xP - 1, yP + nPSH - 1).

[0584] Spatial candidate B0 may be a block adjacent to the upper right corner of the target block. B0 may be a block occupying the pixels at coordinates (xP + nPSW, yP - 1).

[0585] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Optionally, B1 may be a block adjacent to the left of B0. B1 may be a block occupying the pixels at coordinates (xP + nPSW - 1, yP - 1).

[0586] Spatial candidate B2 may be a block adjacent to the upper left corner of the target block. B2 may be a block occupying the pixels at coordinates (xP - 1, yP - 1).

[0587] Determination of the availability of spatial and temporal candidates In order to include motion information of a spatial candidate or motion information of a temporal candidate in a list, it is necessary to determine whether the motion information of the spatial candidate or the motion information of the temporal candidate is available.

[0588] Hereinafter, a candidate block may include a spatial candidate and a temporal candidate.

[0589] For example, the determination may be performed by sequentially applying the following steps 1) to 4).

[0590] Step 1) When the PU including the candidate block is outside the boundary of the picture, the availability of the candidate block may be set to "false". The expression "the availability is set to false" may have the same meaning as "set to unavailable".

[0591] Step 2) When the PU including the candidate block is outside the boundary of the strip, the availability of the candidate block may be set to "false". When the target block and the candidate block are in different strips, the availability of the candidate block may be set to "false".

[0592] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block may be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block may be set to "false".

[0593] Step 4) When the prediction mode of the PU including the candidate block is an intra prediction mode, the availability of the candidate block may be set to "false". When the PU including the candidate block does not use inter prediction, the availability of the candidate block may be set to "false".

[0594] Figure 11 Shows the order of adding motion information of a spatial candidate to the merge list according to an embodiment.

[0595] As Figure 11 shown, when multiple pieces of motion information of a spatial candidate are added to the merge list, the order of A1, B1, B0, A0, and B2 may be used. That is, multiple pieces of motion information of available spatial candidates may be added to the merge list in the order of A1, B1, B0, A0, and B2.

[0596] Method for deriving a merge list in merge mode and skip mode As described above, the maximum number of merge candidates in the merge list may be set. The set maximum number may be indicated by "N". The set number may be sent from the encoding device 100 to the decoding device 200. The strip header of the strip may include N. In other words, the maximum number of merge candidates in the merge list for the target block of the strip may be set by the strip header. For example, the value of N may be substantially 5.

[0597] Multiple pieces of motion information (i.e., merge candidates) can be added to the merge list in the order of the following steps 1) to 4).

[0598] Step 1) Among the spatial candidates, available spatial candidates can be added to the merge list. Multiple pieces of motion information of the available spatial candidates can be added to the merge list in the order shown in Figure 11 Here, when the motion information of the available spatial candidate overlaps with other motion information already existing in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list can be simply referred to as "overlap check".

[0599] The maximum number of motion information added can be N.

[0600] Step 2) When the number of motion information in the merge list is less than N and time candidates are available, the motion information of the time candidates can be added to the merge list. Here, when the motion information of the available time candidate overlaps with other motion information already existing in the merge list, the motion information of the available time candidate may not be added to the merge list.

[0601] Step 3) When the number of motion information in the merge list is less than N and the type of the target strip is "B", the combined motion information generated by combining bidirectional prediction (bi-prediction) can be added to the merge list.

[0602] The target strip can be a strip including the target block.

[0603] The combined motion information can be a combination of L0 motion information and L1 motion information. The L0 motion information can be motion information that only refers to the reference picture list L0. The L1 motion information can be motion information that only refers to the reference picture list L1.

[0604] In the merge list, there can be one or more pieces of L0 motion information. In addition, in the merge list, there can be one or more pieces of L1 motion information.

[0605] The combined motion information can include one or more pieces of combined motion information. When generating the combined motion information, the L0 motion information and L1 motion information of the steps to be used for generating the combined motion information among the one or more pieces of L0 motion information and the one or more pieces of L1 motion information can be predefined. One or more pieces of combined motion information can be generated in a predefined order through combined bidirectional prediction using a combination of a pair of different motion information in the merge list. One piece of motion information in the pair of different motion information can be L0 motion information, and the other piece of motion information in the pair of different motion information can be L1 motion information.

[0606] For example, the combined motion information with the highest priority added thereto may be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information or when the motion information with a merge index of 1 is not L1 motion information, the combined motion information may neither be generated nor added. Next, the combined motion information with the next highest priority added thereto may be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. Subsequent detailed combinations may conform to other combinations in the field of video encoding / decoding.

[0607] Here, when the combined motion information overlaps with other motion information already present in the merge list, the combined motion information may not be added to the merge list.

[0608] Step 4) When the number of motion information in the merge list is less than N, motion information of a zero vector may be added to the merge list.

[0609] The motion information of a zero vector may be motion information whose motion vector is a zero vector.

[0610] The number of motion information of a zero vector may be one or more. The reference picture indices of one or more motion information of a zero vector may be different from each other. For example, the value of the reference picture index of the first motion information of a zero vector may be 0. The value of the reference picture index of the second motion information of a zero vector may be 1.

[0611] The number of motion information of a zero vector may be the same as the number of reference pictures in the reference picture list.

[0612] The reference direction of the motion information of a zero vector may be bidirectional. Both motion vectors may be zero vectors. The number of motion information of a zero vector may be the smaller one of the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1. Optionally, when the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1 are different from each other, a unidirectional reference direction may be used for the reference picture index that can be applied only to a single reference picture list.

[0613] The encoding device 100 and / or the decoding device 200 may then add the motion information of a zero vector to the merge list while changing the reference picture index.

[0614] When the motion information of a zero vector overlaps with other motion information already present in the merge list, the motion information of a zero vector may not be added to the merge list.

[0615] The order of the above steps 1) to 4) is merely exemplary and can be changed. In addition, some of the above steps can be omitted according to predefined conditions.

[0616] Method for deriving a list of predicted motion vector candidates in AMVP mode The maximum number of predicted motion vector candidates in the predicted motion vector candidate list can be predefined. The predefined maximum number can be indicated by N. For example, the predefined maximum number can be 2.

[0617] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of the following steps 1) to 3).

[0618] Step 1) Available spatial candidates among the spatial candidates can be added to the predicted motion vector candidate list. The spatial candidates can include a first spatial candidate and a second spatial candidate.

[0619] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.

[0620] Multiple pieces of motion information of the available spatial candidates can be added to the predicted motion vector candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, when the motion information of the available spatial candidates overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available spatial candidates may not be added to the predicted motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.

[0621] The maximum number of the added motion information can be N.

[0622] Step 2) When the number of pieces of motion information in the predicted motion vector candidate list is less than N and the temporal candidate is available, the motion information of the temporal candidate can be added to the predicted motion vector candidate list. In this case, when the motion information of the available temporal candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information of the available temporal candidate may not be added to the predicted motion vector candidate list.

[0623] Step 3) When the number of pieces of motion information in the predicted motion vector candidate list is less than N, the zero vector motion information can be added to the predicted motion vector candidate list.

[0624] The zero vector motion information may include one or more pieces of zero vector motion information. The reference picture indices of the one or more pieces of zero vector motion information may be different from each other.

[0625] The encoding device 100 and / or the decoding device 200 may sequentially add a plurality of pieces of zero vector motion information to the prediction motion vector candidate list while changing the reference picture index.

[0626] When the zero vector motion information overlaps with other motion information already existing in the prediction motion vector candidate list, the zero vector motion information may not be added to the prediction motion vector candidate list.

[0627] The description of the zero vector motion information made in combination with the merge list above may also be applied to the zero vector motion information. The repeated description thereof will be omitted.

[0628] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some of the steps may be omitted according to predefined conditions.

[0629] Figure 12 Shows the transform and quantization processing according to an example.

[0630] As Figure 12 shown, the quantized levels may be generated by performing transform and / or quantization processing on the residual signal.

[0631] The residual signal may be generated as the difference between the original block and the prediction block. Here, the prediction block may be a block generated via intra prediction or inter prediction.

[0632] The residual signal may be transformed into a signal in the frequency domain by a transform process that is part of the quantization process.

[0633] The transform kernel for the transform may include various DCT kernels, such as discrete cosine transform (DCT) type 2 (DCT-II) and discrete sine transform (DST) kernels.

[0634] These transform kernels may perform a separable transform or a two-dimensional (2D) non-separable transform on the residual signal. The separable transform may be a transform that indicates performing a one-dimensional (1D) transform on the residual signal in each of the horizontal and vertical directions.

[0635] The DCT types and DST types adaptively used for the 1D transform may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II, as shown in each of Table 3 and Table 4 below.

[0636] Table 3

[0637] Table 4

[0638] As shown in Table 3 and Table 4, a transform set can be used when deriving the DCT type or DST type to be used for transformation. Each transform set can include multiple transform candidates. Each transform candidate can be a DCT type or a DST type.

[0639] Table 5 below shows examples of the transform set to be applied to the horizontal direction and the transform set to be applied to the vertical direction according to the intra prediction mode.

[0640] Table 5

[0641] In Table 5, the numbers of the vertical transform set and the horizontal transform set to be applied to the horizontal direction of the residual signal according to the intra prediction mode of the target block are shown.

[0642] As illustrated in Table 5, the transform sets to be applied to the horizontal direction and the vertical direction can be predefined according to the intra prediction mode of the target block. The encoding device 100 can use the transforms included in the transform set corresponding to the intra prediction mode of the target block to perform transformation and inverse transformation on the residual signal. In addition, the decoding device 200 can use the transforms included in the transform set corresponding to the intra prediction mode of the target block to perform inverse transformation on the residual signal.

[0643] In the transformation and inverse transformation, as illustrated in Table 3, Table 4, and Table 5, the transform set to be applied to the residual signal can be determined and not signaled. The transform indication information can be signaled from the encoding device 100 to the decoding device 200. The transform indication information can be information indicating which one of the multiple transform candidates included in the transform set to be applied to the residual signal is used.

[0644] For example, when the size of the target block is 64×64 or smaller, transform sets each having three transforms can be configured according to the intra prediction mode. The optimal transform method can be selected from a total of nine multi-transform methods generated by the combination of three transforms in the horizontal direction and three transforms in the vertical direction. By such an optimal transform method, the residual signal can be encoded and / or decoded, and thus the encoding efficiency can be improved.

[0645] Here, the information indicating which one of the multiple transforms belonging to each transform set has been used in at least one of the vertical transformation and the horizontal transformation can be entropy encoded and / or entropy decoded. Here, truncated binary coding can be used to encode and / or decode such information.

[0646] As described above, the method using various transforms can be applied to the residual signal generated via intra prediction or inter prediction.

[0647] The transformation may include at least one of a first transformation and a secondary transformation. The transformation coefficients may be generated by performing the first transformation on the residual signal, and the secondary transformation coefficients may be generated by performing the secondary transformation on the transformation coefficients.

[0648] The first transformation may be referred to as the "primary transformation". In addition, the first transformation may also be referred to as the "Adaptive Multi-Transform (AMT) scheme". As described above, AMT may represent applying different transformations to each 1D direction (i.e., the vertical direction and the horizontal direction).

[0649] The secondary transformation may be a transformation for enhancing the energy concentration of the transformation coefficients generated by the first transformation. Similar to the first transformation, the secondary transformation may be a separable transformation or a non-separable transformation. Such a non-separable transformation may be a Non-Separable Secondary Transform (NSST).

[0650] At least one of a plurality of predefined transformation methods may be used to perform the first transformation. For example, the plurality of predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), etc.

[0651] In addition, according to the kernel functions defining the Discrete Cosine Transform (DCT) or the Discrete Sine Transform (DST), the first transformation may be a transformation of various types.

[0652] For example, the type of transformation may be determined based on at least one of the following items: 1) the prediction mode of the target block (e.g., one of intra-frame prediction and inter-frame prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra-frame prediction mode of the target block, 5) the component of the target block (e.g., one of the luminance component and the chrominance component), and 6) the partition type applied to the target block (e.g., one of quadtree, binary tree, and ternary tree).

[0653] For example, according to the transformation kernels presented in Table 6 below, the first transformation may include transformations such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8. In Table 6 below, various transformation types and transformation kernel functions for Multi-Transform Selection (MTS) are illustrated.

[0654] MTS may refer to the selection of a combination of one or more DCT and / or DST kernels for transforming the residual signal in the horizontal and / or vertical directions.

[0655] Table 6

[0656] In Table 6, i and j may be integer values equal to or greater than 0 and less than or equal to N - 1.

[0657] A secondary transformation may be performed on the transformation coefficients generated by performing a first transformation.

[0658] As in the first transformation, a set of transformations may also be defined in the secondary transformation. The method for deriving and / or determining the above set of transformations may be applied not only to the first transformation but also to the secondary transformation.

[0659] The first transformation and the secondary transformation may be determined for a specific target.

[0660] For example, the first transformation and the secondary transformation may be applied to signal components corresponding to one or more of the luma component and the chroma component. Whether to apply the first transformation and / or the secondary transformation may be determined according to at least one of the coding parameters for the target block and / or neighboring blocks. For example, whether to apply the first transformation and / or the secondary transformation may be determined according to the size and / or shape of the target block.

[0661] In the encoding device 100 and the decoding device 200, transformation information indicating the transformation method to be used for the target may be derived by using specified information.

[0662] For example, the transformation information may include a transformation index to be used for the primary transformation and / or the secondary transformation. Optionally, the transformation information may indicate that the primary transformation and / or the secondary transformation is not used.

[0663] For example, when the targets of the primary transformation and the secondary transformation are the target block, the transformation method indicated by the transformation information to be applied to the primary transformation and / or the secondary transformation may be determined according to at least one of the coding parameters for the target block and / or the block adjacent to the target block.

[0664] Optionally, the transformation information indicating the transformation method for a specific target may be signaled from the encoding device 100 to the decoding device 200.

[0665] For example, for a single CU, the decoding device 200 may derive as transformation information whether to use the primary transformation, the index indicating the primary transformation, whether to use the secondary transformation, and the index indicating the secondary transformation. Optionally, for a single CU, the transformation information indicating the following items may be signaled: whether to use the primary transformation, the index indicating the primary transformation, whether to use the secondary transformation, and the index indicating the secondary transformation.

[0666] Quantized transformation coefficients (i.e., quantized levels) may be generated by performing quantization on the result generated by performing the first transformation and / or the secondary transformation or by performing quantization on the residual signal.

[0667] Figure 13 A diagonal scan according to an example is shown.

[0668] Figure 14 Shows a horizontal scan according to an example.

[0669] Figure 15 Shows a vertical scan according to an example.

[0670] The quantized transform coefficients can be scanned via at least one of (upper right) diagonal scan, vertical scan, and horizontal scan according to at least one of the intra prediction mode, block size, and block shape. The block can be a transform unit (TU).

[0671] Each scan can be initiated at a specific start point and terminated at a specific end point.

[0672] For example, by using Figure 13 the diagonal scan of to scan the coefficients of a block to change the quantized transform coefficients into a 1D vector form. Optionally, according to the size of the block and / or the intra prediction mode, Figure 14 the horizontal scan of or Figure 15 the vertical scan of can be used without using the diagonal scan.

[0673] The vertical scan can be an operation of scanning 2D block type coefficients in the column direction. The horizontal scan can be an operation of scanning 2D block type coefficients in the row direction.

[0674] In other words, which one of the diagonal scan, vertical scan, and horizontal scan will be used can be determined according to the size of the block and / or the inter prediction mode.

[0675] As Figure 13 , Figure 14 and Figure 15 shown, the quantized transform coefficients can be scanned along the diagonal direction, horizontal direction, or vertical direction.

[0676] The quantized transform coefficients can be represented by the block shape. Each block can include a plurality of sub-blocks. Each sub-block can be defined according to the minimum block size or the minimum block shape.

[0677] In the scan, the scan order according to the type or direction of the scan can be first applied to the sub-blocks. In addition, the scan order according to the direction of the scan can be applied to the quantized transform coefficients in each sub-block.

[0678] For example, as Figure 13 , Figure 14 and Figure 15 shown, when the size of the target block is 8×8, the quantized transform coefficients can be generated by the first transform, secondary transform, and quantization of the residual signal of the target block. Therefore, one of the three types of scan orders can be applied to four 4×4 sub-blocks, and the quantized transform coefficients can also be scanned for each 4×4 sub-block according to the scan order.

[0679] The encoding device 100 can generate entropy-encoded quantized transform coefficients by performing entropy encoding on the scanned and quantized transform coefficients, and can generate a bitstream including the entropy-encoded quantized transform coefficients.

[0680] The decoding device 200 can extract the entropy-encoded quantized transform coefficients from the bitstream, and can generate quantized transform coefficients by performing entropy decoding on the entropy-encoded quantized transform coefficients. The quantized transform coefficients can be arranged in the form of 2D blocks via inverse scanning. Here, as a method of inverse scanning, at least one of right-up diagonal scanning, vertical scanning, and horizontal scanning can be performed.

[0681] In the decoding device 200, inverse quantization can be performed on the quantized transform coefficients. Depending on whether secondary inverse transformation is performed, secondary inverse transformation can be performed on the result generated by performing inverse quantization. In addition, depending on whether first inverse transformation will be performed, first inverse transformation can be performed on the result generated by performing secondary inverse transformation. A reconstructed residual signal can be generated by performing first inverse transformation on the result generated by performing secondary inverse transformation.

[0682] For the luminance component reconstructed by intra prediction or inter prediction, inverse mapping with a dynamic range can be performed before loop filtering.

[0683] The dynamic range can be divided into 16 equal segments, and the mapping function of the corresponding segment can be signaled. Such a mapping function can be signaled at the slice level or the parallel block group level.

[0684] An inverse mapping function for performing inverse mapping can be derived based on the mapping function.

[0685] Loop filtering, storage of reference pictures, and motion compensation can be performed in the inverse mapping region.

[0686] The prediction block generated by inter prediction can be transformed to the mapping region by using the mapping of the mapping function, and the transformed prediction block can be used to generate a reconstructed block. However, since intra prediction is performed in the mapping region, the prediction block generated by intra prediction can be used to generate a reconstructed block without the need for mapping and / or inverse mapping.

[0687] For example, when the target block is a residual block of the chrominance component, the residual block can be transformed to the inverse mapping region by scaling the chrominance component of the mapping region.

[0688] Whether scaling is available can be signaled at the slice level or the parallel block group level.

[0689] For example, scaling can be applied only to the case where mapping is available for the luminance component and the partitions of the luminance component and the chrominance component follow the same tree structure.

[0690] Scaling may be performed based on the average value of the samples in the luma prediction block corresponding to the chroma prediction block. Here, when the target block uses inter prediction, the luma prediction block may represent the mapped luma prediction block.

[0691] The value required for scaling may be derived by referring to a lookup table using the index of the segment to which the average value of the samples of the luma prediction block belongs.

[0692] The residual block may be transformed to the inverse mapped region by scaling the residual block using the finally derived value. Thereafter, for the blocks of the chroma component, reconstruction, intra prediction, inter prediction, loop filtering, and storage of reference pictures may be performed in the inverse mapped region.

[0693] For example, information indicating whether mapping and / or inverse mapping of the luma component and the chroma component are available may be signaled by the sequence parameter set.

[0694] A prediction block of the target block may be generated based on a block vector. The block vector may indicate the displacement between the target block and the reference block. The reference block may be a block in the target picture.

[0695] In this way, the prediction mode of generating a prediction block by referring to the target picture may be referred to as an "intra block copy (IBC) mode".

[0696] The IBC mode may be applied to a CU having a specific size. For example, the IBC mode may be applied to a CU of M×N. Here, M and N may be less than or equal to 64.

[0697] The IBC mode may include a skip mode, a merge mode, an AMVP mode, etc. In the case of the skip mode or the merge mode, a merge candidate list may be configured, and a merge index may be signaled, and thus a single merge candidate may be specified among the merge candidates existing in the merge candidate list. The block vector of the specified merge candidate may be used as the block vector of the target block.

[0698] In the case of the AMVP mode, a differential block vector may be signaled. In addition, a prediction block vector may be derived from the left neighboring block and the upper neighboring block of the target block. In addition, an index indicating which neighboring block will be used may be signaled.

[0699] The prediction block in the IBC mode may be included in the target CTU or the left CTU, and may be limited to the blocks within the previously reconstructed region. For example, the value of the block vector may be restricted such that the prediction block of the target block is located in a specific region. The specific region may be a region defined by three 64×64 blocks that have been encoded and / or decoded before the 64×64 block including the target block. Restricting the value of the block vector in this way, the memory consumption and device complexity caused by the implementation of the IBC mode may be reduced.

[0700] Figure 16 It is a configuration diagram of an encoding device according to an embodiment.

[0701] The encoding device 1600 may correspond to the encoding device 100 described above.

[0702] The encoding device 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a storage 1640 that communicate with each other via a bus 1690. The encoding device 1600 may further include a communication unit 1620 connected to a network 1699.

[0703] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1630 or the storage 1640. The processing unit 1610 may be at least one hardware processor.

[0704] The processing unit 1610 may generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1610.

[0705] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transformation unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transformation unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.

[0706] At least some of the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transformation unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transformation unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190 may be program modules and may communicate with an external device or system. The program modules may be included in the encoding device 1600 in the form of an operating system, an application program module, or other program modules.

[0707] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the encoding device 1600.

[0708] A program module may include, but is not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing an abstract data type according to an embodiment.

[0709] The program module may be implemented using instructions or code run by at least one processor of the encoding device 1600.

[0710] The processing unit 1610 may run instructions or code in the inter-frame prediction unit 110, intra-frame prediction unit 120, switch 115, subtractor 125, transform unit 130, quantization unit 140, entropy encoding unit 150, inverse quantization unit 160, inverse transform unit 170, adder 175, filter unit 180, and reference picture buffer 190.

[0711] The storage unit may represent the memory 1630 and / or the storage 1640. Each of the memory 1630 and the storage 1640 may be any of various types of volatile or non-volatile storage media. For example, the memory 1630 may include at least one of a read-only memory (ROM) 1631 and a random access memory (RAM) 1632.

[0712] The storage unit may store data or information for the operation of the encoding device 1600. In an embodiment, the data or information of the encoding device 1600 may be stored in the storage unit.

[0713] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.

[0714] The encoding device 1600 may be implemented in a computer system including a computer-readable storage medium.

[0715] The storage medium may store at least one module required for the operation of the encoding device 1600. The memory 1630 may store at least one module and may be configured such that the at least one module is run by the processing unit 1610.

[0716] Functions related to the communication of data or information with the encoding device 1600 may be performed by the communication unit 1620.

[0717] For example, the communication unit 1620 may send a bitstream to the decoding device 1700 to be described later.

[0718] Figure 17 is a configuration diagram of a decoding device according to an embodiment.

[0719] The decoding device 1700 may correspond to the decoding device 200 described above.

[0720] The decoding device 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage 1740 that communicate with each other via a bus 1790. The decoding device 1700 may further include a communication unit 1720 connected to a network 1799.

[0721] The processing unit 1710 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1730 or the storage 1740. The processing unit 1710 may be at least one hardware processor.

[0722] The processing unit 1710 may generate and process signals, data, or information input to, output from, or used in the decoding device 1700, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1710.

[0723] The processing unit 1710 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, a switch 245, an adder 255, a filter unit 260, and a reference picture buffer 270.

[0724] At least some of the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra prediction unit 240, the inter prediction unit 250, the adder 255, the switch 245, the filter unit 260, and the reference picture buffer 270 of the decoding device 200 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding device 1700 in the form of an operating system, an application program module, or other program modules.

[0725] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the decoding device 1700.

[0726] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.

[0727] The program modules may be implemented using instructions or code run by at least one processor of the decoding device 1700.

[0728] The processing unit 1710 can execute instructions or code in the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra prediction unit 240, the inter prediction unit 250, the switch 245, the adder 255, the filter unit 260, and the reference picture buffer 270.

[0729] The storage unit can represent the memory 1730 and / or the storage 1740. Each of the memory 1730 and the storage 1740 can be any one of various types of volatile or non-volatile storage media. For example, the memory 1730 can include at least one of the ROM 1731 and the RAM 1732.

[0730] The storage unit can store data or information for the operation of the decoding device 1700. In an embodiment, the data or information of the decoding device 1700 can be stored in the storage unit.

[0731] For example, the storage unit can store pictures, blocks, lists, motion information, inter prediction information, bitstreams, etc.

[0732] The decoding device 1700 can be implemented in a computer system including a computer-readable storage medium.

[0733] The storage medium can store at least one module required for the operation of the decoding device 1700. The memory 1730 can store at least one module and can be configured such that the at least one module is executed by the processing unit 1710.

[0734] Functions related to the communication of data or information of the decoding device 1700 can be executed through the communication unit 1720.

[0735] For example, the communication unit 1720 can receive a bitstream from the encoding device 1600.

[0736] Hereinafter, the processing unit may represent the processing unit 1610 of the encoding device 1600 and / or the processing unit 1710 of the decoding device 1700. For example, regarding functions related to prediction, the processing unit may represent switch 115 and / or switch 245. Regarding functions related to inter-frame prediction, the processing unit may represent the inter-frame prediction unit 110, subtractor 125, and adder 175, and may represent the inter-frame prediction unit 250 and adder 255. Regarding functions related to intra-frame prediction, the processing unit may represent the intra-frame prediction unit 120, subtractor 125, and adder 175, and may represent the intra-frame prediction unit 240 and adder 255. Regarding functions related to transformation, the processing unit may represent the transformation unit 130 and the inverse transformation unit 170, and may represent the inverse transformation unit 230. Regarding functions related to quantization, the processing unit may represent the quantization unit 140 and the inverse quantization unit 160, and may indicate the inverse quantization unit 220. Regarding functions related to entropy encoding and / or entropy decoding, the processing unit may represent the entropy encoding unit 150 and / or the entropy decoding unit 210. Regarding functions related to filtering, the processing unit may represent the filter unit 180 and / or the filter unit 260. Regarding functions related to reference pictures, the processing unit may indicate the reference picture buffer 190 and / or the reference picture buffer 270.

[0737] Techniques based on artificial neural networks are being proposed for video compression. Video compression techniques using artificial neural networks require much more computational power than existing video compression techniques, but provide better performance compared to existing video compression techniques that have reached the performance limit. Therefore, attempts are being made to adopt artificial neural networks across various technical fields of traditional video compression.

[0738] Embodiments of the present disclosure may provide a method for obtaining an image closer to the original image or a method for obtaining an image that increases image encoding / decoding efficiency from one or more intermediate images by adopting an artificial neural network.

[0739] The artificial neural network according to an embodiment of the present disclosure may be used in the process of generating a predicted image in inter-picture prediction, the process of transforming an image into a high-resolution image, and the process of adding an image to a reference picture list to be used for inter-picture prediction.

[0740] The artificial neural network according to an embodiment of the present disclosure may use two or more images as inputs to effectively extract the correlation between the input images as spatial features, and additionally use information related to the encoding / decoding of the input images as inputs.

[0741] An artificial neural network may alternatively be referred to as a deep neural network, a neural network, or simply a network. Hereinafter, the artificial neural network will be simply referred to as a neural network, and a filter based on the artificial neural network will also be referred to as a neural network-based filter.

[0742] Also, terms such as learning or training may be used in the same context in view of updating weights of nodes that constitute layers of a neural network.

[0743] In the following embodiments, an image may represent units of input and output of a neural network. An image may refer to a picture, a frame, a picture block, a block, or a parallel block. In other words, various forms including a picture, a frame, a picture block, a block, or a parallel block may be defined as an image.

[0744] Figure 18 is a flowchart showing a process of processing an image using a neural network according to an embodiment of the present disclosure.

[0745] In step S1810, the type of the neural network for the output image and the input image to be input to the neural network are determined. The type of the neural network may be determined according to the purpose of using the output image. In addition, two or more images may be used as the input to the neural network so that the neural network can easily extract spatial features or feature maps required for a high-quality output image.

[0746] For example, two or more prediction images, two or more recovery images, a prediction image and a recovery image, two or more reference images included in a reference picture list, a chrominance image and a luminance image, etc. may be input to the neural network as a combination consisting of two or more images.

[0747] In the above example, if there is only one input image, after generating another image by copying or modifying the input image, multiple images may be used as the input to the neural network. In addition, when using multiple different input images, at least one image among the multiple images may be selected and used as the input to the neural network instead of selecting and using all available images. For example, a low-quality image among the multiple images may be excluded from the neural network input. For example, among the multiple input images, an image having a quantization parameter (QP) greater than a predetermined threshold compared to the quantization parameter (QP) of the current image may be excluded from the neural network input.

[0748] Since a neural network may process an image in a predetermined coding unit such as a block, a parallel block, a stripe, and a picture, the input image of the neural network may take the form of a block, a parallel block, a stripe, or a picture.

[0749] In addition, the neural network input may further include encoding / decoding information for encoding or decoding an image used as the input of the neural network, such as block partitioning information, quantization parameter information, boundary strength (BS) information, and picture or slice type information.

[0750] The network structure of the neural network may vary according to the type or number of input images; or may vary according to the type or number of encoding information.

[0751] Step S1820 processes the input image using the neural network determined in step S1810 and outputs image data.

[0752] Step S1830 transforms the image data output by the neural network into an image in a predetermined form, for example, a block, a parallel block, a slice, or a picture.

[0753] By Figure 18 The image generated by the process can be used as a predicted image in an inter-picture prediction mode, or upsampled to a high-resolution restored image, or added to the reference picture list as a reference image.

[0754] In other words, Figure 18 The operation flowchart of can be used for the process of generating a predicted image in an inter-picture prediction mode, the process of upsampling an image restored at a low resolution to a high-resolution image, and the process of generating a new reference picture using the reference pictures included in the reference picture list.

[0755] First, an embodiment of using Figure 18 the process to generate a predicted image in an inter-picture prediction mode will be described.

[0756] Figure 19 is a flowchart showing an encoding method for generating a bitstream by prediction of a target block according to an embodiment.

[0757] The encoding method for generating a bitstream by prediction of a target block according to an embodiment may be executed by an encoding device 1600. The embodiment may be part of an encoding method for a target block or a video encoding method.

[0758] In step S1910, the processing unit 1610 may determine prediction information to be applied to the encoding of the target block.

[0759] The prediction information may include information for the above prediction. For example, the prediction information may include one or more of inter-picture prediction information and intra-picture prediction information.

[0760] In step S1920, the processing unit 1610 may perform prediction on the target block using the determined prediction information.

[0761] A predicted block can be generated by predicting a target block.

[0762] A residual block, which is the difference between the target block and the predicted block, can be generated. By applying a transform and quantization to the residual block, information about the target block can be generated.

[0763] The information about the target block can include transform and quantization coefficients for the target block. The information about the target block can include prediction information.

[0764] Furthermore, a reconstructed block can be generated by summing the predicted block and the recovered residual block.

[0765] As described later, the prediction for the target block can be a neural network-based unidirectional prediction. Here, the neural network-based unidirectional prediction can improve intermediate predicted blocks.

[0766] The neural network-based unidirectional prediction can refer to a unidirectional prediction performed by a neural network.

[0767] The neural network-based unidirectional prediction can be performed based on the attributes and / or coding parameters of the target block.

[0768] It can be determined whether to perform the neural network-based unidirectional prediction based on the attributes and / or coding parameters of the target block.

[0769] Here, the attributes of the target block can include the block size of the target block and the component of the target block. For example, it can be determined whether to perform the neural network-based unidirectional prediction based on the block size of the target block.

[0770] As described later, the prediction for the target block can be a neural network-based bidirectional prediction. Here, the neural network-based bidirectional prediction can be used to generate one final predicted block using multiple intermediate predicted blocks.

[0771] Here, the multiple intermediate predicted blocks can be blocks improved by the neural network-based unidirectional prediction.

[0772] The neural network-based bidirectional prediction can refer to a bidirectional prediction performed by a neural network.

[0773] The neural network-based bidirectional prediction can be performed based on the attributes and / or coding parameters of the target block.

[0774] It can be determined whether to perform the neural network-based bidirectional prediction based on the attributes and / or coding parameters of the target block.

[0775] Here, the attributes of the target block can include the block size of the target block and the component of the target block.

[0776] The prediction of the target block can be performed using a bidirectional prediction method selected from multiple bidirectional prediction methods. Here, the multiple bidirectional prediction methods may include neural network-based bidirectional prediction. In other words, the neural network-based bidirectional prediction can be selected from the multiple bidirectional prediction methods according to specific conditions.

[0777] The multiple bidirectional prediction methods may also include one or more of average bidirectional prediction, optical flow-based bidirectional prediction, and weighted average-based bidirectional prediction, which will be described later.

[0778] When determining the bidirectional prediction method, the attributes of the target block and / or coding parameters can be used. For example, the bidirectional prediction method can be determined based on the block size of the target block.

[0779] When determining the bidirectional prediction method, the attributes of the bidirectional prediction and / or coding parameters related to the bidirectional prediction can be used. For example, the bidirectional prediction method can be determined based on the weight of the bidirectional prediction.

[0780] In step S1930, the processing unit 1610 can generate a bitstream.

[0781] The bitstream may include information about the target block. In addition, the bitstream can include the information described in the above embodiments. For example, the bitstream can include coding parameters related to the target block and / or the attributes of the target block.

[0782] The information included in the bitstream can be generated in step S1930, or can be at least partially generated in steps S1910 and S1920.

[0783] The processing unit 1610 can store the generated bitstream in the storage device 1640. Optionally, the communication unit 1620 can send the bitstream to the decoding device 1700.

[0784] The bitstream can include information about the encoded target block. The processing unit 1610 can generate the information about the encoded target block by performing entropy coding on the information about the target block.

[0785] Figure 20 It is a flowchart showing a method for predicting a target block using a bitstream according to an embodiment.

[0786] The method for predicting a target block using a bitstream according to an embodiment can be executed by the decoding device 1700. The embodiment can be part of a decoding method for the target block or a video decoding method.

[0787] In step S2010, the communication unit 1720 can obtain the bitstream. The communication unit 1720 can receive the bitstream from the encoding device 1600.

[0788] The bitstream may include information about the target block.

[0789] The information about the target block may include transform and quantization coefficients for the target block. The information about the target block may include prediction information.

[0790] In addition, the bitstream may include the information described in the embodiments. For example, the bitstream may include coding parameters related to the target block and / or attributes of the target block.

[0791] The computer-readable recording medium may include the bitstream, and the information about the target block included in the bitstream may be used to perform prediction for the target block and decode the target block.

[0792] The bitstream may include information about the encoded target block. The processing unit 1710 may generate information about the target block by performing entropy decoding on the information about the encoded target block.

[0793] The processing unit 1710 may store the obtained bitstream in the storage device 1740.

[0794] In step S2020, the processing unit 1710 may determine the prediction information to be applied to the decoding of the target block.

[0795] The processing unit 1710 may use the method used in the above embodiments to determine the prediction information.

[0796] The processing unit 1710 may determine the prediction information of the target block based on the information related to the prediction method obtained from the bitstream.

[0797] The prediction information may include inter-picture prediction information and intra-picture prediction information.

[0798] In step S2030, the processing unit 1710 may perform prediction on the target block using the information about the target block and the determined prediction information.

[0799] As described later, the prediction of the target block may be based on the unidirectional prediction of the neural network. Here, the unidirectional prediction based on the neural network may improve the intermediate prediction block.

[0800] The unidirectional prediction based on the neural network may be performed based on the attributes and / or coding parameters of the target block.

[0801] It may be determined whether to perform the unidirectional prediction based on the neural network based on the attributes and / or coding parameters of the target block.

[0802] Here, the attributes of the target block may include the block size of the target block and the component of the target block. For example, it may be determined whether to perform the unidirectional prediction based on the neural network based on the block size of the target block.

[0803] In an embodiment, as described later, the prediction for a target block may be based on bidirectional prediction of a neural network. Here, the bidirectional prediction based on the neural network can be used to generate a final prediction block using multiple intermediate prediction blocks.

[0804] Here, the multiple intermediate prediction blocks may be blocks improved by unidirectional prediction based on a neural network.

[0805] The bidirectional prediction based on the neural network may be performed based on the attributes and / or coding parameters of the target block.

[0806] It may be determined whether to perform the bidirectional prediction based on the neural network based on the attributes and / or coding parameters of the target block.

[0807] Here, the attributes of the target block may include the block size of the target block and the components of the target block.

[0808] The prediction for the target block may be performed using a bidirectional prediction method selected from multiple bidirectional prediction methods. Here, the multiple bidirectional prediction methods may include the bidirectional prediction based on the neural network. In other words, the bidirectional prediction based on the neural network may be selected from multiple bidirectional prediction methods according to specific conditions.

[0809] The multiple bidirectional prediction methods may also include one or more of average bidirectional prediction, optical flow-based bidirectional prediction, and weighted average-based bidirectional prediction, which will be described later.

[0810] When determining the bidirectional prediction method, the attributes of the target block and / or coding parameters may be used. For example, the bidirectional prediction method may be determined based on the block size of the target block.

[0811] When determining the bidirectional prediction method, the attributes of the bidirectional prediction and / or coding parameters related to the bidirectional prediction may be used. For example, the bidirectional prediction method may be determined based on the weight of the bidirectional prediction.

[0812] In step S2030, a prediction block may be generated by performing prediction on the target block using prediction information.

[0813] Furthermore, a reconstructed block may be generated by summing the prediction block and the recovered residual block.

[0814] Unidirectional prediction and bidirectional prediction In an embodiment, unidirectional prediction may refer to a method for generating a prediction image (i.e., a prediction picture, a prediction frame, a prediction block, a prediction tile, or a prediction parallel block) using one reference image (i.e., one reference picture, one reference frame, one reference block, one reference tile, or one reference parallel block).

[0815] In an embodiment, bidirectional prediction may refer to a method for generating a predicted image (i.e., a predicted picture, a predicted frame, a predicted block, a predicted tile, or a predicted parallel block) using one or more reference images (i.e., one or more reference pictures, one or more reference frames, one or more reference blocks, one or more reference tiles, or one or more reference parallel blocks).

[0816] To generate a predicted image, not only can a reference image be selected from adjacent images of the target image, but also a reference image can be selected from within the target image.

[0817] In an embodiment, the selection of one or more reference images may also be applied to a predicted image generated by intra prediction.

[0818] In neural network-based unidirectional prediction according to an embodiment, one reference image defined in the unidirectional prediction can be used as an input to the neural network. The neural network can use one reference image defined in the unidirectional prediction to improve the predicted image. The predicted image improved from the neural network-based unidirectional prediction can be used as an input to neural network-based bidirectional prediction.

[0819] In neural network-based bidirectional prediction according to an embodiment, one or more reference images defined in the bidirectional prediction can be used as an input to the neural network. The neural network can use one or more reference images defined in the bidirectional prediction to generate a predicted image.

[0820] The reference image used as an input to neural network-based unidirectional prediction or neural network-based bidirectional prediction may refer to a region of already recovered pixels. For example, a reference image can be obtained from a recovered (or reconstructed) picture stored in a decoded picture buffer (DPB), or alternatively, a reference image can be obtained from a region of already recovered pixels within the current picture. Thus, the input for neural network-based unidirectional prediction or bidirectional prediction may correspond to a recovered image or a reconstructed image.

[0821] Examples of reference images At least one image can be used as a reference image.

[0822] An image can be processed into various forms, and each of the images processed into various forms can be used as a reference image.

[0823] As described above, in the present embodiment, an image may represent a unit of an input and an output of a neural network. And, a reference image for inter-picture prediction may correspond to a unit to which inter-picture prediction is applied, e.g., a block.

[0824] Each generated reference image can be used as an input to the neural network.

[0825] A neighboring image may be an image that is temporally adjacent to the target image. The neighboring image may be an image within a reference picture that is before the target picture including the target block or an image within a reference picture that is after the target picture including the target block. The position of the reference image within the reference picture may correspond to the position of the target image within the target picture.

[0826] The neighboring image may be used as a reference image without involving scaling.

[0827] A scaled neighboring image may be generated by applying scaling to the neighboring image according to various conditions. The scaled neighboring image may be used as a reference image.

[0828] For example, scaling using picture order count (POC) may be applied to specify the position of the neighboring image.

[0829] Another image within the target picture including the target image may be used as a reference image.

[0830] For example, in the case where the target image is a target block, a block adjacent to the target block may be used as a reference image. For example, the reference image may be a (spatially) neighboring block of the target block. Alternatively, the reference image may be a reference block adjacent to the target block. Here, the reference block being adjacent to the target block may mean that the reference block is within a region determined based on the position and size of the target block.

[0831] In one embodiment, scaling may not be applied to the reference block.

[0832] In one embodiment, scaling according to various conditions may be applied to the reference block. For example, the various conditions may include 1) the distance between the target block and the reference block and 2) the prediction mode for the target image.

[0833] Figure 21 is a flowchart showing a prediction method according to an embodiment.

[0834] According to Figure 21 the prediction method of the embodiment may include selecting a prediction method S2110, configuring a prediction mode S2120, and performing prediction S2130.

[0835] From an encoding perspective, the steps S2110 and S2120 may be performed in the step S1910 of Figure 19 showing an encoding method, and the step S2130 may be performed in Figure 19 the step S1920 of.

[0836] From a decoding perspective, the steps S2110 and S2120 may be performed in the step S2020 of the above Figure 20 and the step S2130 may be performed in Figure 20It is executed in step S2030.

[0837] In step S2110, the processing unit may select a prediction method for the target block.

[0838] Here, the prediction method may include a one-way prediction method based on neural network technology.

[0839] Use information indicating whether to perform one-way prediction based on neural network technology can be signaled through the bitstream. If the use information indicates not to use one-way prediction based on neural network technology, then one-way prediction based on neural network technology is not performed, and conventional one-way prediction can be performed.

[0840] One-way prediction based on neural network technology can be performed after conventional one-way prediction. One-way prediction based on neural network technology can improve the predicted block predicted by conventional one-way prediction.

[0841] The processing unit may apply neural network technology to one-way prediction for the target image using the one-way prediction method. The one-way prediction method can determine whether to perform one-way prediction based on neural network technology.

[0842] In step S2110, the prediction method may adopt a method of applying neural network technology for bidirectional prediction. The processing unit can select how to apply neural network technology to bidirectional prediction.

[0843] The method of applying neural network technology to bidirectional prediction may include: 1) The first bidirectional prediction method, which applies neural network technology depending on the weighted average prediction mode index; 2) The second bidirectional prediction method, which applies neural network technology depending on the equal weighted average prediction mode index; 3) The third bidirectional prediction method, which applies neural network technology based on an independent coding structure; and 4) The fourth bidirectional prediction method, which applies an adaptive neural network technology based on an independent coding structure.

[0844] The processing unit may use at least one of the first bidirectional prediction method, the second bidirectional prediction method, the third bidirectional prediction method, and the fourth bidirectional prediction method to select a method of applying neural network technology to bidirectional prediction of the target image.

[0845] The first bidirectional prediction method is a method that applies neural network-based bidirectional prediction based on whether the weights to be used for bidirectional prediction of the target block belong to a set of multiple predefined weights.

[0846] The second bi - directional prediction method is a method in which neural - network - based bi - directional prediction is applied based on whether the weights to be applied to the target block are equal weights. In addition, it is possible to determine whether to apply the first bi - directional prediction method and the second bi - directional prediction method based on the block size. Additionally, the weights for the target block can be signaled (e.g.) at the block level or the strip level.

[0847] In addition, the third bi - directional prediction method is a method in which it is determined whether to apply neural - network - based bi - directional prediction without depending on the weights of the target block. For example, the third bi - directional prediction method determines whether to apply neural - network - based bi - directional prediction according to whether the size of the target block is included in a set of predefined block sizes.

[0848] The fourth bi - directional prediction method is the same as the third bi - directional prediction method; the fourth bi - directional prediction method can be additionally applied not only to the case where a set of multiple predefined weights is used for the weights of the target block, but also to the case where the weights of the target block do not use the set.

[0849] In other words, the fourth bi - directional prediction method can not only perform bi - directional prediction using predefined weights according to specific conditions, but also perform neural - network - based bi - directional prediction (determining weights based on a neural network). At this time, the specific conditions can be methods predefined in the encoder and the decoder, or can be signaled from the encoder to the decoder.

[0850] For example, two methods can be executed to generate bi - directional prediction blocks, and the generated bi - directional prediction blocks can be used to compare encoding costs (e.g., SAT, SSE, SATD, and rate - distortion cost) to determine the optimal bi - directional prediction method. At this time, the determined optimal bi - directional prediction method can be sent from the encoder to the decoder.

[0851] As another example, the encoding / decoding information of neighboring blocks can be used to determine the bi - directional prediction method. At this time, the bi - directional prediction method can be determined according to the prediction method for neighboring blocks. For example, when at least one neighboring block is determined by a neural - network - based bi - directional prediction method or when the number of neighboring blocks determined as the neural - network - based bi - directional prediction method is greater than or equal to a predetermined number or greater than the number of neighboring blocks determined as the weight - based bi - directional prediction method, the neural - network - based method can be determined as the bi - directional prediction method for the corresponding block.

[0852] will be referred to later Figures 24 to 27 to describe the process of determining whether to apply neural - network - based bi - directional prediction according to the first to fourth bi - directional prediction methods.

[0853] The processing unit can determine the range of neural - network techniques for applying bi - directional prediction to the target block.

[0854] When applying bidirectional prediction based on neural network technology to a current block, a processing unit may select one of four ranges for applying the neural network technology and may perform bidirectional prediction based on the neural network technology on the current block according to the selected range. The four ranges for applying the neural network technology may refer to the methods for applying the neural network technology for bidirectional prediction described above.

[0855] Usage information indicating whether to perform unidirectional prediction and bidirectional prediction based on neural network technology may be signaled via a bitstream. If the usage information indicates not to use unidirectional prediction or bidirectional prediction based on neural network technology, then unidirectional prediction or bidirectional prediction based on neural network may not be performed, but instead, conventional unidirectional prediction or bidirectional prediction may be selectively performed.

[0856] Unidirectional prediction or bidirectional prediction based on neural network technology may replace conventional unidirectional prediction or bidirectional prediction. Conventional unidirectional prediction or bidirectional prediction may refer to other unidirectional prediction or bidirectional prediction described above in the embodiments. Conventional unidirectional prediction or bidirectional prediction may refer to other unidirectional prediction or bidirectional prediction for video encoding / decoding. Conventional bidirectional prediction may include average bidirectional prediction, optical flow-based bidirectional prediction, and weighted average-based bidirectional prediction.

[0857] In addition, bidirectional prediction based on neural network technology may be adaptively performed together with conventional bidirectional prediction.

[0858] In the weighted average prediction mode, bidirectional prediction of a target block may be performed for each of the candidates for search.

[0859] In the bidirectional prediction of a target block, two intermediate prediction blocks may be derived from two reference pictures respectively. The two intermediate prediction blocks may be motion compensated prediction blocks generated by motion compensation.

[0860] Motion compensated prediction blocks may be generated by unidirectional prediction or bidirectional prediction.

[0861] The two intermediate prediction blocks may be improved by unidirectional prediction based on neural network technology.

[0862] The two reference pictures may be different from each other. Optionally, the two reference pictures may be the same as each other.

[0863] In the bidirectional prediction for a target block, the two intermediate prediction blocks may be used to generate one (final) prediction block.

[0864] The following Equation 1 describes the generation of a prediction block according to one embodiment.

[0865] [Equation 1]

[0866] Two intermediate prediction blocks may include and .

[0867] may be a first motion-compensated intermediate prediction block. may be a second motion-compensated intermediate prediction block.

[0868] w may be weight information indicating weights for the intermediate prediction blocks. Optionally, w may represent a weight pattern in which a specific weight within a set of weights is used. Optionally, w may be an adaptive weight for the intermediate prediction block.

[0869] Inter-picture prediction may include weight information. w Weights for the intermediate prediction blocks may be determined.

[0870] As described in Equation 1, in the weighted average prediction mode, a weighted average of two intermediate prediction blocks (using w ) may be used to generate a prediction block for the target block.

[0871] may be set for a specific unit w . For example, the specific unit may be the unit described in the above embodiments. Optionally, the specific unit may be a coding unit (CU).

[0872] In the weighted average prediction mode, one weight may be selected from available weights as the weight. The available weights may be, for example, {-2, 3, 4, 5, 10}.

[0873] When the target picture is a low-latency picture, all five available weights may be used as candidates for the search. In other words, one may be selected from {-2, 3, 4, 5, 10} as the weight.

[0874] When the target picture is not a low-latency picture, for example, three of the five weights may be used as candidates for the search. In other words, one may be selected from {3, 4, 5} as the weight.

[0875] When the prediction method selected in step S2110 includes at least one of a first bidirectional prediction method, a second bidirectional prediction method, a third bidirectional prediction method, and a fourth bidirectional prediction method, the processing unit may configure a bidirectional prediction mode for the target block in step S2120.

[0876] The bidirectional prediction mode may be a bidirectional prediction weighted average prediction mode for applying neural network techniques.

[0877] The processing unit can configure a bidirectional prediction mode for applying neural network techniques using at least one of the following: 1) a first search method that searches all weighted average prediction modes for all block sizes, 2) a second search method that searches only equal weighted average prediction modes for all block sizes, 3) a third search method that searches all weighted average prediction modes for a specific block size, 4) a fourth search method that searches only equal weighted average prediction modes for a specific block size, 5) a fifth search method that does not search all weighted average prediction modes for a specific block size, and 6) a sixth search method that does not search all weighted average prediction modes for all block sizes. Here, the block size can refer to the block size of the target block.

[0878] The first search method can be a method that performs bidirectional prediction based on neural network techniques for all weights within all block sizes and weight sets of the target block.

[0879] The second search method can be a method that performs bidirectional prediction based on neural network techniques for all block sizes and equal weights of the target block. Equal weights can mean that the weights for two intermediate prediction blocks are the same. Optionally, equal weights can mean, for example, 4.

[0880] The third search method can be a method that performs bidirectional prediction based on neural network techniques for all weights within the weight set when the target block has a specific block size.

[0881] The fourth search method can be a method that performs bidirectional prediction based on neural network techniques with equal weights when the target block has a specific block size.

[0882] The fifth search method can be to perform bidirectional prediction based on neural network techniques on the target block when the target block has a specific block size.

[0883] The sixth search method can be a method that performs bidirectional prediction based on neural network techniques on the target block.

[0884] Reference will be made to Figures 24 to 27 Describe an example of a process for determining whether to perform neural network-based bidirectional prediction by applying one or more of the first to sixth search methods to each of the above first to fourth bidirectional prediction methods.

[0885] In step S2130, the processing unit can perform unidirectional prediction and bidirectional prediction on the target block.

[0886] The processing unit can determine whether to perform unidirectional prediction using neural network techniques based on the size of the target block.

[0887] When performing a neural network-based unidirectional prediction mode, one motion compensation prediction block after performing conventional unidirectional prediction can be input into the neural network.

[0888] As will be described later with reference to Figure 22 and Figure 23 the neural network technology can be selectively applied to unidirectional prediction according to conditions.

[0889] The processing unit can perform bidirectional prediction on the target block using at least one of the following: 1) bidirectional prediction based on neural network technology, 2) bidirectional prediction based on weighted average, 3) average bidirectional prediction, and 4) bidirectional prediction based on optical flow.

[0890] In an embodiment, when performing bidirectional prediction mode based on neural network, two motion-compensated intermediate prediction blocks can be input into the neural network. In other words, the combination of two motion-compensated prediction blocks can be used as the input of the neural network.

[0891] In an embodiment, when performing bidirectional prediction based on weighted average, adaptive weights can be applied to two motion-compensated intermediate prediction blocks respectively. The weighted average (using adaptive weights) of two intermediate prediction blocks can be used to generate the prediction block of the target block. The prediction block for the target block can be a bidirectional prediction block.

[0892] In an embodiment, when performing average bidirectional prediction, the average value of two motion-compensated intermediate prediction blocks can be used to generate the prediction block of the target block. The prediction block for the target block can be a bidirectional prediction block.

[0893] In an embodiment, when performing bidirectional prediction based on optical flow, the motion vector prediction technology based on optical flow can be applied to the target block.

[0894] The bidirectional motion vector can be obtained in units of blocks through the motion vector prediction technology based on optical flow. Here, the block can be a coding unit (CU). The optimal motion vector can be searched in units of sub-blocks based on the bidirectional motion vector. The optimal motion vector can be derived in units of sub-blocks with a specific size. For example, the specific size can be 4×4. The change of pixels can be estimated from the reference block based on the optimal motion vector. The estimated change can be used to perform correction (or refinement) on the prediction block. Here, the estimation of the change can be performed to improve the prediction error caused by block-by-block prediction.

[0895] The prediction mode candidates can be determined within the search structure of the weighted average prediction mode through steps S2110 and S2120. The prediction mode candidates can be one or more of the following: 1) bidirectional prediction mode based on neural network technology, 2) bidirectional prediction mode based on weighted average, 3) average bidirectional prediction mode, and 4) bidirectional prediction mode based on optical flow.

[0896] Once a candidate prediction mode is determined, the processing unit can select a bidirectional prediction mode from 1) a bidirectional prediction mode based on neural network technology, 2) a bidirectional prediction mode based on weighted average, 3) an average bidirectional prediction mode, and 4) a bidirectional prediction mode based on optical flow; and can perform prediction on the target block according to the selected bidirectional prediction mode.

[0897] As will be referred to later Figure 22 、 Figure 24 、 Figure 25 、 Figure 26 and Figure 27 As described, four prediction methods corresponding to four ways of applying neural network technology to bidirectional prediction can be selected according to conditions, and the selected prediction method can be performed.

[0898] Figure 22 FIG. shows a method for applying neural network technology according to an embodiment.

[0899] Referring to Figure 21 The step S2110 described can include steps S2210, S2215, S2220, S2225, S2230, S2235, S2240, S2245, and S2255.

[0900] In step S2210, a unidirectional prediction method can be performed on the target block. The unidirectional prediction method can be a unidirectional prediction method applying neural network technology.

[0901] In step S2215, it can be determined whether to perform a bidirectional prediction method on the target block.

[0902] In S2220, it can be determined whether to perform a first bidirectional prediction method depending on the weighted average prediction mode index on the target block.

[0903] According to the above determination, step S2225 can be performed when the first bidirectional prediction method is applied, and step S2230 can be performed when the second bidirectional prediction method is not applied.

[0904] In step S2225, a first bidirectional prediction method can be performed on the target block. The first bidirectional prediction method can include applying neural network technology according to the weighted average prediction mode index.

[0905] When performing the first bidirectional prediction method applying neural network technology depending on the weighted average prediction mode index, the bidirectional prediction mode based on neural network can be correspondingly applied to the search for the weighted average prediction mode. The bidirectional prediction mode based on neural network can replace the existing bidirectional prediction mode.

[0906] The bidirectional prediction mode based on neural network can be independently performed within the search structure of the weighted average prediction mode.

[0907] The inter - picture prediction information may include weighted - average prediction mode information.

[0908] When performing a first bi - directional prediction method that applies neural - network technology depending on a weighted - average prediction mode index, the weighted - average prediction mode information may be signaled. The weighted - average prediction mode information may indicate the execution of a first bi - directional prediction method that applies neural - network technology depending on a weighted - average prediction mode index. For example, the weighted - average prediction mode information may be a weighted - average prediction mode coding bit.

[0909] In step S2230, whether to apply a second bi - directional prediction method depending on an equal - weighted - average prediction mode index to a target block.

[0910] According to the above determination, step S2235 may be executed when the second bi - directional prediction method is applied, and step S2240 may be executed when the second bi - directional prediction method is not applied.

[0911] In step S2235, the second bi - directional prediction method may be executed on the target block. The second bi - directional prediction method may include applying neural - network technology depending on an equal - weighted - average prediction mode index.

[0912] When performing a second bi - directional prediction method depending on an equal - weighted - average prediction mode index, a neural - network - based bi - directional prediction mode may be correspondingly applied to the equal - weighted - average prediction mode in the weighted - average prediction mode. The neural - network - based bi - directional prediction mode may replace the existing bi - directional prediction mode.

[0913] The neural - network - based bi - directional prediction mode may be correspondingly executed within a search structure of equal weights in the weighted - average prediction mode.

[0914] Here, equal weights may mean that the weights for two intermediate prediction blocks are the same. Optionally, equal weights may mean, for example, 4.

[0915] The inter - picture prediction information may contain equal - weighted - average prediction mode information.

[0916] When performing a second bi - directional prediction method that applies neural - network technology depending on an equal - weight - average prediction mode index, the equal - weight - average prediction mode information may be signaled. The equal - weighted - average prediction mode information may indicate the execution of a second bi - directional prediction method that applies neural - network technology depending on an equal - weighted - average prediction mode index. For example, the equal - weighted - average prediction mode information may be an equal - weighted - average prediction mode coding bit.

[0917] In step S2240, it may be determined whether to apply a third bi - directional prediction method using an independent coding structure to the target block.

[0918] According to the above determination, when applying the third bidirectional prediction method, step S2245 can be executed, and when not applying the third bidirectional prediction method, step S2255 of applying the fourth bidirectional prediction method to the target block can be executed. The fourth bidirectional prediction method can be an adaptive neural network technology based on an independent coding structure.

[0919] In step S2245, the third bidirectional prediction method can be executed on the target block. The third bidirectional prediction method can include a method for applying a neural network technology using an independent coding structure.

[0920] The neural network-based bidirectional prediction mode can operate independently of the existing bidirectional prediction mode and can replace the existing bidirectional prediction mode.

[0921] The neural network-based bidirectional prediction mode can be executed separately from the weighted average prediction mode search structure.

[0922] The inter-picture prediction information can include independent bidirectional prediction mode information.

[0923] When executing the neural network technology using an independent coding structure, the independent bidirectional prediction mode information can be signaled. The independent bidirectional prediction mode information can indicate the execution of the third bidirectional prediction method that applies the neural network technology using an independent coding structure. The independent bidirectional prediction mode information can be independent bidirectional prediction mode coding bits.

[0924] In step S2255, the fourth bidirectional prediction method can be executed on the target block. The fourth bidirectional prediction method can include a method for applying an adaptive neural network technology based on an independent coding structure.

[0925] The neural network-based bidirectional prediction mode can operate independently of the existing bidirectional prediction mode and can be adaptively executed together with the existing bidirectional prediction mode.

[0926] The neural network-based bidirectional prediction mode can be executed separately from the weighted average prediction mode search structure.

[0927] The inter-picture prediction information can include adaptive bidirectional prediction mode information.

[0928] When executing the adaptive neural network technology based on an independent coding structure, the adaptive bidirectional prediction mode information can be signaled. The adaptive bidirectional prediction mode information can indicate that the fourth bidirectional prediction method applying the adaptive neural network technology based on an independent coding structure is executed. The adaptive bidirectional prediction mode information can be adaptive bidirectional prediction mode coding bits.

[0929] Figure 23 A unidirectional prediction method according to an embodiment is shown.

[0930] Reference Figure 22 The described step S2210 may include Figure 23 The steps S2310, S2315, and S2320 shown.

[0931] In step S2310, a conventional one-way prediction mode may be performed.

[0932] Step S2315 may determine whether there is a block size of the target block among the block sizes to which the neural network technology is applied.

[0933] In an embodiment, cnn_size may represent the block size to which the neural network technology is applied. For example, cnn_size may be one of 128×128, 64×64, or 32×32. Alternatively, cnn_size may have a block size different from the block sizes described in the embodiment.

[0934] In an embodiment, b_size may represent the block size of the target block.

[0935] In an embodiment, the fact that the block size of the target block exists among the block sizes to which the neural network technology is applied may mean that the block size of the target block is the same as one of the block sizes to which the neural network technology is applied.

[0936] If the block size of the target block exists within the block sizes to which the neural network technology is applied, step S2320 may be performed.

[0937] In step S2320, one-way prediction based on neural network technology may be performed on the target block.

[0938] The one-way prediction mode based on neural network technology may improve the predicted block predicted by the existing one-way prediction mode for the target block.

[0939] If the block size of the target block does not correspond to a specific block size, the one-way prediction mode based on neural network may not be used, but only the existing one-way prediction mode may be used.

[0940] Hereinafter, the configuration of the weighted average prediction mode in step S2120 will be described.

[0941] As described above, the weighted average prediction mode may be configured to perform bidirectional prediction for the target block based on the neural network, and one of the above six methods may be selected to configure the weighted average prediction mode.

[0942] Figure 24 A first bidirectional prediction method according to an embodiment is shown.

[0943] Above reference Figure 22The described step S2225 may include Figure 24 the steps shown in Figure 24 (S2410 to S2480). Optionally, steps S2430, S2450, S2470, and S2480 may be included in step S2130. When one of steps S2430, S2450, S2470, and S2480 is executed, step S2225 may terminate.

[0944] In step S2410, it may be determined w whether exists within the weight set.

[0945] In an embodiment, for example, the weight set may be {-2, 3, 4, 5, 10}.

[0946] In an embodiment,[[]] w the fact that exists within the weight set may indicate w has the same value as one of the values within the weight set.

[0947] If w exists within the weight set, step S2420 may be executed. If w does not exist within the weight set, S2480 may be executed.

[0948] Step S2420 may determine whether a first search method of searching all weighted average prediction modes for all block sizes has been executed.

[0949] If it is determined that the first search method is executed, step S2430 may be executed. If it is determined that the first search method is not executed, step S2440 may be executed, which executes a third search method of searching all weighted average prediction modes for a specific block size.

[0950] In step S2430, bidirectional prediction based on neural network technology may be performed on the target block.

[0951] In step S2440, it may be determined whether the block size of the target block exists among the block sizes to which neural network technology is applied.

[0952] In an implementation, cnn_size may represent the block size to which neural network technology is applied. For example, cnn_size may be one of 128×128, 64×64, or 32×32. Optionally, cnn_size may have a block size different from the block sizes described in the embodiment.

[0953] In an embodiment, b_size may represent the block size of the target block.

[0954] In an embodiment, the fact that the block size of the target block exists among the block sizes to which neural network techniques are applied may mean that the block size of the target block is the same as one of the block sizes to which neural network techniques are applied.

[0955] If the block size of the target block exists within the block sizes to which neural network techniques are applied, step S2450 can be executed. If the block size of the target block does not exist within the block sizes to which neural network techniques are applied, step S2460 can be executed.

[0956] In step S2450, bidirectional prediction based on neural network techniques can be performed on the target block.

[0957] In step S2460, it can be determined w whether it is 4.

[0958] In an embodiment, when w is 4, it can be indicated that the weights for two intermediate prediction blocks are the same as each other.

[0959] In an embodiment, the value 4 can be replaced with another value indicating that the weights for two intermediate prediction blocks are the same as each other.

[0960] In an embodiment, when w is 4, it can be indicated that the same weights are used for the weight pattern of two intermediate prediction blocks.

[0961] In an embodiment, when w is 4, it can be indicated that an equal-weighted average prediction mode is used.

[0962] When w is 4, step S2470 can be executed. When w is not 4, step S2480 can be executed.

[0963] In step S2470, average bidirectional prediction or optical flow-based bidirectional prediction can be performed on the target block.

[0964] In step S2480, bidirectional prediction based on weighted average values can be performed on the target block.

[0965] In reference Figure 21 to the S2110 step described, a first bidirectional prediction method of applying neural network techniques depending on the weighted average prediction mode index can be selected. As Figure 22 shown, in step S2220, a first bidirectional prediction method for executing step S2225 can be determined.

[0966] If it is determined that the first bidirectional prediction method is to be executed, the step S2120 may select the first search method of step S2420 or the third search method of step S2440. In the first search method, all weighted average prediction modes are searched for all block sizes. In the third search method, all weighted average prediction modes are searched for a specific block size.

[0967] For example, if in step S2110, the first bidirectional prediction method that applies neural network technology depending on the weighted average prediction mode index of the target block is selected, and in step S2120, a method of searching all weighted average prediction modes is selected for three block sizes (e.g., block sizes of {128×128, 64×64, 32×32}), then the neural network-based bidirectional prediction mode can replace the existing bidirectional prediction modes for the three block sizes.

[0968] When the target picture is a low-latency picture, one weighted average prediction mode can be selected from five weighted average prediction modes (in other words, w ∈ -2, 3, 4, 5, 10). The target block can be encoded and / or decoded according to the selected weighted average prediction mode.

[0969] When the target picture is not a low-latency picture, one weighted average prediction mode can be selected from three weighted average prediction modes (in other words, w ∈ 3, 4, 5). The target block can be encoded and / or decoded according to the selected weighted average prediction mode.

[0970] If the block size of the target block does not correspond to the three block sizes (e.g., block sizes of 128×128, 64×64, or 32×32), then the neural network-based bidirectional prediction mode may not be executed, and only the existing bidirectional prediction mode can be executed.

[0971] Figure 25 The second bidirectional prediction method according to an embodiment is shown.

[0972] Referring to Figure 22 The step S2235 described may include Figure 25 The steps shown (steps S2510 to S2555). Optionally, steps S2525, S2540, S2550, and S2555 may be included in step S2130. If one of steps S2525, S2540, S2550, and S2555 is executed, then step S2235 may terminate. If step S2235 terminates without executing any of steps S2525, S2540, S2550, and S2555, then the existing bidirectional prediction can be performed on the target block.

[0973] In step S2510, it can be determinedw Whether it exists in the weight set.

[0974] If w it exists in the weight set, step S2515 can be executed. If w it does not exist in the weight set, S2235 can be terminated.

[0975] Step S2515 can determine whether to execute a second search method that only searches for the equal weighted average prediction mode for all block sizes.

[0976] If it is determined to execute the second search method, step S2520 can be executed. If it is determined not to execute the second search method, step S2530 can be executed, which executes a fourth search method that only searches for the equal weighted average prediction mode for a specific block size.

[0977] In step S2520, it can be determined w whether it is 4.

[0978] When w it is 4, step S2525 can be executed. When w it is not 4, step S2525 can be terminated.

[0979] In step S2525, bidirectional prediction based on neural network technology can be performed on the target block.

[0980] In step S2530, it can be determined whether the block size of the target block exists in the block sizes to which neural network technology is applied.

[0981] If the block size of the target block exists in the block sizes to which neural network technology is applied, step S2535 can be executed. If the block size of the target block does not exist in the block sizes to which neural network technology is applied, step S2545 can be executed.

[0982] In step S2535, it can be determined w whether it is 4.

[0983] When w it is 4, step S2540 can be executed. When w it is not 4, step S2235 can be terminated.

[0984] In step S2540, bidirectional prediction based on neural network technology can be performed on the target block.

[0985] In step S2545, it can be determined w whether it is 4.

[0986] When w it is 4, step S2550 can be executed. Whenw When it is not 4, step S2555 can be executed.

[0987] In step S2550, average bidirectional prediction or optical flow-based bidirectional prediction can be performed on the target block.

[0988] In step S2555, bidirectional prediction based on a weighted average can be performed on the target block.

[0989] In the reference Figure 21 described step S2110, a second bidirectional prediction method that applies neural network technology depending on the equally weighted average prediction mode index can be selected. As Figure 22 shown, in step S2230, a second bidirectional prediction method for executing step S2235 can be determined.

[0990] If it is determined that the second bidirectional prediction method is to be executed, step S2120 can select the second search method of step S2515 or the fourth search method of step S2530. In the second search method, only the equally weighted average prediction mode is searched for all block sizes. In the fourth search method, only the equally weighted average prediction mode is searched for a specific block size.

[0991] For example, if in step S2110, a second bidirectional prediction method that applies neural network technology depending on the equally weighted average prediction mode index is selected, and in step S2120, the second search method that only searches for the equally weighted average prediction mode for all block sizes is selected, then the neural network-based bidirectional prediction mode can replace the existing bidirectional prediction mode for all block sizes, and the equally weighted average prediction mode (in other words, w the weighted average prediction mode with a weight of 4) can be used to encode / decode the target block. At this time, the existing bidirectional prediction mode cannot be executed.

[0992] Figure 26 A third bidirectional prediction method according to an embodiment is shown.

[0993] Reference Figure 22 described step S2245 can include Figure 26 the steps shown (steps S2610 to S2680). Optionally, steps S2630, S2660, S2670, and S2680 can be included in step S2130. If one of steps S2630, S2660, S2670, and S2680 is executed, step S2245 can terminate. If step S2245 terminates without any of steps S2630, S2660, S2670, and S2680 being executed, the existing bidirectional prediction can be performed on the target block.

[0994] In step S2610, it is possible to determine whether to execute a fifth search method that does not search all weighted average prediction modes for a specific block size.

[0995] If the fifth search method is executed according to the above determination, step S2620 can be executed. If the fifth search method is not executed according to the above determination, step S2680 can be executed, which executes a sixth search method that does not search all weighted average prediction modes for all block sizes.

[0996] In step S2620, it is possible to determine whether the block size of the target block exists among the block sizes to which the neural network technology is applied.

[0997] If the block size of the target block exists within the block sizes to which the neural network technology is applied, step S2630 can be executed. If the block size of the target block does not exist within the block sizes to which the neural network technology is applied, step S2640 can be executed.

[0998] In step S2630, bidirectional prediction based on neural network technology can be performed on the target block.

[0999] In step S2640, it is possible to determine w whether it exists within the weight set.

[1000] If w exists within the weight set, step S2650 can be executed. If w does not exist within the weight set, step S2245 can be terminated.

[1001] In step S2650, it is possible to determine w whether it is 4.

[1002] When w is 4, step S2660 can be executed. When w is not 4, step S2670 can be terminated.

[1003] In step S2660, average bidirectional prediction or optical flow-based bidirectional prediction can be performed on the target block.

[1004] In step S2670, bidirectional prediction based on weighted average values can be performed on the target block.

[1005] In step S2680, bidirectional prediction based on neural network technology can be performed on the target block. [...

Claims

1. An image decoding method, comprising: Performing a prediction operation; Generating a restored image based on the image generated by the prediction operation; And Generating an image to be used for the prediction operation based on the restored image, wherein at least one or more steps of performing the prediction operation and generating the image to be used for the prediction operation include: generating an image by using a neural network with two or more images as inputs.

2. The method according to claim 1, wherein, The step of performing the prediction operation includes: generating a prediction block by using two intermediate prediction blocks as inputs of the neural network, and the two intermediate prediction blocks are derived by performing motion compensation on one or more reference pictures.

3. The method according to claim 2, wherein, The step of performing the prediction operation further includes: when only one of the intermediate prediction blocks is available, generating a second intermediate prediction block based on the available first intermediate prediction block.

4. The method according to claim 1, wherein The step of generating an image to be used for prediction includes: generating a reference picture by using two restored images as inputs of the neural network.

5. The method according to claim 4, wherein, The step of generating an image to be used for prediction further includes: adding the generated reference picture to reference picture list 0 and reference picture list 1 that use the same POC as the current picture for performing the prediction operation.

6. The method according to claim 1, further comprising: Obtaining an upsampled image through a neural network, and the neural network uses the restored image and an image obtained by upsampling the restored image through an interpolation filter as inputs.

7. The method according to claim 6, wherein, The step of obtaining the upsampled image additionally uses a prediction image generated by performing the prediction operation as an input of the neural network.

8. The method according to claim 6, wherein The neural network for obtaining the upsampled image uses one or more of quantization parameters, slice information, and block partition information as additional inputs.

9. The method according to claim 6, wherein The neural network for obtaining the upsampled image performs pixel de-shuffling on an image upsampled to the resolution of the restored image through the interpolation filter for use as an input, or incorporates the upsampled image into the output of the neural network.

10. The method according to claim 9, wherein, An image generated by decomposing the image upsampled through the interpolation filter into high-frequency components and low-frequency components in the horizontal and vertical directions is used as an input of the neural network for obtaining the upsampled image.

11. The method according to claim 1, wherein, The step of generating the image through the neural network includes: when generating an image of a chrominance component, using an image of a luminance component corresponding to the chrominance component as an additional input.

12. An image encoding method, comprising: Downsampling an input image based on a scaling factor; Performing a prediction operation on the downsampled image; Generating a residual image based on the image generated by the prediction operation and encoding the residual image; Decoding the encoded residual image to generate a restored image; And Generating an image to be used for the prediction operation based on the restored image, wherein at least one or more steps of performing the prediction operation and generating the image to be used for the prediction operation include: generating an image by using a neural network with two or more images as inputs.

13. The method according to claim 12, wherein, The step of downsampling includes: Determine the scaling factor based on a rescaled picture, the rescaled picture being obtained by encoding a first picture in a GOP (Group of Pictures) composed of multiple pictures while reducing the resolution of the first picture, decoding the encoded first picture, and then upsampling the decoded first picture; and Downsample all the pictures included in the GOP with the scaling factor determined for the first picture.

14. The method according to claim 13, wherein, The scaling factor is determined as one of multiple values.

15. A computer-readable storage medium storing a bitstream of image information, wherein, The image information is generated by an image encoding method, the image encoding method comprising: Downsample an input image based on a scaling factor; Perform a prediction operation on the downsampled image; Generate a residual image based on the image generated by the prediction operation and encode the residual image; Decode the encoded residual image to generate a restored image; and Generate an image to be used for the prediction operation based on the restored image, wherein performing at least one or more of the steps of the prediction operation and generating the image to be used for the prediction operation includes: generating an image by a neural network that uses two or more images as inputs.