Method and apparatus for image encoding / decoding, and recording medium
By performing a fill operation on the external area of the image and using prediction information to fill the target block, the problem of inefficient encoding/decoding of high resolution and high definition image is solved, and the accuracy and quality of image boundary processing are improved.
Patent Information
- Application Number
- CN202380085901.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-13
- Publication Date
- 2025-07-22
AI Technical Summary
Existing image encoding/decoding technologies are difficult to effectively process the boundary areas of high-resolution and high-definition images, resulting in ineffective encoding/decoding.
By performing a fill operation on an external area of the image, the target block is filled with prediction information, including filling using fixed values, pixel values or motion information, and is suitable for prediction operations before and after in-loop filtering or at boundaries.
Improves encoding/decoding efficiency of high resolution and high definition images, and enhances the accuracy and quality of image boundary processing.
Smart Images

Figure CN120359755A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to methods, devices, and storage media for image encoding / decoding. More specifically, the present disclosure relates to methods, devices, and storage media for image encoding / decoding using prediction.
[0002] This application claims the benefit of Korean Patent Application Nos. 10-2022-0131595, filed on October 13, 2022, and 10-2023-0136967, filed on October 13, 2023, the entire contents of which are incorporated herein by reference. Background Art
[0003] With the continuous development of the information and communication industry, broadcast services supporting high-definition (HD) resolution have become widespread worldwide. Through this widespread use, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.
[0004] To meet the users' demand for high definition, a large number of institutions have accelerated the development of next-generation imaging devices. In addition to high-definition TVs (HDTVs) and full high-definition (FHD) TVs, users' interest in UHD TVs has also increased, where the resolution of UHD TVs is more than four times that of FHD TVs. With the increase in their interest, there is now a need for image encoding / decoding techniques for images with higher resolution and higher clarity.
[0005] As image compression techniques, there are various techniques such as inter-frame prediction techniques, intra-frame prediction techniques, transformation, quantization techniques, and entropy encoding techniques.
[0006] The inter-frame prediction technique is a technique for predicting the values of pixels included in the current picture using the picture before the current picture and / or the picture after the current picture. The intra-frame prediction technique is a technique for predicting the values of pixels included in the current picture using information about the pixels in the current picture. The transformation and quantization techniques may be techniques for compressing the energy of the residual signal. The entropy encoding technique is a technique for assigning short codewords to frequently occurring values and long codewords to less frequently occurring values.
[0007] By using these image compression techniques, data about images can be effectively compressed, transmitted, and stored. Summary of the Invention
[0008] Technical Problem
[0009] Embodiments are intended to provide a device, method, and storage medium for performing encoding / decoding on an image using prediction.
[0010] An embodiment aims to provide a device, a method, and a storage medium for performing padding on a region outside an image in inter-frame prediction.
[0011] Technical solution
[0012] According to one aspect, there is provided an image decoding method, the image decoding method including: determining prediction information of a target block; and performing prediction on the target block using the prediction information, wherein padding is performed on a target image related to the prediction.
[0013] The padding may be performed immediately before performing in-loop filtering on the target image, immediately after applying a specific filter among a plurality of filters in the in-loop filter to the target image, immediately after performing in-loop filtering on the target image, or when performing the prediction on a block located at the boundary of the target image.
[0014] Padding may be performed on a region outside the target image.
[0015] The region outside the target image may be used as a reference block for a block in an additional image, a part of a reference block for a block in an additional image, or a target to which an interpolation filter is applied in motion compensation in inter-frame prediction of a block in an additional image.
[0016] The padding may be performed in the following manner: performing padding using a fixed value, performing padding based on pixel values in the target image, or performing padding based on motion information in the target image.
[0017] Correction of pixel values of pixels in a padding block of the target image may be performed.
[0018] In the correction, a boundary strength at the boundary of the padding block may be calculated.
[0019] According to another aspect, there is provided an image encoding method, the image encoding method including: determining prediction information of a target block; and performing prediction on the target block using the prediction information, wherein padding is performed on a target image related to the prediction.
[0020] The padding may be performed immediately before performing in-loop filtering on the target image, immediately after applying a specific filter among a plurality of filters in the in-loop filter to the target image, immediately after performing in-loop filtering on the target image, or when performing the prediction on a block located at the boundary of the target image.
[0021] Padding may be performed on a region outside the target image.
[0022] The region outside the target image can be used as a reference block for a block in an additional image, a part of a reference block for a block in an additional image, or a target to which an interpolation filter is applied in motion compensation in inter prediction of a block in an additional image.
[0023] The padding can be performed in the following manner: perform padding using a fixed value, perform padding based on pixel values in the target image, or perform padding based on motion information in the target image.
[0024] Correction of pixel values of pixels in a padding block of the target image can be performed.
[0025] In the correction, a boundary strength at a boundary of the padding block can be calculated.
[0026] According to another aspect, there is provided a computer-readable storage medium for storing a bitstream for image decoding, where the bitstream includes prediction information, prediction for a target block is performed using the prediction information, and padding is performed on a target image related to the prediction.
[0027] The padding can be performed immediately before in-loop filtering of the target image, immediately after applying a specific filter among a plurality of filters in the in-loop filter to the target image, immediately after in-loop filtering of the target image, or when the prediction is performed on a block located at a boundary of the target image.
[0028] Padding can be performed on a region outside the target image.
[0029] The region outside the target image can be used as a reference block for a block in an additional image, a part of a reference block for a block in an additional image, or a target to which an interpolation filter is applied in motion compensation in inter prediction of a block in an additional image.
[0030] The padding can be performed in the following manner: perform padding using a fixed value, perform padding based on pixel values in the target image, or perform padding based on motion information in the target image.
[0031] Correction of pixel values of pixels in a padding block of the target image can be performed.
[0032] Advantageous Effects
[0033] There is provided an apparatus, method, and storage medium for performing encoding / decoding of an image using prediction.
[0034] There is provided an apparatus, method, and storage medium for performing padding on a region outside an image in inter prediction. Description of the Drawings
[0035] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present disclosure is applied;
[0036] Figure 2 is a block diagram showing the configuration of an embodiment of a decoding device to which the present disclosure is applied;
[0037] Figure 3 is a diagram schematically showing the partitioning structure of an image when the image is encoded and decoded;
[0038] Figure 4 is a diagram showing the forms of prediction units (PUs) that an encoding unit (CU) can include;
[0039] Figure 5 is a diagram showing the forms of transform units (TUs) that can be included in a CU;
[0040] Figure 6 shows the division of blocks according to an example;
[0041] Figure 7 is a diagram for explaining an embodiment of the intra prediction process;
[0042] Figure 8 is a diagram showing reference sample points used in the intra prediction process;
[0043] Figure 9 is a diagram for explaining an embodiment of the inter prediction process;
[0044] Figure 10 shows spatial candidates according to an embodiment;
[0045] Figure 11 shows the order of adding the motion information of spatial candidates to the merge list according to an embodiment;
[0046] Figure 12 shows the transform and quantization processing according to an example;
[0047] Figure 13 shows the diagonal scan according to an example;
[0048] Figure 14 shows the horizontal scan according to an example;
[0049] Figure 15 shows the vertical scan according to an example;
[0050] Figure 16 is a configuration diagram of an encoding device according to an embodiment;
[0051] Figure 17 is a configuration diagram of a decoding device according to an embodiment;
[0052] Figure 18 is a flowchart showing a target block prediction method and a bitstream generation method according to an embodiment;
[0053] Figure 19 is a flowchart showing a target block prediction method using a bitstream according to an embodiment;
[0054] Figure 20 shows inter - prediction including OOP pixels according to an example;
[0055] Figure 21 shows an OOP boundary in a target block according to an example;
[0056] Figure 22 shows a case where an OOP boundary in a target block is divided into two or more independent OOP boundaries according to an example;
[0057] Figure 23 shows a padding region of a target image according to an example;
[0058] Figure 24 shows another padding region of a target image according to an example;
[0059] Figure 25 shows repeated padding of a padding region of a target image according to an example;
[0060] Figure 26 shows a first mirror padding of a padding region of a target image according to an example;
[0061] Figure 27 shows a second mirror padding of a padding region of a target image according to an example;
[0062] Figure 28 shows a first determination method of a padding block size according to an example;
[0063] Figure 29 shows a second determination method of a padding block size according to an example;
[0064] Figure 30 shows the execution of two adjacent padding blocks according to an example;
[0065] Figure 31 shows another execution of two adjacent padding blocks according to an example;
[0066] Figure 32 shows a first division method for dividing a padding region into two regions according to an example;
[0067] Figure 33Illustrates a second partitioning method for partitioning a filled region into two regions according to an example;
[0068] Figure 34 Illustrates the direction perpendicular to the image boundary of a specific filling block and / or specific pixels in the filled region of the target image according to an example;
[0069] Figure 35 Illustrates a bidirectional motion compensation block according to an example;
[0070] Figure 36 Illustrates a refined bidirectional motion compensation block according to an example;
[0071] Figure 37 Illustrates a filling method for a region outside the target picture according to an example;
[0072] Figure 38 Illustrates a motion compensation filling method according to an example; and
[0073] Figure 39 Illustrates the derivation of an M×4 filling block with a left filling direction. Detailed Description
[0074] The present invention can be variously changed and can have various embodiments. Specific embodiments will be described in detail below with reference to the drawings. However, it should be understood that these embodiments are not intended to limit the present invention to a specific disclosed form, and they include all changes, equivalent forms, or modified forms included within the spirit and scope of the present invention.
[0075] The following exemplary embodiments will be described in detail with reference to the drawings showing specific embodiments. These embodiments are described so that those of ordinary skill in the art to which the present disclosure pertains can easily implement these embodiments. It should be noted that the various embodiments are different from each other but do not need to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be implemented as the other embodiments without departing from the spirit and scope of the other embodiments related to one embodiment. In addition, it should be understood that the positions or arrangements of the respective components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Therefore, the appended detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is only defined by the appended claims and their equivalents (as long as they are properly described).
[0076] In the drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, sizes, etc. of the components in the drawings may be exaggerated for clarity of description.
[0077] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by such terms. The terms are only used to distinguish one component from another. For example, without departing from the scope of this specification, the first component may be referred to as the second component. Similarly, the second component may be referred to as the first component. The term "and / or" may include a combination of multiple related description items or any one of the multiple related description items.
[0078] It will be understood that when a component is referred to as "connected" or "coupled" to another component, the two components may be directly connected or coupled to each other, or there may be an intermediate component between the two components. On the other hand, it will be understood that when a component is referred to as "directly connected or coupled", there is no intermediate component between the two components.
[0079] The components described in the embodiments are independently shown to indicate different characteristic functions, but this does not mean that each component is formed by a single piece of hardware or software. That is, for convenience of description, multiple components are separately arranged and included. For example, at least two of the multiple components may be integrated into a single component. On the contrary, one component may be divided into multiple components. As long as it does not depart from the essence of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included in the scope of this specification.
[0080] The terms used in the embodiments are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless specifically stated to the contrary in the context. In the embodiments, it should be understood that terms such as "comprising" or "having" are only intended to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. That is, in the embodiments, the expression describing a component "comprising" a specific component means that additional components may be included within the scope of the practice of the present invention or the technical spirit of the present invention, but does not exclude the presence of components other than the specific component.
[0081] In the embodiments, the term "at least one" may mean one of one or more quantities (such as 1, 2, 3, and 4). In the embodiments, the term "a plurality" may mean one of two or more quantities (such as 2, 3, and 4).
[0082] Some components of the embodiments are not essential components for performing necessary functions, but may be optional components only for improving performance. An embodiment may be implemented only using the essential components for realizing the essence of the embodiment. For example, a structure including only essential components (excluding optional components only for improving performance) is also included in the scope of the embodiments.
[0083] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings, so that those of ordinary skill in the art to which the embodiments belong can easily implement the embodiments. In the following description of the embodiments, a detailed description of well-known functions or configurations that are considered to obscure the gist of the present specification will be omitted. In addition, the same reference numerals are used throughout the drawings to designate the same components, and repeated descriptions of the same components will be omitted.
[0084] Hereinafter, an "image" may represent a single frame constituting a video, or may represent the video itself. For example, "encoding and / or decoding of an image" may represent "encoding and / or decoding of a video", and may also represent "encoding and / or decoding of any one of a plurality of images constituting a video".
[0085] Hereinafter, the terms "video" and "moving picture" may be used with the same meaning and may be used interchangeably with each other.
[0086] Hereinafter, a target image may be an encoding target image that is a target to be encoded and / or a decoding target image that is a target to be decoded. In addition, the target image may be an input image input to an encoding device or an input image input to a decoding device. And, the target image may be a current image, that is, a target currently to be encoded and / or decoded. For example, the terms "target image" and "current image" may be used with the same meaning and may be used interchangeably with each other.
[0087] Hereinafter, the terms "image", "frame", "picture" and "screen" may be used with the same meaning and may be used interchangeably with each other.
[0088] Hereinafter, a target block may be an encoding target block (i.e., a target to be encoded) and / or a decoding target block (i.e., a target to be decoded). In addition, the target block may be a current block, that is, a target currently to be encoded and / or decoded. Here, the terms "target block" and "current block" may be used with the same meaning and may be used interchangeably with each other. The current block may represent an encoding target block that is an encoding target during encoding and / or a decoding target block that is a decoding target during decoding. In addition, the current block may be at least one of an encoding block, a prediction block, a residual block, and a transform block.
[0089] Hereinafter, the terms "block" and "unit" may be used with the same meaning and may be used interchangeably with each other. Alternatively, a "block" may represent a specific unit.
[0090] Hereinafter, the terms "region" and "segment" may be used interchangeably with each other.
[0091] In the following embodiments, specific information, data, flags, indexes, elements, and attributes may have their respective values. The value "0" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate false, logical false, or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value may be used interchangeably with each other. The value "1" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate true, logical true, or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value may be used interchangeably with each other.
[0092] When a variable such as i or j is used to indicate a row, column, or index, the value of i may be an integer 0 or an integer greater than 0, or may be an integer 1 or an integer greater than 1. In other words, in an embodiment, each of the row, column, and index may be counted starting from 0, or may be counted starting from 1.
[0093] In an embodiment, the term "one or more" or the term "at least one" may represent the term "multiple". The term "one or more" or the term "at least one" may be used interchangeably with "multiple".
[0094] Hereinafter, terms to be used in the embodiments will be described.
[0095] Encoder: An encoder represents a device for performing encoding. That is, an encoder may represent an encoding device.
[0096] Decoder: A decoder represents a device for performing decoding. That is, a decoder may represent a decoding device.
[0097] Unit: A unit may represent a unit of image encoding and decoding. The terms "unit" and "block" may be used with the same meaning and may be used interchangeably with each other.
[0098] – A unit may be an M×N sample array. Each of M and N may be a positive integer. A unit generally may represent a two-dimensional sample array.
[0099] – In the process of image encoding and decoding, a "unit" may be a region generated by partitioning an image. In other words, a "unit" may be a region specified in an image. A single image may be partitioned into multiple units. Alternatively, an image may be partitioned into sub-parts, and a unit may represent each of the partitioned sub-parts when encoding or decoding is performed on the partitioned sub-parts.
[0100] – In the process of image encoding and decoding, predefined processing may be performed on each unit according to the type of the unit.
[0101] – According to functions, unit types can be classified as macro units, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, according to functions, a unit can represent a block, a macro block, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc. For example, a target unit that is a target of encoding and / or decoding can be at least one of a CU, a PU, a residual unit, and a TU.
[0102] – The term “unit” can represent a block including a luma component block, a chroma component block corresponding to the luma component block, and information on syntax elements for each block, such that the unit is designated as being distinguishable from a block.
[0103] – The size and shape of a unit can be implemented differently. In addition, a unit can have any one of various sizes and shapes. Specifically, the shape of a unit can include not only a square but also geometric shapes (such as a rectangle, a trapezoid, a triangle, and a pentagon) that can be represented in two dimensions (2D).
[0104] – In addition, unit information can include one or more of the type of the unit, the size of the unit, the depth of the unit, the encoding order of the unit, and the decoding order of the unit, etc. For example, the type of a unit can indicate one of a CU, a PU, a residual unit, and a TU.
[0105] – A unit can be partitioned into sub-units, each sub-unit having a size smaller than that of the relevant unit.
[0106] Depth: The depth can represent the degree to which a unit is partitioned. In addition, the depth of a unit can indicate the level at which the corresponding unit exists when the unit is represented by a tree structure.
[0107] – Unit partition information can include a depth indicating the depth of the unit. The depth can indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned.
[0108] – In a tree structure, it can be considered that the depth of the root node is the smallest and the depth of the leaf node is the largest. The root node can be the highest (top) node. The leaf node can be the lowest node.
[0109] – A single unit can be hierarchically partitioned into multiple sub-units, while the single unit has depth information based on a tree structure. In other words, a unit and sub-units generated by partitioning the unit can respectively correspond to a node and the children nodes of the node. Each partitioned sub-unit can have a unit depth. Since the depth indicates the number of times the unit is partitioned and / or the degree to which the unit is partitioned, the partition information of the sub-unit can include information on the size of the sub-unit.
[0110] In a tree structure, the top node may correspond to the initial node before partitioning. The top node may be referred to as the "root node". In addition, the root node may have the minimum depth value. Here, the depth of the top node may be level "0".
[0111] – A node with a depth of level "1" may represent the unit generated when the initial unit is partitioned once. A node with a depth of level "2" may represent the unit generated when the initial unit is partitioned twice.
[0112] – A leaf node with a depth of level "n" may represent the unit generated when the initial unit is partitioned n times.
[0113] – A leaf node may be the bottom node that cannot be further partitioned. The depth of the leaf node may be the maximum level. For example, the predefined value for the maximum level may be 3.
[0114] – QT depth may represent the depth for quad-partitioning. BT depth may represent the depth for binary partitioning. TT depth may represent the depth for ternary partitioning.
[0115] – Sample point: A sample point may be the basic unit that constitutes a block. The sample point may be represented by values from 0 to 2 Bd -1 according to the bit depth (Bd).
[0116] – A sample point may be a pixel or a pixel value.
[0117] – Hereinafter, the terms "pixel" and "sample point" may be used with the same meaning and may be used interchangeably with each other.
[0118] Coded tree unit (CTU): A CTU may be composed of a single luma component (Y) coded tree block and two chroma component (i.e., Cb, Cr) coded tree blocks related to the luma component coded tree block. In addition, a CTU may represent the information including the above blocks and the syntax elements for each block.
[0119] – Each coded tree unit (CTU) may be partitioned using one or more partitioning methods such as quadtree (QT), binary tree (BT), and ternary tree (TT) to configure sub-units such as coding units, prediction units, and transform units. A quadtree may represent a quaternary tree. In addition, each coded tree unit may be partitioned using one or more partitioning methods with a multi-type tree (MTT).
[0120] – "CTU" may be used as a term to specify the pixel block that is the processing unit in image decoding and encoding processes (such as in the case of partitioning an input image).
[0121] Coding Tree Block (CTB): "CTB" can be used as a term to specify any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.
[0122] Neighboring block: A neighboring block (or adjacent block) can represent a block adjacent to a target block. A neighboring block can represent a reconstructed neighboring block.
[0123] Hereinafter, the terms "neighboring block" and "adjacent block" can be used with the same meaning and can be used interchangeably with each other.
[0124] A neighboring block can represent a reconstructed neighboring block.
[0125] Spatial neighboring block: A spatial neighboring block can be a block that is spatially adjacent to a target block. A neighboring block can include a spatial neighboring block.
[0126] – The target block and the spatial neighboring block can be included in the target picture.
[0127] – A spatial neighboring block can represent a block whose boundary touches the target block or a block located within a predetermined distance from the target block.
[0128] – A spatial neighboring block can represent a block adjacent to a vertex of the target block. Here, a block adjacent to a vertex of the target block can represent a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block.
[0129] Temporal neighboring block: A temporal neighboring block can be a block that is temporally adjacent to a target block. A neighboring block can include a temporal neighboring block.
[0130] – A temporal neighboring block can include a col block.
[0131] – A col block can be a block in a previously reconstructed collocated picture (col picture). The position of the col block in the col picture can correspond to the position of the target block in the target picture. Alternatively, the position of the col block in the col picture can be equal to the position of the target block in the target picture. The col picture can be a picture included in a reference picture list.
[0132] – A temporal neighboring block can be a block that is temporally adjacent to a spatial neighboring block of the target block.
[0133] Prediction mode: A prediction mode can be information indicating a mode for intra prediction or a mode for inter prediction.
[0134] Prediction unit: A prediction unit can be a basic unit for prediction (such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation).
[0135] – A single prediction unit can be divided into multiple partitions or sub-prediction units with smaller sizes. The multiple partitions can also be the basic units during prediction or compensation. The partitions generated by dividing a prediction unit can also be prediction units.
[0136] Prediction unit partition: A prediction unit partition can be the shape into which a prediction unit is divided.
[0137] Reconstructed neighboring unit: A reconstructed neighboring unit can be a unit that has been decoded and reconstructed and is neighboring to a target unit.
[0138] – A reconstructed neighboring unit can be a unit that is spatially adjacent to the target unit or temporally adjacent to the target unit.
[0139] – A reconstructed spatial neighboring unit can be a unit that has been reconstructed through encoding and / or decoding and is included in the target picture.
[0140] – A reconstructed temporal neighboring unit can be a unit that has been reconstructed through encoding and / or decoding and is included in a reference picture. The position of the reconstructed temporal neighboring unit in the reference picture can be the same as the position of the target unit in the target picture, or can correspond to the position of the target unit in the target picture. In addition, a reconstructed temporal neighboring unit can be a block neighboring to a corresponding block in the reference picture. Here, the position of the corresponding block in the reference picture can correspond to the position of the target block in the target picture. Here, the fact that the positions of the blocks correspond to each other can mean that the positions of the blocks are the same as each other, can mean that one block is included in another block, or can mean that one block occupies a specific position in another block.
[0141] Sub-picture: A picture can be divided into one or more sub-pictures. A sub-picture can be composed of one or more parallel block rows and one or more parallel block columns.
[0142] – A sub-picture can be an area in the picture with a square shape or a rectangular (i.e., non-square rectangle) shape. In addition, a sub-picture can include one or more CTUs.
[0143] – A sub-picture can be a rectangular area of one or more stripes in the picture.
[0144] – A sub-picture can include one or more parallel blocks, one or more bricks, and / or one or more stripes.
[0145] Parallel block: A parallel block can be an area in the picture with a square shape or a rectangular (i.e., non-square rectangle) shape.
[0146] – A parallel block can include one or more CTUs.
[0147] – A parallel block can be partitioned into one or more chunks.
[0148] Chunk: A chunk can represent one or more CTU rows in a parallel block.
[0149] – A parallel block can be partitioned into one or more chunks. Each chunk can include one or more CTU rows.
[0150] – A parallel block that is not partitioned into two parts can also represent a chunk.
[0151] Strip: A strip can include one or more parallel blocks in a picture. Optionally, a strip can include one or more chunks in a parallel block.
[0152] – A sub-picture can include one or more strips that jointly cover a rectangular area of the picture. Thus, each sub-picture boundary is always a strip boundary, and each vertical sub-picture boundary is always a vertical parallel block boundary.
[0153] Parameter set: A parameter set can correspond to header information in the internal structure of a bitstream.
[0154] – A parameter set can include at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a decoding parameter set (DPS), etc.
[0155] – The information signaled by each parameter set can be applied to the pictures that reference the corresponding parameter set. For example, the information in the VPS can be applied to the pictures that reference the VPS. The information in the SPS can be applied to the pictures that reference the SPS. The information in the PPS can be applied to the pictures that reference the PPS.
[0156] – Each parameter set can reference a higher parameter set. For example, the PPS can reference the SPS. The SPS can reference the VPS.
[0157] – In addition, a parameter set can include a parallel block group, strip header information, and parallel block header information. A parallel block group can be a group that includes multiple parallel blocks. In addition, the meaning of “parallel block group” can be the same as the meaning of “strip”.
[0158] Rate-distortion optimization: An encoding device can use rate-distortion optimization to provide high encoding efficiency by utilizing a combination of the following: the size of a coding unit (CU), a prediction mode, the size of a prediction unit (PU), motion information, and the size of a transform unit (TU).
[0159] – A rate-distortion optimization scheme can calculate the rate-distortion cost of each combination to select an optimal combination from these combinations. The equation “D + λ * R” can be used to calculate the rate-distortion cost. Generally, the combination that minimizes the rate-distortion cost can be selected as the optimal combination under the rate-distortion optimization scheme.
[0160] – D may represent distortion. D can be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in a transform unit (i.e., the mean square error).
[0161] – R may represent the rate, which can represent the bit rate using relevant context information.
[0162] – λ represents the Lagrange multiplier. R can include not only coding parameter information (such as prediction mode, motion information, and coding block flags), but also the bits generated due to coding the transform coefficients.
[0163] – The encoding device can perform processes such as inter prediction and / or intra prediction, transformation, quantization, entropy coding, inverse quantization (dequantization), and / or inverse transformation to calculate accurate D and R. These processes greatly increase the complexity of the encoding device.
[0164] – Bitstream: The bitstream can represent a stream of bits including encoded image information.
[0165] Parsing: Parsing can be a determination of the value of a syntax element made by performing entropy decoding on the bitstream. Optionally, the term "parsing" can represent this entropy decoding itself.
[0166] Symbol: A symbol can be at least one of a syntax element, a coding parameter, and a transform coefficient of an encoding target unit and / or a decoding target unit. In addition, a symbol can be the target of entropy coding or the result of entropy decoding.
[0167] Reference picture: A reference picture can be an image that is referenced by a unit to perform inter prediction or motion compensation. Optionally, a reference picture can be an image including a reference unit that is referenced by a target unit to perform inter prediction or motion compensation.
[0168] Hereinafter, the terms "reference picture" and "reference image" can be used with the same meaning and can be used interchangeably with each other.
[0169] Reference picture list: A reference picture list can be a list including one or more reference images used for inter prediction or motion compensation.
[0170] – The types of reference picture lists can include a combined list (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.
[0171] – For inter prediction, one or more reference picture lists can be used.
[0172] Inter - frame prediction indicator: The inter - frame prediction indicator can indicate the inter - frame prediction direction for the target unit. The inter - frame prediction can be one of unidirectional prediction and bidirectional prediction. Optionally, the inter - frame prediction indicator can represent the number of reference pictures used to generate the prediction unit for the target unit. Optionally, the inter - frame prediction indicator can represent the number of prediction blocks used for the inter - frame prediction or motion compensation of the target unit.
[0173] Prediction list utilization flag: The prediction list utilization flag can indicate whether at least one reference picture in a specific reference picture list is used to generate the prediction unit.
[0174] – The prediction list utilization flag can be used to derive the inter - frame prediction indicator. Conversely, the inter - frame prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag indicates "0" (as the first value), it can indicate that for the target unit, the reference pictures in the reference picture list are not used to generate the prediction block. When the prediction list utilization flag indicates "1" (as the second value), it can indicate that for the target unit, the reference picture list is used to generate the prediction unit.
[0175] Reference picture index: The reference picture index can be an index indicating a specific reference picture in the reference picture list.
[0176] Picture Order Count (POC): The POC value of a picture can represent the order of displaying the corresponding picture.
[0177] Motion Vector (MV): The motion vector can be a 2D vector used for inter - frame prediction or motion compensation. The motion vector can represent the offset between the target image and the reference image.
[0178] – For example, the MV can be represented in the form such as (mv x , mv y ). mv x can indicate the horizontal component, and mv y can indicate the vertical component.
[0179] – Search range: The search range can be a 2D area where the search for the MV is performed during inter - frame prediction. For example, the size of the search range can be M×N. M and N can be positive integers respectively.
[0180] Motion vector candidate: The motion vector candidate can be a block that is a prediction candidate when the motion vector is predicted or the motion vector of a block that is a prediction candidate.
[0181] – The motion vector candidate can be included in the motion vector candidate list.
[0182] Motion vector candidate list: The motion vector candidate list can be a list configured using one or more motion vector candidates.
[0183] Motion vector candidate index: The motion vector candidate index may be an indicator for indicating a motion vector candidate in a motion vector candidate list. Optionally, the motion vector candidate index may be an index of a motion vector predictor.
[0184] Motion information: The motion information may be information including at least one of a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index, and a motion vector, a reference picture index, and an inter prediction indicator.
[0185] Merge candidate list: The merge candidate list may be a list configured using one or more merge candidates.
[0186] Merge candidate: The merge candidate may be a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-prediction merge candidate, a history-based candidate, an average merge candidate based on an average of two candidates, a zero merge candidate, etc. The merge candidate may include an inter prediction indicator and may include motion information such as prediction type information, a reference picture index for each list, a motion vector, a prediction list utilization flag, and an inter prediction indicator.
[0187] Merge index: The merge index may be an indicator for indicating a merge candidate in a merge candidate list.
[0188] – The merge index may indicate a reconstructed unit among reconstructed units adjacent to a target unit spatially and temporally for deriving a merge candidate.
[0189] – The merge index may indicate at least one of multiple pieces of motion information of a merge candidate.
[0190] Transform unit: The transform unit may be a basic unit for residual signal encoding and / or residual signal decoding (such as transform, inverse transform, quantization, inverse quantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit may be partitioned into multiple sub-transform units having smaller sizes. Here, the transform may include one or more of a primary transform and a secondary transform, and the inverse transform may include one or more of a primary inverse transform and a secondary inverse transform.
[0191] Scaling: Scaling may represent a process of multiplying a factor by a transform coefficient level.
[0192] – As a result of scaling a transform coefficient level, transform coefficients may be generated. Scaling may also be referred to as "inverse quantization".
[0193] Quantization Parameter (QP): The quantization parameter can be a value used to generate transform coefficient levels for transform coefficients in quantization. Optionally, the quantization parameter can also be a value used to generate transform coefficients by scaling transform coefficient levels in inverse quantization. Optionally, the quantization parameter can be a value mapped to a quantization step size.
[0194] Delta Quantization Parameter: The delta quantization parameter can represent the difference between the quantization parameter of a target unit and the predicted quantization parameter.
[0195] Scanning: Scanning can represent a method of arranging the order of coefficients in a unit, block, or matrix. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array can be referred to as "scanning". Optionally, a method for arranging a 1D array in the form of a 2D array can also be referred to as "scanning" or "inverse scanning".
[0196] Transform Coefficient: The transform coefficient can be a coefficient value generated when a coding device performs a transform. Optionally, the transform coefficient can be a coefficient value generated when a decoding device performs at least one of entropy decoding and inverse quantization.
[0197] – Quantized levels or quantized transform coefficient levels generated by applying quantization to transform coefficients or residual signals can also be included in the meaning of the term "transform coefficient".
[0198] Quantized Level: The quantized level can be a value generated when a coding device performs quantization on transform coefficients or residual signals. Optionally, the quantized level can be a value targeted for inverse quantization when a decoding device performs inverse quantization.
[0199] – Quantized transform coefficient levels that are the result of transform and quantization can also be included in the meaning of the quantized level.
[0200] Non-Zero Transform Coefficient: The non-zero transform coefficient can be a transform coefficient having a value other than 0, or can be a transform coefficient level having a value other than 0. Optionally, the non-zero transform coefficient can be a transform coefficient whose value magnitude is not 0, or can be a transform coefficient level whose value magnitude is not 0.
[0201] Quantization Matrix: The quantization matrix can be a matrix used in the quantization process or inverse quantization process to improve the subjective or objective image quality of an image. The quantization matrix can also be referred to as a "scaling list".
[0202] Quantization Matrix Coefficient: The quantization matrix coefficient can be each element in the quantization matrix. The quantization matrix coefficient can also be referred to as a "matrix coefficient".
[0203] Default matrix: The default matrix can be a quantization matrix predefined by the encoding device and the decoding device.
[0204] Non-default matrix: The non-default matrix can be a quantization matrix not predefined by the encoding device and the decoding device. The non-default matrix can represent a quantization matrix signaled by a user from the encoding device to the decoding device.
[0205] Most Probable Mode (MPM): The MPM can represent an intra prediction mode that is highly probable to be used for intra prediction of a target block.
[0206] The encoding device and the decoding device can determine one or more MPMs based on encoding parameters related to the target block and attributes of entities related to the target block.
[0207] The encoding device and the decoding device can determine one or more MPMs based on the intra prediction mode of a reference block. The reference block can include multiple reference blocks. The multiple reference blocks can include a spatially adjacent block adjacent to the left of the target block and a spatially adjacent block adjacent to the above of the target block. In other words, one or more different MPMs can be determined according to which intra prediction modes have been used for the reference block.
[0208] – One or more MPMs can be determined in the same way in both the encoding device and the decoding device. That is, the encoding device and the decoding device can share the same MPM list including one or more MPMs.
[0209] MPM list: The MPM list can be a list including one or more MPMs. The number of one or more MPMs in the MPM list can be predefined.
[0210] MPM indicator: The MPM indicator can indicate the MPM among one or more MPMs in the MPM list that will be used for intra prediction of the target block. For example, the MPM indicator can be an index for the MPM list.
[0211] – Since the MPM list is determined in the same way in both the encoding device and the decoding device, it may not be necessary to send the MPM list itself from the encoding device to the decoding device.
[0212] – The MPM indicator can be signaled from the encoding device to the decoding device. Since the MPM indicator is signaled, the decoding device can determine the MPM among the MPMs in the MPM list that will be used for intra prediction of the target block.
[0213] MPM usage indicator: The MPM usage indicator can indicate whether the MPM usage mode will be used for prediction of the target block. The MPM usage mode can be a mode of using the MPM list to determine the MPM that will be used for intra prediction of the target block.
[0214] – The MPM usage indicator can be signaled from an encoding device to a decoding device.
[0215] Signaling: "Signaling" can mean that information is sent from an encoding device to a decoding device. Optionally, "signaling" can mean that the encoding device includes the information in a bitstream or a recording medium. The information signaled by the encoding device can be used by the decoding device.
[0216] – The encoding device can generate encoded information by performing encoding on the information to be signaled. The encoded information can be sent from the encoding device to the decoding device. The decoding device can obtain the information by decoding the sent encoded information. Here, the encoding can be entropy encoding, and the decoding can be entropy decoding.
[0217] Selective signaling: Information can be selectively signaled. Selective signaling for information can mean that the encoding device selectively includes the information in the bitstream or the recording medium (according to specific conditions). Selective signaling for information can mean that the decoding device selectively extracts the information from the bitstream (according to specific conditions).
[0218] Omission of signaling: Signaling for information can be omitted. Omission of signaling for information regarding information can mean that the encoding device does not include the information in the bitstream or the recording medium (according to specific conditions). Omission of signaling for information can mean that the decoding device does not extract the information from the bitstream (according to specific conditions).
[0219] Statistical value: A variable, an encoding parameter, a constant, etc. can have a computable value. A statistical value can be a value generated by performing a calculation (operation) on the value of a specified target. For example, a statistical value can indicate one or more of an average value, a weighted average value, a weighted sum, a minimum value, a maximum value, a mode, a median, and an interpolation of the value of a specific variable, a specific encoding parameter, a specific constant, etc.
[0220] Figure 1 is a block diagram showing the configuration of an embodiment of an encoding device to which the present disclosure is applied.
[0221] The encoding device 100 can be an encoder, a video encoding device, or an image encoding device. The video can include one or more images (frames). The encoding device 100 can sequentially encode one or more images of the video.
[0222] Refer to Figure 1, the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.
[0223] The encoding device 100 can perform encoding on a target image using the intra-frame mode and / or the inter-frame mode. In other words, the prediction mode of the target block can be one of the intra-frame mode and the inter-frame mode.
[0224] Hereinafter, the terms "intra-frame mode", "intra-frame prediction mode", "intra-picture mode", and "intra-frame prediction mode" can be used with the same meaning and can be used interchangeably with each other.
[0225] Hereinafter, the terms "inter-frame mode", "inter-frame prediction mode", "inter-picture mode", and "inter-picture prediction mode" can be used with the same meaning and can be used interchangeably with each other.
[0226] Hereinafter, the term "image" can only indicate a partial image or can indicate a block. In addition, the processing of an "image" can indicate the sequential processing of multiple blocks.
[0227] In addition, the encoding device 100 can generate a bitstream including encoded information by encoding a target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed through a wired and / or wireless transmission medium.
[0228] When the intra-frame mode is used as the prediction mode, the switch 115 can switch to the intra-frame mode. When the inter-frame mode is used as the prediction mode, the switch 115 can switch to the inter-frame mode.
[0229] The encoding device 100 can generate a prediction block for a target block. In addition, after the prediction block has been generated, the encoding device 100 can encode the residual block for the target block using the residual between the target block and the prediction block.
[0230] When the prediction mode is the intra-frame mode, the intra-frame prediction unit 120 can use the pixels of previously encoded / decoded neighboring blocks adjacent to the target block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the target block using the reference samples and can generate prediction samples for the target block via spatial prediction. The prediction samples can represent the samples in the prediction block.
[0231] The inter-frame prediction unit 110 can include a motion prediction unit and a motion compensation unit.
[0232] When the prediction mode is an inter-frame mode, the motion prediction unit may search for the region in the reference image that best matches the target block during the motion prediction process, and may derive a motion vector for the target block and the found region based on the found region. Here, the motion prediction unit may use the search range as the target region for the search.
[0233] The reference image may be stored in the reference picture buffer 190. More specifically, when the encoding and / or decoding of the reference image has been processed, the encoded and / or decoded reference image may be stored in the reference picture buffer 190.
[0234] Since the decoded pictures are stored, the reference picture buffer 190 may be a decoded picture buffer (DPB).
[0235] The motion compensation unit may generate a prediction block for the target block by performing motion compensation using the motion vector. Here, the motion vector may be a two-dimensional (2D) vector for inter-frame prediction. In addition, the motion vector may indicate the offset between the target image and the reference image.
[0236] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit may generate a prediction block by applying an interpolation filter to a partial region of the reference image. To perform inter-frame prediction or motion compensation, it may be determined which of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the method for predicting and compensating for the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation may be performed according to the mode.
[0237] The subtractor 125 may generate a residual block, where the residual block is the difference between the target block and the prediction block. The residual block may also be referred to as a "residual signal".
[0238] The residual signal may be the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing the difference between the original signal and the prediction signal or a signal generated by transforming and quantizing the difference. The residual block may be the residual signal for a block unit.
[0239] The transform unit 130 may generate transform coefficients by transforming the residual block and may output the generated transform coefficients. Here, the transform coefficients may be the coefficient values generated by transforming the residual block.
[0240] The transform unit 130 may use one of a plurality of predefined transform methods when performing the transform.
[0241] The multiple predefined transformation methods may include a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), etc.
[0242] The transformation method for transforming the residual block may be determined according to at least one of the coding parameters for the target block and / or neighboring blocks. For example, the transformation method may be determined based on at least one of an inter prediction mode for a PU, an intra prediction mode for a PU, the size of a TU, and the shape of a TU. Alternatively, transformation information indicating the transformation method may be signaled from the coding device 100 to the decoding device 200.
[0243] When using the transform skip mode, the transform unit 130 may omit the operation of transforming the residual block.
[0244] By performing quantization on the transform coefficients, quantized transform coefficient levels or quantized levels may be generated. Hereinafter, in the embodiments, each of the quantized transform coefficient levels and the quantized levels may also be referred to as "transform coefficients".
[0245] The quantization unit 140 may generate quantized transform coefficient levels (i.e., quantized levels or quantized coefficients) by quantizing the transform coefficients according to quantization parameters. The quantization unit 140 may output the generated quantized transform coefficient levels. In this case, the quantization unit 140 may use a quantization matrix to quantize the transform coefficients.
[0246] The entropy coding unit 150 may generate a bitstream by performing entropy coding based on probability distribution on the values calculated by the quantization unit 140 and / or the coding parameter values calculated during the coding process. The entropy coding unit 150 may output the generated bitstream.
[0247] The entropy coding unit 150 may perform entropy coding on information about pixels of an image and information required for decoding the image. For example, the information required for decoding the image may include syntax elements and the like.
[0248] When applying entropy coding, fewer bits may be allocated to more frequently occurring symbols, and more bits may be allocated to less frequently occurring symbols. Since the symbols are represented by this allocation, the size of the bit string for the target symbols to be encoded may be reduced. Therefore, the compression performance of video coding may be improved by entropy coding.
[0249] In addition, for entropy encoding, the entropy encoding unit 150 may use encoding methods such as exponential Golomb, context - adaptive variable - length coding (CAVLC), or context - adaptive binary arithmetic coding (CABAC). For example, the entropy encoding unit 150 may use a variable - length coding / table (VLC) to perform entropy encoding. For example, the entropy encoding unit 150 may derive a binarization method for a target symbol. In addition, the entropy encoding unit 150 may derive a probability model for the target symbol / binary bit. The entropy encoding unit 150 may use the derived binarization method, probability model, and context model to perform arithmetic coding.
[0250] The entropy encoding unit 150 may transform the coefficients in 2D block form into 1D vector form through a transform coefficient scanning method to encode the quantized transform coefficient levels.
[0251] Coding parameters may be information required for encoding and / or decoding. The coding parameters may include information encoded by the encoding device 100 and sent from the encoding device 100 to the decoding device, and may also include information that can be derived during the encoding or decoding process. For example, the information sent to the decoding device may include syntax elements.
[0252] Coding parameters may include not only information such as syntax elements (or flags or indices) encoded by an encoding device and signaled by the encoding device to a decoding device, but also information derived during the encoding or decoding process. Additionally, the coding parameters may include information required for encoding or decoding an image. For example, the coding parameters may include at least one value, a combination of the following items, or statistics: the size of a unit / block, the shape / form of a unit / block, the depth of a unit / block, the partitioning information of a unit / block, the partitioning structure of a unit / block, information indicating whether a unit / block is partitioned in a quadtree structure, information indicating whether a unit / block is partitioned in a binary tree structure, the partitioning direction (horizontal or vertical) of a binary tree structure, the partitioning form (symmetric partitioning or asymmetric partitioning) of a binary tree structure, information indicating whether a unit / block is partitioned in a ternary tree structure, the partitioning direction (horizontal or vertical) of a ternary tree structure, the partitioning form (symmetric partitioning or asymmetric partitioning, etc.) of a ternary tree structure, information indicating whether a unit / block is partitioned in a multi-type tree structure, the combination and direction (horizontal or vertical, etc.) of partitioning in a multi-type tree structure, the partitioning form (symmetric partitioning or asymmetric partitioning, etc.) of a multi-type tree structure, the partitioning tree (binary tree or ternary tree) of a multi-type tree form, the prediction type (intra prediction or inter prediction), the intra prediction mode / direction, the intra luminance prediction mode / direction, the intra chrominance prediction mode / direction, the intra partitioning information, the inter partitioning information, the coding block partitioning flag, the prediction block partitioning flag, the transform block partitioning flag, the reference sample filtering method, the reference sample filter taps, the reference sample filter coefficients, the prediction block filtering method, the prediction block filter taps, the prediction block filter coefficients, the prediction block boundary filtering method, the prediction block boundary filter taps, the prediction block boundary filter coefficients, the inter prediction mode, the motion information, the motion vector, the motion vector difference, the reference picture index, the inter prediction direction, the inter prediction indicator, the prediction list utilization flag, the reference picture list, the reference image, the POC, the motion vector prediction factor, the motion vector prediction index, the motion vector prediction candidate, the motion vector candidate list, information indicating whether the merge mode is used, the merge index, the merge candidate, the merge candidate list, information indicating whether the skip mode is used, the type of interpolation filter, the taps of the interpolation filter, the filter coefficients of the interpolation filter, the size of the motion vector, the precision of the motion vector representation, the transform type, the transform size, information indicating whether the first transform is used, information indicating whether an additional (second) transform is used, the first transform selection information (or the first transform index), the second transform selection information (or the second transform index), information indicating the presence or absence of a residual signal, the coding block style, the coding block flag, the quantization parameter, the residual quantization parameter, the quantization matrix, information about the loop filter, information indicating whether the loop filter is applied, the coefficients of the loop filter, the taps of the loop filter, the shape / form of the loop filter, information indicating whether the deblocking filter is applied,Coefficients of the deblocking filter, taps of the deblocking filter, deblocking filter strength, shape / form of the deblocking filter, information indicating whether adaptive sample offset is applied, value of the adaptive sample offset, category of the adaptive sample offset, type of the adaptive sample offset, information indicating whether the adaptive loop filter is applied, coefficients of the adaptive loop filter, taps of the adaptive loop filter, shape / form of the adaptive loop filter, binarization / de-binarization method, context model, context model determination method, context model update method, information indicating whether the normal mode is executed, information indicating whether the bypass mode is executed, valid coefficient flag, last valid coefficient flag, coding flag of the coefficient group, position of the last valid coefficient, information indicating whether the value of the coefficient is greater than 1, information indicating whether the value of the coefficient is greater than 2, information indicating whether the value of the coefficient is greater than 3, remaining coefficient value information, positive / negative sign information, reconstructed luma samples, reconstructed chroma samples, context bits, bypass bits, residual luma samples, residual chroma samples, transform coefficients, luma transform coefficients, chroma transform coefficients, quantization levels, luma quantization levels, chroma quantization levels, transform coefficient levels, transform coefficient level scanning method, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of times of motion vector search on the decoding device side, size of the CTU, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, slice identification information, slice type, slice partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, picture type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luma signal, information about the chroma signal, color space of the target block and color space of the residual block. In addition, the above information related to the coding parameters may also be included in the coding parameters. Information used to calculate and / or derive the above coding parameters may also be included in the coding parameters. Information calculated or derived using the above coding parameters may also be included in the coding parameters.
[0253] The first transform selection information may indicate the first transform applied to the target block.
[0254] The second transform selection information may indicate the second transform applied to the target block.
[0255] The residual signal may represent the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be the residual signal for the block.
[0256] Here, sending information by a signal may indicate that the encoding device 100 includes in the bitstream the entropy-encoded information generated by performing entropy encoding on a flag or an index, and may indicate that the decoding device 200 obtains information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. Here, the information may include a flag, an index, and the like.
[0257] The signal may mean the information to be sent by the signal. Hereinafter, the information for an image and a block may be referred to as a "signal". In addition, hereinafter, the terms "information" and "signal" may be used to have the same meaning and may be used interchangeably with each other. For example, a specific signal may be a signal representing a specific block. The original signal may be a signal representing a target block. The prediction signal may be a signal representing a prediction block. The residual signal may be a signal representing a residual block.
[0258] The bitstream may include information based on a specific syntax. The encoding device 100 may generate a bitstream including information according to the specific syntax. The decoding device 200 may obtain information from the bitstream according to the specific syntax.
[0259] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image may be used as a reference image for another image to be subsequently processed. Therefore, the encoding device 100 may reconstruct or decode the encoded target image and store the reconstructed or decoded image in the reference picture buffer 190 as a reference image. For decoding, inverse quantization and inverse transformation of the encoded target image may be performed.
[0260] The quantization level may be inverse-quantized by the inverse quantization unit 160 and may be inverse-transformed by the inverse transformation unit 170. The inverse quantization unit 160 may generate inverse-quantized coefficients by performing an inverse transformation on the quantization level. The inverse transformation unit 170 may generate coefficients that have been inverse-quantized and inverse-transformed by performing an inverse transformation on the inverse-quantized coefficients.
[0261] The coefficients that have been inverse-quantized and inverse-transformed may be added to the prediction block by the adder 175. Adding the coefficients that have been inverse-quantized and inverse-transformed and the prediction block may then generate a reconstructed block. Here, the coefficients that have been inverse-quantized and / or inverse-transformed may represent coefficients on which one or more of inverse quantization and inverse transformation have been performed, and may also represent a reconstructed residual block. Here, the reconstructed block may represent a restored block or a decoded block.
[0262] The reconstructed block can be filtered by the filter unit 180. The filter unit 180 can apply one or more of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), and a non-local filter (NLF) to the reconstructed samples, the reconstructed block, or the reconstructed picture. The filter unit 180 may also be referred to as a "loop filter".
[0263] The deblocking filter can remove block distortion that appears at the boundaries between blocks in the reconstructed picture. To determine whether to apply the deblocking filter, the number of columns or rows that are included in the block and that include the pixels based on which it is determined whether to apply the deblocking filter to the target block can be determined.
[0264] When the deblocking filter is applied to the target block, the filter applied can be different according to the intensity of the required deblocking filtering. In other words, among different filters, the filter determined in consideration of the intensity of the deblocking filtering can be applied to the target block. When the deblocking filter is applied to the target block, one or more of a long-tap filter, a strong filter, a weak filter, and a Gaussian filter can be applied to the target block according to the intensity of the required deblocking filtering.
[0265] In addition, when performing vertical filtering and horizontal filtering on the target block, the horizontal filtering and the vertical filtering can be performed in parallel.
[0266] SAO can add an appropriate offset to the pixel value to compensate for the coding error. SAO can perform correction on the image to which deblocking is applied based on pixels, where the correction uses the offset of the difference between the original image and the image to which deblocking is applied. To perform offset correction on the image, a method of dividing the pixels included in the image into a specific number of regions, determining the regions to which the offset will be applied among the divided regions, and applying the offset to the determined regions can be used, and a method of applying the offset in consideration of the edge information of each pixel can also be used.
[0267] ALF can perform filtering based on the value obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether to apply the adaptive loop filter can be signaled for each CU. Such information can be signaled for the luminance signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, an ALF having a fixed form can be applied to the block regardless of the characteristics of the block.
[0268] The non-local filter may perform filtering based on a reconstructed block similar to a target block. A region similar to the target block may be selected from the reconstructed picture, and statistical attributes of the selected similar region may be used to perform filtering of the target block. Information on whether to apply the non-local filter may be signaled for a coding unit (CU). In addition, the shape and filter coefficients of the non-local filter applied to a block may vary according to the block.
[0269] The reconstructed block or the reconstructed image filtered by the filter unit 180 may be stored in the reference picture buffer 190 as a reference picture. The reconstructed block filtered by the filter unit 180 may be part of the reference picture. In other words, the reference picture may be a reconstructed picture composed of the reconstructed blocks filtered by the filter unit 180. The stored reference picture may then be used for inter-frame prediction or motion compensation.
[0270] Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied.
[0271] The decoding device 200 may be a decoder, a video decoding device, or an image decoding device.
[0272] Referring to Figure 2 , the decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, a switch 245, an adder 255, a filter unit 260, and a reference picture buffer 270.
[0273] The decoding device 200 may receive a bitstream output from the encoding device 100. The decoding device 200 may receive a bitstream stored in a computer-readable storage medium, and may receive a bitstream streamed through a wired / wireless transmission medium.
[0274] The decoding device 200 may perform decoding on the bitstream in an intra mode and / or an inter mode. In addition, the decoding device 200 may generate a reconstructed image or a decoded image via decoding, and may output the reconstructed image or the decoded image.
[0275] For example, the operation of switching to the intra mode or the inter mode based on the prediction mode for decoding may be performed by the switch 245. When the prediction mode for decoding is the intra mode, the switch 245 may be operated to switch to the intra mode. When the prediction mode for decoding is the inter mode, the switch 245 may be operated to switch to the inter mode.
[0276] The decoding device 200 can obtain the reconstructed residual block by decoding the input bitstream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as the target to be decoded by adding the reconstructed residual block and the prediction block.
[0277] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols can include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method can be similar to the entropy encoding method described above. That is, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0278] The entropy decoding unit 210 can change the coefficients in the form of a one-dimensional (1D) vector to a 2D block shape by a transform coefficient scanning method in order to decode the quantized transform coefficient levels.
[0279] For example, the coefficients of a block can be changed to a 2D block shape by scanning the block coefficients using a right upper diagonal scan. Alternatively, which one of the right upper diagonal scan, vertical scan, and horizontal scan will be used can be determined according to the size of the corresponding block and / or the intra prediction mode.
[0280] The quantized coefficients can be dequantized by the dequantization unit 220. The dequantization unit 220 can generate dequantized coefficients by performing dequantization on the quantized coefficients. In addition, the dequantized coefficients can be inverse-transformed by the inverse transform unit 230. The inverse transform unit 230 can generate a reconstructed residual block by performing an inverse transform on the dequantized coefficients. As a result of performing dequantization and inverse transform on the quantized coefficients, a reconstructed residual block can be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 can apply a quantization matrix to the quantized coefficients.
[0281] When using the intra mode, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the target block, where the spatial prediction uses the pixel values of previously decoded neighboring blocks adjacent to the target block.
[0282] The inter prediction unit 250 can include a motion compensation unit. Alternatively, the inter prediction unit 250 can be designated as the "motion compensation unit".
[0283] When using the inter mode, the motion compensation unit can generate a prediction block by performing motion compensation on the target block, where the motion compensation uses a motion vector and a reference image stored in the reference picture buffer 270.
[0284] The motion compensation unit may apply an interpolation filter to a partial area of a reference image when the motion vector has a value other than an integer, and may generate a prediction block using the reference image to which the interpolation filter has been applied. To perform motion compensation, the motion compensation unit may determine, based on the CU, which one of the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to the motion compensation method for the PU included in the CU, and may perform motion compensation according to the determined mode.
[0285] The reconstructed residual block and the prediction block may be added to each other by adder 255. Adder 255 may generate a reconstructed block by adding the reconstructed residual block and the prediction block.
[0286] The reconstructed block may be filtered by filter unit 260. Filter unit 260 may apply at least one of a deblocking filter, SAO filter, ALF, and NLF to the reconstructed block or the reconstructed image. The reconstructed image may be a picture including the reconstructed block.
[0287] The filter unit may output the reconstructed image.
[0288] The reconstructed image and / or the reconstructed block filtered by filter unit 260 may be stored as a reference picture in reference picture buffer 270. The reconstructed block filtered by filter unit 260 may be a part of the reference picture. In other words, the reference picture may be an image constituted by the reconstructed blocks filtered by filter unit 260. The stored reference picture may then be used for inter prediction or motion compensation.
[0289] Figure 3 is a diagram schematically showing the partitioning structure of an image when the image is encoded and decoded.
[0290] Figure 3 An example in which a single unit is partitioned into a plurality of sub-units may be schematically shown.
[0291] To partition an image effectively, a coding unit (CU) may be used in encoding and decoding. The term "unit" may be used to commonly specify 1) a block including image samples and 2) syntax elements. For example, "partitioning of a unit" may mean "partitioning of a block corresponding to the unit".
[0292] The CU may be used as a basic unit for image encoding / decoding. The CU may be used as a unit to which one mode selected from an intra mode and an inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it may be determined which one of the intra mode and the inter mode will be applied to each CU.
[0293] In addition, the CU can be a basic unit for predicting, transforming, quantizing, inverse-transforming, dequantizing, and encoding / decoding transform coefficients.
[0294] Referring to Figure 3 , image 300 can be sequentially partitioned into units corresponding to the largest coding unit (LCU), and the partitioning structure can be determined for each LCU. Here, the LCU can be used to have the same meaning as the coding tree unit (CTU).
[0295] Partitioning a unit can represent partitioning the block corresponding to the unit. The block partitioning information can include depth information regarding the depth of the unit. The depth information can indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned. A single unit can be hierarchically partitioned into multiple sub-units while the single unit has depth information based on a tree structure.
[0296] Each partitioned sub-unit can have depth information. The depth information can be information indicating the size of the CU. The depth information can be stored for each CU.
[0297] Each CU can have depth information. When a CU is partitioned, the depth of the CU generated from the partition can be increased by 1 from the depth of the partitioned CU.
[0298] The partitioning structure can represent the distribution of coding units (CUs) in the LCU 310 for efficiently encoding an image. This distribution can be determined based on whether a single CU will be partitioned into multiple CUs. The number of CUs generated by partitioning can be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, etc.
[0299] According to the number of CUs generated by partitioning, the horizontal size and vertical size of each CU generated by partitioning can be smaller than the horizontal size and vertical size of the CU before partitioning. For example, the horizontal size and vertical size of each CU generated by partitioning can be half of the horizontal size and vertical size of the CU before partitioning.
[0300] Each partitioned CU can be recursively partitioned into four CUs in the same way. Compared with at least one of the horizontal size and vertical size of the CU before partitioning, at least one of the horizontal size and vertical size of each partitioned CU can be reduced via recursive partitioning.
[0301] The partitioning of the CU can be recursively performed until a predefined depth or a predefined size.
[0302] For example, the depth of the CU can have a value ranging from 0 to 3. The size range of the CU can be from a size of 64×64 to a size of 8×8 depending on the depth of the CU.
[0303] For example, the depth of the LCU 310 can be 0, and the depth of the smallest coding unit (SCU) can be a predefined maximum depth. Here, as described above, the LCU can be a CU having the maximum coding unit size, and the SCU can be a CU having the smallest coding unit size.
[0304] Partitioning can start at the LCU 310, and whenever the horizontal size and / or vertical size of a CU is reduced by partitioning, the depth of the CU can be incremented by 1.
[0305] For example, for each depth, a non-partitioned CU can have a size of 2N×2N. Additionally, in the case where a CU is partitioned, a CU of size 2N×2N can be partitioned into four CUs each of size N×N. Whenever the depth is incremented by 1, the value of N can be halved.
[0306] Referring to Figure 3 , the LCU with a depth of 0 can have 64×64 pixels or a block of 64×64. 0 can be the minimum depth. The SCU with a depth of 3 can have 8×8 pixels or a block of 8×8. 3 can be the maximum depth. Here, the CU having a block of 64×64 as the LCU can be represented by a depth of 0. The CU having a block of 32×32 can be represented by a depth of 1. The CU having a block of 16×16 can be represented by a depth of 2. The CU having a block of 8×8 as the SCU can be represented by a depth of 3.
[0307] Information regarding whether a corresponding CU is partitioned can be represented by the partitioning information of the CU. The partitioning information can be 1-bit information. All CUs except the SCU can include the partitioning information. For example, the value of the partitioning information of a non-partitioned CU can be a first value. The value of the partitioning information of a partitioned CU can be a second value. When the partitioning information indicates whether a CU is partitioned, the first value can be "0" and the second value can be "1".
[0308] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by partitioning can be half of the horizontal size and vertical size of the CU before partitioning. When a CU of size 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned in a quadtree structure. In other words, it can be considered that quadtree partitioning has been applied to the CU.
[0309] For example, when a single CU is partitioned into two CUs, the horizontal or vertical size of each of the two CUs generated by the partitioning can be half of the horizontal or vertical size of the CU before partitioning. When a CU with a size of 32×32 is vertically partitioned into two CUs, the size of each of the two partitioned CUs can be 16×32. When a CU with a size of 32×32 is horizontally partitioned into two CUs, the size of each of the two partitioned CUs can be 32×16. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned in a binary tree structure. In other words, it can be considered that binary tree partitioning has been applied to the CU.
[0310] For example, when a single CU is partitioned (or divided) into three CUs, the original CU before partitioning is partitioned such that its horizontal or vertical size is divided in a ratio of 1:2:1, thus enabling the generation of three sub-CUs. For example, when a CU with a size of 16×32 is horizontally partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have sizes of 16×8, 16×16, and 16×8 respectively in the top-to-bottom direction. For example, when a CU with a size of 32×32 is vertically partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have sizes of 8×32, 16×32, and 8×32 respectively in the left-to-right direction. When a single CU is partitioned into three CUs, it can be considered that the CU is partitioned in a ternary tree form. In other words, it can be considered that ternary tree partitioning has been applied to the CU.
[0311] Both quadtree partitioning and binary tree partitioning are applied to Figure 3 the LCU 310.
[0312] In the encoding device 100, a coding tree unit (CTU) with a size of 64×64 can be partitioned into multiple smaller CUs through a recursive quadtree structure. A single CU can be partitioned into four CUs with the same size. Each CU can be recursively partitioned and can have a quadtree structure.
[0313] Through the recursive partitioning of the CU, an optimal partitioning method that incurs the minimum rate-distortion cost can be selected.
[0314] Figure 3 The coding tree unit (CTU) 320 in
[0315] As described above, in order to partition the CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning can be applied to the CTU. The partitioning can be applied based on a specific priority.
[0316] For example, quadtree partitioning can be preferentially applied to CTUs. A CU that cannot be further partitioned in the form of a quadtree can correspond to a leaf node of the quadtree. A CU corresponding to a leaf node of the quadtree can be the root node of a binary tree and / or a ternary tree. That is, a CU corresponding to a leaf node of the quadtree can be partitioned in the form of a binary tree or a ternary tree, or may not be further partitioned. In this case, it is prevented that each CU generated by applying binary tree partitioning or ternary tree partitioning to a CU corresponding to a leaf node of the quadtree is again quadtree partitioned, thereby effectively performing the operation of partitioning a block and / or signaling block partition information.
[0317] Quad-partition information can be used to signal the partitioning of the CU corresponding to each node of the quadtree. Quad-partition information having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a quadtree. Quad-partition information having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in the form of a quadtree. The quad-partition information can be a flag having a specific length (e.g., 1 bit).
[0318] There may be no priority between binary tree partitioning and ternary tree partitioning. That is, a CU corresponding to a leaf node of the quadtree can be partitioned in the form of a binary tree or a ternary tree. In addition, a CU generated by binary tree partitioning or ternary tree partitioning can be further partitioned in the form of a binary tree or a ternary tree, or may not be further partitioned.
[0319] The partitioning performed when there is no priority between binary tree partitioning and ternary tree partitioning can be referred to as "multi-type tree partitioning". That is, a CU corresponding to a leaf node of the quadtree can be the root node of a multi-type tree. At least one of information indicating whether a CU is partitioned according to a multi-type tree, partitioning direction information, and partitioning tree information can be used to signal the partitioning of the CU corresponding to each node of the multi-type tree. For the partitioning of the CU corresponding to each node of the multi-type tree, information indicating whether the partitioning according to the multi-type tree is performed, partitioning direction information, and partitioning tree information can be signaled sequentially.
[0320] For example, information indicating whether a CU is partitioned according to a multi-type tree and having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in the form of a multi-type tree. Information indicating whether a CU is partitioned according to a multi-type tree and having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in the form of a multi-type tree.
[0321] When a CU corresponding to each node of the multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU can further include partitioning direction information.
[0322] The partition direction information may indicate the partition direction of multiple types of tree partitions. The partition direction information having a first value (e.g., "1") may indicate that the corresponding CU is partitioned in the vertical direction. The partition direction information having a second value (e.g., "0") may indicate that the corresponding CU is partitioned in the horizontal direction.
[0323] When the CU corresponding to each node of the multiple-type tree is partitioned in the form of the multiple-type tree, the corresponding CU may further include partition tree information. The partition tree information may indicate the tree used for the multiple-type tree partition.
[0324] For example, the partition tree information having a first value (e.g., "1") may indicate that the corresponding CU is partitioned in the form of a binary tree. The partition tree information having a second value (e.g., "0") may indicate that the corresponding CU is partitioned in the form of a ternary tree.
[0325] Here, each of the above information indicating whether the partition of the multiple-type tree is performed, the partition tree information, and the partition direction information may be a flag having a specific length (e.g., 1 bit).
[0326] At least one of the above-mentioned quad-partition information, the information indicating whether the partition of the multiple-type tree is performed, the partition direction information, and the partition tree information may be entropy-coded and / or entropy-decoded. To perform the entropy coding / decoding of such information, the information of neighboring CUs adjacent to the target CU may be used.
[0327] For example, it can be considered that the probability that the partition form (i.e., partition / non-partition, partition tree, and / or partition direction) of the left CU and / or the upper CU is similar to that of the target CU is very high. Therefore, based on the information of the neighboring CUs, the context information for the entropy coding and / or entropy decoding of the information for the target CU may be derived. Here, the information of the neighboring CUs may include at least one of the following: 1) the quad-partition information of the neighboring CUs, 2) the information indicating whether the neighboring CUs are partitioned in the form of a multiple-type tree, 3) the partition direction information of the neighboring CUs, and 4) the partition tree information of the neighboring CUs.
[0328] In another embodiment of the binary tree partition and the ternary tree partition, the binary tree partition may be preferentially performed. That is, the binary tree partition may be first applied, and then the CU corresponding to the leaf node of the binary tree may be set as the root node of the ternary tree. In this case, the quad-tree partition or the binary-tree partition may not be performed on the CU corresponding to the node of the ternary tree.
[0329] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning can be a unit for encoding, prediction, and / or transformation. That is, the CU may not be further partitioned for prediction and / or transformation. Therefore, a partitioning structure for partitioning the CU into prediction units (PUs) / or transform units (TUs), its partitioning information, etc. may not exist in the bitstream.
[0330] However, when the size of the CU as a partitioning unit is larger than the size of the maximum transform block, the CU can be recursively partitioned until the size of the CU becomes less than or equal to the size of the maximum transform block. For example, when the size of the CU is 64×64 and the size of the maximum transform block is 32×32, the CU can be partitioned into four 32×32 blocks to perform transformation. For example, when the size of the CU is 32×64 and the size of the maximum transform block is 32×32, the CU can be partitioned into two 32×32 blocks.
[0331] In this case, information indicating whether the CU is partitioned for transformation may not be signaled separately. Without signaling, it can be determined whether the CU is partitioned via a comparison between the horizontal size (and / or vertical size) of the CU and the horizontal size (and / or vertical size) of the maximum transform block. For example, when the horizontal size of the CU is larger than the horizontal size of the maximum transform block, the CU can be bisected vertically. In addition, when the vertical size of the CU is larger than the vertical size of the maximum transform block, the CU can be bisected horizontally.
[0332] Information about the maximum size and / or minimum size of the CU and information about the maximum size and / or minimum size of the transform block can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be the sequence level, picture level, parallel block level, parallel block group level, or slice level. For example, the minimum size of the CU can be set to 4×4. For example, the maximum size of the transform block can be set to 64×64. For example, the maximum size of the transform block can be set to 4×4.
[0333] Information about the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or information about the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be the sequence level, picture level, slice level, parallel block group level, or parallel block level. Information about the minimum size of the quadtree and / or information about the maximum depth of the multi-type tree can be signaled or determined separately at each of the intra-slice level and inter-slice level.
[0334] Information about the difference between the size of a CTU and the maximum size of a transform block can be signaled or determined at a level higher than the level of the CU. For example, the higher level can be the sequence level, the picture level, the slice level, the parallel block group level, or the parallel block level. Information about the maximum size of the CU (i.e., the maximum size of the binary tree) corresponding to each node of the binary tree can be determined based on the size of the CTU and the information about the difference. The maximum size of the CU (i.e., the maximum size of the ternary tree) corresponding to each node of the ternary tree can have different values according to the type of the slice. For example, the maximum size of the ternary tree at the intra-slice level can be 32×32. For example, the maximum size of the ternary tree at the inter-slice level can be 128×128. For example, the minimum size of the CU (i.e., the minimum size of the binary tree) corresponding to each node of the binary tree and / or the minimum size of the CU (i.e., the minimum size of the ternary tree) corresponding to each node of the ternary tree can be set to the minimum size of the CU.
[0335] In another example, the maximum size of the binary tree and / or the maximum size of the ternary tree can be signaled or determined at the slice level. In addition, the minimum size of the binary tree and / or the minimum size of the ternary tree can be signaled or determined at the slice level.
[0336] Based on the various block sizes and depths described above, the quad-partition information, the information indicating whether the partitioning by the multi-type tree is performed, the partition tree information, and / or the partition direction information may or may not be present in the bitstream.
[0337] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include the quad-partition information, and the quad-partition information of the CU can be inferred as a second value.
[0338] For example, when the size (horizontal size and vertical size) of the CU corresponding to each node of the multi-type tree is greater than the maximum size (horizontal size and vertical size) of the binary tree and / or the maximum size (horizontal size and vertical size) of the ternary tree, the CU may not be partitioned in the form of the binary tree and / or the ternary tree. By this determination method, the information indicating whether the partitioning by the multi-type tree is performed may not be signaled, but can be inferred as a second value.
[0339] Optionally, when the size (horizontal size and vertical size) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal size and vertical size) of the binary tree, or when the size (horizontal size and vertical size) of the CU is equal to twice the minimum size (horizontal size and vertical size) of the ternary tree, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. By this determination, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value. The reason is that when the CU is partitioned in the form of a binary tree and / or a ternary tree, CUs smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree are generated.
[0340] Optionally, the binary tree partitioning or the ternary tree partitioning may be restricted based on the size of the virtual pipeline data unit (i.e., the size of the pipeline buffer). For example, when the CU is partitioned into sub-CUs that are not suitable for the size of the pipeline buffer by binary tree partitioning or ternary tree partitioning, the binary tree partitioning or the ternary tree partitioning may be restricted. The size of the pipeline buffer may be equal to the maximum size of the transform block (e.g., 64×64).
[0341] For example, when the size of the pipeline buffer is 64×64, the following partitionings may be restricted.
[0342] - Ternary tree partitioning for an N×M CU (where N and / or M is 128)
[0343] - Horizontal binary tree partitioning for a 128×N CU (where N <= 64)
[0344] - Vertical binary tree partitioning for an N×128 CU (where N <= 64)
[0345] Optionally, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. By this determination, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.
[0346] Optionally, the information indicating whether the partitioning according to the multi-type tree is performed may be signaled only when at least one of the vertical binary tree partitioning, the horizontal binary tree partitioning, the vertical ternary tree partitioning, and the horizontal ternary tree partitioning is possible for the CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in the form of a binary tree and / or a ternary tree. By this determination, the information indicating whether the partitioning according to the multi-type tree is performed may not be signaled, but may be inferred as a second value.
[0347] Optionally, for a CU corresponding to each node of the multi-type tree, the partitioning direction information may be signaled only when both the vertical binary tree partitioning and the horizontal binary tree partitioning are feasible or only when both the vertical ternary tree partitioning and the horizontal ternary tree partitioning are feasible. Otherwise, the partitioning direction information may not be signaled, but may be inferred as a value indicating the direction in which the CU may be partitioned.
[0348] Optionally, for a CU corresponding to each node of the multi-type tree, the partitioning tree information may be signaled only when both the vertical binary tree partitioning and the vertical ternary tree partitioning are feasible or only when both the horizontal binary tree partitioning and the horizontal ternary tree partitioning are feasible. Otherwise, the partitioning tree information may not be signaled, but may be inferred as a value indicating the tree of the partitioning applicable to the CU.
[0349] Figure 4 is a diagram showing the forms of prediction units that a coding unit can include.
[0350] Among the CUs partitioned from an LCU, the CUs that are no longer partitioned may be divided into one or more prediction units (PUs). This division is also referred to as "partitioning".
[0351] A PU may be a basic unit for prediction. A PU may be encoded and decoded in any one of the skip mode, the inter-frame mode, and the intra-frame mode. A PU may be partitioned into various shapes according to each mode. For example, the target blocks described above with reference to Figure 1 and the target blocks described above with reference to Figure 2 may both be PUs.
[0352] A CU may not be divided into PUs. When a CU is not divided into PUs, the size of the CU and the size of the PU may be equal to each other.
[0353] In the skip mode, there may be no partitioning in a CU. In the skip mode, the 2N×2N mode 410 may be supported without partitioning, where in the 2N×2N mode 410, the size of the PU and the size of the CU are the same.
[0354] In the inter-frame mode, there may be 8 types of partitioning shapes in a CU. For example, in the inter-frame mode, the 2N×2N mode 410, the 2N×N mode 415, the N×2N mode 420, the N×N mode 425, the 2N×nU mode 430, the 2N×nD mode 435, the nL×2N mode 440, and the nR×2N mode 445 may be supported.
[0355] In the intra-frame mode, the 2N×2N mode 410 and the N×N mode 425 may be supported.
[0356] In the 2N×2N mode 410, a PU with a size of 2N×2N can be encoded. A PU with a size of 2N×2N can represent a PU having the same size as the CU. For example, a PU with a size of 2N×2N can have a size of 64×64, 32×32, 16×16, or 8×8.
[0357] In the N×N mode 425, a PU with a size of N×N can be encoded.
[0358] For example, in intra prediction, when the size of the PU is 8×8, the PUs divided into four partitions can be encoded. The size of each partitioned PU can be 4×4.
[0359] When encoding a PU in the intra mode, any one of multiple intra prediction modes can be used to encode the PU. For example, the HEVC technology can provide 35 intra prediction modes, and the PU can be encoded under any one of the 35 intra prediction modes.
[0360] It can be determined based on the rate-distortion cost which one of the 2N×2N mode 410 and the N×N mode 425 will be used to encode the PU.
[0361] The encoding device 100 can perform an encoding operation on a PU with a size of 2N×2N. Here, the encoding operation can be an operation of encoding the PU under each of the multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra prediction mode for the PU with a size of 2N×2N can be derived. The optimal intra prediction mode can be the intra prediction mode that incurs the minimum rate-distortion cost when encoding the PU with a size of 2N×2N among the multiple intra prediction modes that can be used by the encoding device 100.
[0362] In addition, the encoding device 100 can sequentially perform an encoding operation on each PU obtained by performing N×N partitioning. Here, the encoding operation can be an operation of encoding the PU under each of the multiple intra prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra prediction mode for the PU with a size of N×N can be derived. The optimal intra prediction mode can be the intra prediction mode that incurs the minimum rate-distortion cost when encoding the PU with a size of N×N among the multiple intra prediction modes that can be used by the encoding device 100.
[0363] The encoding device 100 can determine which one of the PU with a size of 2N×2N and the PU with a size of N×N will be encoded based on a comparison between the rate-distortion cost of the PU with a size of 2N×2N and the rate-distortion cost of the PU with a size of N×N.
[0364] A single CU can be partitioned into one or more PUs, and a PU can be partitioned into multiple PUs.
[0365] For example, when a single PU is partitioned into four PUs, the horizontal size and vertical size of each of the four PUs generated by the partitioning can be half of the horizontal size and vertical size of the PU before partitioning. When a PU with a size of 32×32 is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, it can be considered that the PU has been partitioned in a quadtree structure.
[0366] For example, when a single PU is partitioned into two PUs, the horizontal size or vertical size of each of the two PUs generated by the partitioning can be half of the horizontal size or vertical size of the PU before partitioning. When a PU with a size of 32×32 is vertically partitioned into two PUs, the size of each of the two partitioned PUs can be 16×32. When a PU with a size of 32×32 is horizontally partitioned into two PUs, the size of each of the two partitioned PUs can be 32×16. When a single PU is partitioned into two PUs, it can be considered that the PU has been partitioned in a binary tree structure.
[0367] Figure 5 It is a diagram showing the form of transform units that can be included in a coding unit.
[0368] A transform unit (TU) can be the basic unit in a CU that is used for processes such as transformation, quantization, inverse transformation, dequantization, entropy coding, and entropy decoding.
[0369] A TU can have a square shape or a rectangular shape. The shape of the TU can be determined based on the size and / or shape of the CU.
[0370] In a CU partitioned from an LCU, a CU that is no longer partitioned into CUs can be partitioned into one or more TUs. Here, the partitioning structure of the TU can be a quadtree structure. For example, as Figure 5 shown, a single CU 510 can be partitioned one or more times according to the quadtree structure. Through this partitioning, a single CU 510 can be composed of TUs with various sizes.
[0371] It can be considered that when a single CU is divided two or more times, the CU is recursively divided. Through the division, a single CU can be composed of transform units (TUs) with various sizes.
[0372] Optionally, a single CU can be divided into one or more TUs based on the number of vertical lines and / or horizontal lines for partitioning the CU.
[0373] A CU can be divided into a symmetric TU or an asymmetric TU. To divide into an asymmetric TU, information about the size and / or shape of each TU can be signaled from the encoding device 100 to the decoding device 200. Optionally, the size and / or shape of each TU can be derived from the information about the size and / or shape of the CU.
[0374] A CU may not be divided into TUs. When a CU is not divided into TUs, the size of the CU and the size of the TU may be equal to each other.
[0375] A single CU can be partitioned into one or more TUs, and a TU can be partitioned into multiple TUs.
[0376] For example, when a single TU is partitioned into four TUs, the horizontal size and vertical size of each of the four TUs generated by the partitioning can be half of the horizontal size and vertical size of the TU before partitioning. When a TU with a size of 32×32 is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, it can be considered that the TU has been partitioned in a quadtree structure.
[0377] For example, when a single TU is partitioned into two TUs, the horizontal size or vertical size of each of the two TUs generated by the partitioning can be half of the horizontal size or vertical size of the TU before partitioning. When a TU with a size of 32×32 is vertically partitioned into two TUs, the size of each of the two partitioned TUs can be 16×32. When a TU with a size of 32×32 is horizontally partitioned into two TUs, the size of each of the two partitioned TUs can be 32×16. When a single TU is partitioned into two TUs, it can be considered that the TU has been partitioned in a binary tree structure.
[0378] The CU can be divided in a way different from Figure 5 the way shown in
[0379] For example, a single CU can be divided into three CUs. The horizontal size or vertical size of the three CUs generated by the division can be 1 / 4, 1 / 2, and 1 / 4 of the horizontal size or vertical size of the original CU before division, respectively.
[0380] For example, when a CU with a size of 32×32 is vertically divided into three CUs, the sizes of the three CUs generated by the division can be 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is divided into three CUs, it can be considered that the CU is divided in the form of a ternary tree.
[0381] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) can be applied to the partitioning of a CU, and multiple partitioning schemes can be combined and used together for the partitioning of the CU. Here, the case of combining and using multiple partitioning schemes together can be referred to as "compound tree form partitioning".
[0382] Figure 6 Shows the partitioning of a block according to an example.
[0383] In video encoding and / or decoding processing, as Figure 6 shown, the target block can be partitioned. For example, the target block can be a CU.
[0384] For the partitioning of the target block, an indicator indicating the partitioning information can be signaled from the encoding device 100 to the decoding device 200. The partitioning information can be information indicating how the target block is partitioned.
[0385] The partitioning information can be one or more of a partitioning flag (hereinafter referred to as "split_flag"), a quadtree-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").
[0386] "split_flag" can be a flag indicating whether the block is partitioned. For example, a split_flag value of 1 can indicate that the corresponding block is partitioned. A split_flag value of 0 can indicate that the corresponding block is not partitioned.
[0387] "QB_flag" can be a flag indicating which of the quadtree form and the binary tree form corresponds to the shape in which the block is partitioned. For example, a QB_flag value of 0 can indicate that the block is partitioned in the quadtree form. A QB_flag value of 1 can indicate that the block is partitioned in the binary tree form. Alternatively, a QB_flag value of 0 can indicate that the block is partitioned in the binary tree form. A QB_flag value of 1 can indicate that the block is partitioned in the quadtree form.
[0388] "quadtree_flag" can be a flag indicating whether the block is partitioned in the quadtree form. For example, a quadtree_flag value of 1 can indicate that the block is partitioned in the quadtree form. A quadtree_flag value of 0 can indicate that the block is not partitioned in the quadtree form.
[0389] "binarytree_flag" can be a flag indicating whether a block is divided in a binary tree form. For example, a binarytree_flag value of 1 can indicate that the block is divided in a binary tree form. A binarytree_flag value of 0 can indicate that the block is not divided in a binary tree form.
[0390] "Btype_flag" can be a flag indicating which of the vertical division and the horizontal division corresponds to the division direction when the block is divided in a binary tree form. For example, a Btype_flag value of 0 can indicate that the block is divided in the horizontal direction. A Btype_flag value of 1 can indicate that the block is divided in the vertical direction. Optionally, a Btype_flag value of 0 can indicate that the block is divided in the vertical direction. A Btype_flag value of 1 can indicate that the block is divided in the horizontal direction.
[0391] For example, the division information of the block in Figure 6 can be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag, as shown in Table 1 below.
[0392] Table 1
[0393]
[0394] For example, the division information of the block in Figure 6 can be derived by signaling at least one of split_flag, QB_flag, and Btype_flag, as shown in Table 2 below.
[0395] Table 2
[0396]
[0397] The division method may be limited to quadtree or binary tree according to the size and / or shape of the block. When this limitation is applied, split_flag can be a flag indicating whether the block is divided in a quadtree form or a flag indicating whether the block is divided in a binary tree form. The size and shape of the block can be derived based on the depth information of the block, and the depth information can be signaled from the encoding device 100 to the decoding device 200.
[0398] When the size of the block falls within a specific range, division is only possible in the quadtree form. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can be divided only in the quadtree form.
[0399] Information indicating the maximum block size and the minimum block size that can be partitioned only in the form of a quadtree can be signaled from the encoding device 100 to the decoding device 200 via a bitstream. In addition, this information can be signaled for at least one of units such as video, sequence, picture, parameter, parallel block group, and slice (or segment).
[0400] Optionally, the maximum block size and / or the minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the size of a block is greater than 64×64 and less than 256×256, partitioning only in the form of a quadtree is possible. In this case, split_flag can be a flag indicating whether to perform partitioning in the form of a quadtree.
[0401] When the size of a block is greater than the maximum size of a transform block, partitioning only in the form of a quadtree is possible. Here, the sub-blocks generated by partitioning can be at least one of a CU and a TU.
[0402] In this case, split_flag can be a flag indicating whether the CU is partitioned in the form of a quadtree.
[0403] When the size of a block falls within a specific range, partitioning only in the form of a binary tree or a ternary tree is possible. For example, the specific range can be defined by at least one of the maximum block size and the minimum block size that can be partitioned only in the form of a binary tree or a ternary tree.
[0404] Information indicating the maximum block size and / or the minimum block size that can be partitioned only in the form of a binary tree or in the form of a ternary tree can be signaled from the encoding device 100 to the decoding device 200 via a bitstream. In addition, this information can be signaled for at least one of units such as sequence, picture, and slice (or segment).
[0405] Optionally, the maximum block size and / or the minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the size of a block is greater than 8×8 and less than 16×16, partitioning only in the form of a binary tree is possible. In this case, split_flag can be a flag indicating whether to perform partitioning in the form of a binary tree or a ternary tree.
[0406] The above description of partitioning in the form of a quadtree can be equally applied to the form of a binary tree and / or the form of a ternary tree.
[0407] The partitioning of a block may be restricted by a previous partition. For example, when a block is partitioned in a specific binary tree form and multiple sub-blocks are generated from the partition, each sub-block may be further partitioned only in a specific tree form. Here, the specific tree form may be at least one of a binary tree form, a ternary tree form, and a quaternary tree form.
[0408] When the horizontal size or the vertical size of a partitioned block is a size that cannot be further divided, the above indicator may not be signaled.
[0409] Figure 7 is a diagram for explaining an embodiment of intra prediction processing.
[0410] From Figure 7 The arrow extending radially from the center of the diagram in indicates the prediction direction of the intra prediction mode. In addition, the numbers appearing near the arrow indicate examples of the mode values assigned to the intra prediction mode or the prediction direction of the intra prediction mode.
[0411] In Figure 7 , the number 0 may represent the planar mode as a non-directional intra prediction mode. The number 1 may represent the DC mode as a non-directional intra prediction mode.
[0412] Intra coding and / or decoding may be performed using reference samples of neighboring blocks of a target block. The neighboring blocks may be reconstructed neighboring blocks. The reference samples may represent neighboring samples.
[0413] For example, intra coding and / or decoding may be performed using the values of the reference samples included in the reconstructed neighboring blocks or the coding parameters of the reconstructed neighboring blocks.
[0414] The encoding device 100 and / or the decoding device 200 may generate a prediction block by performing intra prediction on a target block based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may generate a prediction block for the target block by performing intra prediction based on information about samples in the target image. When intra prediction is performed, the encoding device 100 and / or the decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.
[0415] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.
[0416] The unit of the prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The size N×N may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.
[0417] Optionally, the prediction block may be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc., or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.
[0418] Intra prediction may be performed considering the intra prediction mode for the target block. The number of intra prediction modes that the target block may have may be a predefined fixed value and may be a value determined differently according to the attributes of the prediction block. For example, the attributes of the prediction block may include the size of the prediction block, the type of the prediction block, etc. In addition, the attributes of the prediction block may indicate the coding parameters for the prediction block.
[0419] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to N. Optionally, the number of intra prediction modes may be, for example, 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95.
[0420] The intra prediction mode may be a non-directional mode or a directional mode.
[0421] For example, the intra prediction mode may include two non-directional modes corresponding to the numbers 0 to 66 shown in Figure 7 and 65 directional modes.
[0422] For example, in the case of using a specific intra prediction method, the intra prediction mode may include two non-directional modes corresponding to the numbers -14 to 80 shown in Figure 7 and 93 directional modes.
[0423] The two non-directional modes may include the DC mode and the planar mode.
[0424] The directional mode may be a prediction mode with a specific direction or a specific angle. The directional mode may also be referred to as an "angle mode".
[0425] The intra prediction mode may be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "(mode) number of the intra prediction mode", "(mode) value of the intra prediction mode", "(mode) angle of the intra prediction mode", and "(mode) direction of the intra prediction mode" may be used with the same meaning and may be used interchangeably with each other.
[0426] The number of intra prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.
[0427] The number of intra prediction modes can be fixed to M regardless of the size of the block and / or color component. For example, the number of intra prediction modes can be fixed to either 35 or 67 regardless of the size of the block.
[0428] Optionally, the number of intra prediction modes can vary according to the shape, size, and / or type of color component of the block.
[0429] For example, in Figure 7 the directional prediction mode shown by the dashed line can be applied only to the prediction for non-square blocks.
[0430] For example, the larger the size of the block, the more intra prediction modes there are. Optionally, the larger the size of the block, the fewer intra prediction modes there are. When the size of the block is 4×4 or 8×8, the number of intra prediction modes can be 67. When the size of the block is 16×16, the number of intra prediction modes can be 35. When the size of the block is 32×32, the number of intra prediction modes can be 19. When the size of the block is 64×64, the number of intra prediction modes can be 7.
[0431] For example, the number of intra prediction modes can vary according to whether the color component is a luminance signal or a chrominance signal. Optionally, the number of intra prediction modes corresponding to the luminance component block can be greater than the number of intra prediction modes corresponding to the chrominance component block.
[0432] For example, in the vertical mode with a mode value of 50, prediction can be performed in the vertical direction based on the pixel values of the reference samples. For example, in the horizontal mode with a mode value of 18, prediction can be performed in the horizontal direction based on the pixel values of the reference samples.
[0433] Even in a direction mode other than the above modes, the encoding device 100 and the decoding device 200 can still perform intra prediction on the target unit using the reference samples according to the angle corresponding to the direction mode.
[0434] The intra prediction mode located to the right of the vertical mode can be referred to as the "vertical - right mode". The intra prediction mode located below the horizontal mode can be referred to as the "horizontal - below mode". For example, in Figure 7 the intra prediction mode with a mode value being one of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 can be the vertical - right mode. The intra prediction mode with a mode value being one of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be the horizontal - below mode.
[0435] The non-directional mode may include a DC mode and a planar mode. For example, the value of the DC mode may be 1. The value of the planar mode may be 0.
[0436] The directional mode may include an angular mode. Among the multiple intra-frame prediction modes, the remaining modes other than the DC mode and the planar mode may be the directional mode.
[0437] When the intra-frame prediction mode is the DC mode, a prediction block may be generated based on the average value of the pixel values of multiple reference pixels. For example, the pixel values of the prediction block may be determined based on the average value of the pixel values of multiple reference pixels.
[0438] The number of the above-described intra-frame prediction modes and the mode values of each intra-frame prediction mode are merely exemplary. The number of the above-described intra-frame prediction modes and the mode values of each intra-frame prediction mode may be defined differently according to embodiments, implementations, and / or requirements.
[0439] To perform intra-frame prediction on a target block, a step of checking whether the samples included in the reconstructed neighboring blocks can be used as reference samples for the target block may be executed. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, the values generated by interpolation and / or copying of at least one of the sample values included in the reconstructed neighboring blocks may replace the sample values of the samples that cannot be used as reference samples. When the values generated by copying and / or interpolation replace the sample values of the existing samples, the samples may be used as reference samples for the target block.
[0440] When using intra-frame prediction, a filter may be applied to at least one of the reference samples and the prediction samples based on at least one of the size of the target block and the intra-frame prediction mode.
[0441] The type of the filter to be applied to at least one of the reference samples and the prediction samples may vary according to at least one of the intra-frame prediction mode of the target block, the size of the target block, and the shape of the target block. The type of the filter may be classified according to one or more of the length of the filter taps, the values of the filter coefficients, and the filter strength. The length of the filter taps may represent the number of the filter taps. In addition, the number of the filter taps may represent the length of the filter.
[0442] When the intra-frame prediction mode is the planar mode, when generating the prediction block of the target block, the sample values of the prediction target block may be generated using the weighted sum of the upper reference sample of the target block, the left reference sample of the target block, the upper-right reference sample of the target block, and the lower-left reference sample of the target block according to the position of the prediction target sample in the prediction block.
[0443] When the intra prediction mode is the DC mode, the average value of the reference samples above the target block and the reference samples to the left of the target block can be used when generating the prediction block of the target block. In addition, filtering using the values of the reference samples can be performed on specific rows or specific columns in the target block. The specific row(s) can be one or more upper rows adjacent to the reference sample(s). The specific column(s) can be one or more left columns adjacent to the reference sample(s).
[0444] When the intra prediction mode is the direction mode, the upper reference sample, the left reference sample, the upper-right reference sample, and / or the lower-left reference sample of the target block can be used to generate the prediction block.
[0445] To generate the above-mentioned prediction samples, interpolation based on real numbers can be performed.
[0446] The intra prediction mode of the target block can be predicted from the intra prediction modes of neighboring blocks adjacent to the target block, and the information used for the prediction can be entropy-coded / entropy-decoded.
[0447] For example, when the intra prediction modes of the target block and the neighboring block are the same, a predefined flag can be used to signal that the intra prediction modes of the target block and the neighboring block are the same.
[0448] For example, an indicator indicating the intra prediction mode among the intra prediction modes of multiple neighboring blocks that is the same as the intra prediction mode of the target block can be signaled.
[0449] When the intra prediction modes of the target block and the neighboring block are different, entropy coding and / or entropy decoding can be used to encode and / or decode the information about the intra prediction mode of the target block.
[0450] Figure 8 is a diagram showing the reference samples used in the intra prediction process.
[0451] The reconstructed reference samples for intra prediction of the target block can include the lower-left reference sample, the left reference sample, the upper-left reference sample, the upper reference sample, and the upper-right reference sample.
[0452] For example, the left reference sample can represent the reconstructed reference pixel adjacent to the left side of the target block. The upper reference sample can represent the reconstructed reference pixel adjacent to the top of the target block. The upper-left reference sample can represent the reconstructed reference pixel located at the upper-left corner of the target block. The lower-left reference sample can represent the reference sample located below the left-side sample line among the samples on the same line as the left-side sample line composed of the left reference samples. The upper-right reference sample can represent the reference sample located to the right of the upper-side sample line among the samples on the same line as the upper-side sample line composed of the upper reference samples.
[0453] When the size of the target block is N×N, the number of the lower-left reference sample points, left reference sample points, upper reference sample points, and upper-right reference sample points can all be N.
[0454] By performing intra prediction on the target block, a prediction block can be generated. The process of generating the prediction block may include determining the values of the pixels in the prediction block. The sizes of the target block and the prediction block can be the same.
[0455] The reference sample points used for intra prediction of the target block can be changed according to the intra prediction mode of the target block. The direction of the intra prediction mode can represent the dependency relationship between the reference sample points and the pixels of the prediction block. For example, the values of the specified reference sample points can be used as the values of one or more specified pixels in the prediction block. In this case, the specified reference sample points and the one or more specified pixels in the prediction block can be the sample points and pixels located on a straight line along the direction of the intra prediction mode. In other words, the values of the specified reference sample points can be copied as the values of the pixels located in the direction opposite to the direction of the intra prediction mode. Alternatively, the values of the pixels in the prediction block can be the values of the reference sample points located in the direction of the intra prediction mode with respect to the position of the pixel.
[0456] In an example, when the intra prediction mode of the target block is the vertical mode, the upper reference sample points can be used for intra prediction. When the intra prediction mode is the vertical mode, the values of the pixels in the prediction block can be the values of the reference sample points vertically located above the position of the pixel. Therefore, the upper reference sample points adjacent to the top of the target block can be used for intra prediction. In addition, the values of the pixels in a row of the prediction block can be the same as the values of the pixels of the upper reference sample points.
[0457] In an example, when the intra prediction mode of the target block is the horizontal mode, the left reference sample points can be used for intra prediction. When the intra prediction mode is the horizontal mode, the values of the pixels in the prediction block can be the values of the reference sample points horizontally located to the left of the position of the pixel. Therefore, the left reference sample points adjacent to the left side of the target block can be used for intra prediction. In addition, the values of the pixels in a column of the prediction block can be the same as the values of the pixels of the left reference sample points.
[0458] In an example, when the mode value of the intra prediction mode of the current block is 34, at least some of the left reference sample points, the upper-left reference sample points, and at least some of the upper reference sample points can be used for intra prediction. When the mode value of the intra prediction mode is 34, the values of the pixels in the prediction block can be the values of the reference sample points diagonally located at the upper left corner of the pixel.
[0459] In addition, in the case of the intra prediction mode with the mode value ranging from 52 to 66, at least a part of the upper-right reference sample points can be used for intra prediction.
[0460] In addition, in the case of an intra prediction mode with a mode value ranging from 2 to 17, at least a part of the lower left reference samples can be used for intra prediction.
[0461] In addition, in the case of an intra prediction mode with a mode value ranging from 19 to 49, the upper left reference sample can be used for intra prediction.
[0462] The number of reference samples for determining the pixel value of a pixel in a prediction block can be 1 or 2 or more.
[0463] As described above, the pixel value of a pixel in a prediction block can be determined according to the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of one reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.
[0464] When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are not integer positions, an interpolated reference sample based on the two reference samples closest to the position of the reference sample can be generated. The value of the interpolated reference sample can be used to determine the pixel value of the pixel in the prediction block. In other words, when the position of the pixel in the prediction block and the position of the reference sample indicated by the direction of the intra prediction mode indicate a position between two reference samples, an interpolation based on the values of these two samples can be generated.
[0465] The prediction block generated through prediction can be different from the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.
[0466] Hereinafter, the terms "difference", "error" and "residual" can be used with the same meaning and can be used interchangeably with each other.
[0467] For example, in the case of directional intra prediction, the longer the distance between the pixels of the prediction block and the reference sample, the greater the possible prediction error. Such a prediction error can lead to discontinuity between the generated prediction block and neighboring blocks.
[0468] To reduce the prediction error, a filtering operation for the prediction block can be used. The filtering operation can be configured to adaptively apply a filter to the area in the prediction block that is considered to have a larger prediction error. For example, the area considered to have a larger prediction error can be the boundary of the prediction block. In addition, the area in the prediction block considered to have a larger prediction error can be different according to the intra prediction mode, and the characteristics of the filter can also be different according to the intra prediction mode.
[0469] AsFigure 8 As shown, for intra prediction of a target block, at least one of reference lines 0 to 3 can be used.
[0470] In Figure 8 each reference line can indicate a reference sample line including one or more reference samples. When the number of the reference line is small, it can indicate a reference sample line closer to the target block.
[0471] The samples in segment A and segment F can be obtained by padding instead of from reconstructed neighboring blocks, where the padding uses the samples in segment B and segment E that are closest to the target block.
[0472] Index information indicating the reference sample line to be used for intra prediction of the target block can be signaled. The index information can indicate the reference sample line among a plurality of reference sample lines to be used for intra prediction of the target block. For example, the index information can have a value corresponding to any one of 0 to 3.
[0473] When the upper boundary of the target block is the boundary of the CTU, only reference line 0 can be available. Therefore, in this case, the index information may not be signaled. When additional reference sample lines other than reference line 0 are used, filtering of the prediction block described later may not be performed.
[0474] In the case of inter-color intra prediction, a prediction block of the target block of the second color component can be generated based on the corresponding reconstructed block of the first color component.
[0475] For example, the first color component can be a luminance component, and the second color component can be a chrominance component.
[0476] To perform inter-color intra prediction, parameters of a linear model between the first color component and the second color component can be derived based on a template.
[0477] The template can include reference samples (upper reference samples) above the target block and / or reference samples (left reference samples) to the left of the target block, and can include upper reference samples and / or left reference samples of the reconstructed blocks of the first color component corresponding to the reference samples.
[0478] For example, the following values can be used to derive the parameters of the linear model: 1) the value of the sample of the first color component having the maximum value among the samples in the template, 2) the value of the sample of the second color component corresponding to the sample of the first color component, 3) the value of the sample of the first color component having the minimum value among the samples in the template, and 4) the value of the sample of the second color component corresponding to the sample of the first color component.
[0479] When exporting the parameters of a linear model, a prediction block of a target block can be generated by applying a corresponding reconstruction block to the linear model.
[0480] According to the image format, subsampling can be performed on the samples adjacent to the reconstruction block of the first color component and the corresponding reconstruction block of the first color component. For example, when one sample of the second color component corresponds to four samples of the first color component, one corresponding sample can be calculated by performing subsampling on the four samples of the first color component. When performing subsampling, the derivation of the parameters of the linear model and the intra-color prediction can be performed based on the corresponding samples to be subsampled.
[0481] Information about whether to perform intra-color prediction and / or the range of the template can be signaled in the intra-prediction mode.
[0482] The target block can be partitioned into two or four sub-blocks in the horizontal direction and / or the vertical direction.
[0483] The sub-blocks generated by partitioning can be reconstructed sequentially. That is, when performing intra-prediction on each sub-block, a sub-prediction block of the sub-block can be generated. In addition, when performing inverse quantization and / or inverse transformation on each sub-block, a sub-residual block for the corresponding sub-block can be generated. The reconstructed sub-block can be generated by adding the sub-prediction block and the sub-residual block. The reconstructed sub-block can be used as a reference sample for intra-prediction of the sub-block with the next priority.
[0484] A sub-block can be a block including a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block can be partitioned into two sub-blocks. In addition, when the target block is a 4×4 block, the target block cannot be partitioned into sub-blocks. When the target block has another size, the target block can be partitioned into four sub-blocks.
[0485] Information about whether to perform intra-prediction based on these sub-blocks and / or information about the partitioning direction (horizontal direction or vertical direction) can be signaled.
[0486] This intra-prediction based on sub-blocks can be restricted so that it is only performed when the reference sample line 0 is used. When performing intra-prediction based on sub-blocks, the filtering of the prediction block described below may not be performed.
[0487] The final prediction block can be generated by performing filtering on the prediction block generated via intra-prediction.
[0488] Filtering can be performed by applying specific weights to the filtering target samples, the left reference sample, the upper reference sample, and / or the upper left reference sample that are the targets to be filtered.
[0489] The weights for filtering and / or the reference samples (e.g., the range of reference samples, the positions of reference samples, etc.) can be determined based on at least one of the block size, the intra prediction mode, and the positions of the filtered target samples in the prediction block.
[0490] For example, filtering can be performed only in specific intra prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).
[0491] The adjacent diagonal mode can be a mode having a number obtained by adding k to the number of the diagonal mode, and can be a mode having a number obtained by subtracting k from the number of the diagonal mode. In other words, the number of the adjacent diagonal mode can be the sum of the number of the diagonal mode and k, or can be the difference between the number of the diagonal mode and k. For example, k can be a positive integer of 8 or less.
[0492] The intra prediction mode of a target block can be derived using the intra prediction modes of neighboring blocks that appear around the target block, and such derived intra prediction mode can be entropy encoded and / or entropy decoded.
[0493] For example, when the intra prediction mode of a target block is the same as the intra prediction mode of a neighboring block, specific flag information can be used to signal information indicating that the intra prediction mode of the target block is the same as the intra prediction mode of the neighboring block.
[0494] In addition, for example, indicator information of neighboring blocks in which the intra prediction mode among the intra prediction modes of multiple neighboring blocks is the same as the intra prediction mode of the target block can be signaled.
[0495] For example, when the intra prediction mode of a target block is different from the intra prediction mode of a neighboring block, entropy encoding and / or entropy decoding of information regarding the intra prediction mode of the target block can be performed by performing entropy encoding and / or entropy decoding based on the intra prediction mode of the neighboring block.
[0496] Figure 9 is a diagram for explaining an embodiment of an inter prediction process.
[0497] Figure 9 The rectangle shown in can represent an image (or picture). In addition, in Figure 9 the arrow can represent the prediction direction. The arrow pointing from the first picture to the second picture indicates that the second picture references the first picture. That is, each image can be encoded and / or decoded according to the prediction direction.
[0498] An image can be classified as an intra picture (I picture), a single predicted picture or a predicted coded picture (P picture), and a bi-predicted picture or a bi-predicted coded picture (B picture) according to the coding type. Each picture can be coded and / or decoded according to the coding type of each picture.
[0499] When the target image to be coded is an I picture, the target image can be coded using the data included in the image itself without performing inter prediction with reference to other images. For example, an I picture can be coded only via intra prediction.
[0500] When the target image is a P picture, the target image can be coded via inter prediction using a reference picture existing in one direction. Here, the one direction can be a forward direction or a backward direction.
[0501] When the target image is a B picture, the image can be coded via inter prediction using reference pictures existing in two directions, or the image can be coded via inter prediction using a reference picture existing in one of the forward direction and the backward direction. Here, the two directions can be the forward direction and the backward direction.
[0502] P pictures and B pictures coded and / or decoded using reference pictures can be regarded as images using inter prediction.
[0503] Hereinafter, inter prediction in an inter mode according to an embodiment will be described in detail.
[0504] Inter prediction or motion compensation can be performed using a reference image and motion information.
[0505] In the inter mode, the coding device 100 can perform inter prediction and / or motion compensation on a target block. The decoding device 200 can perform inter prediction and / or motion compensation corresponding to the inter prediction and / or motion compensation performed by the coding device 100 on the target block.
[0506] Motion information of a target block can be separately derived by the coding device 100 and the decoding device 200 during inter prediction. Motion information can be derived using motion information of a reconstructed neighboring block, motion information of a col block, and / or motion information of a block adjacent to the col block.
[0507] For example, the coding device 100 or the decoding device 200 can perform prediction and / or motion compensation by using motion information of a spatial candidate and / or a temporal candidate as the motion information of the target block. The target block can represent a PU and / or a PU partition.
[0508] The spatial candidate can be a reconstructed block spatially adjacent to the target block.
[0509] A temporal candidate may be a reconstructed block corresponding to a target block in a previously reconstructed co-located picture (col picture).
[0510] In inter prediction, the encoding device 100 and the decoding device 200 may improve the encoding efficiency and the decoding efficiency by using motion information of spatial candidates and / or temporal candidates. The motion information of spatial candidates may be referred to as "spatial motion information". The motion information of temporal candidates may be referred to as "temporal motion information".
[0511] Hereinafter, the motion information of spatial candidates may be the motion information of a PU including a spatial candidate. The motion information of temporal candidates may be the motion information of a PU including a temporal candidate. The motion information of a candidate block may be the motion information of a PU including a candidate block.
[0512] Inter prediction may be performed using a reference picture.
[0513] The reference picture may be at least one of a picture before the target picture and a picture after the target picture. The reference picture may be an image for prediction of a target block.
[0514] In inter prediction, a reference picture index (or refIdx) for indicating a reference picture, a motion vector to be described later, etc. may be used to specify a region in the reference picture. Here, the region specified in the reference picture may indicate a reference block.
[0515] Inter prediction may select a reference picture, and may also select a reference block corresponding to a target block from the reference picture. In addition, inter prediction may use the selected reference block to generate a prediction block for the target block.
[0516] Motion information may be derived by each of the encoding device 100 and the decoding device 200 during inter prediction.
[0517] A spatial candidate may be a block that 1) exists in the target picture, 2) has been previously reconstructed via encoding and / or decoding, and 3) is adjacent to or located at a corner of the target block. Here, a "block located at a corner of the target block" may be a block that is vertically adjacent to a neighboring block horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block vertically adjacent to the target block. In addition, a "block located at a corner of the target block" may have the same meaning as a "block adjacent to a corner of the target block". The meaning of a "block located at a corner of the target block" may be included in the meaning of a "block adjacent to the target block".
[0518] For example, a spatial candidate may be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower left corner of the target block, a reconstructed block located at the upper right corner of the target block, or a reconstructed block located at the upper left corner of the target block.
[0519] Each of the encoding device 100 and the decoding device 200 can identify a block at a position in the col picture that spatially corresponds to the target block. The position of the target block in the target picture and the position of the identified block in the col picture can correspond to each other.
[0520] Each of the encoding device 100 and the decoding device 200 can determine a col block at a predefined relative position for the identified block as a temporal candidate. The predefined relative position can be a position inside and / or outside the identified block.
[0521] For example, the col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at the coordinates (xP + nPSW, yP + nPSH). The second col block can be the block located at the coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first col block is not available, the second col block can be selectively used.
[0522] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. In addition, the motion vector of the motion information of the temporal candidate stored in the list can be a scaled motion vector.
[0523] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first temporal distance to the second temporal distance. The first temporal distance can be the distance between the reference picture and the target picture of the target block. The second temporal distance can be the distance between the reference picture and the col picture of the col block.
[0524] The scheme for deriving motion information can be changed according to the inter-frame prediction mode of the target block. For example, as the inter-frame prediction mode applied to inter-frame prediction, there can be an Advanced Motion Vector Predictor (AMVP) mode, a merge mode, a skip mode, a merge mode with motion vector difference, a sub-block merge mode, a triangular partitioning mode, an inter-frame / intra-frame combined prediction mode, an affine inter-frame mode, a current picture reference mode, etc. The merge mode can also be referred to as the "motion merge mode". Each mode will be described in detail below.
[0525] 1) AMVP Mode
[0526] When using the AMVP mode, the encoding device 100 may search for similar blocks in the neighboring area of the target block. The encoding device 100 may obtain a predicted block by performing prediction on the target block by using the motion information of the found similar blocks. The encoding device 100 may encode the residual block that is the difference between the target block and the predicted block.
[0527] 1-1) Create a list of predicted motion vector candidates
[0528] When the AMVP mode is used as a prediction mode, each of the encoding device 100 and the decoding device 200 may use a motion vector of a spatial candidate, a motion vector of a temporal candidate, and a zero vector to create a list of predicted motion vector candidates. The list of predicted motion vector candidates may include one or more predicted motion vector candidates. At least one of the motion vector of the spatial candidate, the motion vector of the temporal candidate, and the zero vector may be determined and used as a predicted motion vector candidate.
[0529] Hereinafter, the terms "predicted motion vector (candidate)" and "motion vector (candidate)" may be used to have the same meaning and may be used interchangeably with each other.
[0530] Hereinafter, the terms "predicted motion vector candidate" and "AMVP candidate" may be used to have the same meaning and may be used interchangeably with each other.
[0531] Hereinafter, the terms "list of predicted motion vector candidates" and "list of AMVP candidates" may be used to have the same meaning and may be used interchangeably with each other.
[0532] Spatial candidates may include reconstructed spatially neighboring blocks. In other words, the motion vector of the reconstructed neighboring block may be referred to as a "spatial predicted motion vector candidate".
[0533] Temporal candidates may include col blocks and blocks adjacent to the col blocks. In other words, the motion vector of the col block or the motion vector of the block adjacent to the col block may be referred to as a "temporal predicted motion vector candidate".
[0534] The zero vector may be a (0,0) motion vector.
[0535] A predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. In addition, in the encoding device 100, each predicted motion vector candidate may be an initial search position for the motion vector.
[0536] 1-2) Search for motion vectors using the list of predicted motion vector candidates
[0537] The encoding device 100 may determine a motion vector to be used for encoding a target block within a search range by using a list of predicted motion vector candidates. Further, the encoding device 100 may determine, among the predicted motion vector candidates present in the predicted motion vector candidate list, a predicted motion vector candidate to be used as the predicted motion vector of the target block.
[0538] The motion vector to be used for encoding the target block may be a motion vector that can be encoded at the minimum cost.
[0539] Further, the encoding device 100 may determine whether to use the AMVP mode to encode the target block.
[0540] 1-3) Transmission of inter-frame prediction information
[0541] The encoding device 100 may generate a bitstream including inter prediction information required for inter prediction. The decoding device 200 may perform inter prediction on the target block by using the inter prediction information of the bitstream.
[0542] The inter prediction information may include 1) mode information indicating whether the AMVP mode is used, 2) a predicted motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.
[0543] Hereinafter, the terms “predicted motion vector index” and “AMVP index” may be used with the same meaning and may be used interchangeably with each other.
[0544] Further, the inter prediction information may include a residual signal.
[0545] When the mode information indicates that the AMVP mode is used, the decoding device 200 may obtain the predicted motion vector index, MVD, reference direction, and reference picture index from the bitstream by entropy decoding.
[0546] The predicted motion vector index may indicate a predicted motion vector candidate to be used for predicting the target block among the predicted motion vector candidates included in the predicted motion vector candidate list.
[0547] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information
[0548] The decoding device 200 may use the predicted motion vector candidate list to derive predicted motion vector candidates and may determine the motion information of the target block based on the derived predicted motion vector candidates.
[0549] The decoding device 200 may determine a motion vector candidate for a target block among the predicted motion vector candidates included in the predicted motion vector candidate list using the predicted motion vector index. The decoding device 200 may select, as the predicted motion vector of the target block, the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list.
[0550] The encoding device 100 may generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and may generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index may be signaled from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 may extract the entropy-coded predicted motion vector index from the bitstream, and may obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.
[0551] The motion vector actually to be used for inter prediction of the target block may not match the predicted motion vector. To indicate the difference between the motion vector actually to be used for inter prediction of the target block and the predicted motion vector, an MVD may be used. The encoding device 100 may derive a predicted motion vector similar to the motion vector actually to be used for inter prediction of the target block so as to use as small an MVD as possible.
[0552] The motion vector difference (MVD) may be the difference between the motion vector of the target block and the predicted motion vector. The encoding device 100 may calculate the MVD, and may generate an entropy-coded MVD by applying entropy coding to the MVD. The encoding device 100 may generate a bitstream including the entropy-coded MVD.
[0553] The MVD may be sent from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 may extract the entropy-coded MVD from the bitstream, and may obtain the MVD by applying entropy decoding to the entropy-coded MVD.
[0554] The decoding device 200 may derive the motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 may be the sum of the MVD and the motion vector candidate.
[0555] In addition, the encoding device 100 may generate an entropy-coded MVD resolution information by applying entropy coding to the calculated MVD resolution information, and may generate a bitstream including the entropy-coded MVD resolution information. The decoding device 200 may extract the entropy-coded MVD resolution information from the bitstream, and may obtain the MVD resolution information by applying entropy decoding to the entropy-coded MVD resolution information. The decoding device 200 may use the MVD resolution information to adjust the resolution of the MVD.
[0556] In addition, the encoding device 100 may calculate the MVD based on an affine model. The decoding device 200 may derive the affine control motion vector of the target block by the sum of the MVD and the affine control motion vector candidate, and may derive the motion vector of the sub-block using the affine control motion vector.
[0557] The reference direction may indicate a list of reference pictures to be used for predicting the target block. For example, the reference direction may indicate one of the reference picture list L0 and the reference picture list L1.
[0558] The reference direction only indicates the list of reference pictures to be used for predicting the target block, and does not necessarily mean that the direction of the reference picture is limited to the forward direction or the backward direction. In other words, each of the reference picture list L0 and the reference picture list L1 may include pictures in the forward direction and / or the backward direction.
[0559] The reference direction being unidirectional may mean using a single reference picture list. The reference direction being bidirectional may mean using two reference picture lists. In other words, the reference direction may indicate one of the following cases: the case of using only the reference picture list L0, the case of using only the reference picture list L1, and the case of using two reference picture lists.
[0560] The reference picture index may indicate the reference picture for predicting the target block among the reference pictures existing in the reference picture list. The encoding device 100 may generate an entropy-coded reference picture index by applying entropy coding to the reference picture index, and may generate a bitstream including the entropy-coded reference picture index. The entropy-coded reference picture index may be signaled from the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-coded reference picture index from the bitstream, and may obtain the reference picture index by applying entropy decoding to the entropy-coded reference picture index.
[0561] When two reference picture lists are used for predicting the target block, a single reference picture index and a single motion vector may be used for each of the reference picture lists. In addition, when two reference picture lists are used for predicting the target block, it may be that the target block designates two prediction blocks. For example, the average value or weighted sum of the two prediction blocks for the target block may be used to generate the (final) prediction block of the target block.
[0562] The motion vector of the target block may be derived by the predicted motion vector index, the MVD, the reference direction, and the reference picture index.
[0563] The decoding device 200 can generate a prediction block for a target block based on the derived motion vector and reference picture index. For example, the prediction block can be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.
[0564] Since the predicted motion vector index and MVD are encoded while the motion vector of the target block itself is not encoded, the number of bits sent from the encoding device 100 to the decoding device 200 can be reduced and the encoding efficiency can be improved.
[0565] For a target block, the motion information of the reconstructed neighboring blocks can be used. In a specific inter-frame prediction mode, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the motion information of the target block, additional information can be encoded, where the additional information enables the motion information of the target block to be derived using the motion information of the reconstructed neighboring blocks. Since the additional information is encoded, the number of bits sent to the decoding device 200 can be reduced and the encoding efficiency can be improved.
[0566] For example, as inter-frame prediction modes in which the motion information of the target block is not directly encoded, there may be a skip mode and / or a merge mode. Here, each of the encoding device 100 and the decoding device 200 can use an identifier and / or index of a unit indicating that the motion information of which among the reconstructed neighboring units will be used as the motion information of the target unit.
[0567] 2) Merge Mode
[0568] As a scheme for deriving the motion information of a target block, there is merge. The term "merge" may mean merging the motions of multiple blocks. "Merge" may also mean that the motion information of one block is also applied to other blocks. In other words, the merge mode can be a mode of deriving the motion information of the target block from the motion information of neighboring blocks.
[0569] When using the merge mode, the encoding device 100 can use the motion information of spatial candidates and / or temporal candidates to predict the motion information of the target block. Spatial candidates can include reconstructed spatial neighboring blocks that are spatially adjacent to the target block. Spatial neighboring blocks can include a left neighboring block and an upper neighboring block. Temporal candidates can include col blocks. The terms "spatial candidate" and "spatial merge candidate" can be used with the same meaning and can be used interchangeably with each other. The terms "temporal candidate" and "temporal merge candidate" can be used with the same meaning and can be used interchangeably with each other.
[0570] The encoding device 100 can obtain a prediction block via prediction. The encoding device 100 can encode a residual block that is the difference between the target block and the prediction block.
[0571] 2-1) Create a merge candidate list
[0572] When using the merge mode, each of the encoding device 100 and the decoding device 200 can create a merge candidate list using the motion information of the spatial candidate and / or the motion information of the temporal candidate. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional. The reference direction may indicate an inter-frame prediction indicator.
[0573] The merge candidate list may include merge candidates. A merge candidate may be motion information. In other words, the merge candidate list may be a list storing multiple pieces of motion information.
[0574] A merge candidate may be the motion information of multiple temporal candidates and / or spatial candidates. In other words, the merge candidate list may include the motion information of temporal candidates and / or spatial candidates, etc.
[0575] In addition, the merge candidate list may include new merge candidates generated by combining the merge candidates already existing in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple pieces of motion information previously existing in the merge candidate list.
[0576] In addition, the merge candidate list may include history-based merge candidates. A history-based merge candidate may be the motion information of a block that has been encoded and / or decoded before the target block.
[0577] In addition, the merge candidate list may include merge candidates based on the average value of two merge candidates.
[0578] A merge candidate may be a specific mode for deriving inter-frame prediction information. A merge candidate may be information indicating a specific mode for deriving inter-frame prediction information. The inter-frame prediction information of the target block can be derived according to the specific mode indicated by the merge candidate. In addition, the specific mode may include a process for deriving a series of inter-frame prediction information. Such a specific mode may be an inter-frame prediction information derivation mode or a motion information derivation mode.
[0579] The inter-frame prediction information of the target block can be derived according to the mode indicated by the merge candidate selected from the merge candidates in the merge candidate list through a merge index.
[0580] For example, the motion information derivation mode in the merge candidate list may be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit and 2) an affine motion information derivation mode.
[0581] In addition, the merge candidate list may include the motion information of a zero vector. The zero vector may also be referred to as a "zero merge candidate".
[0582] In other words, multiple pieces of motion information in the merge candidate list may be at least one of the following information: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple pieces of motion information that previously existed in the merge candidate list, and 4) a zero vector.
[0583] The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter - frame prediction indicator". The reference direction may be unidirectional or bidirectional. The unidirectional reference direction may indicate L0 prediction or L1 prediction.
[0584] The merge candidate list may be created before performing prediction in the merge mode.
[0585] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, such that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made the same as each other using the predefined scheme and the predefined priority.
[0586] Merge may be applied based on a CU or a PU. When performing merge based on a CU or a PU, the encoding device 100 may send a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merge is performed for each block partition, and 2) information about the block to be merged among the blocks that are spatial candidates and / or temporal candidates for the target block.
[0587] 2-2) Search for motion vectors using the merge candidate list
[0588] The encoding device 100 may determine a merge candidate to be used for encoding a target block. For example, the encoding device 100 may perform prediction on the target block using a merge candidate in the merge candidate list, and may generate a residual block for the merge candidate. The encoding device 100 may use the merge candidate that generates the minimum cost in the prediction and the encoding of the residual block to encode the target block.
[0589] In addition, the encoding device 100 may determine whether to encode the target block using the merge mode.
[0590] 2-3) Transmission of inter-frame prediction information
[0591] The encoding device 100 may generate a bitstream including inter-prediction information required for inter-prediction. The encoding device 100 may generate entropy-coded inter-prediction information by performing entropy coding on the inter-prediction information, and may send the bitstream including the entropy-coded inter-prediction information to the decoding device 200. The entropy-coded inter-prediction information may be signaled by the encoding device 100 to the decoding device 200 through the bitstream. The decoding device 200 may extract the entropy-coded inter-prediction information from the bitstream, and may obtain the inter-prediction information by applying entropy decoding to the entropy-coded inter-prediction information.
[0592] The decoding device 200 may perform inter-prediction on a target block using the inter-prediction information of the bitstream.
[0593] The inter-prediction information may include 1) mode information indicating whether the merge mode is used, 2) a merge index, and 3) correction information.
[0594] In addition, the inter-prediction information may include a residual signal.
[0595] The decoding device 200 may obtain the merge index from the bitstream only when the mode information indicates that the merge mode is used.
[0596] The mode information may be a merge flag. The unit of the mode information may be a block. Information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.
[0597] The merge index may indicate a merge candidate among the merge candidates included in the merge candidate list that will be used for predicting the target block. Alternatively, the merge index may indicate a block among the neighboring blocks that are spatially or temporally adjacent to the target block and that will be merged with the target block.
[0598] The encoding device 100 may select a merge candidate having the highest coding performance among the merge candidates included in the merge candidate list, and may set the value of the merge index to indicate the selected merge candidate.
[0599] The correction information may be information for correcting a motion vector. The encoding device 100 may generate the correction information. The decoding device 200 may correct the motion vector of the merge candidate selected by the merge index based on the correction information.
[0600] The correction information may include at least one of information indicating whether correction will be performed, correction direction information, and correction size information. A prediction mode in which the motion vector is corrected based on the signaled correction information may be referred to as a "merge mode with a motion vector difference".
[0601] 2-4) Inter-frame prediction in merge mode using inter-frame prediction information
[0602] The decoding device 200 may perform prediction on a target block using a merge candidate indicated by a merge index among the merge candidates included in the merge candidate list.
[0603] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the merge index.
[0604] 3) Skip Mode
[0605] The skip mode may be a mode in which the motion information of a spatial candidate or a temporal candidate is applied to the target block without change. Additionally, the skip mode may be a mode in which a residual signal is not used. In other words, when the skip mode is used, the reconstructed block may be the same as the predicted block.
[0606] The difference between the merge mode and the skip mode lies in whether a residual signal is sent or used. That is, the skip mode may be similar to the merge mode except that a residual signal is not sent or used.
[0607] When the skip mode is used, the encoding device 100 may send information related to a block among the blocks that are spatial candidates or temporal candidates and whose motion information will be used as the motion information of the target block to the decoding device 200 through a bitstream. The encoding device 100 may generate entropy-coded information by performing entropy coding on this information, and may signal the entropy-coded information to the decoding device 200 through a bitstream. The decoding device 200 may extract the entropy-coded information from the bitstream, and may obtain the information by applying entropy decoding to the entropy-coded information.
[0608] Furthermore, when the skip mode is used, the encoding device 100 may not send other syntax information (such as MVD) to the decoding device 200. For example, when the skip mode is used, the encoding device 100 may not signal syntax elements related to at least one of MVD, coded block flag, and transform coefficient level to the decoding device 200.
[0609] 3-1) Create a merge candidate list
[0610] The skip mode may also use the merge candidate list. In other words, the merge candidate list may be used in both the merge mode and the skip mode. In this regard, the merge candidate list may also be referred to as the "skip candidate list" or the "merge / skip candidate list".
[0611] Optionally, the skip mode may use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate may be replaced with the skip candidate list and the skip candidate, respectively.
[0612] A merge candidate list may be created before performing prediction in the execution skip mode.
[0613] 3-2) Search for motion vectors using the merge candidate list
[0614] The encoding device 100 may determine merge candidates to be used for encoding a target block. For example, the encoding device 100 may perform prediction on the target block using the merge candidates in the merge candidate list. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in the prediction.
[0615] In addition, the encoding device 100 may determine whether to encode the target block using the skip mode.
[0616] 3-3) Transmission of inter-frame prediction information
[0617] The encoding device 100 may generate a bitstream including inter prediction information required for inter prediction. The decoding device 200 may perform inter prediction on the target block using the inter prediction information of the bitstream.
[0618] The inter prediction information may include 1) mode information indicating whether the skip mode is used and 2) a skip index.
[0619] The skip index may be the same as the merge index described above.
[0620] When using the skip mode, the target block may be encoded without using a residual signal. The inter prediction information may not include a residual signal. Optionally, the bitstream may not include a residual signal.
[0621] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index may be the same as each other. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.
[0622] The skip index may indicate the merge candidate among the merge candidates included in the merge candidate list that will be used for predicting the target block.
[0623] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information
[0624] The decoding device 200 may perform prediction on the target block using the merge candidate indicated by the skip index among the merge candidates included in the merge candidate list.
[0625] The motion vector of the target block may be specified by the motion vector, reference picture index, and reference direction of the merge candidate indicated by the skip index.
[0626] 4) Current Picture Reference Mode
[0627] The current picture reference mode can represent such a prediction mode: the prediction mode uses a previously reconstructed region in the target picture to which the target block belongs.
[0628] Motion vectors for specifying the previously reconstructed region can be used. The reference picture index of the target block can be used to determine whether the target block has been encoded in the current picture reference mode.
[0629] A flag or index indicating whether the target block is a block encoded in the current picture reference mode can be signaled by the encoding device 100 to the decoding device 200. Optionally, it can be inferred whether the target block is a block encoded in the current picture reference mode by the reference picture index of the target block.
[0630] When the target block is encoded in the current picture reference mode, the current picture can be present at a fixed position or an arbitrary position in the reference picture list for the target block.
[0631] For example, the fixed position can be the position where the value of the reference picture index is 0 or the last position.
[0632] When the target picture is present at an arbitrary position in the reference picture list, an additional reference picture index indicating such an arbitrary position can be signaled by the encoding device 100 to the decoding device 200.
[0633] 5) Sub-Block Merge Mode
[0634] The sub-block merge mode can be a mode for deriving motion information from sub-blocks of a CU.
[0635] When the sub-block merge mode is applied, motion information of a co-located sub-block (i.e., time-based merge candidate based on sub-blocks) of a target sub-block in a reference image and / or affine control point motion vector merge candidates can be used to generate a sub-block merge candidate list.
[0636] 6) Triangular Partition Mode
[0637] In the triangular partitioning mode, the target block can be partitioned in the diagonal direction, and sub-target blocks generated by the partitioning can be generated. For each sub-target block, motion information corresponding to the sub-target block can be derived, and the derived motion information can be used to derive prediction samples for each sub-target block. The prediction samples of the target block can be derived by a weighted sum of the prediction samples of the sub-target blocks generated via the partitioning.
[0638] 7) Combined Inter-Frame - Intra-Frame Prediction Mode
[0639] The combined inter-intra prediction mode can be a mode for deriving the prediction samples of the target block by using a weighted sum of the prediction samples generated via inter-frame prediction and the prediction samples generated via intra-frame prediction.
[0640] In the above mode, the decoding device 200 can autonomously correct the derived motion information. For example, the decoding device 200 can search for motion information with the minimum sum of absolute differences (SAD) in a specific region based on the reference block indicated by the derived motion information, and can export the found motion information as corrected motion information.
[0641] In the above mode, the decoding device 200 can use optical flow to compensate for the predicted samples derived via inter-frame prediction.
[0642] In the AMVP mode, merge mode, skip mode, etc. described above, the index information of the list can be used to specify the motion information among multiple pieces of motion information in the list that will be used for predicting the target block.
[0643] To improve the encoding efficiency, the encoding device 100 can signal only the index of the element among the elements in the list that generates the minimum cost in the inter-frame prediction of the target block. The encoding device 100 can encode the index and signal the encoded index.
[0644] Therefore, it must be possible for the encoding device 100 and the decoding device 200 to derive the above-described lists (i.e., the predicted motion vector candidate list and the merge candidate list) based on the same data using the same scheme. Here, the same data may include the reconstructed picture and the reconstructed block. In addition, to specify an element using an index, the order of the elements in the list must be fixed.
[0645] Figure 10 Shows spatial candidates according to an embodiment.
[0646] In Figure 10 the position of the spatial candidate is shown.
[0647] The large block at the center of the figure may represent the target block. The five small blocks may represent spatial candidates.
[0648] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).
[0649] Spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying the pixels at coordinates (xP - 1, yP + nPSH).
[0650] Spatial candidate A1 may be a block adjacent to the left side of the target block. A1 may be the lowermost block among the blocks adjacent to the left side of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block occupying the pixels at coordinates (xP - 1, yP + nPSH - 1).
[0651] Spatial candidate B0 can be a block adjacent to the upper right corner of the target block. B0 can be a block that occupies the pixel at coordinates (xP + nPSW, yP - 1).
[0652] Spatial candidate B1 can be a block adjacent to the top of the target block. B1 can be the rightmost block among the blocks adjacent to the top of the target block. Optionally, B1 can be a block adjacent to the left side of B0. B1 can be a block that occupies the pixel at coordinates (xP + nPSW - 1, yP - 1).
[0653] Spatial candidate B2 can be a block adjacent to the upper left corner of the target block. B2 can be a block that occupies the pixel at coordinates (xP - 1, yP - 1).
[0654] Determination of the availability of spatial and temporal candidates
[0655] In order to include the motion information of the spatial candidate or the motion information of the temporal candidate in the list, it is necessary to determine whether the motion information of the spatial candidate or the motion information of the temporal candidate is available.
[0656] Hereinafter, the candidate block may include a spatial candidate and a temporal candidate.
[0657] For example, the determination can be performed by sequentially applying the following steps 1) to 4).
[0658] Step 1) When the PU including the candidate block is outside the boundary of the picture, the availability of the candidate block can be set to "false". The expression "the availability is set to false" may have the same meaning as "set to unavailable".
[0659] Step 2) When the PU including the candidate block is outside the boundary of the slice, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different slices, the availability of the candidate block can be set to "false".
[0660] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block can be set to "false".
[0661] Step 4) When the prediction mode of the PU including the candidate block is the intra prediction mode, the availability of the candidate block can be set to "false". When the PU including the candidate block does not use inter prediction, the availability of the candidate block can be set to "false".
[0662] Figure 11 Illustrates the order of adding the motion information of the spatial candidate to the merge list according to an embodiment.
[0663] As Figure 11As shown, when multiple pieces of motion information of spatial candidates are added to the merge list, the order of A1, B1, B0, A0, and B2 can be used. That is, multiple pieces of motion information of available spatial candidates can be added to the merge list in the order of A1, B1, B0, A0, and B2.
[0664] Method for deriving a merge list in merge mode and skip mode
[0665] As described above, the maximum number of merge candidates in the merge list can be set. The set maximum number can be indicated by "N". The set number can be sent from the encoding device 100 to the decoding device 200. The strip header of the strip can include N. In other words, the maximum number of merge candidates in the merge list for the target block of the strip can be set through the strip header. For example, the value of N can basically be 5.
[0666] Multiple pieces of motion information (i.e., merge candidates) can be added to the merge list in the order of the following steps 1) to 4).
[0667] Step 1) Among the spatial candidates, available spatial candidates can be added to the merge list. Multiple pieces of motion information of the available spatial candidates can be added to the merge list in the order shown in Figure 11 Here, when the motion information of the available spatial candidate overlaps with other motion information already existing in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information existing in the list can be simply referred to as "overlap check".
[0668] The maximum number of added motion information can be N.
[0669] Step 2) When the number of motion information in the merge list is less than N and temporal candidates are available, the motion information of the temporal candidates can be added to the merge list. Here, when the motion information of the available temporal candidate overlaps with other motion information already existing in the merge list, the motion information of the available temporal candidate may not be added to the merge list.
[0670] Step 3) When the number of motion information in the merge list is less than N and the type of the target strip is "B", the combined motion information generated by combining bidirectional prediction (bi-prediction) can be added to the merge list.
[0671] The target strip can be the strip including the target block.
[0672] The combined motion information can be a combination of L0 motion information and L1 motion information. The L0 motion information can be the motion information that only refers to the L0 reference picture list. The L1 motion information can be the motion information that only refers to the L1 reference picture list.
[0673] In the merge list, there may be one or more L0 motion information. In addition, in the merge list, there may be one or more L1 motion information.
[0674] The combined motion information may include one or more combined motion information. When generating the combined motion information, the L0 motion information and the L1 motion information among the one or more L0 motion information and the one or more L1 motion information that will be used to generate the steps of the combined motion information may be predefined. One or more combined motion information may be generated in a predefined order by using a combination of bidirectional prediction of a pair of different motion information in the merge list. One of the pair of different motion information may be L0 motion information, and the other of the pair of different motion information may be L1 motion information.
[0675] For example, the combined motion information with the highest priority added may be a combination of the L0 motion information with merge index 0 and the L1 motion information with merge index 1. When the motion information with merge index 0 is not L0 motion information or when the motion information with merge index 1 is not L1 motion information, the combined motion information may neither be generated nor added. Next, the combined motion information with the next priority added may be a combination of the L0 motion information with merge index 1 and the L1 motion information with merge index 0. Subsequent detailed combinations may conform to other combinations in the field of video coding / decoding.
[0676] Here, when the combined motion information overlaps with other motion information already existing in the merge list, the combined motion information may not be added to the merge list.
[0677] Step 4) When the number of motion information in the merge list is less than N, the motion information of the zero vector may be added to the merge list.
[0678] The zero vector motion information may be motion information whose motion vector is a zero vector.
[0679] The number of zero vector motion information may be one or more. The reference picture indices of one or more zero vector motion information may be different from each other. For example, the value of the reference picture index of the first zero vector motion information may be 0. The value of the reference picture index of the second zero vector motion information may be 1.
[0680] The number of zero vector motion information may be the same as the number of reference pictures in the reference picture list.
[0681] The reference direction of the zero vector motion information can be bidirectional. Both motion vectors can be zero vectors. The number of zero vector motion information can be the smaller one of the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1. Optionally, when the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1 are different from each other, a unidirectional reference direction can be used for the reference picture index that can be applied only to a single reference picture list.
[0682] Encoding device 100 and / or decoding device 200 can then add the zero vector motion information to the merge list while changing the reference picture index.
[0683] When the zero vector motion information overlaps with other motion information already existing in the merge list, the zero vector motion information may not be added to the merge list.
[0684] The order of the above steps 1) to 4) is only exemplary and can be changed. In addition, some of the above steps can be omitted according to predefined conditions.
[0685] Method for deriving a list of predicted motion vector candidates in AMVP mode
[0686] The maximum number of predicted motion vector candidates in the predicted motion vector candidate list can be predefined. The predefined maximum number can be indicated by N. For example, the predefined maximum number can be 2.
[0687] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of the following steps 1) to 3).
[0688] Step 1) The available spatial candidates among the spatial candidates can be added to the predicted motion vector candidate list. The spatial candidates can include a first spatial candidate and a second spatial candidate.
[0689] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.
[0690] Multiple pieces of motion information of available space candidates may be added to the predicted motion vector candidate list in the order of the first space candidate and the second space candidate. In this case, when the motion information of an available space candidate overlaps with other motion information already present in the predicted motion vector candidate list, the motion information of the available space candidate may not be added to the predicted motion vector candidate list. In other words, when the value of N is 2, if the motion information of the second space candidate is the same as that of the first space candidate, the motion information of the second space candidate may not be added to the predicted motion vector candidate list.
[0691] The maximum number of pieces of motion information to be added may be N.
[0692] Step 2) When the number of pieces of motion information in the predicted motion vector candidate list is less than N and time candidates are available, the motion information of the time candidates may be added to the predicted motion vector candidate list. In this case, when the motion information of an available time candidate overlaps with other motion information already present in the predicted motion vector candidate list, the motion information of the available time candidate may not be added to the predicted motion vector candidate list.
[0693] Step 3) When the number of pieces of motion information in the predicted motion vector candidate list is less than N, zero vector motion information may be added to the predicted motion vector candidate list.
[0694] The zero vector motion information may include one or more pieces of zero vector motion information. The reference picture indices of the one or more pieces of zero vector motion information may be different from each other.
[0695] The encoding device 100 and / or the decoding device 200 may sequentially add multiple pieces of zero vector motion information to the predicted motion vector candidate list while changing the reference picture index.
[0696] When the zero vector motion information overlaps with other motion information already present in the predicted motion vector candidate list, the zero vector motion information may not be added to the predicted motion vector candidate list.
[0697] The description of the zero vector motion information made in conjunction with the merge list above may also apply to the zero vector motion information. The repetitive description thereof will be omitted.
[0698] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some of the steps may be omitted according to predefined conditions.
[0699] Figure 12 Illustrates the transform and quantization processing according to an example.
[0700] As Figure 12As shown, a quantized level can be generated by performing a transform and / or quantization process on a residual signal.
[0701] The residual signal can be generated as the difference between an original block and a predicted block. Here, the predicted block can be a block generated via intra prediction or inter prediction.
[0702] The residual signal can be transformed into a signal in the frequency domain by a transform process that is part of the quantization process.
[0703] The transform kernels used for the transform can include various DCT kernels, such as the discrete cosine transform (DCT) type 2 (DCT-II) and the discrete sine transform (DST) kernels.
[0704] These transform kernels can perform a separable transform or a two-dimensional (2D) non-separable transform on the residual signal. A separable transform can be a transform that indicates performing a one-dimensional (1D) transform on the residual signal in each of the horizontal and vertical directions.
[0705] The DCT types and DST types adaptively used for the 1D transform can include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II, as shown in each of Table 3 and Table 4 below.
[0706] Table 3
[0707] Transform Set Transform Candidate 0 DST-VII, DCT-VIII 1 DST-VII, DST-I 2 DST-VII, DCT-V
[0708] Table 4
[0709] Transform Set Transform Candidate 0 DST-VII, DCT-VIII, DST-I 1 DST-VII, DST-I, DCT-VIII 2 DST-VII, DCT-V, DST-I
[0710] As shown in Table 3 and Table 4, when deriving the DCT type or DST type to be used for the transform, a transform set can be used. Each transform set can include multiple transform candidates. Each transform candidate can be a DCT type or a DST type.
[0711] Table 5 below shows examples of the transform set to be applied to the horizontal direction and the transform set to be applied to the vertical direction according to the intra prediction mode.
[0712] Table 5
[0713]
[0714]
[0715] In Table 5, the numbers of the vertical transform set and the horizontal transform set to be applied to the horizontal direction of the residual signal according to the intra prediction mode of the target block are shown.
[0716] As Figure 4 and Figure 5As illustrated, a set of transforms to be applied to the horizontal and vertical directions can be predefined according to the intra prediction mode of the target block. The encoding device 100 can perform transform and inverse transform on the residual signal using the transforms included in the transform set corresponding to the intra prediction mode of the target block. In addition, the decoding device 200 can perform inverse transform on the residual signal using the transforms included in the transform set corresponding to the intra prediction mode of the target block.
[0717] In the transform and inverse transform, as illustrated in Tables 3, 4, and 5, the set of transforms to be applied to the residual signal can be determined and not signaled. Transform indication information can be signaled from the encoding device 100 to the decoding device 200. The transform indication information can be information indicating which one of a plurality of transform candidates included in the set of transforms to be applied to the residual signal is used.
[0718] For example, when the size of the target block is 64×64, a transform set each having three transforms can be configured according to the intra prediction mode. The optimal transform method can be selected from a total of nine multi-transform methods generated by the combination of three transforms in the horizontal direction and three transforms in the vertical direction. By such an optimal transform method, the residual signal can be encoded and / or decoded, and thus the encoding efficiency can be improved.
[0719] Here, information indicating which one of the plurality of transforms belonging to each transform set has been used for at least one of the vertical transform and the horizontal transform can be entropy encoded and / or entropy decoded. Here, truncated binary unary coding can be used to encode and / or decode such information.
[0720] As described above, methods using various transforms can be applied to the residual signal generated via intra prediction or inter prediction.
[0721] The transform can include at least one of a first transform and a secondary transform. The transform coefficients can be generated by performing the first transform on the residual signal, and the secondary transform coefficients can be generated by performing the secondary transform on the transform coefficients.
[0722] The first transform can be referred to as the "primary transform". In addition, the first transform can also be referred to as the "adaptive multi-transform (AMT) scheme". As described above, AMT can represent applying different transforms to each 1D direction (i.e., the vertical direction and the horizontal direction).
[0723] The secondary transform can be a transform for improving the energy concentration of the transform coefficients generated by the first transform. Similar to the first transform, the secondary transform can be a separable transform or a non-separable transform. Such a non-separable transform can be a non-separable secondary transform (NSST).
[0724] The first transformation can be performed using at least one of a plurality of predefined transformation methods. For example, the plurality of predefined transformation methods may include a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen - Loeve transform (KLT), etc.
[0725] In addition, according to the kernel functions that define the discrete cosine transform (DCT) or the discrete sine transform (DST), the first transformation can be a transformation of various types.
[0726] For example, the type of transformation can be determined based on at least one of the following: 1) the prediction mode of the target block (e.g., one of intra - prediction and inter - prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra - prediction mode of the target block, 5) the component of the target block (e.g., one of the luminance component and the chrominance component), and 6) the partition type applied to the target block (e.g., one of quadtree, binary tree, and ternary tree).
[0727] For example, according to the transformation kernels presented in Table 6 below, the first transformation may include transformations such as DCT - 2, DCT - 5, DCT - 7, DST - 7, DST - 1, DST - 8, and DCT - 8. In Table 6 below, various transformation types and transformation kernel functions for multi - transformation selection (MTS) are illustrated.
[0728] MTS may refer to the selection of a combination of one or more DCT and / or DST kernels in order to transform the residual signal in the horizontal and / or vertical directions.
[0729] Table 6
[0730]
[0731] In Table 6, i and j can be integer values equal to or greater than 0 and less than or equal to N - 1.
[0732] A secondary transformation can be performed on the transform coefficients generated by performing the first transformation.
[0733] As in the first transformation, a set of transformations can also be defined in the secondary transformation. The methods for deriving and / or determining the above - mentioned set of transformations can be applied not only to the first transformation but also to the secondary transformation.
[0734] The first transformation and the secondary transformation can be determined for a specific target.
[0735] For example, a first transform and a secondary transform may be applied to signal components corresponding to one or more of a luma component and a chroma component. Whether to apply the first transform and / or the secondary transform may be determined according to at least one of coding parameters for a target block and / or neighboring blocks. For example, whether to apply the first transform and / or the secondary transform may be determined according to the size and / or shape of the target block.
[0736] In the encoding device 100 and the decoding device 200, transform information indicating a transform method to be used for a target may be derived by using specified information.
[0737] For example, the transform information may include a transform index to be used for a primary transform and / or a secondary transform. Alternatively, the transform information may indicate that the primary transform and / or the secondary transform is not used.
[0738] For example, when the target of the primary transform and the secondary transform is a target block, the transform method indicated by the transform information to be applied to the primary transform and / or the secondary transform may be determined according to at least one of coding parameters for the target block and / or blocks neighboring the target block.
[0739] Alternatively, the transform information indicating the transform method for a specific target may be signaled from the encoding device 100 to the decoding device 200.
[0740] For example, for a single CU, whether to use the primary transform, an index indicating the primary transform, whether to use the secondary transform, and an index indicating the secondary transform may be derived as transform information by the decoding device 200. Alternatively, for a single CU, the transform information indicating the following items may be signaled: whether to use the primary transform, an index indicating the primary transform, whether to use the secondary transform, and an index indicating the secondary transform.
[0741] Quantized transform coefficients (i.e., quantized levels) may be generated by performing quantization on the result generated by performing the first transform and / or the secondary transform or by performing quantization on a residual signal.
[0742] Figure 13 Diagonal scanning according to an example is shown.
[0743] Figure 14 Horizontal scanning according to an example is shown.
[0744] Figure 15 Vertical scanning according to an example is shown.
[0745] The quantized transform coefficients may be scanned via at least one of (upper right) diagonal scanning, vertical scanning, and horizontal scanning according to at least one of an intra prediction mode, a block size, and a block shape. The block may be a transform unit (TU).
[0746] Each scan can be initiated at a specific starting point and terminated at a specific ending point.
[0747] For example, by using Figure 13 a diagonal scan of Figure 14 to scan the coefficients of a block, the quantized transform coefficients can be changed into a 1D vector form. Alternatively, a Figure 15 horizontal scan of
[0748] or a
[0749] vertical scan of
[0750] can be used according to the size of the block and / or the intra prediction mode without using a diagonal scan. Figure 13 、 Figure 14 and Figure 15 As shown in
[0751] The quantized transform coefficients can be represented by a block shape. Each block can include a plurality of sub-blocks. Each sub-block can be defined according to the minimum block size or the minimum block shape.
[0752] In the scan, the scan order according to the type or direction of the scan can be first applied to the sub-blocks. In addition, the scan order according to the direction of the scan can be applied to the quantized transform coefficients in each sub-block.
[0753] For example, as shown in Figure 13 、 Figure 14 and Figure 15 when the size of the target block is 8×8, the quantized transform coefficients can be generated by the first transform, secondary transform and quantization of the residual signal of the target block. Therefore, one of the three types of scan orders can be applied to four 4×4 sub-blocks, and the quantized transform coefficients can also be scanned for each 4×4 sub-block according to the scan order.
[0754] The encoding device 100 can generate entropy-coded quantized transform coefficients by performing entropy coding on the scanned quantized transform coefficients, and can generate a bitstream including the entropy-coded quantized transform coefficients.
[0755] The decoding device 200 can extract entropy-coded quantized transform coefficients from a bitstream and generate quantized transform coefficients by performing entropy decoding on the entropy-coded quantized transform coefficients. The quantized transform coefficients can be arranged in the form of 2D blocks via inverse scanning. Here, as a method of inverse scanning, at least one of right upper diagonal scanning, vertical scanning, and horizontal scanning can be performed.
[0756] In the decoding device 200, inverse quantization can be performed on the quantized transform coefficients. Depending on whether secondary inverse transformation is performed, secondary inverse transformation can be performed on the result generated by performing inverse quantization. In addition, depending on whether primary inverse transformation will be performed, primary inverse transformation can be performed on the result generated by performing secondary inverse transformation. A reconstructed residual signal can be generated by performing primary inverse transformation on the result generated by performing secondary inverse transformation.
[0757] For a luminance component reconstructed by intra prediction or inter prediction, inverse mapping with a dynamic range can be performed before loop filtering.
[0758] The dynamic range can be divided into 16 equal segments, and the mapping function of the corresponding segment can be signaled. Such a mapping function can be signaled at the slice level or the parallel block group level.
[0759] An inverse mapping function for performing inverse mapping can be derived based on the mapping function.
[0760] Loop filtering, storage of reference pictures, and motion compensation can be performed in the inverse mapping region.
[0761] A prediction block generated by inter prediction can be transformed into the mapping region by mapping using the mapping function, and the transformed prediction block can be used to generate a reconstructed block. However, since intra prediction is performed in the mapping region, a prediction block generated by intra prediction can be used to generate a reconstructed block without the need for mapping and / or inverse mapping.
[0762] For example, when the target block is a residual block of a chrominance component, the residual block can be transformed into the inverse mapping region by scaling the chrominance component of the mapping region.
[0763] Whether scaling is available can be signaled at the slice level or the parallel block group level.
[0764] For example, scaling can be applied only to cases where mapping is available for the luminance component and the partitions of the luminance component and the chrominance component follow the same tree structure.
[0765] Scaling can be performed based on the average value of the samples in the luminance prediction block corresponding to the chrominance prediction block. Here, when the target block uses inter prediction, the luminance prediction block can represent the mapped luminance prediction block.
[0766] The values required for scaling can be derived by referring to a lookup table using the index of the segment to which the average value of the samples of the luminance prediction block belongs.
[0767] The residual block can be transformed into the inverse mapping region by scaling the residual block using the finally derived values. Thereafter, for the blocks of the chrominance components, reconstruction, intra prediction, inter prediction, loop filtering, and storage of the reference pictures can be performed in the inverse mapping region.
[0768] For example, information indicating whether mapping and / or inverse mapping of the luminance component and the chrominance component are available can be signaled by the sequence parameter set.
[0769] A prediction block of the target block can be generated based on the block vector. The block vector can indicate the displacement between the target block and the reference block. The reference block can be a block in the target image.
[0770] In this way, the prediction mode of generating the prediction block by referring to the target image can be referred to as the "intra block copy (IBC) mode".
[0771] The IBC mode can be applied to a CU having a specific size. For example, the IBC mode can be applied to a CU of M×N. Here, M and N can be less than or equal to 64.
[0772] The IBC mode can include a skip mode, a merge mode, an AMVP mode, etc. In the case of the skip mode or the merge mode, a merge candidate list can be configured, and a merge index can be signaled, and thus a single merge candidate can be specified among the merge candidates present in the merge candidate list. The block vector of the specified merge candidate can be used as the block vector of the target block.
[0773] In the case of the AMVP mode, a differential block vector can be signaled. In addition, a prediction block vector can be derived from the left neighboring block and the upper neighboring block of the target block. In addition, an index indicating which neighboring block will be used can be signaled.
[0774] The prediction block in the IBC mode can be included in the target CTU or the left CTU, and can be limited to the blocks within the previously reconstructed region. For example, the value of the block vector can be restricted such that the prediction block of the target block is located in a specific region. The specific region can be a region defined by three 64×64 blocks that have been encoded and / or decoded before the 64×64 block including the target block. By restricting the value of the block vector in this way, the memory consumption and device complexity caused by the implementation of the IBC mode can be reduced.
[0775] Figure 16 It is a configuration diagram of an encoding device according to an embodiment.
[0776] The encoding device 1600 can correspond to the encoding device 100 described above.
[0777] The encoding device 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a storage 1640 that communicate with each other via a bus 1690. The encoding device 1600 may also include a communication unit 1620 connected to a network 1699.
[0778] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1630 or the storage 1640. The processing unit 1610 may be at least one hardware processor.
[0779] The processing unit 1610 may generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1610.
[0780] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transformation unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transformation unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.
[0781] At least some of the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transformation unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transformation unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190 may be program modules and may communicate with an external device or system. The program modules may be included in the encoding device 1600 in the form of an operating system, an application program module, or other program modules.
[0782] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the encoding device 1600.
[0783] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.
[0784] The program modules may be implemented using instructions or code run by at least one processor of the encoding device 1600.
[0785] The processing unit 1610 can execute instructions or codes in the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190.
[0786] The storage unit can represent the memory 1630 and / or the storage 1640. Each of the memory 1630 and the storage 1640 can be any of various types of volatile or non-volatile storage media. For example, the memory 1630 can include at least one of a read-only memory (ROM) 1631 and a random access memory (RAM) 1632.
[0787] The storage unit can store data or information for the operation of the encoding device 1600. In an embodiment, the data or information of the encoding device 1600 can be stored in the storage unit.
[0788] For example, the storage unit can store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.
[0789] The encoding device 1600 can be implemented in a computer system including a computer-readable storage medium.
[0790] The storage medium can store at least one module required for the operation of the encoding device 1600. The memory 1630 can store at least one module and can be configured such that the at least one module is executed by the processing unit 1610.
[0791] Functions related to the communication of data or information of the encoding device 1600 can be performed through the communication unit 1620.
[0792] For example, the communication unit 1620 can send a bitstream to the decoding device 1700 to be described later.
[0793] Figure 17 is a configuration diagram of a decoding device according to an embodiment.
[0794] The decoding device 1700 can correspond to the decoding device 200 described above.
[0795] The decoding device 1700 can include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage 1740 that communicate with each other through a bus 1790. The decoding device 1700 can also include a communication unit 1720 connected to a network 1799.
[0796] The processing unit 1710 may be a central processing unit (CPU) or a semiconductor device for running processing instructions stored in the memory 1730 or the storage 1740. The processing unit 1710 may be at least one hardware processor.
[0797] The processing unit 1710 may generate and process signals, data, or information that are input to, output from, or used in the decoding device 1700, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in an embodiment, the generation and processing of data or information and the checks, comparisons, and determinations related to the data or information may be performed by the processing unit 1710.
[0798] The processing unit 1710 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, an inter prediction unit 250, a switch 245, an adder 255, a filter unit 260, and a reference picture buffer 270.
[0799] At least some of the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra prediction unit 240, the inter prediction unit 250, the adder 255, the switch 245, the filter unit 260, and the reference picture buffer 270 of the decoding device 200 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding device 1700 in the form of an operating system, an application module, or other program modules.
[0800] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device capable of communicating with the decoding device 1700.
[0801] The program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.
[0802] The program modules may be implemented using instructions or code run by at least one processor of the decoding device 1700.
[0803] The processing unit 1710 may run the instructions or code in the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra prediction unit 240, the inter prediction unit 250, the switch 245, the adder 255, the filter unit 260, and the reference picture buffer 270.
[0804] The storage unit may represent memory 1730 and / or storage 1740. Each of memory 1730 and storage 1740 may be any of various types of volatile or non-volatile storage media. For example, memory 1730 may include at least one of ROM 1731 and RAM 1732.
[0805] The storage unit may store data or information for the operation of decoding device 1700. In an embodiment, the data or information of decoding device 1700 may be stored in the storage unit.
[0806] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.
[0807] Decoding device 1700 may be implemented in a computer system including a computer-readable storage medium.
[0808] The storage medium may store at least one module required for the operation of decoding device 1700. Memory 1730 may store at least one module and may be configured such that the at least one module is run by processing unit 1710.
[0809] Functions related to the communication of data or information with decoding device 1700 may be performed by communication unit 1720.
[0810] For example, communication unit 1720 may receive a bitstream from encoding device 1700.
[0811] Hereinafter, the processing unit may represent processing unit 1610 of encoding device 1600 and / or processing unit 1710 of decoding device 1700. For example, with respect to functions related to prediction, the processing unit may represent switch 115 and / or switch 245. With respect to functions related to inter-frame prediction, the processing unit may represent inter-frame prediction unit 110, subtractor 125, and adder 175, and may represent inter-frame prediction unit 250 and adder 255. With respect to functions related to intra-frame prediction, the processing unit may represent intra-frame prediction unit 120, subtractor 125, and adder 175, and may represent intra-frame prediction unit 240 and adder 255. With respect to functions related to transformation, the processing unit may represent transformation unit 130 and inverse transformation unit 170, and may represent inverse transformation unit 230. With respect to functions related to quantization, the processing unit may represent quantization unit 140 and dequantization unit 160, and may indicate dequantization unit 220. With respect to functions related to entropy encoding and / or entropy decoding, the processing unit may represent entropy encoding unit 150 and / or entropy decoding unit 210. With respect to functions related to filtering, the processing unit may represent filter unit 180 and / or filter unit 260. With respect to functions related to reference pictures, the processing unit may indicate reference picture buffer 190 and / or reference picture buffer 270.
[0812] Figure 18 is a flowchart showing a target block prediction method and a bitstream generation method according to an embodiment.
[0813] The target block prediction method and the bitstream generation method according to an embodiment can be executed by an encoding device 1600. This embodiment can be part of a target block encoding method or a video encoding method.
[0814] In step 1810, the processing unit 1610 may determine prediction information to be applied to the encoding of a target block.
[0815] The prediction information may include information for the above prediction. For example, the prediction information may include inter prediction information. For example, the prediction information may include intra prediction information.
[0816] The prediction information may include encoding information.
[0817] In step 1820, the processing unit 1610 may perform prediction on the target block using information about the target block and the determined prediction information.
[0818] The prediction may be inter prediction.
[0819] A prediction block may be generated through the prediction of the target block.
[0820] When performing prediction on the target block, padding for an image related to the prediction may be performed.
[0821] For example, the image may include a target image and a reference image.
[0822] For example, padding may be performed on an area outside the image.
[0823] A residual block that is the difference between the target block and the prediction block may be generated. Information about the target block may be generated by applying a transform and quantization to the residual block.
[0824] The information about the target block may include transform coefficients and quantization coefficients for the target block. The information about the target block may include prediction information.
[0825] In addition, a reconstructed block that is the sum of the prediction block and the reconstruction residual block may be generated.
[0826] In step 1830, the processing unit 1610 may generate a bitstream.
[0827] The bitstream may include information about the target block. In addition, the bitstream may include the information described above in the embodiment. For example, the bitstream may include encoding parameters related to the target block and / or attributes of the target block.
[0828] The information included in the bitstream may be generated at step 1830, or may be at least partially generated at steps 1810 and 1820.
[0829] The processing unit 1610 may store the generated bitstream in the memory 1640. Alternatively, the communication unit 1620 may send the bitstream to the decoding device 1700.
[0830] The bitstream may include encoding information about the target block. The processing unit 1610 may generate the encoding information about the target block by performing entropy encoding on the information about the target block.
[0831] Figure 19 is a flowchart of a target block prediction method using a bitstream according to an embodiment.
[0832] The target block prediction method using a bitstream according to an embodiment may be performed by the decoding device 1700. This embodiment may be part of a target block decoding method or a video decoding method.
[0833] At step 1910, the communication unit 1720 may obtain the bitstream. The communication unit 1720 may receive the bitstream from the encoding device 1600.
[0834] The bitstream may include information about the target block.
[0835] The information about the target block may include transform coefficients and quantization coefficients for the target block. The information about the target block may include prediction information.
[0836] In addition, the bitstream may include the information described above in the embodiment. For example, the bitstream may include encoding parameters related to the target block and / or the attributes of the target block.
[0837] The computer-readable storage medium may include the bitstream, and may use the information about the target block included in the bitstream to perform prediction and decoding of the target block.
[0838] The bitstream may include encoding information about the target block. The processing unit 1710 may generate the information about the target block by performing entropy decoding on the encoding information about the target block.
[0839] The processing unit 1710 may store the obtained bitstream in the memory 1740.
[0840] At step 1920, the processing unit 1710 may determine the prediction information to be applied to the decoding of the target block.
[0841] The processing unit 1710 may use the method used in the above embodiment to determine the prediction information.
[0842] The processing unit 1710 can determine prediction information for a target block based on information related to a prediction method obtained from a bitstream.
[0843] The prediction information can include inter-frame prediction information. The prediction information can include intra-frame prediction information.
[0844] The prediction information can include coding information.
[0845] In step 1930, the processing unit 1710 can perform prediction on the target block using information about the target block and the determined prediction information.
[0846] The prediction can be inter-frame prediction.
[0847] In step 1930, a prediction block can be generated by performing prediction on the target block using the prediction information.
[0848] When performing prediction on a target block, padding for an image related to the prediction can be performed.
[0849] For example, the image can include a target image and a reference image.
[0850] For example, padding can be performed on an area outside the image.
[0851] In addition, a reconstructed block can be generated as the sum of the prediction block and the reconstructed residual block.
[0852] Padding of the image in inter-frame prediction
[0853] In inter-frame prediction, reference to an area outside the reference image can be restricted. To make an area outside the reference image available, padding can be performed on the area outside the image.
[0854] Padding can refer to an operation of deriving pixel values of pixels in an unavailable area. Through padding, the unavailable area can be filled with values. When the unavailable area is filled with values generated by padding, the filled values can be provided for reference of the area.
[0855] In the case where a reference block includes an area outside the reference image when performing inter-frame prediction on a target block, a prediction block similar to the original block can be generated by padding the area outside the image, as described in the embodiments. The original block can be a block in the original image that occupies a position corresponding to the position of the target block.
[0856] Through this padding, reference, and generation, the number of bits of the residual signal used to signal the target block can be reduced, and thus the coding efficiency can be improved.
[0857] The original block can refer to a block in the original image corresponding to the position of the target block.
[0858] OOP Pixel
[0859] In an embodiment, out-of-picture (OOP) pixels may refer to at least one of the following 1) to 4):
[0860] 1) OOP pixels may refer to pixels in a region outside the reference image among the pixels of a reference block when performing inter-frame prediction.
[0861] In an embodiment, the reference block and the reference image may refer to 1) the reference block in the L0 direction and the reference image in the L0 direction or 2) the reference block in the L1 direction and the reference image in the L1 direction.
[0862] 2) OOP pixels may refer to pixels in a region where the corresponding reference block pixels are outside the reference image among the pixels of a prediction block in the L0 direction and / or the L1 direction when performing inter-frame prediction.
[0863] In an embodiment, the prediction block, the reference block, and the reference image may be 1) the prediction block in the L0 direction, the reference block in the L0 direction, and the reference image in the L0 direction, or 2) the prediction block in the L1 direction, the reference block in the L1 direction, and the reference image in the L1 direction.
[0864] 3) OOP pixels may refer to pixels among the inter-frame prediction pixels of a target block (or target image) where the corresponding reference block pixels in the L0 direction and / or the reference block pixels in the L1 direction are in a region outside the reference image.
[0865] 4) OOP pixels may refer to pixels in an OOP region.
[0866] In an embodiment, non-OOP pixels may refer to pixels other than OOP pixels.
[0867] OOP Region
[0868] In an embodiment, an out-of-picture (OOP) region may refer to at least one of the following 1) to 4).
[0869] 1) The OOP region may refer to a region in a region outside the reference image among the regions of a reference block when performing inter-frame prediction.
[0870] 2) The OOP region may refer to a region in a region where the corresponding reference block region is outside the reference image among the regions of a prediction block in the L0 direction and / or the L1 direction when performing inter-frame prediction.
[0871] 3) The OOP region may refer to a region among the inter-frame prediction pixels of a target block (or target image) where the corresponding reference block region in the L0 direction and / or the corresponding reference block region in the L1 direction is a region outside the reference image.
[0872] 4) The OOP region may refer to a group of OOP pixels.
[0873] In an embodiment, the non-OOP region may be a region other than the OOP region.
[0874] Determination of OOP Pixel
[0875] When a specific pixel in a reference block (or prediction block) in a specific LX direction satisfies at least one of [Equation 1] to [Equation 4], the specific pixel may be determined as an OOP pixel.
[0876] When a specific pixel does not satisfy any of the following [Equation 1] to [Equation 4], the specific pixel may not be an OOP pixel.
[0877] [Equation 1]
[0878] (Pos i,j .x + Mv i,j .hor) > (Pos RightB .x + half_pixel)
[0879] [Equation 2]
[0880] (Pos i,j .x + Mv i,j .hor) > (Pos LeftB .x + half_pixel)
[0881] [Equation 3]
[0882] (Pos i,j .y + Mv i,j .ver) > (Pos BottomB .y + half_pixel)
[0883] [Equation 4]
[0884] (Pos i,j .y + Mv i,j .ver) > (Pos TopB .y + half_pixel)
[0885] The specific pixel may refer to the coordinates of the pixel corresponding to the i-th position starting from the left and the j-th position starting from the top within the reference block (or prediction block).
[0886] For example, i may be at least one of the values from 1 to the value corresponding to the width of the reference block.
[0887] For example, j may be at least one of the values from 1 to the value corresponding to the height of the reference block.
[0888] Pos i,j can represent the coordinates of the pixel corresponding to the i-th position starting from the left and the j-th position starting from the top within a reference block (or a predicted block). That is, Pos i,j can represent the coordinates of a specific pixel.
[0889] Pos RightB can represent the coordinates of the pixel corresponding to the right boundary of the image. In other words, Pos RightB can represent the coordinates of the pixel adjacent to the right boundary of the image within the image.
[0890] Pos LeftB can represent the coordinates of the pixel corresponding to the left boundary of the image. In other words, Pos LeftB can represent the coordinates of the pixel adjacent to the left boundary of the image within the image.
[0891] Pos BottomB can represent the coordinates of the pixel corresponding to the bottom boundary of the image. In other words, Pos BottomB can represent the coordinates of the pixel adjacent to the bottom boundary of the image within the image.
[0892] Pos TopB can represent the coordinates of the pixel corresponding to the top boundary of the image. In other words, Pos TopB can represent the coordinates of the pixel adjacent to the top boundary of the image within the image.
[0893] Mv i,j can represent the motion vector corresponding to the pixel corresponding to the i-th position starting from the left and the j-th position starting from the top within the target block.
[0894] Mv i,j can represent the motion vector of the pixel corresponding to the i-th position starting from the left and the j-th position starting from the top within the target block.
[0895] For a specific motion information Mv, Mv.hor can represent the horizontal component of Mv.
[0896] For a specific motion information Mv, Mv.ver can refer to the vertical component of Mv.
[0897] For a specific coordinate Post, Pos.x can refer to the x component of Pos.
[0898] For a specific coordinate Pos, Pos.y can refer to the y component of Pos.
[0899] X can be 0 or 1.
[0900] half_pixel can be 0, 1, 2, 4, 8, 16, 32, or a positive integer.
[0901] For example, half_pixel can be a value corresponding to 1 / 2-pel.
[0902] For example, half_pixel can be determined based on the minimum unit representing a motion vector.
[0903] For example, when the minimum unit representing a motion vector is 1 / 16-pel, half_pixel can be 8.
[0904] Figure 20 Inter-frame prediction including OOP pixels according to an example is shown.
[0905] In Figure 20 Inter-frame prediction including OOP pixels or an OOP region is shown.
[0906] In Figure 20 Pixels in the region indicated by the shaded part can be OOP pixels.
[0907] In Figure 20 The region indicated by the shaded part can be an OOP region.
[0908] In an embodiment, at least one of the following prediction methods can be used to predict OOP pixels (or an OOP region) of a target block and / or a reference block (or a prediction block) in the LX direction:
[0909] First OOP prediction method: When performing prediction for a target block, filling can be performed on a region outside the image. Pixel values in the filled region outside the image can be used as prediction values for OOP pixels (or an OOP region).
[0910] For example, when performing filling, filling in the horizontal direction is first performed, and then filling in the vertical direction can be performed.
[0911] For example, when performing filling, filling in the vertical direction is first performed, and then filling in the horizontal direction can be performed.
[0912] Second OOP prediction method: For each OOP pixel in a reference block (or a prediction block) in the LX direction, the value of the pixel closest to the OOP pixel among the pixels in the LX direction reference block (or prediction block) excluding the OOP pixel can be used as the prediction value.
[0913] Third OOP prediction method: For each OOP pixel in a reference block (or a prediction block) in the LX direction, the value of the pixel corresponding to the OOP pixel in the reference block (or prediction block) in the L(1 - X) direction can be used as the prediction value.
[0914] X can be 0 or 1.
[0915] For each motion information storage unit, at least one of determination, change, and storage of motion information can be performed.
[0916] For example, a motion information storage unit can refer to the smallest unit of a group of pixels having the same motion information. That is, pixels in a specific motion information storage unit in an image and / or a block can have the same motion information.
[0917] For example, a motion information storage unit can refer to a block having a size corresponding to MISAVE_SIZE × MISAVE_SIZE.
[0918] For example, MISAVE_SIZE can be 4, 8, or a positive integer.
[0919] For example, in a specific motion information storage unit in a target block, at least one of 1) determination of motion information, 2) change of motion information, and 3) storage of motion information can be performed based on the number of OOP pixels in the specific motion information storage unit.
[0920] For example, motion information can include at least one of an inter-frame prediction indicator, a weight index for inter-frame bi-directional prediction using weights, and a weight for inter-frame bi-directional prediction using weights.
[0921] For example, bi-directional inter-frame prediction using motion information MI0 in the L0 direction and motion information MI1 in the L1 direction can be performed for a target block.
[0922] For example, when none of the pixels belonging to a specific motion information storage unit in a target block are OOP pixels, the motion information of the specific motion information storage unit can be determined as bi-directional inter-frame prediction motion information using motion information MI0 in the L0 direction and motion information MI1 in the L1 direction.
[0923] In an embodiment, determination of motion information can refer to change and / or storage of motion information.
[0924] For example, when there are one or more OOP pixels among the pixels belonging to a specific motion information storage unit in a target block, and all pixels in the reference block in the L0 direction and all pixels in the reference block in the L1 direction corresponding to all OOP pixels in the specific motion information storage unit are OOP pixels, the motion information of the specific motion information storage unit can be determined as bi-directional inter-frame prediction motion information using motion information MI0 in the L0 direction and motion information MI1 in the L1 direction.
[0925] For example, when there are one or more OOP pixels among the pixels belonging to a specific motion information storage unit in a target block, and all the pixels in the reference block in the L0 direction and all the pixels in the reference block in the L1 direction corresponding to all the OOP pixels in the specific motion information storage unit are pixels in an area outside the reference image, the motion information of the specific motion information storage unit can be determined as bidirectional inter-frame prediction motion information using the motion information MI0 in the L0 direction and the motion information MI1 in the L1 direction.
[0926] For example, when there are one or more OOP pixels among the pixels belonging to a specific motion information storage unit in a target block, and among the pixels in the reference block in the L0 direction and the pixels in the reference block in the L1 direction corresponding to all the OOP pixels in the specific motion information storage unit, only the pixels in the reference block in the LX direction are OOP pixels, the motion information of the specific motion information storage unit can be determined as unidirectional inter-frame prediction motion information using the motion information MI(1-X) in the L(1-X) direction. Otherwise, the motion information of the specific motion information storage unit can be determined as bidirectional inter-frame prediction motion information using the motion information MI0 in the L0 direction and the motion information MI1 in the L1 direction.
[0927] For example, when there are one or more OOP pixels among the pixels belonging to a specific motion information storage unit in a target block, and among the pixels in the reference block in the L0 direction and the pixels in the reference block in the L1 direction corresponding to all the OOP pixels in the specific motion information storage unit, only the pixels in the reference block in the LX direction are pixels in an area outside the reference image, the motion information of the specific motion information storage unit can be determined as unidirectional inter-frame prediction motion information using the motion information MI(1-X) in the L(1-X) direction. Otherwise, the motion information of the specific motion information storage unit can be determined as bidirectional inter-frame prediction motion information using the motion information MI0 in the L0 direction and the motion information MI1 in the L1 direction.
[0928] For example, when there are OOP_NUM_THRES or more OOP pixels among the pixels belonging to a specific motion information storage unit in a target block, and among the pixels in the reference block in the L0 direction and the pixels in the reference block in the L1 direction corresponding to all the OOP pixels in the specific motion information storage unit, only the pixels in the reference block in the LX direction are OOP pixels, the motion information of the specific motion information storage unit can be determined as unidirectional inter-frame prediction motion information using the motion information MI(1-X) in the L(1-X) direction. Otherwise, the motion information of the specific motion information storage unit can be determined as bidirectional inter-frame prediction motion information using the motion information MI0 in the L0 direction and the motion information MI1 in the L1 direction.
[0929] For example, when there are OOP_NUM_THRES or more OOP pixels among the pixels belonging to a specific motion information storage unit in a target block, and among the pixels in the reference block in the L0 direction and the pixels in the reference block in the L1 direction corresponding to all the OOP pixels in the specific motion information storage unit, only the pixels in the reference block in the LX direction are pixels in an area outside the reference image, the motion information of the specific motion information storage unit can be determined as unidirectional inter-frame prediction motion information using the motion information MI(1-X) in the L(1-X) direction. Otherwise, the motion information of the specific motion information storage unit can be determined as bidirectional inter-frame prediction motion information using the motion information MI0 in the L0 direction and the motion information MI1 in the L1 direction.
[0930] For example, X can be 0 or 1.
[0931] For example, OOP_NUM_THRES can be 0, 1, 2, 4, 8, or a positive integer.
[0932] For example, OOP_NUM_THRES can be determined based on MISAVE_SIZE.
[0933] OOP_NUM_THRES can be a result obtained by dividing MISAVE_SIZE or MISAVE_SIZE × MISAVE_SIZE by a predetermined value. For example, the predetermined value can be 1, 2, 4, 8, 16, or a positive integer.
[0934] For example, when performing unidirectional inter-frame prediction in the LX direction for a target block and the target block includes one or more OOP pixels (or OOP regions), prediction for the OOP pixels (or OOP regions) in the target block can be performed according to the first OOP prediction method or the second OOP prediction method.
[0935] X can be 0 or 1.
[0936] For example, the case where a specific pixel is an OOP pixel only in the LX direction when performing bidirectional inter-frame prediction for a target block can mean that the pixel corresponding to the specific pixel in the reference block (or prediction block) in the LX direction is an OOP pixel, and the pixel corresponding to the specific pixel in the reference block (or prediction block) in the L(1-X) direction is not an OOP pixel.
[0937] For example, the case where a specific region is an OOP region only in the LX direction when performing bidirectional inter-frame prediction for a target block can mean that the region corresponding to the specific region in the reference block (or prediction block) in the LX direction is an OOP region, and the region corresponding to the specific region in the reference block (or prediction block) in the L(1-X) direction is not an OOP region.
[0938] For example, the case where a specific pixel is an OOP pixel only in the LX direction when performing bidirectional inter-frame prediction on a target block may mean that: the pixel corresponding to the specific pixel in the reference block (or prediction block) in the LX direction is a pixel in an area outside the reference image in the LX direction, and the pixel corresponding to the specific pixel in the reference block (or prediction block) in the L(1-X) direction is a pixel in an area inside the reference image in the L(1-X) direction.
[0939] For example, the case where a specific area is an OOP area only in the LX direction when performing bidirectional inter-frame prediction on a target block may mean that: the area corresponding to the specific area in the reference block (or prediction block) in the LX direction is an OOP area, and the area corresponding to the specific area in the reference block (or prediction block) in the L(1-X) direction is not an OOP area.
[0940] For example, the case where a specific area is an OOP area only in the LX direction when performing bidirectional inter-frame prediction on a target block may mean that: the area corresponding to the specific area in the reference block (or prediction block) in the LX direction is an area outside the reference image in the LX direction, and the area corresponding to the specific area in the reference block (or prediction block) in the L(1-X) direction is an area inside the reference image in the L(1-X) direction.
[0941] For example, in the case where a pixel at a specific coordinate POS in a target block is an OOP pixel only in the LX direction when performing bidirectional inter-frame prediction on the target block, at least one of a first OOP prediction method, a second OOP prediction method, and a third OOP prediction method may be used to determine the value of the OOP pixel corresponding to the position POS in the reference block (or prediction block) in the LX direction.
[0942] X can be 0 or 1.
[0943] For example, it can be considered that for pixels and / or areas in a target block that are OOP pixels only in the LX direction, unidirectional inter-frame prediction in the L(1-X) direction is performed.
[0944] For example, the inter-frame prediction indicator for those (or that) pixels and / or areas can be determined as the inter-frame prediction indicator corresponding to unidirectional inter-frame prediction in the L(1-X) direction.
[0945] For example, the inter-frame prediction indicator for those (or that) pixels and / or areas can be changed to the inter-frame prediction indicator corresponding to unidirectional inter-frame prediction in the L(1-X) direction.
[0946] For example, the inter-frame prediction indicator for those (or that) pixels and / or regions may be stored as an inter-frame prediction indicator corresponding to unidirectional inter-frame prediction in the L(1-X) direction.
[0947] For example, the weights for weighted inter-frame bi-prediction for those (or that) pixels and / or regions may be determined to be the same for the L0 direction and the L1 direction.
[0948] For example, the weights for weighted inter-frame bi-prediction for those (or that) pixels and / or regions may be changed to be the same for the L0 direction and the L1 direction.
[0949] For example, the weights for weighted inter-frame bi-prediction for those (or that) pixels and / or regions may be stored as being the same for the L0 direction and the L1 direction.
[0950] For example, the index for weighted inter-frame bi-prediction for those (or that) pixels and / or regions may be determined to be BCW_DEFAULT.
[0951] For example, the index for weighted inter-frame bi-prediction for those (or that) pixels and / or regions may be changed to BCW_DEFAULT.
[0952] For example, the index for weighted inter-frame bi-prediction for those (or that) pixels and / or regions may be stored as BCW_DEFAULT.
[0953] For example, the inter-frame prediction indicator for at least one of the motion information storage units may be determined to be an inter-frame prediction indicator indicating unidirectional inter-frame prediction in the L(1-X) direction, where the motion information storage unit includes at least one of the pixels that are only OOP pixels in the LX direction and / or the regions that are only OOP regions in the LX direction in the target block.
[0954] For example, the inter-frame prediction indicator for at least one of the motion information storage units may be changed to an inter-frame prediction indicator indicating unidirectional inter-frame prediction in the L(1-X) direction, where the motion information storage unit includes at least one of the pixels that are only OOP pixels in the LX direction and / or the regions that are only OOP regions in the LX direction in the target block.
[0955] For example, the inter-frame prediction indicator for at least one of the motion information storage units may be stored as an inter-frame prediction indicator corresponding to unidirectional inter-frame prediction in the L(1-X) direction, where the motion information storage unit includes at least one of the pixels that are only OOP pixels in the LX direction and / or the regions that are only OOP regions in the LX direction in the target block.
[0956] For example, the weights for weighted inter - prediction for at least one of the motion information storage units may be determined to be the same for the L0 direction and the L1 direction, where the motion information storage unit includes at least one of pixels that are OOP pixels only in the LX direction and / or regions that are OOP regions only in the LX direction in a target block.
[0957] For example, the weights for weighted inter - prediction for at least one of the motion information storage units may be changed to be the same for the L0 direction and the L1 direction, where the motion information storage unit includes at least one of pixels that are OOP pixels only in the LX direction and / or regions that are OOP regions only in the LX direction in a target block.
[0958] For example, the weights for weighted inter - prediction for at least one of the motion information storage units may be stored to be the same for the L0 direction and the L1 direction, where the motion information storage unit includes at least one of pixels that are OOP pixels only in the LX direction and / or regions that are OOP regions only in the LX direction in a target block.
[0959] For example, the index for weighted inter - prediction for at least one of the motion information storage units may be determined to be BCW_DEFAULT, where the motion information storage unit includes at least one of pixels that are OOP pixels only in the LX direction and / or regions that are OOP regions only in the LX direction in a target block.
[0960] For example, the index for weighted inter - prediction for at least one of the motion information storage units may be changed to BCW_DEFAULT, where the motion information storage unit includes at least one of pixels that are OOP pixels only in the LX direction and / or regions that are OOP regions only in the LX direction in a target block.
[0961] For example, the index for weighted inter - prediction for at least one of the motion information storage units may be stored to be BCW_DEFAULT, where the motion information storage unit includes at least one of pixels that are OOP pixels only in the LX direction and / or regions that are OOP regions only in the LX direction in a target block.
[0962] When generating a prediction block of a target block using a reference block in the L0 direction and a reference block in the L1 direction, weighted inter - prediction may determine a combination of weights of the reference blocks in units of coded blocks.
[0963] For example, the weights of each reference block may be determined by signaling / encoding / decoding an index of a predefined table including weights.
[0964] An index for inter - frame dual prediction with a utilization weight equal to BCW_DEFAULT may mean that the L0 - direction weight and the L1 - direction weight in the inter - frame dual prediction with a utilization weight for the target block are the same as each other. Optionally, an index for inter - frame dual prediction with a utilization weight equal to BCW_DEFAULT may mean that no inter - frame dual prediction with a utilization weight is performed on the target block.
[0965] For example, in a case where pixels corresponding to a position POS in a reference block (or prediction block) in the L0 direction and pixels corresponding to the position POS in a reference block (or prediction block) in the L1 direction are both OOP pixels when performing bi - directional inter - frame prediction on a target block, at least one of the first OOP prediction method and the second OOP prediction method can be used to determine the value of the pixels corresponding to the position POS in the reference block (or prediction block) in the LX direction. The position POS may be the position of a specific coordinate POS in the target block.
[0966] For example, in a case where pixels corresponding to a position POS in a reference block (or prediction block) in the L0 direction and pixels corresponding to the position POS in a reference block (or prediction block) in the L1 direction are both pixels in an area outside the reference image when performing bi - directional inter - frame prediction on a target block, at least one of the first OOP prediction method and the second OOP prediction method can be used to determine the value of the pixels corresponding to the position POS in the reference block (or prediction block) in the LX direction. The position POS may be the position of a specific coordinate POS in the target block.
[0967] For example, when there is one or more OOP boundaries in a target block, pixel value correction can be performed on pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to at least one OOP boundary in the target block.
[0968] In an embodiment, each OOP boundary may refer to a boundary between an OOP region and a non - OOP region.
[0969] In an embodiment, an OOP boundary may refer to a boundary between: 1) a region that is an OOP region only in the L0 direction; 2) a region that is an OOP region only in the L1 direction; 3) a region that is an OOP region in both the L0 direction and the L1 direction; and 4) a region that is a non - OOP region.
[0970] The fact that there is an OOP boundary in a target block may mean that there are at least two of the following in the target block: 1) a region that is an OOP region only in the L0 direction; 2) a region that is an OOP region only in the L1 direction; 3) a region that is an OOP region in both the L0 direction and the L1 direction; and 4) a region that is a non - OOP region.
[0971] In other words, the OOP boundary can be a boundary between two different types of regions. The different types of regions can be determined based on the direction in which each of the multiple directions is an OOP region. The multiple directions can include the L0 direction and the L1 direction.
[0972] The types of regions can include: 1) a first type of region that is an OOP region only in the L0 direction, 2) a second type of region that is an OOP region only in the L1 direction, 3) a third type of region that is an OOP region in both the L0 direction and the L1 direction, and 4) a fourth type of region that is a non-OOP region.
[0973] The "boundary between the OOP region and the non-OOP region" described in the embodiments can be interpreted as the "boundary between regions of different types".
[0974] For example, when there is one or more OOP regions and one or more non-OOP regions in the target block, there may be one or more OOP boundaries in the target block.
[0975] For example, for a specific OOP boundary existing in the target block, correction can be performed on the pixel values of the pixels among the pixels in the OOP region of the target block that belong to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the specific OOP boundary.
[0976] Optionally, for example, for a specific OOP boundary existing in the target block, correction can be performed on the pixel values of the pixels among the pixels in the OOP region and the non-OOP region of the target block that belong to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the specific OOP boundary.
[0977] Optionally, for example, for a specific OOP boundary existing in the target block, correction can be performed on the pixel values of the pixels among the pixels in the first region of the target block and the pixels in the second region of the target block that belong to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the OOP boundary between the first region and the second region of the target block. Each of the first region and the second region can be one of the following: 1) a region that is an OOP region only in the L0 direction; 2) a region that is an OOP region only in the L1 direction; 3) a region that is an OOP region in both the L0 direction and the L1 direction; and 4) a region that is a non-OOP region. The first region and the second region can be different from each other. The type of the first region and the type of the second region can be different from each other.
[0978] The specific OOP boundary can refer to the OOP boundary between the first region and the second region.
[0979] By correcting the pixel values of the BOUNDARY_REFINE_PIX_NUM columns and / or rows belonging to the vicinity of a specific OOP boundary, blocking artifacts at the specific OOP boundary can be removed (or reduced).
[0980] For example, it can be implicitly determined whether to perform correction on the pixel values of the pixels belonging to the BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary.
[0981] For example, an indicator indicating whether to perform correction on the pixel values of the pixels belonging to the BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary can be signaled / encoded / decoded.
[0982] In an embodiment, the indicator can be signaled / encoded / decoded in at least one of the units described in the embodiment such as Video Parameter Set (VPS), Decoding Parameter Set (DPS), Sequence Parameter Set (SPS), Adaptive Parameter Set (APS), Picture Parameter Set (PPS), picture header, sub-picture header, slice header, parallel block group header, parallel block header, coding tree block, Coding Tree Unit (CTU), Coding Unit (CU), Prediction Unit (PU), Transform Unit (TU), Coding Block (CB), Prediction Block (PB), and Transform Block (TB).
[0983] For example, an indicator indicating whether to perform correction on the pixel values of the pixels belonging to the BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the OOP boundary of a target image can be signaled / encoded / decoded.
[0984] In an embodiment, the indicator can be signaled / encoded / decoded in at least one of the units described in the embodiment such as VPS, DPS, SPS, APS, PPS, picture header, sub-picture header, slice header, parallel block group header, parallel block header, coding tree block, CTU, CU, PU, TU, CB, PB, and TB.
[0985] For example, an indicator indicating whether to perform correction on the pixel values of the pixels belonging to the BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the OOP boundary of a target block can be signaled / encoded / decoded.
[0986] In an embodiment, the indicator can be signaled / encoded / decoded in at least one of the units described in the embodiment such as VPS, DPS, SPS, APS, PPS, picture header, sub-picture header, slice header, parallel block group header, parallel block header, coding tree block, CTU, CU, PU, TU, CB, PB, and TB.
[0987] For example, for a specific OOP boundary present in a target block, correction may be performed on the pixel values of only the pixels belonging to the OOP region among the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the specific OOP boundary. That is, for the pixels belonging to the non-OOP region among the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the specific OOP boundary, correction of the pixel values of the pixels may not be performed.
[0988] For example, when performing correction on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary in a target block, one or more specific processes for the target block may be determined based on one or more coding parameters in the coding parameters of the target block. For example, the coding parameters may include the motion information of the target block, the coding parameters of the target block, the boundary strength at the specific OOP boundary, the size of the target block, etc. The specific processes may include: 1) whether to perform correction on the pixel values at the specific OOP boundary, 2) the coefficients of the filter for correcting the pixel values at the specific OOP boundary, 3) the number of taps of the filter for correcting the pixel values at the specific OOP boundary, 4) the strength of the filter for correcting the pixel values at the specific OOP boundary, 5) the form / shape of the filter for correcting the pixel values at the specific OOP boundary, 6) whether to perform correction on the pixel values in the OOP region adjacent to the specific OOP boundary, and 7) whether to perform correction on the pixel values in the non-OOP region adjacent to the specific OOP boundary, and may include other processes described in the embodiments.
[0989] For example, correction of the pixel values at the OOP boundary may be performed only when bidirectional prediction is performed for the target block.
[0990] For example, when performing correction on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary in a target block, the boundary strength at the specific OOP boundary may be calculated.
[0991] The boundary strength at the specific OOP boundary may be determined based on the information about the target block described in the embodiments. The information about the target block may include: 1) the pixel values and / or motion information of the pixels belonging to TOCALC_BS_PIX_NUM columns and / or rows adjacent to the specific OOP boundary among the pixels in the OOP region of the target block, 2) the pixel values and / or motion information of the pixels belonging to TOCALC_BS_PIX_NUM columns and / or rows adjacent to the specific OOP boundary among the pixels in the non-OOP region of the target block, 3) the motion information of the target block, and 4) the coding parameters of the target block.
[0992] For example, when a specific OOP boundary is a vertical boundary, the boundary strength at the specific OOP boundary can be determined based on the information about the target block as described in the embodiments. The information about the target block may include: 1) the pixel values and / or motion information of the pixels among the pixels in the OOP region of the target block that belong to the TOCALC_BS_PIX_NUM columns adjacent to the specific OOP boundary, 2) the pixel values and / or motion information of the pixels among the pixels in the non-OOP region of the target block that belong to the TOCALC_BS_PIX_NUM columns adjacent to the specific OOP boundary, 3) the motion information of the target block, and 4) the coding parameters of the target block.
[0993] For example, when a specific OOP boundary is a horizontal boundary, the boundary strength at the specific OOP boundary can be determined based on the information about the target block as described in the embodiments. The information about the target block may include: 1) the pixel values and / or motion information of the pixels among the pixels in the OOP region of the target block that belong to the TOCALC_BS_PIX_NUM columns adjacent to the specific OOP boundary, 2) the pixel values and / or motion information of the pixels among the pixels in the non-OOP region of the target block that belong to the TOCALC_BS_PIX_NUM columns adjacent to the specific OOP boundary, 3) the motion information of the target block, and 4) the coding parameters of the target block.
[0994] TOCALC_BS_PIX_NUM can be a positive integer.
[0995] For example, TOCALC_BS_PIX_NUM can be at least one of 1, 2, 4, 8, or 16.
[0996] TOCALC_BS_PIX_NUM can be a predefined value.
[0997] For example, the TOCALC_BS_PIX_NUM values at the OOP boundaries can be the same as each other.
[0998] For example, the TOCALC_BS_PIX_NUM values at the OOP boundaries can be different from each other.
[0999] For example, the TOCALC_BS_PIX_NUM values at each OOP boundary can be determined based on the information about the target block as described in the embodiments. The information about the target block may include the motion information of the target block, the coding parameters of the target block, and the size of the target block.
[1000] The information about the TOCALC_BS_PIX_NUM value can be encoded / decoded / signaled.
[1001] For example, when correcting the pixel values of the pixels included in BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to each OOP boundary in a target block, at least one filter can be used.
[1002] When correcting the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary in a target block, a filter to be applied to the correction of the values can be determined based on the information about the target block described in the embodiments. The information about the target block can include the motion information of the target block, the coding parameters of the target block, the boundary strength at the specific OOP boundary, and the size of the target block. In other words, among different filters, the filter determined based on the information about the target block can be applied to the specific OOP boundary.
[1003] For example, when correcting the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary, one or more filters among a plurality of filters can be applied to the specific OOP boundary based on the information about the target block described in the embodiments. The information about the target block can include the motion information of the target block, the coding parameters of the target block, the boundary strength at the specific OOP boundary, and the size of the target block.
[1004] In an embodiment, the plurality of filters can include a long tap filter, a strong filter, a weak filter, and a Gaussian filter.
[1005] For example, when correcting the pixel values of the pixels included in BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary in a target block, filter information to be used for the correction of the pixel values can be predefined. Optionally, when correcting the pixel values of the pixels included in BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to a specific OOP boundary in a target block, the filter information to be used for the correction of the pixel values can be encoded / decoded / signaled.
[1006] The filter information can be information related to the filter described in the embodiments.
[1007] Although the filter information can include at least one of the coefficients of the corresponding filter, the number of taps of the filter, the strength of the filter, and the form / shape of the filter, the filter information is not limited to the items listed above. The filter information can include information related to the filter described in the embodiments.
[1008] For example, when performing correction on the pixel values of pixels adjacent to a specific OOP boundary in a target block and the specific OOP boundary in the target block is a vertical boundary, correction can be performed on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns adjacent to the specific OOP boundary among the pixels in the OOP region and the pixels in the non-OOP region.
[1009] Optionally, for example, when performing correction on the pixel values of pixels at a specific OOP boundary and the specific OOP boundary is a vertical boundary, correction can be performed on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns adjacent to the specific OOP boundary among the pixels in the OOP region.
[1010] For example, when performing correction on the pixel values of pixels adjacent to a specific OOP boundary in a target block and the specific OOP boundary in the target block is a vertical boundary, correction can be performed on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns adjacent to the specific OOP boundary among the pixels in the first region and the pixels in the second region. Each of the first region and the second region can be one of the following: 1) a region that is an OOP region only in the L0 direction; 2) a region that is an OOP region only in the L1 direction; 3) a region that is an OOP region in both the L0 direction and the L1 direction; and 4) a region that is a non-OOP region. The first region and the second region can be different from each other. The type of the first region and the type of the second region can be different from each other.
[1011] The specific OOP boundary can refer to the OOP boundary between the first region and the second region. For example, when performing correction on the pixel values of pixels adjacent to a specific OOP boundary in a target block and the specific OOP boundary in the target block is a horizontal boundary, correction can be performed on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM rows adjacent to the specific OOP boundary among the pixels in the OOP region and the pixels in the non-OOP region.
[1012] Optionally, for example, when performing correction on the pixel values of pixels at a specific OOP boundary and the specific OOP boundary is a horizontal boundary, correction can be performed on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM rows adjacent to the specific OOP boundary among the pixels in the OOP region.
[1013] For example, when correcting the pixel values of pixels adjacent to a specific OOP boundary in a target block and the specific OOP boundary in the target block is a horizontal boundary, the pixel values of the pixels belonging to the BOUNDARY_REFINE_PIX_NUM rows adjacent to the specific OOP boundary among the pixels in the first region and the pixels in the second region can be corrected. Each of the first region and the second region can be one of the following: 1) a region that is an OOP region only in the L0 direction; 2) a region that is an OOP region only in the L1 direction; 3) a region that is an OOP region in both the L0 direction and the L1 direction; and 4) a region that is a non-OOP region. The first region and the second region can be different from each other. The type of the first region and the type of the second region can be different from each other.
[1014] The specific OOP boundary can refer to the OOP boundary between the first region and the second region. BOUNDARY_REFINE_PIX_NUM can be a positive integer.
[1015] For example, BOUNDARY_REFINE_PIX_NUM can be at least one of 1, 2, 4, 8, or 16.
[1016] BOUNDARY_REFINE_PIX_NUM can be a predefined value.
[1017] For example, the BOUNDARY_REFINE_PIX_NUM values at the OOP boundary can be the same as each other.
[1018] For example, the BOUNDARY_REFINE_PIX_NUM values at the OOP boundary can be different from each other.
[1019] For example, the BOUNDARY_REFINE_PIX_NUM values at each OOP boundary can be determined based on information about the target block. The information about the target block can include boundary strength, motion information of the target block, coding parameters of the target block, and the size of the target block.
[1020] The information about the BOUNDARY_REFINE_PIX_NUM value can be encoded / decoded / signaled.
[1021] Figure 21 An OOP boundary in a target block according to an example is shown.
[1022] Figure 22 A case where the OOP boundary in a target block according to an example is divided into two or more independent OOP boundaries is shown.
[1023] It can be considered that a single OOP boundary is composed of two or more independent OOP boundaries. Different correction methods can be respectively performed on the pixels adjacent to two or more independent OOP boundaries.
[1024] For example, the method for correcting the pixel values of the pixels adjacent to a specific OOP boundary can be classified based on the OOP-related information described above in the embodiments. The OOP-related information can include: 1) the direction of a specific OOP boundary, 2) the boundary strength at a specific OOP boundary, 3) the filter information for correcting the pixel values at a specific OOP boundary, 4) the value of BOUNDARY_REFINE_PIX_NUM for correcting the pixel values at a specific OOP boundary, 5) whether to perform correction on the pixel values of the pixels in the OOP region adjacent to a specific OOP boundary, and 6) whether to perform correction on the pixel values of the pixels in the non-OOP region adjacent to a specific OOP boundary. However, the criteria for classifying the method for correcting the pixel values of the pixels adjacent to a specific OOP boundary are not limited to the OOP-related information described above.
[1025] In Figure 21 and Figure 22 it can be seen that a single OOP boundary in the target block is composed of two or more independent OOP boundaries.
[1026] In Figure 21 and Figure 22 the first OOP boundary can be regarded as the second OOP boundary and the third OOP boundary that are independent of each other.
[1027] For example, when there are two or more OOP boundaries in the target block, the methods for correcting the pixel values of the pixels adjacent to each OOP boundary at the OOP boundary can be the same as each other.
[1028] Optionally, for example, when there are two or more OOP boundaries in the target block, the methods for correcting the pixel values of the pixels adjacent to each OOP boundary at the OOP boundary can be different from each other.
[1029] For example, after performing correction on the pixel values of the pixels adjacent to the OOP boundary in the target block, correction can be performed on all or some of the pixel values of the pixels in the OOP region.
[1030] For example, after performing correction on all or some of the pixel values of the pixels in the OOP region of the target block, correction can be performed on the pixel values of the pixels adjacent to the OOP boundary.
[1031] For example, when there are one or more OOP regions and one or more non-OOP regions in the target block, correction can be performed on all or some of the pixel values of the pixels in each OOP region.
[1032] For example, based on information about a target block, at least one of the following can be determined: whether to perform correction on the pixel values of all pixels in each OOP region, and whether to perform correction on the pixel values of some pixels in each OOP region. The information about the target block may include motion information of the target block, coding parameters of the target block, and the size of the target block.
[1033] For example, correction may be performed on the pixel values of all or some pixels in each OOP region only when bidirectional prediction is performed on the target block.
[1034] For example, it may be implicitly determined whether to perform correction on the pixel values of all or some pixels in a specific OOP region.
[1035] For example, an indicator indicating whether to perform correction on the pixel values of all and / or some pixels in a specific OOP region may be signaled / encoded / decoded.
[1036] In an embodiment, the indicator may be signaled / encoded / decoded in at least one of the units described in the embodiments such as VPS, DPS, SPS, APS, PPS, picture header, sub - picture header, slice header, parallel block group header, parallel block header, block, CTU, CU, PU, TU, CB, PB, and TB.
[1037] For example, an indicator indicating whether to perform correction on the pixel values of all and / or some pixels in the OOP region of a target image may be signaled / encoded / decoded.
[1038] In an embodiment, the indicator may be signaled / encoded / decoded in at least one of the units described in the embodiments such as VPS, DPS, SPS, APS, PPS, picture header, sub - picture header, slice header, parallel block group header, parallel block header, block, CTU, CU, PU, TU, CB, PB, and TB.
[1039] For example, an indicator indicating whether to perform correction on the pixel values of all and / or some pixels in the OOP region of a target block may be signaled / encoded / decoded.
[1040] In an embodiment, the indicator may be signaled / encoded / decoded in at least one of the units described in the embodiments such as VPS, DPS, SPS, APS, PPS, picture header, sub - picture header, slice header, parallel block group header, parallel block header, block, CTU, CU, PU, TU, CB, PB, and TB.
[1041] For example, when there are one or more OOP regions and one or more non - OOP regions in a target block, correction may be performed on the pixel values of some pixels in each OOP region.
[1042] For example, when the OOP boundary between a specific OOP region and a non-OOP region in a target block is a vertical boundary when performing correction on the pixel values of some pixels in each OOP region in the target block, correction can be performed on the pixel values of the pixels belonging to REGION_REFINE_PIX_NUM columns adjacent to the OOP boundary among the pixels in the OOP region and the pixels in the non-OOP region.
[1043] Optionally, when the OOP boundary between a specific OOP region and a non-OOP region in a target block is a vertical boundary when performing correction on the pixel values of some pixels in each OOP region in the target block, correction can be performed on the pixel values of the pixels belonging to REGION_REFINE_PIX_NUM columns adjacent to the OOP boundary among the pixels in the OOP region.
[1044] Optionally, for example, for a specific OOP boundary existing in a target block, when the specific OOP boundary is a vertical boundary, correction can be performed on the pixel values of the pixels belonging to BOUNDARY_REFINE_PIX_NUM columns and / or rows adjacent to the OOP boundary between the first region and the second region of the target block among the pixels in the first region and the second region of the target block. Each of the first region and the second region can be one of the following: 1) a region that is an OOP region only in the L0 direction; 2) a region that is an OOP region only in the L1 direction; 3) a region that is an OOP region in both the L0 direction and the L1 direction; and 4) a region that is a non-OOP region. The first region and the second region can be different from each other. The type of the first region and the type of the second region can be different from each other.
[1045] The specific OOP boundary can refer to the OOP boundary between the first region and the second region.
[1046] REGION_REFINE_PIX_NUM can be a positive integer.
[1047] For example, REGION_REFINE_PIX_NUM can be at least one of 1, 2, 4, 8, or 16.
[1048] REGION_REFINE_PIX_NUM can be a predefined value.
[1049] For example, the REGION_REFINE_PIX_NUM in the target block can be determined based on the OOP-related information described in the embodiments. The OOP-related information may include the boundary strength at a specific OOP boundary in the target block, the motion information of the target block, the coding parameters of the target block, the pixels in the OOP region, the motion information storage unit in the OOP region, and the size of the target block.
[1050] For example, the REGION_REFINE_PIX_NUM values at the OOP boundaries in the target block can be the same as each other.
[1051] For example, the REGION_REFINE_PIX_NUM values at the OOP boundaries in the target block can be different from each other.
[1052] For example, the REGION_REFINE_PIX_NUM at each OOP boundary in the target block can be determined based on the information about the target block described in the embodiments. The information about the target block may include the boundary strength at a specific boundary, the coding parameters of the target block, the motion information of the target block, the pixels in the OOP region forming a specific OOP boundary, the motion information storage unit in the OOP region forming a specific OOP boundary, and the size of the target block.
[1053] The information about the REGION_REFINE_PIX_NUM value can be encoded / decoded / signaled.
[1054] For example, the pixel values of all or some pixels in a specific OOP region of the target block can be added with an offset BETA_FOR_OOP.
[1055] For example, the value of BETA_FOR_OOP can be (AVGVAL_NONOOP - AVGVAL_OOP). Optionally, for example, the value of BETA_FOR_OOP can be determined based on the value of (AVGVAL_NONOOP - AVGVAL_OOP).
[1056] AVGVAL_OOP can represent the average value of all or some pixels belonging to a specific OOP region in the target block.
[1057] Some pixels can refer to the pixels belonging to AVG_PIX_NUM_FOR_OOP rows (or columns) adjacent to the OOP boundary between the specific OOP region and the non-OOP region in the target block among the pixels in the specific OOP region.
[1058] For example, when the OOP boundary between a specific OOP region and a non-OOP region in a target block is a horizontal boundary, some pixels may refer to the pixels among the pixels in the specific OOP region that belong to AVG_PIX_NUM_FOR_OOP rows adjacent to the OOP boundary between the specific OOP region and the non-OOP region in the target block.
[1059] For example, when the OOP boundary between a specific OOP region and a non-OOP region in a target block is a vertical boundary, some pixels may refer to the pixels among the pixels in the specific OOP region that belong to AVG_PIX_NUM_FOR_OOP columns adjacent to the OOP boundary between the specific OOP region and the non-OOP region in the target block.
[1060] AVGVAL_NONOOP may represent the average value of all or some pixels belonging to the non-OOP region in the target block.
[1061] Some pixels may refer to the pixels among the pixels in the non-OOP region that belong to AVG_PIX_NUM_FOR_OOP rows (or columns) adjacent to the OOP boundary between the specific non-OOP region and the specific OOP region.
[1062] For example, when the OOP boundary between the non-OOP region and a specific OOP region is a horizontal boundary, some pixels may refer to the pixels among the pixels in the non-OOP region that belong to AVG_PIX_NUM_FOR_OOP rows adjacent to the OOP boundary.
[1063] For example, when the OOP boundary between the non-OOP region and a specific OOP region is a vertical boundary, some pixels may refer to the pixels among the pixels in the non-OOP region that belong to AVG_PIX_NUM_FOR_OOP columns adjacent to the OOP boundary.
[1064] AVG_PIX_NUM_FOR_OOP can be a positive integer.
[1065] For example, AVG_PIX_NUM_FOR_OOP can be at least one value among 1, 2, 4, 8, 16, and 32.
[1066] For example, AVG_PIX_NUM_FOR_OOP in a specific OOP region in the target block can be determined based on the information about the target block described in the embodiments. The information about the target block can include: 1) the motion information of the target block, 2) the coding parameters of the target block, 3) the size of the target block, 4) the values of the pixels in the specific OOP region, and 5) the number of pixels in the specific OOP region.
[1067] For example, the pixel values of all or some of the pixels in the first region of the target block can be added with the offset BETA_FOR_OOP.
[1068] For example, the first region can be an OOP region. For example, the first region can be one of the following regions: a region that is an OOP region only in the L0 direction, a region that is an OOP region only in the L1 direction, and a region that is an OOP region in both the L0 direction and the L1 direction.
[1069] For example, the value of BETA_FOR_OOP can be (AVGVAL_NONOOP - AVGVAL_OOP). Optionally, for example, the value of BETA_FOR_OOP can be determined based on the value of (AVGVAL_NONOOP - AVGVAL_OOP).
[1070] AVGVAL_OOP can represent the average value of all or some of the pixels belonging to the first region.
[1071] Some pixels can refer to the pixels among the pixels in the first region that belong to AVG_PIX_NUM_FOR_OOP rows (or columns) adjacent to the OOP boundary between the first region and the second region in the target block.
[1072] For example, the second region can be a non-OOP region or an OOP region. For example, the second region can be one of the following regions: a region that is a non-OOP region, a region that is an OOP region only in the L0 direction, a region that is an OOP region only in the L1 direction, and a region that is an OOP region in both the L0 direction and the L1 direction. However, when the second region is an OOP region, the first region and the second region can be different types of OOP regions. That is, when the first region is a region that is an OOP region only in the L0 direction, the second region can be a region that is an OOP region only in the L1 direction, or a region that is an OOP region in both the L0 direction and the L1 direction.
[1073] For example, when the OOP boundary between the first region and the second region is a horizontal boundary, some pixels can refer to the pixels among the pixels in the first region that belong to AVG_PIX_NUM_FOR_OOP rows adjacent to the OOP boundary between the first region and the second region.
[1074] For example, when the OOP boundary between the first region and the second region is a vertical boundary, some pixels can refer to the pixels among the pixels in the first region that belong to AVG_PIX_NUM_FOR_OOP columns adjacent to the OOP boundary between the first region and the second region.
[1075] AVGVAL_NONOOP may represent the average value of all or some of the pixels belonging to the second region in the target block.
[1076] Some pixels may refer to the pixels belonging to AVG_PIX_NUM_FOR_OOP rows (or columns) adjacent to the OOP boundary between the first region and the second region among the pixels in the second region.
[1077] Some pixels may refer to the pixels belonging to AVG_PIX_NUM_FOR_OOP rows (or columns) adjacent to the OOP boundary between the first region and t...
Claims
1. An image decoding method, comprising: determining prediction information of a target block; and performing prediction on the target block using the prediction information, wherein padding is performed on a target image related to the prediction.
2. The image decoding method according to claim 1, wherein, The padding is performed immediately before performing in-loop filtering on the target image, immediately after applying a specific filter among a plurality of filters in the in-loop filter to the target image, immediately after performing in-loop filtering on the target image, or when performing the prediction on a block located at the boundary of the target image.
3. The image decoding method according to claim 1, wherein, Padding is performed on a region outside the target image.
4. The image decoding method according to claim 3, wherein, The region outside the target image is used as a reference block for a block in an additional image, a part of a reference block for a block in an additional image, or an object to which an interpolation filter is applied in motion compensation in inter-frame prediction of a block in an additional image.
5. The image decoding method according to claim 1, wherein, The padding is performed in the following manner: performing padding using a fixed value, performing padding based on pixel values in the target image, or performing padding based on motion information in the target image.
6. The image decoding method according to claim 1, wherein, Pixel value correction is performed on pixels in a padding block of the target image.
7. The image decoding method according to claim 6, wherein, In the correction, boundary strength at the boundary of the padding block is calculated.
8. An image encoding method, comprising: determining prediction information of a target block; and performing prediction on the target block using the prediction information, wherein padding is performed on a target image related to the prediction.
9. The image encoding method according to claim 8, wherein, The padding is performed immediately before performing in-loop filtering on the target image, immediately after applying a specific filter among a plurality of filters in the in-loop filter to the target image, immediately after performing in-loop filtering on the target image, or when performing the prediction on a block located at the boundary of the target image.
10. The image encoding method according to claim 8, wherein, Padding is performed on a region outside the target image.
11. The image encoding method according to claim 10, wherein, The region outside the target image is used as a reference block for a block in an additional image, a part of a reference block for a block in an additional image, or an object to which an interpolation filter is applied in motion compensation in inter-frame prediction of a block in an additional image.
12. The image encoding method according to claim 8, wherein, The padding is performed in the following manner: performing padding using a fixed value, performing padding based on pixel values in the target image, or performing padding based on motion information in the target image.
13. The image encoding method according to claim 8, wherein, Pixel value correction is performed on pixels in a padding block of the target image.
14. The image encoding method according to claim 13, wherein, In the correction, boundary strength at the boundary of the padding block is calculated.
15. A computer-readable storage medium for storing a bitstream for image decoding, wherein the bitstream includes prediction information, prediction for a target block is performed using the prediction information, padding is performed on a target image related to the prediction.
16. The computer-readable storage medium according to claim 15, wherein, The padding is performed immediately before performing in-loop filtering on the target image, immediately after applying a specific filter among a plurality of filters in the in-loop filter to the target image, immediately after performing in-loop filtering on the target image, or when performing the prediction on a block located at the boundary of the target image.
17. The computer-readable storage medium according to claim 15, wherein, Padding is performed on a region outside the target image.
18. The computer-readable storage medium according to claim 17, wherein, The area outside the target image is used as a reference block for a block in the additional image, a part of a reference block for a block in the additional image, or an object to which an interpolation filter is applied in motion compensation in inter prediction of a block in the additional image.
19. The computer-readable storage medium according to claim 15, wherein, The padding is performed in the following manner: padding using a fixed value is performed, padding based on pixel values in the target image is performed, or padding based on motion information in the target image is performed.
20. The computer-readable storage medium according to claim 15, wherein, Correction of the pixel values of the pixels in the padding block of the target image is performed.
Citation Information
Patent Citations
Apparatus and method for measuring elasticity of skin
KR1020220131595A
A ballast water treatment system and a operating method thereof
KR1020230136967A