Method and apparatus for encoding / decoding image and recording medium for storing bitstream

KR103022755B1Active Publication Date: 2026-09-21ELECTRONICS & TELECOMM RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020250047813
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-28
Filing Date
2025-04-14
Publication Date
2026-09-21
Estimated Expiration
2039-12-20

Smart Images

  • Figure 112025041535043-PAT00216_ABST
    Figure 112025041535043-PAT00216_ABST
Patent Text Reader

Abstract

The present specification discloses an image decoding method. The image decoding method of the present invention comprises a step of determining whether a current block is in a bidirectional optical flow mode, a step of calculating gradient information of prediction samples of the current block when the current block is in a bidirectional optical flow mode, and a step of generating a prediction block of the current block using the calculated gradient information. The step of calculating gradient information of prediction samples of the current block may be characterized by calculating gradient information using at least one neighbor sample adjacent to the prediction sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for image encoding / decoding, an apparatus, and a recording medium storing a bitstream. Specifically, the present invention relates to a method and apparatus for encoding / decoding images on a block basis using bidirectional optical flow. Background Technology

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields. As video data becomes higher in resolution and quality, the relative volume of data increases compared to conventional video data; consequently, transmission and storage costs increase when video data is transmitted using existing wired or wireless broadband lines or stored using existing storage media. To address these issues arising from the increase in video data resolution and quality, high-efficiency video encoding and decoding technologies for video with higher resolution and quality are required.

[0003] Various video compression technologies exist, such as inter-frame prediction technology that predicts pixel values ​​in the current picture from previous or subsequent pictures, intra-frame prediction technology that predicts pixel values ​​in the current picture using pixel information within the current picture, transformation and quantization technology for compressing the energy of residual signals, and entropy coding technology that assigns short codes to values ​​with high frequency and long codes to values ​​with low frequency. By utilizing these video compression technologies, video data can be effectively compressed for transmission or storage.

[0004] In conventional image encoding / decoding methods and devices using bidirectional optical flow, BIO (bi-directional optical flow) could only be applied when there were two pieces of motion information, so BIO could not be applied to blocks containing only one piece of motion information.

[0005] In addition, in conventional image encoding / decoding methods and devices using bidirectional light flow, there was a problem where the gradient value was calculated using pixel values ​​outside the target block area, which increased memory bandwidth or computational load. The problem to be solved

[0006] The present invention can provide a method and apparatus for applying BIO by inducing second motion information when a block to be encoded / decoded has only one motion information under conditions where bidirectional prediction is possible.

[0007] In addition, a method / device capable of providing a variable unit size of a subgroup for calculating a BIO offset to reduce complexity, a method / device for calculating a BIO offset in subgroup units, and a method / device capable of encoding / decoding by selecting whether to apply BIO in block units can be provided.

[0008] In addition, when applying BIO, a method / device for calculating gradients and BIO parameters for reducing memory bandwidth can be provided.

[0010] * means of solving the problem

[0011] An image decoding method according to one embodiment of the present invention comprises: a step of determining whether a current block is in a bidirectional optical flow mode; a step of calculating gradient information of prediction samples of the current block when the current block is in a bidirectional optical flow mode; and a step of generating a prediction block of the current block using the calculated gradient information, wherein the step of calculating gradient information of prediction samples of the current block may calculate gradient information using at least one neighbor sample adjacent to the prediction sample.

[0012] In the image decoding method of the present invention, when the neighbor sample is located outside the area of ​​the current block, the value of the neighbor sample may use the sample value of an integer pixel location close to the neighbor sample.

[0013] In the image decoding method of the present invention, the gradient information can be calculated in units of sub-blocks of a pre-defined size.

[0014] In the image decoding method of the present invention, the step of determining whether the current block is in a bidirectional light flow mode may be determined based on the distance between the first reference picture of the current block and the current picture, and the distance between the second reference picture of the current block and the current picture.

[0015] In the image decoding method of the present invention, the step of determining whether the current block is in a bidirectional optical flow mode may determine that the current block is not in a bidirectional optical flow mode if the distance between the first reference picture and the current picture and the distance between the second reference picture and the current picture are not the same.

[0016] In the image decoding method of the present invention, the step of determining whether the current block is in a bidirectional optical flow mode may be determined based on the type of the reference picture of the current block.

[0017] In the image decoding method of the present invention, the step of determining whether the current block is in a bidirectional optical flow mode may determine that the current block is not in a bidirectional optical flow mode if at least one of the type of the first reference picture of the current block and the type of the second reference picture of the current block is not a short-term reference picture.

[0018] In the image decoding method of the present invention, the step of determining whether the current block is in a bidirectional optical flow mode can be determined based on the size of the current block.

[0019] An image encoding method according to one embodiment of the present invention comprises: a step of determining whether a current block is in a bidirectional optical flow mode; a step of calculating gradient information of prediction samples of the current block when the current block is in a bidirectional optical flow mode; and a step of generating a prediction block of the current block using the calculated gradient information, wherein the step of calculating gradient information of prediction samples of the current block may calculate gradient information using at least one neighbor sample adjacent to the prediction sample.

[0020] In the image encoding method of the present invention, when the neighbor sample is located outside the area of ​​the current block, the value of the neighbor sample may use the sample value of an integer pixel position close to the neighbor sample.

[0021] In the image encoding method of the present invention, the gradient information can be calculated in units of sub-blocks of a pre-defined size.

[0022] In the image encoding method of the present invention, the step of determining whether the current block is in a bidirectional light flow mode may be determined based on the distance between the first reference picture of the current block and the current picture and the distance between the second reference picture of the current block and the current picture.

[0023] In the image encoding method of the present invention, the step of determining whether the current block is in a bidirectional light flow mode may determine that the current block is not in a bidirectional light flow mode if the distance between the first reference picture and the current picture and the distance between the second reference picture and the current picture are not the same.

[0024] In the image encoding method of the present invention, the step of determining whether the current block is in a bidirectional light flow mode may be determined based on the type of the reference picture of the current block.

[0025] In the image encoding method of the present invention, the step of determining whether the current block is in a bidirectional optical flow mode may determine that the current block is not in a bidirectional optical flow mode if at least one of the type of the first reference picture of the current block and the type of the second reference picture of the current block is not a short-term reference picture.

[0026] In the image encoding method of the present invention, the step of determining whether the current block is in a bidirectional optical flow mode can be determined based on the size of the current block.

[0027] A computer-readable recording medium according to one embodiment of the present invention stores a bitstream generated by an image encoding method, wherein the image encoding method comprises: a step of determining whether a current block is in a bidirectional optical flow mode; a step of calculating gradient information of prediction samples of the current block when the current block is in a bidirectional optical flow mode; and a step of generating a prediction block of the current block using the calculated gradient information, wherein the step of calculating gradient information of prediction samples of the current block may calculate gradient information using at least one neighbor sample adjacent to the prediction sample. Effects of the invention

[0029] According to the present invention, an image encoding / decoding method and apparatus with improved compression efficiency can be provided.

[0030] According to the present invention, when a block to be encoded / decoded has only one motion information under conditions where bidirectional prediction is possible, a method and apparatus for inducing a second motion information and applying BIO may be provided.

[0031] According to the present invention, a method / device capable of variably providing a unit size of a subgroup for obtaining a BIO offset to reduce complexity, a method / device for calculating a BIO offset in subgroup units, and a method / device capable of selecting whether to apply BIO in block units for encoding / decoding may be provided.

[0032] In addition, according to the present invention, memory bandwidth and computational load can be reduced in slope calculation used in BIO. Brief explanation of the drawing

[0034] FIG. 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied. FIG. 2 is a block diagram showing the configuration according to one embodiment of a decoding device to which the present invention is applied. Figure 3 is a diagram schematically showing the segmentation structure of an image when encoding and decoding an image. Figure 4 is a diagram illustrating an example of an in-screen prediction process. Figure 5 is a diagram illustrating an example of an inter-frame prediction process. Figure 6 is a diagram illustrating the process of transformation and quantization. Figure 7 is a diagram illustrating reference samples available for in-screen prediction. FIG. 8 is a diagram illustrating various embodiments for inducing second motion information based on first motion information. Figure 9 is a diagram illustrating an example of calculating the gradient values ​​of vertical and horizontal components. FIG. 10 is a diagram illustrating various embodiments of subgroups that serve as units for calculating BIO offsets. Figure 11 is a diagram illustrating the weights that can be applied to each BIO correlation parameter value within a subgroup. FIG. 12 is a diagram illustrating an example of weighted summing only the BIO correlation parameter values ​​at specific locations within a subgroup. FIG. 13 is a diagram illustrating an example of calculating BIO correlation parameter values. Figure 14 shows the BIO correlation parameter S when the subgroup size is 4x4. group This is a drawing for explaining an example of calculating. FIG. 15 is a diagram illustrating an embodiment for deriving a motion vector of a color difference component based on a luminance component. Figure 16 is an example diagram illustrating the motion compensation process for color difference components. FIGS. 17 to 20 are drawings for illustrating various embodiments of inducing gradient values ​​in BIO by padding unavailable pixels outside the block boundary into pixels inside the block boundary. FIG. 21 is a flowchart illustrating an image decoding method according to an embodiment of the present invention. Specific details for implementing the invention

[0035] The present invention is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention. Similar reference numerals in the drawings refer to the same or similar functions across various aspects. The shapes and sizes of elements in the drawings may be exaggerated for clearer explanation. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments as examples. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the invention in relation to one embodiment. Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of exemplary embodiments is limited only by the appended claims, together with all equivalents to those claimed therein, provided they are properly described.

[0036] In the present invention, terms such as "first," "second," etc. may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0037] When it is stated that a component of the present invention is “connected” or “connected” to another component, it should be understood that it may be directly connected to or connected to the other component, or that other components may exist in between. On the other hand, when it is stated that a component is “directly connected” or “directly connected” to another component, it should be understood that no other components exist in between.

[0038] The components shown in the embodiments of the present invention are illustrated independently to represent different characteristic functions and do not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for the convenience of explanation; however, at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separated embodiments of each component are included within the scope of the present invention as long as they do not deviate from the essence of the invention.

[0039] The terms used in this invention are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this invention, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, the description in this invention that a specific configuration "comprising" does not exclude configurations other than that configuration, but means that additional configurations may be included within the scope of the practice or technical concept of this invention.

[0040] Some components of the present invention may not be essential components performing an essential function in the present invention, but may be optional components merely for enhancing performance. The present invention may be implemented by including only the components essential for realizing the essence of the present invention, excluding components used merely for performance enhancement, and a structure including only the essential components, excluding optional components used merely for performance enhancement, is also included within the scope of the rights of the present invention.

[0041] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of this specification, such detailed description is omitted; similar reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.

[0042] In the following, "image" may refer to a single picture constituting a video, or it may refer to the video itself. For example, "encoding and / or decoding of an image" may mean "encoding and / or decoding of an image," and may also mean "encoding and / or decoding of one of the images constituting the video."

[0043] In the following, the terms "video" and "video" may be used interchangeably with the same meaning.

[0044] In the following, the target image may be an image to be encoded and / or an image to be decoded. Additionally, the target image may be an input image input to an encoding device and an input image input to a decoding device. Here, the target image may have the same meaning as the current image.

[0045] In the following, the terms "image," "picture," "frame," and "screen" may be used interchangeably with the same meaning.

[0046] In the following, the target block may be an encoding target block that is the target of encoding and / or a decoding target block that is the target of decoding. Additionally, the target block may be a current block that is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably with the same meaning.

[0047] In the following, the terms "block" and "unit" may be used interchangeably with the same meaning. Alternatively, "block" may refer to a specific unit.

[0048] In the following, the terms "region" and "segment" may be used interchangeably.

[0049] In the following, a specific signal may be a signal representing a specific block. For example, the original signal may be a signal representing the target block. The prediction signal may be a signal representing the prediction block. The residual signal may be a signal representing the residual block.

[0050] In the embodiments, each of the specified information, data, flag, index and element, attribute, etc., may have a value. A value "0" of the information, data, flag, index and element, attribute, etc., may represent logical false or a first predefined value. That is to say, the value "0", false, logical false, and the first predefined value may be used interchangeably. A value "1" of the information, data, flag, index and element, attribute, etc., may represent logical true or a second predefined value. That is to say, the value "1", true, logical true, and the second predefined value may be used interchangeably.

[0051] When a variable such as i or j is used to represent a row, column, or index, the value of i may be an integer greater than or equal to 0, or an integer greater than or equal to 1. That is to say, in the embodiments, the row, column, and index, etc. may be counted from 0, or from 1.

[0053] Glossary of Terms

[0054] Encoder: Refers to a device that performs encoding. In other words, it can mean an encoding device.

[0055] Decoder: Refers to a device that performs decoding. In other words, it can mean a decoding device.

[0056] Block: An MxN array of samples. Here, M and N may represent positive integer values, and a block may commonly represent a two-dimensional array of samples. A block may represent a unit. The current block may represent a block to be encoded during encoding, or a block to be decoded during decoding. Additionally, the current block may be at least one of an encoding block, a prediction block, a residual block, or a transformation block.

[0057] Sample: The basic unit that makes up a block. Bit depth (B d From 0 to 2 depending on ) Bd - It can be expressed as a value up to 1. In the present invention, the term "sample" can be used interchangeably with "pixel" or "pixel." That is, "sample," "pixel," and "pixel" can have the same meaning.

[0058] Unit: This may refer to a unit of image encoding and decoding. In image encoding and decoding, a unit may be a region into which a single image is divided. Additionally, when an image is divided into subdivided units for encoding or decoding, a unit may refer to the divided unit. In other words, a single image can be divided into multiple units. In image encoding and decoding, predefined processing may be performed for each unit. A single unit may be further subdivided into sub-units that have a smaller size than the unit. Depending on the function, a unit may refer to a Block, Macroblock, Coding Tree Unit, Coding Tree Block, Coding Unit, Coding Block, Prediction Unit, Prediction Block, Residual Unit, Residual Block, Transform Unit, Transform Block, etc. Additionally, to distinguish it from a block, a unit may refer to a block of luminance (Luma) components, a corresponding block of chroma (Chroma) components, and syntactic elements for each block. A unit may have various sizes and shapes, and in particular, the shape of a unit may include not only squares but also geometric shapes that can be represented in two dimensions, such as rectangles, trapezoids, triangles, and pentagons. Additionally, unit information may include at least one of the following: the type of unit indicating an encoding unit, a prediction unit, a residual unit, a transformation unit, etc., the size of the unit, the depth of the unit, and the encoding and decoding order of the unit.

[0059] Coding Tree Unit: Consists of a single luminance component (Y) coding tree block and two chrominance component (Cb, Cr) coding tree blocks associated with it. It may also refer to the blocks and the syntactic elements for each block. Each coding tree unit may be partitioned using one or more partitioning methods, such as a quad tree, binary tree, or ternary tree, to form sub-units such as a coding unit, a prediction unit, and a transform unit. It may be used as a term to refer to a sample block that serves as a processing unit in the image decoding process, such as the partitioning of an input image. Here, a quad tree may refer to a quaternary tree.

[0060] If the size of the encoding block falls within a predetermined range, it may be possible to split it into a quadtree only. Here, the predetermined range may be defined as at least one of the maximum size and minimum size of the encoding block that can be split into a quadtree only. Information indicating the maximum / minimum size of the encoding block for which quadtree-type splitting is allowed may be signaled via a bitstream, and such information may be signaled in at least one unit among a sequence, picture parameter, tile group, or slice (segment). Alternatively, the maximum / minimum size of the encoding block may be a fixed size pre-set in the encoder / decoder. For example, if the size of the encoding block corresponds to 256x256 to 64x64, it may be possible to split it into a quadtree only. Or, if the size of the encoding block is larger than the maximum size of the conversion block, it may be possible to split it into a quadtree only. In this case, the block being split may be at least one of the encoding block or the conversion block. In such cases, information indicating the splitting of the encoding block (e.g., split_flag) may be a flag indicating whether to split it into a quadtree. If the size of the encoding block falls within a predetermined range, it may be divided only into a binary tree or a triad tree. In this case, the above description regarding the quad tree may be applied equally to the binary tree or the triad tree.

[0061] Coding Tree Block: This term may be used to refer to any one of the Y coding tree block, Cb coding tree block, or Cr coding tree block.

[0062] Neighbor block: This may refer to a block adjacent to the current block. A block adjacent to the current block may refer to a block whose boundary meets the current block or a block located within a certain distance from the current block. A neighbor block may refer to a block adjacent to a vertex of the current block. Here, a block adjacent to a vertex of the current block may be a block vertically adjacent to a neighbor block horizontally adjacent to the current block, or a block horizontally adjacent to a neighbor block vertically adjacent to the current block. A neighbor block may also refer to a restored neighbor block.

[0063] Reconstructed Neighbor Block: This may refer to a neighbor block that has already been encoded or decoded spatially or temporally around the current block. In this case, a reconstructed neighbor block may refer to a reconstructed neighbor unit. A reconstructed spatial neighbor block may be a block within the current picture that has already been reconstructed through encoding and / or decoding. A reconstructed temporal neighbor block may be a reconstructed block or its neighbor block located at a position corresponding to the current block of the current picture within the reference image.

[0064] Unit Depth: This refers to the degree to which a unit is divided. In a tree structure, the topmost node (Root Node) corresponds to the initial, undivided unit. This topmost node can be referred to as the root node. Additionally, the topmost node can have a minimum depth value. In this case, the topmost node can have a depth of Level 0. A node with a depth of Level 1 can represent a unit created as the initial unit is divided once. A node with a depth of Level 2 can represent a unit created as the initial unit is divided twice. A node with a depth of Level n can represent a unit created as the initial unit is divided n times. A Leaf Node can be the lowest node and can be a node that cannot be further divided. The depth of a Leaf Node can be the maximum level. For example, the predefined value for the maximum level can be 3. It can be said that the Root Node has the shallowest depth, and the Leaf Node has the deepest depth. Additionally, when units are represented as a tree structure, the level at which a unit exists can represent the unit depth.

[0065] Bitstream: Can refer to a sequence of bits containing encoded image information.

[0066] Parameter Set: This corresponds to header information within the structure of the bitstream. At least one of the video parameter set, sequence parameter set, picture parameter set, and adaptation parameter set may be included in the parameter set. Additionally, the parameter set may include tile group, slice header, and tile header information. Furthermore, the tile group may refer to a group containing multiple tiles and may have the same meaning as a slice.

[0067] An adaptive parameter set may refer to a set of parameters that can be referenced and shared across different pictures, subpictures, slices, tile groups, tiles, or bricks. Additionally, subpictures, slices, tile groups, tiles, or bricks within a picture may reference different adaptive parameter sets to utilize information within those sets.

[0068] Additionally, within a picture, different adaptation parameter sets can be referenced using the identifiers of different adaptation parameter sets in subpictures, slices, tile groups, tiles, or bricks.

[0069] Additionally, within a slice, tile group, tile, or brick in a subpicture, different adaptation parameter sets can be referenced using the identifiers of different adaptation parameter sets.

[0070] Additionally, an adaptive parameter set can refer to different adaptive parameter sets within a tile or brick using the identifier of a different adaptive parameter set.

[0071] Additionally, within a brick in a tile, different adaptation parameter sets can be referenced using the identifiers of different adaptation parameter sets.

[0072] Information regarding an adaptive parameter set identifier is included in the parameter set or header of the above subpicture, so that an adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the subpicture.

[0073] Information regarding an adaptive parameter set identifier is included in the parameter set or header of the above tile, so that an adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the tile.

[0074] The header of the above brick includes information regarding an adaptive parameter set identifier, so that an adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the brick.

[0075] The above picture can be divided into one or more rows of tiles and one or more columns of tiles.

[0076] The above subpicture may be divided into one or more tile rows and one or more tile columns within the picture. The above subpicture is an area having a rectangular / square shape within the picture and may include one or more CTUs. Additionally, at least one tile / brick / slice may be included within a single subpicture.

[0077] The above tile is an area within the picture that has a rectangular or square shape and may include one or more CTUs. Additionally, the tile may be divided into one or more bricks.

[0078] The above brick may refer to one or more CTU rows within a tile. A tile may be divided into one or more bricks, and each brick may have at least one CTU row. A tile that is not divided into two or more may also refer to a brick.

[0079] The above slice may include one or more tiles within the picture and one or more bricks within the tile.

[0080] Parsing: This refers to determining the value of a syntax element by entropy decoding a bitstream, or it may refer to entropy decoding itself.

[0081] Symbol: May represent at least one of the following: a syntactic element of the unit to be encoded / decoded, a coding parameter, or a value of a transform coefficient. Additionally, the symbol may represent the target of entropy encoding or the result of entropy decoding.

[0082] Prediction Mode: This may be information indicating a mode of encoding / decoding by intra-frame prediction or a mode of encoding / decoding by inter-frame prediction.

[0083] Prediction Unit: This refers to the basic unit used when performing predictions, such as cross-frame prediction, intra-frame prediction, cross-frame reward, intra-frame reward, and motion reward. A single prediction unit may be divided into multiple partitions or multiple sub-prediction units of smaller sizes. Multiple partitions may also serve as basic units for performing prediction or reward. A partition created by the division of a prediction unit may also be a prediction unit.

[0084] Prediction Unit Partition: This can refer to a form in which prediction units are divided.

[0085] Reference Picture List: This may refer to a list containing one or more reference pictures used for cross-frame prediction or motion compensation. The types of reference picture lists may include LC (List Combined), L0 (List 0), L1 (List 1), L2 (List 2), L3 (List 3), etc., and one or more reference picture lists may be used for cross-frame prediction.

[0086] Inter Prediction Indicator: May indicate the inter-frame prediction direction (unidirectional prediction, bidirectional prediction, etc.) of the current block. Alternatively, it may indicate the number of reference images used when generating the prediction blocks for the current block. Alternatively, it may indicate the number of prediction blocks used when performing inter-frame prediction or motion compensation for the current block.

[0087] Prediction list utilization flag: Indicates whether a prediction block is generated using at least one reference image within a specific reference image list. A prediction list utilization flag can be used to derive a prediction indicator between frames, and conversely, a prediction list utilization flag can be used to derive a prediction indicator between frames. For example, if the prediction list utilization flag indicates a first value of 0, it may indicate that a prediction block is not generated using a reference image within the reference image list, and if it indicates a second value of 1, it may indicate that a prediction block can be generated using the reference image list.

[0088] Reference Picture Index: This can refer to an index in a reference picture list that points to a specific reference picture.

[0089] Reference Picture: This may refer to an image referenced by a specific block for inter-frame prediction or motion compensation. Alternatively, the reference picture may be an image containing a reference block referenced by the current block for inter-frame prediction or motion compensation. Hereinafter, the terms "reference picture" and "reference image" may be used interchangeably with the same meaning.

[0090] Motion Vector: This can be a 2D vector used for cross-frame prediction or motion compensation. A motion vector can represent the offset between the block to be encoded / decoded and the reference block. For example, (mvX, mvY) can represent a motion vector. mvX can represent the horizontal component, and mvY can represent the vertical component.

[0091] Search Range: The search range may be a 2-dimensional area where a search for motion vectors takes place during cross-frame prediction. For example, the size of the search range may be MxN. M and N may each be positive integers.

[0092] Motion Vector Candidate: This may refer to a block that serves as a prediction candidate when predicting a motion vector, or the motion vector of that block. Additionally, a motion vector candidate may be included in the motion vector candidate list.

[0093] Motion Vector Candidate List: This can refer to a list composed of one or more motion vector candidates.

[0094] Motion Vector Candidate Index: May refer to an indicator pointing to a motion vector candidate within the motion vector candidate list. May be the index of a Motion Vector Predictor.

[0095] Motion Information: This may refer to information including at least one of motion vectors, reference image indices, cross-frame prediction indicators, as well as prediction list utilization flags, reference image list information, reference images, motion vector candidates, motion vector candidate indices, merge candidates, merge indices, etc.

[0096] Merge Candidate List: Can refer to a list composed of one or more merge candidates.

[0097] Merge Candidate: This may refer to spatial merge candidates, temporal merge candidates, combined merge candidates, combined positive prediction merge candidates, zero merge candidates, etc. A merge candidate may include motion information such as cross-frame prediction indicators, reference image indices for each list, motion vectors, prediction list utilization flags, and cross-frame prediction indicators.

[0098] Merge Index: This may refer to an indicator pointing to a merge candidate within a merge candidate list. Additionally, the merge index may indicate the block that induced the merge candidate among the blocks restored spatially or temporally adjacent to the current block. Furthermore, the merge index may indicate at least one of the movement information possessed by the merge candidate.

[0099] Transform Unit: This may refer to a basic unit for performing residual signal encoding / decoding, such as transform, inverse transform, quantization, inverse quantization, and transform coefficient encoding / decoding. A single transform unit may be divided into multiple sub-transform units of smaller sizes. Here, the transform / inverse transform may include at least one of a first-order transform / inverse transform and a second-order transform / inverse transform.

[0100] Scaling: This can refer to the process of multiplying a factor by a quantized level. Transformation coefficients can be generated as a result of scaling the quantized level. Scaling can also be called dequantization.

[0101] Quantization Parameter: This may refer to a value used to generate a quantized level using a transform factor in quantization. Alternatively, it may refer to a value used to generate a transform factor by scaling the quantized level in inverse quantization. The quantization parameter may be a value mapped to the quantization step size.

[0102] Delta Quantization Parameter: This may refer to the difference between the predicted quantization parameter and the quantization parameter of the unit to be encoded / decoded.

[0103] Scan: This can refer to a method of sorting the order of units, blocks, or coefficients within a matrix. For example, sorting a 2D array into a 1D array is called a scan. Alternatively, sorting a 1D array into a 2D array can also be called a scan or inverse scan.

[0104] Transform Coefficient: This may refer to the coefficient value generated after performing a transformation in the encoder. Alternatively, it may refer to the coefficient value generated after performing at least one of entropy decoding and inverse quantization in the decoder. Quantized levels or quantized transform coefficient levels obtained by applying quantization to the transform coefficient or residual signal may also be included in the meaning of transform coefficient.

[0105] Quantized Level: This may refer to a value generated by performing quantization on transform coefficients or residual signals in an encoder. Alternatively, it may refer to the value subject to inverse quantization before it is performed in a decoder. Similarly, the quantized transform coefficient level resulting from transform and quantization may also be included within the meaning of quantized level.

[0106] Non-zero Transform Coefficient: This may refer to a transform coefficient whose magnitude is not zero, a transform coefficient level whose magnitude is not zero, or a quantized level.

[0107] Quantization Matrix: This refers to a matrix used in the quantization or inverse quantization process to improve the subjective or objective quality of an image. A quantization matrix can also be called a scaling list.

[0108] Quantization Matrix Coefficient: This can refer to each element within the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.

[0109] Default Matrix: This may refer to a predetermined quantization matrix predefined in the encoder and decoder.

[0110] Non-default Matrix: This can refer to a quantization matrix that is not predefined in the encoder and decoder and is signaled by the user.

[0111] Statistic value: A statistical value for at least one variable, encoding parameter, constant, etc., having specific values ​​that can be computed, may be at least one of the average value, weighted average value, weighted sum value, minimum value, maximum value, mode, median value, and interpolation value of said specific values.

[0112] FIG. 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.

[0113] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images. The encoding device (100) may sequentially encode one or more images.

[0114] Referring to FIG. 1, the encoding device (100) may include a motion prediction unit (111), a motion compensation unit (112), an intra prediction unit (120), a switch (115), a subtractor (125), a converter (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse converter (170), an adder (175), a filter unit (180), and a reference picture buffer (190).

[0115] The encoding device (100) can perform encoding on an input image in intra mode and / or inter mode. Additionally, the encoding device (100) can generate a bitstream containing encoded information through encoding of the input image and can output the generated bitstream. The generated bitstream can be stored on a computer-readable recording medium or streamed via a wired / wireless transmission medium. When intra mode is used as the prediction mode, the switch (115) can be switched to intra, and when inter mode is used as the prediction mode, the switch (115) can be switched to inter. Here, intra mode may refer to an intra-frame prediction mode, and inter mode may refer to an inter-frame prediction mode. The encoding device (100) can generate a prediction block for an input block of the input image. Additionally, after the prediction block is generated, the encoding device (100) can encode a residual block using the difference (residual) of the input block and the prediction block. The input image may be referred to as the current image that is the target of the current encoding. The input block may be referred to as the current block or the block to be encoded, which is the target of the current encoding.

[0116] When the prediction mode is an intra mode, the intra prediction unit (120) may use a sample of a block that has already been encoded / decoded around the current block as a reference sample. The intra prediction unit (120) may perform spatial prediction for the current block using the reference sample and generate prediction samples for the input block through spatial prediction. Here, intra prediction may mean intra-frame prediction.

[0117] When the prediction mode is an inter mode, the motion prediction unit (111) can search for the region that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched region. At this time, the search region can be used as the region. The reference image can be stored in the reference picture buffer (190). Here, the reference image can be stored in the reference picture buffer (190) when encoding / decoding of the reference image is processed.

[0118] The motion compensation unit (112) can generate a prediction block for the current block by performing motion compensation using a motion vector. Here, inter-prediction may mean inter-frame prediction or motion compensation.

[0119] The motion prediction unit (111) and motion compensation unit (112) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. To perform inter-frame prediction or motion compensation, based on the encoding unit, it can determine whether the motion prediction and motion compensation method of the prediction unit included in the corresponding encoding unit is a Skip Mode, Merge Mode, Advanced Motion Vector Prediction (AMVP) Mode, or Current Picture Reference Mode, and can perform inter-frame prediction or motion compensation according to each mode.

[0120] The subtractor (125) can generate a residual block using the difference between the input block and the prediction block. The residual block may also be referred to as a residual signal. The residual signal may represent the difference between the original signal and the prediction signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the prediction signal. The residual block may be a residual signal in block units.

[0121] The transformation unit (130) can generate a transform coefficient by performing a transform on the remaining block and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the remaining block. When a transform skip mode is applied, the transformation unit (130) may skip the transform on the remaining block.

[0122] A quantized level can be generated by applying quantization to a conversion coefficient or a residual signal. In the following embodiments, the quantized level may also be referred to as a conversion coefficient.

[0124] * The quantization unit (140) can generate a quantized level by quantizing a conversion coefficient or residual signal according to a quantization parameter and can output the generated quantized level. At this time, the quantization unit (140) can quantize the conversion coefficient using a quantization matrix.

[0125] The entropy encoding unit (150) can generate a bitstream and output a bitstream by performing entropy encoding according to a probability distribution on values ​​calculated by the quantization unit (140) or coding parameter values ​​calculated during the encoding process. The entropy encoding unit (150) can perform entropy encoding on information regarding a sample of an image and information for decoding an image. For example, information for decoding an image may include syntax elements, etc.

[0126] When entropy coding is applied, a small number of bits are allocated to symbols with a high probability of occurrence and a large number of bits are allocated to symbols with a low probability of occurrence, thereby representing the symbols and reducing the size of the bit sequence for the symbols to be encoded. The entropy coding unit (150) may use encoding methods such as exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding) for entropy coding. For example, the entropy coding unit (150) may perform entropy coding using a Variable Length Coding (VLC) table. In addition, the entropy encoding unit (150) may perform arithmetic encoding using the derived binarization method, probability model, and context model after deriving a binarization method of the target symbol and a probability model of the target symbol / bin.

[0127] The entropy encoding unit (150) can convert a 2-dimensional block form coefficient into a 1-dimensional vector form through a transform coefficient scanning method to encode a transform coefficient level (quantized level).

[0128] Coding parameters may include not only information (flags, indices, etc.) that is encoded in the encoder and signaled to the decoder, such as syntax elements, but also information derived during the encoding or decoding process, and may refer to information required when encoding or decoding images. For example, unit / block size, unit / block depth, unit / block partitioning information, unit / block shape, unit / block partitioning structure, whether to partition in quadtree form, whether to partition in binary tree form, binary tree partitioning direction (horizontal or vertical), binary tree partitioning type (symmetrical or asymmetrical), whether to partition in triad tree form, triad tree partitioning direction (horizontal or vertical), triad tree partitioning type (symmetrical or asymmetrical), whether to partition in complex tree form, complex tree partitioning direction (horizontal or vertical), complex tree partitioning type (symmetrical or asymmetrical), complex tree partitioning tree (binary tree or triad tree), prediction mode (intra-frame prediction or inter-frame prediction), intra-frame luminance prediction mode / direction, intra-frame chrominance prediction mode / direction, intra-frame partitioning information, inter-frame partitioning information, encoded block partitioning flag, predicted block partitioning flag, transform block partitioning flag, reference sample filtering method, reference sample filter tab, reference sample filter coefficients, predicted block filtering method, predicted block filter tab, predicted block Filter coefficients, prediction block boundary filtering method, prediction block boundary filter tab, prediction block boundary filter coefficients, intra-frame prediction mode, inter-frame prediction mode, motion information, motion vector, motion vector difference, reference image index, inter-frame prediction direction, inter-frame prediction indicator, prediction list utilization flag, reference image list, reference image, motion vector prediction index, motion vector prediction candidate, motion vector candidate list, whether to use merge mode, merge index, merge candidate, merge candidate list, whether to use skip mode,Interpolation filter type, Interpolation filter tab, Interpolation filter coefficients, Motion vector magnitude, Motion vector representation accuracy, Transform type, Transform magnitude, Info on whether to use 1st-order transform, Info on whether to use 2nd-order transform, 1st-order transform index, 2nd-order transform index, Info on presence of residual signal, Coded Block Pattern, Coded Block Flag, Quantization parameters, Residual quantization parameters, Quantization matrix, In-frame loop filter application status, In-frame loop filter coefficients, In-frame loop filter tab, In-frame loop filter shape / form, Deblocking filter application status, Deblocking filter coefficients, Deblocking filter tab, Deblocking filter strength, Deblocking filter shape / form, Adaptive sample offset application status, Adaptive sample offset value, Adaptive sample offset category, Adaptive sample offset type, Adaptive loop filter application status, Adaptive loop filter coefficients, Adaptive loop filter tab, Adaptive loop filter shape / form, Binarization / Debinarization method, Context model determination method, Context model update method, Regular mode execution status, Bypass mode execution Status, Context Bin, Bypass Bin, Important Factor Flag, Last Important Factor Flag, Factor Group Unit Encoding Flag, Last Important Factor Position, Flag for whether the factor value is greater than 1, Flag for whether the factor value is greater than 2, Flag for whether the factor value is greater than 3, Remaining Factor Value Information, Sign Information, Recovered Luminance Sample, Recovered Chromaticity Sample, Residual Luminance Sample, Residual Chromaticity Sample, Luminance Conversion Factor, Chromaticity Conversion Factor, Luminance Quantized Level, Chromaticity Quantized Level, Conversion Factor Level Scanning Method, Size of Decoder Side Motion Vector Search Area, Shape of Decoder Side Motion Vector Search Area, Number of Decoder Side Motion Vector Searches, CTU Size Information, Minimum Block Size Information, Maximum Block Size Information, Maximum Block Depth Information, Minimum Block Depth Information, Image Display / Output Order, Slice Identification Information, Slice Type,At least one value or a combined form of slice splitting information, tile group identification information, tile group type, tile group splitting information, tile identification information, tile type, tile splitting information, picture type, input sample bit depth, restored sample bit depth, residual sample bit depth, transform factor bit depth, quantized level bit depth, information about the luminance signal, and information about the chrominance signal may be included in the encoding parameter.

[0129] Here, signaling a flag or index may mean that in an encoder, the corresponding flag or index is entropy encoded and included in a bitstream, and in a decoder, the corresponding flag or index is entropy decoded from the bitstream.

[0130] When the encoding device (100) performs encoding through inter-prediction, the encoded current image can be used as a reference image for another image to be processed later. Accordingly, the encoding device (100) can restore or decode the encoded current image again, and can store the restored or decoded image as a reference image in the reference picture buffer (190).

[0131] The quantized level can be dequantized in the dequantization unit (160) and inverse transformed in the inverse transform unit (170). The dequantized and / or inverse transformed coefficients can be added to the prediction block through the adder (175). A reconstructed block can be generated by adding the dequantized and / or inverse transformed coefficients and the prediction block. Here, the dequantized and / or inverse transformed coefficients refer to coefficients for which at least one of dequantization and inverse transformation has been performed, and may refer to the reconstructed residual block.

[0132] The restoration block may pass through a filter section (180). The filter section (180) may apply at least one of a deblocking filter, a Sample Adaptive Offset (SAO), an Adaptive Loop Filter (ALF), etc., to the restoration sample, restoration block, or restoration image. The filter section (180) may also be referred to as an in-loop filter.

[0133] Deblocking filters can remove block distortion occurring at the boundaries between blocks. To determine whether to perform deblocking, the decision to apply the filter to the current block can be made based on samples contained in a few columns or rows within the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering intensity.

[0134] To compensate for encoding errors using a sample adaptive offset, an appropriate offset value can be added to the sample value. The sample adaptive offset can correct the offset from the original image on a sample-by-sample basis for the deblocked image. One method may be to divide the samples included in the image into a certain number of regions, determine the region to be offset, and apply the offset to that region, or to apply the offset by considering the edge information of each sample.

[0135] An adaptive loop filter can perform filtering based on a comparison of the reconstructed image and the original image. After dividing the samples included in the image into predetermined groups, a filter to be applied to each group can be determined, thereby performing filtering differently for each group. Information regarding whether to apply an adaptive loop filter can be signaled per coding unit (CU), and the shape and filter coefficients of the adaptive loop filter to be applied may vary depending on each block.

[0136] The restored block or restored image that has passed through the filter unit (180) can be stored in the reference picture buffer (190). The restored block that has passed through the filter unit (180) may be part of the reference image. That is to say, the reference image may be a restored image composed of the restored blocks that have passed through the filter unit (180). The stored reference image may subsequently be used for inter-frame prediction or motion compensation.

[0137] FIG. 2 is a block diagram showing the configuration according to one embodiment of a decoding device to which the present invention is applied.

[0138] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.

[0139] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), a motion compensation unit (250), an adder (255), a filter unit (260), and a reference picture buffer (270).

[0140] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium or a bitstream stream streamed through a wired / wireless transmission medium. The decoding device (200) can perform decoding on the bitstream in intra mode or inter mode. Additionally, the decoding device (200) can generate a restored image or a decoded image through decoding and can output the restored image or the decoded image.

[0141] If the prediction mode used for decoding is intra mode, the switch can be switched to intra. If the prediction mode used for decoding is inter mode, the switch can be switched to inter.

[0142] The decoding device (200) can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded may be referred to as the current block.

[0143] The entropy decoding unit (210) can generate symbols by performing entropy decoding according to the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the inverse process of the entropy encoding method described above.

[0144] The entropy decoding unit (210) can convert a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a conversion coefficient scanning method to decode a conversion coefficient level (quantized level).

[0145] The quantized level can be inversely quantized in the inverse quantization unit (220) and inversely transformed in the inverse transformation unit (230). The quantized level can be generated as a restored residual block as a result of inverse quantization and / or inverse transformation being performed. At this time, the inverse quantization unit (220) can apply a quantization matrix to the quantized level.

[0146] When intra mode is used, the intra prediction unit (240) can generate a prediction block by performing a spatial prediction on the current block using sample values ​​of already decoded blocks around the block to be decoded.

[0147] When an inter mode is used, the motion compensation unit (250) can generate a prediction block by performing motion compensation on the current block using a motion vector and a reference image stored in the reference picture buffer (270). The motion compensation unit (250) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. To perform motion compensation, it can determine whether the motion compensation method of the prediction unit included in the corresponding encoding unit is a skip mode, merge mode, AMVP mode, or current picture reference mode based on the encoding unit, and can perform motion compensation according to each mode.

[0148] The adder (255) can generate a restored block by adding the restored residual block and the prediction block. The filter unit (260) can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the restored block or the restored image. The filter unit (260) can output the restored image. The restored block or the restored image can be stored in a reference picture buffer (270) and used for inter-prediction. The restored block that has passed through the filter unit (260) may be part of the reference image. That is to say, the reference image may be a restored image composed of the restored blocks that have passed through the filter unit (260). The stored reference image may subsequently be used for inter-frame prediction or motion compensation.

[0149] FIG. 3 is a diagram schematically illustrating the segmentation structure of an image when encoding and decoding an image. FIG. 3 schematically illustrates an embodiment in which a single unit is divided into a plurality of sub-units.

[0150] To efficiently segment the image, a coding unit (CU) may be used in encoding and decoding. The coding unit may be used as the basic unit of image encoding / decoding. Additionally, the coding unit may be used as a unit to distinguish between intra-frame prediction mode and inter-frame prediction mode during image encoding / decoding. The coding unit may be the basic unit used for the processes of prediction, transform, quantization, inverse transform, inverse quantization, or encoding / decoding of transform coefficients.

[0151] Referring to FIG. 3, the image (300) is sequentially divided into Largest Coding Units (LCUs), and the division structure is determined in LCU units. Here, LCU can be used with the same meaning as Coding Tree Unit (CTU). The division of a unit may refer to the division of a block corresponding to the unit. The block division information may include information regarding the depth of the unit. The depth information may indicate the number and / or degree of division of the unit. A unit may be hierarchically divided into multiple sub-units based on a tree structure and having depth information. That is to say, the unit and the sub-units generated by the division of the unit may correspond to a node and a child node of the node, respectively. Each divided sub-unit may have depth information. The depth information may be information indicating the size of the CU and may be stored for each CU. Since the unit depth indicates the number and / or degree of division of the unit, the division information of the sub-unit may include information regarding the size of the sub-unit.

[0152] The partitioning structure may refer to the distribution of coding units (CUs) within the CTU (310). This distribution may be determined by whether to partition a single CU into multiple CUs (positive integers of 2 or more, including 2, 4, 8, 16, etc.). The width and height of the CUs generated by partitioning may be half the width and height of the CUs before partitioning, respectively, or may have a size smaller than the width and height of the CUs before partitioning, depending on the number of partitions. The CUs may be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the width and height of the partitioned CUs may be reduced compared to at least one of the width and height of the CUs before partitioning. The partitioning of the CUs may be performed recursively up to a predefined depth or a predefined size. For example, the depth of the CTU may be 0, and the depth of the Smallest Coding Unit (SCU) may be a predefined maximum depth. Here, the CTU may be a coding unit having the maximum coding unit size as described above, and the SCU may be a coding unit having the minimum coding unit size. Division begins from the CTU (310), and the depth of the CU increases by 1 each time the horizontal and / or vertical size of the CU is reduced by division. For example, for each depth, the CU that is not divided may have a size of 2Nx2N. Also, for the CU that is divided, the CU of size 2Nx2N may be divided into 4 CUs of size NxN. The size of N may be reduced by half each time the depth increases by 1.

[0153] Additionally, information regarding whether a CU is divided can be expressed through the division information of the CU. The division information may be 1 bit of information. All CUs except the SCU may include division information. For example, if the value of the division information is a first value, the CU may not be divided, and if the value of the division information is a second value, the CU may be divided.

[0154] Referring to FIG. 3, a CTU with a depth of 0 can be a 64x64 block. 0 can be the minimum depth. A SCU with a depth of 3 can be an 8x8 block. 3 can be the maximum depth. CUs of 32x32 blocks and 16x16 blocks can be represented as depth 1 and depth 2, respectively.

[0155] For example, if a single encoding unit is divided into four encoding units, the width and height of the four divided encoding units may each have half the size of the encoding unit before division. For example, if a 32x32 encoding unit is divided into four encoding units, the four divided encoding units may each have a size of 16x16. When a single encoding unit is divided into four encoding units, the encoding unit can be said to have been divided into a quad-tree form (quad-tree partition).

[0156] For example, if a single encoding unit is divided into two encoding units, the width or height of the two divided encoding units may be half the size of the encoding unit before division. For example, if a 32x32 encoding unit is divided vertically into two encoding units, the two divided encoding units may each have a size of 16x32. For example, if an 8x32 encoding unit is divided horizontally into two encoding units, the two divided encoding units may each have a size of 8x16. When a single encoding unit is divided into two encoding units, the encoding unit can be said to have been partitioned into a binary tree form (binary-tree partition).

[0157] For example, when a single encoding unit is divided into three encoding units, the encoding unit can be divided into three encoding units by dividing the horizontal or vertical dimensions of the encoding unit before division in a ratio of 1:2:1. For example, if a 16x32 encoding unit is divided horizontally into three encoding units, the three divided encoding units may have dimensions of 16x8, 16x16, and 16x8, respectively, starting from the top. For example, if a 32x32 encoding unit is divided vertically into three encoding units, the three divided encoding units may have dimensions of 8x32, 16x32, and 8x32, respectively, starting from the left. When a single encoding unit is divided into three encoding units, the encoding unit can be said to have been partitioned in the form of a ternary-tree (ternary-tree partition).

[0158] The CTU (320) of Fig. 3 is an example of a CTU to which quadtree splitting, binary tree splitting and 3-part tree splitting are all applied.

[0159] As described above, to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and triad tree partitioning may be applied. Each partitioning may be applied based on a predetermined priority. For example, quadtree partitioning may be applied preferentially to the CTU. A coding unit that can no longer be quadtree partitioned may correspond to a leaf node of a quadtree. A coding unit corresponding to a leaf node of a quadtree may become a root node of a binary tree and / or a triad tree. That is, a coding unit corresponding to a leaf node of a quadtree may be binary tree partitioned, triad tree partitioned, or not partitioned further. At this time, by ensuring that quadtree partitioning is not performed again on the coding unit created by binary tree partitioning or triad tree partitioning of the coding unit corresponding to a leaf node of a quadtree, the partitioning of the block and / or signaling of partitioning information can be effectively performed.

[0160] The division of a coding unit corresponding to each node of a quadtree can be signaled using quad division information. Quad division information having a first value (e.g., '1') can indicate that the corresponding coding unit is quadtree divided. Quad division information having a second value (e.g., '0') can indicate that the corresponding coding unit is not quadtree divided. Quad division information may be a flag having a predetermined length (e.g., 1 bit).

[0161] There may be no priority between binary tree splitting and triad tree splitting. That is, encoding units corresponding to the leaf nodes of a quadtree can be binary tree split or triad tree split. Additionally, encoding units generated by binary tree splitting or triad tree splitting may be binary tree split or triad tree split again, or may not be split any further.

[0162] A partition in which there is no priority between binary tree partitioning and triad tree partitioning can be referred to as a multi-type tree partition. That is, the encoding unit corresponding to the leaf node of a quadtree can become the root node of a multi-type tree. The partition of the encoding unit corresponding to each node of the multi-type tree can be signaled using at least one of the partition status information, partition direction information, and partition tree information of the multi-type tree. For the partition of the encoding unit corresponding to each node of the multi-type tree, the partition status information, partition direction information, and partition tree information may be signaled sequentially.

[0163] Information on whether a composite tree is split with a first value (e.g., '1') may indicate that the corresponding encoding unit is split into a composite tree. Information on whether a composite tree is split with a second value (e.g., '0') may indicate that the corresponding encoding unit is not split into a composite tree.

[0164] When a encoding unit corresponding to each node of a composite tree is split, the corresponding encoding unit may further include splitting direction information. The splitting direction information may indicate the splitting direction of the composite tree split. Splitting direction information having a first value (e.g., '1') may indicate that the corresponding encoding unit is split in the vertical direction. Splitting direction information having a second value (e.g., '0') may indicate that the corresponding encoding unit is split in the horizontal direction.

[0165] When a encoding unit corresponding to each node of a composite tree is partitioned, the encoding unit may further include partition tree information. The partition tree information may indicate the tree used for the composite tree partition. Partition tree information having a first value (e.g., '1') may indicate that the encoding unit is partitioned into a binary tree. Partition tree information having a second value (e.g., '0') may indicate that the encoding unit is partitioned into a triad tree.

[0166] The splitting information, the splitting tree information, and the splitting direction information may each be flags having a predetermined length (e.g., 1 bit).

[0167] At least one of quad splitting information, information on whether a composite tree is split, splitting direction information, and splitting tree information can be entropy encoded / decoded. For the entropy encoding / decoding of the above information, information from a neighboring encoding unit adjacent to the current encoding unit may be used. For example, the splitting form (segmentation status, splitting tree, and / or splitting direction) of the left encoding unit and / or the upper encoding unit is highly likely to be similar to the splitting form of the current encoding unit. Therefore, context information for the entropy encoding / decoding of the information of the current encoding unit can be derived based on the information of the neighboring encoding unit. At this time, the information of the neighboring encoding unit may include at least one of the quad splitting information, information on whether a composite tree is split, splitting direction information, and splitting tree information of the corresponding encoding unit.

[0168] In another embodiment, among binary tree partitioning and three-part tree partitioning, binary tree partitioning may be performed first. That is, binary tree partitioning is applied first, and a encoding unit corresponding to a leaf node of the binary tree may be set as the root node of the three-part tree. In this case, quadtree partitioning and binary tree partitioning may not be performed for the encoding unit corresponding to a node of the three-part tree.

[0169] A encoding unit that is no longer divided by quadtree splitting, binary tree splitting, and / or ternary tree splitting can be a unit of encoding, prediction, and / or conversion. That is, the encoding unit may no longer be divided for prediction and / or conversion. Therefore, a splitting structure, splitting information, etc., for splitting the encoding unit into a prediction unit and / or conversion unit may not exist in the bitstream.

[0170] However, if the size of the encoding unit serving as the unit of division is larger than the size of the maximum conversion block, the encoding unit may be recursively divided until it becomes equal to or smaller than the size of the maximum conversion block. For example, if the size of the encoding unit is 64x64 and the size of the maximum conversion block is 32x32, the encoding unit may be divided into four 32x32 blocks for conversion. For example, if the size of the encoding unit is 32x64 and the size of the maximum conversion block is 32x32, the encoding unit may be divided into two 32x32 blocks for conversion. In this case, whether the encoding unit is divided for conversion is not separately signaled, but may be determined by comparing the width or height of the encoding unit with the width or height of the maximum conversion block. For example, if the width of the encoding unit is larger than the width of the maximum conversion block, the encoding unit may be divided vertically into two. In addition, if the vertical dimension of the encoding unit is greater than the vertical dimension of the maximum conversion block, the encoding unit can be divided horizontally into two halves.

[0171] Information regarding the maximum and / or minimum size of the encoding unit and information regarding the maximum and / or minimum size of the conversion block may be signaled or determined at an upper level of the encoding unit. The upper level may be, for example, a sequence level, a picture level, a tile level, a tile group level, a slice level, etc. For example, the minimum size of the encoding unit may be determined to be 4x4. For example, the maximum size of the conversion block may be determined to be 64x64. For example, the minimum size of the conversion block may be determined to be 4x4.

[0172] Information regarding the minimum size of an encoding unit corresponding to a leaf node of a quadtree (quadtree minimum size) and / or information regarding the maximum depth from the root node to a leaf node of a composite tree (composite tree maximum depth) may be signaled or determined at an upper level of the encoding unit. The upper level may be, for example, a sequence level, a picture level, a slice level, a tile group level, a tile level, etc. Information regarding the quadtree minimum size and / or information regarding the composite tree maximum depth may be signaled or determined for each of the in-frame slice and the inter-frame slice.

[0173] Difference information regarding the size of the CTU and the maximum size of the transform block may be signaled or determined at an upper level of the encoding unit. The upper level may be, for example, a sequence level, a picture level, a slice level, a tile group level, a tile level, etc. Information regarding the maximum size of the encoding unit corresponding to each node of the binary tree (binary tree maximum size) may be determined based on the size of the encoding tree unit and the difference information. The maximum size of the encoding unit corresponding to each node of the triad tree (triad tree maximum size) may have different values ​​depending on the type of slice. For example, in the case of an in-frame slice, the triad tree maximum size may be 32x32. Also, for example, in the case of an inter-frame slice, the triad tree maximum size may be 128x128. For example, the minimum size of the encoding unit corresponding to each node of the binary tree (binary tree minimum size) and / or the minimum size of the encoding unit corresponding to each node of the triad tree (triad tree minimum size) can be set as the minimum size of the encoding block.

[0174] As another example, the maximum size of a binary tree and / or the maximum size of a triad tree can be signaled or determined at the slice level. Also, the minimum size of a binary tree and / or the minimum size of a triad tree can be signaled or determined at the slice level.

[0175] Based on the size and depth information of the various blocks mentioned above, quad splitting information, information on whether a composite tree is split, splitting tree information and / or splitting direction information, etc., may or may not exist in the bitstream.

[0176] For example, if the size of the encoding unit is not larger than the minimum size of the quadtree, the encoding unit does not include quad splitting information, and the said quad splitting information can be inferred as a second value.

[0177] For example, if the size (width and height) of a encoding unit corresponding to a node of a composite tree is larger than the maximum size (width and height) of a binary tree and / or the maximum size (width and height) of a three-part tree, the encoding unit may not be divided into a binary tree and / or a three-part tree. Accordingly, information on whether the composite tree is divided is not signaled and can be inferred as a second value.

[0178] Alternatively, if the size (width and height) of the encoding unit corresponding to the node of the composite tree is equal to the minimum size (width and height) of the binary tree, or if the size (width and height) of the encoding unit is equal to twice the minimum size (width and height) of the 3-partition tree, the encoding unit may not be divided into a binary tree and / or a 3-partition tree. Accordingly, information regarding whether the composite tree is divided is not signaled and can be inferred as a second value. This is because if the encoding unit is divided into a binary tree and / or a 3-partition tree, an encoding unit smaller than the minimum size of the binary tree and / or the minimum size of the 3-partition tree is generated.

[0179] Alternatively, binary tree splitting or three-part tree splitting may be limited based on the size of a virtual pipeline data unit (hereinafter referred to as the pipeline buffer size). For example, if a encoding unit is divided into sub-coding units that are not suitable for the pipeline buffer size by binary tree splitting or three-part tree splitting, said binary tree splitting or three-part tree splitting may be limited. The pipeline buffer size may be the size of a maximum transform block (e.g., 64X64). For example, when the pipeline buffer size is 64X64, the following splitting may be limited.

[0180] - 3-partition tree partitioning for NxM (N and / or M are 128) encoding units

[0181] - Horizontal binary tree partitioning for 128xN (N <= 64) encoding units

[0182] - Vertical binary tree partitioning for Nx128 (N <= 64) encoding units

[0183] Alternatively, if the depth of the encoding unit within the composite tree corresponding to the node of the composite tree is equal to the maximum depth of the composite tree, the encoding unit may not be divided into a binary tree and / or a three-way tree. Accordingly, information regarding whether the composite tree is divided is not signaled and can be inferred as a second value.

[0184] Alternatively, information on whether the composite tree is divided can be signaled only when at least one of vertical binary tree splitting, horizontal binary tree splitting, vertical three-way splitting, and horizontal three-way splitting is possible for the encoding unit corresponding to the node of the composite tree. Otherwise, the encoding unit may not be divided into a binary tree and / or a three-way splitting. Accordingly, information on whether the composite tree is divided is not signaled and can be inferred as a second value.

[0185] Alternatively, the division direction information may be signaled only when both vertical binary tree division and horizontal binary tree division are possible for the encoding unit corresponding to the node of the composite tree, or when both vertical 3-division tree division and horizontal 3-division tree division are possible. Otherwise, the division direction information is not signaled and may be inferred as a value indicating a direction in which division is possible.

[0186] Alternatively, the split tree information may be signaled only when both vertical binary tree splitting and vertical triad tree splitting are possible for the encoding unit corresponding to the node of the composite tree, or when both horizontal binary tree splitting and horizontal triad tree splitting are possible. Otherwise, the split tree information is not signaled and may be inferred as a value indicating a splittable tree.

[0187] Figure 4 is a diagram illustrating an example of an in-screen prediction process.

[0188] The arrows from the center to the outer edge of Fig. 4 may indicate the prediction directions of the prediction modes within the screen.

[0189] In-frame encoding and / or decoding may be performed using reference samples from neighboring blocks of the current block. Neighboring blocks may be restored neighboring blocks. For example, in-frame encoding and / or decoding may be performed using the values ​​of reference samples or encoding parameters contained in the restored neighboring blocks.

[0190] A prediction block may refer to a block generated as a result of performing a prediction within the screen. A prediction block may correspond to at least one of CU, PU, ​​and TU. The unit of a prediction block may be at least one of the sizes of CU, PU, ​​and TU. A prediction block may be a square-shaped block with sizes such as 2x2, 4x4, 16x16, 32x32, or 64x64, or a rectangular-shaped block with sizes such as 2x8, 4x8, 2x16, 4x16, and 8x16.

[0191] In-frame prediction can be performed according to the in-frame prediction mode for the current block. The number of in-frame prediction modes that the current block may have can be a predefined fixed value or a value determined differently based on the attributes of the prediction block. For example, the attributes of the prediction block may include the size and shape of the prediction block.

[0192] The number of in-screen prediction modes may be fixed at N regardless of the block size. Or, for example, the number of in-screen prediction modes may be 3, 5, 9, 17, 34, 35, 36, 65, or 67. Or, the number of in-screen prediction modes may vary depending on the block size and / or the type of color component. For example, the number of in-screen prediction modes may differ depending on whether the color component is a luminance signal or a chroma signal. For example, as the block size increases, the number of in-screen prediction modes may increase. Or, the number of in-screen prediction modes for a luminance component block may be greater than the number of in-screen prediction modes for a chroma component block.

[0193] The in-frame prediction mode may be a non-directional mode or a directional mode. The non-directional mode may be a DC mode or a Planar mode, and the angular mode may be a prediction mode having a specific direction or angle. The in-frame prediction mode may be represented by at least one of a mode number, a mode value, a mode number, a mode angle, or a mode direction. The number of in-frame prediction modes may be one or more M, including the non-directional and directional modes. A step of checking whether samples included in the restored surrounding blocks can be used as reference samples for the current block to predict the current block in-frame may be performed. If there are samples that cannot be used as reference samples for the current block, the sample value of the sample that cannot be used as a reference sample may be replaced with a value obtained by copying and / or interpolating at least one sample value among the samples included in the restored surrounding blocks, and then used as a reference sample for the current block.

[0194] Figure 7 is a diagram illustrating reference samples available for in-screen prediction.

[0195] As illustrated in FIG. 7, at least one of reference sample lines 0 to 3 may be used for in-frame prediction of the current block. In FIG. 7, samples of segment A and segment F may be padded with the nearest samples of segment B and segment E, respectively, instead of being taken from restored neighboring blocks. Index information indicating the reference sample line to be used for in-frame prediction of the current block may be signaled. If the top boundary of the current block is the boundary of the CTU, only reference sample line 0 may be available. Therefore, in this case, the index information may not be signaled. If a reference sample line other than reference sample line 0 is used, filtering for the prediction block described below may not be performed.

[0196] When making an in-screen prediction, a filter may be applied to at least one of the reference sample or the prediction sample based on at least one of the in-screen prediction mode and the size of the current block.

[0197] In Planner mode, when generating a prediction block for the current block, the sample value of the target sample can be generated using the weighted sum of the top and left reference samples of the current sample and the top-right and bottom-left reference samples of the current block, depending on the position of the target sample within the prediction block. Additionally, in DC mode, when generating a prediction block for the current block, the average value of the top and left reference samples of the current block can be used. Furthermore, in Directional mode, a prediction block can be generated using the top, left, top-right, and / or bottom-left reference samples of the current block. Real-valued interpolation may also be performed to generate the prediction sample value.

[0198] In the case of in-frame prediction between color components, a prediction block for the current block of the second color component can be generated based on the corresponding restoration block of the first color component. For example, the first color component may be a luminance component, and the second color component may be a chrominance component. For in-frame prediction between color components, parameters of a linear model between the first color component and the second color component may be derived based on a template. The template may include upper and / or left peripheral samples of the current block and corresponding upper and / or left peripheral samples of the restoration block of the first color component. For example, the parameters of the linear model may be derived using the sample value of the first color component having the maximum value among the samples in the template and the corresponding sample value of the second color component, and the sample value of the first color component having the minimum value among the samples in the template and the corresponding sample value of the second color component. Once the parameters of the linear model are derived, the corresponding restoration block can be applied to the linear model to generate a prediction block for the current block. Depending on the image format, subsampling may be performed on the peripheral samples of the restoration block of the first color component and the corresponding restoration block. For example, if one sample of the second color component corresponds to four samples of the first color component, one corresponding sample can be calculated by subsampling the four samples of the first color component. In this case, parameter derivation of the linear model and intra-frame prediction between color components can be performed based on the subsampled corresponding sample. Whether to perform intra-frame prediction between color components and / or the range of the template can be signaled as an intra-frame prediction mode.

[0199] The current block can be divided into two or four sub-blocks in the horizontal or vertical direction. The divided sub-blocks can be restored sequentially. That is, an in-frame prediction can be performed on the sub-blocks to generate sub-predicted blocks. Additionally, inverse quantization and / or inverse transformation can be performed on the sub-blocks to generate sub-residual blocks. A restored sub-block can be generated by adding the sub-predicted block to the sub-residual block. The restored sub-block can be used as a reference sample for the in-frame prediction of the lower-ranked sub-blocks. A sub-block may be a block containing a predetermined number (e.g., 16) or more samples. Thus, for example, if the current block is an 8x4 block or a 4x8 block, the current block can be divided into two sub-blocks. Also, if the current block is a 4x4 block, the current block cannot be divided into sub-blocks. If the current block has other sizes, the current block can be divided into four sub-blocks. Information regarding whether the sub-block-based in-screen prediction is performed and / or the division direction (horizontal or vertical) may be signaled. The sub-block-based in-screen prediction may be restricted to be performed only when using reference sample line 0. When the sub-block-based in-screen prediction is performed, filtering for the prediction block described below may not be performed.

[0200] A final prediction block can be generated by performing filtering on the predicted prediction blocks within the screen. The filtering can be performed by applying a predetermined weight to the filtering target sample, the left reference sample, the top reference sample, and / or the top-left reference sample. The weight and / or reference samples (range, position, etc.) used for the filtering can be determined based on at least one of the block size, the prediction mode within the screen, and the position of the filtering target sample within the prediction block. The filtering can be performed only in the case of a predetermined prediction mode within the screen (e.g., DC, planar, vertical, horizontal, diagonal, and / or adjacent diagonal mode). The adjacent diagonal mode may be a mode obtained by adding or subtracting k from the diagonal mode. For example, k may be a positive integer less than or equal to 8.

[0201] The in-frame prediction mode of the current block can be entropy encoded / decoded by predicting it from the in-frame prediction mode of blocks existing in the vicinity of the current block. If the in-frame prediction modes of the current block and the surrounding blocks are identical, information indicating that the in-frame prediction modes of the current block and the surrounding blocks are identical can be signaled using predetermined flag information. Additionally, indicator information regarding the in-frame prediction mode that is identical to the in-frame prediction mode of the current block among multiple in-frame prediction modes of surrounding blocks can be signaled. If the in-frame prediction modes of the current block and the surrounding blocks are different, the in-frame prediction mode information of the current block can be entropy encoded / decoded by performing entropy encoding / decoding based on the in-frame prediction modes of the surrounding blocks.

[0202] Figure 5 is a diagram illustrating an example of an inter-frame prediction process.

[0203] The rectangle shown in Fig. 5 can represent an image. Additionally, the arrow in Fig. 5 can indicate the prediction direction. Each image can be classified into I-picture (Intra Picture), P-picture (Predictive Picture), B-picture (Bi-predictive Picture), etc., depending on the encoding type.

[0204] Picture I can be encoded / decoded through intra-frame prediction without inter-frame prediction. Picture P can be encoded / decoded through inter-frame prediction using only reference images existing in a unidirectional direction (e.g., forward or reverse). Picture B can be encoded / decoded through inter-frame prediction using reference images existing in both directions (e.g., forward and reverse). Additionally, in the case of Picture B, it can be encoded / decoded through inter-frame prediction using reference images existing in both directions, or through inter-frame prediction using reference images existing in either the forward or reverse direction. Here, the bidirectional direction may be the forward and reverse directions. Here, when inter-frame prediction is used, the encoder may perform inter-frame prediction or motion compensation, and the decoder may perform corresponding motion compensation.

[0205] Below, the inter-screen prediction according to the embodiment is described in detail.

[0206] Inter-frame prediction or motion compensation can be performed using reference images and motion information.

[0207] Motion information for the current block can be derived during inter-frame prediction by each of the encoding device (100) and the decoding device (200). Motion information can be derived using motion information of restored surrounding blocks, motion information of a collocated block, and / or a block adjacent to the collocated block. A collocated block may be a block corresponding to the spatial position of the current block within an already restored collocated picture. Here, the collocated picture may be one picture among at least one reference picture included in a reference picture list.

[0208] The method of deriving motion information may vary depending on the prediction mode of the current block. For example, prediction modes applied for inter-frame prediction may include AMVP mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, triangulation mode, inter-intra combined prediction mode, and affine inter mode. Here, the merge mode can be referred to as the motion merge mode.

[0209] For example, when AMVP is applied as a prediction mode, a motion vector candidate list can be generated by determining at least one of the motion vector of a restored surrounding block, the motion vector of a call block, the motion vector of a block adjacent to the call block, and the (0, 0) motion vector as a motion vector candidate. Motion vector candidates can be derived using the generated motion vector candidate list. Motion information of the current block can be determined based on the derived motion vector candidates. Here, the motion vector of the call block or the motion vector of a block adjacent to the call block can be referred to as a temporal motion vector candidate, and the motion vector of a restored surrounding block can be referred to as a spatial motion vector candidate.

[0210] The encoding device (100) can calculate the Motion Vector Difference (MVD) between the motion vector of the current block and a motion vector candidate, and can entropy-encode the MVD. Additionally, the encoding device (100) can generate a bitstream by entropy-encoding a motion vector candidate index. The motion vector candidate index can indicate the optimal motion vector candidate selected from among the motion vector candidates included in the motion vector candidate list. The decoding device (200) entropy-decodes the motion vector candidate index from the bitstream and can select a motion vector candidate for the block to be decoded from among the motion vector candidates included in the motion vector candidate list using the entropy-decoded motion vector candidate index. Additionally, the decoding device (200) can derive the motion vector of the block to be decoded through the sum of the entropy-decoded MVD and the motion vector candidate.

[0211] Meanwhile, the encoding device (100) can entropy-encode the resolution information of the calculated MVD. The decoding device (200) can adjust the resolution of the entropy-decoded MVD using the MVD resolution information.

[0212] Meanwhile, the encoding device (100) can calculate the Motion Vector Difference (MVD) between the motion vector of the current block and motion vector candidates based on an affine model, and can entropy encode the MVD. The decoding device (200) can derive the affine control motion vector of the block to be decoded by deriving the affine control motion vector of the block to be decoded through the sum of the entropy decoded MVD and the affine control motion vector candidates, thereby deriving motion vectors in sub-block units.

[0213] The bitstream may include a reference image index indicating a reference image. The reference image index may be entropy encoded and signaled from the encoding device (100) to the decoding device (200) via the bitstream. The decoding device (200) may generate a prediction block for a block to be decoded based on the induced motion vector and the reference image index information.

[0214] Another example of a method for deriving motion information is merge mode. Merge mode may refer to the merging of motions for multiple blocks. Merge mode may refer to a mode in which motion information of the current block is derived from the motion information of surrounding blocks. When merge mode is applied, a merge candidate list can be generated using the restored motion information of surrounding blocks and / or the motion information of the call block. Motion information may include at least one of 1) a motion vector, 2) a reference image index, and 3) an inter-frame prediction indicator. The prediction indicator may be unidirectional (L0 prediction, L1 prediction) or bidirectional.

[0215] The merge candidate list may represent a list in which motion information is stored. The motion information stored in the merge candidate list may be at least one of motion information of neighboring blocks adjacent to the current block (spatial merge candidate), motion information of a block collocated with the current block in a reference image (temporal merge candidate), new motion information generated by a combination of motion information already existing in the merge candidate list, motion information of a block encoded / decoded prior to the current block (history-based merge candidate), and zero merge candidate.

[0216] The encoding device (100) can generate a bitstream by entropy encoding at least one of a merge flag and a merge index and then signal it to the decoding device (200). The merge flag may be information indicating whether to perform a merge mode on a block-by-block basis, and the merge index may be information regarding which block among the surrounding blocks adjacent to the current block will be merged with. For example, the surrounding blocks of the current block may include at least one of the left adjacent block, the top adjacent block, and the temporally adjacent block of the current block.

[0217] Meanwhile, the encoding device (100) can entropy-encode correction information for correcting the motion vector among the motion information of the merge candidate and signal it to the decoding device (200). The decoding device (200) can correct the motion vector of the merge candidate selected by the merge index based on the correction information. Here, the correction information may include at least one of correction status information, correction direction information, and correction magnitude information. As described above, the prediction mode for correcting the motion vector of the merge candidate based on the signaled correction information can be referred to as a merge mode having a motion vector difference.

[0218] Skip mode may be a mode that applies the motion information of surrounding blocks directly to the current block. When skip mode is used, the encoding device (100) may entropy-encode information regarding which block's motion information to use as the motion information of the current block and signal it to the decoding device (200) via a bitstream. At this time, the encoding device (100) may not signal to the decoding device (200) any syntax elements regarding at least one of motion vector difference information, encoding block flags, and transform coefficient levels (quantized levels).

[0219] The subblock merge mode may refer to a mode that derives motion information at the subblock level of a coding block (CU). When the subblock merge mode is applied, a subblock merge candidate list may be generated using motion information of the subblock collocated to the current subblock in the reference image (subblock-based temporal merge candidate) and / or affine control point motion vector merge candidate.

[0220] The triangle partition mode may refer to a mode in which the current block is divided diagonally to derive movement information for each, each prediction sample is derived using the derived movement information, and each prediction sample is derived by weighted summing the derived prediction samples to derive the prediction sample of the current block.

[0221] The inter-intra combined prediction mode may refer to a mode that derives the prediction sample of the current block by weighting the prediction sample generated by inter-frame prediction and the prediction sample generated by intra-frame prediction.

[0222] The decoding device (200) can self-correct the derived motion information. The decoding device (200) can derive the motion information having the minimum SAD into corrected motion information by searching a predefined area based on the reference block indicated by the derived motion information.

[0223] The decoding device (200) can compensate for prediction samples derived through inter-frame prediction using optical flow.

[0224] Figure 6 is a diagram illustrating the process of transformation and quantization.

[0225] As illustrated in FIG. 6, a quantized level may be generated by performing a transformation and / or quantization process on the residual signal. The residual signal may be generated as the difference between the original block and the prediction block (in-frame prediction block or inter-frame prediction block). Here, the prediction block may be a block generated by in-frame prediction or inter-frame prediction. Here, the transformation may include at least one of a first transformation and a second transformation. Transformation coefficients may be generated by performing a first transformation on the residual signal, and second transformation coefficients may be generated by performing a second transformation on the transformation coefficients.

[0226] The primary transform may be performed using at least one of a plurality of predefined transform methods. For example, the plurality of predefined transform methods may include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a Karhunen-Loeve Transform (KLT)-based transform. A secondary transform may be performed on the transform coefficients generated after the primary transform is performed. The transform method applied during the primary transform and / or secondary transform may be determined based on at least one of the encoding parameters of the current block and / or surrounding blocks. Alternatively, transform information indicating the transform method may be signaled. The DCT-based transform may include, for example, DCT2, DCT-8, etc. The DST-based transform may include, for example, DST-7.

[0228] Quantized levels can be generated by performing quantization on the result of a first transformation and / or a second transformation, or on the residual signal. The quantized levels can be scanned according to at least one of an up-right diagonal scan, a vertical scan, and a horizontal scan based on at least one of an in-frame prediction mode or a block size / shape. For example, the coefficients of a block can be converted into a one-dimensional vector form by scanning them using an up-right diagonal scan. Depending on the size of the transformed block and / or the in-frame prediction mode, a vertical scan that scans the two-dimensional block shape coefficients in the column direction, or a horizontal scan that scans the two-dimensional block shape coefficients in the row direction, may be used instead of an up-right diagonal scan. The scanned quantized levels can be entropy-encoded and included in a bitstream.

[0229] In the decoder, the bitstream can be entropy decoded to generate quantized levels. The quantized levels can be inverse scanned and aligned into a two-dimensional block shape. At this time, at least one of an upper-right diagonal scan, a vertical scan, and a horizontal scan can be performed as a method of inverse scanning.

[0230] Inverse quantization can be performed on the quantized level, and depending on whether a second inverse transform is performed, a second inverse transform can be performed, and depending on whether a first inverse transform is performed on the result of the second inverse transform, a first inverse transform can be performed to generate a restored residual signal.

[0231] Inverse mapping of the dynamic range can be performed on the luminance component restored through intra-frame prediction or inter-frame prediction before in-loop filtering. The dynamic range can be divided into 16 equal pieces, and a mapping function for each piece can be signaled. The mapping function can be signaled at the slice level or the tile group level. An inverse mapping function for performing the inverse mapping can be derived based on the mapping function. In-loop filtering, saving of the reference picture, and motion compensation are performed in the inversely mapped area, and the prediction block generated through inter-frame prediction can be used to generate the restoration block after being converted to the mapped area by mapping using the mapping function. However, since intra-frame prediction is performed in the mapped area, the prediction block generated by intra-frame prediction can be used to generate the restoration block without mapping / inverse mapping.

[0232] If the current block is a residual block of a chrominance component, the residual block can be converted into an inversely mapped area by performing scaling on the chrominance component of the mapped area. The availability of the scaling can be signaled at the slice level or the tile group level. The scaling may be applied only when the mapping for the luminance component is available and the division of the luminance component and the division of the chrominance component follow the same tree structure. The scaling may be performed based on the average of the sample values ​​of the luminance prediction block corresponding to the chrominance block. In this case, if the current block uses cross-frame prediction, the luminance prediction block may refer to the mapped luminance prediction block. By referencing a lookup table using the index of the piece to which the average of the sample values ​​of the luminance prediction block belongs, the value required for the scaling can be derived. Finally, by scaling the residual block using the derived value, the residual block can be converted into an inversely mapped area. Subsequent restoration of color difference component blocks, intra-frame prediction, inter-frame prediction, in-loop filtering, and saving of the reference picture can be performed in the inversely mapped area.

[0233] Information indicating whether mapping / inverse mapping of the above luminance component and color difference component is available can be signaled through a sequence parameter set.

[0234] The predicted block of the current block can be generated based on a block vector representing the displacement between the current block and the reference block within the current picture. In this way, the prediction mode that generates the predicted block by referencing the current picture can be named the Intra Block Copy (IBC) mode. The IBC mode can be applied to an MxN (M <= 64, N <= 64) encoding unit. The IBC mode may include skip mode, merge mode, AMVP mode, etc. In the case of skip mode or merge mode, a merge candidate list is constructed, and a merge index is signaled to identify a single merge candidate. The block vector of the identified merge candidate can be used as the block vector of the current block. The merge candidate list may include at least one of the following: a spatial candidate, a history-based candidate, a candidate based on the average of two candidates, or a zero-merge candidate. In the case of AMVP mode, a difference block vector may be signaled. Additionally, the predicted block vector can be derived from the left neighbor block and the top neighbor block of the current block. An index regarding which neighbor block to use can be signaled. The predicted block in IBC mode may be limited to a block within a previously restored area that is included in the current CTU or the left CTU. For example, the value of the block vector may be restricted so that the predicted block of the current block is located within the three 64x64 block areas that precede the 64x64 block to which the current block belongs in terms of encoding / decoding order. By restricting the value of the block vector in this way, memory consumption and device complexity associated with the implementation of IBC mode can be reduced.

[0236] Hereinafter, specific embodiments according to the present invention will be described with reference to FIGS. 8 to 21.

[0238] Bi-directional optical flow (BIO) can refer to a motion correction technology at the pixel or sub-block level performed based on block-based motion compensation. In other words, BIO can refer to a technology that corrects bidirectional prediction signals at the pixel or sub-block level.

[0239] For example, the pixel value at time t (I t Given ), Equation 1 can be obtained through first-order Taylor expansion.

[0240]

[0242] I t0 Ga I t Assuming that it is located on the motion trajectory and that optical flow is valid along the motion trajectory, the following mathematical equation 2 can be established.

[0243]

[0244] Based on the above mathematical formula 2, the following mathematical formula 3 can be derived from mathematical formula 1.

[0245]

[0247] and If we consider as the movement speed, V x0 and V y0 It can be expressed as such. Therefore, the following mathematical formula 4 can be derived from the above mathematical formula 3.

[0248]

[0250] With a forward reference picture at time t0 and a backward reference picture at time t1, if (t-t0) = (t-t1) = Δt=1, the pixel value at time t can be calculated based on the following mathematical formula 5.

[0251]

[0253] Also, because the movement follows a trajectory, It can be assumed that... Therefore, the following mathematical formula 6 can be derived from the above mathematical formula 5.

[0254]

[0256] In the above mathematical formula, It can be obtained from restored reference pictures.

[0258] In the above mathematical formula 6, can correspond to a general bidirectional prediction. Also, can mean a BIO offset.

[0259] Motion correction vector It can be obtained in the encoder and decoder using the following mathematical formula 7.

[0260]

[0262] In the above mathematical formula 7, r and m can have a value of 0, and the threshold limit can be determined according to the bit depth (Bidepth) of the luminance component.

[0263] In the above mathematical equation 7, s1, s2, s3, s5, and s6 are time pixel value of ( , )class , , , It can be calculated as shown in mathematical formula 8 below using .

[0265]

[0266] In the above mathematical formula 8, s1, s2, s3, s5, and s6 represent BIO correlation parameters at pixel locations (i, j) within the target block area for pixel-unit-based motion correction.

[0267] In the above mathematical formula 8 is the pixel gradient value of the horizontal component at the position (i, j) of the L1 reference image prediction block ( It represents ). is the pixel gradient value of the vertical component at the position (i, j) of the L1 reference image prediction block ( It represents ). is the pixel gradient value of the horizontal component at position (i, j) of the L0 reference image prediction block ( It represents ). is the pixel gradient value of the vertical component at the position (i, j) of the L0 reference image prediction block ( It represents ). and can represent the predicted pixel value at the L1 reference image prediction block (i, j) and the predicted pixel value at the L0 reference image prediction block (i, j). Here, the gradient value may mean the slope value.

[0268] When performing BIO-based motion correction on a sub-block basis, the above mathematical formula 8 can be derived on a sub-block basis as shown in the mathematical formula 9 below.

[0269]

[0270] In the above mathematical formula 9 can be expressed as a BIO parameter in the present invention, and This can be expressed as a BIO correlation parameter in the present invention.

[0271] In the above mathematical formula 9, the range of values ​​for i and j can be determined by the size of the sub-block, the position of the sub-block within the block, and the window size applied to the sub-block.

[0272] For example, if there are 4 4x4 subblocks for an 8x8 block and the window size applied to the subblocks is 6x6, for the first 4x4 subblock at the top-left Is It can have a value, and for the second 4x4 sub-block from the top right Is It can have a value, and for the third 4x4 sub-block from the bottom left Is It can have a value, and for the fourth 4x4 sub-block from the bottom right Is It can have a value.

[0273] In the above mathematical formula 9 and bit depth It can have a positive integer value as a parameter for control.

[0274] for example It can have a value of.

[0275] In calculating sub-block unit BIO correlation parameters, unavailable pixel values ​​and gradient values ​​outside the boundaries of the block area to be decoded may be padded with pixel values ​​and gradient values ​​at block area locations close to the boundaries. That is, unavailable pixel values ​​and gradient values ​​outside the boundaries of the block area to be decoded may be replaced with pixel values ​​and gradient values ​​at block area locations close to the boundaries and used to calculate sub-block unit BIO correlation parameters.

[0276] Using the BIO correlation parameters (s1, s2, s3, s5, s6) obtained at the sub-block level through the above mathematical formula 9 In sub-block units as in mathematical formula 10 can be calculated.

[0277]

[0279] In the above mathematical formula 10 , , limit= Can have.

[0280] Meanwhile, in the present invention, s1, s2, s3, s5, and s6 can be represented as s.

[0281] Calculated in pixel units or sub-block units Based on Equation 6, which uses pixel values ​​and gradient values ​​at each pixel location, the BIO offset value at each pixel location is obtained and then added to the bidirectionally weighted prediction signal to calculate the final prediction signal of the block.

[0282] If two different reference pictures are located temporally before or after the current picture, the predicted signal at time t can be calculated by considering the time distance between the current picture and the reference pictures as in Equation 11.

[0283]

[0285] According to one embodiment of the present invention, if the picture and / or slice to which the current block belongs can be encoded / decoded through inter-frame prediction using reference pictures existing in a bidirectional reference picture list, the current block can be encoded / decoded by applying BIO even if it has only first motion information. In the present invention, the first motion information may refer to motion information in the L0 direction or motion information in the L1 direction of the current block.

[0286] According to the present invention, if the current block has only the first movement information, it can be encoded / decoded by applying BIO after inducing the second movement information.

[0287] In deriving the second motion information, whether to derive the second motion information can be determined based on the first motion vector of the current block.

[0288] For example, whether to induce second motion information may be determined based on the result of comparing the first motion vector value of the current block with a predetermined threshold value. For example, as shown in Equation 12 below, the first of the current block Directional motion vector ( ) and y-direction motion vector ( Depending on the size of ), it is possible to determine whether the second movement information is induced. According to one embodiment, MV x0 Wow, the MV y0 When all are below the threshold, it can be determined that the second movement information is induced.

[0289]

[0291] If it is determined that second motion information is to be derived, for example, if the above conditions are satisfied, the second motion information can be derived based on the first motion information of the current block. Additionally, BIO can be applied to the current block using the first motion information and the derived second motion information.

[0292] The threshold value (Th) for determining whether to induce second motion information may use a predetermined value or be transmitted included in the bitstream. The threshold value may be adaptively determined based on encoding parameters of the current block, such as the size and / or shape of the current block.

[0293] For example, threshold values ​​can be transmitted as sequence parameters, picture parameters, slice headers, and block-level syntax data.

[0294] The second motion information can be derived based on the time distance between the current picture and the reference pictures.

[0295] FIG. 8 is a diagram illustrating various embodiments for inducing second motion information based on first motion information.

[0296] For example, as in FIG. 8(a), the time distance between the current picture (CurPic) and the first reference picture (Ref0) indicated by the first motion information (MV0) of the current block ( Only when a picture that has the same time distance as ) and has a different POC from the first reference picture (Ref0) exists in the second reference picture list can the second motion information (MV1) be derived by making that picture the second reference picture (Ref1).

[0297] When a picture with the same time distance is used as a second reference picture, the second motion vector can be derived from the first motion vector as shown in Equation 13 below.

[0298]

[0300] For example, as shown in FIG. 8(b), the time distance between the current picture (CurPic) and the first reference picture (Ref0) indicated by the first movement information (MV0) of the current block ( If there is no picture in the second reference picture list that has the same time distance as the current picture (CurPic), the picture that has the shortest time distance from the current picture (CurPic) and has a different POC from the first reference picture (Ref0) can be used as the second reference picture (Ref1) to derive the second motion information (MV1).

[0301] In the above example, the second motion vector (MV1) is the time distance between the first motion vector (MV0), the current picture (CurPic), and the first reference picture (Ref0). ), time distance between current picture (CurPic) and second reference picture (Ref1) It can be derived as shown in mathematical formula 14 below using ).

[0302]

[0304] For example, the time distance between the current picture (CurPic) and the first reference picture (Ref0) indicated by the first motion information (MV0) of the current block ( Regardless of ), the second motion information (MV1) can be derived by making the picture with the shortest time distance from the current picture (CurPic) and a different POC from the first reference picture (Ref0) within the second reference picture list the second reference picture (Ref1).

[0305] In the above example, the second motion vector (MV1) is the time distance between the first motion vector (MV0), the current picture (CurPic), and the first reference picture (Ref0). ), time distance between current picture (CurPic) and second reference picture (Ref1) It can be derived as shown in mathematical formula 15 below using ).

[0306]

[0308] In deriving the second motion information (MV1), the first motion information of the current picture (CurPic) and the current block ( A second motion information (MV1) can be derived through motion prediction using reference pictures in a reference picture list in a different direction from ) and reference pictures.

[0309] For example, based on the prediction block (P0) generated from the first movement information of the current block as in Fig. 8 (c), the prediction block ( A block that minimizes the distortion value with ) can be found. As an initial motion vector for motion detection, a (0,0) motion vector indicating the same location corresponding to the current block may be used, and as in FIG. 8 (a) or FIG. 8 (b), the first motion vector ( By using the motion vector derived based on ) as the initial motion vector, it is possible to find a block that minimizes the distortion value within a predetermined search area.

[0310] Prediction block ( A block that minimizes the distortion value with ) ) and prediction block( Represents the distance offset between ). and current block and predicted block( Represents the distance offset between ). Using this, the second movement information of the current block can be derived as shown in Equation 16 below. Also, e.g. The reference picture index of the included reference picture (Ref1) can be used as the reference picture index of the second motion information.

[0311]

[0313] If the second motion information obtained from the above embodiments is available, the first reference picture of the current block ( ) and first movement information and second reference picture( A final prediction signal can be generated by applying pixel-level or sub-block-level BIO to the current block using the second motion information.

[0314] In applying BIO to the current block using the first motion information and the second motion information, the first reference picture indicated by the first motion information ( ) and the second reference picture indicated by the second motion information( ) exists on different time axes with respect to the current picture at time t, and the time distance between the current picture and the first reference picture and the time distance between the current picture and the second reference picture ( If ) are different from each other, the BIO offset can be calculated by considering the time distance between the current picture and the reference pictures.

[0315] For example, if the following conditions are satisfied, the final predicted signal of the current block can be obtained by calculating the BIO offset as in Equation 17 below, taking into account the time distance between the current picture and the reference pictures.

[0316]

[0317] When applying BIO to the current block using the first motion information and the second motion information of the current block, the first reference picture (Ref0) indicated by the first motion information and the second reference picture (Ref1) indicated by the second motion information exist on different time axes with respect to the current picture at time t, and only when the time distance (TD0) between the current picture and the first reference picture and the time distance (TD1) between the current picture and the second reference picture are the same, can the final predicted signal be obtained by applying BIO and adding a BIO offset to the predicted signal of the current block. Here, the first motion information may refer to motion information in the first predicted direction, and the second motion information may refer to motion information in the second predicted direction. Furthermore, the fact that the two reference pictures exist on different time axes may mean that they are located in different directions with respect to the current picture. For example, if the relationship between the POC of the first reference picture (POC_ref0), the POC of the current picture (POC_cur), and the POC of the second picture (POC_ref) satisfies the following condition, it may mean that the two reference pictures exist on different time axes with respect to the current picture, and that the time distance between the current picture and the first reference picture (TD_0) and the time distance between the current picture and the second reference picture (TD_1) are the same.

[0318] Condition: (POC_ref0 - POC_cur) == (POC_cur - POC_ref1)

[0320] When applying BIO to the current block using the first motion information and the second motion information, the first reference picture (Ref0) indicated by the first motion information and the second reference picture (Ref1) indicated by the second motion information exist on different time axes relative to the current picture at time t, and if the motion vector of the first motion information is (0,0) and the motion vector of the second motion information is (0,0), BIO may not be applied to the current block.

[0321] In applying BIO to the current block using the first motion information and the second motion information, the first reference picture (Ref0) indicated by the first motion information and the second reference picture (Ref1) indicated by the second motion information exist on different time axes with respect to the current picture at time t, and the time distance between the current picture and the first reference picture and the time distance between the current picture and the second reference picture ( If ) are identical to each other, and the motion vector of the first motion information is (0,0) and the motion vector of the second motion information is (0,0), then BIO may not be applied to the current block.

[0322] Whether two reference pictures exist on different time axes relative to the current picture can be determined by the difference in the POC between the current picture and the reference picture. For example, if the relationship between the POC of the first reference picture (POC_ref0), the POC of the current picture (POC_cur), and the POC of the second picture (POC_ref) satisfies the following condition, it may mean that the two reference pictures exist on different time axes relative to the current picture.

[0323] Condition: (POC_ref0 - POC_cur) × (POC_cur - POC_ref1) > 0

[0325] The gradient used in the BIO offset calculation process ( , , , ) can be calculated using the first movement information and the second movement information.

[0326] When a motion vector points to a subpixel location within a reference picture, the gradient values ​​of the vertical and horizontal components at that subpixel location can be calculated by applying filters using the values ​​of surrounding integer pixel locations. Tables 1 and 2 below show the filter coefficients of the interpolation filter.

[0327]

[0328]

[0330] When a motion vector points to a subpixel location, the gradient values ​​of the vertical and horizontal components can be calculated by rounding to an integer pixel location close to that subpixel location and using the surrounding integer pixel values. That is, when a motion vector points to a subpixel location, the motion vector is rounded, and the pixel value of the integer pixel location indicated by the rounded motion vector can be used for gradient calculation. In this case, the gradient values ​​of the vertical and horizontal components can be calculated using only the filter coefficients at pixel location 0 in Table 1.

[0331] For example, in the case where the horizontal and vertical motion vector magnitudes are (15, 15) with 1 / 16 motion vector precision, the motion vector (16, 16) is rounded as shown in Equation 18 below, and then the gradient value of the horizontal component can be calculated using the pixel values ​​of integer pixel positions and filter coefficients (8, -39, -3, 46, -17, 5). In the case of 1 / 16 motion vector precision, the shift can be 4, and in the case of 1 / 8 motion vector precision, the shift can be 3.

[0332]

[0334] When a motion vector points to a subpixel location, interpolated pixel values ​​are generated at the corresponding subpixel location, and then the gradient values ​​of the vertical and horizontal components can be calculated using the interpolated pixel values. The gradient can be calculated using a [-1, 0, 1] filter on the interpolated pixel values. The gradient values ​​of the vertical and horizontal components can be calculated at the corresponding location within the prediction block for the motion-compensated prediction signal.

[0335] For example, when the [-1, 0, 1] filter is applied, the gradients of the horizontal and vertical components can be calculated for the motion-compensated prediction signals of the L0 reference picture as shown in Equation 19. can mean motion-compensated predicted signal values ​​at position (i, j).

[0337]

[0340] FIG. 9 is a diagram illustrating an example of generating motion-compensated predicted pixel values ​​at the above-mentioned subpixel locations and then calculating gradient values ​​of vertical and horizontal components using the pixel values.

[0341] The embodiment illustrated in FIG. 9(a) is an example of a case where gradients are calculated at all pixel locations for a 4x4 block, that is, a case where pixel values ​​for an area filled with a pattern are required.

[0342] For example, when applying a [-1, 0, 1] filter to calculate gradient values ​​at each position within a 4x4 block, e.g., the horizontal gradient (G) at the top-left (0, 0) position 0, 0 To calculate ), the pixel value at the (-1, 0) position outside the 4x4 block area is required. Additionally, the horizontal gradient (G) at the top-right (3, 0) position 3, 0To calculate ), a pixel value at the (4, 0) position outside the 4x4 block area is required. When an 8-tab interpolation filter is applied to generate interpolated pixel values ​​for a block of size W(width)xH(height), (W+7)x(H+7) pixel values ​​from the reference picture are required, but in the above method, a total of (W+7+2)x(H+7+2) pixel values, each increased by 2 pixels in the horizontal and vertical directions respectively, are required to calculate the gradient value.

[0343] To reduce memory bandwidth for the reference picture, a bi-linear interpolation filter may be used for pixel values ​​outside the target block area that are additionally required for gradient calculation at block boundary locations.

[0344] To reduce memory bandwidth for the reference picture, pixel values ​​outside the target block area may use pixel values ​​at integer pixel locations close to the left or right of the subpixel location indicated by the motion vector within the reference picture, without a separate interpolation process.

[0345] To reduce memory bandwidth for the reference picture, pixel values ​​outside the target block area may use pixel values ​​at the nearest integer pixel location to the subpixel location indicated by the motion vector within the reference picture without a separate interpolation process. For example, in FIG. 9 (a), pixel values ​​in the pattern-filled area (i.e., pixel values ​​outside the target block area) may use pixel values ​​at the nearest integer pixel location to the subpixel location indicated by the motion vector within the reference picture without a separate interpolation process.

[0346] That is, when a motion vector indicates a subpixel location outside the target block area within the reference picture, the motion vector can be rounded to a nearby integer location without a separate interpolation process, and then the gradient can be calculated using the pixel value of the rounded integer pixel location.

[0347] For example, pixel values ​​outside the target block area can be replaced with pixel values ​​at the locations indicated by (xIntL + (xFracL >> 3) 1) and (yIntL + (yFracL >> 3) 1) within the reference picture, and gradient calculation can be performed. Here, (xIntL, yIntL) may represent integer pixel-unit positions of the motion vector, and (xFracL, yFracL) may represent subpixel-unit positions of the motion vector.

[0349] The above method can be used in the process of calculating gradient values ​​at each position within a 4x4 sub-block required for prediction signal correction in an affine prediction mode that derives affine control motion vectors of target blocks based on an affine model and derives motion vectors in sub-block units.

[0350] That is, when an affine control motion vector of a target block is derived based on an affine model and the motion vector derived at the sub-block level indicates a subpixel location outside the area of ​​the target block (4x4 sub-block) within the reference picture, the motion vector can be rounded to a nearby integer location without a separate interpolation process, and then a gradient calculation can be performed using the pixel value of the rounded integer pixel location.

[0351] For example, pixel values ​​outside the area of ​​the target block (4x4 sub-block) can be replaced with pixel values ​​at the locations indicated by (xIntL + (xFracL >> 3) 1) and (yIntL + (yFracL >> 3) 1) within the reference picture, and gradient calculation can be performed. Here, (xIntL, yIntL) may represent integer pixel-unit positions of the motion vector, and (xFracL, yFracL) may represent subpixel-unit positions of the motion vector.

[0353] To generate pixel values ​​outside the target block area without increasing memory bandwidth for the reference picture, pixel values ​​outside the block area can be generated by padding with pixel values ​​at integer pixel locations close to the encoding location indicated by the motion vector within the reference picture and then using an existing 8-tab interpolation filter.

[0354] To reduce memory bandwidth for the reference picture, gradient values ​​can be calculated only at the inner positions of the 4x4 block area as in Fig. 9 (b) and used to calculate the BIO offset.

[0355] For example, in the case of a 4x4 block, the gradient can be calculated only at the inner positions of the block (1, 1), (2, 1), (1, 2), and (2, 2) and used to calculate the BIO offset. FIG. 9(c) shows an example of calculating the gradient only at the inner positions of an 8x8 block.

[0356] In addition, for locations where the gradient is not calculated, the gradient calculated inside the block can be copied and used. FIG. 9 (d) shows an example of a case where the gradient calculated inside the block is copied and used. For example, as shown in FIG. 9 (d), for locations where the gradient value is not calculated, the gradient value of an adjacent location can be used as the gradient value of that location.

[0358] As another example of reducing reference picture access memory bandwidth, the gradient at each pixel position within the block can be calculated using only available pixel values ​​within the block area as in Equation 20.

[0359]

[0361] In mathematical formula 20, (b) and (c) are formulas for calculating the gradient when pixel values ​​outside the block boundary are not available at the block boundary location, and (a) is a formula for calculating the gradient when surrounding pixels to the left / right or up / down are available within the block area. That is, different gradient calculation formulas can be applied based on the location where the gradient within the block is calculated.

[0362] The embodiment illustrated in FIG. 9(a) calculates the gradient at every pixel position in units of 4x4 sub-blocks for a 4x4 block, for example, the horizontal gradient (G) at the top-left (0, 0) position. 0,0 ) can be calculated using only the pixel values ​​at positions (0, 0) and (1, 0) if the pixel value at position (-1, 0) is not available. Additionally, the horizontal gradient (G) at the top-right position (3, 0) 3, 0 ) can be calculated using only the pixel values ​​at positions (3, 0) and (2, 0) if the pixel value at position (4, 0) is not available. In addition, the vertical gradient (G) at the top-left position (0, 0) 0,0 ) can be calculated using only the pixel values ​​at positions (0, 0) and (0, 1) if the pixel value at position (0, -1) is not available. Also, the vertical gradient (G) at the top-right (3, 0) position 3, 0 ) can be calculated using only the pixel values ​​at positions (3, 0) and (3, 1) if the pixel value at position (3, -1) is not available.

[0363] In the above Equation 20, (b) and (c) can achieve the same effect as padding the pixel values ​​at locations outside the 4x4 block boundary with the pixel values ​​at the nearest block boundary in the embodiment illustrated in FIG. 9 (a), and then calculating the gradient values ​​at all locations within the block using Equation 20(a). For example, the horizontal gradient (G) at the top-left (0, 0) position0, 0 ) can be calculated by Equation 20(a) using the pixel values ​​at positions (-1, 0) and (1, 0) after padding the pixel value at position (0, 0) to (-1, 0) if the pixel value at position (-1, 0) is unavailable. Additionally, the horizontal gradient (G) at the top-right position (3, 0) 3, 0 ) can be calculated by Equation 20(a) using pixel values ​​at positions (4, 0) and (2, 0) after padding the pixel value at position (3, 0) to the (4, 0) position if the pixel value at position (4, 0) is not available. Additionally, the vertical gradient (G) at the top-left (0, 0) position 0,0 ) can be calculated by Equation 20(a) using the pixel values ​​at positions (0, -1) and (0, 1) after padding the pixel value at position (0, 0) to the (0, -1) position if the pixel value at position (0, -1) is not available. Additionally, the vertical gradient (G) at the top-right (3, 0) position 3, 0 If the pixel value at position (3, -1) is not available, the pixel value at position (3, 0) can be padded to position (3, -1), and then calculated by Equation 20(a) using the pixel values ​​at positions (3, -1) and (3, 1). For example, the gradient value in BIO can be derived by padding unavailable pixels outside the block boundary into pixels inside the block boundary, as shown in FIGS. 17 and 18.

[0365] As another example of reducing reference picture access memory bandwidth, the gradient value at each pixel location within the block can be calculated by Equation 21 using only the available pixel values ​​within the block.

[0366]

[0368] In mathematical formula 21, (b) and (c) are formulas for calculating the gradient when pixel values ​​outside the block boundary are unavailable at the block boundary location, and (a) is a formula for calculating the gradient when surrounding pixels to the left / right or up / down are available inside the block. That is, different gradient calculation formulas can be applied based on the location where the gradient within the block is calculated. For example, the gradient value in BIO can be derived by padding unavailable pixels outside the block boundary into pixels inside the block boundary, as shown in FIG. 19.

[0369] The embodiment illustrated in FIG. 9(a) calculates the gradient at every pixel position in units of 4x4 sub-blocks for a 4x4 block, for example, the horizontal gradient (G) at the top-left (0, 0) position. 0,0 ) can be calculated using only the pixel values ​​at positions (0, 0) and (1, 0) if the pixel value at position (-1, 0) is not available. Additionally, the horizontal gradient (G) at the top-right position (3, 0) 3, 0 ) can be calculated using only the pixel values ​​at positions (3, 0) and (2, 0) if the pixel value at position (4, 0) is not available. In addition, the vertical gradient (G) at the top-left (0, 0) position 0,0 ) can be calculated using only the pixel values ​​at positions (0, 0) and (0, 1) if the pixel value at position (0, -1) is not available. Also, the vertical gradient (G) at the top-right (3, 0) position 3, 0) can be calculated using only the pixel values ​​at positions (3, 0) and (3, 1) if the pixel value at position (3, -1) is unavailable. In the above Equation 21, (b) and (c) achieve the same effect as padding the unavailable pixel values ​​at positions outside the 4x4 block boundary of the embodiment illustrated in FIG. 9 (a) with values ​​calculated from the block boundary pixel values ​​as in Equation 22, and then calculating the gradient values ​​at all positions within the block using Equation 21 (a). For example, the horizontal gradient (G) at the top-left position (0, 0) 0, 0 ) can be calculated using the pixel values ​​at positions (-1, 0) and (1, 0) after padding the pixel value at position (-1, 0) with the value calculated from Equation 22(a) using the pixel values ​​at positions (0, 0) and (1, 0) if the pixel value at position (-1, 0) is not available. Additionally, the horizontal gradient (G) at the top-right position (3, 0) 3, 0 ) can be calculated using the pixel values ​​at positions (4, 0) and (2, 0) after padding the pixel value at position (4, 0) with the value calculated from Equation 22(b) using the pixel values ​​at positions (3, 0) and (2, 0) if the pixel value at position (4, 0) is not available. Additionally, the vertical gradient (G) at the top-left (0, 0) position 0,0 ) can be calculated using the pixel values ​​at positions (0, -1) and (0, 1) after padding the pixel value at position (0, -1) with the value calculated from Equation 22(c) using the pixel values ​​at positions (0, 0) and (0, 1) if the pixel value at position (0, -1) is not available. Additionally, the vertical gradient (G) at the top-right (3, 0) position 3, 0) can be calculated using the pixel values ​​at positions (3, -1) and (3, 1) after padding the pixel value at position (3, -1) with the value calculated from Equation 22(d) using the pixel value at position (3, 0) and the pixel value at position (3, 1) when the pixel value at position (3, -1) is not available. For example, the gradient value in BIO can be derived by padding the unavailable pixels outside the block boundary into the block boundary pixels as shown in FIG. 20.

[0371]

[0373] Motion correction vector for calculating the BIO offset of the current block , It can be calculated in pixel units or in units of at least one or more subgroups.

[0374] In calculating on a subgroup basis, the size of the subgroup may be determined based on the ratio of the width to the height of the current target block, or information regarding the subgroup size may be entropy encoded / decoded. Additionally, predefined fixed-size subgroup units may be used depending on the size and / or shape of the current block.

[0375] FIG. 10 is a diagram illustrating various embodiments of subgroups that serve as units for calculating BIO offsets.

[0376] For example, if the current target block size is 16x16 as in FIG. 10(a), in 4x4 subgroup units , It can be calculated.

[0377] For example, if the current target block size is 8x16 as in FIG. 10(b), in 2x4 subgroup units , It can be calculated.

[0378] For example, if the current target block size is 16x8 as in Fig. 10 (c), in 4x2 subgroup units , It can be calculated.

[0379] For example, as in FIG. 10(d), if the current target block size is 8x16, in units of one 8x8 subgroup, two 4x4 subgroups, and eight 2x2 subgroups , It can be calculated.

[0380] Or, the width and height of the current target block, , The size of a subgroup unit can be defined using at least one of a minimum depth information value and a pre-defined minimum subgroup size to induce the minimum depth information value. The minimum depth information value can be transmitted in entropy-encoded form.

[0381] For example, if the current target block size is 64x64, the minimum depth information value is 3, and the predefined minimum subgroup size is 4, the size of the subgroup unit can be determined as 8x8 by the following mathematical formula 23.

[0382]

[0384] For example, if the current target block size is 128x64, the minimum depth information value is 3, and the predefined minimum subgroup size is 4, the subgroup unit size can be determined as 8x8 by the following mathematical formula 24.

[0385]

[0387] If BIO at the subgroup level is applied to the current target block, deblocking filtering for the current target block can be performed after determining whether to apply deblocking filtering at the subgroup level.

[0388] For example, if the size of the subgroup unit is 2x4 as in (b) of FIG. 10, deblocking filtering can be performed after determining whether to apply deblocking filtering to the boundary between subgroups with a horizontal length greater than 4 and the boundary between subgroups with a vertical length greater than 4 within the current target block.

[0389] If BIO at the subgroup level is applied to the current target block, transformation and inverse transformation can be performed at the subgroup level.

[0390] For example, if the size of the subgroup unit is 4x4 as in (a) of Fig. 10, the conversion and inverse conversion can be performed in 4x4 units.

[0392] Subgroup-based motion correction vector , is calculated at the subgroup level It can be calculated from the value.

[0394] Refers to subgroup-level BIO correlation parameters is the gradient value at each pixel position within the block area without expanding the current block. It can be calculated from the BIO correlation parameter at each pixel location calculated from Equation 8 using only the pixel values.

[0395] For example, if the size of the current block's subgroup is 4x4, the motion correction vector at the subgroup level can be calculated as shown in Equation 25 below, using only the gradient value and pixel value at each pixel position without expanding the block area, as the sum of the BIO correlation parameters at each pixel position calculated from Equation 8. In the above, S represents the pixel-unit BIO correlation parameter calculated from the pixel value and gradient value at each pixel location as in mathematical formula 8.

[0396]

[0398] Fig. 11 shows subgroup Sgroup This is a diagram illustrating the weights that can be applied to each S value within a subgroup in order to calculate.

[0399] Subgroup unit In the process of calculating the value, each within the subgroup as shown in FIG. 11 (a) The sum of values ​​obtained by applying equal weights to the subgroups It can be used as a value.

[0400] As shown in Fig. 9(b), when the gradient is calculated only at inner positions within the subgroup, the S value calculated by weighting the S values ​​calculated using the corresponding gradient is the subgroup's It can be used as a value.

[0401] As shown in Fig. 9(d), when gradients are calculated only at inner positions within a subgroup and the gradients at outer positions are copied and used, the S value is calculated by weighted summing the S values ​​calculated using only the gradient values ​​at inner positions, and the subgroup It can be used as a value

[0402] As shown in Fig. 9(d), when gradients are calculated only at inner positions within a subgroup and the calculated gradients at outer positions are copied, the S value obtained by the weighted sum of the S values ​​calculated from the gradient values ​​at all positions within the subgroup is the subgroup It can be used as a value.

[0404] Subgroup unit In the process of calculating the value, each within the subgroup as shown in FIG. 11 (b) The sum of values ​​obtained by applying different weights to the values ​​of the subgroups It can be used as a value

[0405] Or, of a specific location within a subgroup The value of the subgroup It can be used as follows. For example, if the size of the current block's subgroup is 4x4 as in Fig. 11 (c), The s value of the location of the corresponding subgroup It can be used as such. Information regarding the specific location above may be predetermined in the encoder / decoder. Alternatively, it may be signaled through a bitstream or derived based on the encoding parameters (size, shape, etc.) of the current block.

[0406] Fig. 12 shows subgroup S group This is a drawing for explaining an example of weighted summing only the S values ​​at specific locations within a subgroup to obtain the value.

[0407] As illustrated in FIG. 12 (a) to (d), by weighting only the S values ​​at specific locations within a subgroup It can be used as.

[0409] Subgroup unit obtained by the above method Using the value, the motion correction vector at the sub-group level as in Equation 7 ,, The above motion correction vector can be obtained. ,, and using the gradient value at each pixel position within the subgroup, in mathematical formula 6 at each pixel position within the subgroup A BIO offset value corresponding to can be calculated. When calculating the BIO offset value, for pixel locations where the gradient is not calculated, the gradient value calculated inside the block can be copied and used in the calculation as shown in (d) of FIG. 9.

[0410] By using motion correction vectors derived at the subgroup level, representative values ​​for gradient values ​​at each pixel location within the subgroup, and representative values ​​for pixel values ​​within the subgroup, a BIO offset can be calculated at the subgroup level and the same BIO offset can be applied to each pixel location within the subgroup. The representative values ​​for gradient values ​​within the subgroup may refer to at least one of the minimum value, maximum value, average value, weighted average value, mode, interpolated value, and median value of the gradient values.

[0411] The representative value of the pixel values ​​within the above subgroup may mean at least one of the minimum value, maximum value, average value, weighted average value, mode, interpolated value, and median value of the pixel values.

[0412] A representative value can be calculated for the BIO offset values ​​obtained at each pixel location within a subgroup, and the same BIO offset value can be applied at each pixel location within the subgroup. The representative value may represent at least one of the minimum value, maximum value, average value, weighted average value, mode, interpolated value, and median value of the BIO offset values.

[0413] Subgroup-based motion correction vector , is a motion correction vector calculated in pixel units within the subgroup , It can be calculated using

[0414] For example, if the size of the subgroup is 2x2, the subgroup as shown in Equation 26 below , It can be calculated.

[0415]

[0417] subgroup unit , The derivation of can be determined based on the size of the current block. The subgroup unit can be determined based on a comparison between the size of the current block and a predetermined threshold value. The predetermined threshold value is , It may refer to a reference size that determines the derivation unit. This may be expressed in the form of at least one of a minimum value and a maximum value. A predetermined threshold value may be a fixed value pre-agreed upon in the encoder / decoder, or it may be variably derived based on the encoding parameters of the current block (e.g., motion vector magnitude, etc.). Alternatively, it may be signaled through a bitstream (e.g., sequence, picture, slice, block level, etc.).

[0418] For example, blocks whose product of width and height is 256 or greater are classified into subgroups , Blocks that are not calculated can be calculated in pixel units.

[0419] For example, blocks where the minimum length between the width and height is 8 or greater are classified as subgroups , It calculates, and blocks that do not can be calculated in pixel units.

[0420] subgroup unit , through comparing the magnitudes of the horizontal and vertical gradient values or Only the derivative can be used to calculate the BIO offset.

[0421] For example, if the sum of the absolute values ​​of the horizontal gradient values ​​for the L0 reference prediction block and the horizontal gradient values ​​for the L1 reference prediction block is greater than the sum of the absolute values ​​of the vertical gradient values ​​for the L0 reference prediction block and the vertical gradient values ​​for the L1 reference prediction block, Only the value can be calculated and used for BIO offset calculation. In this case, The value can mean 0.

[0422] For example, if the sum of the absolute values ​​of the vertical gradient values ​​for the L0 reference prediction block and the vertical gradient values ​​for the L1 reference prediction block is greater than the sum of the absolute values ​​of the horizontal gradient values ​​for the L0 reference prediction block and the horizontal gradient values ​​for the L1 reference prediction block, Only the value can be calculated and used for BIO offset calculation. In this case, The value can mean 0.

[0424] Subgroup unit The value is calculated in pixel units by considering the gradient values ​​of surrounding pixel locations for the current block. It can be calculated from the values.

[0425] FIG. 13 is a diagram illustrating an example of calculating the S value.

[0426] At the top-left position (0,0) within the current block The value can be calculated by applying a 5x5 window to the corresponding location and considering the gradient values ​​of surrounding pixel locations together with the gradient value at the current location. The gradient value at a location outside the current block can be used by utilizing the gradient value within the current block, as shown in FIG. 13 (a), or by calculating it directly. At other locations within the current block The value can also be calculated in the same way.

[0427] Subgroup units within the current block The value can be calculated by applying different weights depending on the position. In this case, only the gradient value within the current block can be used without expanding the block.

[0428] For example, if the subgroup size is 2x2, when a 5x5 window is applied to each pixel location, the calculated gradient values ​​of surrounding pixel locations By applying a 6x6 weighting table as shown in Fig. 13(b) to the values ​​of the subgroup The value can be calculated.

[0429] For example, if the subgroup size is 4x4, when a 5x5 window is applied to each pixel location, the calculated gradient values ​​of surrounding pixel locations are considered. By applying an 8x8 weighting table as in Fig. 13 (c) to the values ​​of the subgroup The value can be calculated.

[0430] For example, if the subgroup size is 8x8, when a 5x5 window is applied to each pixel location, the calculated gradient values ​​of surrounding pixel locations are considered. By applying a 12x12 weighting table as shown in Fig. 13 (d) to the values ​​of the subgroup The value can be calculated.

[0431] Figure 14 shows S when the size of the subgroup is 4x4. group This is a drawing for explaining an example of calculating. FIG. 14(a) is a drawing illustrating the gradient at a pixel location within a 4x4 block and at surrounding pixel locations. FIG. 14(b) is a drawing at a pixel location within a 4x4 block and at surrounding pixel locations This is a drawing showing the values.

[0432] For example, if the size of the subgroup is 4x4, as shown in Fig. 14, the sum of all S values ​​at each position calculated from the gradient obtained from the current block area pixel position as well as the surrounding pixel positions and the pixel values It can be calculated. For example, the following mathematical formula 27 may be used. In this case, the weight at each location may be the same as a predetermined value (e.g., 1), or different weights may be applied. If the S value at a surrounding pixel location is not available, only the available surrounding S values ​​are summed to the S values ​​within the current subgroup. ...can be calculated. If the surrounding pixel location is outside the boundary of the block area and is not available, the S value of the nearest block area boundary can be padded and used. If the gradient and pixel value are not available because the surrounding pixel location is outside the boundary of the block area, the gradient and pixel value of the nearest block area boundary can be padded and the S value at that location can be calculated.

[0433]

[0434] In the above-described embodiment, the size and weights of the weight table may vary depending on the size of the MxN window applied to each pixel location. M and N are natural numbers greater than 0, and M and N may be the same or different from each other.

[0435] V calculated at the subgroup level x , V y By reflecting this in the first and second motion information, the motion information of the current block is updated and saved in subgroup units, and then used for the next target block. When updating the motion vector in subgroup units, the motion correction vector (V) of a predefined predetermined subgroup position is used. x , V y The motion information of the current block can be updated using only ). As shown in FIG. 10 (a), when the target block is 16x16 and the size of the subgroup is 4x4, the motion vector reflecting only the motion correction vector of the first subgroup in the upper left corner to the first motion vector and the second motion vector of the current block can be stored as the motion vector of the current block.

[0437] Below, the derivation of the motion vector of the chrominance component is explained.

[0438] According to one embodiment, a motion correction vector (V) calculated in subgroup units from the luminance component. x , V y) can be reflected in the motion vector of the chrominance component and used in the motion compensation process for the chrominance component.

[0439] Alternatively, a motion vector in which the motion correction vector of a predefined relative position subgroup is reflected in the first and second motion vectors of the current block can be used as the motion vector of the chrominance component.

[0440] FIG. 15 is a diagram illustrating an embodiment for deriving a motion vector of a color difference component based on a luminance component.

[0441] As shown in FIG. 15, when the current target block size is 8x8 and the subgroup size is 4x4, the motion vector of the chrominance block is the motion correction vector (V) of sub-block ④. x , V y ) is the first motion vector of the current target block( ) and the second motion vector( It can be a motion vector reflected in ). That is, the motion vector of the chrominance component can be derived as shown in Equation 28 below.

[0442]

[0443] As a motion correction vector for calculating the motion vector of the color difference block above, a motion correction vector of another sub-block may be used instead of the motion correction vector of sub-block ④. Alternatively, at least one of the maximum value, minimum value, median value, average value, weighted average value, and mode of the motion correction vectors of two or more sub-blocks among sub-blocks ① to ④ may be used.

[0445] Motion correction vector (V) calculated from luminance components in subgroup units x , V y ) can be reflected in the color difference components (Cb, Cr) and used in the motion compensation process for the color difference components.

[0446] Figure 16 is an example diagram illustrating the motion compensation process for color difference components.

[0447] As shown in FIG. 16, when the size of the subgroup unit of the luminance block is 4x4, the corresponding chrominance block (Cb, Cr) has a subgroup of size 2x2, and the motion correction vector of each subgroup within the chrominance block can use the motion correction vector of the subgroup of the corresponding luminance block.

[0448] For example, the motion correction vector (Vcx1, Vcy1) of the first subgroup of the color difference block (Cb, Cr) can use the motion correction vector (Vx1, Vy1) of the first subgroup of the corresponding luminance block.

[0449] Using the motion correction vector values ​​obtained in subgroup units of the chrominance components (Cb, Cr) and the pixel values ​​of the restored chrominance components (Cb, Cr), the BIO offset can be calculated in subgroup units of the chrominance components (Cb, Cr), just as in the case of the luminance components. In this case, the following Equation 29 can be used.

[0450]

[0452] In the above mathematical formula, It can be obtained from the restored pixels of the reference pictures of the color difference components (Cb, Cr).

[0453] For example, at the pixel value location within the second subgroup of the color difference component shown in FIG. 16 It can be calculated as follows.

[0454] G of the x component at position P2 c It can be calculated through the difference between the pixel value at position P1 and the pixel value at position P3.

[0455] G of the x component at position P3 c It can be calculated through the difference between the pixel value at position P2 and the pixel value at position P3.

[0456] G of the y component at position P2 c It can be calculated through the difference between the pixel value at position P6 and the pixel value at position P2.

[0457] The y-component Gc at position P6 is the pixel value at position P2 and P 10 It can be calculated through the difference with the pixel value at the location.

[0459] The encoder can determine whether to perform BIO on the current block and then encode information indicating whether to perform it (e.g., entropy encoding). Whether to perform BIO can be determined by comparing distortion values ​​between the predicted signal before applying BIO and the predicted signal after applying BIO. The decoder can decode information indicating whether to perform BIO from the bitstream (e.g., entropy decoding) and perform BIO according to the received information. Additionally, information indicating whether to enable BIO can be signaled at a higher level (sequence, picture, slice, CTU, etc.). For example, information indicating whether to perform BIO can be determined only when the information indicating whether to enable BIO indicates BIO activation.

[0460] Information indicating whether to perform BIO can be entropy encoded / decoded based on the encoding parameters of the current block.

[0461] Alternatively, information indicating whether to perform BIO may be omitted from encoding / decoding based on the encoding parameters of the current block. That is, information indicating whether to perform BIO may be determined based on the encoding parameters of the current block. Here, the encoding parameters may include at least one of a prediction mode, accuracy of motion compensation, size of the current block, shape, partitioning type (whether it is a quad-tree partitioning, a binary tree partitioning, or a ternary tree partitioning), picture type, global motion compensation mode, and motion correction mode in the decoder.

[0462] For example, the encoder / decoder can determine the accuracy of motion compensation by using a prediction signal generated by performing motion compensation based on the first motion information of the current block and a prediction signal generated by performing motion compensation based on the second motion information. For instance, the encoder / decoder can determine the accuracy of motion compensation based on the difference signal between the two prediction signals, or by comparing the difference signal with a predetermined threshold value. The difference signal may refer to the SAD value between the two prediction signals.

[0463] A predetermined threshold value refers to a reference value that determines whether to perform BIO by determining the accuracy of the difference signal. This can be expressed in at least one form of a minimum value and a maximum value. The predetermined threshold value may be a fixed value pre-agreed upon in the encoder / decoder, a value determined by encoding parameters such as the size, shape, and bit depth of the current block, and a value signaled at the SPS, PPS, Slice header, Tile, CTU, and CU levels.

[0464] Additionally, the encoder / decoder can determine whether to perform BIO on a sub-block basis for the current block. The encoder / decoder can determine whether to perform BIO on a sub-block basis based on the comparison of the difference signal between two predicted signals corresponding to the sub-block and a predetermined threshold value at each sub-block level. The predetermined threshold value used at the sub-block level may be the same as or different from the threshold value used at the block level. It may be expressed in at least one form of a minimum value and a maximum value, and may be a fixed value pre-agreed upon by the encoder / decoder, a value determined by encoding parameters such as the size, shape, and bit depth of the current block, or a value signaled at the SPS, PPS, Slice header, Tile, CTU, and CU levels. For example, if the SAD of a subblock for the current block is smaller than a threshold value determined based on the size of the current block (e.g., 2 x width of the subblock x height of the subblock), BIO may not be applied to the current subblock.

[0465] For example, when the current block is in merge mode, the encoder / decoder does not entropy-encode / decode information indicating whether to perform BIO, and can always apply BIO.

[0466] For example, if the current block is in AMVP mode, the encoder / decoder can entropy-encode / decode information indicating whether to perform BIO, and perform BIO according to that information.

[0467] For example, when the current block is in AMVP mode, the encoder / decoder does not entropy-encode / decode information indicating whether to perform BIO, and can always apply BIO.

[0468] For example, when the current block is in AMVP mode, the encoder / decoder may not entropy-encode / decode information indicating whether to perform BIO, and may not always apply BIO.

[0469] For example, if the current block is in merge mode, the encoder / decoder can entropy-encode / decode information indicating whether to perform BIO, and perform BIO according to that information.

[0470] For example, if the current block is in AMVP mode and motion compensation is performed in quarter-pixel units, the encoder / decoder can always apply BIO without entropy-encoding information indicating whether to perform BIO. Additionally, if the current block is in AMVP mode and motion compensation is performed in integer-pixel units (1 pixel or 4 pixels), the encoder / decoder can entropy-encode information indicating whether to perform BIO and perform BIO according to that information.

[0471] For example, if the current block is in AMVP mode and motion compensation is performed in quarter pixel units, the encoder / decoder may not always apply BIO without entropy-decoding the BIO performance indicator information. Additionally, if the current block is in AMVP mode and motion compensation is performed in integer pixel units (1 pixel or 4 pixels), the encoder / decoder may not always apply BIO without entropy-decoding the BIO performance indicator information.

[0472] For example, when the current block is in AMVP mode and motion compensation is performed in integer pixel units (1 pixel or 4 pixels), the encoder / decoder can always apply BIO without entropy-encoding information indicating whether to perform BIO, and when motion compensation is performed in quarter pixel units, it can entropy-encode information indicating whether to perform BIO and perform BIO according to that information.

[0473] For example, the encoder / decoder can perform BIO based on entropy encoding / decoding of information indicating whether to perform BIO when the current block is in AMVP mode and its size is smaller than or equal to 256 luminance pixels. If the condition is not satisfied, BIO can always be performed.

[0474] For example, the encoder / decoder may not always perform BIO if the current block is in inter-intra combined prediction mode.

[0476] For example, the encoder / decoder may not always perform BIO if the current block is in an affine motion model-based motion prediction mode. The above affine motion model-based motion prediction mode may correspond to the case where the encoding parameter MotionModelIdc has a non-zero value.

[0477] For example, the encoder / decoder may not always perform BIO when the current block is in a symmetric motion vector difference mode. The symmetric motion vector difference mode may mean a mode in which the motion vector difference value in the L1 direction is not entropy encoded / decoded, and the values ​​obtained by mirroring the horizontal and vertical component values ​​(MVD0x, MVD0y) of the motion vector difference value in the L0 direction in the L1 direction (-MVD0x, -MVD0y) are used as the motion vector difference value in the L1 direction.

[0478] For example, the encoder / decoder may not perform BIO if the current block size is less than or equal to a predefined size.

[0479] For example, the encoder / decoder may not perform BIO if the vertical length of the current block is 4.

[0480] For example, the encoder / decoder may not perform BIO if the width of the current block is 4 and the height is 8.

[0481] For example, the encoder / decoder may not perform BIO if the vertical length of the current block is less than 8.

[0482] For example, the encoder / decoder may not perform BIO if the width of the current block is less than 8.

[0483] For example, the encoder / decoder may not perform BIO if the area of ​​the current block is less than 128.

[0484] For example, the encoder / decoder may not perform BIO if the current block size is less than or equal to a predefined size and is partitioned into a binary tree.

[0485] For example, the encoder / decoder may not perform BIO if the current block size is less than or equal to a predefined size and is partitioned into a ternary tree.

[0486] For example, if the current block is in illumination compensation mode, affine mode, subblock merge mode, or a mode that corrects motion information in the decoder (e.g., PMMVD (Pattern matched motion vector derivation), DMVR (Decoder-side motion vector refinement), or Current Picture Referencing (CPR) mode that performs inter-frame prediction by referencing the current image containing the current block or the reconstructed pixels within the CTU containing the current block), the encoder / decoder may not always apply BIO.

[0487] The encoder / decoder can determine whether to perform BIO based on the reference picture of the current block.

[0488] For example, the encoder / decoder may apply BIO if all reference pictures in the current block are short-term reference pictures. Conversely, the encoder / decoder may not always apply BIO if at least one of the reference pictures in the current block is not a short-term reference picture.

[0489] Whether to apply BIO to the current target block can be determined based on entropy-decrypted flag information in at least one unit among the CTU unit and the sub-units of the CTU. In this case, the sub-unit may include at least one of the sub-units of the CTU, the CU unit, and the PU unit.

[0490] For example, if the CTU block size is 128x128 and information regarding BIO is entropy decoded in a 32x32 block unit, which is a sub-unit of the CTU, the encoder / decoder can perform BIO based on the entropy decoded BIO-related information in the 32x32 block unit for blocks that belong to the 32x32 block and are smaller in size than the 32x32 block unit.

[0491] For example, if the block depth of a CTU is 0 and the information regarding BIO is entropy decoded in a subunit of a CTU with a block depth of 1, the encoder / decoder is included in the subunit of the CTU, and for blocks with a block depth of 1 or more, BIO can be performed based on the entropy decoded BIO-related information in the subunit of a CTU with a block depth of 1.

[0492] The final predicted sample signal of the current block is the predicted sample signal (P) obtained through existing bidirectional prediction. conventional bi-prediction ) and predicted sample signals obtained through BIO (P optical flow It can be generated using the weighted sum of ). In this case, the following mathematical formula 30 can be used.

[0493]

[0495] In the above mathematical formula, the weight applied to each block ( ) may be identical to each other, or may be determined differently depending on the encoding parameters of the current block. The encoding parameters may include at least one of the following: a prediction mode, the accuracy of motion compensation, the size, shape, and partitioning type of the current block (whether it is a quad-tree partition, a binary tree partition, or a ternary tree partition), a global motion compensation mode, a motion correction mode in the decoder, and the hierarchy of the current picture to which the current block belongs.

[0496] For example, depending on whether the current block is in merge mode or AMVP mode, the weight ( ) can change.

[0497] For example, if the current block is in AMVP mode, the weight ( ) can change.

[0498] For example, if the current block is in merge mode, weights are applied according to affine mode, illumination compensation mode, and modes that correct motion information in the decoder (e.g., PMMVD, DMVR). ) can change.

[0499] For example, weights based on the size and / or shape of the current block ( ) can change.

[0500] For example, weights according to the temporal layer of the current picture to which the current block belongs ( ) can change.

[0501] For example, if BIO is applied to the current block on a subgroup basis, weights on a subgroup basis ( ) can change.

[0503] FIG. 21 is a flowchart illustrating an image decoding method according to an embodiment of the present invention.

[0504] Referring to FIG. 21, the decoder can determine whether the current block is in a bidirectional optical flow mode (S2110). Specifically, it can be determined based on the distance between the first reference picture of the current block and the current picture, and the distance between the second reference picture of the current block and the current picture.

[0505] For example, if the distance between the first reference picture and the current picture and the distance between the second reference picture and the current picture are not the same, the decoder can determine that the current block is not in bidirectional optical flow mode.

[0506] Meanwhile, the decoder can determine whether the current block is in bidirectional optical flow mode based on the type of the reference picture of the current block.

[0507] For example, if at least one of the type of the first reference picture of the current block and the type of the second reference picture of the current block is not a short-term reference picture, the decoder may determine that the current block is not in a bidirectional optical flow mode.

[0508] Meanwhile, the decoder can determine whether the current block is in bidirectional optical flow mode based on the size of the current block.

[0509] And, if the current block is in bidirectional optical flow mode (S2110-Yes), the decoder can calculate gradient information of the prediction samples of the current block (S2120). Specifically, the decoder can calculate gradient information using at least one neighbor sample adjacent to the prediction sample. In this case, if the neighbor sample is located outside the area of ​​the current block, the value of the neighbor sample can use the sample value of an integer pixel location adjacent to the neighbor sample.

[0510] Meanwhile, gradient information can be calculated in units of sub-blocks of a predefined size.

[0511] And, the decoder can generate a predicted block of the current block using the calculated gradient information (S2130).

[0512] In order to obtain the same prediction result as the decoder in the encoder, the image decoding method of FIG. 21 can also be performed as an image encoding method.

[0514] *

[0515] The bitstream generated by the image encoding method of the present invention may be temporarily stored in a computer-readable non-transient recording medium and may be a bitstream encoded by the image encoding method described above.

[0516] Specifically, as a computer-readable recording medium storing a bitstream generated by an image encoding method, the image encoding method may include a step of determining whether a current block is in a bidirectional optical flow mode; a step of calculating gradient information of prediction samples of a current block if the current block is in a bidirectional optical flow mode; and a step of generating a prediction block of a current block using the calculated gradient information. Here, the step of calculating gradient information of prediction samples of a current block may be characterized by calculating gradient information using at least one neighbor sample adjacent to a prediction sample.

[0518] The above embodiments can be performed in the same way in the encoder and decoder.

[0519] An image can be encoded / decoded using at least one of the above embodiments or a combination of at least one.

[0520] The order of applying the above embodiments may differ between the encoder and the decoder, and the order of applying the above embodiments may be the same between the encoder and the decoder.

[0521] The above example can be performed for each of the luminance and color difference signals, and the above example can be performed in the same way for the luminance and color difference signals.

[0522] The shape of the block to which the above embodiments of the present invention are applied may be square or non-square.

[0523] The above embodiments of the present invention may be applied according to the size of at least one of an encoding block, a prediction block, a conversion block, a block, a current block, an encoding unit, a prediction unit, a conversion unit, a unit, and a current unit. The size here may be defined as a minimum size and / or a maximum size for the application of the above embodiments, or it may be defined as a fixed size for the application of the above embodiments. Furthermore, the above embodiments may be applied as a first embodiment at a first size, and as a second embodiment at a second size. That is, the above embodiments may be applied in combination according to the size. Additionally, the above embodiments of the present invention may be applied only when the size is greater than or equal to the minimum size and less than or equal to the maximum size. That is, the above embodiments may be applied only when the block size falls within a certain range.

[0524] For example, the above embodiments may be applied only when the current block size is 8x8 or larger. For example, the above embodiments may be applied only when the current block size is 4x4. For example, the above embodiments may be applied only when the current block size is 16x16 or smaller. For example, the above embodiments may be applied only when the current block size is 16x16 or larger and 64x64 or smaller.

[0525] The embodiments of the present invention may be applied according to a temporal layer. A separate identifier is signaled to identify the temporal layer to which the embodiments are applicable, and the embodiments may be applied to the temporal layer specified by the identifier. The identifier may be defined as the lowest layer and / or the highest layer to which the embodiments are applicable, or it may be defined as indicating a specific layer to which the embodiments are applied. Additionally, a fixed temporal layer to which the embodiments are applied may be defined.

[0526] For example, the above embodiments may be applied only when the temporal layer of the current image is the lowest layer. For example, the above embodiments may be applied only when the temporal layer identifier of the current image is 1 or greater. For example, the above embodiments may be applied only when the temporal layer of the current image is the highest layer.

[0527] The slice type or tile group type to which the above embodiments of the present invention are applied is defined, and the above embodiments of the present invention may be applied according to the said slice type or tile group type.

[0528] In the embodiments described above, methods are described based on flowcharts as a series of steps or units; however, the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, that other steps may be included, or that one or more steps of the flowcharts may be omitted without affecting the scope of the present invention.

[0529] The embodiments described above include examples of various aspects. While it is not possible to describe all possible combinations for representing various aspects, those skilled in the art will recognize that other combinations are possible. Accordingly, the present invention shall be deemed to include all other substitutions, modifications, and changes falling within the scope of the following claims.

[0530] The embodiments according to the present invention described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.

[0531] Although the present invention has been described above with specific details such as specific components, limited embodiments, and drawings, this is provided only to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments, and a person skilled in the art to which the invention belongs can make various modifications and variations from this description.

[0532] Accordingly, the scope of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims set forth below, as well as the claims described below, shall be considered to fall within the scope of the concept of the present invention.

Claims

Claim 1 A step of constructing a merge candidate list for the current block; a step of deriving motion information of the current block based on the merge candidate list; a step of determining whether a bidirectional optical flow mode is applied to the current block; if the bidirectional optical flow mode is applied to the current block, a step of obtaining prediction samples for an extended sub-block based on the motion information of the current block, wherein the extended sub-block is composed of a sub-block and an extended region surrounding the sub-block; a step of calculating slope information for each predicted position within the extended sub-block; a step of obtaining a refinement vector of the sub-block based on the slope information for each predicted position within the extended sub-block; a step of deriving a prediction offset for a predicted position within the sub-block; A method for image decoding, comprising the step of obtaining a final prediction sample within a sub-block using the prediction offset and prediction sample for the prediction position, wherein the slope information of the prediction position is calculated using at least one neighbor prediction sample adjacent to the prediction position, and if the integer sample position closest to the reference position for the first prediction sample within the extended area is available, the first prediction sample is obtained from a sample at the integer sample position, and the reference position is determined by the motion vector of the current block. Claim 2 A video decoding method according to claim 1, characterized in that whether the bidirectional light flow mode is applied to the current block is determined based on a first distance between a first reference picture of the current block and a current picture and a second distance between a second reference picture of the current block and a current picture. Claim 3 A video decoding method according to paragraph 2, characterized in that when the first distance and the second distance are not the same, the bidirectional optical flow mode is not applied to the current block. Claim 4 A video decoding method according to claim 1, characterized in that whether the bidirectional light flow mode is applied to the current block is determined based on the type of reference pictures of the current block. Claim 5 A video decoding method according to claim 4, characterized in that when at least one of the type of the first reference picture of the current block and the type of the second reference picture of the current block is not a short-term reference picture, the bidirectional light flow mode is not applied to the current block. Claim 6 An image decoding method according to claim 1, characterized in that whether the bidirectional optical flow mode is applied to the current block is determined based on the size of the current block. Claim 7 A step of constructing a merge candidate list for the current block; a step of deriving motion information of the current block based on the merge candidate list; a step of determining whether a bidirectional optical flow mode is applied to the current block; if the bidirectional optical flow mode is applied to the current block, a step of obtaining prediction samples for an extended sub-block based on the motion information of the current block, wherein the extended sub-block is composed of a sub-block and an extended region surrounding the sub-block; a step of calculating slope information for each predicted position within the extended sub-block; a step of obtaining a refinement vector of the sub-block based on the slope information for each predicted position within the extended sub-block; a step of deriving a prediction offset for a predicted position within the sub-block; A video encoding method comprising the step of obtaining a final prediction sample for the prediction position within the sub-block using the prediction offset and prediction sample for the prediction position, wherein the slope information of the prediction position is calculated using at least one neighbor prediction sample adjacent to the prediction position, and if the integer sample position closest to the reference position for the first prediction sample within the extension area is available, the first prediction sample is obtained from a sample at the integer sample position, and the reference position is determined by the motion vector of the current block. Claim 8 A computer-readable recording medium storing a bitstream generated by a video encoding method, wherein the video encoding method comprises: a step of configuring a merge candidate list of a current block; a step of deriving motion information of the current block based on the merge candidate list; a step of determining whether a bidirectional optical flow mode is applied to the current block; a step of, if the bidirectional optical flow mode is applied to the current block, a step of acquiring prediction samples for an extended sub-block based on the motion information of the current block, wherein the extended sub-block is composed of a sub-block and an extended region surrounding the sub-block; a step of calculating slope information of each predicted position within the extended sub-block; a step of acquiring a refinement vector of the sub-block based on the slope information of each predicted position within the extended sub-block; and a step of deriving a prediction offset for a predicted position within the sub-block. A computer-readable recording medium comprising the step of obtaining a final prediction sample for the prediction position within the sub-block using the prediction offset and prediction sample for the prediction position, wherein the slope information of the prediction position is calculated using at least one neighbor prediction sample adjacent to the prediction position, and if the integer sample position closest to the reference position for the first prediction sample within the extension area is available, the first prediction sample is obtained from a sample at the integer sample position, and the reference position is determined by the motion vector of the current block.

Citation Information

Patent Citations

  • Inter-prediction method and apparatus in image coding system

    KR1020180129860A

  • Method and apparatus for processing video signal by using improved optical flow motion vector

    WO2018048265A1

  • Encoding device, decoding device, encoding method and decoding method

    WO2018212110A1