Method, apparatus and recording medium for encoding / decoding image using affine transform
Patent Information
- Application Number
- KR1020200075832
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-06-22
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2040-06-22
Smart Images

Figure 112020064014146-PAT00027_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method, apparatus, and recording medium for image encoding / decoding. Specifically, the present invention relates to an image encoding / decoding method and apparatus utilizing affine conversion, and a recording medium storing a bitstream. Background Technology
[0002] With the continuous development of the information and communications industry, broadcasting services featuring HD (High Definition) resolution have spread globally. Through this expansion, many users have become accustomed to high-resolution and high-quality images and / or videos.
[0003] In order to satisfy users' demand for high image quality, many organizations are accelerating the development of next-generation video devices. User interest has increased not only in High Definition TV (HDTV) and Full HD (FHD) TVs, but also in Ultra High Definition (UHD) TVs, which have more than four times the resolution of FHD TVs. With this increased interest, video encoding / decoding technology for video with higher resolution and image quality is required.
[0004] As video compression technologies, various techniques exist, such as inter-prediction, intra-prediction, transformation and quantization, and entropy encoding.
[0005] Inter-prediction technique is a technique that predicts the values of pixels included in the current picture by using the previous and / or subsequent pictures of the current picture. Intra-prediction technique is a technique that predicts the values of pixels included in the current picture by using information about pixels within the current picture. Transformation and quantization techniques are methods for compressing the energy of residual signals. Entropy coding technique is a technique that assigns short codes to values with high occurrence frequency and long codes to values with low occurrence frequency.
[0006] Using such image compression technology, data regarding images can be effectively compressed, transmitted, and stored. The problem to be solved
[0007] One embodiment may provide an apparatus and method for performing inter prediction through an affine transformation suitable for the distance of a reference picture.
[0008] One embodiment may provide an apparatus and method for performing inter-prediction through appropriate affine transformation according to conditions in encoding / decoding.
[0009] One embodiment may provide an apparatus and method for generating a prediction block through affine transformation advanced motion vector prediction (AMVP) and / or affine transformation merge based on the distance between a target block and a reference block. means of solving the problem
[0010] A video decoding method is provided, comprising: a step of determining a prediction mode for a target block in one aspect; and a step of performing a prediction for the target block using the prediction mode, wherein if the prediction mode is an intermode using an affine transformation, derivation of an affine transformation model and decoding of the difference value of the motion vector of a control point are performed, configuration of an affine transformation advanced motion vector prediction (AMVP) list and derivation of the control point motion vector are performed, and a motion vector of the target block is generated using the control point motion vector. Effects of the invention
[0011] An apparatus and method for performing inter prediction through an affine transformation suitable for the distance of a reference picture are provided.
[0012] An apparatus and method are provided for performing inter-prediction through appropriate affine transformation according to conditions in encoding / decoding.
[0013] An apparatus and a method for generating a prediction block through affine transformation AMVP and / or affine transformation merging based on the distance between a target block and a reference block are provided. Brief explanation of the drawing
[0014] FIG. 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied. FIG. 2 is a block diagram showing the configuration according to one embodiment of a decoding device to which the present invention is applied. Figure 3 is a diagram schematically showing the segmentation structure of an image when encoding and decoding an image. Figure 4 is a diagram illustrating the shape of a prediction unit that a coding unit may include. Figure 5 is a drawing illustrating the form of a conversion unit that can be included in a coding unit. Figure 6 shows the division of a block according to one example. Figure 7 is a diagram illustrating an example of an intra-prediction process. Figure 8 is a diagram illustrating a reference sample used in the intra-prediction process. Figure 9 is a diagram illustrating an example of an inter prediction process. Figure 10 shows spatial candidates according to one example. Figure 11 shows the order of addition of motion information of spatial candidates to a merge list according to one example. Figure 12 illustrates the process of transformation and quantization according to one example. Figure 13 shows diagonal scanning according to one example. Figure 14 shows horizontal scanning according to one example. Figure 15 shows vertical scanning according to one example. FIG. 16 is a structural diagram of an encoding device according to one embodiment. FIG. 17 is a structural diagram of a decoding device according to one embodiment. Figures 18 and 19 show control point motion vectors according to the number of parameters of an affine transformation model according to one example. Figure 18 shows the control point motion vector of a 4-parameter affine transformation model according to one example. Figure 19 shows the control point motion vector of a 6-parameter affine transformation model according to one example. Figure 20 shows an encoding method using an affine conversion AMVP mode according to one example. Figure 21 shows a decoding method using an affine conversion AMVP mode according to one example. Figure 22 illustrates an encoding method using an affine transform merge mode according to one example. Figure 23 shows a decoding method using an affine transform merge mode according to one example. FIG. 24 shows the location of adjacent reference pixels that can identify spatial candidates for a target block according to one example. Figure 25 illustrates a method for calculating the predicted value of an affine transformation control point according to one example. Figure 26 illustrates the process of constructing an affine transformation control point prediction list according to one example. FIG. 27 is a flowchart of a method for predicting a target block and a method for generating a bitstream according to one embodiment. FIG. 28 is a flowchart of a method for predicting a target block using a bitstream according to one embodiment. Specific details for implementing the invention
[0015] The present invention is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the invention to specific embodiments, and it should be understood that the invention includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the invention.
[0016] The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein with respect to one embodiment may be implemented in other embodiments without departing from the spirit and scope of the invention. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the detailed description below is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments is limited only by the appended claims, with all equivalents to those claimed therein, provided appropriately described.
[0017] In drawings, similar reference numerals refer to the same or similar functions across multiple aspects. The shapes and sizes of elements in the drawings may be exaggerated for clearer explanation.
[0018] In the present invention, terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" may include a combination of a plurality of related described items or any of a plurality of related described items.
[0019] When it is stated that a component is "connected" or "connected" to another component, it should be understood that the two components may be directly connected or connected to each other, or that another component may exist between the two components. On the other hand, when it is stated that a component is "directly connected" or "directly connected" to another component, it should be understood that no other component exists between the two components.
[0020] The components shown in the embodiments of the present invention are illustrated independently to represent different characteristic functions and do not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for the convenience of explanation; however, at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separated embodiments of each component are included within the scope of the present invention as long as they do not deviate from the essence of the invention.
[0021] Furthermore, the description that a specific configuration in the exemplary embodiments is "included" does not exclude configurations other than the specific configuration mentioned above, and means that additional configurations may be included within the scope of the practice of the exemplary embodiments or the technical concept of the exemplary embodiments.
[0022] The terms used in this invention are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this invention, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, the description in this invention that a specific configuration "comprising" does not exclude configurations other than that configuration, and means that additional configurations may also be included within the scope of the practice or technical concept of this invention.
[0023] Some components of the present invention may not be essential components performing an essential function in the present invention, but may be optional components merely for enhancing performance. The present invention may be implemented by including only the essential components for realizing the essence of the present invention, excluding the components used merely for performance enhancement. A structure including only the essential components, excluding the optional components used merely for performance enhancement, is also included within the scope of the present invention.
[0024] In the following, embodiments are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the embodiments. In describing the embodiments, detailed descriptions of related known components or functions are omitted if it is determined that such detailed descriptions may obscure the gist of this specification. Additionally, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0025] In the following, "image" may refer to a picture constituting a video, or it may refer to the video itself. For example, "encoding and / or decoding of an image" may mean "encoding and / or decoding of a video," and may also mean "encoding and / or decoding of one of the images constituting a video."
[0026] In the following, the terms "video" and "motion picture(s)" may be used interchangeably with the same meaning.
[0027] In the following, the target image may be an image to be encoded and / or an image to be decoded. Additionally, the target image may be an input image input to an encoding device and an input image input to a decoding device. Additionally, the target image may be a current image currently being encoded and / or decoded. For example, the terms "target image" and "current image" may be used interchangeably.
[0028] In the following, the terms "image," "picture," "frame," and "screen" may be used interchangeably with the same meaning.
[0029] In the following, the target block may be an encoding target block that is the target of encoding and / or a decoding target block that is the target of decoding. Additionally, the target block may be a current block that is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably with the same meaning. The current block may mean an encoding target block that is the target of encoding during encoding and / or a decoding target block that is the target of decoding during decoding. Additionally, the current block may be at least one of a coding block, a prediction block, a residual block, and a transformation block.
[0030] In the following, the terms "block" and "unit" may be used interchangeably with the same meaning. Alternatively, "block" may refer to a specific unit.
[0031] In the following, the terms "region" and "segment" may be used interchangeably.
[0032] In the following, a specific signal may be a signal representing a specific block. For example, the original signal may be a signal representing the target block. The prediction signal may be a signal representing the prediction block. The residual signal may be a signal representing the residual block.
[0033] In the embodiments, each of the specified information, data, flag, index, element, attribute, etc., may have a value. A value "0" of the information, data, flag, index, element, and attribute, etc., may represent logical false or a first predefined value. That is to say, the value "0", false, logical false, and the first predefined value may be used interchangeably. A value "1" of the information, data, flag, index, element, and attribute, etc., may represent logical true or a second predefined value. That is to say, the value "1", true, logical true, and the second predefined value may be used interchangeably.
[0034] When a variable such as i or j is used to represent a row, column, or index, the value of i may be an integer greater than or equal to 0, or an integer greater than or equal to 1. That is to say, in the embodiments, the row, column, and index, etc., may be counted from 0 and may be counted from 1.
[0035] In the embodiments, the term "one or more" or the term "at least one" may mean the term "plural." "One or more" or "at least one" may be replaced with "plural."
[0036] Below, the terms used in the examples are explained.
[0037] Encoder: An encoder can refer to a device that performs encoding. In other words, an encoder can refer to an encoding device.
[0038] Decoder: A decoder can refer to a device that performs decoding. In other words, a decoder can refer to a decoding device.
[0039] Unit: A unit may represent a unit of image encoding and / or decoding. The terms "unit" and "block" may be used interchangeably with the same meaning.
[0040] - A unit can be an MxN array of samples. M and N can each be positive integers. A unit can commonly refer to an array of samples in a two-dimensional form.
[0041] - In the encoding and decoding of an image, a unit may be a region created by the division of a single image. That is to say, a unit may be a specific region within a single image. A single image may be divided into multiple units. Alternatively, a unit may refer to the divided portions when a single image is divided into subdivided parts and encoding or decoding is performed on the divided portions.
[0042] - In the encoding and decoding of images, predefined processing for the unit may be performed depending on the type of unit.
[0043] - Depending on the function, the unit type may be classified into a Macro Unit, Coding Unit (CU), Prediction Unit (PU), Residual Unit, and Transform Unit (TU). Alternatively, depending on the function, the unit may refer to a block, Macroblock, Coding Tree Unit, Coding Tree Block, Coding Unit, Coding Block, Prediction Unit, Prediction Block, Residual Unit, Residual Block, Transform Unit, and Transform Block. For example, the target unit may be at least one of a CU, PU, Residual Unit, and TU that are the target of encoding and / or decoding.
[0044] - To distinguish it from blocks, a unit may refer to information including a lumina component block and a corresponding chroma component block, and a syntax element for each block.
[0045] - The size and shape of the unit may vary. In addition, the unit may have various sizes and shapes. In particular, the shape of the unit may include not only squares but also geometric shapes that can be represented in two dimensions, such as rectangles, trapezoids, triangles, and pentagons.
[0046] - In addition, unit information may include at least one of the unit type, unit size, unit depth, unit encoding order and unit decoding order. For example, the unit type may refer to one of CU, PU, residual unit and TU.
[0047] - A single unit can be further divided into sub-units that have a smaller size relative to the unit.
[0048] Depth: Depth can refer to the degree of subdivision of a unit. Additionally, the depth of a unit can indicate the level at which the unit exists when the unit(s) are represented as a tree structure.
[0049] - Unit division information may include depth regarding the depth of the unit. The depth may indicate the number and / or degree to which the unit is divided.
[0050] - In a tree structure, the root node can be considered to have the shallowest depth, and the leaf nodes the deepest depth. The root node can be the highest node. The leaf nodes can be the lowest nodes.
[0051] - A single unit can be hierarchically divided into multiple sub-units based on a tree structure and possess depth information. That is to say, the unit and the sub-units generated by the division of the unit can correspond to a node and a child node of the node, respectively. Each divided sub-unit can have depth. Since depth indicates the number and / or degree of division of the unit, the division information of the sub-unit may include information regarding the size of the sub-unit.
[0052] - In a tree structure, the topmost node can correspond to the first undivided unit. The topmost node can be referred to as the root node. Additionally, the topmost node can have a minimum depth value. In this case, the topmost node can have a depth of level 0.
[0053] - A node with a depth of level 1 can represent a unit created as the initial unit is divided once. A node with a depth of level 2 can represent a unit created as the initial unit is divided twice.
[0054] - A node with a depth of level n can represent a unit created as the initial unit is divided n times.
[0055] - A leaf node can be the lowest node and can be a node that cannot be further divided. The depth of a leaf node can be the maximum level. For example, the predefined value for the maximum level can be 3.
[0056] - QT depth can represent the depth for quad partitioning. BT depth can represent the depth for binary partitioning. TT depth can represent the depth for ternary partitioning.
[0057] Sample: A sample can be a base unit that constitutes a block. Samples range from 0 to 2 depending on the bit depth (Bd). Bd It can be expressed as values up to -1.
[0058] - A sample can be a pixel or a pixel value.
[0059] - In the following, the terms "pixel," "pixel," and "sample" may be used interchangeably with the same meaning.
[0060] Coding Tree Unit (CTU): A CTU may consist of a single luminance component (Y) coding tree block and two chroma component (Cb, Cr) coding tree blocks associated with the said luminance component coding tree block. Additionally, a CTU may refer to the said blocks and the syntax elements for each of the said blocks.
[0061] - Each coding tree unit may be partitioned using one or more partitioning methods, such as a Quad Tree (QT), Binary Tree (BT), and Ternary Tree (TT), to form sub-units such as coding units, prediction units, and transformation units. A Quad Tree may refer to a quaternary tree. Additionally, each coding tree unit may be partitioned using a MultiType Tree (MTT) that uses one or more partitioning methods.
[0062] - CTU can be used as a term to refer to a pixel block, which is a processing unit in the decoding and encoding processes of an image, as in the segmentation of an input image.
[0063] Coding Tree Block (CTB): Coding Tree Block may be used as a term to refer to any one of the Y Coding Tree Block, Cb Coding Tree Block, and Cr Coding Tree Block.
[0064] Neighbor block: A neighbor block may refer to a block adjacent to the target block. A neighbor block may also refer to a reconstructed neighbor block.
[0065] - In the following, the terms "neighbor block" and "adjacent block" may be used interchangeably with the same meaning.
[0066] - A neighbor block may also refer to a reconstructed neighbor block.
[0067] Spatial neighbor block: A spatial neighbor block may be a block that is spatially adjacent to the target block. A neighbor block may include a spatial neighbor block.
[0068] - The target block and spatial neighbor blocks can be included within the target picture.
[0069] - Spatial neighbor blocks may refer to blocks whose boundaries meet the target block or blocks located within a predetermined distance from the target block.
[0070] - Spatial neighbor blocks may refer to blocks adjacent to a vertex of the target block. Here, a block adjacent to a vertex of the target block may be a block vertically adjacent to a neighbor block horizontally adjacent to the target block, or a block horizontally adjacent to a neighbor block vertically adjacent to the target block.
[0071] Temporal neighbor block: A temporal neighbor block may be a block that is temporally adjacent to the target block. A neighbor block may include a temporal neighbor block.
[0072] - Temporal neighbor blocks may include co-located blocks (col blocks).
[0073] - A call block may be a block within an already reconstructed co-located picture (col picture). The location of the call block within the co-located picture may correspond to the location of the target block within the target picture. Alternatively, the location of the call block within the co-located picture may be the same as the location of the target block within the target picture. A call picture may be a picture included in a reference picture list.
[0074] - Temporal neighbor blocks may be blocks that are temporally adjacent to the spatial neighbor blocks of the target block.
[0075] Prediction mode: The prediction mode may be information indicating a mode that is encoded and / or decoded for intra-prediction or a mode that is encoded and / or decoded for inter-prediction.
[0076] Prediction unit: A prediction unit may refer to a base unit for predictions such as inter-prediction, intra-prediction, inter-compensation, intra-compensation, and motion compensation.
[0077] - A single prediction unit may be divided into multiple partitions or sub-prediction units of smaller sizes. Multiple partitions may also serve as base units for performing prediction or reward. Partitions generated by the division of a prediction unit may also be prediction units.
[0078] Prediction unit partition: A prediction unit partition can refer to a form in which prediction units are divided.
[0079] Reconstructed neighboring unit: The reconstructed neighboring unit may be a unit that has already been decoded and reconstructed in the neighboring part of the target unit.
[0080] - The reconstructed neighbor unit may be a spatially adjacent unit or a temporally adjacent unit to the target unit.
[0081] - Reconstructed spatial neighbor units may be units within the target picture that have already been reconstructed through encoding and / or decoding.
[0082] - The reconstructed temporal neighbor unit may be a unit within the reference image that has already been reconstructed through encoding and / or decoding. The position of the reconstructed temporal neighbor unit within the reference image may be the same as the position of the target unit within the target picture, or may correspond to the position of the target unit within the target picture. Alternatively, the reconstructed temporal neighbor unit may be a neighbor block of a corresponding block within the reference image. Here, the position of the corresponding block within the reference image may correspond to the position of the target block within the target image. Here, the correspondence of the block positions may mean that the positions of the blocks are identical, that one block is contained within another block, or that one block occupies a specific position within another block.
[0083] Sub-picture: A picture can be divided into one or more sub-pictures. A sub-picture can consist of one or more rows of tiles and one or more columns of tiles.
[0084] - A sub-picture may be an area within the picture that is square or rectangular (i.e., non-square). Additionally, a sub-picture may include one or more CTUs.
[0085] - A single sub-picture may include one or more tiles, one or more bricks and / or one or more slices.
[0086] Tile: A tile can be a square or rectangular (i.e., non-square) area within the picture.
[0087] - A tile can contain one or more CTUs.
[0088] - A tile can be divided into one or more bricks.
[0089] Brick: A brick can refer to one or more CTU rows within a tile.
[0090] - A tile can be divided into one or more bricks. Each brick can contain one or more CTU rows.
[0091] - Tiles that are not divided into two or more can also mean bricks.
[0092] Slice: A slice may include one or more tiles within a picture. Alternatively, a slice may include one or more bricks within a tile.
[0093] Parameter set: The parameter set can correspond to header information within the structure of the bitstream.
[0094] - The parameter set may include at least one of a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), and a Decoding Parameter Set (DPS).
[0095] Information signaled through a parameter set can be applied to pictures referencing the parameter set. For example, information within a VPS can be applied to pictures referencing the VPS. Information within an SPS can be applied to pictures referencing the SPS. Information within a PPS can be applied to pictures referencing the PPS.
[0096] A parameter set can reference a parent parameter set. For example, PPS can reference SPS. SPS can reference VPS.
[0097] - Additionally, the parameter set may include a tile group, slice header information, and tile header information. A tile group may refer to a group containing multiple tiles. Furthermore, the meaning of a tile group may be the same as the meaning of a slice.
[0098] Rate-distortion optimization: An encoding device may use rate-distortion optimization to provide high encoding efficiency by using a combination of the size of the coding unit, the prediction mode, the size of the prediction unit, motion information, and the size of the transform unit.
[0099] - The rate-distortion optimization method can calculate the rate-distortion cost of each combination to select the optimal combination among the above combinations. The rate-distortion cost can be calculated using the formula "D+λ*R". In general, the combination that minimizes the rate-distortion cost calculated by the formula "D+λ*R" can be selected as the optimal combination in the rate-distortion optimization method.
[0100] - D may represent distortion. D may be the mean square error of the squares of the differences between the original transformation coefficients and the reconstructed transformation coefficients within the transformation unit.
[0101] - R can represent a rate. R can represent a bit rate using related context information.
[0102] - λ can represent a Lagrangian multiplier. R can include coding parameter information such as prediction mode, motion information, and coded block flag, as well as bits generated by the encoding of transform coefficients.
[0103] - The encoding device can perform processes such as inter-prediction, intra-prediction, transform, quantization, entropy encoding, inverse quantization, and / or inverse transform to calculate accurate D and R. These processes can significantly increase the complexity of the encoding device.
[0104] Bitstream: A bitstream can refer to a sequence of bits containing encoded image information.
[0105] Parameter set: The parameter set can correspond to header information within the structure of the bitstream.
[0106] Parsing: Parsing can refer to determining the value of a syntax element by entropy decoding a bitstream. Alternatively, parsing can refer to entropy decoding itself.
[0107] Symbol: May represent at least one of the syntax elements, coding parameters, and transform coefficients of a unit to be encoded and / or a unit to be decoded. Additionally, Symbol may represent the target of entropy encoding or the result of entropy decoding.
[0108] Reference picture: A reference picture may refer to an image that a unit references for inter prediction or motion compensation. Alternatively, a reference picture may be an image containing a reference unit that a target unit references for inter prediction or motion compensation.
[0109] Hereinafter, the terms "reference picture" and "reference image" may be used interchangeably with the same meaning.
[0110] Reference picture list: The reference picture list may be a list containing one or more reference images used for inter-prediction or motion compensation.
[0111] - The types of reference picture lists may include List Combined (LC), List 0 (L0), List 1 (L1), List 2 (L2), and List 3 (L3).
[0112] - One or more reference picture lists may be used for inter prediction.
[0113] Inter Prediction Indicator: The inter prediction indicator may indicate the direction of inter prediction for the target unit. Inter prediction may be one of unidirectional prediction or bidirectional prediction. Alternatively, the inter prediction indicator may indicate the number of reference pictures used when generating the prediction units of the target unit. Alternatively, the inter prediction indicator may signify the number of prediction blocks used for inter prediction or motion compensation for the target unit.
[0114] Prediction list utilization flag: The prediction list utilization flag may indicate whether to generate a prediction unit using at least one reference picture within a specific reference picture list.
[0115] - An inter-prediction indicator can be derived using the prediction list utilization flag. Conversely, an inter-prediction indicator can be derived using the prediction list utilization flag. For example, if the prediction list utilization flag indicates a first value of 0, it may indicate that for the target unit, a prediction block is not generated using the reference picture in the reference picture list. If the prediction list utilization flag indicates a second value of 1, it may indicate that for the target unit, a prediction unit is generated using the reference picture list.
[0116] Reference picture index: The reference picture index can be an index that points to a specific reference picture in the reference picture list.
[0117] Picture Order Count (POC): The POC of a picture can indicate the display order of the picture.
[0118] Motion Vector (MV): A motion vector can be a two-dimensional vector used in inter-prediction or motion compensation. A motion vector can represent an offset between a target image and a reference image.
[0119] - For example, MV is (mv x , mv y It can be expressed in the form of ). mv x can represent a horizontal component, and mv y It can represent a vertical component.
[0120] Search range: The search range may be a two-dimensional area where a search for MV is performed during inter-prediction. For example, the size of the search range may be MxN. M and N may each be positive integers.
[0121] Motion vector candidate: A motion vector candidate can refer to a block that is a prediction candidate or the motion vector of a prediction candidate block when predicting a motion vector.
[0122] - Motion vector candidates can be included in the motion vector candidate list.
[0123] Motion vector candidate list: A motion vector candidate list can refer to a list composed of one or more motion vector candidates.
[0124] Motion vector candidate index: A motion vector candidate index may refer to an indicator pointing to a motion vector candidate within a motion vector candidate list. Alternatively, a motion vector candidate index may be an index of a motion vector predictor.
[0125] Motion information: Motion information may refer to information including at least one of a motion vector, a reference picture index, and an inter prediction indicator, as well as a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index.
[0126] Merge candidate list: A merge candidate list can refer to a list composed of one or more merge candidates.
[0127] Merge candidate: A merge candidate may refer to a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-prediction merge candidate, a history-based candidate, a candidate based on the average of two candidates, and a zero-merge candidate. A merge candidate may include an inter-prediction indicator and may include motion information such as a reference picture index for each list, motion vectors, a prediction list utilization flag, and an inter-prediction indicator.
[0128] Merge index: A merge index can be an indicator pointing to a merge candidate within a merge candidate list.
[0129] - The merge index can indicate a rebuilt unit that induced a merge candidate among rebuilt units spatially adjacent to the target unit and rebuilt units temporally adjacent to the target unit.
[0130] - The merge index can indicate at least one of the movement information of the merge candidate.
[0131] Transform unit: The transform unit may be a basic unit in residual signal encoding and / or residual signal decoding, such as transform, inverse transform, quantization, inverse quantization, transform factor encoding and transform factor decoding. A single transform unit may be divided into a plurality of sub-transform units having a smaller size. Here, the transform may include one or more of a first-order transform and a second-order transform, and the inverse transform may include one or more of a first-order inverse transform and a second-order inverse transform.
[0132] Scaling: Scaling can refer to the process of multiplying a factor by the transformation factor level.
[0133] - As a result of scaling over the transformation coefficient levels, transformation coefficients can be generated. Scaling may also be referred to as dequantization.
[0134] Quantization Parameter (QP): A quantization parameter may refer to a value used to generate transform coefficient levels for transform coefficients in quantization. Alternatively, a quantization parameter may refer to a value used to generate transform coefficients by scaling transform coefficient levels in inverse quantization. Alternatively, a quantization parameter may be a value mapped to a quantization step size.
[0135] Delta quantization parameter: The delta quantization parameter may represent the difference between the predicted quantization parameter and the quantization parameter of the target unit.
[0136] Scan: Scan can refer to a method of arranging the order of coefficients within a unit, block, or matrix. For example, arranging a two-dimensional array into a one-dimensional array form can be called a scan. Alternatively, arranging a one-dimensional array into a two-dimensional array form can also be called a scan or an inverse scan.
[0137] Transform coefficient: The transform coefficient may be a coefficient value generated as a transform is performed in the encoding device. Alternatively, the transform coefficient may be a coefficient value generated as at least one of entropy decoding and inverse quantization is performed in the decoder.
[0138] - Quantized levels or quantized transformation coefficient levels generated by applying quantization to transformation coefficients or residual signals may also be included in the meaning of transformation coefficients.
[0139] Quantized level: A quantized level may refer to a value generated by performing quantization on a transform coefficient or residual signal in an encoding device. Alternatively, a quantized level may refer to a value that is the subject of inverse quantization when performing inverse quantization in a decoder.
[0140] - Quantized transformation coefficient levels resulting from transformation and quantization can also be included in the meaning of quantized levels.
[0141] Non-zero transform coefficient: A non-zero transform coefficient may refer to a transform coefficient having a non-zero value or a transform coefficient level having a non-zero value. Alternatively, a non-zero transform coefficient may refer to a transform coefficient having a non-zero magnitude or a transform coefficient level having a non-zero magnitude.
[0142] Quantization matrix: A quantization matrix refers to a matrix used in the quantization or inverse quantization process to improve the subjective or objective quality of an image. A quantization matrix can also be referred to as a scaling list.
[0143] Quantization matrix coefficient: A quantization matrix coefficient can refer to each element within the quantization matrix. Quantization matrix coefficients can also be referred to as matrix coefficients.
[0144] Default matrix: The default matrix may be a predefined quantization matrix in the encoding and decoding devices.
[0145] Non-default matrix: A non-default matrix may be a quantization matrix that is not predefined in the encoding and decoding devices. A non-default matrix may refer to a quantization matrix signaled from the encoding device to the decoding device by the user.
[0146] Most Probable Mode (MPM): MPM can represent the intra prediction mode most likely to be used for intra prediction of the target block.
[0147] - The encoding device and the decoding device can determine one or more MPMs based on coding parameters related to the target block and attributes of objects related to the target block.
[0148] - The encoding device and the decoding device can determine one or more MPMs based on the intra prediction mode of the reference block. There may be multiple reference blocks. Multiple reference blocks may include spatial neighbor blocks adjacent to the left of the target block and spatial neighbor blocks adjacent to the top of the target block. That is to say, one or more different MPMs may be determined depending on which intra prediction modes are used for the reference blocks.
[0149] - One or more MPMs can be determined in the same way in the encoding device and the decoder. That is to say, the encoding device and the decoder can share an MPM list containing the same one or more MPMs.
[0150] MPM List: An MPM list may be a list containing one or more MPMs. The number of one or more MPMs in an MPM list may be predefined.
[0151] MPM Indicator: The MPM Indicator may indicate one or more MPMs in the MPM list that are used for intra prediction of the target block. For example, the MPM Indicator may be an index to the MPM list.
[0152] Since the MPM list is determined in the same way in the encoding and decoding units, the MPM list itself may not need to be transmitted from the encoding unit to the decoding unit.
[0153] - An MPM indicator can be signaled from an encoder to a decoder. As the MPM indicator is signaled, the decoder can determine which MPM among the MPMs in the MPM list will be used for intra prediction of the target block.
[0154] MPM Usage Indicator: The MPM Usage Indicator can indicate whether the MPM usage mode is used for prediction of a target block. The MPM usage mode may be a mode that uses an MPM list to determine the MPM to be used for intra prediction of a target block.
[0155] - The MPM usage indicator can be signaled from the encoding device to the decoder.
[0156] Signaling: Signaling may indicate the transmission of information from an encoding device to a decoder. Alternatively, signaling may mean the inclusion of information within a bitstream or a recording medium. Information signaled by an encoding device may be used by a decoder.
[0157] - An encoding device can generate encoded information by performing encoding on signaling information. The encoded information can be transmitted from the encoding device to a decoder. The decoder can obtain information by performing decoding on the transmitted encoded information. Here, encoding may be entropy encoding, and decoding may be entropy decoding.
[0158] Statistical value: Variables, coding parameters, constants, etc., may have values that can be computed. A statistical value may be a value generated by an operation on the values of these specified objects. For example, a statistical value may be one or more of the average value, weighted average value, weighted sum, minimum value, maximum value, mode, median value, and interpolation value for the values of the specified variables, specified coding parameters, and specified constants.
[0159] FIG. 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.
[0160] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. The video may include one or more images. The encoding device (100) may sequentially encode one or more images of the video.
[0161] Referring to FIG. 1, the encoding device (100) may include an inter prediction unit (110), an intra prediction unit (120), a switch (115), a subtractor (125), a converter (130), a quantizer (140), an entropy encoding unit (150), an inverse quantizer (160), an inverse converter (170), an adder (175), a filter unit (180), and a reference picture buffer (190).
[0162] The encoding device (100) can perform encoding for a target image using an intra mode and / or an inter mode. That is to say, the prediction mode for the target block can be one of the intra mode and the inter mode.
[0163] In the following, the terms "intra mode," "intra prediction mode," "in-screen mode," and "in-screen prediction mode" may be used interchangeably with the same meaning.
[0164] In the following, the terms "inter mode," "inter prediction mode," "inter-frame mode," and "inter-frame prediction mode" may be used interchangeably with the same meaning.
[0165] In the following, the term "image" may refer merely to a part of an image and may refer to a block. Additionally, processing of an "image" may refer to sequential processing of multiple blocks.
[0166] Additionally, the encoding device (100) can generate a bitstream containing encoded information through encoding of a target image, and can output and store the generated bitstream. The generated bitstream can be stored on a computer-readable recording medium and can be streamed via a wired and / or wireless transmission medium.
[0167] As a prediction mode, if the intra mode is used, the switch (115) can be switched to intra. As a prediction mode, if the inter mode is used, the switch (115) can be switched to inter.
[0168] The encoding device (100) can generate a prediction block for a target block. Additionally, after the prediction block is generated, the encoding device (100) can encode a residual block for a target block using the residuals of the target block and the prediction block.
[0169] When the prediction mode is an intra mode, the intra prediction unit (120) may use pixels of already encoded and / or decoded blocks located in the neighborhood of the target block as reference samples. The intra prediction unit (120) may perform spatial prediction for the target block using the reference samples and generate prediction samples for the target block through spatial prediction. A prediction sample may refer to a sample within a prediction block.
[0170] The inter prediction unit (110) may include a motion prediction unit and a motion compensation unit.
[0171] When the prediction mode is an inter mode, the motion prediction unit can search for the region that best matches the target block from the reference image during the motion prediction process, and can derive motion vectors for the target block and the searched region using the searched region. In this case, the motion prediction unit can use a search region as the region to be searched.
[0172] A reference image may be stored in a reference picture buffer (190), and when encoding and / or decoding of the reference image is processed, the encoded and / or decoded reference image may be stored in the reference picture buffer (190).
[0173] As the decoded picture is stored, the reference picture buffer (190) may be a decoded picture buffer (DPB).
[0174] The motion compensation unit can generate a prediction block for a target block by performing motion compensation using a motion vector. Here, the motion vector may be a 2D vector used for inter-prediction. Additionally, the motion vector may represent an offset between a target image and a reference image.
[0175] The motion prediction unit and the motion compensation unit can generate a prediction block by applying an interpolation filter to a portion of the reference image when the motion vector has a non-integer value. To perform inter prediction or motion compensation, based on the CU, it can be determined whether the motion prediction and motion compensation method of the PU included in the CU is a skip mode, merge mode, advanced motion vector prediction (AMVP) mode, or current picture reference mode, and inter prediction or motion compensation can be performed according to each mode.
[0176] The subtractor (125) can generate a residual block which is the difference between the target block and the prediction block. The residual block may also be referred to as a residual signal.
[0177] A residual signal can refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or both transforming and quantizing the difference between the original signal and the predicted signal. A residual block may be a residual signal for a block unit.
[0178] The transformation unit (130) can generate transformation coefficients by performing a transformation on the residual block and output the generated transformation coefficients. Here, the transformation coefficients may be coefficient values generated by performing a transformation on the residual block.
[0179] The conversion unit (130) may use one of a plurality of predefined conversion methods when performing the conversion.
[0180] Multiple predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and Karhunen-Loeve Transform (KLT)-based transformations.
[0181] A transformation method used for transformation of a residual block may be determined according to at least one of coding parameters for a target block and / or neighbor blocks. For example, the transformation method may be determined based on at least one of an inter-prediction mode for a PU, an intra-prediction mode for a PU, the size of a TU, and the shape of a TU. Alternatively, transformation information indicating the transformation method may be signaled from an encoding device (100) to a decoding device (200).
[0182] When the transform skip mode is applied, the transform unit (130) may skip the transformation for the residual block.
[0183] By applying quantization to the transform coefficient, a quantized transform coefficient level or a quantized level can be generated. In the following embodiments, the quantized transform coefficient level and the quantized level may also be referred to as transform coefficients.
[0184] The quantization unit (140) can generate a quantized transform coefficient level (i.e., a quantized level or a quantized coefficient) by quantizing the transform coefficients according to the quantization parameters. The quantization unit (140) can output the generated quantized transform coefficient level. At this time, the quantization unit (140) can quantize the transform coefficients using a quantization matrix.
[0185] The entropy encoding unit (150) can generate a bitstream by performing entropy encoding according to a probability distribution based on values calculated in the quantization unit (140) and / or coding parameter values calculated during the encoding process. The entropy encoding unit (150) can output the generated bitstream.
[0186] The entropy encoding unit (150) can perform entropy encoding for information regarding pixels of an image and information for decoding an image. For example, the information for decoding an image may include syntax elements, etc.
[0187] When entropy coding is applied, a small number of bits can be allocated to symbols with a high probability of occurrence, and a large number of bits can be allocated to symbols with a low probability of occurrence. As symbols are represented through this allocation, the size of the bitstring for the symbols being encoded can be reduced. Therefore, the compression performance of video encoding can be improved through entropy coding.
[0188] Additionally, the entropy encoding unit (150) may use encoding methods such as exponential golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) for entropy encoding. For example, the entropy encoding unit (150) may perform entropy encoding using a Variable Length Coding (VLC) table. For example, the entropy encoding unit (150) may derive a binarization method for a target symbol. Additionally, the entropy encoding unit (150) may derive a probability model of a target symbol / bin. The entropy encoding unit (150) may also perform arithmetic encoding using the derived binarization method, probability model and context model.
[0189] The entropy encoding unit (150) can change the coefficients of a two-dimensional block form into a one-dimensional vector form through a transform coefficient scanning method to encode quantized transform coefficient levels.
[0190] Coding parameters may be information required for encoding and / or decoding. Coding parameters may include information encoded in an encoding device (100) and transmitted from the encoding device (100) to a decoding device, and may include information that can be derived during the encoding or decoding process. For example, as information transmitted to the decoding device, there are syntax elements.
[0191] Coding parameters may include information (or flags and indices, etc.) that is encoded in an encoding device, such as syntax elements, and signaled from the encoding device to a decoder, as well as information derived during the encoding or decoding process. Additionally, coding parameters may include information required for encoding or decoding an image. For example, the size of the unit / block, the shape of the unit / block, the depth of the unit / block, the partitioning information of the unit / block, the partitioning structure of the unit / block, information indicating whether the unit / block is partitioned in a quad tree form, information indicating whether the unit / block is partitioned in a binary tree form, the partitioning direction of the binary tree form (horizontal or vertical), the partitioning type of the binary tree form (symmetrical or asymmetrical), information indicating whether the unit / block is partitioned in a ternary tree form, the partitioning direction of the ternary tree form (horizontal or vertical), the partitioning type of the ternary tree form (symmetrical or asymmetrical partitioning, etc.), information indicating whether the unit / block is partitioned in a multi-type tree form, the combination and direction of the multi-type tree partitioning (horizontal or vertical, etc.), the partitioning type of the multi-type tree partitioning (symmetrical or asymmetrical partitioning), the partitioning tree of the multi-type tree form (binary tree or ternary tree), the type of prediction mode (intra-prediction or inter-prediction), the intra-prediction mode / direction, the intra-luminary prediction mode / direction, the intra-chroma prediction mode / direction, the intra-partitioning information, Inter-segment information, coding block segmentation flag, predict block segmentation flag, transform block segmentation flag, reference sample filtering method, reference sample filter tab, reference sample filter coefficients, predict block filtering method, predict block filter tab, predict block filter coefficients, predict block boundary filtering method, predict block boundary filter tab, predict block boundary filter coefficients, inter-prediction mode, motion information, motion vector, motion vector difference, reference picture index,Inter-prediction direction, Inter-prediction indicator, Prediction list utilization flag, Reference picture list, Reference image, POC, Motion vector predictor, Motion vector prediction index, Motion vector prediction candidate, Motion vector candidate list, Information indicating whether merge mode is used, Merge index, Merge candidate, Merge candidate list, Information indicating whether skip mode is used, Type of interpolation filter, Filter tab of interpolation filter, Filter coefficient of interpolation filter, Motion vector magnitude, Motion vector representation accuracy, Transform type, Transform size, Information indicating whether first-order transformation is used, Information indicating whether additional (secondary) transformation is used, First-order transformation selection information (or, First-order transformation index), Second-order transformation selection information (or, Second-order transformation index), Information indicating the presence or absence of residual signals, Coded block pattern, Coded block flag, Quantization parameter, Residual quantization parameter, Quantization matrix, Information about intra-loop filter, Information indicating whether intra-loop filter is applied, Intra-loop filter coefficients, Intra-loop filter tab, shape / form of the intra-loop filter, information indicating whether to apply a deblocking filter, coefficients of the deblocking filter, filter tab of the deblocking filter, strength of the deblocking filter, shape / form of the deblocking filter, information indicating whether to apply an adaptive sample offset, adaptive sample offset value, adaptive sample offset category, type of adaptive sample offset, information indicating whether to apply an adaptive in-loop filter, coefficients of the adaptive in-loop filter, filter tab of the adaptive in-loop filter, shape / form of the adaptive in-loop filter, binarization / debinarization method, context model, method for determining the context model, method for updating the context model, information indicating whether to perform regular mode, information indicating whether to perform bypass mode,Significant factor flag, last significant factor flag, factor group unit coding flag, last significant factor position, flag indicating whether the factor value is greater than 1, flag indicating whether the factor value is greater than 2, flag indicating whether the factor value is greater than 3, remaining factor value information, sign information, reconstructed luminance sample, reconstructed chroma sample, context bin, bypass bin, residual luminance sample, residual chroma sample, transform factor, luminance transform factor, chroma transform factor, quantized level, luminance quantized level, chroma quantized level, transform factor level, luminance transform factor level, chroma transform factor level, transform factor level scanning method, size of motion vector search area on the side of the decoder, shape of motion vector search area on the side of the decoder, number of motion vector searches on the side of the decoder, CTU size, minimum block size, maximum block size, maximum block depth, minimum block depth, display / output order of the image, slice identification information, slice type, slice splitting information, tile group identification information, tile group type, tile group Partition information, tile identification information, tile type, tile partitioning information, picture type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform factor bit depth, quantized level bit depth, information regarding the luminance signal and information regarding the chroma signal, at least one value among the color space of the target block and the color space of the residual block, combined form or statistics may be included in the coding parameters. Additionally, information related to the aforementioned coding parameters may also be included in the coding parameters. Information used to calculate and / or derive the aforementioned coding parameters may also be included in the coding parameters. Information calculated or derived using the aforementioned coding parameters may also be included in the coding parameters.
[0192] The prediction method can represent one of the prediction modes: intra prediction mode and inter prediction mode.
[0193] The first transformation selection information can indicate the first transformation applied to the target block.
[0194] The secondary transformation selection information can indicate the secondary transformation applied to the target block.
[0195] A residual signal may represent the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. A residual block may be a residual signal for a block.
[0196] Here, signaling information may mean that the encoding device (100) includes entropy-encoded information generated by performing entropy encoding on a flag or index into a bitstream, and the decoding device (200) obtains information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. Here, the information may include flags and indices, etc.
[0197] The bitstream may contain information according to a specified syntax. The encoding device (100) may generate a bitstream containing information according to a specified syntax. The encoding device (200) may obtain information from the bitstream according to a specified syntax.
[0198] Since encoding is performed through inter-prediction by the encoding device (100), the encoded target image can be used as a reference image for other image(s) that are subsequently processed. Accordingly, the encoding device (100) can reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference picture buffer (190). Inverse quantization and inverse transform of the encoded target image can be processed for decoding.
[0199] The quantized level can be inversely quantized in the inverse quantization unit (160) and inversely transformed in the inverse transformation unit (170). The inverse quantization unit (160) can generate inversely quantized coefficients by performing inverse quantization on the quantized level. The inverse transformation unit (170) can generate inversely quantized and inversely transformed coefficients by performing inverse transformation on the inversely quantized coefficient.
[0200] The inverse quantized and inverse-transformed coefficients can be combined with the prediction block through an adder (175). By combining the inverse quantized and inverse-transformed coefficients with the prediction block, a reconstructed block can be generated. Here, the inverse quantized and / or inverse-transformed coefficients may refer to coefficients for which at least one of inverse quantization and inverse transformation has been performed, and may refer to a reconstructed residual block. Here, the reconstructed block may refer to a recovered block or a decoded block.
[0201] The reconstructed block may pass through a filter section (180). The filter section (180) may apply at least one of a deblocking filter, a Sample Adaptive Offset (SAO), an Adaptive Loop Filter (ALF), and a Non Local Filter (NLF) to the reconstructed sample, the reconstructed block, or the reconstructed picture. The filter section (180) may also be referred to as an in-loop filter.
[0202] A deblocking filter can remove block distortion that occurs at the boundaries between blocks. To determine whether to apply a deblocking filter, the decision to apply the filter to a target block may be made based on the pixel(s) contained in a number of columns or rows included in the block.
[0203] When applying a deblocking filter to a target block, the applied filter may vary depending on the required strength of deblocking filtering. In other words, among different filters, a filter determined by the strength of deblocking filtering may be applied to the target block. When a deblocking filter is applied to a target block, either a strong filter or a weak filter may be applied, depending on the required strength of deblocking filtering.
[0204] In addition, if vertical filtering and horizontal filtering are performed on the target block, horizontal filtering and vertical filtering can be processed in parallel.
[0205] SAO can add an appropriate offset to the pixel value of a pixel to compensate for coding errors. For a deblocked image, SAO can perform correction using an offset on the difference between the original image and the deblocked image in pixel units. To perform offset correction on an image, a method may be used in which the pixels included in the image are divided into a certain number of regions, the region to which the offset is to be performed is determined among the divided regions, and the offset is applied to the determined region, or a method may be used in which the offset is applied by considering the edge information of each pixel of the image.
[0206] ALF can perform filtering based on a comparison of the reconstructed image and the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each divided group can be determined, and filtering can be performed differentially for each group. Information regarding whether to apply an adaptive loop filter can be signaled per CU. This information can be signaled for the luminance signal. The shape of the ALF and filter coefficients to be applied to each block may differ for each block. Alternatively, a fixed form of ALF may be applied to the block regardless of the characteristics of the block.
[0207] Non-local filters can perform filtering based on reconstructed blocks similar to the target block. Regions similar to the target block in the reconstructed image can be selected, and filtering of the target block can be performed using the statistical properties of the selected similar regions. Information regarding whether to apply non-local filters can be signaled to the CU. Additionally, the shapes and filter coefficients of the non-local filters to be applied to the blocks may differ depending on the block.
[0208] The reconstructed block or reconstructed image that has passed through the filter unit (180) can be stored in the reference picture buffer (190) as a reference picture. The reconstructed block that has passed through the filter unit (180) may be part of the reference picture. That is to say, the reference picture may be a reconstructed picture composed of the reconstructed blocks that have passed through the filter unit (180). The stored reference picture may subsequently be used for inter prediction or motion compensation.
[0209] FIG. 2 is a block diagram showing the configuration according to one embodiment of a decoding device to which the present invention is applied.
[0210] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.
[0211] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), an inter prediction unit (250), a switch (245), an adder (255), a filter unit (260), and a reference picture buffer (270).
[0212] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium and can receive a bitstream stream streamed through a wired / wireless transmission medium.
[0213] The decoding device (200) can perform intra-mode and / or inter-mode decoding on the bitstream. Additionally, the decoding device (200) can generate a reconstructed image or a decoded image through decoding, and can output the generated reconstructed image or decoded image.
[0214] For example, switching to an intra mode or an inter mode according to the prediction mode used for decoding can be done by a switch (245). If the prediction mode used for decoding is an intra mode, the switch (245) can be switched to intra. If the prediction mode used for decoding is an inter mode, the switch (245) can be switched to inter.
[0215] The decoding device (200) can obtain a reconstructed residual block and generate a prediction block by decoding the input bitstream. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by combining the reconstructed residual block and the prediction block.
[0216] The entropy decoding unit (210) can generate symbols by performing entropy decoding on the bitstream based on a probability distribution on the bitstream. The generated symbols may include symbols in the form of quantized transform coefficient levels (i.e., quantized levels or quantized coefficients). Here, the entropy decoding method may be similar to the entropy encoding method described above. For example, the entropy decoding method may be the inverse process of the entropy encoding method described above.
[0217] The entropy decoding unit (210) can convert the coefficients in the form of a one-dimensional vector into the form of a two-dimensional block through a transform coefficient scanning method to decode the quantized transform coefficient levels.
[0218] For example, the coefficients of a block can be changed into a two-dimensional block form by scanning the coefficients of the block using an upper-right diagonal scan. Alternatively, depending on the block size and / or intra-prediction mode, it may be determined which scan to use among an upper-right diagonal scan, a vertical scan, and a horizontal scan.
[0219] Quantized coefficients can be dequantized in the dequantization unit (220). The dequantization unit (220) can generate dequantized coefficients by performing dequantization on the quantized coefficients. Additionally, the dequantized coefficients can be inversely transformed in the inverse transformation unit (230). The inverse transformation unit (230) can generate a reconstructed residual block by performing an inverse transformation on the dequantized coefficients. As a result of performing dequantization and inverse transformation on the quantized coefficients, a reconstructed residual block can be generated. At this time, the dequantization unit (220) can apply a quantization matrix to the quantized coefficients when generating the reconstructed residual block.
[0220] When an intra mode is used, the intra prediction unit (240) can generate a prediction block by performing a spatial prediction on the target block using the pixel values of the already decoded blocks of the neighbors of the target block.
[0221] The inter prediction unit (250) may include a motion compensation unit. Alternatively, the inter prediction unit (250) may be named a motion compensation unit.
[0222] When an inter mode is used, the motion compensation unit can generate a prediction block by performing motion compensation on the target block using a motion vector and a reference image stored in the reference picture buffer (270).
[0223] The motion compensation unit can apply an interpolation filter to a portion of a reference image when the motion vector has a non-integer value, and can generate a prediction block using the reference image with the interpolation filter applied. To perform motion compensation, the motion compensation unit can determine which mode among skip mode, merge mode, AMVP mode, and current picture reference mode the motion compensation method used for the PU included in the CU based on the CU is, and can perform motion compensation according to the determined mode.
[0224] The reconstructed residual block and the prediction block can be added through an adder (255). The adder (255) can generate a reconstructed block by adding the reconstructed residual block and the prediction block.
[0225] The reconstructed block may pass through a filter unit (260). The filter unit (260) may apply at least one of a deblocking filter, SAO, ALF, and a non-local filter to the reconstructed block or the reconstructed image. The reconstructed image may be a picture containing the reconstructed block.
[0226] The filter unit (260) can output the reconstructed image.
[0227] The reconstructed block and / or reconstructed image passed through the filter unit (260) can be stored as a reference picture in the reference picture buffer (270). The reconstructed block passed through the filter unit (260) may be part of the reference picture. That is to say, the reference picture may be a reconstructed image composed of the reconstructed blocks passed through the filter unit (260). The stored reference picture may subsequently be used for inter prediction and / or motion compensation.
[0228] Figure 3 is a diagram schematically showing the segmentation structure of an image when encoding and decoding an image.
[0229] Figure 3 schematically illustrates an example in which a single unit is divided into multiple sub-units.
[0230] To efficiently partition an image, a coding unit (CU) may be used in encoding and decoding. A unit may be a term referring to a combination of 1) a block containing image samples and 2) a syntax element. For example, "partition of a unit" may mean "partition of a block corresponding to a unit."
[0231] A CU can be used as a base unit for video encoding and / or decoding. Additionally, a CU can be used as a unit to which one of an intra mode and an inter mode is applied in video encoding and / or decoding. That is to say, in video encoding and / or decoding, it can be determined which mode, either an intra mode or an inter mode, will be applied to each CU.
[0232] Additionally, CU may be a base unit for prediction, transformation, quantization, inverse transformation, inverse quantization, and encoding and / or decoding of transformation coefficients.
[0233] Referring to FIG. 3, the image (300) can be sequentially divided into units of Largest Coding Units (LCU). For each LCU, a division structure can be determined. Here, LCU can be used with the same meaning as Coding Tree Unit (CTU).
[0234] The division of a unit may refer to the division of a block corresponding to the unit. Block division information may include depth information regarding the depth of the unit. The depth information may indicate the number and / or degree to which the unit is divided. A single unit may be hierarchically divided into multiple sub-units based on a tree structure and depth information.
[0235] Each divided sub-unit may have depth information. The depth information may be information indicating the size of the CU. The depth information may be stored for each CU.
[0236] Each CU can have depth information. When a CU is divided, the CUs created by the division can have a depth increased by 1 from the depth of the divided CU.
[0237] The partitioning structure may refer to a distribution of CUs for efficiently encoding images within the LCU (310). This distribution may be determined by whether to partition a single CU into multiple CUs. The number of partitioned CUs may be positive integers greater than or equal to 2, including 2, 4, 8, and 16.
[0238] The width and height of the CU created by the division may be smaller than the width and height of the CU before the division, depending on the number of CUs created by the division. For example, the width and height of the CU created by the division may be half the width and half the height of the CU before the division.
[0239] A divided CU can be recursively divided into multiple CUs in the same way. Through recursive division, at least one of the width and height of the divided CU can be reduced compared to at least one of the width and height of the CU before division.
[0240] The partitioning of CU can be performed recursively up to a predefined depth or predefined size.
[0241] For example, the depth of the CU can have a value from 0 to 3. The size of the CU can be from 64x64 to 8x8 depending on the depth of the CU.
[0242] For example, the depth of the LCU (310) may be 0, and the depth of the Smallest Coding Unit (SCU) may be a predefined maximum depth. Here, the LCU may be a CU with a maximum coding unit size as described above, and the SCU may be a CU with a minimum coding unit size.
[0243] A division can be started from LCU (310), and the depth of CU can be increased by 1 each time the width and / or height of CU is reduced by the division.
[0244] For example, for each depth, an undivided CU can have a size of 2Nx2N. Also, for a divided CU, a CU of size 2Nx2N can be divided into 4 CUs of size NxN. The size of N can be halved for every 1 increase in depth.
[0245] Referring to FIG. 3, an LCU with a depth of 0 can be 64x64 pixels or 64x64 blocks. 0 can be the minimum depth. An SCU with a depth of 3 can be 8x8 pixels or 8x8 blocks. 3 can be the maximum depth. In this case, the CU of 64x64 blocks, which is the LCU, can be represented as depth 0. The CU of 32x32 blocks can be represented as depth 1. The CU of 16x16 blocks can be represented as depth 2. The CU of 8x8 blocks, which is the SCU, can be represented as depth 3.
[0246] Information regarding whether a CU is divided can be expressed through the division information of the CU. The division information may be 1 bit of information. All CUs except the SCU may include division information. For example, the value of the division information of a CU that is not divided may be a first value, and the value of the division information of a CU that is divided may be a second value. When the division information indicates whether the CU is divided, the first value may be 0 and the second value may be 1.
[0247] For example, if a single CU is divided into four CUs, the width and height of each of the four CUs created by the division may be half the width and half the height of the CU before the division, respectively. If a CU of size 32x32 is divided into four CUs, the sizes of the four divided CUs may be 16x16. When a single CU is divided into four CUs, it can be said that the CU has been divided into a quad-tree form. In other words, it can be seen that a quad-tree partition has been applied to the CU.
[0248] For example, if a single CU is divided into two CUs, the width or height of each of the two CUs created by the division may be half the width or half the height of the CU before the division, respectively. If a 32x32 CU is divided vertically into two CUs, the sizes of the two divided CUs may be 16x32. If a 32x32 CU is divided horizontally into two CUs, the sizes of the two divided CUs may be 32x16. When a single CU is divided into two CUs, it can be said that the CU has been divided into a binary-tree form. In other words, it can be seen that a binary-tree partition has been applied to the CU.
[0249] For example, when a single CU is divided into three CUs, the three divided CUs can be generated by dividing the width or height of the original CU in a ratio of 1:2:1. For example, if a 16x32 CU is divided horizontally into three CUs, the three divided CUs can have sizes of 16x8, 16x16, and 16x8 from top to bottom, respectively. For example, if a 32x32 CU is divided vertically into three CUs, the three divided CUs can have sizes of 8x32, 16x32, and 8x32 from left to right, respectively. When a single CU is divided into three CUs, it can be said that the CU has been divided in a ternary-tree form. In other words, it can be considered that a ternary-tree partition has been applied to the CU.
[0250] In the LCU (310) of Fig. 3, both quad-tree type partitioning and binary-tree type partitioning were applied.
[0251] In the encoding device (100), a 64x64 coding tree unit (CTU) can be divided into a plurality of smaller CUs by a recursive quad-tree structure. One CU can be divided into four CUs of the same size. The CUs can be recursively divided, and each CU can have a quad-tree structure.
[0252] Through recursive partitioning of the CU, the optimal partitioning method that generates the minimum rate-distortion ratio can be selected.
[0253] The CTU (320) in Fig. 3 is an example of a CTU to which quad tree splitting, binary tree splitting and ternary tree splitting are all applied.
[0254] As described above, to partition a CTU, at least one of quad tree partitioning, binary tree partitioning, and ternary tree partitioning may be applied to the CTU. The partitions may be applied based on a specified priority.
[0255] For example, quadtree splitting may be applied preferentially to a CTU. A CU that can no longer be quadtree split may correspond to a leaf node of the quadtree. A CU corresponding to a leaf node of the quadtree may become a root node of a binary tree and / or a ternary tree. That is, a CU corresponding to a leaf node of the quadtree may be split into a binary tree form or a ternary tree form, or may not be split further. In this case, by ensuring that quadtree splitting is not applied again to a CU created by applying binary tree splitting or ternary tree splitting to a CU corresponding to a leaf node of the quadtree, block splitting and / or signaling of block splitting information can be performed effectively.
[0256] The partitioning of a CU corresponding to each node of a quad tree can be signaled using quad partition information. Quad partition information having a first value (e.g., "1") can indicate that the CU is partitioned into a quad tree form. Quad partition information having a second value (e.g., "0") can indicate that the CU is not partitioned into a quad tree form. The quad partition information may be a flag having a specified length (e.g., 1 bit).
[0257] There may be no priority between binary tree splitting and ternary tree splitting. That is, CUs corresponding to the leaf nodes of a quad tree can be split into a binary tree or a ternary tree. Additionally, CUs generated by binary tree splitting or ternary tree splitting can be split again into a binary tree or a ternary tree, or they may not be split any further.
[0258] A partition in which there is no priority between binary tree partitioning and ternary tree partitioning can be referred to as a multi-type tree partition. That is, a CU corresponding to a leaf node of a quad tree can be the root node of a multi-type tree. For the partition of a CU corresponding to each node of a multi-type tree, at least one of information indicating whether the multi-type tree is partitioned, information on the direction of partitioning, and information on the partition tree may be signaled. For the partition of a CU corresponding to each node of a multi-type tree, information indicating whether the partition is partitioned, information on the direction of partitioning, and information on the partition tree may be signaled sequentially.
[0259] For example, information indicating whether a multi-type tree has a first value (e.g., "1") may indicate that the corresponding CU is divided into a multi-type tree form. Information indicating whether a multi-type tree has a second value (e.g., "0") may indicate that the corresponding CU is not divided into a multi-type tree form.
[0260] When a CU corresponding to each node of a multi-type tree is split into a multi-type tree form, the CU may include additional splitting direction information.
[0261] Split direction information can indicate the split direction of a multi-type tree split. Split direction information having a first value (e.g., "1") can indicate that the corresponding CU is split in the vertical direction. Split direction information having a second value (e.g., "0") can indicate that the corresponding CU is split in the horizontal direction.
[0262] When a CU corresponding to each node of a multi-type tree is split into a multi-type tree, that CU may include additional split tree information. The split tree information may indicate the tree used for the multi-type tree split.
[0263] For example, partition tree information having a first value (e.g., "1") may indicate that the corresponding CU is partitioned into a binary tree form. Partition tree information having a second value (e.g., "0") may indicate that the corresponding CU is partitioned into a ternary tree form.
[0264] Here, each of the information indicating whether the aforementioned division exists, the division tree information, and the division direction information may be a flag having a specific length (e.g., 1 bit).
[0265] At least one of the aforementioned quad splitting information, information indicating whether a multi-type tree is split, splitting direction information, and splitting tree information may be entropy encoded and / or entropy decoded. For the entropy encoding / decoding of such information, information of neighboring CUs adjacent to the target CU may be used.
[0266] For example, the partitioning form of the left CU and / or upper CU (i.e., whether it is partitioned, the partition tree and / or the partition direction) and the partitioning form of the target CU may be considered to have a high probability of being similar to each other. Accordingly, context information for entropy encoding and / or entropy decoding of the information of the target CU can be derived based on the information of the neighboring CU. In this case, the information of the neighboring CU may include at least one of 1) quad partitioning information, 2) information indicating whether the multi-type tree is partitioned, 3) partition direction information, and 4) partition tree information of the neighboring CU.
[0267] In another embodiment, among binary tree partitioning and ternary tree partitioning, binary tree partitioning may be performed first. That is, binary tree partitioning is applied first, and a CU corresponding to a leaf node of the binary tree may be set as the root node of the ternary tree. In this case, quad tree partitioning and binary tree partitioning may not be performed for a CU corresponding to a node of the ternary tree.
[0268] A CU that is no longer divided by quad tree division, binary tree division, and / or ternary tree division can be a unit of encoding, prediction, and / or transformation. That is, for prediction and / or transformation, the CU may no longer be divided. Therefore, a division structure and division information, etc., for dividing the CU into prediction units and / or transformation units may not exist within the bitstream.
[0269] However, if the size of the CU that serves as the unit of division is larger than the size of the maximum transformation block, such CU may be recursively divided until the size of the CU becomes less than or equal to the size of the maximum transformation block. For example, if the size of the CU is 64x64 and the size of the maximum transformation block is 32x32, the CU may be divided into 4 32x32 blocks for transformation. For example, if the size of the CU is 32x64 and the size of the maximum transformation block is 32x32, the CU may be divided into 2 32x32 blocks for transformation.
[0270] In such cases, information regarding whether the CU is split for transformation may not be signaled separately. Without signaling, whether the CU is split may be determined by comparing the width (and / or height) of the CU with the width (and / or height) of the maximum transformation block. For example, if the width of the CU is greater than the width of the maximum transformation block, the CU may be split vertically. Additionally, if the height of the CU is greater than the height of the maximum transformation block, the CU may be split horizontally.
[0271] Information regarding the maximum and / or minimum size of a CU, and information regarding the maximum and / or minimum size of a transform block, may be signaled or determined at a higher level for the CU. For example, the higher level may be the sequence level, picture level, tile level, tile group level, and slice level, etc. For example, the minimum size of the CU may be determined to be 4x4. For example, the maximum size of the transform block may be determined to be 64x64. For example, the minimum size of the transform block may be determined to be 4x4.
[0272] Information regarding the minimum size of a CU corresponding to a leaf node of a quad tree (i.e., quad tree minimum size) and / or information regarding the maximum depth of a path from a root node to a leaf node of a multi-type tree (i.e., multi-type tree maximum depth) may be signaled or determined at a higher level for the CU. For example, the higher level may be the sequence level, picture level, slice level, tile group level, and tile level. Information regarding the minimum size of the quad tree and / or information regarding the maximum depth of the multi-type tree may be signaled or determined separately for each of the intra-slice and inter-slice.
[0273] Differential information regarding the size of the CTU and the maximum size of the transformation block can be signaled or determined at a higher level for the CU. For example, the higher level may be the sequence level, picture level, slice level, tile group level, and tile level. Information regarding the maximum size of the CU corresponding to each node of the binary tree (i.e., binary tree maximum size) can be determined based on the size of the CTU and differential information. The maximum size of the CU corresponding to each node of the ternary tree (i.e., ternary tree maximum size) may have different values depending on the type of slice. For example, within an intra-slice, the ternary tree maximum size may be 32x32. Also, for example, within an inter-slice, the ternary tree maximum size may be 128x128. For example, the minimum size of the CU corresponding to each node of the binary tree (i.e., binary tree minimum size) and / or the minimum size of the CU corresponding to each node of the ternary tree (i.e., ternary tree minimum size) can be set as the minimum size of the CU.
[0274] As another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the slice level. Additionally, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the slice level.
[0275] Based on the various block sizes and various depths described above, quad splitting information, information indicating whether a multi-type tree is split, splitting tree information and / or splitting direction information, etc., may or may not exist within the bitstream.
[0276] For example, if the size of the CU is not larger than the minimum size of the quad tree, the CU may not contain quad partition information, and the quad partition information for the CU can be inferred as a second value.
[0277] For example, if the size (width and height) of a CU corresponding to a node of a multi-type tree is larger than the maximum size (width and height) of a binary tree and / or the maximum size (width and height) of a ternary tree, the CU may not be split into a binary tree form and / or a ternary tree form. Depending on this decision method, information indicating whether the multi-type tree is split may not be signaled and may be inferred as a second value.
[0278] Alternatively, if the size (width and height) of a CU corresponding to a node of a multi-type tree is equal to the binary tree minimum size (width and height), or if the size (width and height) of a CU is equal to twice the ternary tree minimum size (width and height), the CU may not be split into a binary tree shape and / or a ternary tree shape. Depending on this decision method, information indicating whether the multi-type tree is split may not be signaled and may be inferred as a second value. This is because if the CU is split into a binary tree shape and / or a ternary tree shape, a CU smaller than the binary tree minimum size and / or the ternary tree minimum size is generated.
[0279] Alternatively, binary tree splitting or ternary tree splitting may be limited based on the size of a virtual pipeline data unit (i.e., pipeline buffer size). For example, binary tree splitting or ternary tree splitting may be limited if, by binary tree splitting or ternary tree splitting, a CU is split into sub-CUs that do not fit the pipeline buffer size. The pipeline buffer size may be equal to the size of the maximum transform block (e.g., 64X64).
[0280] For example, when the pipeline buffer size is 64X64, the following partitions may be limited.
[0281] - Ternary tree partitioning for NxM (N and / or M are 128) CUs
[0282] - Horizontal binary tree partitioning for 128xN (N <= 64) CUs
[0283] - Vertical binary tree partitioning for Nx128 (N <= 64) CUs
[0284] Alternatively, if the depth of the CU within the multi-type tree corresponding to the node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be split into a binary tree form and / or a ternary tree form. Depending on this decision method, information indicating whether the multi-type tree is split may not be signaled and may be inferred as a second value.
[0285] Alternatively, for a CU corresponding to a node of a multi-type tree, information indicating whether the multi-type tree is split may be signaled only if at least one of vertical binary tree splitting, horizontal binary tree splitting, vertical ternary tree splitting, and horizontal ternary tree splitting is possible. Otherwise, the CU may not be split into a binary tree form and / or a ternary tree form. Depending on this decision method, information indicating whether the multi-type tree is split may not be signaled and may be inferred as a second value.
[0286] Alternatively, for a CU corresponding to a node of a multi-type tree, split direction information may be signaled only if both vertical binary tree splitting and horizontal binary tree splitting are possible, or if both vertical ternary tree splitting and horizontal ternary tree splitting are possible. Otherwise, split direction information may not be signaled and may be inferred as a value indicating the direction in which the CU can be split.
[0287] Alternatively, for a CU corresponding to a node of a multi-type tree, split tree information may be signaled only if both vertical binary tree splitting and vertical ternary tree splitting are possible, or if both horizontal binary tree splitting and horizontal ternary tree splitting are possible. Otherwise, split tree information may not be signaled and may be inferred as a value indicating a tree that can be applied to the splitting of the CU.
[0288] Figure 4 is a diagram illustrating the shape of a prediction unit that a coding unit may include.
[0289] Among the CUs split from the LCU, the CUs that are no longer split may be split into one or more Prediction Units (PUs).
[0290] A PU may be a basic unit for prediction. A PU may be encoded and decoded in any one of skip mode, inter mode, and intra mode. A PU may be divided into various forms depending on each mode. For example, the target block described with reference to FIG. 1 and the target block described with reference to FIG. 2 may be a PU.
[0291] A CU may not be divided into PUs. If a CU is not divided into PUs, the size of the CU and the size of the PU may be the same.
[0292] In skip mode, there may be no division within the CU. In skip mode, a 2Nx2N mode (410) in which the sizes of the PU and CU are the same without division may be supported.
[0293] In the inter mode, eight types of divided forms within the CU may be supported. For example, in the inter mode, 2Nx2N mode (410), 2NxN mode (415), Nx2N mode (420), NxN mode (425), 2NxnU mode (430), 2NxnD mode (435), nLx2N mode (440), and nRx2N mode (445) may be supported.
[0294] In intra mode, 2Nx2N mode (410) and NxN mode (425) may be supported.
[0295] In 2Nx2N mode (410), a PU of size 2Nx2N can be encoded. A PU of size 2Nx2N can mean a PU of the same size as the CU. For example, a PU of size 2Nx2N can have a size of 64x64, 32x32, 16x16, or 8x8.
[0296] In NxN mode (425), NxN size PUs can be encoded.
[0297] For example, in intra prediction, when the size of the PU is 8x8, 4 divided PUs can be encoded. The size of the divided PU can be 4x4.
[0298] When a PU is encoded by an intra mode, the PU may be encoded using one of a plurality of intra prediction modes. For example, High Efficiency Video Coding (HEVC) technology may provide 35 intra prediction modes, and the PU may be encoded using one of the 35 intra prediction modes.
[0299] Whether the PU will be encoded by the 2Nx2N mode (410) or the NxN mode (425) can be determined by the rate-distortion cost.
[0300] The encoding device (100) can perform encoding operations on a PU of size 2Nx2N. Here, the encoding operation may involve encoding the PU using each of a plurality of intra prediction modes that the encoding device (100) can use. Through the encoding operation, an optimal intra prediction mode for a PU of size 2Nx2N can be derived. The optimal intra prediction mode may be an intra prediction mode among the plurality of intra prediction modes that the encoding device (100) can use that generates the minimum rate-distortion cost for encoding a PU of size 2Nx2N.
[0301] Additionally, the encoding device (100) can sequentially perform encoding operations on each PU of the NxN divided PUs. Here, the encoding operation may involve encoding the PUs using each of the multiple intra prediction modes available to the encoding device (100). Through the encoding operation, an optimal intra prediction mode for the NxN size PUs may be derived. The optimal intra prediction mode may be an intra prediction mode among the multiple intra prediction modes available to the encoding device (100) that generates the minimum rate-distortion cost for encoding the NxN size PUs.
[0302] The encoding device (100) can determine which of the 2Nx2N size PU and the NxN size PU to encode based on a comparison of the rate-distortion costs of the 2Nx2N size PU and the rate-distortion costs of the NxN size PUs.
[0303] One CU can be divided into one or more PUs, and a PU can also be divided into multiple PUs.
[0304] For example, if a single PU is divided into four PUs, the width and height of each of the four PUs created by the division may be half the width and half the height of the PU before the division, respectively. If a PU of size 32x32 is divided into four PUs, the sizes of the four divided PUs may be 16x16. If a single PU is divided into four PUs, it can be said that the PU has been divided into a quad-tree form.
[0305] For example, if a single PU is divided into two PUs, the width or height of each of the two PUs created by the division may be half the width or half the height of the PU before the division, respectively. If a 32x32 PU is divided vertically into two PUs, the sizes of the two divided PUs may be 16x32. If a 32x32 PU is divided horizontally into two PUs, the sizes of the two divided PUs may be 32x16. If a single PU is divided into two PUs, it can be said that the PU has been divided into a binary-tree form.
[0306] Figure 5 is a drawing illustrating the form of a conversion unit that can be included in a coding unit.
[0307] A Transform Unit (TU) may be a basic unit used for the processes of transformation, quantization, inverse transformation, inverse quantization, entropy encoding, and entropy decoding within a CU.
[0308] TU can have a square or rectangular shape. The shape of TU can be determined depending on the size and / or shape of CU.
[0309] Among the CUs divided from the LCU, the CUs that are no longer divided into CUs may be divided into one or more TUs. In this case, the division structure of the TUs may be a quad-tree structure. For example, as shown in FIG. 5, one CU (510) may be divided once or more times according to the quad-tree structure. Through division, one CU (510) may be composed of TUs of various sizes.
[0310] If a CU is partitioned two or more times, the CU can be viewed as being partitioned recursively. Through partitioning, a CU can be composed of TUs of various sizes.
[0311] Alternatively, a single CU may be divided into one or more TUs based on the number of vertical and / or horizontal lines dividing the CU.
[0312] A CU can be divided into symmetric TUs or into asymmetric TUs. For division into asymmetric TUs, information regarding the size and / or shape of the TUs can be signaled from the encoding device (100) to the decoding device (200). Alternatively, the size and / or shape of the TUs can be derived from information regarding the size and / or shape of the CUs.
[0313] A CU may not be divided into TUs. If a CU is not divided into TUs, the size of the CU and the size of the TU may be the same.
[0314] One CU can be divided into one or more TUs, and a TU can also be divided into multiple TUs.
[0315] For example, if a single TU is divided into four TUs, the width and height of each of the four TUs generated by the division may be half the width and half the height of the TU before the division, respectively. If a TU of size 32x32 is divided into four TUs, the sizes of the four divided TUs may be 16x16. If a single TU is divided into four TUs, it can be said that the TU has been divided into a quad-tree.
[0316] For example, if a single TU is divided into two TUs, the width or height of each of the two TUs generated by the division may be half the width or half the height of the TU before the division, respectively. If a TU of size 32x32 is divided vertically into two TUs, the sizes of the two divided TUs may be 16x32. If a TU of size 32x32 is divided horizontally into two TUs, the sizes of the two divided TUs may be 32x16. If a single TU is divided into two TUs, it can be said that the TU has been divided into a binary-tree form.
[0317] The CU may be divided in a way other than that shown in Fig. 5.
[0318] For example, one CU can be divided into three CUs. The width or height of the three divided CUs can be 1 / 4, 1 / 2, and 1 / 4 of the width or height of the CU before division, respectively.
[0319] For example, if a CU of size 32x32 is vertically divided into 3 CUs, the sizes of the 3 divided CUs can be 8x32, 16x32, and 8x32, respectively. In this way, when a CU is divided into 3 CUs, the CU can be seen as being divided into the form of a ternary tree.
[0320] One of the exemplified quad tree-type partitioning, binary tree-type partitioning, and ternary tree-type partitioning can be applied to partition the CU, and multiple partitioning methods may be combined and used to partition the CU. In this case, the case where multiple partitioning methods are combined and used can be referred to as a composite tree-type partitioning.
[0321] Figure 6 shows the division of a block according to one example.
[0322] In the process of encoding and / or decoding the image, the target block may be divided as shown in FIG. 6. For example, the target block may be a CU.
[0323] For the division of a target block, an indicator representing division information may be signaled from an encoding device (100) to a decoding device (200). The division information may be information indicating how the target block is divided.
[0324] The splitting information may be one or more of the splitting flag (hereinafter referred to as "split_flag"), quad-binary flag (hereinafter referred to as "QB_flag"), quad tree flag (hereinafter referred to as "quadtree_flag"), binary tree flag (hereinafter referred to as "binarytree_flag"), and binary type flag (hereinafter referred to as "Btype_flag").
[0325] split_flag can be a flag indicating whether a block is split. For example, a value of split_flag 1 may indicate that the block is split. A value of split_flag 0 may indicate that the block is not split.
[0326] QB_flag may be a flag indicating whether the block is partitioned into a quadtree or a binary tree. For example, a value of 0 for QB_flag may indicate that the block is partitioned into a quadtree. A value of 1 for QB_flag may indicate that the block is partitioned into a binary tree. Alternatively, a value of 0 for QB_flag may indicate that the block is partitioned into a binary tree. A value of 1 for QB_flag may indicate that the block is partitioned into a quadtree.
[0327] quadtree_flag may be a flag indicating whether a block is divided into a quadtree. For example, a value of 1 for quadtree_flag may indicate that the block is divided into a quadtree. A value of 0 for quadtree_flag may indicate that the block is not divided into a quadtree.
[0328] binarytree_flag may be a flag indicating whether the block is partitioned into a binary tree. For example, a value of binarytree_flag 1 may indicate that the block is partitioned into a binary tree. A value of binarytree_flag 0 may indicate that the block is not partitioned into a binary tree.
[0329] Btype_flag may be a flag indicating whether the block was split vertically or horizontally when the block is split into a binary tree. For example, a value of 0 for Btype_flag may indicate that the block is split horizontally. A value of 1 for Btype_flag may indicate that the block is split vertically. Alternatively, a value of 0 for Btype_flag may indicate that the block was split vertically. A value of 1 for Btype_flag may indicate that the block was split horizontally.
[0330] For example, partitioning information for the block of Fig. 6 can be induced by signaling at least one of quadtree_flag, binarytree_flag and Btype_flag as shown in Table 1 below.
[0331] [Table 1]
[0332]
[0333] For example, splitting information for the block of FIG. 6 can be induced by signaling at least one of split_flag, QB_flag and Btype_flag as shown in Table 2 below.
[0334] [Table 2]
[0335]
[0336] The splitting method may be limited to a quad tree only, or to a binary tree only, depending on the size and / or shape of the block. When such a limitation applies, split_flag may be a flag indicating whether to split into a quad tree form or a flag indicating whether to split into a binary tree form. The size and shape of the block may be derived based on the depth information of the block, and the depth information may be signaled from the encoding device (100) to the decoding device (200).
[0337] If the block size falls within a specified range, only quadtree-shaped partitioning may be possible. For example, the specified range may be defined by at least one of a maximum block size and a minimum block size where only quadtree-shaped partitioning is possible.
[0338] Information indicating the maximum block size and / or minimum block size, which is possible only through a quotate tree-shaped partition, can be signaled from the encoding device (100) to the decoding device (200) via a bitstream. Additionally, this information can be signaled for at least one unit among a video, sequence, picture, parameter, tile group, and slice (or segment).
[0339] Alternatively, the maximum block size and / or minimum block size may be fixed sizes defined in the encoding device (100) and the decoding device (200). For example, if the block size is 64x64 or larger and 256x256 or smaller, only quad tree-shaped splitting may be possible. In this case, split_flag may be a flag indicating whether to split into quad tree shapes.
[0340] If the block size is larger than the maximum transformation block size, only quadtree-type partitioning may be possible. In this case, the partitioned block may be at least one of CU and TU.
[0341] In this case, split_flag may be a flag indicating whether to split into a quad tree.
[0342] If the size of the block falls within a specified range, only binary tree or ternary tree partitioning may be possible. Here, for example, the specified range may be defined by at least one of a maximum block size and a minimum block size that allow only binary tree or ternary tree partitioning.
[0343] Information indicating the maximum block size and / or minimum block size, which is possible only in the form of a binary tree or a ternary tree, can be signaled from the encoding device (100) to the decoding device (200) via a bitstream. Additionally, this information can be signaled for at least one unit among a sequence, a picture, and a slice (or, a segment).
[0344] Alternatively, the maximum block size and / or minimum block size may be fixed sizes defined in the encoding device (100) and the decoding device (200). For example, if the block size is 8x8 or larger and 16x16 or smaller, only binary tree-shaped splitting may be possible. In this case, split_flag may be a flag indicating whether to split into a binary tree or a ternary tree.
[0345] The description regarding the aforementioned code tree-shaped partitioning can be applied equally to binary tree-shaped and / or ternary tree-shaped partitioning.
[0346] The division of a block may be limited by a previous division. For example, if a block is divided into a specific binary tree shape to generate multiple divided blocks, each divided block may be further divided only into the specific tree shape. Here, the specific tree shape may be at least one of a binary tree shape, a ternary tree shape, and a quad tree shape.
[0347] If the width or height of a divided block corresponds to a size that can no longer be divided, the aforementioned indicator may not be signaled.
[0348] The arrows from the center of the graph in Fig. 7 outward may indicate the prediction directions of the directional intra prediction modes. Additionally, the numbers displayed near the arrows may represent an example of a mode value assigned to an intra prediction mode or the prediction direction of the intra prediction mode.
[0349] In FIG. 7, the number 0 may represent the Planar mode, which is a non-directional intra prediction mode. The number 1 may represent the DC mode, which is a non-directional intra prediction mode.
[0350] Intra-coding and / or decoding may be performed using a reference sample of a neighbor unit of the target block. The neighbor block may be a reconstructed neighbor block. The reference sample may refer to a neighbor sample.
[0351] For example, intra-coding and / or decoding can be performed using the value of a reference sample or coding parameter contained in a reconstructed neighbor block.
[0352] The encoding device (100) and / or the decoding device (200) can generate a prediction block by performing intra prediction on a target block based on information of a sample within a target image. When performing intra prediction, the encoding device (100) and / or the decoding device (200) can generate a prediction block for a target block by performing intra prediction based on information of a sample within a target image. When performing intra prediction, the encoding device (100) and / or the decoding device (200) can perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.
[0353] A prediction block may refer to a block generated as a result of performing intra-prediction. A prediction block may correspond to at least one of CU, PU, and TU.
[0354] The unit of the prediction block may be at least one of the sizes of CU, PU, and TU. The prediction block may have a square shape with a size of 2Nx2N or NxN. The NxN size may include 4x4, 8x8, 16x16, 32x32, and 64x64, etc.
[0355] Alternatively, the prediction block may be a square block with dimensions such as 2x2, 4x4, 8x8, 16x16, 32x32, or 64x64, or a rectangular block with dimensions such as 2x8, 4x8, 2x16, 4x16, and 8x16.
[0356] Intra prediction can be performed according to the intra prediction mode for the target block. The number of intra prediction modes that the target block may have can be a predefined fixed value or a value determined differently based on the attributes of the prediction block. For example, the attributes of the prediction block may include the size and type of the prediction block. Additionally, the attributes of the prediction block may refer to coding parameters for the prediction block.
[0357] For example, the number of intra prediction modes can be fixed at N regardless of the size of the prediction block. Or, for example, the number of intra prediction modes can be 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95, etc.
[0358] The intra prediction mode can be a non-directional mode or a directional mode.
[0359] For example, the intra prediction mode may include two non-directional modes and 65 directional modes corresponding to numbers 0 to 66 shown in FIG. 7.
[0360] For example, when a specific intra-prediction method is used, the intra-prediction mode may include two non-directional modes and 93 directional modes corresponding to numbers 14 to 80 shown in FIG. 7.
[0361] The two non-directional modes may include DC mode and Planar mode.
[0362] A directional mode can be a prediction mode that has a specific direction or a specific angle. A directional mode may also be referred to as an argular mode.
[0363] An intra-prediction mode can be expressed as at least one of a mode number, a mode value, a mode angle, and a mode direction. That is to say, the terms “(mode) number of the intra-prediction mode,” “(mode) value of the intra-prediction mode,” “(mode) angle of the intra-prediction mode,” and “(mode) direction of the intra-prediction mode” can be used with the same meaning and can be used interchangeably.
[0364] The number of intra prediction modes may be M. M may be 1 or more. That is to say, the number of intra prediction modes may be M, including the number of non-directional modes and the number of directional modes.
[0365] The number of intra prediction modes can be fixed at M regardless of the block size and / or color component. For example, the number of intra prediction modes can be fixed at either 35 or 67 regardless of the block size.
[0366] Alternatively, the number of intra prediction modes may vary depending on the shape, size, and / or type of color component of the block.
[0367] For example, in Fig. 7, the directional prediction modes illustrated by dashed lines can be applied only to predictions for non-square blocks.
[0368] For example, as the block size increases, the number of intra prediction modes may increase. Or, as the block size increases, the number of intra prediction modes may decrease. If the block size is 4x4 or 8x8, the number of intra prediction modes may be 67. If the block size is 16x16, the number of intra prediction modes may be 35. If the block size is 32x32, the number of intra prediction modes may be 19. If the block size is 64x64, the number of intra prediction modes may be 7.
[0369] For example, the number of intra prediction modes may differ depending on whether the color component is a lumina signal or a chroma signal. Alternatively, the number of intra prediction modes in a lumina component block may be greater than the number of intra prediction modes in a chroma component block.
[0370] For example, in the case of a vertical mode with a mode value of 50, prediction can be performed in the vertical direction based on the pixel value of a reference sample. For example, in the case of a horizontal mode with a mode value of 18, prediction can be performed in the horizontal direction based on the pixel value of a reference sample.
[0371] Even in the case of a directional mode other than the aforementioned mode, the encoding device (100) and the decoding device (200) can perform intra-prediction for the target unit using a reference sample according to an angle corresponding to the directional mode.
[0372] An intra-prediction mode located to the right of a vertical mode can be named a vertical-right mode. An intra-prediction mode located at the bottom of a horizontal mode can be named a horizontal-below mode. For example, in FIG. 7, intra-prediction modes with mode values of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 may be vertical-right modes. Intra-prediction modes with mode values of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 may be horizontal-below modes.
[0373] The non-directional mode may include DC mode and planar mode. For example, the mode value of DC mode may be 1. The mode value of planar mode may be 0.
[0374] Directional modes may include angular modes. Among the multiple intra prediction modes, the remaining modes, excluding the DC mode and the planner mode, may be directional modes.
[0375] When the intra prediction mode is DC mode, a prediction block can be generated based on the average of the pixel values of multiple reference samples. For example, the pixel value of the prediction block can be determined based on the average of the pixel values of multiple reference samples.
[0376] The number of intra prediction modes and the mode values of each intra prediction mode described above may be merely exemplary. The number of intra prediction modes and the mode values of each intra prediction mode described above may be defined differently depending on the embodiment, implementation, and / or as necessary.
[0377] A step of examining whether samples included in a reconstructed neighbor block can be used as reference samples for the target block to perform intra prediction for the target block may be performed. If there are samples among the neighbor blocks that cannot be used as reference samples for the target block, a value generated by copying and / or interpolation using at least one sample value among the samples included in the reconstructed neighbor block may be replaced with the sample value of the sample that cannot be used as a reference sample. If the value generated by copying and / or interpolation is replaced with the sample value of the sample, the sample may be used as a reference sample for the target block.
[0378] When intra prediction is used, a filter may be applied to at least one of the reference sample or prediction sample based on at least one of the intra prediction mode and the size of the target block.
[0379] The type of filter applied to at least one of the reference sample or the prediction sample may vary depending on at least one of the intra-prediction mode of the target block, the size of the target block, and the shape of the target block. The type of filter may be classified according to one or more of the length of the filter tap, the value of the filter coefficient, and the filter strength. The length of the filter tap mentioned above may refer to the number of filter taps. Additionally, the number of filter taps may refer to the length of the filter.
[0380] When the intra prediction mode is planner mode, when generating the prediction block of the target block, the sample value of the prediction target sample can be generated using a weighted sum of the top reference sample of the target sample, the left reference sample of the target sample, the upper right reference sample of the target block, and the lower left reference sample of the target block, depending on the position of the prediction target sample within the prediction block.
[0381] When the intra prediction mode is DC mode, the average value of the top reference samples and left reference samples of the target block may be used to generate the prediction block of the target block. Additionally, filtering using the values of the reference samples may be performed on specific rows or specific columns within the target block. The specific rows may be one or more top rows adjacent to the reference sample. The specific columns may be one or more left columns adjacent to the reference sample.
[0382] When the intra prediction mode is a directional mode, a prediction block can be generated using the top reference sample, left reference sample, top-right reference sample, and / or bottom-left reference sample of the target block.
[0383] Real-valued interpolation may be performed to generate the aforementioned prediction samples.
[0384] The intra prediction mode of a target block can be predicted from the intra prediction mode of a neighbor block of the target block, and the information used for prediction can be entropy encoded / decoded.
[0385] For example, if the intra prediction modes of the target block and neighbor blocks are the same, a predefined flag can be used to signal that the intra prediction modes of the target block and neighbor blocks are the same.
[0386] For example, among the intra prediction modes of multiple neighboring blocks, an indicator pointing to an intra prediction mode identical to the intra prediction mode of the target block may be signaled.
[0387] If the intra-prediction modes of the target block and neighboring blocks are different from each other, information of the intra-prediction mode of the target block can be encoded and / or decoded using entropy encoding and / or decoding.
[0388] Figure 8 is a diagram illustrating a reference sample used in the intra-prediction process.
[0389] The reconstructed reference samples used for intra-prediction of the target block may include below-left reference samples, left reference samples, above-left corner reference samples, above reference samples, and above-right reference samples.
[0390] For example, left reference samples may refer to reconstructed reference pixels adjacent to the left side of the target block. Top reference samples may refer to reconstructed reference pixels adjacent to the top side of the target block. Top left corner reference samples may refer to reconstructed reference pixels located at the top left corner of the target block. Additionally, bottom left reference samples may refer to reference samples located at the bottom of the left sample line among samples aligned with the left sample line composed of left reference samples. Top right reference samples may refer to reference samples located to the right of the top pixel line among samples aligned with the top sample line composed of top reference samples.
[0391] When the size of the target block is NxN, the bottom left reference samples, left reference samples, top reference samples, and top right reference samples can each be N.
[0392] A prediction block can be generated through intra-prediction on a target block. The generation of the prediction block may include determining the values of the pixels of the prediction block. The sizes of the target block and the prediction block may be the same.
[0393] The reference samples used for intra-prediction of the target block may vary depending on the intra-prediction mode of the target block. The direction of the intra-prediction mode may indicate the dependency relationship between the reference samples and the pixels of the prediction block. For example, the value of a specific reference sample may be used as the value of one or more specific pixels of the prediction block. In this case, the specific reference sample and the one or more specific pixels of the prediction block may be samples and pixels designated along a straight line of the direction of the intra-prediction mode. That is to say, the value of the specific reference sample may be copied to the value of a pixel located in the reverse direction of the intra-prediction mode. Alternatively, the value of a pixel in the prediction block may be the value of a reference sample located in the direction of the intra-prediction mode relative to the location of the aforementioned pixel.
[0394] For example, if the intra prediction mode of the target block is vertical mode, the top reference samples can be used for intra prediction. When the intra prediction mode is vertical mode, the value of a pixel in the prediction block may be the value of a reference sample located vertically above the position of the pixel mentioned above. Therefore, the top reference samples adjacent to the top of the target block can be used for intra prediction. Additionally, the values of pixels in a row of the prediction block may be identical to the values of the top reference samples.
[0395] For example, if the intra prediction mode of the target block is horizontal mode, left reference samples can be used for intra prediction. When the intra prediction mode is horizontal mode, the value of a pixel in the prediction block may be the value of a reference sample located horizontally to the left of the pixel. Therefore, left reference samples adjacent to the left of the target block can be used for intra prediction. Additionally, the values of pixels in a column of the prediction block may be identical to the values of the left reference samples.
[0396] For example, when the mode value of the intra prediction mode of the target block is 34, at least some of the left reference samples, the top left corner reference sample, and at least some of the top reference samples may be used for intra prediction. When the mode value of the intra prediction mode is 34, the value of the pixel of the prediction block may be the value of the reference sample located diagonally to the top left with respect to the pixel.
[0397] In addition, when an intra prediction mode with a mode value of 52 to 66 is used, at least some of the upper right reference samples may be used for intra prediction.
[0398] In addition, when an intra prediction mode with a mode value of 2 to 17 is used, at least some of the bottom left reference samples may be used for intra prediction.
[0399] In addition, when an intra prediction mode with a mode value of 19 to 49 is used, the top left corner reference sample can be used for intra prediction.
[0400] The reference samples used to determine the pixel value of a single pixel in a prediction block may be one or two or more.
[0401] As described above, the pixel value of a pixel in a prediction block can be determined according to the position of the pixel and the position of a reference sample indicated by the direction of the intra-prediction mode. If the position of the pixel and the position of the reference sample indicated by the direction of the intra-prediction mode are integer positions, the value of a single reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.
[0402] If the position of a reference sample indicated by the pixel position and the direction of the intra-prediction mode is not an integer position, an interpolated reference sample may be generated based on the two reference samples closest to the position of the reference sample. The value of the interpolated reference sample may be used to determine the pixel value of the pixel of the prediction block. That is to say, when the position of a reference sample indicated by the pixel position of the prediction block and the direction of the intra-prediction mode represents the space between two reference samples, an interpolated value may be generated based on the values of the two samples.
[0403] The prediction block generated by the prediction may not be identical to the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.
[0404] In the following, the terms "difference," "error," and "residual" may be used interchangeably with the same meaning.
[0405] For example, in the case of directional intra-prediction, a larger prediction error may occur as the distance between the pixels of a prediction block and a reference sample increases. Due to this prediction error, discontinuities may occur between the prediction block and neighboring blocks generated by the same.
[0406] Filtering of prediction blocks may be used to reduce prediction errors. Filtering may involve adaptively applying a filter to areas within the prediction blocks that are considered to have large prediction errors. For example, areas considered to have large prediction errors may be the boundaries of the prediction blocks. Furthermore, depending on the intra-prediction mode, the areas considered to have large prediction errors within the prediction blocks may differ, and the characteristics of the filters may also vary.
[0407] As illustrated in FIG. 8, at least one of reference lines 0 to 3 may be used for intra-prediction of a target block. Each reference line may represent a reference sample line. The smaller the number of the reference line, the closer the line of reference samples may be to the target block.
[0408] Samples of segments A and F can be obtained through padding using the nearest samples of segments B and E, respectively, instead of being obtained from reconstructed neighbor blocks.
[0409] Index information indicating a reference sample line to be used for intra prediction of a target block may be signaled. The index information may indicate a reference sample line used for intra prediction of a target block among a plurality of reference sample lines. For example, the index information may have a value from 0 to 3.
[0410] If the top boundary of the target block is the boundary of the CTU, only reference sample line 0 is available. Therefore, in this case, index information may not be signaled. If a reference sample line other than reference sample line 0 is used, filtering for the prediction block described below may not be performed.
[0411] In the case of inter-color intra-prediction, a prediction block for a target block of a second color component can be generated based on a corresponding reconstructed block of a first color component.
[0412] For example, the first color component may be a lumina component, and the second color component may be a chroma component.
[0413] For intra-prediction between color components, the parameters of the linear model between the first color component and the second color component can be derived based on a template.
[0414] The template may include a top reference sample and / or a left reference sample of the target block, and may include a top reference sample and / or a left reference sample of the reconstructed block of the first color component corresponding to these reference samples.
[0415] For example, the parameters of a linear model can be derived using 1) the value of a sample of a first color component having the maximum value among the samples in the template, 2) the value of a sample of a second color component corresponding to this sample of the first color component, 3) the value of a sample of a first color component having the minimum value among the samples in the template, and 4) the value of a sample of a second color component corresponding to this sample of the first color component.
[0416] Once the parameters of the linear model are derived, a prediction block for the target block can be generated by applying the corresponding reconstructed block to the linear model.
[0417] Depending on the image format, subsampling may be performed on the surrounding samples of the reconstructed block of the first color component and the corresponding reconstructed blocks. For example, if one sample of the second color component corresponds to four samples of the first color component, one corresponding sample may be calculated by subsampling the four samples of the first color component. When subsampling is performed, the derivation of parameters of the linear model and intra-prediction between color components may be performed based on the subsampled corresponding samples.
[0418] Whether to perform intra-prediction between color components and / or the range of the template can be signaled as an intra-prediction mode.
[0419] The target block can be divided into 2 or 4 sub-blocks in the horizontal and / or vertical directions.
[0420] The partitioned sub-blocks can be reconstructed sequentially. That is, as intra-prediction is performed on the sub-blocks, sub-prediction blocks for the sub-blocks can be generated. Additionally, as inverse quantization and / or inverse transformation is performed on the sub-blocks, sub-residual blocks for the sub-blocks can be generated. Reconstructed sub-blocks can be generated by adding the sub-prediction blocks to the sub-residual blocks. The reconstructed sub-blocks can be used as reference samples for intra-prediction of subsequent sub-blocks.
[0421] A sub-block may be a block containing more than a specified number (e.g., 16 samples). Thus, for example, if the target block is an 8x4 block or a 4x8 block, the target block may be divided into two sub-blocks. Also, if the target block is a 4x4 block, the target block cannot be divided into sub-blocks. If the target block has any other size, the target block may be divided into four sub-blocks.
[0422] Information regarding whether intra prediction based on these sub-blocks is performed and / or the partitioning direction (horizontal or vertical) may be signaled.
[0423] Such sub-block-based intra prediction may be restricted to be performed only when using reference sample line 0. When sub-block-based intra prediction is performed, filtering for the prediction block described below may not be performed.
[0424] A final prediction block can be generated by performing filtering on the prediction block generated by intra prediction.
[0425] Filtering can be performed by applying specific weights to the filtering target sample, left reference sample, top reference sample, and / or top-left reference sample.
[0426] The weights and / or reference samples used for filtering (or, range of reference samples or location of reference samples, etc.) may be determined based on at least one of the block size, intra-prediction mode, and location of the sample to be filtered within the prediction block.
[0427] For example, filtering may be performed only on specific intra-prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode and / or adjacent diagonal mode).
[0428] An adjacent diagonal mode may be a mode with a number obtained by adding k to the number of the diagonal mode, or a mode with a number obtained by subtracting k from the number of the diagonal mode. That is to say, the number of the adjacent diagonal mode may be the sum of the number of the diagonal mode and k, or the difference between the number of the diagonal mode and k. For example, k may be a positive integer less than or equal to 8.
[0429] The intra prediction mode of a target block can be derived using the intra prediction mode of neighboring blocks existing around the target block, and this derived intra prediction mode can be entropy encoded and / or entropy decoded.
[0430] For example, if the intra prediction mode of the target block and the intra prediction mode of the neighbor block are the same, information that the intra prediction mode of the target block and the intra prediction mode of the neighbor block are the same can be signaled using specific flag information.
[0431] In addition, for example, indicator information for a neighbor block having the same intra prediction mode as the target block among the intra prediction modes of multiple neighbor blocks can be signaled.
[0432] For example, if the intra prediction mode of the target block and the intra prediction mode of the neighbor block are different from each other, entropy encoding and / or entropy decoding of information about the intra prediction mode of the target block can be performed by performing entropy encoding and / or entropy decoding based on the intra prediction mode of the neighbor block.
[0433] Figure 9 is a diagram illustrating an example of an inter prediction process.
[0434] The rectangles shown in FIG. 9 may represent images (or pictures). Additionally, the arrows in FIG. 9 may represent prediction directions. An arrow from the first picture to the second picture may indicate that the second picture refers to the first picture. That is, images may be encoded and / or decoded according to the prediction direction.
[0435] Each image can be classified into I-pictures (Intra Picture), P-pictures (Uni-prediction Picture), and B-pictures (Bi-prediction Picture) according to the encoding type. Each picture can be encoded and / or decoded according to the encoding type of each picture.
[0436] If the target image to be encoded is an I-picture, the target image can be encoded using data within the image itself without inter-prediction referencing other images. For example, an I-picture can be encoded solely by intra-prediction.
[0437] If the target image is a P picture, the target image can be encoded through inter-prediction using only the reference picture existing in a unidirectional direction. Here, the unidirectional direction can be forward or reverse.
[0438] When the target image is a B picture, the target image can be encoded through inter-prediction using reference pictures existing in both directions, or through inter-prediction using a reference picture existing in one of the forward and reverse directions. Here, both directions can be the forward and reverse directions.
[0439] P picture and B picture, which are encoded and / or decoded using a reference picture, can be considered as images in which inter-prediction is used.
[0440] Below, the inter prediction in the inter mode according to the embodiment is described in detail.
[0441] Inter prediction or motion compensation can be performed using reference images and motion information.
[0442] In inter mode, the encoding device (100) can perform inter prediction and / or motion compensation for a target block. The decoding device (200) can perform inter prediction and / or motion compensation for a target block that corresponds to the inter prediction and / or motion compensation in the encoding device (100).
[0443] Motion information for a target block can be derived during inter-prediction by each of the encoding device (100) and the decoding device (200). The motion information can be derived using motion information of a reconstructed neighbor block, motion information of a call block, and / or motion information of a block adjacent to the call block.
[0444] For example, the encoding device (100) or the decoding device (200) can perform prediction and / or motion compensation by using motion information of a spatial candidate and / or temporal candidate as motion information of a target block. The target block may refer to a PU and / or PU partition.
[0445] A spatial candidate may be a reconstructed block spatially adjacent to the target block.
[0446] The temporal candidate may be a reconstructed block corresponding to the target block within an already reconstructed collocated picture.
[0447] In inter prediction, the encoding device (100) and the decoding device (200) can improve encoding efficiency and decoding efficiency by using motion information of spatial candidates and / or temporal candidates. Motion information of spatial candidates may be referred to as spatial motion information. Motion information of temporal candidates may be referred to as temporal motion information.
[0448] In the following, the motion information of a spatial candidate may be the motion information of a PU containing the spatial candidate. The motion information of a temporal candidate may be the motion information of a PU containing the temporal candidate. The motion information of a candidate block may be the motion information of a PU containing the candidate block.
[0449] Inter prediction can be performed using a reference picture.
[0450] The reference picture may be at least one of the picture preceding the target picture or the picture following the target picture. The reference picture may refer to an image used for predicting the target block.
[0451] In inter prediction, an area within a reference picture can be specified by using a reference picture index (or refIdx) indicating the reference picture and a motion vector to be described later. Here, the specified area within the reference picture may represent a reference block.
[0452] Inter prediction can select a reference picture and select a reference block within the reference picture that corresponds to the target block. Additionally, inter prediction can generate a prediction block for the target block using the selected reference block.
[0453] Motion information can be derived during inter-prediction by each of the encoding device (100) and the decoding device (200).
[0454] A spatial candidate may be a block that 1) exists within the target picture, 2) has already been reconstructed through encoding and / or decoding, and 3) is adjacent to the target block or located at a corner of the target block. Here, a block located at a corner of the target block may be a block vertically adjacent to a neighbor block horizontally adjacent to the target block, or a block horizontally adjacent to a neighbor block vertically adjacent to the target block. "A block located at a corner of the target block" may have the same meaning as "a block adjacent to a corner of the target block." "A block located at a corner of the target block" may be included in "a block adjacent to the target block."
[0455] For example, spatial candidates may be a reconstructed block located to the left of the target block, a reconstructed block located at the top of the target block, a reconstructed block located at the bottom-left corner of the target block, a reconstructed block located at the top-right corner of the target block, or a reconstructed block located at the top-left corner of the target block.
[0456] Each of the encoding device (100) and the decoding device (200) can identify a block located at a position spatially corresponding to a target block within a col picture. The position of the target block within the target picture and the position of the identified block within the col picture can correspond to each other.
[0457] Each of the encoding device (100) and the decoding device (200) can determine a col block existing at a predefined relative position with respect to the identified block as a temporal candidate. The predefined relative position may be a position inside and / or outside the identified block.
[0458] For example, a call block may include a first call block and a second call block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is (nPSW, nPSH), the first call block may be a block located at coordinates (xP + nPSW, yP + nPSH). The second call block may be a block located at coordinates (xP + (nPSW >> 1), yP + (nPSH >> 1)). The second call block may be optionally used when the first call block is unavailable.
[0459] The motion vector of the target block can be determined based on the motion vector of the call block. Each of the encoding device (100) and the decoding device (200) can scale the motion vector of the call block. The scaled motion vector of the call block can be used as the motion vector of the target block. Additionally, the motion vector of the motion information of the temporal candidate stored in the list may be the scaled motion vector.
[0460] The ratio of the motion vector of the target block and the motion vector of the call block may be equal to the ratio of the first temporal distance and the second temporal distance. The first temporal distance may be the distance between the reference picture of the target block and the target picture. The second temporal distance may be the distance between the reference picture of the call block and the call picture.
[0461] The method of deriving motion information may vary depending on the inter prediction mode of the target block. For example, the inter prediction modes applied for inter prediction may include the Advanced Motion Vector Predictor (AMVP) mode, merge mode and skip mode, merge mode with motion vector difference, sub-block merge mode, triangulation mode, inter-intra combined prediction mode, affine inter mode, and current picture reference mode. The merge mode may also be referred to as the motion merge mode. Each of these modes is described in detail below.
[0462] 1) AMVP mode
[0463] When AMVP mode is used, the encoding device (100) can search for a similar block in the neighborhood of the target block. The encoding device (100) can obtain a prediction block by performing a prediction on the target block using the movement information of the searched similar block. The encoding device (100) can encode a residual block, which is the difference between the target block and the prediction block.
[0464] 1-1) Creation of the list of candidate predicted motion vectors
[0465] When the AMVP mode is used as the prediction mode, each of the encoding device (100) and the decoder (200) can generate a list of predicted motion vector candidates using the motion vector of a spatial candidate, the motion vector of a temporal candidate, and a zero vector. The list of predicted motion vector candidates may include one or more predicted motion vector candidates. At least one of the motion vector of a spatial candidate, the motion vector of a temporal candidate, and the zero vector may be determined and used as a predicted motion vector candidate.
[0466] In the following, the terms "predicted motion vector (candidate)" and "motion vector (candidate)" may be used interchangeably with the same meaning.
[0467] In the following, the terms "predicted motion vector candidate" and "AMVP candidate" may be used interchangeably with the same meaning.
[0468] In the following, the terms "predicted motion vector candidate list" and "AMVP candidate list" may be used interchangeably with the same meaning.
[0469] Spatial candidates may include reconstructed spatial neighbor blocks. That is to say, the motion vector of a reconstructed neighbor block may be called a spatial prediction motion vector candidate.
[0470] Temporal candidates may include a call block and a block adjacent to the call block. That is to say, the motion vector of the call block or the motion vector of the block adjacent to the call block may be referred to as a temporal prediction motion vector candidate.
[0471] The zero vector can be the (0, 0) motion vector.
[0472] The predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. Additionally, in the encoding device (100), the predicted motion vector candidate may be a motion vector initial search location.
[0473] 1-2) Searching for motion vectors using a list of predicted motion vector candidates
[0474] The encoding device (100) can determine a motion vector to be used for encoding a target block within a search range using a list of predicted motion vector candidates. Additionally, the encoding device (100) can determine a predicted motion vector candidate to be used as the predicted motion vector of the target block among the predicted motion vector candidates in the list of predicted motion vector candidates.
[0475] The motion vector to be used for encoding the target block can be a motion vector that can be encoded at minimum cost.
[0476] Additionally, the encoding device (100) can determine whether to use the AMVP mode in encoding the target block.
[0477] 1-3) Transmission of Inter-prediction Information
[0478] The encoding device (100) can generate a bitstream containing inter prediction information required for inter prediction. The decoding device (200) can perform inter prediction for a target block using the inter prediction information of the bitstream.
[0479] Inter prediction information may include 1) mode information indicating whether AMVP mode is used, 2) prediction motion vector index, 3) motion vector difference (MVD), 4) reference direction, and 5) reference picture index.
[0480] In the following, the terms "Predicted Motion Vector Index" and "AMVP Index" may be used interchangeably with the same meaning.
[0481] In addition, the inter-prediction information may include residual signals.
[0482] The decoding device (200) can obtain the predicted motion vector index, motion vector difference, reference direction, and reference picture index from the bitstream through entropy decoding when the mode information indicates that the AMVP mode is used.
[0483] The predicted motion vector index can point to a predicted motion vector candidate used for predicting the target block among the predicted motion vector candidates included in the predicted motion vector candidate list.
[0484] 1-4) Inter prediction in AMVP mode using inter prediction information
[0485] The decoding device (200) can derive predicted motion vector candidates using a list of predicted motion vector candidates, and can determine motion information of the target block based on the derived predicted motion vector candidates.
[0486] The decoding device (200) can determine a motion vector candidate for a target block from among the predicted motion vector candidates included in the predicted motion vector candidate list using a predicted motion vector index. The decoding device (200) can select a predicted motion vector candidate pointed to by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list as the predicted motion vector for the target block.
[0487] The encoding device (100) can generate an entropy-encoded predicted motion vector index by applying entropy encoding to a predicted motion vector index and can generate a bitstream containing the entropy-encoded predicted motion vector index. The entropy-encoded predicted motion vector index can be signaled from the encoding device (100) to the decoding device (200) through the bitstream. The decoding device (200) can extract the entropy-encoded predicted motion vector index from the bitstream and can obtain the predicted motion vector index by applying entropy decoding to the entropy-encoded predicted motion vector index.
[0488] The motion vector actually used for inter-predicting of the target block may not match the predicted motion vector. An MVD may be used to represent the difference between the motion vector actually used for inter-predicting of the target block and the predicted motion vector. The encoding device (100) may derive a predicted motion vector similar to the motion vector actually used for inter-predicting of the target block in order to use an MVD of the smallest possible size.
[0489] The MVD may be the difference between the motion vector of the target block and the predicted motion vector. The encoding device (100) may calculate the MVD and generate an entropy-encoded MVD by applying entropy encoding to the MVD. The encoding device (100) may generate a bitstream containing the entropy-encoded MDV.
[0490] The MVD can be transmitted from the encoding device (100) to the decoding device (200) via a bitstream. The decoding device (200) can extract the entropy-encoded MVD from the bitstream and can obtain the MVD by applying entropy decoding to the entropy-encoded MVD.
[0491] The decoding device (200) can derive the motion vector of the target block by combining the MVD and the predicted motion vector. That is to say, the motion vector of the target block derived from the decoding device (200) may be the sum of the MVD and the motion vector candidate.
[0492] Additionally, the encoding device (100) can generate entropy-encoded MVD resolution information by applying entropy encoding to the calculated MVD resolution information, and can generate a bitstream containing the entropy-encoded MVD resolution information. The decoding device (200) can extract entropy-encoded MVD resolution information from the bitstream and can obtain MVD resolution information by applying entropy decoding to the entropy-encoded MVD resolution information. The decoding device (200) can adjust the resolution of the MVD using the MVD resolution information.
[0493] Meanwhile, the encoding device (100) can calculate the MVD based on an affine model. The decoding device (200) can derive the affine control motion vector of the target block through the sum of the MVD and the affine control motion vector candidates, and can derive the motion vector for the sub-block using the affine control motion vector.
[0494] The reference direction may point to a reference picture list used for predicting the target block. For example, the reference direction may point to either reference picture list L0 or reference picture list L1.
[0495] The reference direction merely refers to the reference picture list used for predicting the target block, and may not indicate that the directions of the reference pictures are restricted to the forward or backward direction. That is to say, each of reference picture list L0 and reference picture list L1 may contain forward and / or backward pictures.
[0496] A unidirectional reference direction may mean that a single reference picture list is used. A bidirectional reference direction may mean that two reference picture lists are used. That is to say, the reference direction may point to either only reference picture list L0 being used, only reference picture list L1 being used, or one of two reference picture lists.
[0497] A reference picture index may point to a reference picture used for predicting a target block among the reference pictures in a reference picture list. An encoding device (100) may generate an entropy-encoded reference picture index by applying entropy encoding to the reference picture index and may generate a bitstream containing the entropy-encoded reference picture index. The entropy-encoded reference picture index may be signaled from the encoding device (100) to the decoding device (200) via the bitstream. The decoding device (200) may extract the entropy-encoded reference picture index from the bitstream and may obtain the reference picture index by applying entropy decoding to the entropy-encoded reference picture index.
[0498] When two reference picture lists are used for the prediction of a target block, one reference picture index and one motion vector may be used for each reference picture list. Additionally, when two reference picture lists are used for the prediction of a target block, two prediction blocks may be specified for the target block. For example, the (final) prediction block of the target block may be generated through the average or weighted-sum of the two prediction blocks for the target block.
[0499] The motion vector of the target block can be derived by the predicted motion vector index, MVD, reference direction, and reference picture index.
[0500] The decoding device (200) can generate a prediction block for a target block based on an induced motion vector and a reference picture index. For example, the prediction block may be a reference block pointed to by an induced motion vector within a reference picture pointed to by a reference picture index.
[0501] By encoding the predicted motion vector index and MVD instead of encoding the motion vector of the target block itself, the amount of bits transmitted from the encoding device (100) to the decoding device (200) can be reduced and the encoding efficiency can be improved.
[0502] Motion information of a reconstructed neighbor block can be used for the target block. In a specific inter-prediction mode, the encoding device (100) may not separately encode the motion information itself for the target block. Instead of encoding the motion information of the target block, other information capable of inducing the motion information of the target block through the motion information of the reconstructed neighbor block may be encoded instead. As other information is encoded instead, the amount of bits transmitted to the decoding device (200) may be reduced, and encoding efficiency may be improved.
[0503] For example, as an inter-prediction mode in which the motion information of such target blocks is not directly encoded, there may be a skip mode and / or a merge mode. In this case, the encoding device (100) and the decoding device (200) may use an identifier and / or an index indicating which unit's motion information among the reconstructed neighboring units is used as the motion information of the target unit.
[0504] 2) Merge Mode
[0505] Merge is a method for deriving motion information of a target block. Merge can refer to the merging of motions across multiple blocks. It can also refer to applying the motion information of one block to other blocks. In other words, merge mode can refer to a mode where the motion information of a target block is derived from the motion information of neighboring blocks.
[0506] When a merge mode is used, the encoding device (100) can perform a prediction of the motion information of the target block using the motion information of the spatial candidate and / or the motion information of the temporal candidate. The spatial candidate may include a reconstructed spatial neighbor block that is spatially adjacent to the target block. The spatial neighbor block may include a left neighbor block and a top neighbor block. The temporal candidate may include a call block. The terms "spatial candidate" and "spatial merge candidate" may be used interchangeably with the same meaning. The terms "temporal candidate" and "temporal merge candidate" may be used interchangeably with the same meaning.
[0507] The encoding device (100) can obtain a prediction block through prediction. The encoding device (100) can encode a residual block which is the difference between the target block and the prediction block.
[0508] 2-1) Creating the merge candidate list
[0509] When a merge mode is used, each of the encoding device (100) and the decoder (200) can generate a merge candidate list using motion information of a spatial candidate and / or motion information of a temporal candidate. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may be unidirectional or bidirectional. The reference direction may mean an inter-prediction indicator.
[0510] The merge candidate list can contain merge candidates. A merge candidate can be movement information. In other words, the merge candidate list can be a list where movement information is stored.
[0511] Merge candidates can be motion information such as temporal candidates and / or spatial candidates. In other words, the merge candidate list can include motion information such as temporal candidates and / or spatial candidates.
[0512] Additionally, the merge candidate list may include new merge candidates generated by combinations of merge candidates already existing in the merge candidate list. That is to say, the merge candidate list may include new motion information generated by combinations of motion information already existing in the merge candidate list.
[0513] Additionally, the merge candidate list may include history-based merge candidates. History-based merge candidates may be movement information of blocks that were encoded and / or decoded prior to the target block.
[0514] In addition, the merge candidate list may include a merge candidate based on the average of two merge candidates.
[0515] Merge candidates may be specific modes that induce inter-prediction information. A merge candidate may be information pointing to a specific mode that induces inter-prediction information. Inter-prediction information of a target block may be induced according to the specific mode pointed to by the merge candidate. In this case, the specific mode may include a process that induces a series of inter-prediction information. Such a specific mode may be an inter-prediction information inducing mode or a movement information inducing mode.
[0516] Inter-prediction information of the target block can be derived depending on the mode pointed to by the merge candidate selected by the merge index among the merge candidates in the merge candidate list.
[0517] For example, motion information induction modes within the merge candidate list may be at least one of 1) a motion information induction mode at the sub-block level and 2) an affine motion information induction mode.
[0518] In addition, the merge candidate list may include movement information of the zero vector. The zero vector may also be referred to as a zero merge candidate.
[0519] In other words, motion information within the merge candidate list may be at least one of 1) motion information of a spatial candidate, 2) motion information of a temporal candidate, 3) motion information generated by a combination of motion information already existing in the merge candidate list, and 4) a zero vector.
[0520] Motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an inter-prediction indicator. The reference direction may be unidirectional or bidirectional. A unidirectional reference direction may represent L0 prediction or L1 prediction.
[0521] The merge candidate list can be generated before prediction by merge mode is performed.
[0522] The number of merge candidates in the merge candidate list can be predetermined. To ensure that the merge candidate list has a predetermined number of merge candidates, the encoding device (100) and the decoding device (200) may add merge candidates to the merge candidate list according to a predetermined method and a predetermined ranking. Through the predetermined method and the predetermined ranking, the merge candidate list of the encoding device (100) and the merge candidate list of the decoding device (200) can be identical.
[0523] Merging can be applied in units of CU or PU. When merging is performed in units of CU or PU, the encoding device (100) can transmit a bitstream containing predefined information to the decoding device (200). For example, the predefined information may include 1) information indicating whether to perform merging by block partition, and 2) information regarding which block among the blocks that are spatial candidates and / or temporal candidates for the target block will be merged with.
[0524] 2-2) Searching for motion vectors using a merge candidate list
[0525] The encoding device (100) can determine merge candidates to be used for encoding a target block. For example, the encoding device (100) can perform predictions for the target block using merge candidates from a merge candidate list and generate residual blocks for the merge candidates. The encoding device (100) can use a merge candidate that requires the minimum cost for encoding the predictions and residual blocks for encoding the target block.
[0526] Additionally, the encoding device (100) can determine whether to use a merge mode in encoding the target block.
[0527] 2-3) Transmission of Inter-prediction Information
[0528] The encoding device (100) can generate a bitstream containing inter prediction information required for inter prediction. The encoding device (100) can generate entropy-encoded inter prediction information by performing entropy encoding on the inter prediction information, and can transmit the bitstream containing the entropy-encoded inter prediction information to the decoding device (200). Through the bitstream, the entropy-encoded inter prediction information can be signaled from the encoding device (100) to the decoding device (200). The decoding device (200) can extract the entropy-encoded inter prediction information from the bitstream and obtain the inter prediction information by performing entropy decoding on the entropy-encoded inter prediction information.
[0529] The decoding device (200) can perform inter prediction for the target block using inter prediction information of the bitstream.
[0530] Inter prediction information may include 1) mode information indicating whether to use merge mode, 2) a merge index, and 3) correction information.
[0531] In addition, the inter-prediction information may include residual signals.
[0532] The decoding device (200) can obtain a merge index from the bitstream only when the mode information indicates that the merge mode is used.
[0533] Mode information can be a merge flag. The unit of mode information can be a block. Information about a block may include mode information, and the mode information may indicate whether a merge mode is applied to the block.
[0534] The merge index may point to a merge candidate among the merge candidates included in the merge candidate list that is used for predicting the target block. Alternatively, the merge index may point to which of the neighboring blocks spatially or temporally adjacent to the target block the merge is performed with.
[0535] The encoding device (100) can select a merge candidate having the highest encoding performance among the merge candidates included in the merge candidate list, and can set the value of the merge index to point to the selected merge candidate.
[0536] Correction information may be information used to correct motion vectors. The encoding device (100) may generate correction information. The decoding device (200) may correct the motion vector of a merge candidate selected by a merge index based on the correction information.
[0537] Correction information may include at least one of information indicating whether correction is made, correction direction information, and correction magnitude information. A prediction mode that corrects motion vectors based on signaled correction information may be referred to as a merge mode having motion vector differences.
[0538] 2-4) Inter prediction in merge mode using inter prediction information
[0539] The decoding device (200) can perform a prediction for the target block using the merge candidate pointed to by the merge index among the merge candidates included in the merge candidate list.
[0540] The motion vector of the target block can be determined by the motion vector of the merge candidate pointed to by the merge index, the reference picture index, and the reference direction.
[0541] 3) Skip Mode
[0542] Skip mode may be a mode that applies the motion information of spatial or temporal candidates directly to the target block. Additionally, skip mode may be a mode that does not use residual signals. In other words, when skip mode is used, the reconstructed block may be identical to the predicted block.
[0543] The difference between merge mode and skip mode may be whether the residual signal is transmitted or used. In other words, skip mode may be similar to merge mode except that the residual signal is not transmitted or used.
[0544] When a skip mode is used, the encoding device (100) can transmit information indicating which block's motion information among the blocks that are spatial candidates or temporal candidates is used as motion information of the target block to the decoding device (200) via a bitstream. The encoding device (100) can generate entropy-encoded information by performing entropy encoding on this information and can signal the entropy-encoded information to the decoding device (200) via a bitstream. The decoding device (200) can extract the entropy-encoded information from the bitstream and obtain information by performing entropy decoding on the entropy-encoded information.
[0545] Additionally, when a skip mode is used, the encoding device (100) may not transmit other syntax element information, such as MVD, to the decoding device (200). For example, when a skip mode is used, the encoding device (100) may not signal syntax elements regarding at least one of MVD, coded block flags, and transform factor levels to the decoding device (200).
[0546] 3-1) Creation of the Merge Candidate List
[0547] Skip mode can also use a merge candidate list. In other words, the merge candidate list can be used in both merge mode and skip mode. In this respect, the merge candidate list may also be referred to as a "skip candidate list" or a "merge / skip candidate list."
[0548] Alternatively, Skip mode may use a separate candidate list different from Merge mode. In this case, the Merge candidate list and Merge candidate in the description below may be replaced with the Skip candidate list and Skip candidate, respectively.
[0549] The merge candidate list can be generated before prediction by skip mode is performed.
[0550] 3-2) Searching for motion vectors using a merge candidate list
[0551] The encoding device (100) can determine merge candidates to be used for encoding a target block. For example, the encoding device (100) can perform predictions for the target block using merge candidates from a merge candidate list. The encoding device (100) can use a merge candidate that requires the minimum cost in the prediction for encoding the target block.
[0552] Additionally, the encoding device (100) can determine whether to use a skip mode in encoding the target block.
[0553] 3-3) Transmission of Inter-prediction Information
[0554] The encoding device (100) can generate a bitstream containing inter prediction information required for inter prediction. The decoding device (200) can perform inter prediction for a target block using the inter prediction information of the bitstream.
[0555] Inter prediction information may include 1) mode information indicating whether to use skip mode and 2) a skip index.
[0556] The skip index can be the same as the merge index described above.
[0557] When skip mode is used, the target block can be encoded without residual signals. Inter-prediction information may not include residual signals. Alternatively, the bitstream may not include residual signals.
[0558] The decoding device (200) can obtain a skip index from the bitstream only when the mode information indicates that the skip mode is used. As previously described, the merge index and the skip index may be the same. The decoding device (200) can obtain a skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.
[0559] The skip index can point to a merge candidate among the merge candidates included in the merge candidate list that is used for predicting the target block.
[0560] 3-4) Inter prediction in skip mode using inter prediction information
[0561] The decoding device (200) can perform a prediction for the target block using the merge candidate pointed to by the skip index among the merge candidates included in the merge candidate list.
[0562] The motion vector of the target block can be determined by the motion vector of the merge candidate pointed to by the skip index, the reference picture index, and the reference direction.
[0563] 4) Current Picture Reference Mode
[0564] The current picture reference mode may refer to a prediction mode that utilizes a pre-reconstructed area within the target picture to which the target block belongs.
[0565] Motion vectors may be used to identify pre-reconstructed regions. Whether the target block is encoded in the current picture reference mode can be determined using the target block's reference picture index.
[0566] A flag or index indicating whether the target block is a block encoded in the current picture reference mode may be signaled from the encoding device (100) to the decoding device (200). Alternatively, whether the target block is a block encoded in the current picture reference mode may be inferred through the reference picture index of the target block.
[0567] If the target block is encoded in the current picture reference mode, the target picture may exist at a fixed position or at any position within the reference picture list for the target block.
[0568] For example, a fixed position can be a position where the value of the reference picture index is 0 or the very last position.
[0569] If the target picture exists at any location within the reference picture list, a separate reference picture index indicating such any location may be signaled from the encoding device (100) to the decoding device (200).
[0570] 5) Subblock merge mode
[0571] The sub-block merge mode may refer to a mode that induces movement information for the sub-blocks of the CU.
[0572] For example, the name of the merge subblock information can be "merge_subblock_flag". merge_subblock_flag can indicate the subblock merge mode. The merge subblock information can be a flag.
[0573] For example, a merge_subblock_flag value of 1 (e.g., "0") may indicate that the subblock merge mode is not applied.
[0574] For example, a second value of merge_subblock_flag (e.g., "1") may indicate that subblock merge mode is applied.
[0575] When the subblock merge mode is applied, a subblock merge candidate list can be generated using motion information of the call subblock of the target subblock in the reference image (i.e., subblock-based temporal merge candidate) and / or affine control point motion vector merge candidate.
[0576] 6) Triangle partition mode
[0577] In triangular partitioning mode, partitioned target blocks can be generated by partitioning the target block diagonally. For each partitioned target block, motion information of that partitioned target block can be derived, and a prediction sample for each partitioned target block can be derived using the derived motion information. A prediction sample of the target block can be derived through a weighted sum of the prediction samples of the partitioned target blocks.
[0578] 7) Inter-intra combined prediction mode
[0579] The inter-intra combined prediction mode may be a mode that derives a prediction sample of a target block using a weighted sum of a prediction sample generated by inter-prediction and a prediction sample generated by intra-prediction.
[0580] In the aforementioned modes, the decoding device (200) can perform self-correction on the derived motion information. For example, the decoding device (200) can search a specific area based on a reference block indicated by the derived motion information to find motion information having a sum of absolute differences (SAD), and can derive the found motion information as corrected motion information.
[0581] In the aforementioned modes, the decoder (200) can perform compensation for the prediction sample derived through inter-prediction using an optical flow.
[0582] In the aforementioned AMVP mode, merge mode, and skip mode, etc., the movement information to be used for predicting the target block among the movement information within the list can be specified through the index for the list.
[0583] To improve encoding efficiency, the encoding device (100) may signal only the index of the element that causes the minimum cost in inter-predicting of the target block among the elements of the list. The encoding device (100) may encode the index and signal the encoded index.
[0584] Accordingly, the aforementioned lists (i.e., the predicted motion vector candidate list and the merge candidate list) may be derived in the same manner based on the same data in the encoding device (100) and the decoding device (200). Here, the same data may include reconstructed pictures and reconstructed blocks. Additionally, in order to identify elements by index, the order of elements within the list may be constant.
[0585] Figure 10 shows spatial candidates according to one example.
[0586] In Fig. 10, the locations of the spatial candidates are shown.
[0587] The large block in the center can represent the target block. The five small blocks can represent spatial candidates.
[0588] The coordinates of the target block can be (xP, yP), and the size of the target block can be (nPSW, nPSH).
[0589] Spatial candidate A0 may be a block adjacent to the bottom-left corner of the target block. A0 may be a block occupying pixels at coordinates (xP - 1, yP + nPSH + 1).
[0590] Spatial candidate A1 may be a block adjacent to the left of the target block. A1 may be the bottommost block among the blocks adjacent to the left of the target block. Or, A1 may be a block adjacent to the top of A0. A1 may be a block occupying pixels at coordinates (xP - 1, yP + nPSH).
[0591] Spatial candidate B0 may be a block adjacent to the top-right corner of the target block. B0 may be a block occupying pixels at coordinates (xP + nPSW + 1, yP - 1).
[0592] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Or, B1 may be a block adjacent to the left of B0. B1 may be a block occupying pixels at coordinates (xP + nPSW, yP - 1).
[0593] Spatial candidate B2 may be a block adjacent to the top-left corner of the target block. B2 may be a block occupying pixels at coordinates (xP - 1, yP - 1).
[0594] Determination of the availability of spatial and temporal candidates
[0595] In order to include movement information of spatial candidates or movement information of temporal candidates in the list, it must be determined whether the movement information of spatial candidates or movement information of temporal candidates is available.
[0596] In the following, candidate blocks may include spatial candidates and temporal candidates.
[0597] For example, the above judgment can be made by sequentially applying steps 1) to 4) below.
[0598] Step 1) If the PU containing the candidate block is outside the boundaries of the picture, the availability of the candidate block can be set to false. "Availability is set to false" may be synonymous with "is set to unavailable."
[0599] Step 2) If the PU containing the candidate block is outside the boundaries of the slice, the availability of the candidate block can be set to false. If the target block and the candidate block are located within different slices, the availability of the candidate block can be set to false.
[0600] Step 3) If the PU containing the candidate block is outside the tile boundary, the availability of the candidate block can be set to false. If the target block and the candidate block are located within different tiles, the availability of the candidate block can be set to false.
[0601] Step 4) If the prediction mode of the PU containing the candidate block is intra prediction mode, the availability of the candidate block can be set to false. If the PU containing the candidate block does not use inter prediction, the availability of the candidate block can be set to false.
[0602] Figure 11 shows the order of addition of motion information of spatial candidates to a merge list according to one example.
[0603] As illustrated in FIG. 11, the order A1, B1, B0, A0, and B2 may be used when adding motion information of spatial candidates to the merge list. That is, motion information of available spatial candidates may be added to the merge list in the order A1, B1, B0, A0, and B2.
[0604] Method for deriving a merge list in merge mode and skip mode
[0605] As described above, the maximum number of merge candidates in the merge list can be set. The set maximum number is denoted as N. The set number can be transmitted from the encoding device (100) to the decoding device (200). The slice header of the slice may include N. That is to say, the maximum number of merge candidates in the merge list for the target block of the slice can be set by the slice header. For example, the value of N can be 5 by default.
[0606] Movement information (i.e., merge candidates) can be added to the merge list in the order of steps 1) through 4) below.
[0607] Step 1) Available spatial candidates among the spatial candidates may be added to the merge list. The movement information of the available spatial candidates may be added to the merge list in the order shown in FIG. 10. At this time, if the movement information of an available spatial candidate overlaps with other movement information already existing in the merge list, the movement information may not be added to the merge list. Checking whether there is overlap with other movement information existing in the list may be abbreviated as "redundancy check."
[0608] The number of additional movement information items can be up to N.
[0609] Step 2) If the number of motion information items in the merge list is less than N and a temporal candidate is available, the motion information of the temporal candidate can be added to the merge list. In this case, if the available motion information of the temporal candidate overlaps with other motion information already existing in the merge list, the motion information may not be added to the merge list.
[0610] Step 3)If the number of motion information items in the merge list is less than N and the type of the target slice is "B", combined motion information generated by combined bi-prediction can be added to the merge list.
[0611] The target slice may be a slice containing the target block.
[0612] The combined motion information may be a combination of L0 motion information and L1 motion information. The L0 motion information may be motion information that references only the reference picture list L0. The L1 motion information may be motion information that references only the reference picture list L1.
[0613] Within the merge list, there may be one or more L0 motion information items. Additionally, within the merge list, there may be one or more L1 motion information items.
[0614] There may be one or more combined motion information. In generating the combined motion information, which L0 motion information and which L1 motion information among one or more L0 motion information and one or more L1 motion information will be used may be predefined. One or more combined motion information may be generated in a predefined order by combined bidirectional prediction using pairs of different motion information within a merge list. One of the pairs of different motion information may be L0 motion information and the other may be L1 motion information.
[0615] For example, the combined motion information added first may be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. If the motion information with a merge index of 0 is not L0 motion information, or if the motion information with a merge index of 1 is not L1 motion information, the above combined motion information may not be generated or added. The motion information added next may be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. The specific combinations below may follow other combinations in the field of video encoding / decoding.
[0616] In this case, if the combined motion information overlaps with other motion information already existing in the merge list, the above combined motion information may not be added to the merge list.
[0617] Step 4) If the number of motion information items in the merge list is less than N, zero vector motion information can be added to the merge list.
[0618] Zero vector motion information can be motion information where the motion vector is a zero vector.
[0619] There may be one or more zero vector motion information. The reference picture indices of one or more zero vector motion information may be different from each other. For example, the value of the reference picture index of the first zero vector motion information may be 0. The value of the reference picture index of the second zero vector motion information may be 1.
[0620] The number of zero vector motion information items can be equal to the number of reference pictures in the reference picture list.
[0621] The reference direction of zero vector motion information can be bidirectional. Both motion vectors can be zero vectors. The number of zero vector motion information can be the smaller of the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1. Alternatively, if the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1 are different, a unidirectional reference direction may be used for reference picture indices that can be applied to only one reference picture list.
[0622] The encoding device (100) and / or the decoding device (200) can sequentially add zero vector movement information to the merge list while changing the reference picture index.
[0623] If zero vector motion information overlaps with other motion information already existing in the merge list, the above zero vector motion information may not be added to the merge list.
[0624] The order of steps 1) through 4) described above is merely illustrative, and the order of the steps may be changed. Additionally, some of the steps may be omitted according to predefined conditions.
[0625] Method for deriving the list of predicted motion vector candidates in AMVP mode
[0626] The maximum number of predicted motion vector candidates in the list of predicted motion vector candidates can be predefined. Let N be the predefined maximum number. For example, the predefined maximum number can be 2.
[0627] Motion information (i.e., predicted motion vector candidates) can be added to the list of predicted motion vector candidates in the order of steps 1) to 3) below.
[0628] Step 1)Available spatial candidates among the spatial candidates may be added to the list of predicted motion vector candidates. Spatial candidates may include a first spatial candidate and a second spatial candidate.
[0629] The first spatial candidate may be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate may be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.
[0630] Motion information of available spatial candidates can be added to the predicted motion vector candidate list in the order of the first spatial candidate and the second spatial candidate. At this time, if the motion information of an available spatial candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information may not be added to the predicted motion vector candidate list. That is to say, when the value of N is 2, if the motion information of the second spatial candidate is identical to the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.
[0631] The number of additional movement information items can be up to N.
[0632] Step 2) If the number of motion information items in the predicted motion vector candidate list is less than N and a temporal candidate is available, the motion information of the temporal candidate can be added to the predicted motion vector candidate list. In this case, if the available motion information of the temporal candidate overlaps with other motion information already existing in the predicted motion vector candidate list, the motion information may not be added to the predicted motion vector candidate list.
[0633] Step 3) If the number of motion information items in the predicted motion vector candidate list is less than N, zero vector motion information can be added to the predicted motion vector candidate list.
[0634] There may be one or more zero vector motion information. The reference picture indices of one or more zero vector motion information may be different from each other.
[0635] The encoding device (100) and / or the decoding device (200) can sequentially add zero vector motion information to the predicted motion vector candidate list while changing the reference picture index.
[0636] If zero vector motion information overlaps with other motion information already existing in the predicted motion vector candidate list, the above zero vector motion information may not be added to the predicted motion vector candidate list.
[0637] The description of zero vector motion information described above for the merge list may also be applied to zero vector motion information. Redundant descriptions are omitted.
[0638] The order of steps 1) through 3) described above is merely illustrative, and the order of the steps may be changed. Additionally, some of the steps may be omitted according to predefined conditions.
[0639] Figure 12 illustrates the process of transformation and quantization according to one example.
[0640] As shown in FIG. 12, a quantized level can be generated by performing a conversion and / or quantization process on the residual signal.
[0641] The residual signal can be generated as the difference between the original block and the prediction block. Here, the prediction block may be a block generated by intra-prediction or inter-prediction.
[0642] The residual signal can be converted into the frequency domain through a transformation process that is part of the quantization process.
[0643] The transformation kernels used for the transformation may include various DCT kernels such as the Discrete Cosine Transform (DCT) type 2 (DCT-II) and the Discrete Sine Transform (DST) kernel.
[0644] These transformation kernels can perform a separable transform or a 2D non-separable transform on the residual signal. A separable transform may be a transformation that performs a 1D transformation on the residual signal in the horizontal direction and the vertical direction, respectively.
[0645] The DCT types and DST types adaptively used for 1D transformation may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II, as shown in Tables 3 and 4 below, respectively.
[0646] [Table 3]
[0647]
[0648] [Table 4]
[0649]
[0650] As shown in Tables 3 and 4, a transform set may be used to derive the DCT type or DST type to be used for the transformation. Each transform set may include multiple transformation candidates. Each transformation candidate may be a DCT type or a DST type, etc.
[0651] Table 5 below shows an example of a set of transformations applied to the horizontal direction and a set of transformations applied to the vertical direction according to the intra prediction mode.
[0652] [Table 5]
[0653]
[0654] In Table 5, the number of vertical direction transformation sets and the number of horizontal direction transformation sets applied to the horizontal direction of the residual signal according to the intra-prediction mode of the target block are shown.
[0655] As exemplified in Table 5, sets of transformations applied to the horizontal and vertical directions may be predefined according to the intra-prediction mode of the target block. The encoding device (100) can perform transformations and inverse transformations on the residual signal using transformations included in the sets of transformations corresponding to the intra-prediction mode of the target block. Additionally, the decoding device (200) can perform inverse transformations on the residual signal using transformations included in the sets of transformations corresponding to the intra-prediction mode of the target block.
[0656] In these transformations and inverse transformations, the transformation set applied to the residual signal may be determined as exemplified in Tables 3, 4, and 5 and may not be signaled. Transformation instruction information may be signaled from the encoding device (100) to the decoding device (200). The transformation instruction information may be information indicating which transformation candidate among a plurality of transformation candidates included in the transformation set applied to the residual signal is used.
[0657] For example, if the size of the target block is 64x64 or less, a set of transformations consisting of three in total can be configured according to the intra prediction mode. An optimal transformation method can be selected from a total of nine multiple transformation methods resulting from the combination of three transformations in the horizontal direction and three transformations in the vertical direction. Encoding efficiency can be improved by encoding and / or decoding the residual signal using this optimal transformation method.
[0658] In this case, for at least one of the vertical transformation and the horizontal transformation, information regarding which of the transformations belonging to the transformation set was used can be entropy encoded and / or decoded. Truncated unary binarization can be used for encoding and / or decoding this information.
[0659] As described above, methods using various transformations can be applied to residual signals generated by intra-prediction or inter-prediction.
[0660] The transformation may include at least one of a first-order transformation and a second-order transformation. Transformation coefficients may be generated by performing a first-order transformation on the residual signal, and second-order transformation coefficients may be generated by performing a second-order transformation on the transformation coefficients.
[0661] The first transformation can be named the primary transformation. Additionally, the first transformation can be named the Adaptive Multiple Transform (AMT). As previously mentioned, AMT may mean that different transformations are applied to each of the 1D directions (i.e., the vertical direction and the horizontal direction).
[0662] A second-order transformation may be a transformation intended to enhance the energy concentration of the transformation factors generated by a first-order transformation. Like the first-order transformation, the second-order transformation may be a separable transformation or a non-separable transformation. A non-separable transformation may be a non-separable secondary transformation (NSST).
[0663] The first transformation can be performed using at least one of a plurality of predefined transformation methods. For example, the plurality of predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and Karhunen-Loeve Transform (KLT) based transformations.
[0664] In addition, the first transformation can be a transformation of various types depending on the kernel function defining the DCT or DST.
[0665] For example, the first transformation may include transformations such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8 according to the transformation kernels presented in Table 6 below. Table 6 provides examples of various transformation types and transformation kernel functions for Multiple Transform Selection (MTS).
[0666] MTS may mean that a combination of one or more DCT and / or DST transformation kernels is selected for the transformation of the residual signal in the horizontal and / or vertical directions.
[0667] [Table 6]
[0668]
[0669] In Table 6, i and j can be integer values greater than or equal to 0 and less than or equal to N-1.
[0670] A secondary transform can be performed on the transformation coefficients generated by the performance of a first transformation.
[0671] As in first-order transformations, a set of transformations can be defined in second-order transformations. Methods for deriving and / or determining such a set of transformations as described above can be applied to second-order transformations as well as first-order transformations.
[0672] The first transformation and the second transformation can be determined for a specific target.
[0673] For example, first-order and second-order transformations may be applied to one or more signal components among the luminance component and the chroma component. Whether to apply the first-order and / or second-order transformations may be determined by at least one of the coding parameters for the target block and / or neighboring blocks. For example, whether to apply the first-order and / or second-order transformations may be determined by the size and / or shape of the target block.
[0674] In the encoding device (100) and the decoding device (200), conversion information indicating the conversion method used for the target can be derived by using specific information.
[0675] For example, the transformation information may include an index of the transformation to be used for the first transformation and / or the second transformation. Alternatively, the transformation information may indicate that the first transformation and / or the second transformation is not used.
[0676] For example, when the target of the first transformation and the second transformation is a target block, the transformation method(s) applied to the first transformation and / or the second transformation indicated by the transformation information may be determined according to at least one of the coding parameters for the target block and / or neighboring blocks.
[0677] Alternatively, conversion information indicating a conversion method for a specific target may be signaled from the encoding device (100) to the decoding device (200).
[0678] For example, for one CU, whether a first transformation is used, an index pointing to the first transformation, whether a second transformation is used, and an index pointing to the second transformation can be derived as transformation information in the decoding device (200). Alternatively, transformation information indicating whether a first transformation is used, an index pointing to the first transformation, whether a second transformation is used, and an index pointing to the second transformation can be signaled for one CU.
[0679] Quantized transformation coefficients (i.e., quantized levels) can be generated by performing quantization on the result or residual signal generated by performing a first transformation and / or a second transformation.
[0680] Figure 13 shows diagonal scanning according to one example.
[0681] Figure 14 shows horizontal scanning according to one example.
[0682] Figure 15 shows vertical scanning according to one example.
[0683] Quantized transform coefficients can be scanned according to at least one of an intra prediction mode, block size, and block shape, and according to at least one of (up-right) diagonal scanning, vertical scanning, and horizontal scanning. A block may be a transform unit.
[0684] Each scanning can start at a specific starting point and end at a specific ending point.
[0685] For example, the quantized transformation coefficients can be changed into a one-dimensional vector form by scanning the coefficients of the block using the diagonal scanning of FIG. 13. Alternatively, depending on the size of the block and / or the intra-prediction mode, the horizontal scanning of FIG. 14 or the vertical scanning of FIG. 15 may be used instead of the diagonal scanning.
[0686] Vertical scanning may be scanning two-dimensional block-shaped coefficients in the column direction. Horizontal scanning may be scanning two-dimensional block-shaped coefficients in the row direction.
[0687] In other words, depending on the block size and / or the inter-prediction mode, it can be determined which scanning method among diagonal scanning, vertical scanning, and horizontal scanning will be used.
[0688] As illustrated in FIGS. 13, 14, and 15, quantized transformation coefficients can be scanned along the diagonal, horizontal, or vertical directions.
[0689] Quantized transform coefficients can be represented in the form of blocks. A block may include multiple sub-blocks. Each sub-block can be defined according to a minimum block size or minimum block shape.
[0690] In scanning, the scanning order according to the type or direction of scanning can first be applied to the sub-blocks. Additionally, the scanning order according to the direction of scanning can be applied to the quantized transform coefficients within the sub-blocks.
[0691] For example, as illustrated in FIGS. 13, 14, and 15, when the size of the target block is 8x8, quantized transformation coefficients can be generated by first transformation, second transformation, and quantization of the residual signal of the target block. Subsequently, one of three scanning sequences can be applied to four 4x4 sub-blocks, and quantized transformation coefficients can be scanned according to the scanning sequence for each 4x4 sub-block.
[0692] The encoding device (100) can generate entropy-encoded quantized conversion coefficients by performing entropy encoding on scanned quantized conversion coefficients, and can generate a bitstream containing entropy-encoded quantized conversion coefficients.
[0693] The decoding device (200) can extract entropy-encoded quantized transformation coefficients from a bitstream and generate quantized transformation coefficients by performing entropy decoding on the entropy-encoded quantized transformation coefficients. The quantized transformation coefficients can be aligned into a two-dimensional block form through inverse scanning. At this time, as a method of inverse scanning, at least one of a (top-right) diagonal scan, a vertical scan, and a horizontal scan may be performed.
[0694] In the decoding device (200), inverse quantization can be performed on the quantized transformation coefficients. Depending on whether a second inverse transformation is performed, a second inverse transformation can be performed on the result generated by the performance of inverse quantization. Also, depending on whether a first inverse transformation is performed, a first inverse transformation can be performed on the result generated by the performance of the second inverse transformation. A reconstructed residual signal can be generated by performing a first inverse transformation on the result generated by the performance of the second inverse transformation.
[0695] For the luminance component reconstructed through intra-prediction or inter-prediction, inverse mapping of the dynamic range can be performed before in-loop filtering.
[0696] The dynamic range can be divided into 16 equal pieces, and a mapping function for each piece can be signaled. The mapping function can be signaled at the slice level or the tile group level.
[0697] A reverse mapping function for performing reverse mapping can be derived based on a mapping function.
[0698] In-loop filtering, saving of the reference picture, and motion compensation can be performed in the inversely mapped area.
[0699] Prediction blocks generated through inter-prediction can be converted to a mapped region by mapping using a mapping function, and the converted prediction blocks can be used to generate reconstructed blocks. However, since intra-prediction is performed in a mapped region, prediction blocks generated by intra-prediction can be used to generate reconstructed blocks without mapping and / or inverse mapping.
[0700] For example, if the target block is a residual block of chroma components, the residual block can be converted into an inversely mapped region by performing scaling on the chroma components of the mapped region.
[0701] Whether scaling is available can be signaled at the slice level or tile group level.
[0702] For example, scaling can be applied only when mapping for the luminance component is available and the splitting of the luminance component and the splitting of the chroma component follow the same tree structure.
[0703] Scaling can be performed based on the average of the values of the samples of the luminance prediction block corresponding to the chroma prediction block. In this case, if the target block uses inter-prediction, the luminance prediction block may refer to the mapped luminance prediction block.
[0704] The value required for scaling can be derived by referring to a lookup table using the index of the piece to which the average of the values of the samples of the Luma prediction block belongs.
[0705] By performing scaling on the residual block using the finally derived value, the residual block can be converted into an inversely mapped region. Subsequently, for the chroma component block, reconstruction, intra-prediction, inter-prediction, in-loop filtering, and saving of the reference picture can be performed in the inversely mapped region.
[0706] For example, information indicating whether mapping and / or inverse mapping of these luminal and chromal components is available can be signaled through a sequence parameter set.
[0707] The predicted block of the target block can be generated based on a block vector. The block vector can represent the displacement between the target block and the reference block. The reference block can be a block within the target image.
[0708] In this way, a prediction mode that generates a prediction block by referencing a target image can be called an Intra Block Copy (IBC) mode.
[0709] The IBC mode can be applied to CUs of a specific size. For example, the IBC mode can be applied to MxN CUs. Here, M and N can be less than or equal to 64.
[0710] IBC mode may include skip mode, merge mode, AMVP mode, etc. In the case of skip mode or merge mode, a merge candidate list may be configured, and one merge candidate from among the merge candidates in the merge candidate list may be specified by signaling a merge index. The block vector of the specified merge candidate may be used as the block vector of the target block.
[0711] In AMVP mode, the difference block vector can be signaled. Additionally, the predicted block vector can be derived from the left and top neighbor blocks of the target block. Furthermore, an index regarding which neighbor block will be used can be signaled.
[0712] The predicted block in IBC mode may be included in the target CTU or the left CTU and may be limited to a block within a pre-reconstructed area. For example, the value of the block vector may be restricted so that the predicted block of the target block is located within a specified area. The specified area may be an area of three 64x64 blocks that are encoded and / or decoded before the 64x64 block containing the target block. By restricting the value of the block vector in this way, memory consumption and device complexity associated with the implementation of IBC mode can be reduced.
[0713] FIG. 16 is a structural diagram of an encoding device according to one embodiment.
[0714] The encoding device (1600) can correspond to the aforementioned encoding device (100).
[0715] The encoding device (1600) may include a processing unit (1610), a memory (1630), a user interface (UI) input device (1650), a UI output device (1660), and a storage (1640) that communicate with each other via a bus (1690). Additionally, the encoding device (1600) may further include a communication unit (1620) connected to a network (1699).
[0716] The processing unit (1610) may be a semiconductor device that executes processing instructions stored in a Central Processing Unit (CPU), memory (1630), or storage (1640). The processing unit (1610) may be at least one hardware processor.
[0717] The processing unit (1610) can perform the generation and processing of signals, data, or information that are input to the encoding device (1600), output from the encoding device (1600), or used internally within the encoding device (1600), and can perform inspection, comparison, and judgment related to the signals, data, or information. That is to say, in the embodiment, the generation and processing of data or information, and the inspection, comparison, and judgment related to the data or information can be performed by the processing unit (1610).
[0718] The processing unit (1610) may include an inter prediction unit (110), an intra prediction unit (120), a switch (115), a subtractor (125), a conversion unit (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse conversion unit (170), an adder (175), a filter unit (180), and a reference picture buffer (190).
[0719] At least some of the inter-prediction unit (110), intra-prediction unit (120), switch (115), subtractor (125), converter (130), quantizer (140), entropy encoding unit (150), inverse quantizer (160), inverse converter (170), adder (175), filter unit (180), and reference picture buffer (190) may be program modules and may communicate with an external device or system. The program modules may be included in the encoding device (1600) in the form of an operating system, an application program module, and other program modules.
[0720] Program modules may be physically stored on various known memory devices. Additionally, at least some of these program modules may be stored in a remote memory device capable of communicating with the encoding device (1600).
[0721] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures that perform functions or operations according to one embodiment or implement abstract data types according to one embodiment.
[0722] Program modules may consist of instructions or code executed by at least one processor of the encoding device (1600).
[0723] The processing unit (1610) can execute instructions or code of the inter prediction unit (110), intra prediction unit (120), switch (115), subtractor (125), converter (130), quantization unit (140), entropy encoding unit (150), inverse quantization unit (160), inverse conversion unit (170), adder (175), filter unit (180), and reference picture buffer (190).
[0724] The storage unit may represent memory (1630) and / or storage (1640). Memory (1630) and storage (1640) may be various forms of volatile or non-volatile storage media. For example, memory (1630) may include at least one of ROM (1631) and RAM (1632).
[0725] The storage unit may store data or information used for the operation of the encoding device (1600). In an embodiment, the data or information possessed by the encoding device (1600) may be stored in the storage unit.
[0726] For example, the storage unit can store pictures, blocks, lists, motion information, inter-prediction information, and bitstreams.
[0727] The encoding device (1600) can be implemented in a computer system including a recording medium that can be read by a computer.
[0728] The recording medium can store at least one module required for the encoding device (1600) to operate. The memory (1630) can store at least one module and can be configured so that at least one module is executed by the processing unit (1610).
[0729] Functions related to the communication of data or information of the encoding device (1600) can be performed through the communication unit (1620).
[0730] For example, the communication unit (1620) can transmit the bitstream to the decoding device (1700) to be described later.
[0731] FIG. 17 is a structural diagram of a decoding device according to one embodiment.
[0732] The decoding device (1700) can correspond to the aforementioned decoding device (200).
[0733] The decoding device (1700) may include a processing unit (1710), a memory (1730), a user interface (UI) input device (1750), a UI output device (1760), and a storage (1740) that communicate with each other via a bus (1790). Additionally, the decoding device (1700) may further include a communication unit (1720) connected to a network (1799).
[0734] The processing unit (1710) may be a semiconductor device that executes processing instructions stored in a Central Processing Unit (CPU), memory (1730), or storage (1740). The processing unit (1710) may be at least one hardware processor.
[0735] The processing unit (1710) can perform the generation and processing of signals, data, or information that are input to the decoding device (1700), output from the decoding device (1700), or used internally within the decoding device (1700), and can perform inspection, comparison, and judgment related to the signals, data, or information. That is to say, in the embodiment, the generation and processing of data or information, and the inspection, comparison, and judgment related to the data or information can be performed by the processing unit (1710).
[0736] The processing unit (1710) may include an entropy decoder (210), an inverse quantizer (220), an inverse transform unit (230), an intra prediction unit (240), an inter prediction unit (250), a switch (245), an adder (255), a filter unit (260), and a reference picture buffer (270).
[0737] At least some of the entropy decoder (210), inverse quantizer (220), inverse transform (230), intra prediction (240), inter prediction (250), switch (245), adder (255), filter (260), and reference picture buffer (270) may be program modules and may communicate with an external device or system. The program modules may be included in the decoder (1700) in the form of an operating system, an application program module, and other program modules.
[0738] The program modules may be physically stored on various known memory devices. Additionally, at least some of these program modules may be stored in a remote memory device capable of communicating with the decoding device (1700).
[0739] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures that perform functions or operations according to one embodiment or implement abstract data types according to one embodiment.
[0740] Program modules may consist of instructions or code executed by at least one processor of the decoding device (1700).
[0741] The processing unit (1710) can execute instructions or code of the entropy decoder (210), inverse quantizer (220), inverse transform unit (230), intra prediction unit (240), inter prediction unit (250), switch (245), adder (255), filter unit (260), and reference picture buffer (270).
[0742] The storage unit may represent memory (1730) and / or storage (1740). Memory (1730) and storage (1740) may be various forms of volatile or non-volatile storage media. For example, memory (1730) may include at least one of ROM (1731) and RAM (1732).
[0743] The storage unit may store data or information used for the operation of the decoding device (1700). In an embodiment, the data or information possessed by the decoding device (1700) may be stored in the storage unit.
[0744] For example, the storage unit can store pictures, blocks, lists, motion information, inter-prediction information, and bitstreams.
[0745] The decoding device (1700) can be implemented in a computer system including a recording medium that can be read by a computer.
[0746] The recording medium can store at least one module required for the decoding device (1700) to operate. The memory (1730) can store at least one module and can be configured so that at least one module is executed by the processing unit (1710).
[0747] Functions related to the communication of data or information of the decoding device (1700) can be performed through the communication unit (1720).
[0748] For example, the communication unit (1720) can receive a bitstream from the encoding device (1600).
[0749] Appinea transform for video encoding / decoding
[0750] In the embodiments, an image encoding / decoding method may be provided that performs inter-prediction through an affine transform suitable for the distance of a reference picture.
[0751] In the embodiments, an image encoding / decoding method can be provided that performs inter-prediction using an appropriate affine transform depending on the conditions in encoding / decoding.
[0752] In the embodiments, to improve the prediction method using an affine transform inter-prediction method, an image encoding / decoding method may be provided that generates a prediction block through an affine transform AMVP and / or affine transform merge using the distance of a reference block.
[0753] In the embodiments, a block may mean a unit. Additionally, a block may mean a unit.
[0754] Figures 18 and 19 show control point motion vectors according to the number of parameters of an affine transformation model according to one example.
[0755] Figure 18 shows the control point motion vector of a 4-parameter affine transformation model according to one example.
[0756] Figure 19 shows the control point motion vector of a 6-parameter affine transformation model according to one example.
[0757] Affine transformations can be used to predict various movements that occur in reality. For example, affine transformations can be used to predict zoom in / out, rotation, and irregular movements in images.
[0758] The Control Point Motion Vector (CPMV) can be used to derive motion vectors for a target block, sub-blocks divided into N units, or pixels.
[0759] N can be a positive integer. A sub-block divided into N units can mean an NxN block.
[0760] In Figure 18, CPMVs of a 4-parameter affine transformation model are exemplified. In a 4-parameter affine transformation model, two CPMVs may be used. The two CPMVs may include the CPMV at the top left of the block and the CPMV at the top right of the block.
[0761] In Figure 19, CPMVs of a 6-parameter affine transformation model are exemplified. In a 6-parameter affine transformation model, three CPMVs may be used. The three CPMVs may include the CPMV at the top left of the block, the CPMV at the top right of the block, and the CPMV at the bottom left of the block.
[0762] The 4-parameter affine transformation model used to generate prediction blocks can be expressed as shown in Equations 1 and 2 below.
[0763] [Formula 1]
[0764] mv_x = (mv_1x - mv_0x) * x / W + (mv_1y - mv_0y) * y / W + mv_0x
[0765] [Equation 2]
[0766] mv_y = (mv_1y - mv_0y) * x / W + (mv_1x - mv_0x) * y / W + mv_0y
[0767] The 6-parameter affine transformation model used to generate prediction blocks can be expressed as shown in Equations 3 and 4 below.
[0768] [Equation 3]
[0769] mv_x = (mv_1x - mv_0x) * x / W + (mv_2x - mv_0x) * y / H + mv_0x
[0770] [Equation 4]
[0771] mv_y = (mv_1y - mv_0y) * x / W + (mv_2y - mv_0y) * y / H + mv_0y
[0772] W and H may represent the width and height of the target block, respectively. For example, the target block may be a coding unit (CU), a coding block, a prediction unit, or a prediction block.
[0773] x and y can represent the location (i.e., coordinates) of the target pixel within the target block. The top-left coordinates of the target block can be considered as (0,0), the top-right coordinates as (W-1, 0), the bottom-left coordinates as (0, H-1), and the bottom-right coordinates as (W-1, H-1).
[0774] (mv_0x, mv_0y) can represent the x-direction and y-direction values of the upper-left CPMV represented by v0 in FIG. 18 and FIG. 19, respectively. (mv_1x, mv_1y) can represent the x-direction and y-direction values of the upper-right CPMV represented by v1 in FIG. 18 and FIG. 19, respectively. (mv_2x, mv_2y) can represent the x-direction and y-direction values of the lower-left CPMV represented by v2 in FIG. 19.
[0775] mv_x and mv_y may represent the x-direction and y-direction values of the motion vector used to generate the prediction block. Here, the x-direction values and y-direction values may mean the x-component and y-component.
[0776] Information related to Affine Mode
[0777] Affine conversion available information
[0778] High-level syntax elements such as the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), and Decoding Parameter Set (DPS), picture header, tile header, tile group header, and slice header may include affine conversion availability information indicating whether affine conversion is available.
[0779] If the affine transformation availability information indicates that the affine transformation is not available, the use of the affine transformation may not be determined for individual blocks. If the affine transformation availability information indicates that the affine transformation is available, the use of the affine transformation for the said block may be determined again for individual blocks.
[0780] The fact that affine conversion is not available may mean that affine conversion is not used (in bulk) for lower-level targets affected by higher-level syntactic elements.
[0781] The availability of affine transformations may mean that the use of affine transformations is determined individually for lower-level targets affected by higher-level syntactic elements.
[0782] Affine conversion availability information can be signaled from the encoding device (1600) to the decoding device (1700) as information within a higher-level syntax element. For example, the name of the affine conversion availability information may be "affine_flag", and the affine conversion availability information may be a flag.
[0783] The adaptive parameter set may refer to a parameter set referenced by multiple pictures, multiple sub-pictures, multiple tile groups, multiple tiles, multiple slices, and multiple CTUs, etc.
[0784] Information on using affine conversion
[0785] Affine usage information can indicate whether an affine transformation is used for the target block. Affine usage information can be signaled on a block or unit basis. For example, the name of the affine usage information can be "inter_affine_flag". The affine usage information can be a flag.
[0786] An affine transformation usage information being the first value may indicate that no affine transformation is used for the target block. For example, if the affine transformation usage information is the first value, an inter prediction that does not use an affine transformation for the target block may be performed.
[0787] A second value for the affine transformation usage information may indicate that an affine transformation is used for the target block. For example, if the affine transformation usage information is a second value, an inter prediction using an affine transformation for the target block may be performed.
[0788] In the embodiments, for example, the first value may mean "0" or "false". The second value may mean "1" or "true". In the embodiments, the first value and the second value may be merely examples of different values. The first value for specific information may be different from the first value for other information. The second value for specific information may be different from the second value for other information.
[0789] Affine conversion usage information can be signaled from the encoding device (1600) to the decoding device (1700) for a target block. If the affine conversion is set not to be available in the upper-level syntax element (for example, if the affine conversion availability information is set to the first value), the signaling of the affine conversion usage information may be omitted, and the affine conversion usage information may be set to the first value.
[0790] Affine conversion mode information
[0791] In an affine transformation, among multiple affine transformation modes of inter-prediction using affine transformation, a selected affine transformation mode may be used for the target block.
[0792] Multiple affine transform modes may include an affine transform merge mode and an affine transform AMVP mode. The affine transform merge mode may be a mode in which inter-prediction for a target block is performed using affine transform merge. The affine transform AMVP mode may be a mode in which inter-prediction for a target block is performed using affine transform AMVP.
[0793] Affine transformation mode information may indicate the affine transformation mode used for the target block among multiple affine transformation modes that use affine transformation. Affine transformation mode information may indicate whether the affine transformation merge mode or the affine transformation AMVP mode is used for the target block to which the affine transformation is applied. The name of the affine transformation mode information may be "general_merge_flag". The affine transformation mode information may be a flag.
[0794] For example, if the affine transformation mode information is the first value, the affine transformation AMVP mode can be used. If the affine transformation mode information is the first value, a prediction block for the target block can be generated using the affine transformation AMVP.
[0795] For example, if the affine transformation mode information is the second value, the affine transformation merge mode may be used. If the affine transformation mode information is the second value, a prediction block for the target block may be generated using affine transformation merge.
[0796] Encoding / decoding using the affine conversion AMVP mode is described with reference to FIGS. 20 and 21. Encoding / decoding using the affine conversion merge mode is described with reference to FIGS. 22 and 23.
[0797] Affine Transform Model Information
[0798] In affine transformation, a selected affine transformation model among multiple affine transformation models can be used for the target block.
[0799] For example, the name of the affine transformation model information can be "MotionModelIdc". MotionModelIdc can represent the affine transformation model type. The affine transformation model information can be a flag.
[0800] For example, an affine transformation model information being a first value (e.g., "0") may indicate that translational motion is used as a motion model for motion compensation. An affine transformation model information being a second value (e.g., "1") may indicate that a 4-parameter affine transformation model is used as a motion model for motion compensation. An affine transformation model information being a third value (e.g., "2") may indicate that a 6-parameter affine transformation model is used as a motion model for motion compensation.
[0801] Multiple affine transformation models may include a 4-parameter affine transformation model and a 6-parameter affine transformation model. The 4-parameter affine transformation model may be a model in which an affine transformation using 4 parameters is performed on a target block. The 6-parameter affine transformation model may be a model in which an affine transformation using 6 parameters is performed on a target block.
[0802] Affine transformation model information can indicate the affine transformation model used for the target block among multiple affine transformation models. Affine transformation model information can indicate whether the 4-parameter affine transformation model or the 6-parameter affine transformation model is used for the target block to which the affine transformation mode is applied. The name of the affine transformation model information can be "affine_type_flag". The affine transformation model information can be a flag.
[0803] For example, if the affine transformation model information is the first value, a 4-parameter affine transformation model can be used. If the affine transformation model information is the first value, encoding / decoding for the target block can be performed using a 4-parameter affine transformation model.
[0804] For example, if the affine transform model information is the second value, a 6-parameter affine transform model can be used. If the affine transform model information is the second value, encoding / decoding for the target block can be performed using a 6-parameter affine transform model.
[0805] As another example, the name of the affine transformation type information can be "affine_type_flag". The affine transformation type information can be a flag. The name of the affine transformation mode information can be "general_merge_flag". The affine transformation mode information can be a flag. The name of the affine transformation usage information can be "inter_affine_flag". The affine transformation usage information can be a flag. The name of the merge sub-block information is It can be "merge_subblock_flag". The merge subblock information can be a flag. Depending on the value of the affine transformation mode information, the method of deriving affine transformation model information can be divided.
[0806] For example, if the affine transform mode information is a first value (e.g., "0"), the affine transform model information can be determined according to the value of the merge sub-block information. As one example, if the merge sub-block information is a first value (e.g., "0"), a translational movement can be used. As another example, if the merge sub-block information is a second value (e.g., "1"), encoding / decoding for the target block can be performed using a 4-parameter affine transform model.
[0807] For example, if the affine transformation mode is a first value (e.g., "0"), the affine transformation model information can be determined based on the sum of the affine transformation type information and the affine transformation usage information. As an example, if the affine transformation usage information is a first value (e.g., "0") and / or the affine transformation type information is a first value (e.g., "0"), encoding / decoding of the target block can be performed using translational motion.
[0808] For example, if the affine transformation mode is a first value (e.g., "0"), the affine transformation model information can be determined based on the sum of the affine transformation type information and the affine transformation usage information. As an example, if the affine transformation usage information is a second value (e.g., "1") and / or the affine transformation type information is a first value (e.g., "0"), encoding / decoding for the target block can be performed using a 4-parameter affine transformation model.
[0809] For example, if the affine transformation mode is a first value (e.g., "0"), the affine transformation model information can be determined based on the sum of the affine transformation type information and the affine transformation usage information. As an example, if the affine transformation usage information is a second value (e.g., "1") and / or the affine transformation type information is a second value (e.g., "1"), encoding / decoding for the target block can be performed using a 6-parameter affine transformation model.
[0810] Encoding and decoding using affine transformation
[0811] Figure 20 shows an encoding method using an affine conversion AMVP mode according to one example.
[0812] A prediction block for a target block can be generated using an affine transformation AMVP by steps (2010, 2020, and 2030).
[0813] In step (2010), the processing unit (1610) of the encoding device (1600) can perform the construction of an affine transform AMVP control point prediction list and / or the derivation of a control point motion vector.
[0814] In step (2020), the processing unit (1610) can generate a motion vector of a target block for an affine transformation prediction block using a control point motion vector.
[0815] An affine transformation prediction block can refer to a prediction block generated by an affine transformation.
[0816] In step (2030), the processing unit (1610) can perform derivation of the affine transform model and / or encoding of the motion vector difference of the control point.
[0817] The following detailed descriptions may apply to the actions of the steps (2010, 2020, and 2030).
[0818] Figure 21 shows a decoding method using an affine conversion AMVP mode according to one example.
[0819] A prediction block for a target block can be generated using an affine transformation AMVP by steps (2110, 2120, and 2130).
[0820] In step (2110), the processing unit (1710) of the decoding device (1700) can perform decoding for the derivation of the affine transform model and / or the motion vector difference of the control point.
[0821] In step (2120), the processing unit (1710) can perform the construction of the affine transformation AMVP control point prediction list and / or the derivation of the control point motion vector.
[0822] In step (2130), the processing unit (1710) can generate a motion vector of a target block for an affine transformation prediction block using a control point motion vector.
[0823] The following detailed descriptions may apply to the operations of steps (2110, 2120, and 2130).
[0824] Figure 22 illustrates an encoding method using an affine transform merge mode according to one example.
[0825] A prediction block for a target block can be generated using affine transformation merge by steps (2210, 2220, and 2230).
[0826] In step (2210), the processing unit (1610) of the encoding device (1600) can perform the construction of an affine transform merge control point prediction list and / or the derivation of a control point motion vector.
[0827] In step (2220), the processing unit (1610) can generate a motion vector of a target block for an affine transformation prediction block using a control point motion vector.
[0828] In step (2230), the processing unit (1610) can perform encoding for the information of the affine transform merge mode. The information of the affine transform merge mode may include flags and / or indices.
[0829] The following detailed descriptions may apply to the operations of steps (2210, 2220, and 2230).
[0830] Figure 23 shows a decoding method using an affine transform merge mode according to one example.
[0831] A prediction block for a target block can be generated using affine transformation merging by steps (2310, 2320, and 2330).
[0832] In step (2310), the processing unit (1710) of the decoding device (1700) can perform decoding on the information of the affine transform merge mode. The information of the affine transform merge mode may include flags and / or indices.
[0833] In step (2320), the processing unit (1710) can construct an affine transform merge control point prediction list and / or derive a control point motion vector.
[0834] In step (2330), the processing unit (1610) can generate a motion vector of a target block for an affine transformation prediction block using a control point motion vector.
[0835] The following detailed descriptions may apply to the operations of steps (2310, 2320, and 2330).
[0836] Operations for affine transformation
[0837] Below, the operation in the steps of the aforementioned methods is described in more detail. In the description below, the processing unit may refer to the processing unit (1610) of the encoding device (1600) and / or the processing unit (1710) of the decoding device (1700).
[0838] Configuration of Affine Transformation AMVP Control Point Prediction List
[0839] The following descriptions may be applied to the configuration of the affine transformation AMVP control point prediction list in steps (2010) and (2120).
[0840] To calculate the predicted value of the affine transformation control point, the method described with reference to FIG. 25 may be used.
[0841] In the steps (2510, 2520, and 2530) described with reference to FIG. 25, the predicted value of the affine transformation control point can be calculated as the affine transformation AMVP control point prediction list cpMvpListLX[numCpMvpCandLX][i] is used.
[0842] "cpMvpListL X " and "numCpMvpCandL X "of " X " can represent the reference picture that the target block refers to. X " can be a first value (e.g., "0") or a second value (e.g., "1").
[0843] "numCpMvpCandLX" can represent an index to construct cpMvpListLX. numCpMvpCandLX can represent the position within cpMvpListLX where a new element is added. This index can limit the maximum number of elements in the list.
[0844] "cpMvpListLX[numCpMvpCandLX][ i In ]", "i" may be an index representing the location of an affine transformation control point.
[0845] cpMvpListLX can be configured in steps (2510, 2520, and 2530) of FIG. 25.
[0846] FIG. 24 shows the location of adjacent reference pixels that can identify spatial candidates for a target block according to one example.
[0847] A spatial candidate may be a block containing adjacent reference pixels.
[0848] The adjacent reference pixels may include at least one of 1) A0, which is adjacent to the bottom-left of the target block; 2) A1, which is the bottommost pixel among the pixels adjacent to the left of the target block; 3) A2, which is the topmost pixel among the pixels adjacent to the left of the target block; 4) B0, which is adjacent to the top-right of the target block; 5) B1, which is the rightmost pixel among the pixels adjacent to the top of the target block; 6) B2, which is adjacent to the top-left of the target block; and 6) B3, which is the leftmost pixel among the pixels adjacent to the top of the target block. Alternatively, A1 may be a pixel adjacent to the top of A0. A2 may be a pixel adjacent to the bottom of B2. B3 may be a pixel adjacent to the right of B2. B1 may be a pixel adjacent to the left of B0.
[0849] Information about a reference block occupying the location of an adjacent reference pixel can be used as information about a spatial candidate. Alternatively, the reference block may be a block corresponding to the location of an adjacent reference pixel.
[0850] For example, an encoded / decoded reference block occupying the location of an adjacent reference pixel of a target block can be utilized as a spatial candidate, etc. The encoded / decoded reference block may be a reconstructed block.
[0851] Information about the reference block may include the size of the reference block, the x-coordinate of the reference block, the y-coordinate of the reference block, and inter-prediction information about the reference block.
[0852] Hereinafter, reference block X may refer to a block occupying adjacent reference pixel X. X may be A0, A1, A2, B0, B1, B2, or B3. Alternatively, the aforementioned A0, A1, A2, B0, B1, B2, and B3 may be understood as reference blocks adjacent to the target block, rather than adjacent pixels.
[0853] In the embodiments, the terms "adjacent" and "neighbor" may be used interchangeably. Additionally, the term "adjacent" may mean that the vertical distance and / or horizontal distance between objects is less than or equal to a predetermined value.
[0854] Figure 25 illustrates a method for calculating the predicted value of an affine transformation control point according to one example.
[0855] In step (2510), the processing unit can construct an affine transformation AMVP control point prediction list using spatial candidates.
[0856] In one embodiment, in the configuration of the affine transformation AMVP control point prediction list, a reference block corresponding to the location of an adjacent reference pixel of a target block described with reference to FIG. 24 may be utilized as a spatial candidate.
[0857] In deriving affine transformation control points, reference block A0 and reference block A1 can be considered as a single group.
[0858] An affine transformation control point for a pixel can be derived and used through the "derivation of an affine transformation control point" described later in the order of reference block A0 and reference block A1.
[0859] An affine transformation control point of reference block A0 or reference block A1 may be derived to fit the size of the target block. The derived affine transformation control point may be allocated and / or stored within cpMvpListLX. When one derived affine transformation control point is added to cpMvpListLX, the reference block that follows in the above order may not be used. In this specification, "existent" may mean "available."
[0860] In the embodiment, the fact that the reference block is not used may mean that the reference block is not referenced to derive an affine transformation control point.
[0861] In the example, the fact that the affine transformation control point of the reference block is used may mean that the affine transformation control point is derived to fit the size of the target block, and the derived affine transformation control point is added to cpMvpListLX.
[0862] For example, if an affine transformation control point of reference block A0 exists, reference block A1 may not be used.
[0863] For example, if the affine transformation control point of reference block A0 does not exist, the existence of the affine transformation control point of reference block A1 can be checked, and if the affine transformation control point of reference block A1 exists, the affine transformation control point of reference block A1 can be used.
[0864] In deriving the affine transformation control point, reference block B0, reference block B1, and reference block B2 can be considered as one group.
[0865] An affine transformation control point for a pixel can be derived and used through the "derivation of an affine transformation control point" described later in the order of reference block B0, reference block B1, and reference block B2.
[0866] Affine transformation control points of reference block B0, reference block B1, or reference block B2 can be derived to fit the size of the target block. The derived affine transformation control points can be allocated and / or stored within cpMvpListLX. When one derived affine transformation control point is added to cpMvpListLX, the reference block that follows in the above order may not be used for deriving the affine transformation control point.
[0867] For example, if an affine transformation control point of reference block B0 exists, reference blocks B1 and B2 may not be used.
[0868] For example, if the affine transformation control point of reference block B0 and the affine transformation control point of reference block B1 do not exist, the existence of the affine transformation control point of reference block B2 can be checked, and if the affine transformation control point of reference block B2 exists, reference block B2 can be used.
[0869] In one embodiment, in the configuration of the affine transformation AMVP control point prediction list, whether a reference block is utilized can be determined based on the coding parameters of the reference block.
[0870] Whether the affine transformation control points of a reference block are derived can be determined based on the coding parameters for the reference block. For example, if the picture referenced by the reference block is a long-term reference picture, the affine transformation control points of the reference block may not be derived.
[0871] In deriving affine transformation control points, reference block A0 and reference block A1 can be considered as a single group.
[0872] An affine transformation control point for a pixel can be derived and used through the "derivation of an affine transformation control point" described later in the order of reference block A0 and reference block A1.
[0873] If the picture referenced by reference block A0 is a long-term reference picture, the affine transformation control point of reference block A0 may not be derived. If the picture referenced by reference block A1 is a long-term reference picture, the affine transformation control point of reference block A1 may not be derived.
[0874] If the picture referenced by reference block A0 is a short-term reference picture, the affine transformation control points of reference block A0 can be derived to fit the size of the target block. If the picture referenced by reference block A1 is a short-term reference picture, the affine transformation control points of reference block A1 can be derived to fit the size of the target block.
[0875] Derived affine transformation control points may be allocated and / or stored within cpMvpListLX. When one derived affine transformation control point is added to cpMvpListLX, the reference block that follows in the above order may not be used.
[0876] For example, if an affine transformation control point of reference block A0 exists, reference block A1 may not be used.
[0877] For example, if the affine transformation control point of reference block A0 does not exist, the existence of the affine transformation control point of reference block A1 can be checked, and if the affine transformation control point of reference block A1 exists, the affine transformation control point of reference block A1 can be used.
[0878] For example, if the picture referenced by reference block A0 is a long-term reference picture, the affine transformation control point of reference picture A0 may not be used. If the picture referenced by reference block A1 is a short-term reference picture, the existence of the affine transformation control point of reference block A1 may be checked, and if the affine transformation control point of reference block A1 exists, the affine transformation control point of reference block A1 may be used.
[0879] For example, if the picture referenced by reference block A0 is a long-term reference picture and the picture referenced by reference block A1 is a long-term reference picture, the affine transformation control points of reference block A0 and reference block A1 may not be used.
[0880] In deriving the affine transformation control point, reference block B0, reference block B1, and reference block B2 can be considered as one group.
[0881] An affine transformation control point for a pixel can be derived and used through the "derivation of an affine transformation control point" described later in the order of reference block B0, reference block B1, and reference block B2.
[0882] If the picture referenced by reference block B0 is a long-term reference picture, the affine transformation control point of reference block B0 may not be derived. If the picture referenced by reference block B1 is a long-term reference picture, the affine transformation control point of reference block B1 may not be derived. If the picture referenced by reference block B2 is a long-term reference picture, the affine transformation control point of reference block B2 may not be derived.
[0883] If the picture referenced by reference block B0 is a short-term reference picture, the affine transformation control points of reference block B0 can be derived to fit the size of the target block. If the picture referenced by reference block B1 is a short-term reference picture, the affine transformation control points of reference block B1 can be derived to fit the size of the target block. If the picture referenced by reference block B2 is a short-term reference picture, the affine transformation control points of reference block B2 can be derived to fit the size of the target block.
[0884] Derived affine transformation control points may be allocated and / or stored within cpMvpListLX. When one derived affine transformation control point is added to cpMvpListLX, the reference block that follows in the above order may not be used.
[0885] For example, if an affine transformation control point of reference block B0 exists, reference blocks B1 and B2 may not be used.
[0886] For example, if the affine transformation control point of reference block B0 and the affine transformation control point of reference block B1 do not exist, the existence of the affine transformation control point of reference block B2 can be checked, and if the affine transformation control point of reference block B2 exists, reference block B2 can be used.
[0887] For example, if the picture referenced by reference block B0 is a long-term reference picture and the picture referenced by reference block B1 is a long-term reference picture, the affine transformation control points of reference picture B0 and the affine transformation control points of reference picture B1 may not be used. If the picture referenced by reference block B2 is a short-term reference picture, the existence of the affine transformation control points of reference block B2 may be checked, and if the affine transformation control points of reference block B2 exist, reference block B2 may be used.
[0888] For example, if the picture referenced by reference block B0 is a long-term reference picture, the affine transformation control point of reference block B0 may not be used. Next, if the picture referenced by reference block B1 is a short-term reference picture, the existence of the affine transformation control point of reference block B1 may be checked, and if the affine transformation control point of reference block B1 exists, the affine transformation control point of reference block B1 may be used. At this time, if the affine transformation control point of reference block B1 is used, the affine transformation control point of reference block B2 may not be used.
[0889] In step (2520), the processing unit may construct an affine transformation AMVP control point prediction list in which control point motion vector combinations are determined. "Affine transformation AMVP control point prediction list in which control point motion vector combinations are determined" may mean that control point motion vector combinations have been added to the affine transformation AMVP control point prediction list.
[0890] "Affine Transform AMVP Control Point Prediction List" may be an affine transformation control point prediction list used in affine transformation AMVP mode.
[0891] In one embodiment, the control point motion vector combination may be a pair of ordered CPMVs.
[0892] CPMVk can represent the location of an affine transformation control point containing CPMV. Alternatively, CPMVk can refer to a group of affine transformation control points from which CPMV is extracted. "CPMV k "at " k " can have a value from 1 to 3.
[0893] In CPMV1, a motion vector of one of reference block B2, reference block B3, and reference block A2 may be assigned and / or stored. In this case, the assigned motion vector may be selected in the order of the motion vector of reference block B2, the motion vector of reference block B3, and the motion vector of reference block A2.
[0894] In CPMV2, a motion vector of either reference block B1 or reference block B0 may be assigned and / or stored. In this case, the assigned motion vector may be selected in the order of the motion vector of reference block B1 and the motion vector of reference block B0.
[0895] In CPMV3, a motion vector of either reference block A1 or reference block A0 may be assigned and / or stored. In this case, the assigned motion vector may be selected in the order of the motion vector of reference block A1 and the motion vector of reference block A0.
[0896] When a 6-parameter affine transformation model is used, the control point motion vector combination can be defined as (CPMV1, CPMV2, CPMV3). If at least one of the three CPMVs of the control point motion vector combination is not available, the control point motion vector combination may not be added to the affine transformation AMVP control point prediction list.
[0897] When a 4-parameter affine transformation model is used, the control point motion vector combination can be defined as (CPMV1, CPMV2). If at least one of the two CPMVs of the control point motion vector combination is not available, the control point motion vector combination may not be added to the affine transformation AMVP control point prediction list.
[0898] For example, if reference block B2 is not available and reference block B3 is available, the motion vector of reference block B3 can be assigned and / or stored in CPMV1.
[0899] For example, if reference block B1 is available, the availability of reference block B0 may not be checked, and the motion vector of reference block B1 may be assigned and / or stored in CPMV2.
[0900] For example, if reference block A1 is not available and reference block A0 is available, the motion vector of reference block A0 can be assigned and / or stored in CPMV3.
[0901] For example, when a 6-parameter affine transformation model is used, if reference block B2, reference block B3, and reference block A2 are all unavailable, CPMV1 may not be available. In this case, even if CPMV2 and CPMV3 are available, (CPMV1, CPMV2, CPMV3) may not be added to the affine transformation AMVP control point prediction list.
[0902] For example, when a 4-parameter affine transformation model is used, if CPMV1 is available but reference block B1 and reference block B0, which are referenced for CPMV2, are both not available, (CPMV1, CPMV2) may not be added to the affine transformation AMVP control point prediction list.
[0903] Depending on the affine transformation model type, CPMVk may be added to the affine transformation AMVP control point prediction list, and the added CPMVk may be used for "construction of the affine transformation AMVP control point prediction list and / or derivation of control point motion vectors" of step (2010) and step (2120).
[0904] For example, if MotionModelIdc is the second value, a 4-parameter affine transform model may be used. In this case, the maximum number of candidates (i.e., control point motion vector combinations) in the affine transform AMVP control point prediction list may be 2. If there are no previously stored candidates in the affine transform AMVP control point prediction list, numCPMVpCandL0 may be 0. numCPMVpCandL0 may represent the number of control point motion vector combinations in the affine transform AMVP control point prediction list. Since numCPMVpCandL0 is 0, CPMV1 may be assigned and / or stored in CPMVpListL0[0][0] (the first CPMV of the first control point motion vector combination in the affine transform AMVP control point prediction list). Because a candidate has been added to the affine transform AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0905] For example, if MotionModelIdc is a third value, a 6-parameter affine transformation model may be used. In this case, the maximum number of candidates (i.e., combinations of control point motion vectors) in the affine transformation AMVP control point prediction list may be 2. If there is 1 previously stored candidate in the affine transformation AMVP control point prediction list, numCPMVpCandL0 may be 1. In this case, CPMV1 may be assigned and / or stored in CPMVpListL0[1][0], CPMV2 may be assigned and / or stored in CPMVpListL0[1][1], and CPMV3 may be assigned and / or stored in CPMVpListL0[1][2]. Because a candidate has been added to the affine transformation AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0906] For example, if MotionModelIdc is the second value, a 4-parameter affine transformation model may be used. In this case, the maximum number of candidates (i.e., combinations of control point motion vectors) in the affine transformation AMVP control point prediction list may be 2. If there is 1 previously stored candidate in the affine transformation AMVP control point prediction list, numCPMVpCandL0 may be 1. In this case, CPMV1 may be assigned and / or stored in CPMVpListL0[1][0], and CPMV2 may be assigned and / or stored in CPMVpListL0[1][1]. Because a candidate has been added to the affine transformation AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0907] For example, if MotionModelIdc is a third value, a 6-parameter affine transformation model may be used. In this case, the maximum number of candidates (i.e., combinations of control point motion vectors) in the affine transformation AMVP control point prediction list may be 2. If there are no previously stored candidates in the affine transformation AMVP control point prediction list, numCPMVpCandL0 may be 0. In this case, CPMV1 may be assigned and / or stored in CPMVpListL0[0][0], CPMV2 may be assigned and / or stored in CPMVpListL0[0][1], and CPMV3 may be assigned and / or stored in CPMVpListL0[0][2]. Because a candidate has been added to the affine transformation AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0908] In another embodiment, CPMVk may represent the location of an affine transformation control point having CPMV. Alternatively, CPMVk may mean a group of affine transformation control points from which CPMV is extracted. k "at kcan have a value from 1 to 3. If the picture referenced by the reference block is a long-term reference picture, the motion vector of the reference block may not be added to the affine transform AMVP control point prediction list.
[0909] In CPMV1, one of reference block B2, reference block B3, and reference block A2 may be assigned and / or stored as a motion vector. In this case, the assigned motion vector may be selected in the order of the motion vector of reference block B2, the motion vector of reference block B3, and the motion vector of reference block A2. In this case, the motion vector of the reference block referencing the long-term reference picture may not be selected.
[0910] In CPMV2, a motion vector of either reference block B1 or reference block B0 may be assigned and / or stored. In this case, the assigned motion vector may be selected in the order of the motion vector of reference block B1 and the motion vector of reference block B0. In this case, the motion vector of the reference block referencing the long-term reference picture may not be selected.
[0911] In CPMV3, a motion vector of either reference block A1 or reference block A0 may be assigned and / or stored. In this case, the assigned motion vector may be selected in the order of the motion vector of reference block A1 and the motion vector of reference block A0. In this case, the motion vector of the reference block referencing the long-term reference picture may not be selected.
[0912] When a 6-parameter affine transformation model is used, the control point motion vector combination can be defined as (CPMV1, CPMV2, CPMV3). If at least one of the three CPMVs in the control point motion vector combination is not available, or if the picture referenced by the reference block from which at least one CPMV was derived is a long-term reference picture, the control point motion vector combination may not be added to the affine transformation AMVP control point prediction list.
[0913] When a 4-parameter affine transformation model is used, the control point motion vector combination can be defined as (CPMV1, CPMV2). If at least one of the two CPMVs in the control point motion vector combination is unavailable, or if the picture referenced by the reference block from which at least one CPMV was derived is a long-term reference picture, the control point motion vector combination may not be added to the affine transformation AMVP control point prediction list.
[0914] For example, if reference block B2 is not available and reference block B3 is available, the motion vector of reference block B3 can be assigned and / or stored in CPMV1.
[0915] For example, if reference block B1 is available, the availability of reference block B0 may not be checked, and the motion vector of reference block B1 may be assigned and / or stored in CPMV2.
[0916] If reference block A1 is not available and reference block A0 is available, the motion vector of reference block A0 can be assigned and / or stored in CPMV3.
[0917] For example, if reference block B2 is available but the picture referenced by reference block B2 is a long-term reference picture, reference block B2 may not be used. If reference block B3 is available and the picture referenced by reference block B3 is a short-term reference picture, the motion vector of reference block B3 may be assigned and / or stored in CPMV1. If reference block B1 exists and the picture referenced by reference block B1 is a short-term reference picture, the availability of reference block B0 is not determined and the motion vector of reference block B1 may be assigned and / or stored in CPMV2. If reference block A1 is not available and reference block A0 is available and the picture referenced by reference block A0 is a short-term reference picture, the motion vector of A0 may be assigned and / or stored in CPMV3.
[0918] For example, when a 6-parameter affine transformation model is used, if reference block B2, reference block B3, and reference block A2 are all unavailable, CPMV1 may not be available. In this case, even if CPMV2 and CPMV3 are available, (CPMV1, CPMV2, CPMV3) may not be added to the affine transformation AMVP control point prediction list.
[0919] For example, when a 4-parameter affine transformation model is used, if CPMV1 is available but reference block B1 and reference block B0, which are referenced for CPMV2, are both not available, (CPMV1, CPMV2) may not be added to the affine transformation AMVP control point prediction list.
[0920] For example, when a 6-parameter affine transformation model is used, if reference block B2, reference block B3, and reference block A2 all refer to a long-term reference picture, CPMV1 may not be available. In this case, even if the reference block from which CPMV2 was derived and the reference block from which CPMV3 was derived both refer to a short-term reference picture and CPMV2 and CPMV3 are available, (CPMV1, CPMV2, CPMV3) may not be added to the affine transformation AMVP control point prediction list.
[0921] Depending on the type of affine transformation model, CPMVk may be added to the affine transformation AMVP control point prediction list, and the added CPMVk may be used for "construction of the affine transformation AMVP control point prediction list and / or derivation of the control point motion vector" of step (2010) and step (2120).
[0922] For example, if MotionModelIdc is the second value, a 4-parameter affine transform model may be used. In this case, the maximum number of candidates (i.e., control point motion vector combinations) in the affine transform AMVP control point prediction list may be 2. If there are no previously stored candidates in the affine transform AMVP control point prediction list, numCPMVpCandL0 may be 0. numCPMVpCandL0 may represent the number of control point motion vector combinations in the affine transform AMVP control point prediction list. Since numCPMVpCandL0 is 0, CPMV1 may be assigned and / or stored in CPMVpListL0[0][0] (the first CPMV of the first control point motion vector combination in the affine transform AMVP control point prediction list). Because a candidate has been added to the affine transform AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0923] For example, if MotionModelIdc is a third value, a 6-parameter affine transformation model may be used. In this case, the maximum number of candidates (i.e., combinations of control point motion vectors) in the affine transformation AMVP control point prediction list may be 2. If there is 1 previously stored candidate in the affine transformation AMVP control point prediction list, numCPMVpCandL0 may be 1. In this case, CPMV1 may be assigned and / or stored in CPMVpListL0[1][0], CPMV2 may be assigned and / or stored in CPMVpListL0[1][1], and CPMV3 may be assigned and / or stored in CPMVpListL0[1][2]. Because a candidate has been added to the affine transformation AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0924] For example, if MotionModelIdc is the second value, a 4-parameter affine transformation model may be used. In this case, the maximum number of candidates (i.e., combinations of control point motion vectors) in the affine transformation AMVP control point prediction list may be 2. If there is 1 previously stored candidate in the affine transformation AMVP control point prediction list, numCPMVpCandL0 may be 1. In this case, CPMV1 may be assigned and / or stored in CPMVpListL0[1][0], and CPMV2 may be assigned and / or stored in CPMVpListL0[1][1]. Because a candidate has been added to the affine transformation AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0925] For example, if MotionModelIdc is a third value, a 6-parameter affine transformation model may be used. In this case, the maximum number of candidates (i.e., combinations of control point motion vectors) in the affine transformation AMVP control point prediction list may be 2. If there are no previously stored candidates in the affine transformation AMVP control point prediction list, numCPMVpCandL0 may be 0. In this case, CPMV1 may be assigned and / or stored in CPMVpListL0[0][0], CPMV2 may be assigned and / or stored in CPMVpListL0[0][1], and CPMV3 may be assigned and / or stored in CPMVpListL0[0][2]. Because a candidate has been added to the affine transformation AMVP control point prediction list, numCPMVpCandLX may be increased by 1.
[0926] In step (2530), the processing unit can construct an affine transformation AMVP control point prediction list using one motion vector or zero motion vector.
[0927] In one embodiment, an affine transformation AMVP control point prediction list can be constructed using CPMV1, CPMV2, and CPMV3 derived in step (2520).
[0928] A 4-parameter affine transformation model is used, and if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list and CPMV1 is available, CPMV1 can be added to CPMVpListLX[numCPMVpCandLX][0] and CPMVPListLX[numCPMVpCandLX][1]. As a candidate is added, numCPMVpCandLX can be increased by 1. Next, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list and CPMV2 is available, CPMV2 can be added to CPMVpListLX[numCPMVpCandLX][0] and CPMVPListLX[numCPMVpCandLX][1]. As a candidate is added, numCPMVpCandLX can be increased by 1.
[0929] A 6-parameter affine transformation model is used, and if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list and CPMV1 is available, CPMV1 may be added to CPMVpListLX[numCPMVpCandLX][0], CPMVpListLX[numCPMVpCandLX][1], and CPMVPListLX[numCPMVpCandLX][2]. As a candidate is added, numCPMVpCandLX may be increased by 1. Next, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list and CPMV2 is available, CPMV2 may be added to CPMVpListLX[numCPMVpCandLX][0], CPMVpListLX[numCPMVpCandLX][1], and CPMVPListLX[numCPMVpCandLX][2]. As a candidate is added, numCPMVpCandLX may be increased by 1. Next, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list and CPMV3 is available, CPMV3 may be added to CPMVpListLX[numCPMVpCandLX][0], CPMVpListLX[numCPMVpCandLX][1], and CPMVPListLX[numCPMVpCandLX][2]. As a candidate is added, numCPMVpCandLX can increase by 1.
[0930] In one embodiment, an affine transformation AMVP control point prediction list can be constructed using CPMV1, CPMV2, and CPMV3 derived in step (2520). In this case, if the reference block from which the CPMV was derived does not reference a long-term reference picture, the CPMV can be added to the affine transformation AMVP control point prediction list.
[0931] If a 4-parameter affine transformation model is used and numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, CPMV1 is available, and the reference block from which CPMV1 is derived refers to a short-term reference picture, CPMV1 may be added to CPMVpListLX[numCPMVpCandLX][0] and CPMVPListLX[numCPMVpCandLX][1]. As a candidate is added, numCPMVpCandLX may be increased by 1. Next, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, CPMV2 is available, and the reference block from which CPMV2 is derived refers to a short-term reference picture, CPMV2 may be added to CPMVpListLX[numCPMVpCandLX][0] and CPMVPListLX[numCPMVpCandLX][1]. As a candidate is added, numCPMVpCandLX may be increased by 1.
[0932] If a 6-parameter affine transformation model is used and numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, CPMV1 is available, and the reference block from which CPMV1 is derived refers to a short-term reference picture, CPMV1 may be added to CPMVpListLX[numCPMVpCandLX][0], CPMVPListLX[numCPMVpCandLX][1], and CPMVpListLX[numCPMVpCandLX][2]. As a candidate is added, numCPMVpCandLX may be increased by 1. Next, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, CPMV2 is available, and the reference block from which CPMV2 is derived refers to a short-term reference picture, CPMV2 may be added to CPMVpListLX[numCPMVpCandLX][0], CPMVPListLX[numCPMVpCandLX][1], and CPMVpListLX[numCPMVpCandLX][2]. As a candidate is added, numCPMVpCandLX may be increased by 1. Next, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, CPMV3 is available, and the reference block from which CPMV3 is derived refers to a short-term reference picture, CPMV3 may be added to CPMVpListLX[numCPMVpCandLX][0], CPMVPListLX[numCPMVpCandLX][1], and CPMVpListLX[numCPMVpCandLX][2]. As a candidate is added, numCPMVpCandLX may be increased by 1.
[0933] A 4-parameter affine transformation model is used, and numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, and even if CPMV1 is available, if the reference block from which CPMV1 was derived refers to a long-term reference picture, CPMV1 may not be added to CPMVpListLX[numCPMVpCandLX][0] and CPMVPListLX[numCPMVpCandLX][1]. Next, since numCPMVpCandLX is still smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, CPMV2 is available, and if the reference block from which CPMV2 was derived refers to a short-term reference picture, CPMV2 may be added to CPMVpListLX[numCPMVpCandLX][0] and CPMVPListLX[numCPMVpCandLX][1]. As a candidate is added, numCPMVpCandLX can increase by 1.
[0934] In the embodiments described above, after it is determined whether CPMV1, CPMV2, and CPMV3 are to be added to the affine transformation AMVP control point prediction list, if numCPMVpCandLX is smaller than the maximum number of candidates in the affine transformation AMVP control point prediction list, a zero motion vector may be added to the affine transformation AMVP control point prediction list. At this time, the zero motion vector may be added repeatedly until the number of candidates in the affine transformation AMVP control point prediction list reaches the maximum number. For example, a value of 0 may be added to the x and y directions of CPMVpListLX until numCPMVpCandLX reaches the maximum number.
[0935] In constructing the affine transformation AMVP control point prediction list, the order of the steps (2510, 2520, and 2530) of FIG. 25 may be interchanged.
[0936] Derivation of Control Point Motion Vector (CPMV)
[0937] The following descriptions may be applied to the derivation of affine transformation control points in steps (2010) and (2120), etc.
[0938] In the step of deriving the CPMV, the number of parameters of the affine transformation model may be predetermined. Additionally, v0, v1, and v2 of FIGS. 18 and 19 may be derived based on 1) the height and width of the reference block, 2) the CPMV of the reference block, and 3) the height and width of the target block.
[0939] Through the following process, the CPMV of the target block can be derived using the CPMV of the reference block adjacent to the target block.
[0940] (xNb, yNb) may be the top-left coordinates of the reference block. nNbW and nNbH may be the width and height of the reference block, respectively.
[0941] (xCb, yCb) can be the top-left coordinates of the target block. cbWidth and cbHeight can be the width and height of the target block, respectively.
[0942] If both Condition 1 and Condition 2 below are true, isCTUboundary can be set to the second value. If one or more of Condition 1 and Condition 2 below are not true, isCTUboundary can be set to the first value.
[0943] isCTUboundary can indicate whether the target block is located at the boundary of a Coding Tree Unit (CTU). If isCTUboundary is the first value, the target block may not be located at the boundary of the CTU. If isCTUboundary is the second value, the target block may be located at the boundary of the CTU.
[0944] [Condition 1]
[0945] ((yNb + nNbH) % CtbSizeY) == 0
[0946] [Condition 2]
[0947] yNb + nNbH == yCb
[0948] When isCTUboundary is the first value, the CPMV of the target block can be derived using the following Equations 5 to 8.
[0949] [Formula 5]
[0950] mvScaleHor = MvLX[xNb][yNb + nNbH - 1][0] << CU_MAX_DEPTH
[0951] [Equation 6]
[0952] mvScaleHor = MvLX[xNb][yNb + nNbH - 1][1] << CU_MAX_DEPTH
[0953] [Equation 7]
[0954] dHorX = (MvLX[xNb + nNbW - 1][yNb + nNbH - 1][0] - MvLX[xNb][yNb + nNbH - 1][0] ) << (CU_MAX_DEPTH - log2(NbW))
[0955] [Equation 8]
[0956] dVerX = (MvLX[xNb + nNbW - 1][yNb + nNbH - 1][1] - MvLX[xNb][yNb + nNbH - 1][1] ) << (CU_MAX_DEPTH - log2(NbW))
[0957] "<<" can be a left shift operation.
[0958] CU_MAX_DEPTH can be the maximum depth of a Coding Unit (CU). Alternatively, CU_MAX_DEPTH can be Log2(maximum size of a CU). For example, if the maximum size of a CU is 128x128, CU_MAX_DEPTH can be 7.
[0959] "MvL X "of " X " can be a first value (e.g., "0") or a second value (e.g., "1"). X " can refer to "one of multiple reference pictures referenced by the block" or "one of multiple reference picture lists for the block". For example, MvL0 can represent a motion vector for a reference picture within reference picture list L0.
[0960] MvLX[xNb][yNb][0] may be the x-direction value of the motion vector of the reference block occupying the coordinates (xNb, yNb). MvLX[xNb][yNb][1] may be the y-direction value of the motion vector of the reference block occupying the coordinates (xNb, yNb).
[0961] If isCTUboundary is the second value, the CPMV of the target block can be derived using the following Equations 9 to 12.
[0962] [Formula 9]
[0963] mvScaleHor = CPMVLX[xNb][yNb][0][0] << CU_MAX_DEPTH
[0964] [Formula 10]
[0965] mvScaleVer = CPMVLX[xNb][yNb][0][1] << CU_MAX_DEPTH
[0966] [Equation 11]
[0967] dHorX = (CPMVLX[xNb + nNbW - 1][yNb][1][0] - CPMVLX[xNb][yNb][0][0]) << (CU_MAX_DEPTH - log2(NbW))
[0968] [Equation 12]
[0969] dVerX = (CPMVLX[xNb + nNbW - 1][yNb][1][1] - CPMVLX[xNb][yNb][0][1]) << (CU_MAX_DEPTH - log2(NbW))
[0970] "CPMVL X "of " X " can be a first value (e.g., "0") or a second value (e.g., "1"). X " can refer to "one of multiple reference pictures referenced by the block" or "one of multiple reference picture lists for the block". For example, MvL0 can represent a motion vector for a reference picture within reference picture list L0.
[0971] CPMVLX[xNb][yNb][0][0] may be the x-direction value of CPMV v0 at the top-left of the reference block occupying coordinates (xNb, yNb). CPMVLX[xNb][yNb][0][1] may be the y-direction value of CPMV v0 at the top-left of the reference block occupying coordinates (xNb, yNb).
[0972] CPMVLX[xNb + nNbW - 1][yNb][1][0] may be the x-direction value of CPMV v1 at the top right of the reference block occupying the coordinates (xNb + nNbW - 1, yNb). CPMVLX[xNb + nNbW - 1][yNb][1][1] may be the y-direction value of CPMV v1 at the top right of the reference block occupying the coordinates (xNb + nNbW - 1, yNb).
[0973] If isCTUboundary is the second value or MotionModelIdc is the second value, the following Equations 13 and 14 may be used.
[0974] [Equation 13]
[0975] dHorY = -dVerX
[0976] [Equation 14]
[0977] dVerY = dHorX
[0978] If isCTUboudary is the first value or MotionModelIdc is the third value, the CPMV of the target block can be derived using the following Equations 15 and 16.
[0979] [Formula 15]
[0980] dHorY = (CPMVLX[xNb][yNb + nNbH - 1][2][0] - CPMVLX[xNb][yNb][2][0] )<< (CU_MAX_DEPTH - log2NbH)
[0981] [Equation 16]
[0982] dVerY = (CPMVLX[xNb][yNb + nNbH - 1][2][1] - CPMVLX[xNb][yNb][2][1]) << (CU_MAX_DEPTH - log2NbH)
[0983] CPMVLX[xNb][yNb + nNbH - 1][2][0] may be the x-direction value of CPMV v2 at the bottom-left of the reference block occupying coordinates (xNb, yNb + nNbH - 1). CPMVLX[xNb][yNb + nNbH - 1][2][1] may be the y-direction value of CPMV v2 at the bottom-left of the reference block occupying coordinates (xNb, yNb + nNbH - 1).
[0984] If isCTUboundary is the first value, yNb can be the same as yCb.
[0985] CPMV v0 and CPMV v1 of the target block can be derived using the following formulas 17 to 20.
[0986] [Equation 17]
[0987] CPMVLX[0][0] = (mvScaleHor + dHorX * (xCb - xNb) + dHorY * (yCb - yNb))
[0988] [Equation 18]
[0989] CPMVLX[0][1] = (mvScaleVer + dVerX * (xCb - xNb) + dVerY * (yCb - yNb))
[0990] [Equation 19]
[0991] CPMVLX[1][0] = (mvScaleHor + dHorX * (xCb + cbWidth - xNb) + dHorY * (yCb - yNb))
[0992] [Formula 20]
[0993] CPMVLX[1][1] = (mvScaleVer + dVerX * (xCb + cbWidth - xNb) + dVerY * (yCb - yNb) )
[0994] "CPMVL X "of " X " can be a first value (e.g., "0") or a second value (e.g., "1"). X " can point to one of multiple reference pictures referenced by the block or one of multiple reference picture lists for the block. For example, MvL0 can represent a motion vector for a reference picture within reference picture list L0.
[0995] CPMVLX[0][0] may be the x-direction value of CPMV v0 of the target block. CPMVLX[0][1] may be the y-direction value of CPMV v0 of the target block.
[0996] CPMVLX[1][0] may be the x-direction value of CPMV v1 of the target block. CPMVLX[1][1] may be the y-direction value of CPMV v1 of the target block.
[0997] When MotionModelIdc is the third value, the CPMV v2 of the target block can be derived using the following Equations 21 and 22.
[0998] [Equation 21]
[0999] CPMVLX[2][0] = (mvScaleHor + dHorX * (xCb - xNb) + dHorY * (yCb + cbHeight - yNb))
[1000] [Equation 22]
[1001] CPMVLX[2][1] = (mvScaleVer + dVerX * (xCb - xNb) + dVerY * (yCb + cbHeight - yNb))
[1002] CPMVLX[2][0] may be the x-direction value of CPMV v2 of the target block. CPMVLX[2][1] may be the y-direction value of CPMV v2 of the target block.
[1003] Generation of motion vectors for target blocks for affine transform prediction blocks using CPMV
[1004] The following descriptions can be applied to the generation of motion vectors of target blocks for affine transformation prediction blocks using control point motion vectors in steps (2020) and (2130).
[1005] The motion vector of the target block for the affine transformation prediction block using the control point motion vector can be derived using the following Equations 23 to 26.
[1006] [Equation 23]
[1007] mvScaleHor = CPMVLX[0][0] << CU_MAX_DEPTH
[1008] [Equation 24]
[1009] mvScaleVer = CPMVLX[0][1] << CU_MAX_DEPTH
[1010] [Formula 25]
[1011] dHorX = (CPMVLX [1][0] - CPMVLX [0][0]) << (CU_MAX_DEPTH - log2(CbW))
[1012] [Equation 26]
[1013] dVerX = (CPMVLX [1][1] - CPMVLX [0][1]) << (CU_MAX_DEPTH - log2(CbW))
[1014] MotionModelIdc can represent an affine transformation model. For example, a MotionModelIdc value of 1 can represent the use of inter-prediction that does not use an affine transformation. A MotionModelIdc value of 2 can represent the use of a 4-parameter affine transformation model. A MotionModelIdc value of 3 can represent the use of a 6-parameter affine transformation model.
[1015] If MotionModelIdc represents the use of a 4-parameter affine transformation model, the following Equations 27 and 28 may be used.
[1016] [Equation 27]
[1017] dHorY = -dVerX
[1018] [Equation 28]
[1019] dVerY = dHorX
[1020] If MotionModelIdc represents the use of a 6-parameter affine transformation model, the following Equations 29 and 30 may be used.
[1021] [Equation 29]
[1022] dHorY = (CPMVLX[2][0] - CPMVLX[0][0]) << (CU_MAX_DEPTH - log2(CbH))
[1023] [Formula 30]
[1024] dVerY = (CPMVLX[2][1] - CPMVLX[0][1]) << (CU_MAX_DEPTH - log2(CbH))
[1025] "CPMVL X "of " X" can be a first value (e.g., "0") or a second value (e.g., "1"). X " can refer to "one of multiple reference pictures referenced by the block" or "one of multiple reference picture lists for the block". For example, MvL0 can represent a motion vector for a reference picture within reference picture list L0.
[1026] "CPMVLX[ a ][ b ]"at " a " can indicate the position of CPMV within the block (i.e., number or index), and " b " can point to one of the x and y directions of CPMV. b A value of 0 can point in the x direction, and " b A value of 1 can indicate the y direction. For example, CPMVLX[0][0] and CPMVLX[0][1] can represent the x direction and y direction of CPMV v0 shown in FIG. 19, respectively. CPMVLX[1][0] and CPMVLX[1][1] can represent the x direction and y direction of CPMV v1 shown in FIG. 19, respectively. CPMVLX[2][0] and CPMVLX[2][1] can represent the x direction and y direction of CPMV v2 shown in FIG. 20, respectively.
[1027] CbW and CbH can be the width and height of the target block, respectively.
[1028] In one embodiment, whether a fallback mode is applied to a target block can be determined based on the prediction direction type information inter_pred_idc for the target block. inter_pred_idc may indicate whether bidirectional prediction or unidirectional prediction is applied to the target block.
[1029] For example, if inter_pred_idc is 0 and (cbWidth + cbHeight) >= 12, unidirectional prediction using reference picture list L0 may be used for the target block. If inter_pred_idc is 1 and (cbWidth + cbHeight) >= 12, unidirectional prediction using reference picture list L1 may be used for the target block. For example, if inter_pred_idc is 2 and (cbWidth + cbHeight) > 12, bidirectional prediction using reference picture list L0 and reference picture list L1 may be used for the target block.
[1030] The fallback mode may indicate whether to apply Equations 47, 48, 49, and 50, which will be described later.
[1031] When the target block references two reference pictures, bidirectional prediction is used for the target block, so the value of fallbackModeTriggered can be determined using Equations 31 to 37 below.
[1032] fallbackModeTriggered may be fallback mode usage information indicating whether fallback mode is used. For example, if the value of fallbackModeTriggered is a first value (e.g., "0"), fallback mode may not be triggered. If the value of fallbackModeTriggered is a second value (e.g., "1"), fallback mode may be triggered.
[1033] RefPicList[0] can represent reference picture list L0. RefPicList[1] can represent reference picture list L1.
[1034] refIdxL0 may be an index for reference picture list L0 and may point to a reference picture within reference picture list L0. refIdxL1 may be an index for reference picture list L1 and may point to a reference picture within reference picture list L1.
[1035] RefPicList[0][refIdxL0] may represent a reference picture specified by refIdxL0 among the reference pictures in reference picture list L0. RefPicList[1][refIdxL1] may represent a reference picture specified by refIdxL1 among the reference pictures in reference picture list L1.
[1036] [Equation 31]
[1037] maxW4= Max(0, Max(4 * (2048 + dHorX), Max(4 * dHorY, 4 * (2048 + dHorX) + 4 * dHorY)))
[1038] [Equation 32]
[1039] minW4= Min(0, Min(4 * (2048 + dHorX), Min(4 * dHorY, 4 * (2048 + dHorX) + 4 * dHorY)))
[1040] [Equation 33]
[1041] maxH4= Max(0, Max(4 * dVerX, Max(4* (2048 + dVerY), 4 * dVerX + 4 * (2048 + dVerY))))
[1042] [Equation 34]
[1043] minH4= Min(0, Min(4 * dVerX, Min(4* (2048 + dVerY), 4 * dVerX + 4 * (2048 + dVerY))))
[1044] [Formula 35]
[1045] bxWX4= ((maxW4- minW4) >> 11) + 9
[1046] [Equation 36]
[1047] bxHX4= ((maxH4- minH4) >> 11) + 9
[1048] [Equation 37]
[1049] bxWX4* bxHX4< THR BI
[1050] In one embodiment, THR BI can be a specified integer value. For example, THR BI can be 226. If the result of Equation 37 is true, fallbackModeTriggered can be set to a first value (e.g., "0"). If the result of Equation 37 is not true, fallbackModeTriggered can be set to a second value (e.g., "1").
[1051] In one embodiment, THR BI can be a specified integer value. For example, THR BI can be 226. If the result of Equation 37 is true or if at least one of RefPicList[0][refIdxL0] and RefPicList[1][refIdxL1] is a long-term reference picture, fallbackModeTriggered may be set to a first value. If the result of Equation 37 is not true and RefPicList[0][refIdxL0] and RefPicList[1][refIdxL1] are not both long-term reference pictures, fallbackModeTriggered may be set to a second value.
[1052] When the target block references one reference picture, unidirectional prediction is used for the target block, so the value of fallbackModeTriggered can be determined using Equations 38 to 43 below.
[1053] [Equation 38]
[1054] bxWX h = ((Max(0, 4 * (2048 + dHorX)) - Min(0, 4 * (2048 + dHorX))) >> 11) + 9
[1055] [Equation 39]
[1056] bxHX h = ((Max(0, 4 * dVerX) - Min(0, 4 * dVerX)) >> 11) + 9
[1057] [Formula 40]
[1058] bxWX v = ((Max(0, 4 * dVerY) - Min(0, 4 * dVerY)) >> 11) + 9
[1059] [Equation 41]
[1060] bxHX v = ((Max(0, 4 * (2048 + dHorY)) - Min(0, 4 * (2048 + dHorY))) >> 11) + 9
[1061] [Equation 42]
[1062] bxWX v * bxHX v < THR v
[1063] [Equation 43]
[1064] bxWX v * bxHX h < THR h
[1065] In one embodiment, THR v and THR h can each be a specified integer value. For example, THR v and THR hcan be 166. If at least one of the result of Equation 42 and the result of Equation 43 is true, fallbackModeTriggered can be set to a first value (e.g., "0"). If neither the result of Equation 42 nor the result of Equation 43 is true, fallbackModeTriggered can be set to a second value (e.g., "1").
[1066] In one embodiment, THR v and THR h can each be a specified integer value. For example, THR v and THR h can be 166. If at least one of the result of Equation 42 and the result of Equation 43 is true and RefPicList[X][refIdxLX] is a long-term reference picture, fallbackModeTriggered can be set to the first value. If neither the result of Equation 42 nor the result of Equation 43 is true and RefPicList[X][refIdxLX] is not a long-term reference picture, fallbackModeTriggered can be set to the second value.
[1067] In one embodiment, if the above conditions do not apply, fallbackModeTriggered may be set to a second value. Additionally, fallbackModeTriggered may be set to a second value before the above conditions are applied, and the value of fallbackModeTriggered may be changed according to the above conditions.
[1068] In the description above, CU_MAX_DEPTH may be the maximum depth of a Coding Unit (CU). Alternatively, CU_MAX_DEPTH may be Log2(maximum size of the CU). For example, if the maximum size of the CU is 128x128, CU_MAX_DEPTH may be 7.
[1069] Max(a, b, c) can represent the maximum value among a, b, and c. Min(a, b, c) can represent the minimum value among a, b, and c.
[1070] As described in Equations 44 and 45 below, xSbIdx can have one value from 0 to numSbX-1. ySbIdx can have one value from 0 to numSbY-1.
[1071] [Equation 44]
[1072] numSbX = (cbWidth << 2)
[1073] [Formula 45]
[1074] numSbY = (cbHeight << 2)
[1075] If fallbackModeTriggered is the first value, the following Equations 46 and 47 may be executed (as the fallback mode is not triggered).
[1076] [Equation 46]
[1077] xPosCb = (cbWidth >> 1)
[1078] [Equation 47]
[1079] yPostCb = (cbHeight >> 1)
[1080] If fallbackModeTriggered is the second value, the following Equations 48 and 49 may be executed (as the fallback mode is triggered).
[1081] [Equation 48]
[1082] xPosCb = 2 + (xSbIdx << 2)
[1083] [Equation 49]
[1084] yPosCb = 2 + (ySbIdx << 2)
[1085] mvLX[xSbIdx][ySbIdx] can represent the motion vector of the target block. mvLX[xSbIdx][ySbIdx] can be derived using the following Equations 50 and 51.
[1086] [Formula 50]
[1087] mvLX[xSbIdx][ySbIdx][0] = (mvScaleHor + dHorX * xPosCb + dHorY * yPosCb)
[1088] [Equation 51]
[1089] mvLX[xSbIdx][ySbIdx][1] = (mvScaleVer + dVerX * xPosCb + dVerY * yPosCb)
[1090] Encoding / decoding of the derivation of the affine transform model and / or the motion vector difference of the control point
[1091] The following descriptions may be applied to the derivation of the affine transform model in step (2030) and step (2110) and / or the encoding / decoding of the motion vector difference of the control point.
[1092] MotionModelIdc can represent an affine transformation model. For example, a MotionModelIdc value of 1 can represent the use of inter-prediction that does not use an affine transformation. A MotionModelIdc value of 2 can represent the use of a 4-parameter affine transformation model. A MotionModelIdc value of 3 can represent the use of a 6-parameter affine transformation model.
[1093] MotionModelIdc can be derived using Equation 52 below.
[1094] [Equation 52]
[1095] MotionModelIdc = inter_affine_flag + affine_type_flag
[1096] For example, if encoding / decoding of a target block is performed through inter-prediction without using affine transformation, MotionModelIdc may be set to a first value (e.g., "0"). In this case, since affine transformation is not used, inter_affine_flag and affine_type_flag may be set to a first value (e.g., "0").
[1097] For example, when encoding / decoding of a target block is performed through a 4-parameter affine transform model, MotionModelIdc may be set to a second value (e.g., "1"). In this case, since an affine transform is used, inter_affine_flag may be set to a second value (e.g., "1"), and affine_type_flag may be set to a first value (e.g., "0").
[1098] For example, if encoding / decoding of the target block is performed through a 6-parameter affine transform model, MotionModelIdc may be set to a third value (e.g., "2"). In this case, since an affine transform is used, inter_affine_flag and affine_type_flag may be set to a second value (e.g., "1").
[1099] In one embodiment, when a 4-parameter affine transform model and a 6-parameter affine transform model are used, encoding / decoding of the motion vector difference of the control point can be performed as follows.
[1100] When a 4-parameter affine transform model is used, the motion vector difference for CPMV v0 and CPMV v1 shown in Fig. 18 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, and the motion vector difference for the y-direction value of CPMV v1 can be encoded / decoded.
[1101] When a 6-parameter affine transform model is used, the motion vector difference for CPMV v0, CPMV v1, and CPMV v2 illustrated in FIG. 19 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, the motion vector difference for the y-direction value of CPMV v1, the motion vector difference for the x-direction value of CPMV v2, and the motion vector difference for the y-direction value of CPMV v2 can be encoded / decoded.
[1102] In one embodiment, when a 4-parameter affine transform model and a 6-parameter affine transform model are used, encoding / decoding of the motion vector difference of the control point can be performed as follows based on whether the reference picture of the target block is a long-term reference picture.
[1103] As described above, RefPicList[0][refIdxL0] and RefPicList[1][refIdxL1] can represent reference pictures for the target block, respectively.
[1104] Through RefPicList[0][refIdxL0] and RefPicList[1][refIdxL1], it can be identified whether the reference picture referenced by the target block is a long-term reference picture.
[1105] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, encoding / decoding of the motion vector difference of the control point can be performed as follows.
[1106] When a 4-parameter affine transform model is used, the motion vector difference for CPMV v0 and CPMV v1 shown in Fig. 18 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, and the motion vector difference for the y-direction value of CPMV v1 can be encoded / decoded.
[1107] When a 6-parameter affine transform model is used, the motion vector difference for CPMV v0, CPMV v1, and CPMV v2 illustrated in FIG. 19 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, the motion vector difference for the y-direction value of CPMV v1, the motion vector difference for the x-direction value of CPMV v2, and the motion vector difference for the y-direction value of CPMV v2 can be encoded / decoded.
[1108] In other cases (that is, when RefPicList[0][refIdxL0] is a long-term reference picture or RefPicList[1][refIdxL1] is a long-term reference picture), encoding / decoding of the motion vector difference of the control point may not be performed.
[1109] In one embodiment, when a 4-parameter affine transform model and a 6-parameter affine transform model are used, encoding / decoding of the motion vector difference of the control point can be performed as follows based on whether the reference picture of the target block is a long-term reference picture.
[1110] As described above, RefPicList[0][refIdxL0] and RefPicList[1][refIdxL1] can represent reference pictures for the target block, respectively.
[1111] Through RefPicList[0][refIdxL0] and RefPicList[1][refIdxL1], it can be identified whether the reference picture referenced by the target block is a long-term reference picture.
[1112] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, encoding / decoding of the motion vector difference of the control point can be performed as follows.
[1113] When a 4-parameter affine transform model is used, the motion vector difference for CPMV v0 and CPMV v1 shown in Fig. 18 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, and the motion vector difference for the y-direction value of CPMV v1 can be encoded / decoded.
[1114] When a 6-parameter affine transform model is used, the motion vector difference for CPMV v0, CPMV v1, and CPMV v2 illustrated in FIG. 19 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, the motion vector difference for the y-direction value of CPMV v1, the motion vector difference for the x-direction value of CPMV v2, and the motion vector difference for the y-direction value of CPMV v2 can be encoded / decoded.
[1115] If RefPicList[0][refIdxL0] is a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, encoding / decoding of the motion vector difference of the control point can be performed as follows.
[1116] When CPMV v0 and CPMV v1 refer to RefPicList[1][refIdxL1] and a 4-parameter affine transform model is used, the motion vector difference for CPMV v0 and CPMV v1 shown in FIG. 18 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, and the motion vector difference for the y-direction value of CPMV v1 can be encoded / decoded.
[1117] When CPMV v0, CPMV v1, and CPMV v2 refer to RefPicList[1][refIdxL1] and a 6-parameter affine transform model is used, the motion vector difference for CPMV v0, CPMV v1, and CPMV v2 shown in FIG. 19 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, the motion vector difference for the y-direction value of CPMV v1, the motion vector difference for the x-direction value of CPMV v2, and the motion vector difference for the y-direction value of CPMV v2 can be encoded / decoded.
[1118] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is a long-term reference picture, encoding / decoding of the motion vector difference of the control point can be performed as follows.
[1119] When CPMV v0 and CPMV v1 refer to RefPicList[0][refIdxL0] and a 4-parameter affine transform model is used, the motion vector difference for CPMV v0 and CPMV v1 shown in FIG. 18 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, and the motion vector difference for the y-direction value of CPMV v1 can be encoded / decoded.
[1120] When CPMV v0, CPMV v1, and CPMV v2 refer to RefPicList[0][refIdxL0] and a 6-parameter affine transform model is used, the motion vector difference for CPMV v0, CPMV v1, and CPMV v2 shown in FIG. 19 can be encoded / decoded. The motion vector difference for the x-direction value of CPMV v0, the motion vector difference for the y-direction value of CPMV v0, the motion vector difference for the x-direction value of CPMV v1, the motion vector difference for the y-direction value of CPMV v1, the motion vector difference for the x-direction value of CPMV v2, and the motion vector difference for the y-direction value of CPMV v2 can be ...
Claims
Claim 1 A video decoding method comprising: a step of determining a prediction mode for a target block; and a step of performing a prediction for the target block using the prediction mode, wherein if the prediction mode is an inter mode using an affine transformation, the derivation of an affine transformation model and the derivation of a motion vector difference value of a control point are performed, and the configuration of an affine transformation list and the derivation of a control point motion vector of the control point are performed, and if the picture referenced by the reference block of the target block is a short-term reference picture, the affine transformation control point of the reference block is used, and the affine transformation list is configured based on the affine transformation control point of the reference block, and if the picture referenced by the reference block is a long-term reference picture, the affine transformation control point of the reference block is not used. Claim 2 In claim 1, the image decoding method wherein the affine transformation is determined based on the distance between a target picture including the target block and a reference picture. Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 A video decoding method according to claim 1, wherein the affine transformation control point of the reference block is derived based on the size of the target block, and the affine transformation list is configured based on the derived affine transformation control point. Claim 7 A video decoding method according to claim 1, wherein the affine transform list is an affine transform advanced motion vector prediction (AMVP) list. Claim 8 In claim 1, the image decoding method wherein the affine transform list is a list of affine transform merge modes. Claim 9 A video encoding method comprising: a step of determining a prediction mode for a target block; and a step of performing a prediction for the target block using the prediction mode, wherein if the prediction mode is an inter mode using an affine transformation, the derivation of an affine transformation model and the derivation of a motion vector difference value of a control point are performed, and the configuration of an affine transformation list and the derivation of a control point motion vector of the control point are performed, and if the picture referenced by the reference block of the target block is a short-term reference picture, the affine transformation control point of the reference block is used, and the affine transformation list is configured based on the affine transformation control point of the reference block, and if the picture referenced by the reference block is a long-term reference picture, the affine transformation control point of the reference block is not used. Claim 10 A computer-readable recording medium for storing a bitstream for image decoding, wherein the bitstream includes information related to a prediction method, a prediction mode for a target block is determined using the information related to the prediction method, a prediction for the target block using the prediction mode is performed, and if the prediction mode is an inter mode using an affine transformation, the derivation of an affine transformation model and the derivation of a difference value of a control point's motion vector are performed, the configuration of an affine transformation list and the derivation of a control point motion vector of the control point are performed, and a motion vector of the target block is generated using the control point motion vector, and if the picture referenced by the reference block of the target block is a short-term reference picture, the affine transformation control point of the reference block is used, and the affine transformation list is configured based on the affine transformation control point of the reference block, and if the picture referenced by the reference block is a long-term reference picture, the affine transformation control point of the reference block is not used. Claim 11 A video encoding method according to claim 9, wherein the affine transformation control points of the reference block are derived based on the size of the target block, and the affine transformation list is configured based on the derived affine transformation control points. Claim 12 In claim 9, the above-mentioned affine transform list is an affine transform advanced motion vector prediction (AMVP) list, a video encoding method. Claim 13 In claim 9, the above affine transform list is a video encoding method in which the affine transform merge mode is a list of affine transform merge modes. Claim 14 A computer-readable recording medium according to claim 10, wherein the affine transformation control points of the reference block are derived based on the size of the target block, and the affine transformation list is configured based on the derived affine transformation control points. Claim 15 In paragraph 10, the above-mentioned affine transformation list is a computer-readable recording medium that is an affine transformation advanced motion vector prediction (AMVP) list. Claim 16 In paragraph 10, the above affine conversion list is a computer-readable recording medium that is a list of affine conversion merge modes.
Citation Information
Patent Citations
Inter-prediction mode based image processing method, and apparatus therefor
KR1020180128955A