Image encoding / decoding method and apparatus, and recording medium storing bit stream
By employing block vector prediction mode and intra-frame block copying technology, the problem of low coding efficiency for high-resolution and high-definition images is solved, achieving more efficient image coding and decoding to meet users' needs for high-resolution and high-definition images.
Patent Information
- Application Number
- CN202480025594.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-09
- Filing Date
- 2024-04-15
- Publication Date
- 2025-11-14
AI Technical Summary
Existing image coding techniques struggle to effectively utilize intra-frame prediction modes of block vectors for efficient encoding and decoding, especially in high-resolution and high-definition image processing, resulting in low coding efficiency.
A block vector-based prediction mode is adopted. By determining the block vector resolution and deriving the block vector, a prediction block of the target block is generated. Intra-frame block copy mode is used for image encoding and decoding, including dividing the target block into sub-blocks, deriving the motion information of the reference block, and generating the block vector corresponding to the sub-block based on this.
It improves the encoding efficiency of high-resolution and high-definition images, enhances the accuracy and efficiency of encoding and decoding, and meets users' needs for high-resolution and high-definition images.
Smart Images

Figure CN120958823A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods and apparatus for encoding / decoding images and recording media for storing bit streams, and more specifically, to methods and apparatus for encoding and decoding images based on intra-frame prediction such as intra-block copying (IBC), intra-template matching, etc. Background Technology
[0002] With the continuous development of the information and communication industry, broadcast services supporting high-definition (HD) resolution have become widespread throughout the world. Through this widespread adoption, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.
[0003] To meet user demand for high definition, many organizations have accelerated the development of next-generation imaging devices. User interest in UHD TVs, with resolutions more than four times that of Full HD (FHD) TVs, as well as High Definition TVs (HDTVs) and FHD TVs, has increased. With this growing interest, there is a current need for image encoding / decoding technologies for images with higher resolution and higher definition.
[0004] As an image compression technology, there are various techniques, such as inter-frame prediction, intra-frame prediction, transform, quantization, filtering, and entropy coding.
[0005] Inter-frame prediction is a technique that uses frames preceding and / or following the current frame to predict the values of pixels included in the current frame. Intra-frame prediction is a technique that uses information about pixels in the current frame to predict the values of pixels included in the current frame. Transform and quantization techniques can be used to compress the energy of residual signals. Entropy coding is a technique used to assign short codewords to frequently occurring values and long codewords to less frequent values.
[0006] This image compression technology allows for the efficient compression, transmission, and storage of image data. Summary of the Invention
[0007] Technical issues The purpose of this disclosure is to provide various methods for encoding or decoding images using at least one intra-frame prediction mode based on block vectors.
[0008] Technical solution According to one aspect of this disclosure, an image decoding method is provided for predicting target blocks included in a target image. The method includes: determining a prediction mode for the target block among a plurality of block vector-based prediction modes; determining a block vector resolution for the target block, wherein the block vector resolution is determined among a plurality of block vector resolution candidates including one or more integer unit resolutions and one or more fractional unit resolutions; deriving at least one block vector for the target block based on the determined prediction mode and the block vector resolution; and generating a predicted block of the target block from pre-reconstructed samples in the target image using the at least one block vector.
[0009] When the determined mode is a sub-block-based intra-block copy (IBC) mode, the step of deriving the at least one block vector may include: partitioning the target block into multiple sub-blocks; deriving motion information for specifying a reference block, the reference block being referenced for deriving the at least one block vector of the target block; determining the reference block in the target frame based on the motion information; and generating a block vector corresponding to a sub-block in the target block using the block vectors of the sub-blocks in the reference block, wherein the predicted block of the target block is generated on a sub-block basis using the block vectors of each sub-block in the target block.
[0010] According to another aspect of this disclosure, an image coding method is provided for predicting target blocks included in a target image. The method includes: determining a prediction mode for the target block among a plurality of block vector-based prediction modes; determining a block vector resolution for the target block, wherein the block vector resolution is determined among a plurality of block vector resolution candidates including one or more integer unit resolutions and one or more fractional unit resolutions; deriving at least one block vector for the target block based on the determined prediction mode and the block vector resolution; and generating a predicted block of the target block from pre-reconstructed samples in the target image using the at least one block vector.
[0011] When the determined mode is a sub-block-based intra-block copy (IBC) mode, the step of deriving the at least one block vector includes: partitioning the target block into multiple sub-blocks; deriving motion information for specifying a reference block, the reference block being referenced for deriving the at least one block vector of the target block; determining the reference block in the target frame based on the motion information; and generating a block vector corresponding to a sub-block in the target block using the block vectors of the sub-blocks in the reference block, wherein the predicted block of the target block is generated on a sub-block basis using the block vectors of each sub-block in the target block.
[0012] According to another aspect of this disclosure, a method is provided for generating a bitstream by predicting and encoding target blocks included in a target image and sending the bitstream to an image decoding device. The step of generating the bitstream includes: determining a prediction mode for the target block among a plurality of block vector-based prediction modes; determining a block vector resolution for the target block, wherein the block vector resolution is determined among a plurality of block vector resolution candidates including one or more integer unit resolutions and one or more fractional unit resolutions; deriving at least one block vector for the target block based on the determined prediction mode and the block vector resolution; and generating a predicted block of the target block from pre-reconstructed samples in the target image using the at least one block vector. Attached Figure Description
[0013] Figure 1 This is a block diagram illustrating the configuration of an embodiment of the encoding device applying the present disclosure; Figure 2 This is a block diagram illustrating the configuration of an embodiment of the decoding device applying the present disclosure; Figure 3 It is a schematic diagram illustrating the partitioning structure of an image as it is encoded and decoded; Figure 4 This is a diagram illustrating the form of prediction units (PUs) that a coding unit (CU) may include; Figure 5 This is a diagram illustrating the form of a transformation unit (TU) that may be included in a CU; Figure 6 This shows the block division based on the example; Figure 7 This is a diagram illustrating an embodiment used to explain the intra-frame prediction process; Figure 8 This is a diagram showing the reference samples used in the intra-frame prediction process; Figure 9 This is a diagram illustrating an embodiment used to explain the inter-frame prediction process; Figure 10 Spatial candidates are shown according to an embodiment; Figure 11 The order in which motion information of spatial candidates is added to the merging list is shown according to an embodiment; Figure 12 The transformation and quantization processes are shown based on the example; Figure 13 The diagonal scan is shown according to the example; Figure 14 Showing a horizontal scan based on an example; Figure 15 The vertical scan is shown according to the example; Figure 16This is a configuration diagram of the encoding device according to an embodiment; Figure 17 This is a configuration diagram of the decoding device according to an embodiment; Figure 18 These are exemplary diagrams used to explain template matching according to embodiments of this disclosure; Figure 19 and Figure 20 Various examples of subsampling methods in template matching are shown; Figures 21 to 26 Each of these illustrates the search method in template matching based on the example; Figure 27 This is an exemplary diagram used to explain a method for configuring a target template in affine mode according to embodiments of the present disclosure; Figure 28a and Figure 28b This is an exemplary diagram used to explain a method for configuring a reference template in affine mode according to embodiments of the present disclosure; Figure 29 This is a diagram used to explain bilateral matching according to embodiments of the present disclosure; Figure 30 This is an exemplary diagram used to explain a method for determining the position of a reference block in IBC mode according to embodiments of the present disclosure; Figure 31 This is an exemplary diagram illustrating the grouping of sub-blocks partitioned from a target block according to embodiments of the present disclosure; Figure 32 This is a flowchart for explaining the image encoding method according to embodiments of the present disclosure; Figure 33 An image decoding method according to an embodiment of the present disclosure is shown; Figure 34 This is an exemplary diagram illustrating the positions of adjacent blocks of a target block considered for deriving a spatial block vector candidate according to an embodiment of the present disclosure; Figure 35 This is an exemplary diagram illustrating the location of a non-adjacent block of a target block considered for deriving a spatial block vector candidate according to an embodiment of the present disclosure; Figure 36 This is a diagram used to explain the positions indicated by time block vector candidates according to embodiments of the present disclosure; Figure 37 a and Figure 37 b is an exemplary diagram used to explain the processing of deriving automatic relocation block vector candidates according to embodiments of the present disclosure; Figure 38a and Figure 38bThis is an exemplary diagram used to explain a method for specifying the sub-pixel unit position of a block vector of a target block in an intra-frame template matching mode according to embodiments of the present disclosure; Figure 39 This is an exemplary diagram used to explain the positional relationship between the target block and the reference block indicated by the block vector; Figure 40 This is an exemplary diagram showing a reference area that may include a reference block indicated by a block vector when the size of the CTB is 128×128; Figures 41 to 44 This is an exemplary diagram used to explain a method for managing a buffer for a reference region in block vector-based image encoding or decoding according to embodiments of the present disclosure; Figure 45a and Figure 45b This is an exemplary diagram illustrating the syntax structure of encoding parameters for entropy encoding / decoding according to embodiments of the present disclosure; and Figures 46 to 48 This is an exemplary diagram illustrating the syntax structure of a transformation unit used as a unit for transforming a target block according to an embodiment of the present disclosure. Detailed Implementation
[0014] This disclosure is subject to various changes and may have various embodiments, and specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that those embodiments are not intended to limit this disclosure to a particular form, and they include all changes, equivalents, or modifications included within the spirit and scope of this disclosure.
[0015] The following detailed description of exemplary embodiments will be made with reference to the accompanying drawings, which illustrate specific embodiments. These embodiments are described to enable those skilled in the art to readily practice them. It should be noted that the various embodiments differ from one another but are not necessarily mutually exclusive. For example, the particular shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the embodiments to which they relate. Furthermore, it should be understood that the position or arrangement of the various components in each disclosed embodiment may be changed without departing from the spirit and scope of the embodiments. Therefore, the appended detailed description is not intended to limit the scope of this disclosure, and the scope of the exemplary embodiments is limited only by the appended claims and their equivalents, provided that they are properly described.
[0016] In the accompanying drawings, similar reference numerals are used to specify the same or similar functions in various aspects. The shape, size, etc. of the components in the drawings may be exaggerated for clarity of description.
[0017] Terms such as “first” and “second” may be used to describe various components, but components are not limited by the terms. These terms are used only to distinguish one component from another. For example, without departing from the scope of this specification, a first component may be referred to as a second component. Similarly, a second component may be referred to as a first component. The term “and / or” may include a combination of multiple related descriptive items or any one of multiple related descriptive items.
[0018] It will be understood that when a component is referred to as "connected" or "combined" to another component, the two components can be directly connected or combined with each other, or there can be an intermediate component between the two components. On the other hand, it will be understood that when a component is referred to as "directly connected or combined," there is no intermediate component between the two components.
[0019] Furthermore, to indicate different functional characteristics, the components described in the embodiments are shown independently, but this does not mean that each of the components is formed by separate hardware or software. That is, for ease of description, the components are arranged and included separately. For example, at least two of the components may be integrated into a single component. Conversely, a component may be divided into multiple components. Embodiments of integrated components or embodiments of separated components are included within the scope of this specification, provided that they do not depart from the spirit of this specification.
[0020] The terminology used in the embodiments is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless specifically indicated in the context. In the embodiments, it should be understood that terms such as “comprising” or “having” are intended only to indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. That is, in the embodiments, the expression describing a component as “comprising” a specific component means that additional components may be included within the practice or technical spirit of this disclosure, but does not exclude the presence of components other than the specific component.
[0021] In embodiments, the term "at least one" may refer to one of 1 or more numbers, such as 1, 2, 3, and 4. In embodiments, the term "a plurality of" may refer to one of 2 or more numbers, such as 2, 3, and 4.
[0022] Some components of the embodiments are not essential components for performing necessary functions, but may be optional components used only to improve performance. Embodiments may be implemented using only the essential components necessary to achieve the essence of the embodiments. For example, a structure that includes only essential components and excludes optional components used only to improve performance is also included within the scope of the embodiments.
[0023] The embodiments will now be described in detail with reference to the accompanying drawings, enabling those skilled in the art to readily practice the embodiments. In the following description of the embodiments, detailed descriptions of known functions or configurations that are considered to obscure the spirit of this specification will be omitted. Furthermore, throughout the drawings, the same reference numerals are used to designate the same components, and repeated descriptions of the same components will be omitted.
[0024] In the following text, “image” may refer to a single frame that constitutes a video, or it may refer to the video itself. For example, “encoding and / or decoding of images” may refer to “encoding and / or decoding of a video”, and may also refer to “encoding and / or decoding of any one of the images that constitute a video”.
[0025] In the following text, the terms “video” and “moving footage” may be used to have the same meaning and may be used interchangeably.
[0026] In the following text, the target image can be an encoded target image that is the target to be encoded and / or a decoded target image that is the target to be decoded. Furthermore, the target image can be an input image input to an encoding device or an input image input to a decoding device. Also, the target image can be the current image, i.e., the target currently to be encoded and / or decoded. For example, the terms "target image" and "current image" can be used to have the same meaning and can be used interchangeably.
[0027] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used to have the same meaning and may be used interchangeably with each other.
[0028] In the following text, a target block can be an encoding target block (the target to be encoded) and / or a decoding target block (the target to be decoded). Furthermore, a target block can be the current block, i.e., the target currently to be encoded and / or decoded. Here, the terms "target block" and "current block" can be used to have the same meaning and can be used interchangeably. A current block can represent an encoding target block that is the encoding target during encoding and / or a decoding target block that is the decoding target during decoding. Furthermore, a current block can be at least one of an encoding block, a prediction block, a residual block, and a transform block.
[0029] In the following text, the terms “block” and “unit” may be used to have the same meaning and may be used interchangeably. Optionally, “block” may refer to a specific unit.
[0030] In the following text, the terms “region” and “segment” are used interchangeably.
[0031] In the following embodiments, specific information, data, flags, indexes, elements, and attributes may have their own values. The value "0" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate false, logical false, or a first predefined value. In other words, the values "0," false, logical false, and the first predefined value are interchangeable. The value "1" corresponding to each of the information, data, flags, indexes, elements, and attributes may indicate true, logical true, or a second predefined value. In other words, the values "1," true, logical true, and the second predefined value are interchangeable.
[0032] When variables such as i or j are used to indicate rows, columns, or indices, the value of i can be an integer of 0 or greater, or an integer of 1 or greater. In other words, in the embodiments, each of the rows, columns, and indices can be counted from 0 or from 1.
[0033] In the embodiments, the term "one or more" or the term "at least one" may mean the term "multiple". The terms "one or more" or the term "at least one" may be used interchangeably with "multiple".
[0034] The terminology used in the embodiments will be described below.
[0035] Encoder: An encoder is a device used to perform encoding. In other words, an encoder can refer to an encoding device.
[0036] Decoder: A decoder is a device used to perform decoding. In other words, a decoder can refer to a decoding device.
[0037] Unit: A unit can represent a component of image encoding and decoding. The terms "unit" and "block" can be used interchangeably and have the same meaning.
[0038] - A cell can be an M×N array of samples. Each of M and N can be a positive integer. A cell typically refers to a two-dimensional array of samples.
[0039] In image encoding and decoding, a "unit" can be a region generated by partitioning an image. In other words, a "unit" can be a specified region within an image. A single image can be partitioned into multiple units. Optionally, an image can be partitioned into sub-parts, and when encoding or decoding is performed on the partitioned sub-parts, a unit can represent each partitioned sub-part.
[0040] - In image encoding and decoding, predefined processing can be performed on each unit according to its type.
[0041] Based on function, unit types can be classified as macrounits, coding units (CUs), prediction units (PUs), residual units, transform units (TUs), etc. Optionally, based on function, units can represent blocks, macroblocks, coding tree units, coding tree blocks, coding units, coding blocks, prediction units, prediction blocks, residual units, residual blocks, transform units, transform blocks, etc. For example, the target unit that serves as the object of encoding and / or decoding can be at least one of CUs, PUs, residual units, and TUs.
[0042] - The term "unit" may mean a block of luma components, a block of chroma components corresponding to the luma components, and information about the syntax elements for each block, such that the unit is specified to be distinct from the block.
[0043] The size and shape of the unit can be implemented in a variety of ways. Furthermore, the unit can have any of a variety of sizes and shapes. In particular, the shape of the unit can include not only squares, but also geometric shapes that can be represented in two dimensions (2D), such as rectangles, trapezoids, triangles, and pentagons.
[0044] - In addition, cell information may include one or more of the following: cell type, cell size, cell depth, cell encoding order, and cell decoding order. For example, the cell type may indicate one of CU, PU, residual cell, and TU.
[0045] - A cell can be divided into sub-cells, and the size of each sub-cell is smaller than the size of the related cell.
[0046] Depth: Depth can refer to the degree to which a cell is partitioned. In addition, the depth of a cell can indicate the level at which the corresponding cell exists when (one or more) cells are represented by a tree structure.
[0047] - Cell partitioning information may include depth, which indicates the depth of the cell. Depth may indicate the number of times the cell is partitioned and / or the degree to which the cell is partitioned.
[0048] In a tree structure, the root node can be considered to have the smallest depth, and the leaf nodes have the largest depth. The root node can be the highest (top) node. The leaf nodes can be the lowest nodes.
[0049] A single cell can be hierarchically partitioned into multiple sub-cells, with depth information based on a tree structure. In other words, a cell and the sub-cells generated by partitioning the cell can correspond to a node and a node's child nodes, respectively. Each of the partitioned sub-cells can have a cell depth. Since the depth indicates the number of times the cell is partitioned and / or the degree to which the cell is partitioned, the partitioning information of a sub-cell can include information about the size of the sub-cell.
[0050] In a tree structure, the top node can correspond to the initial node before the partition. The top node can be called the "root node". Furthermore, the root node can have a minimum depth value. Here, the top node can have a depth of level "0".
[0051] - A node with a depth of level "1" can represent a cell generated when the initial cell is partitioned once. A node with a depth of level "2" can represent a cell generated when the initial cell is partitioned twice.
[0052] - Leaf nodes with a depth of level "n" can represent cells generated when the initial cell has been partitioned n times.
[0053] - Leaf nodes can be bottom nodes and cannot be further partitioned. The depth of a leaf node can be the maximum level. For example, a predefined value for the maximum level could be 3.
[0054] -QT depth can represent the depth used for quad-partitioning. BT depth can represent the depth used for binary-partitioning. TT depth can represent the depth used for ternary-partitioning.
[0055] Samples: Samples can be the basic units that make up a block. Depending on the bit depth (Bd), they can range from 0 to 2. Bd The value of -1 represents a sample point.
[0056] A sample point can be a pixel or a pixel value.
[0057] - In the following text, the terms “pixel” and “sample” may be used to have the same meaning and may be used interchangeably.
[0058] Code Tree Unit (CTU): A CTU can consist of a single luma component (Y) code tree block and two chroma component (Cb, Cr) code tree blocks associated with the luma component code tree block. Furthermore, a CTU can refer to information including the aforementioned blocks and the syntax elements used for each of the blocks.
[0059] Each coding tree unit (CTU) can be partitioned using one or more partitioning methods (such as quadtree (QT), binary tree (BT), and ternary tree (TT)) to configure subunits, such as coding units, prediction units, and transform units. A quadtree can refer to a quaternion tree. Additionally, one or more partitioning methods can be used to partition each coding tree unit using a multi-type tree (MTT).
[0060] - "CTU" can be used as a term to specify a pixel block, which is a processing unit in image decoding and encoding, such as in the case of partitioning an input image.
[0061] Coding Tree Block (CTB): "CTB" can be used as a term to specify any one of the Y coding tree block, Cb coding tree block, and Cr coding tree block.
[0062] Neighboring block: A neighboring block (or adjacent block) can refer to a block that is adjacent to the target block. A neighboring block can also refer to a neighboring block that is being reconstructed.
[0063] In the following text, the terms “neighboring block” and “adjacent block” may be used to have the same meaning and may be used interchangeably.
[0064] Neighboring blocks can refer to the neighboring blocks that are being reconstructed.
[0065] Spatial neighbor block: A spatial neighbor block can be a block that is spatially adjacent to the target block. Neighbor blocks can include spatial neighbor blocks.
[0066] - Target blocks and spatially adjacent blocks can be included in the target frame.
[0067] - A spatially adjacent block can refer to a block whose boundary is in contact with the target block, or a block located within a predetermined distance from the target block.
[0068] - A spatial neighboring block can refer to a block that is adjacent to the vertex of the target block. Here, a block that is adjacent to the vertex of the target block can refer to a block that is vertically adjacent to a neighboring block that is horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block that is vertically adjacent to the target block.
[0069] Temporally neighboring blocks: Temporally neighboring blocks can be blocks that are temporally adjacent to the target block. Neighboring blocks can include temporally neighboring blocks.
[0070] -Time-proximity blocks can include col blocks.
[0071] The col block can be a block in a previously reconstructed co-location frame (col frame). The position of the col block in the col frame can correspond to the position of the target block in the target frame. Optionally, the position of the col block in the col frame can be equal to the position of the target block in the target frame. The col frame can be a frame included in the list of reference frames.
[0072] - A temporally neighboring block can be a block that is temporally adjacent to the spatially neighboring block of the target block.
[0073] Prediction mode: Prediction mode can be information indicating the mode used for intra-frame prediction or the mode used for inter-frame prediction.
[0074] Prediction Unit: A prediction unit can be a basic unit used for prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation).
[0075] A single prediction unit can be divided into multiple partitions or sub-prediction units with smaller sizes. Multiple partitions can also be the basic units in the execution of prediction or compensation. Partitions generated by dividing prediction units can also be prediction units.
[0076] Prediction cell partitioning: Prediction cell partitioning can be the shape in which prediction cells are divided.
[0077] Reconstructed neighboring cells: The reconstructed neighboring cells can be cells that have already been decoded and reconstructed and are adjacent to the target cell.
[0078] - The reconstructed neighboring units can be those that are spatially adjacent to the target unit or temporally adjacent to the target unit.
[0079] - The reconstructed spatial neighbor unit can be a unit that is included in the target image and has already been reconstructed through encoding and / or decoding.
[0080] - The reconstructed temporal neighbor unit can be a unit included in the reference image and has been reconstructed through encoding and / or decoding. The position of the reconstructed temporal neighbor unit in the reference image can be the same as, or correspond to, the position of the target unit in the target image. Furthermore, the reconstructed temporal neighbor unit can be a block adjacent to a corresponding block in the reference image. Here, the position of the corresponding block in the reference image can correspond to the position of the target block in the target image. The fact that the positions of the blocks correspond to each other can mean that the positions of the blocks are the same, that one block is included in another block, or that one block occupies a specific position in another block.
[0081] Sub-screen: A screen can be divided into one or more sub-screens. A sub-screen can consist of one or more parallel block rows and one or more parallel block columns.
[0082] - A sub-screen can be an area in the screen that has a square or rectangular shape (i.e., a non-square or rectangular shape). In addition, a sub-screen can include one or more CTUs.
[0083] - A sub-picture can be a rectangular area of one or more strips in the picture.
[0084] A sub-screen may include one or more parallel blocks, one or more sub-blocks, and / or one or more stripes.
[0085] Parallel blocks: Parallel blocks can be areas in the image that have a square or rectangular shape (i.e., non-square or rectangular).
[0086] - A parallel block may contain one or more CTUs.
[0087] - Parallel blocks can be partitioned into one or more blocks.
[0088] Blocking: Blocks can represent one or more CTU lines in a parallel block.
[0089] - A parallel block can be partitioned into one or more blocks. Each block may contain one or more CTU rows.
[0090] - Parallel blocks that are not partitioned into two parts can also be represented as blocks.
[0091] Strip: A strip may include one or more parallel blocks of a frame. Optionally, a strip may include one or more sub-blocks of a parallel block.
[0092] A sub-picture can contain one or more stripes that share a rectangular area covering the entire picture. Therefore, each sub-picture boundary is always a stripe boundary, and each vertical sub-picture boundary is always a vertical parallel block boundary.
[0093] Parameter set: The parameter set can correspond to the header information in the internal structure of the bit stream.
[0094] The parameter set may include at least one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Decoding Parameter Set (DPS).
[0095] Information transmitted via signals for each parameter set can be applied to the screen referencing the corresponding parameter set. For example, information in a VPS can be applied to the screen referencing a VPS. Information in an SPS can be applied to the screen referencing an SPS. Information in a PPS can be applied to the screen referencing a PPS.
[0096] - Each parameter set can reference a higher-level parameter set. For example, PPS can reference SPS. SPS can reference VPS.
[0097] Additionally, the parameter set may include parallel block groups, stripe header information, and parallel block header information. A parallel block group can be a group comprising multiple parallel blocks. Furthermore, the meaning of "parallel block group" can be the same as that of "strip".
[0098] Rate-distortion optimization: Coding devices can use rate-distortion optimization to provide high coding efficiency by utilizing a combination of coding unit (CU) size, prediction mode, prediction unit (PU) size, motion information, and transform unit (TU) size.
[0099] Rate distortion optimization schemes can calculate the rate distortion cost of each combination in order to select the optimal combination. The equation "D+( × RThe rate-distortion cost is calculated using the formula "( )". Typically, the combination that minimizes the rate-distortion cost is chosen as the optimal combination in the rate-distortion optimization scheme.
[0100] -D can represent distortion. D can be the mean of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (i.e., the mean square error).
[0101] -R can represent the rate, which can represent the bit rate using relevant context information.
[0102] - R represents the Lagrange multiplier. It can include not only coding parameter information (such as prediction mode, motion information, and coded block flags), but also bits generated due to the encoding of the transform coefficients.
[0103] - Encoding devices can perform processes such as inter-frame prediction and / or intra-frame prediction, transform, quantization, entropy coding, inverse quantization (dequantization), and / or inverse transform to compute accurate D and R. These processes can significantly increase the complexity of the encoding device.
[0104] - Bitstream: A bitstream can represent a stream of bits that includes encoded image information.
[0105] Parsing: Parsing can be a determination of the values of syntax elements made by performing entropy decoding on a bitstream. Alternatively, the term "parsing" can refer to such entropy decoding itself.
[0106] Symbols: A symbol can be at least one of the syntax elements, encoding parameters, and transform coefficients of the encoding target unit and / or the decoding target unit. Furthermore, a symbol can be the target of entropy encoding or the result of entropy decoding.
[0107] Reference frame: The reference frame can be an image referenced by the cell to perform inter-frame prediction or motion compensation. Optionally, the reference frame can be an image that includes reference cells referenced by the target cell to perform inter-frame prediction or motion compensation.
[0108] In the following text, the terms “reference screen” and “reference image” may be used to have the same meaning and may be used interchangeably.
[0109] Reference frame list: The reference frame list can be a list of one or more reference frames used for inter-frame prediction or motion compensation.
[0110] - The types of reference screen lists may include list combination (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc.
[0111] - For inter-frame prediction, one or more reference frame lists can be used.
[0112] Inter-frame prediction indicator: The inter-frame prediction indicator indicates the direction of inter-frame prediction for the target cell. Inter-frame prediction can be either unidirectional or bidirectional. Optionally, the inter-frame prediction indicator can represent the number of reference frames used to generate the prediction cell for the target cell. Optionally, the inter-frame prediction indicator can represent the number of prediction blocks used for inter-frame prediction or motion compensation of the target cell.
[0113] Prediction list utilization flags: Prediction list utilization flags indicate whether to use at least one reference screen from a specific reference screen list to generate prediction cells.
[0114] - The prediction list utilization flag can be used to derive the inter-frame prediction indicator. Conversely, the inter-frame prediction indicator can be used to derive the prediction list utilization flag. For example, a prediction list utilization flag indicating a value of "0" for the first value indicates that, for the target cell, reference frames from the reference frame list are not used to generate the prediction block. A prediction list utilization flag indicating a value of "1" for the second value indicates that, for the target cell, the reference frame list is used to generate the prediction cell.
[0115] Reference screen index: The reference screen index can be an index that indicates a specific reference screen in the list of reference screens.
[0116] Screen Order Count (POC): The POC value of a screen indicates the order in which the corresponding screens are displayed.
[0117] Motion Vector (MV): A motion vector can be a 2D vector used for inter-frame prediction or motion compensation. A motion vector can refer to the offset between a target image and a reference image.
[0118] - For example, MV can be represented as such as (mv x , mv y In the form of ) mv x It can indicate the horizontal component, and mv y It can indicate the vertical component.
[0119] - Search Range: The search range can be a 2D region where the search for the MV is performed during inter-frame prediction. For example, the size of the search range can be M×N. M and N can be positive integers.
[0120] Motion vector candidates: Motion vector candidates can be blocks that are used as prediction candidates when predicting motion vectors, or motion vectors that are used as prediction candidates.
[0121] - Motion vector candidates can be included in the motion vector candidate list.
[0122] Motion vector candidate list: The motion vector candidate list can be a list using one or more motion vector candidate configurations.
[0123] Motion vector candidate index: The motion vector candidate index can be an indicator used to indicate motion vector candidates in the motion vector candidate list. Optionally, the motion vector candidate index can be an index of motion vector predictors.
[0124] Motion information: Motion information may include at least one of the following: a list of reference frames, a reference image, motion vector candidates, a motion vector candidate index, a merge candidate and a merge index, as well as information on motion vectors, reference frame indexes and inter-frame prediction indicators.
[0125] Merge candidate list: The merge candidate list can be a list that uses one or more merge candidate configurations.
[0126] Merging Candidates: Merging candidates can be spatial merging candidates, temporal merging candidates, combined merging candidates, combined dual-prediction merging candidates, history-based candidates, candidates based on the average of two candidates, zero merging candidates, etc. Merging candidates may include inter-frame prediction indicators and may include motion information, such as prediction type information, reference frame index for each list, motion vectors, prediction list utilization flags, and inter-frame prediction indicators.
[0127] Merge index: A merge index can be an indicator used to indicate merge candidates in a merge candidate list.
[0128] - The merge index can indicate the candidate reconstruction cells for deriving the merge between reconstruction cells that are spatially adjacent to the target cell and reconstruction cells that are temporally adjacent to the target cell.
[0129] - The merge index can indicate at least one of the multiple motion information candidates to be merged.
[0130] Transform Unit: A transform unit can be the basic unit for residual signal encoding and / or residual signal decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding, and transform coefficient decoding). A single transform unit can be partitioned into multiple sub-transform units with smaller sizes. Here, the transform may include one or more primary transforms and secondary transforms, and the inverse transform may include one or more primary inverse transforms and secondary inverse transforms.
[0131] Scaling: Scaling can represent the process of multiplying a factor by a transformation coefficient level.
[0132] - As a result of scaling the transform coefficient levels, transform coefficients can be generated. Scaling can also be referred to as "inverse quantization".
[0133] Quantization parameter (QP): The quantization parameter can be a value used to generate a transform coefficient level for the transform coefficients during quantization. Optionally, the quantization parameter can also be a value used to generate the transform coefficients by scaling the transform coefficient level during dequantization. Optionally, the quantization parameter can be a value mapped to the quantization step size.
[0134] Incremental quantization parameter: The incremental quantization parameter can refer to the difference between the predicted quantization parameter and the quantization parameter of the target cell.
[0135] Scan: A scan can refer to a method used to align the order of coefficients in cells, blocks, or matrices. For example, a method for aligning a 2D array in the form of a one-dimensional (1D) array can be called a "scan". Alternatively, a method for aligning a 1D array in the form of a 2D array can also be called a "scan" or "inverse scan".
[0136] Transform coefficients: Transform coefficients can be coefficient values generated when the encoding device performs a transform. Optionally, transform coefficients can be coefficient values generated when the decoding device performs at least one of entropy decoding and dequantization.
[0137] - The quantization level or quantization transform coefficient level generated by applying quantization to the transform coefficients or residual signal can also be included in the meaning of the term "transform coefficient".
[0138] Quantization level: The quantization level can be a value generated when the encoding device performs quantization on the transform coefficients or residual signal. Optionally, the quantization level can be a value that serves as the target for dequantization when the decoding device performs dequantization.
[0139] - The quantization transformation coefficient level, which is a result of transformation and quantization, can also be included in the meaning of quantization level.
[0140] Non-zero transform coefficients: Non-zero transform coefficients can be transform coefficients with values other than 0 or transform coefficient classes with values other than 0. Optionally, non-zero transform coefficients can be transform coefficients whose values are not zero, or transform coefficient classes whose values are not zero.
[0141] Quantization matrix: A quantization matrix is a matrix used during the quantization or dequantization process to improve the subjective or objective image quality of an image. A quantization matrix can also be referred to as a "scaling list".
[0142] Quantization matrix coefficients: Quantization matrix coefficients can be each element in the quantization matrix. Quantization matrix coefficients are also referred to as "matrix coefficients".
[0143] Default matrix: The default matrix can be a quantization matrix predefined by the encoding and decoding devices.
[0144] Non-default matrix: A non-default matrix can be a quantization matrix that is not predefined by the encoding and decoding devices. A non-default matrix can refer to a quantization matrix that is transmitted by the user from the encoding device to the decoding device.
[0145] Most Probable Mode (MPM): MPM can represent an intra-prediction mode that has a high probability of being used for intra-prediction of the target block.
[0146] Encoding and decoding devices can determine one or more MPMs based on encoding parameters associated with the target block and attributes of entities associated with the target block.
[0147] Encoding and decoding devices can determine one or more MPMs based on the intra-prediction modes of a reference block. A reference block may include multiple reference blocks. These multiple reference blocks may include the spatially adjacent block to the left of the target block and the spatially adjacent block to the top of the target block. In other words, depending on which intra-prediction modes have been used for the reference blocks, one or more distinct MPMs can be determined.
[0148] - One or more MPMs can be identified in the same way in both the encoding and decoding devices. That is, the encoding and decoding devices can share the same list of MPMs, which includes one or more MPMs.
[0149] MPM List: The MPM list can be a list that includes one or more MPMs. The number of one or more MPMs in the MPM list can be predefined.
[0150] MPM Indicator: The MPM indicator can indicate one or more MPMs in the MPM list that will be used for intra-prediction of the target block. For example, the MPM indicator can be an index to the MPM list.
[0151] Since the MPM list is determined in the same way in both the encoding and decoding devices, it is not necessary to send the MPM list itself from the encoding device to the decoding device.
[0152] - An MPM indicator can be signaled from the encoding device to the decoding device. When the MPM indicator is signaled, the decoding device can determine the MPM to be used for intra-frame prediction of the target block from the MPMs in the MPM list.
[0153] MPM Usage Indicator: The MPM usage indicator indicates whether an MPM usage mode will be used for the prediction of the target block. The MPM usage mode can be the mode of the MPM that will be used for intra-frame prediction of the target block, determined using the MPM list.
[0154] - An MPM usage indicator can be transmitted from the encoding device to the decoding device via a signal.
[0155] Transmitted by signal: "Transmitted by signal" can mean that information is transmitted from an encoding device to a decoding device. Alternatively, "transmitted by signal" can mean that information is included in a bitstream or recording medium by an encoding device. Information transmitted by signal by an encoding device can be used by a decoding device.
[0156] An encoding device generates encoded information by encoding the information to be transmitted via a signal. The encoded information can be sent from the encoding device to a decoding device. The decoding device obtains the information by decoding the sent encoded information. Here, the encoding can be entropy encoding, and the decoding can be entropy decoding.
[0157] Selective signal transmission: Information can be selectively transmitted using signals. Selective signal transmission of information can mean that the encoding device (based on specific conditions) selectively includes information in the bitstream or recording medium. Selective signal transmission of information can also mean that the decoding device (based on specific conditions) selectively extracts information from the bitstream.
[0158] The omission of signal transmission: The signal transmission of information can be omitted. Omitting the signal transmission of information may mean that the encoding device (under certain conditions) does not include information in the bitstream or recording medium. The omission of signal transmission of information may also mean that the decoding device (under certain conditions) does not extract information from the bitstream.
[0159] Statistical values: Variables, coding parameters, constants, etc., can have computable values. Statistical values can be values generated by performing calculations (operations) on the values of a specified target. For example, a statistical value can indicate one or more of the following: the mean, weighted average, weighted sum, minimum, maximum, modulo, median, and interpolation of the values of a specific variable, a specific coding parameter, a specific constant, etc.
[0160] Figure 1 This is a block diagram illustrating the configuration of an embodiment of the encoding device applying the present disclosure.
[0161] Encoding device 100 can be an encoder, a video encoding device, or an image encoding device. Video may include one or more images (frames). Encoding device 100 can sequentially encode one or more images of the video.
[0162] Reference Figure 1 The encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.
[0163] The encoding device 100 can perform encoding on the target image using intra-frame mode and / or inter-frame mode. In other words, the prediction mode for the target block can be one of intra-frame mode and inter-frame mode.
[0164] In the following text, the terms "intra-frame mode", "intra-frame prediction mode", "in-frame mode" and "in-frame prediction mode" may be used to have the same meaning and may be used interchangeably.
[0165] In the following text, the terms "inter-frame mode", "inter-frame prediction mode", "inter-picture mode" and "inter-picture prediction mode" may be used to have the same meaning and may be used interchangeably.
[0166] In the following text, the term "image" may refer to only a portion of an image, or it may refer to a block. Furthermore, the processing of an "image" may refer to the sequential processing of multiple blocks.
[0167] Furthermore, the encoding device 100 can generate a bitstream including encoded information by encoding the target image, and can output and store the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed via wired and / or wireless transmission media.
[0168] When intra-frame mode is used as prediction mode, switcher 115 can switch to intra-frame mode. When inter-frame mode is used as prediction mode, switcher 115 can switch to inter-frame mode.
[0169] The encoding device 100 can generate a prediction block of the target block. Furthermore, after generating the prediction block, the encoding device 100 can encode the residual block of the target block using the residual between the target block and the prediction block.
[0170] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use pixels from previously encoded / decoded neighboring blocks adjacent to the target block as reference samples. The intra-frame prediction unit 120 can use the reference samples to perform spatial prediction on the target block and can generate prediction samples for the target block via spatial prediction. Prediction samples can refer to samples in the prediction block.
[0171] The inter-frame prediction unit 110 may include a motion prediction unit and a motion compensation unit.
[0172] When the prediction mode is inter-frame mode, the motion prediction unit can search for the region in the reference image that most closely matches the target block during motion prediction, and can derive motion vectors for the target block and the found region based on the found region. Here, the motion prediction unit can use the search range as the target region for the search.
[0173] A reference image can be stored in a reference frame buffer 190. More specifically, when the encoding and / or decoding of a reference image has been processed, the encoded and / or decoded reference image can be stored in the reference frame buffer 190.
[0174] Since the decoded screen is stored, the reference screen buffer 190 can be the decoded screen buffer (DPB).
[0175] The motion compensation unit can generate a predicted block for the target block by performing motion compensation using motion vectors. Here, the motion vector can be a two-dimensional (2D) vector used for inter-frame prediction. Furthermore, the motion vector can indicate the offset between the target image and the reference image.
[0176] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit can generate prediction blocks by applying an interpolation filter to a portion of the reference image. To perform inter-frame prediction or motion compensation, it can be determined which of the following modes—skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current frame reference mode—corresponds to the method for predicting and compensating for the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation can be performed according to that mode.
[0177] Subtractor 125 generates a residual block, which is the difference between the target block and the prediction block. The residual block can also be called a "residual signal".
[0178] The residual signal can be the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming or quantizing the difference between the original signal and the predicted signal, or by transforming and quantizing the difference. The residual block can be the residual signal for a block unit.
[0179] The transformation unit 130 can generate transformation coefficients by transforming the residual block, and can output the generated transformation coefficients. Here, the transformation coefficients can be coefficient values generated by transforming the residual block.
[0180] Transformation unit 130 may use one of a number of predefined transformation methods when performing a transformation.
[0181] Several predefined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), etc.
[0182] The transformation method for transforming the residual block can be determined based on at least one of the coding parameters used for the target block and / or neighboring blocks. For example, the transformation method can be determined based on at least one of the inter-frame prediction mode for the PU, the intra-frame prediction mode for the PU, the size of the TU, and the shape of the TU. Optionally, transformation information indicating the transformation method can be signaled from the encoding device 100 to the decoding device 200.
[0183] When using the transform skip mode, the transform unit 130 can omit the transformation of the residual block.
[0184] By applying quantization to the transform coefficients, a quantized transform coefficient level or quantization level can be generated. In the following examples, each of the quantized transform coefficient level and quantization level may also be referred to as a "transform coefficient".
[0185] The quantization unit 140 can generate a quantized transform coefficient level (i.e., a quantization level or quantization coefficient) by quantizing the transform coefficients according to quantization parameters. The quantization unit 140 can output the generated quantized transform coefficient level. In this case, the quantization unit 140 can use a quantization matrix to quantize the transform coefficients.
[0186] Entropy coding unit 150 can generate a bitstream by performing probability distribution-based entropy coding based on values calculated by quantization unit 140 and / or coding parameter values calculated during the encoding process. Entropy coding unit 150 can output the generated bitstream.
[0187] The entropy coding unit 150 can perform entropy coding on information about the pixels of the image and information required to decode the image. For example, the information required to decode the image may include syntax elements, etc.
[0188] When applying entropy coding, fewer bits can be allocated to more frequently occurring symbols, and more bits can be allocated to less frequently occurring symbols. Since symbols are represented through this allocation, the size of the bit string of the target symbol to be encoded can be reduced. Therefore, entropy coding can improve the compression performance of video coding.
[0189] Furthermore, for entropy coding, the entropy coding unit 150 can use coding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), or Context Adaptive Binary Arithmetic Coding (CABAC). For example, the entropy coding unit 150 can use a variable-length code / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 can derive a binarization method for the target symbol. Furthermore, the entropy coding unit 150 can derive a probabilistic model for the target symbol / bit. The entropy coding unit 150 can use the derived binarization method, probabilistic model, and context model to perform arithmetic coding.
[0190] The entropy coding unit 150 can transform coefficients in 2D block form into 1D vector form through a transform coefficient scanning method in order to encode the quantization transform coefficient level.
[0191] Encoding parameters can be information required for encoding and / or decoding. Encoding parameters may include information encoded by encoding device 100 and transmitted from encoding device 100 to decoding device, and may also include information that can be deduced during encoding or decoding. For example, information transmitted to decoding device may include syntax elements.
[0192] Encoding parameters can include not only information (or flags or indexes) encoded by the encoding device and transmitted by the encoding device to the decoding device via signals, such as syntax elements, but also information deduced during the encoding or decoding process. Furthermore, encoding parameters can include information required for encoding or decoding an image. For example, encoding parameters can include at least one value, combination, or statistic of the following: unit / block size, unit / block shape / form, unit / block depth, unit / block partitioning information, unit / block partitioning structure, information indicating whether a unit / block is partitioned in a quadtree structure, information indicating whether a unit / block is partitioned in a binary tree structure, partitioning direction (horizontal or vertical) of the binary tree structure, partitioning form (symmetric or asymmetric) of the binary tree structure, information indicating whether a unit / block is partitioned in a ternary tree structure, partitioning direction (horizontal or vertical) of the ternary tree structure, and partitioning form (symmetric or asymmetric) of the ternary tree structure. Information such as: whether the indicator unit / block is partitioned in a multi-type tree structure; the combination and direction of partitions in the multi-type tree structure (horizontal or vertical, etc.); the partition form of the multi-type tree structure (symmetric or asymmetric partitions, etc.); the partition tree in the multi-type tree form (binary or ternary tree); the prediction type (intra-frame prediction or inter-frame prediction); the intra-frame prediction mode / direction; the intra-frame luma prediction mode / direction; the intra-frame chroma prediction mode / direction; intra-frame partition information; inter-frame partition information; coded block partition flag; prediction block partition flag; transform block partition flag; reference sample filtering method; reference sample filter taps; reference sample filter coefficients; prediction block filtering method; pre- Prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, inter-frame prediction mode, motion information, motion vector, motion vector difference, reference frame index, inter-frame prediction direction, inter-frame prediction indicator, prediction list utilization flag, reference frame list, reference image, POC, motion vector prediction factor, motion vector prediction index, motion vector prediction candidate, motion vector candidate list, information indicating whether to use merge mode, merge index, merge candidate, merge candidate list, information indicating whether to use skip mode, type of interpolation filter, interpolation filter taps, interpolation filter filter... Filter coefficients, magnitude of motion vector, precision of motion vector representation, transform type, transform size, information indicating whether a primary transform is used, information indicating whether an additional (secondary) transform is used, primary transform selection information (or primary transform index), secondary transform selection information (or secondary transform index), information indicating the presence or absence of residual signal, code block mode, code block flag, quantization parameters, residual quantization parameters, quantization matrix, information about in-loop filters, information indicating whether an in-loop filter is applied, coefficients of the in-loop filter, taps of the in-loop filter, shape / form of the in-loop filter, information indicating whether a deblocking filter is applied, coefficients of the deblocking filter.Deblocking filter taps, deblocking filter strength, deblocking filter shape / form, information indicating whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, information indicating whether adaptive in-loop filter is applied, adaptive in-loop filter coefficients, adaptive in-loop filter taps, adaptive in-loop filter shape / form, binarization / debinarization method, context model, context model determination method, context model update method, information indicating whether to execute normal mode, information indicating whether to execute bypass mode, valid coefficient flag, last valid coefficient flag, encoding flag for coefficient group, position of last valid coefficient, information indicating whether the coefficient value is greater than 1, information indicating whether the coefficient value is greater than 2, information indicating whether the coefficient value is greater than 3, remaining coefficient value information, sign information, reconstructed luminance sample, reconstructed chrominance sample, context binary bits, bypass binary bits, residual The encoding parameters include: differential luminance samples, residual chrominance samples, transform coefficients, luminance transform coefficients, chrominance transform coefficients, quantization levels, luminance quantization levels, chrominance quantization levels, transform coefficient levels, transform coefficient level scanning method, size of the motion vector search area on the decoding device side, shape / form of the motion vector search area on the decoding device side, number of motion vector searches on the decoding device side, CTU size, minimum block size, maximum block size, maximum block depth, minimum block depth, image display / output order, stripe identification information, stripe type, stripe partition information, parallel block group identification information, parallel block group type, parallel block group partition information, parallel block identification information, parallel block type, parallel block partition information, screen type, bit depth, input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantization level bit depth, information about the luminance signal, information about the chrominance signal, color space of the target block, and color space of the residual block. Furthermore, information related to the above encoding parameters can also be included in the encoding parameters. Information used to calculate and / or derive the above encoding parameters can also be included in the encoding parameters. Information calculated or derived using the above encoding parameters can also be included in the encoding parameters.
[0193] The initial transformation selection information indicates the first transformation to be applied to the target block.
[0194] Secondary transformation selection information indicates the secondary transformations to be applied to the target block.
[0195] The residual signal can represent the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming the difference between the original signal and the predicted signal. Optionally, the residual signal can be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. A residual block can be the residual signal of a block.
[0196] Here, transmitting information via signals may mean that the encoding device 100 includes entropy-encoded information generated by performing entropy encoding on flags or indices in the bitstream, and the decoding device 200 obtains the information by performing entropy decoding on the entropy-encoded information extracted from the bitstream. Here, the information may include flags, indices, etc.
[0197] A signal can refer to information that is transmitted via a signal. In the following text, information concerning images and blocks may be referred to as a signal. Furthermore, in the following text, the terms "information" and "signal" may be used to have the same meaning and are used interchangeably. For example, a specific signal may be a signal representing a specific block. A raw signal may be a signal representing a target block. A prediction signal may be a signal representing a predicted block. A residual signal may be a signal representing a residual block.
[0198] The bitstream may include information based on a specific syntax. Encoding device 100 can generate a bitstream including information according to the specific syntax. Decoding device 200 can obtain information from the bitstream according to the specific syntax.
[0199] Since the encoding device 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for another (one or more) images to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference frame buffer 190. For decoding, inverse quantization and inverse transform of the encoded target image can be performed.
[0200] The quantization level can be dequantized by the dequantization unit 160 and inversely transformed by the inverse transform unit 170. The dequantization unit 160 can generate dequantized coefficients by performing dequantization on the quantization level. The inverse transform unit 170 can generate dequantized and inversely transformed coefficients by performing inverse transform on the dequantized coefficients.
[0201] The coefficients of the inverse quantization and inverse transform can be added to the prediction block using adder 175. Adding the coefficients of the inverse quantization and inverse transform to the prediction block then generates a reconstructed block. Here, the coefficients of the inverse quantization and / or inverse transform can represent one or more coefficients of the dequantization and inverse transform performed, and can also represent the reconstructed residual block. Here, the reconstructed block can mean either the recovered block or the decoded block.
[0202] The reconstructed blocks can be filtered by filter unit 180. Filter unit 180 can apply one or more of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), and a nonlocal filter (NLF) to the reconstructed samples, reconstructed blocks, or reconstructed images. Filter unit 180 may also be referred to as an "in-loop filter".
[0203] Deblocking filters eliminate block distortion that occurs at the boundaries between blocks in a reconstructed image. To determine whether to apply a deblocking filter, the number of rows or columns of pixels included in the block and comprising one or more pixels for the purpose of determining whether to apply the deblocking filter to the target block is determined based on those pixels.
[0204] When a deblocking filter is applied to a target block, the applied filter can vary depending on the desired deblocking strength. In other words, among different filters, the filter chosen based on the desired deblocking strength can be applied to the target block. When a deblocking filter is applied to a target block, one or more filters, such as long-tap filters, strong filters, weak filters, and Gaussian filters, can be applied to the target block according to the desired deblocking strength.
[0205] In addition, when performing vertical and horizontal filtering on the target block, the horizontal and vertical filtering can be processed in parallel.
[0206] SAO can compensate for coding errors by adding an appropriate offset to the pixel value. SAO can perform offset correction on a pixel-by-pixel basis for images that have been deblocked, using the difference between the original image and the deblocked image. To perform offset correction on an image, a method can be used to divide the pixels included in the image into a certain number of regions, determine the regions to which the offset will be applied within the divided regions, and apply the offset to the determined regions. Alternatively, a method can be used to apply the offset taking into account the edge information of each pixel.
[0207] The ALF can perform filtering based on values obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. Information related to whether an adaptive loop filter is applied can be transmitted via a signal for each CU. Such information can be transmitted via a signal for the luminance signal. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, an ALF with a fixed form can be applied to the block regardless of its characteristics.
[0208] Nonlocal filters can perform filtering based on reconstructed blocks similar to the target block. Regions similar to the target block can be selected from the reconstructed image, and the statistical properties of the selected similar regions can be used to perform filtering on the target block. Information about whether a nonlocal filter is applied can be transmitted to the coding unit (CU) via a signal. Furthermore, the shape and filter coefficients of the nonlocal filter applied to the block can vary depending on the block.
[0209] The reconstructed blocks or reconstructed image filtered by filter unit 180 can be stored as a reference frame in reference frame buffer 190. The reconstructed blocks filtered by filter unit 180 can be part of the reference frame. In other words, the reference frame can be a reconstructed frame composed of reconstructed blocks filtered by filter unit 180. The stored reference frame can then be used for inter-frame prediction or motion compensation.
[0210] Figure 2 This is a block diagram illustrating the configuration of an embodiment of the decoding device applying the present disclosure.
[0211] Decoding device 200 can be a decoder, video decoding device, or image decoding device.
[0212] Reference Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference frame buffer 270.
[0213] Decoding device 200 can receive bit streams output from encoding device 100. Decoding device 200 can receive bit streams stored in computer-readable storage media and can also receive bit streams transmitted via wired / wireless transmission media.
[0214] The decoding device 200 can perform decoding on the bitstream in intra-frame mode and / or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image via decoding, and can output the reconstructed image or the decoded image.
[0215] For example, the switcher 245 can perform a switch to intra-frame mode or inter-frame mode based on the prediction mode used for decoding. When the prediction mode used for decoding is intra-frame mode, the switcher 245 can be operated to switch to intra-frame mode. When the prediction mode used for decoding is inter-frame mode, the switcher 245 can be operated to switch to inter-frame mode.
[0216] The decoding device 200 can obtain a reconstructed residual block by decoding the input bitstream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block, which is the target to be decoded, by adding the reconstructed residual block and the prediction block.
[0217] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantization transform coefficient levels (i.e., quantization levels or quantization coefficients). Here, the entropy decoding method can be similar to the entropy encoding method described above. That is, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0218] The entropy decoding unit 210 can transform coefficients in one-dimensional (1D) vector form into 2D block shapes by a transform coefficient scanning method in order to decode the quantized transform coefficient levels.
[0219] For example, by scanning the block coefficients using a top-right diagonal scan, the block coefficients can be transformed into a 2D block shape. Optionally, the choice between a top-right diagonal scan, a vertical scan, and a horizontal scan can be determined based on the size of the corresponding block and / or the intra-frame prediction mode.
[0220] The quantization coefficients can be inversely quantized by the dequantization unit 220. The dequantization unit 220 generates inversely quantized coefficients by performing dequantization on the quantization coefficients. Furthermore, the inversely quantized coefficients can be inversely transformed by the inverse transform unit 230. The inverse transform unit 230 generates a reconstructed residual block by performing an inverse transform on the inversely quantized coefficients. As a result of performing dequantization and inverse transform on the quantization coefficients, a reconstructed residual block can be generated. Here, the dequantization unit 220 can apply the quantization matrix to the quantization coefficients when generating the reconstructed residual block.
[0221] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the target block using the pixel values of previously decoded neighboring blocks adjacent to the target block.
[0222] The inter-frame prediction unit 250 may include a motion compensation unit. Optionally, the inter-frame prediction unit 250 may be designated as a "motion compensation unit".
[0223] When using inter-frame mode, the motion compensation unit can generate a prediction block by performing motion compensation on the target block using motion vectors and a reference image stored in the reference frame buffer 270.
[0224] The motion compensation unit can apply an interpolation filter to a portion of the reference image when the motion vector has a value other than an integer, and can use the reference image with the interpolation filter applied to generate prediction blocks. To perform motion compensation, the motion compensation unit can determine, based on the CU, which of the following modes—skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current frame reference mode—corresponds to the motion compensation method used by the PU included in the CU, and can perform motion compensation according to the determined mode.
[0225] The reconstructed residual block and the prediction block can be added to each other by adder 255. Adder 255 generates the reconstructed block by adding the reconstructed residual block to the prediction block.
[0226] The reconstructed blocks can be filtered by the filter unit 260. The filter unit 260 can apply at least one of a deblocking filter, a SAO filter, an ALF filter, and an NLF filter to the reconstructed blocks or the reconstructed image. The reconstructed image can be a picture that includes the reconstructed blocks.
[0227] The filter unit can output a reconstructed image.
[0228] The reconstructed image and / or reconstructed blocks filtered by filter unit 260 can be stored as a reference image in reference image buffer 270. The reconstructed blocks filtered by filter unit 260 can be part of the reference image. In other words, the reference image can be an image composed of reconstructed blocks filtered by filter unit 260. The stored reference image can then be used for inter-frame prediction or motion compensation.
[0229] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded.
[0230] Figure 3 An example can be illustrated by showing a single cell divided into multiple sub-cells.
[0231] To effectively partition an image, coding units (CUs) can be used in encoding and decoding. The term "unit" can be used to generally specify 1) a block that comprises image samples and 2) a syntax element. For example, "partition of a unit" can mean "partition of a block corresponding to a unit".
[0232] A CU (Frame Controller) can be used as the basic unit for image encoding / decoding. A CU can be used as the unit to which a mode, selected from intra-frame and inter-frame modes, is applied during image encoding / decoding. In other words, during image encoding / decoding, it can be determined which of the intra-frame and inter-frame modes will be applied to each CU.
[0233] Furthermore, the CU can be the basic unit in the prediction, transformation, quantization, inverse transformation, dequantization, and encoding / decoding of transform coefficients.
[0234] Reference Figure 3 The image 300 can be sequentially partitioned into units corresponding to the largest coding unit (LCU), and the partitioning structure can be determined for each LCU. Here, LCU can be used to have the same meaning as coding tree unit (CTU).
[0235] A cell partition can refer to the partition of the block corresponding to that cell. Block partitioning information can include depth information about the depth of the cell. Depth information can indicate the number of times the cell is partitioned and / or the degree to which the cell is partitioned. A single cell can be hierarchically partitioned into multiple sub-cells with depth information based on a tree structure.
[0236] Each sub-unit within a partition can have depth information. This depth information can be information indicating the size of the CU. Depth information can be stored for each CU.
[0237] Each CU can have depth information. When a CU is partitioned, the CU generated through partitioning can have a depth that is 1 greater than the depth of the partitioned CU.
[0238] The partitioning structure refers to the distribution of coding units (CUs) in the LCU 310 used to efficiently encode an image. This distribution can be determined by whether a single CU will be partitioned into multiple CUs. The number of CUs generated through partitioning can be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, etc.
[0239] Depending on the number of CUs generated through partitioning, the horizontal and vertical dimensions of each CU generated through partitioning can be smaller than the horizontal and vertical dimensions of the CU before partitioning. For example, the horizontal and vertical dimensions of each CU generated through partitioning can be half the horizontal and vertical dimensions of the CU before partitioning.
[0240] Each partition's CU can be recursively partitioned into four CUs in the same manner. Through recursive partitioning, at least one of the horizontal and vertical dimensions of the CU in each partition can be reduced compared to at least one of the horizontal and vertical dimensions of the CU before partitioning.
[0241] The partitioning of a CU can be performed recursively up to a predefined depth or predefined size.
[0242] For example, the depth of the CU can have a value ranging from 0 to 3. Depending on the depth of the CU, the size of the CU can range from 64×64 to 8×8.
[0243] For example, the depth of LCU 310 can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, as mentioned above, the LCU can be a CU with the maximum coding unit size, and the SCU can be a CU with the minimum coding unit size.
[0244] Partitioning can begin with LCU 310, and the depth of the CU can be increased by 1 whenever the horizontal and / or vertical dimensions of the CU are reduced by partitioning.
[0245] For example, for each depth, an unpartitioned CU can have a size of 2N×2N. Furthermore, when the CU is partitioned, a CU of size 2N×2N can be partitioned into four CUs, each with a size of N×N. The value of N is halved each time the depth increases by 1.
[0246] Reference Figure 3 An LCU with a depth of 0 can have 64×64 pixels or 64×64 blocks. 0 can be the minimum depth. An SCU with a depth of 3 can have 8×8 pixels or 8×8 blocks. 3 can be the maximum depth. Here, a CU with 64×64 blocks as an LCU can be represented by depth 0. A CU with 32×32 blocks can be represented by depth 1. A CU with 16×16 blocks can be represented by depth 2. A CU with 8×8 blocks as an SCU can be represented by depth 3.
[0247] Information about whether a corresponding CU is partitioned can be represented by the CU's partition information. Partition information can be 1 bit. All CUs except the SCU can include partition information. For example, the partition information value of an unpartitioned CU can be a first value. The partition information value of a partitioned CU can be a second value. When the partition information indicates whether a CU is partitioned, the first value can be "0" and the second value can be "1".
[0248] For example, when a single CU is partitioned into four CUs, the horizontal and vertical dimensions of each of the four CUs generated by the partitioning can be half the horizontal and vertical dimensions of the original CU. When a CU with a size of 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned using a quadtree structure. In other words, it can be considered that quadtree partitioning has been applied to the CU.
[0249] For example, when a single CU is partitioned into two CUs, the horizontal or vertical dimension of each of the two resulting CUs can be half the horizontal or vertical dimension of the original CU. When a CU with a size of 32×32 is vertically partitioned into two CUs, the size of each of the two resulting CUs can be 16×32. When a CU with a size of 32×32 is horizontally partitioned into two CUs, the size of each of the two resulting CUs can be 32×16. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned using a binary tree structure. In other words, it can be considered that binary tree partitioning has been applied to the CU.
[0250] For example, when a single CU is partitioned (or divided) into three CUs, the original CU before partitioning is partitioned such that its horizontal or vertical dimensions are divided in a 1:2:1 ratio, thus enabling the generation of three sub-CUs. For example, when a CU with dimensions of 16×32 is horizontally partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have dimensions of 16×8, 16×16, and 16×8 respectively from top to bottom. For example, when a CU with dimensions of 32×32 is vertically partitioned into three sub-CUs, the three sub-CUs generated by the partitioning can have dimensions of 8×32, 16×32, and 8×32 respectively from left to right. When a single CU is partitioned into three CUs, the CU can be considered to be partitioned in the form of a ternary tree. In other words, ternary tree partitioning can be considered to have been applied to the CU.
[0251] Both quadtree partitioning and binary tree partitioning are used. Figure 3 LCU 310.
[0252] In the encoding device 100, a 64×64 coding tree unit (CTU) can be partitioned into multiple smaller CUs using a recursive quadtree structure. A single CU can be partitioned into four CUs of the same size. Each CU can be recursively partitioned and can have a quadtree structure.
[0253] By using recursive partitioning of the CU, the optimal partitioning method that results in the minimum rate distortion cost can be selected.
[0254] Figure 3 The Code Tree Unit (CTU) 320 in the example is an example of all CTUs that apply quadtree partitioning, binary tree partitioning, and ternary tree partitioning.
[0255] As mentioned above, to partition a CTU, at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning can be applied to the CTU. Partitioning can be applied based on specific priorities.
[0256] For example, quadtree partitioning can be preferentially applied to CTUs. CUs that cannot be further partitioned in quadtree form can correspond to leaf nodes of a quadtree. CUs corresponding to leaf nodes of a quadtree can be the root nodes of a binary tree and / or a ternary tree. That is, CUs corresponding to leaf nodes of a quadtree can be partitioned in binary or ternary tree form, or may not be further partitioned. In this case, it prevents each CU generated by applying binary or ternary tree partitioning to the CU corresponding to the leaf node of the quadtree from undergoing quadtree partitioning again, thereby effectively performing block partitioning and / or signaling of block partitioning information.
[0257] Four-partition information can be used to signal the partitions of a CU corresponding to each node of a quadtree. A four-partition message with a first value (e.g., "1") indicates that the corresponding CU is partitioned in quadtree form. A four-partition message with a second value (e.g., "0") indicates that the corresponding CU is not partitioned in quadtree form. The four-partition message can be a flag with a specific length (e.g., 1 bit).
[0258] There may be no priority between binary tree partitions and ternary tree partitions. That is, a CU corresponding to a leaf node of a quadtree can be partitioned in either binary or ternary tree form. Furthermore, CUs generated by binary or ternary tree partitions can be further partitioned in either binary or ternary tree form, or they may not be further partitioned.
[0259] A partition executed when there is no priority between a binary tree partition and a ternary tree partition can be called a "multi-type tree partition". That is, a CU corresponding to a leaf node of a quadtree can be the root node of a multi-type tree. The partitioning of the CU corresponding to each node of the multi-type tree can be signaled using at least one of the following: information indicating whether the CU is partitioned as a multi-type tree, partitioning direction information, and partitioning tree information. For the partitioning of the CU corresponding to each node of the multi-type tree, the information indicating whether to execute the multi-type tree partitioning, partitioning direction information, and partitioning tree information can be signaled sequentially.
[0260] For example, information indicating whether a CU is partitioned in a multi-type tree and having a first value (e.g., "1") can indicate that the corresponding CU is partitioned in a multi-type tree format. Information indicating whether a CU is partitioned in a multi-type tree and having a second value (e.g., "0") can indicate that the corresponding CU is not partitioned in a multi-type tree format.
[0261] When the CU corresponding to each node of the multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU may further include partitioning direction information.
[0262] Partition direction information indicates the partitioning direction of a multi-type tree partition. Partition direction information with a first value (e.g., "1") indicates that the corresponding CU is partitioned in the vertical direction. Partition direction information with a second value (e.g., "0") indicates that the corresponding CU is partitioned in the horizontal direction.
[0263] When a CU corresponding to each node of a multi-type tree is partitioned in the form of a multi-type tree, the corresponding CU may further include partition tree information. The partition tree information may indicate the tree used for multi-type tree partitioning.
[0264] For example, partition tree information with a first value (e.g., "1") can indicate that the corresponding CU is partitioned in a binary tree format. Partition tree information with a second value (e.g., "0") can indicate that the corresponding CU is partitioned in a ternary tree format.
[0265] Here, each of the above indications regarding whether to execute the partitioning information, partition tree information, and partitioning direction information of the multi-type tree can be a flag with a specific length (e.g., 1 bit).
[0266] Entropy encoding and / or entropy decoding can be performed on at least one of the above four partition information, information indicating whether to perform partitioning with multiple tree types, partition direction information, and partition tree information. To perform such entropy encoding / decoding of information, information from neighboring CUs adjacent to the target CU can be used.
[0267] For example, it can be assumed that the partitioning patterns (i.e., partitioned / non-partitioned, partitioned tree, and / or partitioned orientation) of the left and / or upper CUs are highly similar to the partitioning pattern of the target CU. Therefore, based on the information of neighboring CUs, contextual information for entropy encoding and / or entropy decoding of the information for the target CU can be derived. Here, the information of neighboring CUs may include at least one of the following: 1) four-partition information of neighboring CUs, 2) information indicating whether neighboring CUs are partitioned in a multi-type tree, 3) partitioned orientation information of neighboring CUs, and 4) partitioned tree information of neighboring CUs.
[0268] In another embodiment, binary tree partitioning can be performed first over ternary tree partitioning. That is, binary tree partitioning can be applied first, and then the CU corresponding to the leaf node of the binary tree can be set as the root node of the ternary tree. In this case, quadtree partitioning or binary tree partitioning can be omitted from the CU corresponding to the node of the ternary tree.
[0269] A CU that is not further partitioned by quadtree partitioning, binary tree partitioning, and / or ternary tree partitioning can be a unit for encoding, prediction, and / or transformation. That is, a CU may not be further partitioned for prediction and / or transformation. Therefore, the partitioning structure used to partition a CU into prediction units (PUs) and / or transformation units (TUs), its partitioning information, etc., may not exist in the bitstream.
[0270] However, when the size of a CU (Computer Unit) used as a partitioning unit is larger than the size of the largest transform block, the CU can be recursively partitioned until the size of the CU becomes smaller than or equal to the size of the largest transform block. For example, when the size of the CU is 64×64 and the size of the largest transform block is 32×32, the CU can be partitioned into four 32×32 blocks to perform the transform. Similarly, when the size of the CU is 32×64 and the size of the largest transform block is 32×32, the CU can be partitioned into two 32×32 blocks.
[0271] In this scenario, it is not necessary to separately transmit information indicating whether a CU has been partitioned for transformation. Without signal transmission, partitioning of the CU can be determined by comparing its horizontal (and / or vertical) dimensions with the horizontal (and / or vertical) dimensions of the largest transform block. For example, if the horizontal dimension of the CU is greater than the horizontal dimension of the largest transform block, the CU can be vertically partitioned. Similarly, if the vertical dimension of the CU is greater than the vertical dimension of the largest transform block, the CU can be horizontally partitioned.
[0272] Information regarding the maximum and / or minimum size of the CU, as well as the maximum and / or minimum size of the transform block, can be transmitted or determined at a higher level than the CU level. For example, a higher level could be a sequence level, picture level, parallel block level, parallel block group level, or stripe level. For example, the minimum size of the CU could be set to 4×4. For example, the maximum size of the transform block could be set to 64×64. For example, the maximum size of the transform block could be set to 4×4.
[0273] Information regarding the minimum size of the CU corresponding to the leaf node of the quadtree (i.e., the minimum size of the quadtree) and / or the maximum depth of the path from the root node to the leaf node of the multi-type tree (i.e., the maximum depth of the multi-type tree) can be signaled or determined at a higher level than the level of the CU corresponding to the leaf node of the quadtree. For example, higher levels can be sequence levels, picture levels, stripe levels, parallel block group levels, or parallel block levels. Information regarding the minimum size of the quadtree and / or the maximum depth of the multi-type tree can be signaled or determined individually at each of the intra-strip and inter-strip levels.
[0274] Information about the difference between the size of the CTU and the maximum size of the transform block can be transmitted or determined at a higher level than the CU level. For example, a higher level could be a sequence level, picture level, strip level, parallel block group level, or parallel block level. Information about the maximum size of the CU corresponding to each node of the binary tree (i.e., the maximum size of the binary tree) can be determined based on the size and difference information of the CTU. The maximum size of the CU corresponding to each node of the ternary tree (i.e., the maximum size of the ternary tree) can have different values depending on the stripe type. For example, the maximum size of the ternary tree in an intra-strip level could be 32×32. For example, the maximum size of the ternary tree in an inter-strip level could be 128×128. For example, the minimum size of the CU corresponding to each node of the binary tree (i.e., the minimum size of the binary tree) and / or the minimum size of the CU corresponding to each node of the ternary tree (i.e., the minimum size of the ternary tree) can be set to the minimum size of the CU.
[0275] In another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the stripe level. Additionally, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the stripe level.
[0276] Based on the various block sizes and depths mentioned above, the four partition information, the information indicating whether partitioning with multiple tree types is performed, the partition tree information, and / or the partition direction information may or may not exist in the bitstream.
[0277] For example, when the size of the CU is not greater than the minimum size of the quadtree, the CU may not include the four-partition information, and the four-partition information of the CU can be inferred as the second value.
[0278] For example, when the size (horizontal and vertical dimensions) of the CU corresponding to each node of a multi-type tree is greater than the maximum size (horizontal and vertical dimensions) of a binary tree and / or a ternary tree, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, information indicating whether to perform partitioning in a multi-type tree can be inferred instead of signaling.
[0279] Optionally, when the size (horizontal and vertical dimensions) of the CU corresponding to each node of the multi-type tree is equal to the minimum size (horizontal and vertical dimensions) of the binary tree, or when the size (horizontal and vertical dimensions) of the CU is twice the minimum size (horizontal and vertical dimensions) of the ternary tree, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, information indicating whether to perform partitioning in a multi-type tree form can be inferred instead of signal transmission. This is because when the CU is partitioned in binary and / or ternary tree form, it generates CUs smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree.
[0280] Optionally, binary or ternary partitioning can be limited based on the size of the virtual pipeline data unit (i.e., the size of the pipeline buffer). For example, binary or ternary partitioning can be limited when partitioning a CU into sub-CUs that are not suitable for the size of the pipeline buffer. The size of the pipeline buffer can be equal to the maximum size of the transform block (e.g., 64×64).
[0281] For example, when the size of the pipeline buffer is 64×64, the following partitions can be restricted.
[0282] -N×M CU of ternary tree partitions (where N and / or M are 128). -128×N CU of horizontal binary tree partitions (where N<=64) -N×128 CU of vertical binary tree partitioning (where N<=64) Alternatively, when the depth of the CU corresponding to each node of the multi-type tree is equal to the maximum depth of the multi-type tree, the CU may not be partitioned in binary and / or ternary tree form. With this determination method, information indicating whether to perform partitioning in a multi-type tree can be inferred instead of signaling.
[0283] Optionally, information indicating whether to perform partitioning in a multi-type tree can be signaled only if at least one of the vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is possible for the CU corresponding to each node of the multi-type tree. Otherwise, the CU may not be partitioned in binary tree form and / or ternary tree form. With this determination method, information indicating whether to perform partitioning in a multi-type tree can be inferred instead of signaling it.
[0284] Optionally, for each CU corresponding to a node of a multi-type tree, partitioning direction information may be signaled only when both vertical binary tree partitioning and horizontal binary tree partitioning are possible, or only when both vertical ternary tree partitioning and horizontal ternary tree partitioning are possible. Otherwise, partitioning direction information may not be signaled, but may be inferred as a value indicating the direction in which the CU can be partitioned.
[0285] Optionally, for each CU corresponding to a node of a multi-type tree, partition tree information may be signaled only when both vertical binary tree partitioning and vertical ternary tree partitioning are possible, or only when both horizontal binary tree partitioning and horizontal ternary tree partitioning are possible. Otherwise, partition tree information may not be signaled, but may be inferred as the value of the tree indicating the partitions that can be applied to the CU.
[0286] Figure 4 This is a diagram illustrating the form of prediction units that may be included in a coding unit.
[0287] When a CU that has not been further partitioned from the LCU can be divided into one or more prediction units (PUs), this partitioning is also called "partitioning".
[0288] A prediction unit (PU) can be the basic unit used for prediction. A PU can be encoded and decoded in any of the skip mode, inter-frame mode, and intra-frame mode. A PU can be partitioned into various shapes according to each mode. For example, see the reference above. Figure 1 The target block described and the above references Figure 2 The target blocks described can each be PUs.
[0289] A CU may not be divided into PUs. When a CU is not divided into PUs, the dimensions of the CU and the PU can be equal.
[0290] In skip mode, partitions may not exist in the CU. In skip mode, a 2N×2N mode 410 can be supported without partitions, where the PU and CU have the same size.
[0291] In inter-frame mode, eight types of partition shapes can exist in the CU. For example, in inter-frame mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440 and nR×2N mode 445 are supported.
[0292] In intra-frame mode, 2N×2N mode 410 and N×N mode 425 are supported.
[0293] In 2N×2N mode 410, PUs with 2N×2N dimensions can be encoded. A PU with 2N×2N dimensions can mean a PU with the same dimensions as a CU. For example, a PU with 2N×2N dimensions can have dimensions of 64×64, 32×32, 16×16, or 8×8.
[0294] In N×N mode 425, PUs with N×N dimensions can be encoded.
[0295] For example, in intra-frame prediction, when the PU size is 8×8, four partitioned PUs can be encoded. The size of the PU in each partition can be 4×4.
[0296] When encoding a PU in intra-frame mode, any of several intra-frame prediction modes can be used to encode the PU. For example, HEVC technology provides 35 intra-frame prediction modes, and the PU can be encoded in any of these 35 intra-frame prediction modes.
[0297] Based on the rate-distortion cost, it can be determined which of the 2N×2N modes 410 and N×N modes 425 will be used to encode the PU.
[0298] Encoding device 100 can perform encoding operations on a PU of size 2N×2N. Here, the encoding operation can be an operation of encoding the PU according to each of a plurality of intra-prediction modes that can be used by encoding device 100. Through the encoding operation, the optimal intra-prediction mode for the PU of size 2N×2N can be derived. The optimal intra-prediction mode can be the intra-prediction mode that incurs the minimum rate-distortion cost when encoding the PU of size 2N×2N among the plurality of intra-prediction modes that can be used by encoding device 100.
[0299] Furthermore, the encoding device 100 can sequentially perform encoding operations on each PU obtained from the N×N partition. Here, the encoding operation can be an operation of encoding the PU according to each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra-prediction mode for the N×N PU can be derived. The optimal intra-prediction mode can be the intra-prediction mode that incurs the minimum rate-distortion cost when encoding the N×N PU among the plurality of intra-prediction modes that can be used by the encoding device 100.
[0300] The encoding device 100 can determine which of the two PUs, one with a size of 2N×2N and the other with a size of N×N, will be encoded based on a comparison of the rate-distortion cost of the PU with a size of 2N×2N and the rate-distortion cost of the PU with a size of N×N.
[0301] A single CU can be partitioned into one or more PUs, and a PU can be partitioned into multiple PUs.
[0302] For example, when a single PU is partitioned into four PUs, the horizontal and vertical dimensions of each of the four PUs generated by the partitioning can be half the horizontal and vertical dimensions of the original PU. When a 32×32 PU is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, the PU can be considered to have been partitioned in a quadtree structure.
[0303] For example, when a single PU is partitioned into two PUs, the horizontal or vertical dimension of each of the two PUs created by the partitioning can be half the horizontal or vertical dimension of the original PU. When a 32×32 PU is vertically partitioned into two PUs, the size of each of the two partitioned PUs can be 16×32. When a 32×32 PU is horizontally partitioned into two PUs, the size of each of the two partitioned PUs can be 32×16. When a single PU is partitioned into two PUs, the PU can be considered to have been partitioned in a binary tree structure.
[0304] Figure 5 This is a diagram illustrating the form of a transformation unit that may be included in an encoding unit.
[0305] A transform unit (TU) may have basic units for processes such as transform, quantization, inverse transform, dequantization, entropy coding, and entropy decoding in a CU.
[0306] The TU can be square or rectangular. The shape of the TU can be determined based on the size and / or shape of the CU.
[0307] Within a CU partitioned from an LCU, CUs that are not further partitioned into CUs can be divided into one or more TUs. Here, the partitioning structure of a TU can be a quadtree structure. For example, as... Figure 5 As shown, a single CU 510 can be partitioned once or multiple times according to a quadtree structure. With the help of this partitioning, a single CU 510 can be composed of TUs of various sizes.
[0308] It can be considered that a CU is recursively divided when a single CU is divided two or more times. Through division, a single CU can be composed of transformation units (TUs) of various sizes.
[0309] Optionally, a single CU may be divided into one or more TUs based on the number of vertical and / or horizontal lines that divide the CU.
[0310] The CU can be divided into symmetrical TUs or asymmetrical TUs. For asymmetrical TU division, information about the size and / or shape of each TU can be transmitted from the encoding device 100 to the decoding device 200 via signals. Optionally, the size and / or shape of each TU can be derived from the information about the size and / or shape of the CU.
[0311] A CU may not be divided into TUs. When a CU is not divided into TUs, the size of the CU and the size of the TU can be equal.
[0312] A single CU can be partitioned into one or more TUs, and a TU can be partitioned into multiple TUs.
[0313] For example, when a single TU is partitioned into four TUs, the horizontal and vertical dimensions of each of the four TUs generated by the partitioning can be half the horizontal and vertical dimensions of the original TU. When a TU of size 32×32 is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, the TU can be considered to have been partitioned in a quadtree structure.
[0314] For example, when a single TU is partitioned into two TUs, the horizontal or vertical dimension of each of the two TUs created by the partitioning can be half the horizontal or vertical dimension of the original TU. When a TU of size 32×32 is vertically partitioned into two TUs, the size of each of the two partitioned TUs can be 16×32. When a TU of size 32×32 is horizontally partitioned into two TUs, the size of each of the two partitioned TUs can be 32×16. When a single TU is partitioned into two TUs, the TU can be considered to have been partitioned in a binary tree structure.
[0315] Can be according to different Figure 5The CU is segmented as shown.
[0316] For example, a single CU can be divided into three CUs. The horizontal or vertical dimensions of the three CUs generated from the division can be 1 / 4, 1 / 2, and 1 / 4 of the horizontal or vertical dimensions of the original CU before it was divided.
[0317] For example, when a 32×32 CU is vertically divided into three CUs, the sizes of the three CUs generated from the division can be 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is divided into three CUs, the CUs can be considered to be divided in the form of a ternary tree.
[0318] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) can be applied to the partitioning of the CU, and multiple partitioning schemes can be combined and used together for partitioning the CU. Here, the combination and use of multiple partitioning schemes together can be referred to as "complex tree format partitioning".
[0319] Figure 6 This shows the segmentation of the block based on the example.
[0320] In video encoding and / or decoding processes, target blocks can be segmented, such as... Figure 6 As shown. For example, the target block could be a CU.
[0321] For the segmentation of a target block, an indicator indicating the segmentation information can be transmitted from the encoding device 100 to the decoding device 200 via a signal. The segmentation information can be information indicating how the target block is segmented.
[0322] The segmentation information can be one or more of the following: a split flag (hereinafter referred to as "split_flag"), a quad flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag"), and a binary type flag (hereinafter referred to as "Btype_flag").
[0323] The "split_flag" can be a flag indicating whether a block is split. For example, a split_flag value of 1 indicates that the corresponding block is split, while a split_flag value of 0 indicates that the corresponding block is not split.
[0324] The `QB_flag` can be a flag indicating whether the block is segmented in quadtree or binary tree form. For example, a `QB_flag` value of 0 indicates that the block is segmented in quadtree form, and a `QB_flag` value of 1 indicates that the block is segmented in binary tree form. Alternatively, a `QB_flag` value of 0 indicates that the block is segmented in binary tree form, and a `QB_flag` value of 1 indicates that the block is segmented in quadtree form.
[0325] The "quadtree_flag" can be a flag indicating whether a block is split in a quadtree format. For example, a quadtree_flag value of 1 indicates that the block is split in a quadtree format, while a quadtree_flag value of 0 indicates that the block is not split in a quadtree format.
[0326] The "binarytree_flag" can be a flag indicating whether a block is split in a binary tree format. For example, a binarytree_flag value of 1 indicates that the block is split in a binary tree format, while a binarytree_flag value of 0 indicates that the block is not split in a binary tree format.
[0327] The `Btype_flag` can be a flag indicating which of the vertical or horizontal partitions corresponds to the partitioning direction when a block is divided in a binary tree format. For example, a `Btype_flag` value of 0 indicates that the block is partitioned horizontally, and a `Btype_flag` value of 1 indicates that the block is partitioned vertically. Alternatively, a `Btype_flag` value of 0 indicates that the block is partitioned vertically, and a `Btype_flag` value of 1 indicates that the block is partitioned horizontally.
[0328] For example, as shown in Table 1 below, it can be derived by transmitting at least one of quadtree_flag, binarytree_flag, and Btype_flag via signals. Figure 6 The segmentation information of the blocks in the text.
[0329] [Table 1]
[0330] For example, as shown in Table 2 below, it can be derived by transmitting at least one of split_flag, QB_flag, and Btype_flag using signals. Figure 6 The segmentation information of the blocks in the text.
[0331] [Table 2]
[0332] Depending on the size and / or shape of the block, the segmentation method may be limited to quadtrees or binary trees. When this limitation is applied, the split_flag can be a flag indicating whether the block is segmented in a quadtree or a binary tree format. The size and shape of the block can be deduced from the block's depth information, and the depth information can be transmitted from the encoding device 100 to the decoding device 200 via a signal.
[0333] When the block size falls within a certain range, it may be possible to partition using only a quadtree. For example, the certain range may be defined by at least one of the maximum block size and the minimum block size, where partitioning using only a quadtree is possible under both the maximum and minimum block sizes.
[0334] Information indicating the maximum and minimum block sizes, which can only be divided into quadtree forms, can be transmitted via a bitstream from encoding device 100 to decoding device 200. Furthermore, this information can be transmitted via signal for at least one of units such as video, sequence, frame, parameter, parallel block group, and strip (or segment).
[0335] Optionally, the maximum block size and / or minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the block size is greater than 64×64 and less than 256×256, it may be possible to split only in the form of a quadtree. In this case, split_flag can be a flag indicating whether to perform splitting in the form of a quadtree.
[0336] When the size of a block is larger than the maximum size of a transform block, it is possible to partition it using only a quadtree. Here, the sub-blocks generated by the partition can be at least one of CU and TU.
[0337] In this case, split_flag can be a flag indicating whether the CU is partitioned in the form of a quadtree.
[0338] When the size of a block falls within a certain range, it may be possible to partition it using only a binary tree or a ternary tree. For example, the certain range may be defined by at least one of the maximum block size and the minimum block size, where partitioning using only a binary tree or a ternary tree is possible at both the maximum and minimum block sizes.
[0339] Information indicating the maximum and / or minimum block size, which can only be segmented in a binary tree or ternary tree format, can be transmitted via a bitstream from encoding device 100 to decoding device 200. Furthermore, this information can be transmitted via signal for at least one of units such as sequences, frames, and stripes (or segments).
[0340] Optionally, the maximum block size and / or minimum block size can be fixed sizes predefined by the encoding device 100 and the decoding device 200. For example, when the block size is greater than 8×8 and less than 16×16, it may be possible to split the block only in a binary tree form. In this case, the split_flag can be a flag indicating whether to perform splitting in a binary tree or ternary tree form.
[0341] The above description of partitioning in quadtree form can also be applied to binary tree form and / or ternary tree form.
[0342] The partitioning of a block can be restricted by previous partitioning. For example, when a block is partitioned in a specific binary tree form, and then multiple sub-blocks are generated from the partitions, each sub-block can be further partitioned only in the specific tree form. Here, the specific tree form can be at least one of binary tree, ternary tree, and quadtree forms.
[0343] When the horizontal or vertical dimensions of a partition block are dimensions that cannot be further subdivided, the aforementioned indicator may not be transmitted by signal.
[0344] Figure 7 This is a diagram illustrating an embodiment used to explain the intra-frame prediction process.
[0345] Figure 7 The arrows extending radially from the center of the attached figure indicate the prediction direction of the intra-prediction mode. Furthermore, the numbers appearing near the arrows indicate examples of mode values assigned to the intra-prediction mode or its prediction direction.
[0346] exist Figure 7 In this context, number 0 can represent the Planar mode as a non-directional intra-prediction mode. Number 1 can represent the DC mode as a non-directional intra-prediction mode.
[0347] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the target block. A neighboring block can be a reconstructed neighboring block. A reference sample can refer to a neighboring sample.
[0348] For example, intra-frame coding and / or decoding can be performed using the values of reference samples included in the reconstructed neighboring blocks or the coding parameters of the reconstructed neighboring blocks.
[0349] Encoding device 100 and / or decoding device 200 can generate a prediction block by performing intra-frame prediction on the target block based on information about samples in the target image. When performing intra-frame prediction, encoding device 100 and / or decoding device 200 can generate a prediction block for the target block by performing intra-frame prediction based on information about samples in the target image. When performing intra-frame prediction, encoding device 100 and / or decoding device 200 can perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.
[0350] A prediction block can be a block generated as a result of performing intra-frame prediction. A prediction block can correspond to at least one of CU, PU, and TU.
[0351] The cells of the prediction block may have a size corresponding to at least one of CU, PU, and TU. The prediction block may have a square shape with a size of 2N×2N or N×N. The N×N size may include sizes such as 4×4, 8×8, 16×16, 32×32, 64×64, etc.
[0352] Optionally, the prediction block can be a square block with a size of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc., or a rectangular block with a size of 2×8, 4×8, 2×16, 4×16, 8×16, etc.
[0353] Intra-prediction can be performed using intra-prediction modes for the target block. The number of intra-prediction modes that a target block can have can be a predefined fixed value, or it can be a value determined differently based on the attributes of the prediction block. For example, the attributes of the prediction block can include the size of the prediction block, the type of the prediction block, etc. In addition, the attributes of the prediction block can indicate the coding parameters for the prediction block.
[0354] For example, the number of intra-prediction modes can be fixed at N, regardless of the size of the prediction block. Alternatively, the number of intra-prediction modes can be, for example, 3, 5, 9, 17, 34, 35, 36, 65, 67, or 95.
[0355] Intra-frame prediction mode can be either non-directional or directional.
[0356] For example, intra-frame prediction modes may include... Figure 7 The numbers 0 to 66 shown correspond to two non-directional patterns and 65 directional patterns.
[0357] For example, when using a specific intra-prediction method, the intra-prediction mode may include... Figure 7 The numbers -14 to 80 shown correspond to the two non-directional patterns and 93 directional patterns.
[0358] The two non-directional modes can include DC mode and planar mode.
[0359] Directional patterns can be prediction patterns with a specific direction or angle. Directional patterns can also be called "angle patterns".
[0360] An intra-prediction mode can be represented by at least one of a mode number, a mode value, a mode angle, and a mode direction. In other words, the terms "(mode) number of intra-prediction mode", "(mode) value of intra-prediction mode", "(mode) angle of intra-prediction mode" and "(mode) direction of intra-prediction mode" can be used to have the same meaning and can be used interchangeably with each other.
[0361] The number of intra-prediction modes can be M. The value of M can be 1 or greater. In other words, the number of intra-prediction modes can be M, which includes the number of non-directional modes and the number of directional modes.
[0362] The number of intra-prediction modes can be fixed at M, regardless of the block size and / or color components. For example, the number of intra-prediction modes can be fixed at either 35 or 67, regardless of the block size.
[0363] Optionally, the number of intra-frame prediction modes may vary depending on the shape, size, and / or type of color components of the block.
[0364] For example, in Figure 7 In the diagram, the direction prediction mode shown by the dashed line can only be applied to the prediction of non-square blocks.
[0365] For example, the larger the block size, the greater the number of intra-prediction modes. Alternatively, the larger the block size, the smaller the number of intra-prediction modes. When the block size is 4×4 or 8×8, the number of intra-prediction modes can be 67. When the block size is 16×16, the number of intra-prediction modes can be 35. When the block size is 32×32, the number of intra-prediction modes can be 19. When the block size is 64×64, the number of intra-prediction modes can be 7.
[0366] For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chrominance signal. Optionally, the number of intra-prediction modes corresponding to the luma component block can be greater than the number of intra-prediction modes corresponding to the chrominance component block.
[0367] For example, in the vertical mode with a mode value of 50, prediction can be performed in the vertical direction based on the pixel values of the reference sample. Similarly, in the horizontal mode with a mode value of 18, prediction can be performed in the horizontal direction based on the pixel values of the reference sample.
[0368] Even in directional modes other than those described above, the encoding device 100 and the decoding device 200 can perform intra-frame prediction on the target cell using reference samples based on the angle corresponding to the directional mode.
[0369] Intra-prediction modes located to the right of the vertical mode can be called "vertical-right mode." Intra-prediction modes located below the horizontal mode can be called "horizontal-bottom mode." For example, in... Figure 7 In the frame prediction mode, a mode value of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 can be a vertical-right mode. An intra-frame prediction mode with a mode value of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be a horizontal-downward mode.
[0370] Non-directional modes can include DC mode and planar mode. For example, the value for DC mode can be 1, and the value for planar mode can be 0.
[0371] Orientation modes can include angle modes. Among multiple intra-frame prediction modes, the remaining modes other than DC mode and planar mode can be orientation modes.
[0372] When the intra-frame prediction mode is DC mode, a prediction block can be generated based on the average pixel values of multiple reference pixels. For example, the pixel values of the prediction block can be determined based on the average pixel values of multiple reference pixels.
[0373] The number of intra-prediction modes and the mode values of each intra-prediction mode described above are merely exemplary. The number of intra-prediction modes and the mode values of each intra-prediction mode may be defined differently depending on the embodiment, implementation method, and / or requirements.
[0374] To perform intra-frame prediction on a target block, a step can be performed to check whether samples included in the reconstructed neighboring blocks can be used as reference samples for the target block. When there are samples among the samples in the neighboring blocks that cannot be used as reference samples for the target block, the sample value of the sample that cannot be used as a reference sample can be replaced by a value generated by copying and / or interpolating at least one sample value included in the reconstructed neighboring blocks. When the sample value of an existing sample is replaced by a value generated by copying and / or interpolating, the sample can be used as a reference sample for the target block.
[0375] When using intra-frame prediction, a filter can be applied to at least one of the reference sample and the prediction sample based on at least one of the intra-frame prediction mode and the size of the target block.
[0376] The type of filter to be applied to at least one of the reference sample and the prediction sample can vary depending on at least one of the intra-prediction mode of the target block, the size of the target block, and the shape of the target block. The type of filter can be classified according to one or more of the length of the filter taps, the values of the filter coefficients, and the filter strength. The length of the filter taps can refer to the number of filter taps. Additionally, the number of filter taps can refer to the length of the filter.
[0377] When the intra-frame prediction mode is planar mode, when generating the prediction block of the target block, the sample value of the predicted target block can be generated by using the weighted sum of the upper reference sample, the left reference sample, the upper right reference sample, and the lower left reference sample of the target block, based on the position of the predicted target sample in the prediction block.
[0378] When the intra-frame prediction mode is DC mode, the average of the reference samples above and to the left of the target block can be used when generating the prediction block for the target block. Furthermore, filtering using the values of the reference samples can be performed on specific rows or columns within the target block. A specific row can be one or more rows above the reference sample. A specific column can be one or more columns to the left of the reference sample.
[0379] When the intra-frame prediction mode is directional mode, the prediction block can be generated using the top reference sample, left reference sample, upper right reference sample, and / or lower left reference sample of the target block.
[0380] To generate the above predicted samples, real-number-based interpolation can be performed.
[0381] The intra-prediction mode of the target block can be predicted from the intra-prediction modes of neighboring blocks adjacent to the target block, and the information used for prediction can be entropy encoded / decoded.
[0382] For example, when the intra prediction modes of the target block and neighboring blocks are the same, a predefined flag can be used to signal that the intra prediction modes of the target block and neighboring blocks are the same.
[0383] For example, an indicator can be used to signal an intra prediction mode that is the same as the intra prediction mode of the target block among the intra prediction modes of multiple neighboring blocks.
[0384] When the intra prediction modes of the target block and neighboring blocks are different from each other, entropy coding and / or decoding can be used to encode and / or decode information about the intra prediction mode of the target block.
[0385] Figure 8 This is a diagram showing the reference samples used in the intra-frame prediction process.
[0386] Reconstruction reference samples used for intra-frame prediction of the target block may include lower left reference sample, left reference sample, upper left reference sample, upper reference sample, and upper right reference sample.
[0387] For example, a left reference sample point can refer to the reconstructed reference pixel adjacent to the left side of the target block. A top reference sample point can refer to the reconstructed reference pixel adjacent to the top of the target block. A top-left reference sample point can refer to the reconstructed reference pixel located at the top-left corner of the target block. A bottom-left reference sample point can refer to a reference sample point located below the left reference sample point line formed by the left reference sample points. A top-right reference sample point can refer to a reference sample point to the right of the top reference sample point line formed by the top reference sample points.
[0388] When the size of the target block is N×N, the number of reference points at the lower left, left, top, and upper right can each be N.
[0389] A prediction block can be generated by performing intra-frame prediction on the target block. Generating a prediction block may include determining the values of the pixels in the prediction block. The target block and the prediction block may be the same size.
[0390] The reference sample used for intra-prediction of the target block can vary depending on the intra-prediction mode of the target block. The direction of the intra-prediction mode can represent the dependency between the reference sample and the pixels of the prediction block. For example, the value of a specified reference sample can be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and one or more specified pixels in the prediction block can be samples and pixels on a straight line located in the direction of the intra-prediction mode. In other words, the value of the specified reference sample can be copied as the value of a pixel located in the direction opposite to the direction of the intra-prediction mode. Optionally, the value of a pixel in the prediction block can be the value of a reference sample located in the direction of the intra-prediction mode relative to the position of that pixel.
[0391] In the example, when the intra-prediction mode of the target block is vertical, the upper reference sample can be used for intra-prediction. When the intra-prediction mode is vertical, the value of a pixel in the prediction block can be the value of a reference sample vertically above that pixel. Therefore, the upper reference sample adjacent to the top of the target block can be used for intra-prediction. Furthermore, the value of a pixel in a row of the prediction block can be the same as the value of the aforementioned reference sample.
[0392] In the example, when the intra-prediction mode of the target block is horizontal, the left reference sample can be used for intra-prediction. When the intra-prediction mode is horizontal, the value of a pixel in the prediction block can be the value of a reference sample horizontally to the left of that pixel's position. Therefore, the left reference sample adjacent to the left side of the target block can be used for intra-prediction. Furthermore, the value of a pixel in a column of the prediction block can be the same as the value of the left reference sample.
[0393] In the example, when the mode value of the intra-prediction mode for the current block is 34, at least some of the left-hand reference samples, the top-left reference sample, and the top reference sample can be used for intra-prediction. When the mode value of the intra-prediction mode is 34, the value of a pixel in the prediction block can be the value of the reference sample located diagonally to the top-left corner of that pixel.
[0394] Furthermore, in the case of an intra-prediction mode with mode values in the range of 52 to 66, at least a portion of the upper right reference sample can be used for intra-prediction.
[0395] Furthermore, in the case of intra-prediction modes with mode values in the range of 2 to 17, at least a portion of the lower left reference samples can be used for intra-prediction.
[0396] Furthermore, in the case of intra-prediction modes with mode values in the range of 19 to 49, the upper left reference sample can be used for intra-prediction.
[0397] The number of reference samples used to determine the pixel value of a pixel in the prediction block can be 1, 2 or more.
[0398] As described above, the pixel value of a pixel in a prediction block can be determined based on the pixel's position and the position of a reference sample indicated by the direction of the intra-prediction mode. When both the pixel's position and the position of the reference sample indicated by the direction of the intra-prediction mode are integer positions, the value of a reference sample indicated by the integer position can be used to determine the pixel value of the pixel in the prediction block.
[0399] When the position of a pixel and the position of a reference sample indicated by the direction of the intra-prediction mode are not integer positions, an interpolated reference sample can be generated based on the positions of the two reference samples closest to the reference sample. The value of the interpolated reference sample can be used to determine the pixel value of a pixel in the prediction block. In other words, when the position of a pixel in the prediction block and the position of a reference sample indicated by the direction of the intra-prediction mode indicate the position between two reference samples, an interpolation based on the values of the two samples can be generated.
[0400] The predicted block generated by prediction may be different from the original target block. In other words, there may be prediction error as the difference between the target block and the predicted block, and there may also be prediction error between the pixels of the target block and the pixels of the predicted block.
[0401] In the following text, the terms “difference,” “error,” and “residual” may be used to mean the same thing and may be used interchangeably.
[0402] For example, in the case of intra-frame prediction, the greater the distance between the pixels of the predicted block and the reference sample, the greater the possible prediction error. Such prediction errors can lead to discontinuities between the generated predicted block and its neighboring blocks.
[0403] To reduce prediction error, filtering can be used for prediction blocks. The filtering can be configured to adaptively apply filters to regions within the prediction block that are considered to have large prediction errors. For example, regions considered to have large prediction errors could be the boundaries of the prediction block. Furthermore, the regions within the prediction block considered to have large prediction errors can vary depending on the intra-prediction mode, and the characteristics of the filter can also vary depending on the intra-prediction mode.
[0404] like Figure 8 As shown, for intra-frame prediction of the target block, at least one of reference lines 0 to 3 can be used.
[0405] Figure 8 Each reference line in the diagram can indicate a reference point line that includes one or more reference points. The smaller the reference line number, the closer the line indicating the reference points is to the target block.
[0406] Samples in segments A and F can be obtained by filling in the samples in segments B and E that are closest to the target block, rather than from the reconstructed neighboring blocks.
[0407] The index information of the reference sample lines to be used for intra-prediction of the target block can be transmitted using signals. The index information can indicate the reference sample lines among a plurality of reference sample lines to be used for intra-prediction of the target block. For example, the index information can have a value corresponding to any one of 0 to 3.
[0408] When the top boundary of the target block is the boundary of the CTU, only reference sample line 0 can be available. Therefore, in this case, index information can be omitted from signal transmission. When additional reference sample lines besides reference sample line 0 are used, the filtering of the prediction block, which will be described later, can be omitted.
[0409] In the case of intra-frame prediction between colors, a prediction block for the target block of the second color component can be generated based on the corresponding reconstructed block of the first color component.
[0410] For example, the first color component can be the luminance component, and the second color component can be the chromaticity component.
[0411] To perform inter-color intra-frame prediction, parameters for a linear model between the first and second color components can be derived based on a template.
[0412] The template may include reference points above the target block (top reference points) and / or reference points to the left of the target block (left reference points), and may include top reference points and / or left reference points corresponding to the reference points of the reconstructed block of the first color component.
[0413] For example, the following items can be used to derive the parameters for a linear model: 1) the value of the sample with the maximum value of the first color component among the samples in the template, 2) the value of the sample of the second color component corresponding to the sample of the first color component, 3) the value of the sample with the minimum value of the first color component among the samples in the template, and 4) the value of the sample of the second color component corresponding to the sample of the first color component.
[0414] When the parameters for the linear model are derived, a prediction block for the target block can be generated by applying the corresponding reconstructed block to the linear model.
[0415] Depending on the image format, subsampling can be performed on samples adjacent to the reconstructed block of the first color component and the corresponding reconstructed block of the first color component. For example, when one sample of the second color component corresponds to four samples of the first color component, a corresponding sample can be calculated by subsampling the four samples of the first color component. When subsampling is performed, the derivation of parameters for the linear model and intra-frame prediction between colors can be performed based on the corresponding samples obtained from the subsampling.
[0416] Information about whether to perform inter-color intra-frame prediction and / or the range of templates can be transmitted via signals in intra-frame prediction mode.
[0417] The target block can be divided into two or four sub-blocks in the horizontal and / or vertical directions.
[0418] Sub-blocks generated by partitioning can be reconstructed sequentially. That is, when intra-prediction is performed on each sub-block, a sub-prediction block for that sub-block can be generated. Furthermore, when inverse quantization and / or inverse transform is performed on each sub-block, a sub-residual block for the corresponding sub-block can be generated. Reconstructed sub-blocks can be generated by adding the sub-prediction blocks to the sub-residual blocks. The reconstructed sub-blocks can be used as reference samples for intra-prediction of sub-blocks with the next higher priority.
[0419] A sub-block can be a block containing a specific number (e.g., 16) or more samples. For example, when the target block is an 8×4 block or a 4×8 block, the target block can be divided into two sub-blocks. Furthermore, when the target block is a 4×4 block, it cannot be divided into sub-blocks. When the target block has another size, it can be divided into four sub-blocks.
[0420] Information about whether to perform intra-frame prediction based on such sub-blocks and / or about the partitioning direction (horizontal or vertical) can be transmitted using signals.
[0421] Such sub-block-based intra-prediction can be restricted so that it is performed only when reference sample line 0 is used. When performing sub-block-based intra-prediction, filtering of the prediction block, which will be described below, may be omitted.
[0422] The final prediction block can be generated by filtering the prediction block generated via intra-frame prediction.
[0423] Filtering can be performed by applying specific weights to the target sample, the left reference sample, the top reference sample, and / or the top-left reference sample, where the target sample is the target to be filtered.
[0424] The weights and / or reference samples used for filtering (e.g., the range of reference samples, the location of reference samples, etc.) can be determined based on at least one of the block size, intra-prediction mode, and the location of the target filter sample in the prediction block.
[0425] For example, filtering can be performed only in specific intra-frame prediction modes (e.g., DC mode, planar mode, vertical mode, horizontal mode, diagonal mode, and / or adjacent diagonal mode).
[0426] Adjacent diagonal patterns can be patterns with numbers obtained by adding k to the diagonal pattern's number, or patterns with numbers obtained by subtracting k from the diagonal pattern's number. In other words, the number of an adjacent diagonal pattern can be the sum of the diagonal pattern's number and k, or it can be the difference between the diagonal pattern's number and k. For example, k can be a positive integer of 8 or less.
[0427] The intra prediction mode of the target block can be derived using the intra prediction modes of neighboring blocks that exist near the target block, and the derived intra prediction mode can be entropy encoded and / or entropy decoded.
[0428] For example, when the intra prediction mode of the target block is the same as that of the neighboring blocks, specific flag information can be used to signal and transmit information indicating that the intra prediction mode of the target block is the same as that of the neighboring blocks.
[0429] Additionally, for example, indicator information can be transmitted for neighboring blocks in an intra-prediction mode that is the same as the intra-prediction mode of the target block in an intra-prediction mode with multiple neighboring blocks.
[0430] For example, when the intra prediction mode of the target block is different from that of the neighboring blocks, entropy coding and / or entropy decoding can be performed on the information about the intra prediction mode of the target block by performing entropy coding and / or entropy decoding based on the intra prediction modes of the neighboring blocks.
[0431] Figure 9 This is a diagram illustrating an embodiment used to explain the inter-frame prediction process.
[0432] Figure 9 The rectangles shown can represent images (or screens). Furthermore, in Figure 9 In the image, arrows indicate the prediction direction. An arrow pointing from the first frame to the second frame means that the second frame references the first frame. In other words, each image can be encoded and / or decoded based on the prediction direction.
[0433] Based on the encoding type, images can be classified into intra-frame frames (I-frames), one-way predictive frames or predictive-coded frames (P-frames), and two-way predictive frames or two-way predictive-coded frames (B-frames). Each frame can be encoded and / or decoded according to its encoding type.
[0434] When the target image to be encoded is an I-frame, the target image can be encoded using data contained within the image itself, without inter-frame prediction referencing other images. For example, an I-frame can be encoded solely via intra-frame prediction.
[0435] When the target image is a P-frame, it can be encoded via inter-frame prediction using a reference frame present in one direction. Here, one direction can be a forward direction or a backward direction.
[0436] When the target image is a B-frame, it can be encoded via inter-frame prediction using reference frames present in both directions, or via inter-frame prediction using reference frames present in one of the forward and backward directions. Here, the two directions can be the forward and backward directions.
[0437] P-frames and B-frames encoded and / or decoded using reference frames can be considered as images using inter-frame prediction.
[0438] The following will describe in detail the inter-frame prediction in inter-frame mode according to the embodiments.
[0439] Reference images and motion information can be used to perform inter-frame prediction or motion compensation.
[0440] In inter-frame mode, encoding device 100 may perform inter-frame prediction and / or motion compensation on the target block. Decoding device 200 may perform inter-frame prediction and / or motion compensation on the target block corresponding to the inter-frame prediction and / or motion compensation performed by encoding device 100.
[0441] During inter-frame prediction, motion information of the target block can be derived independently by the encoding device 100 and the decoding device 200. This motion information can be derived using motion information from reconstructed neighboring blocks, motion information from the col block, and / or motion information from blocks adjacent to the col block.
[0442] For example, encoding device 100 or decoding device 200 can perform prediction and / or motion compensation by using motion information of spatial candidates and / or temporal candidates as motion information of the target block. The target block may refer to a PU and / or a PU partition.
[0443] Spatial candidates can be reconstructed blocks that are spatially adjacent to the target block.
[0444] The time candidate can be a reconstructed block corresponding to the target block in a previously reconstructed co-location frame (col frame).
[0445] In inter-frame prediction, the encoding device 100 and the decoding device 200 can improve encoding efficiency and decoding efficiency by utilizing motion information from spatial candidates and / or temporal candidates. The motion information from spatial candidates can be referred to as "spatial motion information." The motion information from temporal candidates can be referred to as "temporal motion information."
[0446] Below, the motion information of spatial candidates can be the motion information of PUs including spatial candidates. The motion information of temporal candidates can be the motion information of PUs including temporal candidates. The motion information of candidate blocks can be the motion information of PUs including candidate blocks.
[0447] Reference frames can be used to perform inter-frame prediction.
[0448] The reference image can be at least one of the images preceding and following the target image. The reference image can be an image used for predicting the target block.
[0449] In inter-frame prediction, regions within a reference frame can be specified using a reference frame index (or refIdx) used to indicate a reference frame, motion vectors (described later), etc. Here, the region specified in the reference frame can indicate a reference block.
[0450] Inter-frame prediction can select a reference frame, and can also select a reference block corresponding to the target block from the reference frame. In addition, inter-frame prediction can use the selected reference block to generate a prediction block for the target block.
[0451] Motion information can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction.
[0452] Spatial candidates can be blocks that satisfy the following conditions: 1) exist in the target image, 2) have been previously reconstructed via encoding and / or decoding, and 3) are adjacent to the target block or located at a corner of the target block. Here, "a block located at a corner of the target block" can be a block that is vertically adjacent to a neighboring block that is horizontally adjacent to the target block, or a block that is horizontally adjacent to a neighboring block that is vertically adjacent to the target block. Furthermore, "a block located at a corner of the target block" can have the same meaning as "a block adjacent to a corner of the target block." The meaning of "a block located at a corner of the target block" can be included within the meaning of "a block adjacent to the target block."
[0453] For example, a spatial candidate can be a reconstruction block located to the left of the target block, a reconstruction block located above the target block, a reconstruction block located at the lower left corner of the target block, a reconstruction block located at the upper right corner of the target block, or a reconstruction block located at the upper left corner of the target block.
[0454] Each of the encoding device 100 and the decoding device 200 can identify a block that exists at a spatially corresponding position in the col frame. The position of the target block in the target frame and the position of the identified block in the col frame can correspond to each other.
[0455] Each of the encoding device 100 and the decoding device 200 can identify a col block existing at a predefined relative position for the identified block as a time candidate. This predefined relative position can be a position existing inside and / or outside the identified block.
[0456] For example, a col block can include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block can be the block located at coordinates (xP + nPSW, yP + nPSH). The second col block can be the block located at coordinates (xP + (nPSW>>1), yP + (nPSH>>1)). When the first col block is unavailable, the second col block can be used selectively.
[0457] The motion vector of the target block can be determined based on the motion vector of the col block. Each of the encoding device 100 and the decoding device 200 can scale the motion vector of the col block. The scaled motion vector of the col block can be used as the motion vector of the target block. Furthermore, the motion vectors for motion information of time candidates stored in a list can be scaled motion vectors.
[0458] The ratio of the motion vector of the target block to the motion vector of the col block can be the same as the ratio of the first time distance to the second time distance. The first time distance can be the distance between the reference frame and the target frame of the target block. The second time distance can be the distance between the reference frame and the col frame of the col block.
[0459] The scheme used to derive motion information can be changed depending on the inter-frame prediction mode of the target block. For example, inter-frame prediction modes applied to inter-frame prediction may include Advanced Motion Vector Prediction (AMVP) mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, triangle partitioning mode, inter-frame / intra-frame combined prediction mode, affine inter-frame mode, and current frame reference mode. The merge mode can also be called "motion merge mode." The following provides a detailed explanation of each mode.
[0460] 1) AMVP mode When using AMVP mode, the encoding device 100 can search for similar blocks in the neighborhood of the target block. The encoding device 100 can obtain a predicted block by performing a prediction on the target block using the motion information of the found similar blocks. The encoding device 100 can encode a residual block, where the residual block is the difference between the target block and the predicted block.
[0461] 1-1) Creation of the list of candidate motion vectors for prediction When the AMVP mode is used as the prediction mode, each of the encoding device 100 and the decoding device 200 can create a prediction motion vector candidate list using spatial candidate motion vectors, temporal candidate motion vectors, and zero vectors. The prediction motion vector candidate list may include one or more prediction motion vector candidates. At least one of the spatial candidate motion vectors, temporal candidate motion vectors, and zero vectors can be determined and used as a prediction motion vector candidate.
[0462] In the following text, the terms “predicted motion vector (candidate)” and “motion vector (candidate)” may be used as having the same meaning and may be used interchangeably with each other.
[0463] In the following text, the terms “predicted motion vector candidate” and “AMVP candidate” may be used as having the same meaning and may be used interchangeably.
[0464] In the following text, the terms “predicted motion vector candidate list” and “AMVP candidate list” may be used as having the same meaning and may be used interchangeably.
[0465] Spatial candidates can include reconstructed spatial neighbor blocks. In other words, the motion vectors of the reconstructed neighbor blocks can be referred to as "spatial prediction motion vector candidates".
[0466] A time candidate can include the col block and the blocks adjacent to the col block. In other words, the motion vector of the col block or the motion vector of the blocks adjacent to the col block can be called a "time prediction motion vector candidate".
[0467] The zero vector can be a (0, 0) motion vector.
[0468] The predicted motion vector candidate can be a motion vector predictor factor used to predict the motion vector. Furthermore, in the encoding device 100, each predicted motion vector candidate can be an initial search position for the motion vector.
[0469] 1-2) Searching for motion vectors using a list of predicted motion vector candidates Encoding device 100 can use a list of predicted motion vector candidates to determine the motion vector to be used for encoding the target block within a search range. Furthermore, encoding device 100 can determine, from among the predicted motion vector candidates existing in the list, a predicted motion vector candidate to be used as the target block.
[0470] The motion vector used to encode the target block can be a motion vector that can be encoded at minimal cost.
[0471] In addition, the encoding device 100 can determine whether to use the AMVP mode to encode the target block.
[0472] 1-3) Transmission of inter-frame prediction information Encoding device 100 can generate a bitstream that includes inter-frame prediction information required for inter-frame prediction. Decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.
[0473] Inter-frame prediction information may include: 1) mode information indicating whether AMVP mode is used, 2) predicted motion vector index, 3) motion vector difference (MVD), 4) reference direction, and 5) reference frame index.
[0474] In the following text, the terms “predicted motion vector index” and “AMVP index” may be used as having the same meaning and may be used interchangeably.
[0475] In addition, inter-frame prediction information may include residual signals.
[0476] When the mode information indicates that AMVP mode is used, the decoding device 200 can obtain the predicted motion vector index, MVD, reference direction and reference frame index from the bitstream through entropy decoding.
[0477] The predicted motion vector index indicates a predicted motion vector candidate that will be used for the prediction of the target block, which is included in the predicted motion vector candidate list.
[0478] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information The decoding device 200 can use the list of predicted motion vector candidates to derive predicted motion vector candidates, and can determine the motion information of the target block based on the derived predicted motion vector candidates.
[0479] The decoding device 200 can use the predicted motion vector index to determine a motion vector candidate for the target block from among the predicted motion vector candidates included in the predicted motion vector candidate list. The decoding device 200 can select the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate candidate list as the predicted motion vector for the target block.
[0480] Encoding device 100 can generate an entropy-coded predicted motion vector index by applying entropy coding to the predicted motion vector index, and can generate a bitstream including the entropy-coded predicted motion vector index. The entropy-coded predicted motion vector index can be signaled from encoding device 100 to decoding device 200 via the bitstream. Decoding device 200 can extract the entropy-coded predicted motion vector index from the bitstream, and can obtain the predicted motion vector index by applying entropy decoding to the entropy-coded predicted motion vector index.
[0481] The motion vectors used for inter-frame prediction in the target block may not match the predicted motion vectors. MVD (Motion Vector Difference) can be used to indicate the difference between the actual inter-frame prediction motion vectors used in the target block and the predicted motion vectors. The coding device 100 can derive predicted motion vectors that are similar to the inter-frame prediction motion vectors used in the target block in order to use the smallest possible MVD.
[0482] Motion Vector Difference (MVD) can be the difference between the motion vector of the target block and the predicted motion vector. Encoding device 100 can compute the MVD and generate an entropy-coded MVD by applying entropy coding to the MVD. Encoding device 100 can generate a bitstream including the entropy-coded MVD.
[0483] The MVD can be sent from the encoding device 100 to the decoding device 200 via a bitstream. The decoding device 200 can extract the entropy-encoded MVD from the bitstream and obtain the MVD by applying entropy decoding to the entropy-encoded MVD.
[0484] The decoding device 200 can derive the motion vector of the target block by summing the MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 can be the sum of the MVD and the motion vector candidate.
[0485] Furthermore, the encoding device 100 can generate entropy-coded MVD resolution information by applying entropy coding to the calculated MVD resolution information, and can generate a bitstream including the entropy-coded MVD resolution information. The decoding device 200 can extract the entropy-coded MVD resolution information from the bitstream, and can obtain the MVD resolution information by applying entropy decoding to the entropy-coded MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the MVD resolution.
[0486] Additionally, the encoding device 100 can calculate the MVD based on an affine model. The decoding device 200 can derive the affine control motion vector of the target block by summing the MVD with the affine control motion vector candidates, and can use the affine control motion vector to derive the motion vector of the sub-block.
[0487] The reference direction can indicate a list of reference frames that will be used for prediction of the target block. For example, the reference direction can indicate one of reference frame list L0 and reference frame list L1.
[0488] The reference direction indicates only the list of reference frames that will be used for prediction of the target block, and does not necessarily imply that the direction of the reference frames is limited to the forward or backward direction. In other words, each of the reference frame lists L0 and L1 can include frames in the forward and / or backward directions.
[0489] A unidirectional reference direction can indicate the use of a single reference screen list. A bidirectional reference direction can indicate the use of two reference screen lists. In other words, the reference direction can indicate one of the following: using only reference screen list L0, using only reference screen list L1, or using both reference screen lists.
[0490] The reference frame index indicates a reference frame in the reference frame list used for predicting the target block. Encoding device 100 can generate an entropy-coded reference frame index by applying entropy coding to the reference frame index, and can generate a bitstream including the entropy-coded reference frame index. The entropy-coded reference frame index can be signaled from encoding device 100 to decoding device 200 via the bitstream. Decoding device 200 can extract the entropy-coded reference frame index from the bitstream, and can obtain the reference frame index by applying entropy decoding to the entropy-coded reference frame index.
[0491] When using two lists of reference frames to predict a target block, a single reference frame index and a single motion vector can be used for each of the reference frame lists. Furthermore, when using two lists of reference frames to predict a target block, two prediction blocks can be specified for the target block. For example, the (final) prediction block for the target block can be generated using the average or weighted sum of the two prediction blocks for the target block.
[0492] The motion vector of the target block can be derived by predicting the motion vector index, MVD, reference direction, and reference screen index.
[0493] The decoding device 200 can generate a predicted block for a target block based on the derived motion vector and the reference frame index. For example, the predicted block can be a reference block in a reference frame indicated by the derived motion vector, which is indicated by the reference frame index.
[0494] Since the predicted motion vector index and MVD are encoded without encoding the motion vector of the target block itself, the number of bits sent from the encoding device 100 to the decoding device 200 can be reduced, and the encoding efficiency can be improved.
[0495] For the target block, motion information from reconstructed neighboring blocks can be used. In a specific inter-frame prediction mode, the encoding device 100 may not encode the actual motion information of the target block separately. Instead of encoding the motion information of the target block, it can encode additional information that allows the motion information of the target block to be derived using the motion information of reconstructed neighboring blocks. Encoding this additional information reduces the number of bits sent to the decoding device 200 and improves encoding efficiency.
[0496] For example, in inter-frame prediction modes where motion information of the target block is not directly encoded, skipping modes and / or merging modes may exist. Here, the motion information of each of the encoding device 100 and the decoding device 200 in the neighboring units that indicate reconstruction will be used as the identifier and / or index of the unit for the motion information of the target unit.
[0497] 2) Merge Mode As a scheme for deriving motion information of a target block, merging exists. The term "merging" can refer to the merging of the motion of multiple blocks. "Merging" can also mean that the motion information of one block is applied to other blocks. In other words, a merging pattern can be a pattern for deriving the motion information of a target block from the motion information of neighboring blocks.
[0498] When using the merging mode, the encoding device 100 can use motion information from spatial candidates and / or temporal candidates to predict motion information of the target block. Spatial candidates may include reconstructed spatially adjacent blocks that are spatially adjacent to the target block. Spatially adjacent blocks may include left-side adjacent blocks and top-side adjacent blocks. Temporal candidates may include col blocks. The terms "spatial candidate" and "spatial merging candidate" are used interchangeably and have the same meaning. The terms "temporal candidate" and "temporal merging candidate" are used interchangeably and have the same meaning.
[0499] The encoding device 100 can obtain a prediction block via prediction. The encoding device 100 can encode a residual block, wherein the residual block is the difference between the target block and the prediction block.
[0500] 2-1) Creation of the merged candidate list When using the merging mode, each of the encoding device 100 and the decoding device 200 can create a merging candidate list using motion information from spatial candidates and / or motion information from temporal candidates. The motion information may include 1) a motion vector, 2) a reference frame index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional. The reference direction may refer to an inter-frame prediction indicator.
[0501] The merge candidate list can include merge candidates. Merge candidates can be motion information. In other words, the merge candidate list can be a list that stores multiple pieces of motion information.
[0502] The merged candidate can be multiple motion information entries from temporal and / or spatial candidates. In other words, the merged candidate list can include motion information from temporal and / or spatial candidates, etc.
[0503] Furthermore, the merge candidate list may include new merge candidates generated by combining merge candidates that already exist in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple motion information items that previously existed in the merge candidate list.
[0504] Additionally, the merge candidate list may include history-based merge candidates. History-based merge candidates may be motion information of blocks encoded and / or decoded prior to the target block.
[0505] In addition, the list of merge candidates may include merge candidates based on the average of two merge candidates.
[0506] Merging candidates can be specific patterns for deriving inter-frame prediction information. Merging candidates can be information indicating specific patterns for deriving inter-frame prediction information. Inter-frame prediction information for a target block can be derived based on the specific patterns indicated by the merging candidates. Furthermore, the specific pattern can include the processing of deriving a series of inter-frame prediction information. This specific pattern can be an inter-frame prediction information derivation pattern or a motion information derivation pattern.
[0507] Inter-frame prediction information for the target block can be derived based on the pattern indicated by the merge candidate selected by the merge index from the merge candidate list.
[0508] For example, the motion information derivation mode in the merged candidate list can be at least one of 1) motion information derivation mode for sub-block units and 2) affine motion information derivation mode.
[0509] In addition, the list of merged candidates may include motion information for the zero vector. The zero vector may also be referred to as a "zero merged candidate".
[0510] In other words, the multiple motion information in the merged candidate list can be at least one of the following: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple motion information that previously existed in the merged candidate list, and 4) zero vector.
[0511] Motion information may include: 1) motion vectors, 2) reference frame indexes, and 3) reference directions. The reference direction can also be referred to as an "inter-frame prediction indicator." The reference direction can be unidirectional or bidirectional. A unidirectional reference direction can indicate L0 prediction or L1 prediction.
[0512] A list of merge candidates can be created before performing predictions in merge mode.
[0513] The number of merge candidates in the merge candidate list can be predefined. Each of the encoding device 100 and the decoding device 200 can add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, such that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 can be made identical using the predefined scheme and predefined priority.
[0514] Merging can be performed based on either CU or PU. When merging is performed based on either CU or PU, the encoding device 100 can send a bitstream including predefined information to the decoding device 200. For example, the predefined information may include: 1) information indicating whether merging is performed for each block partition, and 2) information about the blocks that will be merged with the target block, which are spatial and / or temporal candidates for the target block.
[0515] 2-2) Search for motion vectors using a merged candidate list Encoding device 100 can determine merge candidates to be used for encoding a target block. For example, encoding device 100 can perform prediction on the target block using merge candidates from the merge candidate list and can generate residual blocks for the merge candidates. Encoding device 100 can encode the target block using the merge candidate that causes the minimum cost in encoding the prediction and the residual block.
[0516] In addition, the encoding device 100 can determine whether to use a merge mode to encode the target block.
[0517] 2-3) Transmission of inter-frame prediction information Encoding device 100 can generate a bitstream including inter-frame prediction information required for inter-frame prediction. Encoding device 100 can generate entropy-coded inter-frame prediction information by performing entropy coding on the inter-frame prediction information, and can send the bitstream including the entropy-coded inter-frame prediction information to decoding device 200. The bitstream allows encoding device 100 to transmit the entropy-coded inter-frame prediction information to decoding device 200 as a signal. Decoding device 200 can extract the entropy-coded inter-frame prediction information from the bitstream, and can obtain inter-frame prediction information by applying entropy decoding to the entropy-coded inter-frame prediction information.
[0518] The decoding device 200 can use the inter-frame prediction information of the bitstream to perform inter-frame prediction on the target block.
[0519] Inter-frame prediction information may include: 1) mode information indicating whether to use a merging mode, 2) merging index, and 3) correction information.
[0520] In addition, inter-frame prediction information may include residual signals.
[0521] The decoding device 200 can only obtain the merge index from the bitstream when the mode information indicates that the merge mode is used.
[0522] Pattern information can be merge flags. The unit of pattern information can be a block. Information about a block can include pattern information, and this pattern information can indicate whether a merge pattern is applied to the block.
[0523] The merge index can indicate a merge candidate that will be used for prediction of the target block among the merge candidates included in the merge candidate list. Optionally, the merge index can indicate the block that the target block will merge with among its spatially or temporally adjacent neighboring blocks.
[0524] Encoding device 100 can select the merge candidate with the highest encoding performance from among the merge candidates included in the merge candidate list, and set the value of the merge index to indicate the selected merge candidate.
[0525] The correction information can be information used to correct motion vectors. Encoding device 100 can generate the correction information. Decoding device 200 can correct the motion vectors of the merging candidates selected by the merging index based on the correction information.
[0526] The correction information may include at least one of information indicating whether correction will be performed, correction direction information, and correction magnitude information. A prediction mode that corrects motion vectors based on correction information transmitted by a signal may be referred to as a "merging mode with motion vector difference".
[0527] 2-4) Inter-frame prediction using merging mode with inter-frame prediction information Decoding device 200 can perform prediction on target blocks using merge candidates indicated by merge indexes that are included in the merge candidate list.
[0528] The motion vector of the target block can be specified by the motion vector of the merge candidate indicated by the merge index, the reference screen index, and the reference direction.
[0529] 3) Skip Mode Skip mode can be a mode in which spatial or temporal motion information is applied to the target block without alteration. Furthermore, skip mode can be a mode that does not use the residual signal. In other words, when using skip mode, the reconstructed block can be identical to the predicted block.
[0530] The difference between merge mode and skip mode lies in whether or not residual signals are transmitted or used. In other words, skip mode is similar to merge mode except that residual signals are not transmitted or used.
[0531] When using skip mode, encoding device 100 can send information about blocks whose motion information will be used as motion information for target blocks via a bitstream to decoding device 200. Encoding device 100 can generate entropy-coded information by performing entropy encoding on this information, and can transmit the entropy-coded information as a signal to decoding device 200 via a bitstream. Decoding device 200 can extract the entropy-coded information from the bitstream, and can obtain information by applying entropy decoding to the entropy-coded information.
[0532] Furthermore, when using skip mode, encoding device 100 may not send other syntax information, such as MVD, to decoding device 200. For example, when using skip mode, encoding device 100 may not signal at least one syntax element associated with MVD, code block flag, and transform coefficient level to decoding device 200.
[0533] 3-1) Creation of the merged candidate list Skip mode can also use a merge candidate list. In other words, a merge candidate list can be used in both merge mode and skip mode. In this respect, the merge candidate list can also be referred to as a "skip candidate list" or a "merge / skip candidate list".
[0534] Optionally, the skip mode can use an additional candidate list that differs from the merge candidate list in the merge mode. In this case, in the following description, the merge candidate list and merge candidate can be replaced by the skip candidate list and the skip candidate, respectively.
[0535] A list of merged candidates can be created before performing predictions in skip mode.
[0536] 3-2) Search for motion vectors using a merged candidate list Encoding device 100 can determine merge candidates to be used for encoding a target block. For example, encoding device 100 can perform prediction on the target block using merge candidates from the merge candidate list. Encoding device 100 can use the merge candidate that causes the minimum cost in the prediction to encode the target block.
[0537] In addition, the encoding device 100 can determine whether to use a skip mode to encode the target block.
[0538] 3-3) Transmission of inter-frame prediction information Encoding device 100 can generate a bitstream that includes inter-frame prediction information required for inter-frame prediction. Decoding device 200 can use the inter-frame prediction information in the bitstream to perform inter-frame prediction on the target block.
[0539] Inter-frame prediction information may include: 1) mode information indicating whether to use a skip mode, and 2) a skip index.
[0540] Skipping indexes is the same as merging indexes as described above.
[0541] When using skip mode, the target block can be encoded without using the residual signal. Inter-frame prediction information may not include the residual signal. Optionally, the bitstream may not include the residual signal.
[0542] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates the use of skip mode. As mentioned above, the merge index and skip index may be the same. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates the use of either merge mode or skip mode.
[0543] Skip indexes can indicate which merge candidates will be used for the prediction of the target block, and which are included in the merge candidate list.
[0544] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information Decoding device 200 can perform prediction on target blocks using merge candidates indicated by skip indexes that are included in the merge candidate list.
[0545] The motion vector of the target block can be specified by the motion vector of the merge candidate indicated by the skip index, the reference screen index, and the reference direction.
[0546] 4) Current screen reference mode The current frame reference mode can represent the prediction mode of the previously reconstructed area in the target frame to which the target block belongs.
[0547] Motion vectors can be used to specify previously reconstructed areas. The reference frame index of the target block can be used to determine whether the target block has already been encoded in the current frame reference mode.
[0548] A flag or index indicating whether a target block is encoded in the current screen reference mode can be transmitted by the encoding device 100 to the decoding device 200 via a signal. Optionally, whether a target block is encoded in the current screen reference mode can be inferred from the reference screen index of the target block.
[0549] When a target block is encoded in the current frame reference mode, the target frame can exist in a fixed position or any position in the reference frame list for the target block.
[0550] For example, a fixed position could be the position where the reference screen index value is 0 or the last position.
[0551] When the target image exists at any position in the reference image list, the index of an additional reference image indicating such an arbitrary position can be transmitted by the encoding device 100 to the decoding device 200 via a signal.
[0552] 5) Sub-block merging mode Sub-block merging mode can be a mode that derives motion information from sub-blocks of the CU.
[0553] When applying the sub-block merging mode, the motion information of the col-blocks (col-blocks) of the target sub-block in the reference image (i.e., time-based merging candidates based on sub-blocks) and / or affine control point motion vector merging candidates can be used to generate a list of sub-block merging candidates.
[0554] 6) Triangle partitioning mode In triangular partitioning mode, the target block can be partitioned diagonally, and sub-target blocks generated by the partitioning can be produced. For each sub-target block, the motion information of the corresponding sub-target block can be derived, and the derived motion information can be used to derive prediction samples for each sub-target block. The prediction samples of the target block can be derived by weighted summing of the prediction samples of the sub-target blocks generated by the partitioning.
[0555] 7) Combined inter-frame and intra-frame prediction modes The combined inter-frame-intra-frame prediction mode can be a mode that uses a weighted sum of prediction samples generated via inter-frame prediction and prediction samples generated via intra-frame prediction to derive prediction samples for the target block.
[0556] In the above mode, the decoding device 200 can autonomously correct the derived motion information. For example, the decoding device 200 can search for motion information with the minimum sum of absolute differences (SAD) in a specific region based on a reference block indicated by the derived motion information, and can derive the found motion information as corrected motion information.
[0557] In the above mode, the decoding device 200 can use optical flow to compensate for the prediction samples derived via inter-frame prediction.
[0558] In the aforementioned AMVP mode, merge mode, skip mode, etc., the index information of the list can be used to specify the motion information to be used for the prediction of the target block among multiple motion information in the list.
[0559] To improve coding efficiency, the encoding device 100 may use only the index of the element in the signal transmission list that causes the minimum cost in inter-frame prediction of the target block. The encoding device 100 may encode the index and may transmit the encoded index via signal transmission.
[0560] Therefore, the aforementioned lists (i.e., the candidate lists for predicted motion vectors and the candidate lists for merging) must be able to be derived by the encoding device 100 and the decoding device 200 using the same data and the same scheme. Here, the same data may include reconstructed frames and reconstructed blocks. Furthermore, in order to specify elements using indices, the order of elements in the lists must be fixed.
[0561] Figure 10 Spatial candidates according to an embodiment are shown.
[0562] exist Figure 10 The image shows the location of the spatial candidate.
[0563] The large block in the center of the attached diagram represents the target block. The five smaller blocks represent spatial candidates.
[0564] The coordinates of the target block can be (xP, yP), and the size of the target block can be represented by (nPSW, nPSH).
[0565] Spatial candidate A0 can be a block adjacent to the lower left corner of the target block. A0 can be a block that occupies the pixel located at coordinates (xP - 1, yP + nPSH).
[0566] Spatial candidate A1 can be the block adjacent to the left of the target block. A1 can be the bottommost block among the blocks adjacent to the left of the target block. Alternatively, A1 can be the block adjacent to the top of A0. A1 can be the block occupying the pixel located at coordinates (xP - 1, yP + nPSH - 1).
[0567] Spatial candidate B0 can be the block adjacent to the top-right corner of the target block. B0 can be the block that occupies the pixel located at coordinates (xP + nPSW, yP - 1).
[0568] Spatial candidate B1 can be the block that is top-adjacent to the target block. B1 can be the rightmost block among the blocks that are top-adjacent to the target block. Alternatively, B1 can be the block that is left-adjacent to B0. B1 can be the block that occupies the pixel located at coordinates (xP + nPSW-1, yP -1).
[0569] Spatial candidate B2 can be a block adjacent to the top-left corner of the target block. B2 can be a block that occupies the pixel located at coordinates (xP - 1, yP - 1).
[0570] Determining the availability of spatial and temporal candidates In order to include spatial or temporal motion information in the list, it must be determined whether the spatial or temporal motion information is available.
[0571] In the following text, candidate blocks may include spatial candidates and temporal candidates.
[0572] For example, this determination can be performed by sequentially applying the following steps 1) through 4).
[0573] Step 1) When the PU including the candidate block is outside the boundary of the screen, the availability of the candidate block can be set to "false". The expression "availability is set to false" can have the same meaning as "set to unavailable".
[0574] Step 2) When the PU including the candidate block is outside the stripe boundary, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different stripes, the availability of the candidate block can be set to "false".
[0575] Step 3) When the PU including the candidate block is outside the boundary of the parallel block, the availability of the candidate block can be set to "false". When the target block and the candidate block are in different parallel blocks, the availability of the candidate block can be set to "false".
[0576] Step 4) When the prediction mode of the PU including the candidate block is intra-frame prediction mode, the availability of the candidate block can be set to "false". When the PU including the candidate block does not use inter-frame prediction, the availability of the candidate block can be set to "false".
[0577] Figure 11 The order in which motion information of spatial candidates is added to the merging list is shown according to an embodiment.
[0578] like Figure 11As shown, when adding multiple motion information entries from spatial candidates to the merge list, the order A1, B1, B0, A0, and B2 can be used. That is, multiple available motion information entries from spatial candidates can be added to the merge list in the order A1, B1, B0, A0, and B2.
[0579] Methods for deriving merge lists in merge mode and skip mode As described above, the maximum number of merge candidates in the merge list can be set. The maximum number is indicated by "N". The set number can be sent from the encoding device 100 to the decoding device 200. The stripe header can include N. In other words, the maximum number of merge candidates in the merge list for the target block of the stripe can be set via the stripe header. For example, the value of N can be approximately 5.
[0580] Multiple motion information (i.e., merge candidates) can be added to the merge list in the order of steps 1) to 4).
[0581] Step 1) In the spatial candidates, available spatial candidates can be added to the merge list. This can be done by... Figure 11 The order shown adds multiple motion information entries from available space candidates to the merge list. Here, if the motion information from an available space candidate overlaps with other motion information already existing in the merge list, that motion information may not be added to the merge list. The operation of checking whether corresponding motion information overlaps with other motion information existing in the list can be simply referred to as "overlap check".
[0582] The maximum number of motion information entries that can be added is N.
[0583] Step 2) When the number of motion information entries in the merge list is less than N and time candidates are available, the motion information of the time candidates can be added to the merge list. Here, if the available time candidate motion information overlaps with other motion information already existing in the merge list, that motion information may not be added to the merge list.
[0584] Step 3) When the number of motion information entries in the merge list is less than N and the target strip type is "B", the combined motion information generated by combining bidirectional prediction (dual prediction) can be added to the merge list.
[0585] The target strip can be a strip that includes the target block.
[0586] Combined motion information can be a combination of L0 motion information and L1 motion information. L0 motion information can be motion information that only references the L0 reference frame list. L1 motion information can be motion information that only references the L1 reference frame list.
[0587] The merged list may contain one or more L0 motion entries. Additionally, the merged list may contain one or more L1 motion entries.
[0588] Combined motion information may include one or more combined motion information pieces. When generating combined motion information, the L0 and L1 motion information pieces from one or more L0 motion information pieces and one or more L1 motion information pieces that will be used for generation can be predefined. One or more combined motion information pieces can be generated via combined bidirectional prediction in a predefined order, wherein the combined bidirectional prediction uses a pair of different motion information pieces from a merge list. One of the pair of different motion information pieces can be an L0 motion information piece, and the other of the pair can be an L1 motion information piece.
[0589] For example, the combined motion information added with the highest priority can be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information, or when the motion information with a merge index of 1 is not L1 motion information, neither combined motion information is generated nor added. Next, the combined motion information added with the next highest priority can be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. Subsequent detailed combinations can conform to other combinations in the field of video encoding / decoding.
[0590] Here, when the combined motion information overlaps with other motion information that already exists in the merge list, the combined motion information may not be added to the merge list.
[0591] Step 4) When the number of motion information entries in the merge list is less than N, the motion information of the zero vector can be added to the merge list.
[0592] Zero-vector motion information can be motion information where the motion vector is zero.
[0593] The number of zero-vector motion information entries can be one or more. The reference frame indices for one or more zero-vector motion information entries can be different from each other. For example, the reference frame index value for the first zero-vector motion information entry can be 0. The reference frame index value for the second zero-vector motion information entry can be 1.
[0594] The number of zero-vector motion information entries can be the same as the number of reference frames in the reference frame list.
[0595] The reference direction for zero-vector motion information can be bidirectional. Both motion vectors can be zero vectors. The number of zero-vector motion information entries can be the smaller of the number of reference frames in reference frame list L0 and the number of reference frames in reference frame list L1. Optionally, when the number of reference frames in reference frame list L0 and the number of reference frames in reference frame list L1 are different from each other, a unidirectional reference direction can be used for reference frame indexing that can be applied to only a single reference frame list.
[0596] Encoding device 100 and / or decoding device 200 can sequentially add zero-vector motion information to the merge list while changing the reference screen index.
[0597] When zero-vector motion information overlaps with other motion information that already exists in the merge list, the zero-vector motion information may not be added to the merge list.
[0598] The order of steps 1) to 4) above is merely exemplary and can be changed. Furthermore, some of the above steps may be omitted based on predefined conditions.
[0599] A method for deriving a candidate list of predicted motion vectors in AMVP mode The maximum number of predicted motion vector candidates in the candidate list can be predefined. This predefined maximum number is indicated by N. For example, the predefined maximum number could be 2.
[0600] Multiple pieces of motion information (i.e., predicted motion vector candidates) can be added to the predicted motion vector candidate list in the order of steps 1) to 3).
[0601] Step 1) Available spatial candidates can be added to the list of predicted motion vector candidates. Spatial candidates can include first spatial candidates and second spatial candidates.
[0602] The first spatial candidate can be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate can be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.
[0603] Multiple motion information entries from available spatial candidates can be added to the predicted motion vector candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, if the motion information of an available spatial candidate overlaps with other motion information already existing in the predicted motion vector candidate list, that motion information may not be added to the predicted motion vector candidate list. In other words, when N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the predicted motion vector candidate list.
[0604] The maximum number of motion information entries that can be added is N.
[0605] Step 2) When the number of motion information entries in the predicted motion vector candidate list is less than N and a time candidate is available, the motion information of the time candidate can be added to the predicted motion vector candidate list. In this case, if the motion information of an available time candidate overlaps with other motion information already existing in the predicted motion vector candidate list, that motion information may not be added to the predicted motion vector candidate list.
[0606] Step 3) When the number of motion information entries in the candidate list of predicted motion vectors is less than N, zero-vector motion information can be added to the candidate list of predicted motion vectors.
[0607] Zero-vector motion information may include one or more zero-vector motion information entries. The reference frame indices for one or more zero-vector motion information entries may be different from each other.
[0608] Encoding device 100 and / or decoding device 200 can sequentially add multiple zero-vector motion information to the candidate list of predicted motion vectors while changing the reference frame index.
[0609] When zero-vector motion information overlaps with other motion information that already exists in the candidate list of predicted motion vectors, the zero-vector motion information may not be added to the candidate list of predicted motion vectors.
[0610] The description of zero-vector motion information above, combined with the merged list, can also be applied to zero-vector motion information. Repeated descriptions will be omitted.
[0611] The order of steps 1) to 3) above is merely exemplary and can be changed. Furthermore, some steps may be omitted based on predefined conditions.
[0612] Figure 12 The transformation and quantization processes are shown based on the example.
[0613] like Figure 12 As shown, quantization levels can be generated by performing transformations and / or quantization on the residual signal.
[0614] The residual signal can be generated as the difference between the original block and the predicted block. Here, the predicted block can be a block generated via intra-frame prediction or inter-frame prediction.
[0615] The residual signal can be transformed into a signal in the frequency domain through a transformation process that is part of the quantization process.
[0616] Transform kernels used for transformations can include various DCT kernels, such as Discrete Cosine Transform (DCT) Type 2 (DCT-II) and Discrete Sine Transform (DST) kernels.
[0617] These transform kernels can perform separable or two-dimensional (2D) non-separable transforms on the residual signal. A separable transform can be a transform that indicates performing a one-dimensional (1D) transform on the residual signal in each of the horizontal and vertical directions.
[0618] In addition to DCT-II, the DCT and DST types adaptively used for 1D transformations may also include DCT-V, DCT-VIII, DST-I and DST-VII, as shown in each of Tables 3 and 4 below.
[0619] [Table 3]
[0620] [Table 4]
[0621] As shown in Tables 3 and 4, transform sets can be used when deriving the DCT or DST type to be used for the transform. Each transform set may include multiple transform candidates. Each transform candidate can be a DCT type or a DST type.
[0622] Table 5 below shows examples of the transform sets to be applied in the horizontal direction and the transform sets to be applied in the vertical direction, based on the intra-frame prediction mode.
[0623] [Table 5]
[0624] Table 5 indicates the number of vertical and horizontal transform sets that will be applied to the horizontal direction of the residual signal according to the intra-frame prediction mode of the target block.
[0625] As illustrated in Table 5, the transform sets to be applied in the horizontal and vertical directions can be predefined based on the intra-prediction mode of the target block. Encoding device 100 can use transforms included in the transform set corresponding to the intra-prediction mode of the target block to perform transforms and inverse transforms on the residual signal. Furthermore, decoding device 200 can use transforms included in the transform set corresponding to the intra-prediction mode of the target block to perform an inverse transform on the residual signal.
[0626] In the transform and inverse transform, the set of transforms to be applied to the residual signal can be determined, as illustrated in Tables 3, 4, and 5, and the set of transforms to be applied to the residual signal can be transmitted without signal transmission. Transform indication information can be transmitted from the encoding device 100 to the decoding device 200 by signal transmission. The transform indication information may be information indicating which of a plurality of transform candidates included in the set of transforms to be applied to the residual signal should be used.
[0627] For example, when the target block size is 64×64 or smaller, transform sets, each with three transforms, can be configured according to the intra-frame prediction mode. The optimal transform method can be selected from a total of nine multi-transform methods generated by combinations of three transforms in the horizontal direction and three transforms in the vertical direction. With such an optimal transform method, the residual signal can be encoded and / or decoded, thus improving coding efficiency.
[0628] Here, information indicating which of the transformations belonging to each transform set has been used for at least one of the vertical and horizontal transformations can be entropy encoded and / or decoded. Here, truncated univariate binarization can be used to encode and / or decode such information.
[0629] As mentioned above, various transformation methods can be applied to residual signals generated via intra-frame prediction or inter-frame prediction.
[0630] The transformation may include at least one of a primary transformation and a secondary transformation. Transform coefficients can be generated by performing a primary transformation on the residual signal, and secondary transform coefficients can be generated by performing a secondary transformation on the transform coefficients.
[0631] The initial transformation can be referred to as the "primary transformation." Alternatively, it can be called an "Adaptive Multiple Transformation (AMT) scheme." AMT can refer to applying different transformations to various 1D directions (i.e., the vertical and horizontal directions), as described above.
[0632] A secondary transformation can be a transformation used to improve the energy concentration on the transformation coefficients generated by the primary transformation. Similar to the primary transformation, a secondary transformation can be a separable or non-separable transformation. Such a non-separable transformation can be a non-separable secondary transformation (NSST).
[0633] The initial transformation can be performed using at least one of a variety of predefined transformation methods. For example, the predefined transformation methods may include the Discrete Cosine Transform (DCT), the Discrete Sine Transform (DST), the Karhunen-Loeve Transform (KLT), and so on.
[0634] Furthermore, the initial transformation can be a transformation with various transformation types based on the kernel function of the defined Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST).
[0635] For example, the transform type can be determined based on at least one of the following: 1) the prediction mode of the target block (e.g., one of intra-frame prediction and inter-frame prediction), 2) the size of the target block, 3) the shape of the target block, 4) the intra-frame prediction mode of the target block, 5) the components of the target block (e.g., one of luma component and chroma component), and 6) the partition type applied to the target block (e.g., one of quadtree, binary tree and ternary tree).
[0636] For example, based on the transform kernels presented in Table 6 below, the initial transform may include transforms such as DCT-2, DCT-5, DCT-7, DST-7, DST-1, DST-8, and DCT-8. Table 6 below illustrates various transform types and transform kernel functions used for Multiple Transform Selection (MTS).
[0637] MTS can refer to the selection of a combination of one or more DCT and / or DST cores to transform the residual signal in the horizontal and / or vertical directions.
[0638] [Table 6]
[0639] In Table 6, i and j can be integer values that are equal to or greater than 0 and less than or equal to N-1.
[0640] Secondary transformations can be performed on the transformation coefficients generated by performing the initial transformation.
[0641] For example, the transformation set can also be defined in the secondary transformation, just as it can be defined in the initial transformation. The methods used to derive and / or determine the above transformation set can be applied not only to the initial transformation but also to the secondary transformation.
[0642] The primary and secondary transformations can be determined for a specific target.
[0643] For example, primary and secondary transformations can be applied to one or more signal components corresponding to the luma and chroma components. Whether to apply the primary and / or secondary transformations can be determined based on at least one of the coding parameters used for the target block and / or neighboring blocks. For example, whether to apply the primary and / or secondary transformations can be determined based on the size and / or shape of the target block.
[0644] In the encoding device 100 and the decoding device 200, transformation information indicating the transformation method to be used for the target can be derived by utilizing specified information.
[0645] For example, the transformation information may include transformation indices that will be used for primary and / or secondary transformations. Optionally, the transformation information may indicate that primary and / or secondary transformations are not used.
[0646] For example, when the target of the primary and secondary transforms is a target block, one or more transform methods to be applied to the primary and / or secondary transforms indicated by the transform information can be determined based on at least one of the encoding parameters of the target block and / or blocks adjacent to the target block.
[0647] Optionally, transformation information indicating a transformation method for a specific target can be sent from the encoding device 100 to the decoding device 200.
[0648] For example, for a single CU, the decoding device 200 can deduce whether to use a primary transform, the index indicating the primary transform, whether to use a secondary transform, and the index indicating the secondary transform as transform information. Optionally, for a single CU, transform information indicating whether to use a primary transform, the index indicating the primary transform, whether to use a secondary transform, and the index indicating the secondary transform can be transmitted via signals.
[0649] Quantization transform coefficients (i.e., quantization levels) can be generated by quantizing the results produced by performing the first transform and / or the second transform, or by quantizing the residual signal.
[0650] Figure 13 This shows a diagonal scan based on an example.
[0651] Figure 14 The horizontal scan is shown based on the example.
[0652] Figure 15 The vertical scan is shown according to the example.
[0653] The quantized transform coefficients can be scanned via at least one of (top right) diagonal scan, vertical scan, and horizontal scan, based on at least one of the intra-frame prediction mode, block size, and block shape. The block can be a transform unit (TU).
[0654] Each scan can be initiated at a specific starting point and terminated at a specific ending point.
[0655] For example, by using Figure 13 A diagonal scan is used to scan the coefficients of the block, which can transform the quantized transform coefficients into a 1D vector form. Optionally, the block size and / or intra-frame prediction mode can be used. Figure 14 Horizontal scan or Figure 15 It uses vertical scanning instead of diagonal scanning.
[0656] A vertical scan is an operation that scans 2D block coefficients in the column direction. A horizontal scan is an operation that scans 2D block coefficients in the row direction.
[0657] In other words, the choice between diagonal, vertical, and horizontal scanning can be determined based on the block size and / or inter-frame prediction mode.
[0658] like Figure 13 , Figure 14 and Figure 15 As shown, the quantization transform coefficients can be scanned along the diagonal, horizontal, or vertical directions.
[0659] Quantization transformation coefficients can be represented by block shapes. Each block can include multiple sub-blocks. Each sub-block can be defined based on the minimum block size or the minimum block shape.
[0660] During scanning, the scanning order, based on the type or direction of the scan, can be primarily applied to sub-blocks. Furthermore, the scanning sequence, based on the scanning direction, can be applied to the quantization transform coefficients within each sub-block.
[0661] For example, such as Figure 13 , Figure 14 and Figure 15 As shown, when the target block size is 8×8, quantization transform coefficients can be generated by performing a first transform, a second transform, and quantization on the residual signal of the target block. Therefore, one of the three types of scan sequences can be applied to four 4×4 sub-blocks, and quantization transform coefficients can be scanned for each 4×4 sub-block according to the scan sequence.
[0662] Encoding device 100 can generate entropy-coded quantization transform coefficients by performing entropy coding on the scanned quantization transform coefficients, and can generate a bit stream including the entropy-coded quantization transform coefficients.
[0663] The decoding device 200 can extract entropy-encoded quantization transform coefficients from the bitstream, and can generate quantization transform coefficients by performing entropy decoding on the entropy-encoded quantization transform coefficients. The quantization transform coefficients can be aligned in the form of 2D blocks via inverse scanning. Here, as a method of inverse scanning, at least one of upper right diagonal scanning, vertical scanning, and horizontal scanning can be performed.
[0664] In the decoding device 200, inverse quantization can be performed on the quantization transform coefficients. A secondary inverse transform can be performed on the result generated by inverse quantization, depending on whether a secondary inverse transform is performed. Furthermore, a first inverse transform can be performed on the result generated by the secondary inverse transform, depending on whether a first inverse transform will be performed. The reconstructed residual signal can be generated by performing a first inverse transform on the result generated by the secondary inverse transform.
[0665] For luminance components reconstructed via intra-frame prediction or inter-frame prediction, an inverse mapping with dynamic range can be performed before in-loop filtering.
[0666] The dynamic range can be divided into 16 equal segments, and mapping functions for each segment can be signaled. Such mapping functions can be signaled at the stripe level or the parallel block group level.
[0667] The inverse mapping function can be derived from the mapping function to perform the inverse mapping.
[0668] Intra-loop filtering, reference frame storage, and motion compensation can be performed in the inverse mapping region.
[0669] Predicted blocks generated via inter-frame prediction can be transformed into mapped regions using a mapping function, and the transformed predicted blocks can be used to generate reconstructed blocks. However, since intra-frame prediction is performed within the mapped regions, predicted blocks generated via intra-frame prediction can be used to generate reconstructed blocks without requiring mapping and / or inverse mapping.
[0670] For example, when the target block is a residual block of the chrominance component, the residual block can be changed into an inverse mapping region by scaling the chrominance component of the mapping region.
[0671] Scaling availability can be signaled at the stripe level or parallel block group level.
[0672] For example, scaling can be applied only when the mapping is available for the luminance component and the partitions for the luminance and chrominance components follow the same tree structure.
[0673] Scaling can be performed based on the average value of samples in a luminance prediction block, which corresponds to a chrominance prediction block. Here, when the target block uses inter-frame prediction, the luminance prediction block may refer to the mapped luminance prediction block.
[0674] The required scaling value can be derived by referencing a lookup table using the index of the segment to which the average value of the sample values of the brightness prediction block belongs.
[0675] The residual block can be transformed into an inverse-mapped region by scaling it with the final derived values. Subsequently, for blocks of chroma components, reconstruction, intra-frame prediction, inter-frame prediction, intra-loop filtering, and storage of reference frames can be performed within the inverse-mapped region.
[0676] For example, information indicating whether the mapping and / or inverse mapping of the luminance and chrominance components is available can be transmitted via a sequence parameter set using signals.
[0677] Predicted blocks for a target block can be generated based on block vectors. The block vectors indicate the displacement between the target block and a reference block. The reference block can be a block in the target image.
[0678] In this way, the prediction mode that generates prediction blocks by referencing the target image can be called the "intra-block copy (IBC) mode".
[0679] The IBC mode can be applied to CUs with specific dimensions. For example, the IBC mode can be applied to an M×N CU. Here, M and N can be less than or equal to 64.
[0680] IBC modes can include skip mode, merge mode, AMVP mode, etc. In skip mode or merge mode, a merge candidate list can be configured, and the merge index is signaled, thus allowing a single merge candidate to be specified from among the existing merge candidates in the merge candidate list. The block vector of the specified merge candidate can be used as the block vector of the target block.
[0681] In AMVP mode, differential block vectors can be signaled. Furthermore, predicted block vectors can be derived from the target block's left and top neighboring blocks. Additionally, the index of which neighboring block will be used can be signaled.
[0682] In IBC mode, the predicted block can be included in the target CTU or the left CTU, and can be limited to blocks within the previously reconstructed region. For example, the value of the block vector can be restricted such that the predicted block of the target block is located in a specific region. The specific region can be defined by three 64×64 blocks encoded and / or decoded before the 64×64 block including the target block. Restricting the value of the block vector in this way reduces memory consumption and device complexity caused by the implementation of IBC mode.
[0683] Figure 16 This is a configuration diagram of an encoding device according to an embodiment.
[0684] Encoding device 1600 may correspond to the aforementioned encoding device 100.
[0685] Encoding device 1600 may include processing unit 1610, memory 1630, user interface (UI) input device 1650, UI output device 1660, and storage device 1640, which communicate with each other via bus 1690. Encoding device 1600 may also include communication unit 1620 connected to network 1699.
[0686] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1630 or the storage device 1640. The processing unit 1610 may be at least one hardware processor.
[0687] The processing unit 1610 can generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and can perform checks, comparisons, determinations, etc., related to the signals, data, or information. In other words, in this embodiment, the generation and processing of data or information, as well as the checks, comparisons, and determinations related to the data or information, can be performed by the processing unit 1610.
[0688] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.
[0689] At least some of the following components—inter-frame prediction unit 110, intra-frame prediction unit 120, switcher 115, subtractor 125, transform unit 130, quantization unit 140, entropy coding unit 150, inverse quantization unit 160, inverse transform unit 170, adder 175, filter unit 180, and reference frame buffer 190—may be program modules and capable of communicating with external devices or systems. These program modules may be included in the encoding device 1600 in the form of an operating system, application module, or other program modules.
[0690] The program modules can be physically stored in various types of known storage devices. In addition, at least some of the program modules can also be stored in a remote storage device capable of communicating with the encoding device 1600.
[0691] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.
[0692] The program module can be implemented using instructions or code executed by at least one processor of the encoding device 1600.
[0693] The processing unit 1610 can execute instructions or codes in the inter-frame prediction unit 110, intra-frame prediction unit 120, switcher 115, subtractor 125, transform unit 130, quantization unit 140, entropy coding unit 150, dequantization unit 160, inverse transform unit 170, adder 175, filter unit 180 and reference frame buffer 190.
[0694] The storage cell may represent memory 1630 and / or storage device 1640. Each of memory 1630 and storage device 1640 may be any of a variety of volatile or non-volatile storage media. For example, memory 1630 may include at least one of read-only memory (ROM) 1631 and random access memory (RAM) 1632.
[0695] The storage unit can store data or information used for the operation of the encoding device 1600. In an embodiment, the data or information of the encoding device 1600 can be stored in the storage unit.
[0696] For example, storage units can store images, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.
[0697] The encoding device 1600 can be implemented in a computer system that includes a computer-readable storage medium.
[0698] The storage medium may store at least one module required for the operation of the encoding device 1600. The memory 1630 may store at least one module and may be configured such that at least one module is executed by the processing unit 1610.
[0699] The communication unit 1620 can perform functions related to communication of data or information with the encoding device 1600.
[0700] For example, communication unit 1620 can send a bit stream to decoding device 1600, which will be described later.
[0701] Figure 17 This is a configuration diagram of a decoding device according to an embodiment.
[0702] Decoding device 1700 may correspond to the aforementioned decoding device 200.
[0703] The decoding device 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage device 1740, which communicate with each other via a bus 1790. The decoding device 1700 may also include a communication unit 1720 connected to a network 1799.
[0704] Processing unit 1710 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in memory 1730 or storage device 1740. Processing unit 1710 may be at least one hardware processor.
[0705] The processing unit 1710 can generate and process signals, data, or information input to, output from, or used in the decoding device 1700, and can perform checks, comparisons, determinations, etc., related to the signals, data, or information. In other words, in this embodiment, the generation and processing of data or information, as well as the checks, comparisons, and determinations related to the data or information, can be performed by the processing unit 1710.
[0706] The processing unit 1710 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switcher 245, an adder 255, a filter unit 260, and a reference frame buffer 270.
[0707] At least some of the entropy decoding unit 210, inverse quantization unit 220, inverse transform unit 230, intra-frame prediction unit 240, inter-frame prediction unit 250, adder 255, switcher 245, filter unit 260, and reference frame buffer 270 of decoding device 200 may be program modules and are capable of communicating with external devices or systems. Program modules may be included in decoding device 1700 in the form of an operating system, application module, or other program modules.
[0708] The program modules can be physically stored in various types of known storage devices. Furthermore, at least some of the program modules can also be stored in a remote storage device capable of communicating with the decoding device 1700.
[0709] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to an embodiment or for implementing abstract data types according to an embodiment.
[0710] The program module can be implemented using instructions or code executed by at least one processor of the decoding device 1700.
[0711] The processing unit 1710 can execute instructions or codes in the entropy decoding unit 210, the dequantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the switcher 245, the adder 255, the filter unit 260, and the reference frame buffer 270.
[0712] The storage cell may represent memory 1730 and / or storage device 1740. Each of memory 1730 and storage device 1740 may be any of various types of volatile or non-volatile storage media. For example, memory 1730 may include at least one of ROM 1731 and RAM 1732.
[0713] The storage unit can store data or information used for the operation of the decoding device 1700. In an embodiment, the data or information of the decoding device 1700 can be stored in the storage unit.
[0714] For example, storage units can store images, blocks, lists, motion information, inter-frame prediction information, bitstreams, etc.
[0715] The decoding device 1700 can be implemented in a computer system that includes a computer-readable storage medium.
[0716] The storage medium may store at least one module required for the operation of the decoding device 1700. The memory 1730 may store at least one module and may be configured such that at least one module is executed by the processing unit 1710.
[0717] The communication unit 1720 can be used to perform functions related to communication of data or information with the decoding device 1700.
[0718] For example, communication unit 1720 can receive bit streams from encoding device 1700.
[0719] In the following text, "processing unit" may refer to processing unit 1610 of encoding device 1600 and / or processing unit 1710 of decoding device 1700. For example, regarding prediction-related functions, the processing unit may represent switcher 115 and / or switcher 245. Regarding inter-frame prediction-related functions, the processing unit may represent inter-frame prediction unit 110, subtractor 125, and adder 175, and may also represent inter-frame prediction unit 250 and adder 255. Regarding intra-frame prediction-related functions, the processing unit may represent intra-frame prediction unit 120, subtractor 125, and adder 175, and may also represent intra-frame prediction unit 240 and adder 255. Regarding transform-related functions, the processing unit may represent transform unit 130 and inverse transform unit 170, and may also represent inverse transform unit 230. Regarding quantization-related functions, the processing unit may represent quantization unit 140 and inverse quantization unit 160, and may also indicate inverse quantization unit 220. Regarding functions related to entropy encoding and / or entropy decoding, the processing unit may represent entropy encoding unit 150 and / or entropy decoding unit 210. Regarding functions related to filtering, the processing unit may represent filter unit 180 and / or filter unit 260. Regarding functions related to reference frames, the processing unit may instruct reference frame buffer 190 and / or reference frame buffer 270.
[0720] The following text discloses refined image encoding and decoding techniques implemented in image encoding and image decoding devices.
[0721] 1. Template matching Figure 18 This is an illustration of an embodiment used for template matching.
[0722] In template matching, the motion information of the target block can be determined and / or modified based on the calculation result of the cost function between the target template used for the target block and the reference template used for the reference block.
[0723] In some embodiments, template matching can be used to determine the motion information of the target block. Based on the cost calculation between the target template and each reference template, the motion information corresponding to the displacement from the position of the target block to the position of the reference block with the lowest cost reference template can be used as the motion information of the target block.
[0724] In another embodiment, template matching can be used to modify or correct the motion information of the target block. For example, reference templates are determined for a reference block at a position indicated by the initial motion information of the target block and a reference block at a position at a predetermined interval in a predetermined direction. Furthermore, based on cost calculations between the target template and each reference template, the motion information corresponding to the displacement from the target block to the reference block with the lowest cost reference template can be used as the final motion information of the target block.
[0725] In another embodiment, template matching can be used to reorder motion information candidates included in the motion information candidate list for the target block. For example, reference templates for reference blocks at the locations indicated by the motion information candidates are configured. Furthermore, based on cost calculations between the target template and its reference templates, the motion information candidates are reordered in ascending order of cost. Thus, lower-indexed candidates for motion information that are highly likely to be selected as the target block can be assigned.
[0726] In template matching, a reference block may include at least one of the following: a reference block indicated by initial motion information, a reference block indicated by motion information derived in the template matching search process, and a reference block indicated by motion information finally corrected by template matching.
[0727] Here, the initial motion information can be the motion information of the target block sent from the encoding device to the decoding device via a signal. Furthermore, the motion information corrected by template matching can be the motion information with the lowest matching cost derived during the template matching search process. However, the methods for deriving motion information are not limited to those mentioned above.
[0728] The template matching method may include at least one of inter-frame template matching mode and intra-frame template matching mode.
[0729] Inter-frame template matching mode can represent a template matching method for configuring a reference block based on at least one of the predicted samples, reconstructed samples, and residual samples of a reference frame reconstructed before the target frame.
[0730] Intra-frame template matching mode can represent a template matching method for configuring a reference block based on at least one of the predicted samples, reconstructed samples, and residual samples of the target image.
[0731] Template matching cost can be represented as the result of a cost function calculated using the templates of the target block and the reference block used in the template matching.
[0732] 1.1 Template Configuration The templates used in template matching can include a target template and a reference template.
[0733] The target template can be configured using a reference region that includes neighboring samples of the target block. The reference region of the target block can include at least one of the samples located in the lower left region, left side region, upper left region, top region, and upper right region.
[0734] In some embodiments, the target template may be the same as the reference area of the target block.
[0735] In some other embodiments, when configuring the target template in template matching, some sample points in the reference region of the target block can be selected. Alternatively, the target template can be configured using the selected sample points.
[0736] A reference template can be configured by using a reference region that includes neighboring samples of the reference block. The reference region of the reference block can be a region corresponding to the reference region of the target block. For example, the reference region of the reference block can include at least one of the samples located in the lower left, left, upper left, top, and upper right regions of the reference block.
[0737] In some embodiments, the reference template in template matching may be the same as the reference region of the reference block. For example, the sample points of the reference template based on the reference block may be sample points corresponding to the sample points of the target template based on the target block.
[0738] In some other embodiments, when configuring a reference template in template matching, some sample points in the reference region of a reference block can be selected. Furthermore, the reference template can be configured using the selected sample points. For example, the sample points selected for configuring the reference template based on the reference block can be sample points corresponding to the sample points selected for configuring the template based on the target block.
[0739] In inter-frame template matching mode, each of the predicted samples, reconstructed samples, and residual samples of the reference frame reconstructed before the target frame is configured from at least one of the reference block, reference template, and reference region.
[0740] In intra-frame template matching mode, each of the following is configured from at least one of the predicted samples, reconstructed samples, and residual samples of the target image: a reference block, a reference template, and a reference region. The target template / reference template for template matching may include at least one of the following: 1) at least one sample from the TMSIZE_LEFT bar adjacent to the left of the target block / reference block, and 2) at least one sample from the TMSIZE_ABOVE bar adjacent to the top of the target block / reference block.
[0741] However, the positional relationship between each sample point in the template and the target / reference block and / or the template configuration method are not limited to the relationships or methods described above.
[0742] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be a positive integer of 0, 1, 2, 3, 4 or higher.
[0743] TMSIZE_LEFT and TMSIZE_ABOVE can be the same. Alternatively, TMSIZE_LEFT and TMSIZE_ABOVE can be different.
[0744] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be a predefined value or a value determined based on information sent / encoded / decoded by signals.
[0745] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be determined based on at least one of the target block's motion information, encoding parameters, size, and prediction mode.
[0746] 1.2 Subsampling in Template Matching Subsampling can be performed on template configurations, cost calculations, and other aspects of template matching. Furthermore, subsampling can even be performed on the search range of template matching.
[0747] ■ Subsampling for template configuration For template configuration used for template matching, subsampling can be used. That is, when configuring a template, all samples in the reference region can be used, or only some samples in the reference region can be used. Here, the reference region can refer to at least one of the region referenced by the target block and the region referenced by the reference block for configuring the template.
[0748] When configuring a template for template matching by using only some sample points located in the reference region, subsampling can be performed on all or part of the reference region.
[0749] When configuring a template for template matching by using only some sample points located in the reference region, the reference region can be divided into two or more regions.
[0750] In some embodiments, each partitioned region can be one of the following: 1) a region where subsampling is performed; 2) a region where subsampling is not performed and is used to configure the template; and 3) a region not used to configure the template. The template for template matching can be configured by using samples selected by subsampling in region 1) and samples in region 2).
[0751] As an example, the area corresponding to 1) can be the area belonging to the left and / or upper left of the block among multiple areas in the reference area.
[0752] As another example, the area corresponding to 1) can be the area belonging to the top and / or upper left of the block among multiple areas in the reference area.
[0753] In some other embodiments, each partitioned region can be one of 1) a region where subsampling is performed and 2) a region not used for template configuration. A template for template matching can be configured using the sample points selected by subsampling in 1).
[0754] ■ Subsampling for cost calculation When performing cost calculation between a target template and a reference template in template matching, all samples from each template can be used, or only some samples from each template can be used. That is, the cost between templates can be calculated by using only some samples from each template.
[0755] To calculate the cost between templates using only some samples from the template, subsampling can be performed on all or part of the template region.
[0756] When performing cost calculations between templates using only some sample points from the template, each region of the template used for template matching can be divided into two or more regions.
[0757] In some embodiments, each partitioned region can be one of the following: 1) a region where subsampling is performed; 2) a region where subsampling is not performed and is used to compute the cost function; and 3) a region not used to compute the cost function. The cost function between templates in template matching can be computed by using samples selected by subsampling in region 1) and samples in region 2).
[0758] In another embodiment, each partitioned region can be one of the following: 1) a region where subsampling is performed and 2) a region not used to compute the cost function. The cost function between templates in template matching can be computed using the samples selected by subsampling in 1).
[0759] ■ Subsampling for the search range The search process in template matching can be performed using a reference template included in a predetermined range starting from the sample point position indicated by the initial motion vector of the target block. In this case, all samples / positions within the search range can be used, or only some samples / positions within the search range can be selected. The search and / or matching cost calculation can be performed only for the selected samples / positions or the motion information indicating the selected samples / positions.
[0760] When performing template matching search processing by using only some samples / locations within the search range, subsampling can be performed on all or part of the search range.
[0761] When performing template matching search processing by using only some samples / locations within the search range, each search range can be divided into two or more regions.
[0762] In some embodiments, each segmented region can be one of the following: 1) a region where subsampling is performed; 2) a region where subsampling is not performed and search processing is performed; and 3) a region where search processing is not performed. Search processing in template matching can be performed for pixels and / or locations selected by subsampling in 1) and for samples / locations in the region of 2). Optionally, search processing in template matching can be performed for samples and / or locations selected by subsampling in 1) and for motion information indicating samples / locations in the region of 2).
[0763] In some other embodiments, each partitioned region can be one of 1) a region where subsampling is performed and 2) a region where no search processing is performed. Search processing in template matching can be performed using the sample points / locations selected by subsampling in 1). Alternatively, search processing in template matching can be performed using motion information indicating the sample points / locations selected by subsampling in 1).
[0764] Figure 19 and Figure 20 Various examples of subsampling methods in template matching are shown.
[0765] Figure 19 and Figure 20 The shaded sample (or location) represents the sample (or location) selected through subsampling.
[0766] In the template, it can be like this: Figure 20 The template shown performs subsampling on all or part of the reference area and can be configured by using only the sample points (or locations) selected through subsampling.
[0767] When calculating costs, it can be as follows: Figure 20 The template region can be subsampled in whole or in part, and cost calculation can be performed only on the selected sample points (or locations).
[0768] In addition, it is possible to Figure 19 The template matching performs subsampling on all or part of the search range, and can perform search and / or matching cost calculation only on selected pixels and / or locations or motion information indicating selected pixels and / or locations.
[0769] 1.3 Template Matching Search Method A first search step is performed by using first motion information encoded into or decoded from a bitstream as initial motion information, and second motion information, as a result of correcting the first motion information, can be derived. In a second search step performed after the first search step, the second motion information can be used as the initial motion information.
[0770] When the initial motion information (e.g., the initial motion vector or the initial block vector) is not motion information in integer pixel units (i.e., in the case of fractional pixel units), the search can be performed by using the result of rounding (or rounding down or rounding up) the initial motion information.
[0771] In the search process, to generate a reference template at a location indicated by motion information obtained by adding a specific offset to the initial motion information in fractional pixel units, samples at fractional pixel units should be generated by applying an interpolation filter to samples at integer pixel locations. However, when the initial motion information is limited to integer pixel units, interpolation for fractional pixel locations is not required in the search process, thus reducing complexity.
[0772] ■ Search Definition The search can be performed by using a cost function to determine the similarity between NUM_TEMPLATE_COMPARE templates.
[0773] The search may include the process of determining at least one piece of motion information that meets specific conditions within a specific search range. The motion information of the target block may be determined and / or modified based on the at least one piece of motion information determined through the search.
[0774] Motion information that meets specific conditions can be represented as motion information with the lowest matching cost among motion information within the search range, but is not limited to this.
[0775] ■ Cost Function The cost function used for cost calculation can be a function that determines the similarity between at least one sample in the target template and at least one sample in the reference template.
[0776] The similarity between a first value and a second value can be determined by using at least one of the following: 1) the difference between the two values, 2) the ratio between the two values, and 3) an operation that compares the difference between the two values with a specific value.
[0777] As an example of the operation of comparing the difference between two values with a specific value, a scheme can be used to assign multiple different similarity values based on which interval the difference between the two values belongs to, which is distinguished by multiple thresholds. For example, a scheme can be used where, when using a threshold, if the difference between the two values is equal to or less than the threshold, a first similarity value (e.g., 1) is assigned, otherwise a second similarity value (e.g., 0) is assigned.
[0778] The cost function can be at least one of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), and Sum of Squared Errors (SSE). However, the cost function is not limited to the items listed above.
[0779] The cost function used in template matching can be predefined or determined based on information sent / encoded / decoded by signals.
[0780] The cost function used in template matching can be determined based on at least one of the following: whether bilateral matching is performed, the conditions associated with bilateral matching, whether inter-frame weighted bidirectional prediction is performed, and the size of the target block.
[0781] In some embodiments, MR-SAD can be used as the cost function for template matching when the target block satisfies one or more of the enable conditions for bilateral matching described below, or when bilateral matching is performed in the target block.
[0782] In another embodiment, SAD can be used as the cost function for template matching when the target block does not meet one or more of the enable conditions for bilateral matching or a portion of those enable conditions, or when bilateral matching is not performed in the target block.
[0783] In another embodiment, the type of cost function in template matching can be determined based on specific conditions used to determine the type of cost function in bilateral matching. The type of cost function in bilateral matching can be determined based on whether said specific conditions are met in bilateral matching. In this case, the type of cost function in template matching can be determined based on whether the enabling conditions for bilateral matching are met and the specific conditions used to determine the type of cost function in bilateral matching.
[0784] For example, MR-SAD can be used as the cost function for template matching when the target block meets the enable conditions for bilateral matching and the specific conditions used to determine the type of cost function in bilateral matching; when the target block does not meet these two conditions, SAD can be used as the cost function for template matching.
[0785] In another embodiment, when the target block meets the enable conditions for bilateral matching and inter-frame weighted bidirectional prediction is performed, or when the number of samples in the target block is greater than a certain value, MR-SAD can be used as the cost function in template matching; otherwise, SAD can be used as the cost function in template matching.
[0786] ■ Search Scope (Search Area) The search area can be a specific range centered on the location indicated by the initial motion information. In other words, the center of the search area can be the location indicated by the initial motion information.
[0787] Optionally, the search range can be a specific area with the position indicated by the initial motion information as the upper left. In other words, the upper left of the search range can be the position indicated by the initial motion information.
[0788] Optionally, the search range may be an area including at least one of the following: the lower left, left side, upper left, top, and upper right of the target block, and the location of neighboring sample points.
[0789] The search area can be a rectangle with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search area can be a rhombus with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search area can be a hexagon formed by removing the lower right rectangular area from the rectangle. However, the shape and size of the search area are not limited to the above embodiments.
[0790] Each of SR_X and SR_Y can be a positive integer. Each of SR_X and SR_Y can be a predefined value or a value determined based on information sent / encoded / decoded using signals.
[0791] Initial motion information can be determined based on at least one of the following: motion information of the target block, encoding parameters of the target block, motion vector of the target block, reference image of the target block, block vector of the target block, motion vector predictor of the target block, block vector predictor of the target block, motion information of at least one neighboring block of the target block, merging candidate of the target block, motion vector difference of the target block, and block vector difference of the target block.
[0792] ■ Search Methods The search method can be defined based on at least one of the units of search method, search resolution, search range, initial motion information, and derived motion information.
[0793] The search method can be one of the following: diamond shape, cross shape, or full search. However, the search method is not limited to those listed above.
[0794] When (0, 0) represents the position indicated by the initial motion information, a diamond-shaped search can represent searching for at least one of the positions (0, 2×RR), (RR, RR), (2×RR, 0), (RR, -RR), (0, -RR), (-RR, -RR), (-RR, 0), (-RR, RR), and (0, 0).
[0795] When (0,0) represents the position indicated by the initial motion information, a cross-shaped search can represent at least one of the positions (0, RR), (RR, 0), (0, -RR), (-RR, 0), and (0, 0).
[0796] RR can be the search resolution or a value determined based on the search resolution, and can be a predefined positive number.
[0797] A full search can be used to search all locations within a predefined search range.
[0798] For example, when FS_i has values from -FS_X to FS_X, and FS_j has values from -FS_Y to FS_Y, a full search can be used to search for the position (FS_i × RR, FS_j × RR). Here, (0, 0) can be a position indicated by the initial motion information. However, the search range is not limited to the positions mentioned above. Each of FS_X and FS_Y can be a predefined positive number.
[0799] The search resolution can be one of 4-pixel (4-pel), full-pel, half-pel, or quarter-pel. However, the search resolution is not limited to the pel values mentioned above.
[0800] The search resolution can be predefined, determined based on at least one of the information about the adaptive motion vector resolution, or determined based on the values sent / encoded / decoded by the signal.
[0801] The unit for deriving information can include entire block units and sub-block units.
[0802] To determine the search method defined by the search mode, search resolution, etc., the following can be considered: motion information of the target block, encoding parameters of the target block, size of the target block, prediction mode of the target block, reference image of the target block, at least one sample value in the target block, target template, at least one sample value in the target template, and at least one region of the target template.
[0803] Figures 21 to 26 Each of the examples illustrates the search method in template matching based on the example.
[0804] Based on the target block's motion information, encoding parameters, prediction mode, and adaptive motion vector resolution, it can be used in... Figures 21 to 26 The table shown in the image specifies a particular column. A search can be performed from top to bottom of the specified column, using the search mode and search resolution corresponding to the row indicated by "v".
[0805] For example, in Figure 21 In the process, when the target block is in AMVP mode and the resolution determined by the adaptive motion vector resolution is 4-pel, a diamond-shaped search using the 4-pel search resolution can be performed, followed by a cross-shaped search using the 4-pel search resolution.
[0806] ALT_IF can represent the index of an adaptive interpolation filter. An interpolation filter can be applied to compute the pixel value at a sample location with a specific resolution. The adaptive interpolation filter can be one of several interpolation filters selected according to the index. In other words, when applying an adaptive interpolation filter, different interpolation filters can be used based on the index to compute the pixel value at a sample location with a specific resolution.
[0807] For example, a specific resolution can be half-PE. However, a specific resolution is not limited to half-PE.
[0808] For example, the interpolation filter determined by the index can be one of a 6-tap interpolation filter and an 8-tap interpolation filter. However, the method for determining the interpolation filter is not limited to the methods described above.
[0809] Figure 27 This is an exemplary diagram used to explain the method for configuring a first template (target template) in affine mode. Figure 28a and Figure 28b This is an exemplary diagram used to explain a method for configuring a second template (reference template) in affine mode.
[0810] CPMV can represent the affine control point motion vector (CPMV). The MV can be derived using CPMV on a per-sub-block basis within the target block.
[0811] Reference Figure 27 The target templates A0 to A3 and L0 to L3 can be composed of sub-blocks adjacent to the top and left side of the target block (target CU).
[0812] Reference Figure 28aIn an embodiment, the motion vector (or block vector) determined for each reference template position using the target block-based CPMV can be used to determine the reference templates A0 to A3 and L0 to L3 corresponding to the target template. For example, the reference templates can be used to configure sub-blocks at positions indicated by the motion vectors of the reference templates from the target template positions.
[0813] Reference Figure 28b In another embodiment, reference templates A0 to A3 and L0 to L4 corresponding to the target template can be determined using the motion vector (or block vector) of the sub-block adjacent to each target template in the target block. For example, the motion vector (or block vector) of the sub-block adjacent to the target template can be determined using the CPMV of the target block. The sub-block at the position indicated by the determined motion vector from the target template position can be configured as the reference template.
[0814] When the target block is in affine mode, it can be divided into sub-blocks of width N and height M. The motion information for each sub-block can be determined based on at least one of motion information, encoding parameters, and the size of the target block. The template matching cost for the target block can be determined based on at least one of the template matching costs for the divided sub-blocks. For example, the template matching cost for the target block can be the sum of the template matching costs of the divided sub-blocks, or the average of the template matching costs of the individual sub-blocks.
[0815] Each of N and M can be 2, 4, 8, or a positive integer.
[0816] Each of N and M can be a predefined value, or it can be a value determined based on information sent / encoded / decoded using signals.
[0817] 1.4 Template Matching in Bidirectional Prediction Blocks The motion information of a target block can be determined based on the motion information of neighboring blocks. This can be expressed as "the target block inherits motion information from its neighboring blocks".
[0818] For example, when the prediction mode of the target block is the merge mode, a merge candidate can be specified from the merge candidate list based on the merge index, and the motion information of the specified merge candidate can be used as the motion information of the target block.
[0819] When the prediction mode of the target block is AMVP mode, an MV candidate can be specified from the MV candidate list based on the MV candidate index, and the motion information of the specified MV candidate can be used as the motion information of the target block.
[0820] When a target block is bidirectionally predicted, for example when motion information inherited from neighboring blocks indicates bidirectional prediction, embodiments of performing template matching in the target block may include the following.
[0821] ■ Step 1: Perform template matching for each of directions L0 and L1, and calculate the template matching costs C0 and C1 for the motion information determined for directions L0 and L1.
[0822] At this point, when performing template matching for each direction, template matching can be performed similarly to performing template matching in unidirectional prediction for the corresponding direction, without considering motion information in the other direction.
[0823] When the predefined conditions are met for the target block, MR-SAD can be used as the cost function, and when the predefined conditions are not met, SAD can be used as the cost function.
[0824] Here, the predefined conditions can be based on at least one of the following: whether a local illumination compensation mode is performed in the target block; whether inter-frame weighted bidirectional prediction is performed in the target block; an indicator indicating whether a model-based prediction method is performed in the target block; motion information of the target block; whether bilateral matching is performed in the target block; an indicator indicating whether bilateral matching is performed in the target block; the size of the target block; the encoding parameters of the target block; motion information of the target block's neighboring blocks; the encoding parameters of the target block's neighboring blocks; and the type of cost function in template matching in the target block's neighboring blocks.
[0825] In some embodiments, MR-SAD can be used as the cost function when it is true whether a model-based prediction method is performed in the target block or when an indicator indicating whether a model-based prediction method is performed in the target block is true; otherwise, SAD can be used as the cost function.
[0826] In another embodiment, MR-SAD can be used as the cost function when the indicator indicating whether bilateral matching is performed in the target block is true, and the number of samples in the target block is equal to or greater than a specific value; otherwise, SAD can be used as the cost function.
[0827] The cost function may represent the cost function used when searching for template matching; and / or the cost function used to calculate at least one of C0, C1, and C'. C' will be described below in step 3.
[0828] The cost function used when the search template matches and the cost function used to calculate at least one of C0, C1, and C' can be the same or different. For example, MR-SAD can be used as the cost function when the search template matches, and SAD can be used as the cost function in the case of C0, C1, and C'.
[0829] In another embodiment, when the size of the target block is smaller than a specific value, SAD can be used as the cost function; otherwise, MR-SAD can be used as the cost function in template matching. The size of the block can include at least one of the width of the block, the height of the block, (the sum of the width and height of the block), and (the product of the width and height of the block).
[0830] In another embodiment, when BCW is not performed in the target block or the same weights are used for the reference blocks in the L0 direction and the reference blocks in the L1 direction in BCW, SAD can be used as the cost function; otherwise, MR-SAD can be used as the cost function in template matching.
[0831] In another embodiment, when the local illumination compensation (LIC) mode is not performed in the target block, SAD can be used as the cost function; otherwise, MR-SAD can be used as the cost function in template matching. The local illumination compensation mode can be a mode of deriving at least one of the weight and the offset by calculating the correlation between the template of the target block and the template of the reference block, and applying the derived weight or offset to part or all of the target block (or the reference block of the target block). Here, the weight represents a parameter for the multiplication operation of the target block and the reference block, and the offset represents a parameter for the addition operation of the target block and the reference block. In the local illumination compensation mode, the sample value P can be modified as follows.
[0832] P' = w ×P + o (where w represents the weight and o represents the offset) ■Step 2: In the case of C0 < C1, a new target template T' is generated by using the target template and the L0 direction template.
[0833] For example, when the target template is represented by T, the L0 direction reference template is represented by T0, and the L1 direction reference template is represented by T1, the new target template T' can be determined as follows: T' = w T ×T + w T0 ×T0 w T and w T0 Each of them can be a predefined value.
[0834] w T can be a positive number, and w T0 can be a negative number. w T and w T0 can be values determined based on whether inter-frame weighted bi-directional prediction is performed in the target block; and / or the weights in inter-frame weighted bi-directional prediction.
[0835] For example, w T can be 2, and wT0 It can be -1. When inter-frame weighted bi-prediction is not performed in the target block, or when the weights in the inter-frame weighted bi-prediction for directions L0 and L1 are the same, w T can be 2, and w T0 can be -1.
[0836] For example, when inter-frame weighted bi-prediction is performed, w T can be a value determined based on the weight for direction L1. w T0 can be a value determined based on the following value: the value obtained by dividing the weight for direction L0 by the weight for direction L1.
[0837] In the case of C0 > C1, a new target template T' can be generated by using the target template and the L1-direction template.
[0838] In the case where the values of C0 and C1 are equal to each other, this case can be regarded as a case of C0 < C1 or C1 > C0, and the process of step 2 can be performed.
[0839] ■ Step 3: In the case of C0 < C1, template matching is performed by setting the target template to T' for the L1 direction, and the template matching cost C' of the motion information determined for the L1 direction is calculated.
[0840] In the case of C0 > C1, template matching is performed by setting the target template to T' for the L0 direction, and the template matching cost C' of the motion information determined for the L0 direction is calculated.
[0841] In the case where the values of C0 and C1 are equal to each other, this case can be regarded as a case of C0 < C1 or C1 > C0, and the process of step 3 can be performed.
[0842] ■ Step 4: The motion information of the target block can be changed to unidirectional motion information based on at least one of C', C0, or C1.
[0843] For example, when the value obtained by multiplying C' by a predetermined value w C' is greater than the smaller value of C0 or C1 (or the value obtained by multiplying the smaller value by a predetermined value w CX ), the motion information of the target block can be changed to the motion information indicating unidirectional prediction in the L0 direction or the L1 direction. Here, the result of multiplying the two values can represent A × B, or one of the rounded, ceiling, or floor values of A × B.
[0844] For example, when C0 < C1, the motion information of the target block can be changed to motion information indicating unidirectional prediction in the L0 direction. Optionally, it can be considered that the motion information for the L1 direction is not available in the target block.
[0845] When C0 > C1, the motion information of the target block can be changed to motion information indicating unidirectional prediction in the L1 direction. Optionally, it can be considered that the motion information for the L0 direction is not available in the target block.
[0846] When the values of C0 and C1 are equal to each other, this case can be regarded as a case of C0 < C1 or C1 > C0, and the process of step 4 can be carried out.
[0847] w C' and w CX Each of them can be a predefined value.
[0848] w C' and w CX can be a value determined based on whether inter-frame weighted bi-directional prediction is performed in the target block; and / or the weight in inter-frame weighted bi-directional prediction.
[0849] For example, w C' can be 1 / 2, and w CX can be 9 / 8.
[0850] For example, w C' can be a value determined based on the L0 direction or the L1 direction in inter-frame weighted bi-directional prediction. When C0 < C1, w C' can be the L1 direction weight in inter-frame weighted bi-directional prediction or a value determined based on the L1 direction weight. When C1 < C0, w C' can be the L0 direction weight in inter-frame weighted bi-directional prediction or a value determined based on the L0 direction weight.
[0851] Only when the target block meets the predefined conditions can the above steps 2 to 4 be executed.
[0852] For example, only when the target block is bi-directionally predicted, bilateral matching is not performed in the target block, or the target block does not meet the enabling conditions for bilateral matching can the above steps 2 to 4 be executed.
[0853] 2. Bilateral matching In bilateral matching, the L0 direction reference block and the L1 direction reference block can be used as templates, and the motion information of the target block can be determined and / or changed based on the calculation result of the cost function between the two templates.
[0854] A reference block may include at least one of the following: 1) a reference block indicated by initial motion information; 2) a reference block indicated by motion information derived in the bilateral matching search process; and 3) a reference block indicated by motion information finally corrected by bilateral matching.
[0855] For example, when configuring a template for bilateral matching, the L0 direction reference block and the L1 direction reference block can be used as templates.
[0856] The bilateral matching cost can be represented as the result of a cost function calculated using templates of the L0-direction reference block and the L1-direction reference block used in the bilateral matching.
[0857] When the prediction mode of the target block is IBC mode, and prediction is performed using two or more reference blocks, bilateral matching can be performed by using two different reference blocks from the target block's reference blocks as templates.
[0858] The following describes bilateral prediction in inter-frame prediction other than IBC mode; however, the technical features of bilateral prediction in inter-frame prediction can be similarly applied to IBC mode. In this case, the L0 and L1 direction reference blocks can be replaced with two reference blocks generated in IBC mode.
[0859] 2-1. Subsampling in bilateral matching ■ Subsampling for configuring templates When configuring a template for bilateral matching, you can select only some pixels and / or positions within the L0 and L1 direction reference blocks. The template can be configured using only the selected pixels and / or selected positions.
[0860] The template used for bilateral matching can represent at least one of the L0 direction template and the L1 direction template.
[0861] As some embodiments, when configuring a template for bilateral matching, subsampling for the L0 direction reference block and the L1 direction reference block can be used.
[0862] As another embodiment, when configuring a template for bilateral matching, subsampling can be used for a portion of the L0 direction reference block and a portion of the L1 direction reference block.
[0863] For example, when configuring a template for bilateral matching, each of the L0 direction reference block and the L1 direction reference block can be divided into two or more regions. Each divided region can be one of the following: 1) a region where subsampling is performed; 2) a region where subsampling is not performed and is used to configure the template; and 3) a region not used to configure the template. The template for bilateral matching can be configured using pixels and / or positions selected by subsampling in 1) and pixels / positions in the region in 2).
[0864] Optionally, each partitioned region can be one of 1) a region where subsampling is performed and 2) a region not used for template configuration. The template for bilateral matching can be configured by using pixels and / or selected locations in 1) through subsampling.
[0865] As another embodiment, when configuring a template for bilateral matching, the area used to configure the template can be a portion of the L0 direction reference block and a portion of the L1 direction reference block.
[0866] As another embodiment, when configuring a template for bilateral matching, pixels (or positions) for configuring the template can be selected only in a portion of the L0 direction reference block and a portion of the L1 direction reference block.
[0867] The size of a portion of the L0 direction reference block can be smaller than the size of the L1 direction reference block.
[0868] For example, the height (vertical dimension) of a portion of the L0 direction reference block can be smaller than the height (vertical dimension) of a portion of the L1 direction reference block.
[0869] For example, the width (horizontal dimension) of a portion of the L0 direction reference block can be smaller than the width (horizontal dimension) of a portion of the L1 direction reference block.
[0870] The size of a portion of the L1 direction reference block can be smaller than the size of the L0 direction reference block.
[0871] For example, the height (vertical dimension) of a portion of the L1 direction reference block can be smaller than the height (vertical dimension) of the L0 direction reference block.
[0872] For example, the width (horizontal dimension) of a portion of the L1 direction reference block can be smaller than the width (horizontal dimension) of the L0 direction reference block.
[0873] ■ Subsampling for cost calculation In the step of performing cost calculation between templates in bilateral matching, only some pixels and / or locations within the template region can be selected. Cost calculation can be performed only for the selected pixels and / or selected locations.
[0874] In some embodiments, when performing cost calculation between templates in bilateral matching, subsampling can be performed on template regions in the L0 direction and template regions in the L1 direction.
[0875] In another embodiment, when performing cost calculation between templates in bilateral matching, subsampling can be performed on a portion of the template region in the L0 direction and a portion of the template region in the L1 direction.
[0876] Optionally, when performing cost calculation between templates in bilateral matching, each of the L0 direction template region and the L1 direction template region can be divided into two or more regions. Each divided region can be one of the following: 1) a region where subsampling is performed; 2) a region where subsampling is not performed and is used for cost calculation; and 3) a region not used for cost calculation. Cost calculation between templates in bilateral matching can be configured by using pixels and / or positions selected by subsampling in 1) and pixels / positions in the region in 2).
[0877] Optionally, each partitioned region can be one of the following: 1) the region where subsampling is performed and 2) the region not used to compute the cost function. The computation of the cost function between templates in bilateral matching can be configured by using the pixels and / or locations selected by subsampling in 1).
[0878] In another embodiment, when performing cost calculation between templates in bilateral matching, the region used for cost calculation can be a portion of the template in the L0 direction and a portion of the template in the L1 direction.
[0879] The dimensions of a portion of the template in the L0 direction and the dimensions of a portion of the template in the L1 direction can be different from each other.
[0880] As an example, the size of a portion of the template in the L0 direction can be smaller than the size of the portion of the template in the L1 direction.
[0881] For example, the height (vertical dimension) of a portion of the template in the L0 direction can be smaller than the height (vertical dimension) of the template in the L1 direction.
[0882] For example, the width (horizontal dimension) of a portion of the template in the L0 direction can be smaller than the width (horizontal dimension) of a block in the template in the L1 direction.
[0883] As another example, the size of a portion of the template in the L1 direction can be smaller than the size of the portion of the template in the L0 direction.
[0884] For example, the height (vertical dimension) of a portion of the template in the L1 direction can be less than the height (vertical dimension) of the template in the L0 direction.
[0885] For example, the width (horizontal dimension) of a portion of the template in the L1 direction can be smaller than the width (horizontal dimension) of the template in the L0 direction.
[0886] ■ Subsampling for the search range For example, when performing a bilateral matching search process, only some pixels and / or locations within the search range may be selected. The search and / or matching cost may be calculated only for the selected pixels and / or selected locations. Alternatively, the search and / or matching cost may be calculated only for motion information indicating the selected pixels and / or locations.
[0887] For example, when performing a bilateral matching search process, subsampling of all or part of the search range can be performed.
[0888] Optionally, for example, when performing bilateral matching search processing, each search range can be divided into two or more regions. Each divided region can be one of the following: 1) a region where subsampling is performed; 2) a region where subsampling is not performed but search processing is performed; 3) a region where search processing is not performed. Bilateral matching search processing can be performed for pixels and / or positions selected by subsampling in 1) and pixels / positions in the region of 2). Optionally, bilateral matching search processing can be performed for motion information indicating pixels and / or positions selected by subsampling in 1) and pixels / positions in the region of 2).
[0889] Optionally, for example, when performing bilateral matching search processing, each search area can be divided into two or more regions. Each divided region can be one of the following: 1) a region where subsampling is performed and 2) a region where search processing is not performed. Bilateral matching search processing can be performed using pixels and / or selected locations in 1) through subsampling. Optionally, bilateral matching search processing can be performed with respect to motion information indicating pixels and / or selected locations in 1).
[0890] Figure 19 Various examples of subsampling methods in bilateral matching are shown in the figure.
[0891] Figure 19 and Figure 20 The shaded sample (or location) can represent the sample (or location) selected through subsampling.
[0892] like Figure 19 As shown, subsampling can be performed on the L0 direction reference block region and the L1 direction reference block region, and the template can be configured by using only the selected sample points (or locations).
[0893] like Figure 19As shown, subsampling can be performed on a portion of the L0 direction reference block region and a portion of the L1 direction reference block region, and the template can be configured by using only the selected sample points (or locations).
[0894] Can Figure 19 The diagram shows that subsampling is performed on all or part of the template region for bilateral matching, and cost calculation can be performed only for selected sample points (or locations).
[0895] Can Figure 19 The example shows that subsampling is performed on all or part of the search range for bilateral matching, and the search and / or matching costs can be calculated only for selected pixels and / or locations.
[0896] Can Figure 19 The example shown performs subsampling on all or part of the search range for bilateral matching, and can perform search and / or matching cost calculations only on motion information indicating the selected pixels and / or locations.
[0897] 2-2. Subsampling in bilateral matching Bilateral matching can be performed continuously. Alternatively, bilateral matching can be performed only if predefined enabling conditions are met.
[0898] For example, bilateral matching can be performed when inter-frame prediction mode is used for the target block and two or more reference blocks are used.
[0899] For example, bilateral matching can be performed when the first direction and the second direction are different from each other and the first POC interval and the second POC interval are the same. The first direction can be a direction from the target image toward the L0 direction reference image. The second direction can be a direction from the target image toward the L1 direction reference image. The first POC interval can be the difference between the POC of the target image and the POC of the L0 direction reference image. The second POC interval can be the difference between the POC of the target image and the POC of the L1 direction reference image.
[0900] For example, bilateral matching can only be performed if the first direction and the second direction are different from each other. The first direction can be the direction from the target image toward the L0 direction reference image. The second direction can be the direction from the target image toward the L1 direction reference image.
[0901] Here, the fact that the first direction and the second direction are different from each other can be expressed as satisfying the following equation 1.
[0902] [Equation 1] (POCt-POC0)×(POCt-POC1)<0 Here, the fact that the first direction and the second direction are the same can be expressed as satisfying the following equation 2.
[0903] [Equation 2] (POCt-POC0)×(POCt-POC1)>0 In Equations 1 and 2, POCt represents the POC of the target image, POC0 represents the POC of the reference image in the L0 direction, and POC1 represents the POC of the reference image in the L1 direction.
[0904] 2-3. Two-sided matching search method ■ Definition of Search The search can be performed by using a cost function to determine the similarity between two templates.
[0905] The search may include the process of determining at least one piece of motion information that meets specific conditions within a specific search range. The motion information of the target block may be determined and / or modified based on the at least one piece of motion information determined by the search.
[0906] For example, motion information that meets certain conditions can represent the motion information with the lowest matching cost among the motion information in the search range, but it is not limited to this.
[0907] A search may include the process of identifying at least one block within a specific search range that satisfies specific conditions. Motion information of the blocks identified through the search can be used as motion information for the target block.
[0908] For example, a block that meets certain conditions can represent motion information with the lowest matching cost among reference blocks in the search range, but is not limited to this.
[0909] ■ Cost Function The cost function can represent a function that determines the similarity between at least one sample in a first template and at least one sample in a second template. For example, the cost function can be a function that determines the similarity between at least one sample in the first template and its corresponding sample in the second template.
[0910] The similarity between a first value and a second value can be determined by using at least one of the following: 1) the difference between the two values, 2) the ratio between the two values, and 3) an operation that compares the difference between the two values with a specific value.
[0911] The cost function can be at least one of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), and Sum of Squared Errors (SSE). However, the cost function is not limited to the items listed above.
[0912] The cost function used in bilateral matching can be predefined or determined based on information sent / encoded / decoded by signals.
[0913] As an example, SAD can be used as the cost function when the size of the target block is smaller than a certain value; otherwise, MR-SAD can be used as the cost function in bilateral matching. The size of the block can include at least one of the following: the width of the block, the height of the block, (the sum of the width and height of the block), and (the product of the width and height of the block).
[0914] As another example, SAD or SATD can be used as the cost function when BCW is not performed in the target block or when the same weights are used for the L0 and L1 reference blocks in BCW; otherwise, MR-SAD or MRSATD can be used as the cost function in template matching.
[0915] As another example, SAD can be used as the cost function when Local Illumination Compensation (LIC) mode is not performed in the target block; otherwise, MR-SAD can be used as the cost function in bilateral matching.
[0916] 2-4. Search steps for bilateral matching When the initial motion information (e.g., the initial motion vector or the initial block vector) is not motion information in integer pixel units (i.e., in the case of fractional pixel units), the initial motion information can be modified by performing rounding (or rounding down or rounding up) on the initial motion information, and the search can be performed by using the modified initial motion vector.
[0917] For example, in the search process, in order to generate a template at a location indicated by motion information obtained by adding a specific offset to the initial motion information in fractional pixel units, samples at fractional pixel units should be generated by applying an interpolation filter to samples at integer pixel locations. However, when the initial motion information is limited to integer pixel units, interpolation for fractional pixel locations does not need to be performed in the search process, thus reducing complexity.
[0918] Two-sided matching can include one or more search steps.
[0919] For example, bilateral matching can be configured to sequentially include steps of 1) deriving motion information for the entire block and 2) deriving motion information for sub-blocks of the block. However, the methods for deriving motion information executed in each step and the order of the steps are not limited to the configuration described above.
[0920] In each search step of bilateral matching, the motion information in BM_NUM directions can be corrected between the motion information in the L0 direction and the motion information in the L1 direction. BM_NUM can be 0, 1, 2, or a positive integer. The BM_NUM used in the steps of bilateral matching can be the same or different from each other.
[0921] For example, when BM_NUM is 1 in a specific search step, it can be used only for LX in that specific search step. BM The motion information is corrected using directional motion information. Here, X... BM It can have a value of 0 or 1. That is, when X BM When L0 is 0, set the direction of L0 to LX. BM Direction, when X BM When the value is 1, the L1 direction is set to LX. BM direction.
[0922] For example, when BM_NUM is 1 and X is 1 in a particular search step BM When the value is 0, in a specific search step, a search can be performed only for the L0 direction while keeping the L1 direction template and L1 direction motion information fixed.
[0923] Signals can be used to send / encode / decode information about X. BM Information.
[0924] Optionally, X BM It can be predefined.
[0925] For example, X BM It can be set to a value in either the L0 or L1 direction that has a larger POC difference from the target image. The POC difference can be the difference between the POC of the reference image and the POC of the target image (in a specific direction).
[0926] Optionally or otherwise, X BM It can be set to the direction with a higher template matching cost in the L0 and L1 directions. Here, the template matching cost of a specific direction can be the template matching cost of motion information in that specific direction.
[0927] The determination of X will be described below. BM Various embodiments.
[0928] ■ X BM Determination method In the embodiment, when the first POC difference is greater than the second POC difference, X BM It can be 0, otherwise, X BM It can be 1. Optionally, when the first POC difference is greater than the second POC difference, X BM It can be 1, otherwise, X BM It can be 0. The first POC difference can be the difference between the POC of the target image and the POC of the reference image in the L0 direction. The second POC difference can be the difference between the POC of the target image and the POC of the reference image in the L1 direction.
[0929] In another embodiment, X can be determined based on a context model and / or a probabilistic model used for entropy encoding and entropy decoding of motion information and encoding parameters of the target block. BM .
[0930] For example, X can be determined based on at least one of the context model and / or probability model used when entropy encoding and entropy decoding inter-prediction indicators that indicate the direction of inter-prediction for the target block. BM .
[0931] Based on the context model and / or probability model used when entropy coding and entropy decoding the inter-frame prediction indicator, the more probable direction in the target block can be selected as LX from the L0-direction unidirectional prediction and L1-direction unidirectional prediction. BM Direction (i.e., the direction of the motion information being corrected).
[0932] Optionally, based on the context model and / or probability model used when entropy coding and entropy decoding the inter-frame prediction indicator, the more probable direction in the target block can be selected as L(1-X) from the L0-direction unidirectional prediction and L1-direction unidirectional prediction. BM ) direction (i.e., the direction in which motion information is not corrected).
[0933] A more probable direction may represent a direction that uses fewer bits when entropy encoding is performed using a context model and / or a probabilistic model. Alternatively, a more probable direction may represent a direction with a higher probability indicated by the context model and / or the probabilistic model for that direction.
[0934] As another embodiment, LX can be determined based on the weights in the inter-frame weighted bidirectional prediction of the target block. BM For example, LX BM It can be the direction with higher weight between the L0 and L1 directions. Optionally, LX BM It can be a direction with a lower weight between the L0 and L1 directions.
[0935] As another embodiment, X can be determined based on at least one of the motion information of neighboring blocks and encoding parameters. BM For example, it can be based on... Figure 11 The motion information of at least one of the corresponding neighboring blocks of A0, A1, B0, B1 and B2 and at least one of the encoding parameters are used to determine X in the target block.
[0936] For example, the X of the target block can be determined based on at least one of the inter-frame prediction indicator of neighboring blocks and the inter-frame bidirectional prediction weights. BM .
[0937] For example, one or more context models and / or probability models can be used for X BMEntropy encoding and entropy decoding. In multiple context models and / or probabilistic models, the entropy encoding information for the target block can be determined based on at least one of the motion information and encoding information of neighboring blocks. BM Contextual and / or probabilistic models for entropy encoding and entropy decoding.
[0938] For example, in a block used for X BM The context model and / or probability model used for entropy encoding and entropy decoding can be the same as each other. Optionally, the block used for X BM The context model and / or probability model for entropy coding and entropy decoding can vary based on at least one of the inter-frame prediction direction of neighboring blocks and the inter-frame bidirectional prediction weights.
[0939] The same X can be used in the search step of bilateral matching. BM Alternatively, different X values can be used in the search steps of bilateral matching. BM .
[0940] The following text will describe BM_NUM, which represents the number of directions in which motion information in the L0 and L1 directions is corrected.
[0941] ■ BM_NUM BM_NUM can be 0, 1, 2 or a positive integer.
[0942] BM_NUM can be predefined.
[0943] BM_NUM in each search step of bilateral matching can be determined based on encoding parameters. Alternatively, BM_NUM can be determined based on at least one of motion information, the search step of bilateral matching, the matching cost in the previous search step, the matching cost of the initial motion information of the current search step, and BM_NUM in the previous search step.
[0944] In an embodiment, BM_NUM can be 1 or 2 in the first search step of bilateral matching.
[0945] In an embodiment, BM_NUM in the current search step can be determined based on the matching cost in the previous search step.
[0946] For example, BM_NUM in the current search step can be 0 when the difference between the matching cost for the initial motion information in the previous search step and the matching cost for the corrected motion information in the previous search step is less than COSTDIFF_FORBMNUM.
[0947] COSTDIFF_FORBMNUM can be 0, 1, 2, 4, 8, 16 or a positive integer.
[0948] COSTDIFF_FORBMNUM can be determined based on the size of the target block. COSTDIFF_FORBMNUM can be the product of the number of pixels in the target block and a specific value. This specific value can be 0, 1, 2, 4, 8, or a positive integer.
[0949] In an embodiment, when BM_NUM was 0 in a previous search step, BM_NUM in the current search step can be 0.
[0950] In an embodiment, when the matching cost of the initial motion information in the current search step is less than COSTDIFF_FORBMNUM_INIT, the BM_NUM of the target block can be 0.
[0951] COSTDIFF_FORBMNUM_INIT can be 0, 1, 2, 4, 8, 16 or a positive integer.
[0952] For example, COSTDIFF_FORBMNUM_INIT can be determined based on the size of the target block. COSTDIFF_FORBMNUM_INIT can be the product of the number of pixels in the target block and a specific value. This specific value can be 0, 1, 2, 4, 8, or a positive integer.
[0953] In the embodiment, when motion correction is performed in the search step of bilateral matching, motion information correction for L0 direction motion information can only be performed if the matching cost of L0 direction motion information of the initial motion information in the current search step is greater than COSTDIFF_FORBMNUM_INIT.
[0954] In some embodiments, when motion correction is performed in the search step of bilateral matching, motion information correction for L1 direction motion information can only be performed if the matching cost of L1 direction motion information of the initial motion information in the current search step is greater than COSTDIFF_FORBMNUM_INIT.
[0955] A BM_NUM value of 0 in a specific search step of a bilateral matching operation indicates that motion information correction is not performed in that specific search step. Alternatively, a BM_NUM value of 0 in a specific search step of a bilateral matching operation indicates that that specific search step is not performed.
[0956] For example, when performing bilateral matching, BM_NUM can be 1 in the step of deriving the motion information of the entire block, and BM_NUM can be 2 in the step of deriving the motion information of the sub-blocks. In this case, only LX can be corrected in the step of deriving the motion information of the entire block. BMThe motion information of the direction can be corrected, and in the step of deriving the motion information of the sub-block, both the motion information of the L0 direction and the motion information of the L1 direction can be corrected.
[0957] Figure 29 This is a diagram used to explain bilateral matching according to embodiments of the present disclosure.
[0958] exist Figure 29 The diagram illustrates the case where BM_NUM is 2 during the derivation steps of the motion information for the entire block in bilateral matching.
[0959] MV0 represents the initial motion information in the L0 direction, and MV1 represents the initial motion information in the L1 direction.
[0960] MV diff It can represent the motion information correction value derived through bilateral matching.
[0961] MV0' and MV1' can be motion information derived through bilateral matching.
[0962] In bilateral matching, the magnitudes of the motion information correction values for the L0 direction and the L1 direction can be equal. The directions of the motion information correction values for the L0 direction and the L1 direction can be opposite. That is, equations 3 and 4 below can be established.
[0963] [Equation 3] MV0' = MV0 + MV diff [Equation 4] MV1' = MV1 - MV diff For example, when the target block is in IBC mode and two or more reference blocks are used, during the bilateral matching process in the target block, the values and directions of the motion information correction for the first reference block and the values and directions of the motion information correction for the L1 direction can be the same.
[0964] Subsampling in template matching or bidirectional prediction as described above can be performed based on at least one of the following: whether template matching or bidirectional prediction is performed; an indicator indicating whether template matching or bidirectional prediction is performed; motion information of the target block; encoding parameters of the target block; and a search step in template matching.
[0965] As an example, the subsampling method can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; motion information of the target block; encoding parameters of the target block; and search steps in template matching.
[0966] The subsampling method used in template matching or bilateral matching can be the same for each search step. Alternatively, the subsampling method used when performing template matching and / or bilateral matching can be different for each step.
[0967] As another example, whether to perform subsampling can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; motion information of the target block; encoding parameters of the target block; and search steps in template matching.
[0968] When performing template matching, whether or not subsampling is performed can be the same for each search step. Alternatively, whether or not subsampling is performed can be different for each search step.
[0969] Whether to perform subsampling in the horizontal direction and whether to perform subsampling in the vertical direction can be the same. Alternatively, whether to perform subsampling in the horizontal direction and whether to perform subsampling in the vertical direction can be different.
[0970] 3. Derivation method of motion information on the decoder side In some embodiments, the decoder-side motion information derivation method may include methods such as template matching, bilateral matching, etc. For example, motion information can be generated by template matching, or the already determined first motion information (or initial motion information) can be corrected by template matching or bilateral matching. In other words, the decoder-side motion information derivation method may represent a method for deducing second motion information by performing corrections on the first motion information.
[0971] A predicted block for the target block can be generated by performing a prediction using the derived second motion information. Alternatively, at least one of the reference blocks for the target block can be specified using the second motion information.
[0972] In this embodiment, "correction" of specific information can mean modifying, correcting, or updating the specific information. In this embodiment, the terms "correction," "modification," and "correction" are used interchangeably. Corrected information can be generated by correcting specific information.
[0973] Modifying specific motion information may mean performing at least one of the following methods on the corresponding motion information.
[0974] - By using a predetermined offset, specific information included in the corresponding motion information can be changed. - Specific information included in the corresponding motion information is altered by performing specific operations on specific information and a predetermined offset. Here, the specific operations include at least one of squaring, weighted averaging, weighted summing, arithmetic operations, and filtering.
[0975] The predetermined offset may include at least one of the following: motion vector, reference frame index and inter-frame prediction indicator, reference frame list information, reference image, motion vector candidate, motion vector candidate index, merge candidate and merge index, block vector, block vector candidate and block vector candidate index.
[0976] A predetermined offset can be specified from an offset candidate list, and information used to specify the predetermined offset can be sent / encoded / decoded using signals. For example, the index of at least one offset can be specified from the predetermined offset candidate list using signals.
[0977] In another embodiment, the decoder-side motion information derivation method can be a method of reordering or reconfiguring the motion information list by using matching costs.
[0978] Based on the matchi...
Claims
1. An image decoding method for predicting target blocks included in a target image, comprising: Determine the prediction pattern for the target block from among multiple block vector-based prediction patterns; Determine a block vector resolution for the target block, wherein the block vector resolution is determined from a plurality of block vector resolution candidates, including one or more integer unit resolutions and one or more fractional unit resolutions; Based on the determined prediction pattern and the block vector resolution, at least one block vector for the target block is derived; and The predicted block of the target block is generated from pre-reconstructed samples in the target image using at least one block vector.
2. The image decoding method according to claim 1, wherein, When the determined mode is an intra-block copy (IBC) mode based on sub-blocks, the steps for deriving the at least one block vector include: The target block is divided into multiple sub-blocks; Derive motion information for a specified reference block, which will be referenced to at least one block vector used to derive the target block; Based on the motion information, the reference block is determined in the target image; and A block vector corresponding to a sub-block in the target block is generated by using the block vector of a sub-block in the reference block. The predicted block of the target block is generated on a sub-block basis using the block vectors of each sub-block in the target block.
3. The image decoding method as described in claim 2, wherein, The motion information is derived based on the block vectors of at least one neighboring block of the target block.
4. The image decoding method as described in claim 2, wherein, The motion information is derived based on at least one of a plurality of candidates included in the block vector candidate list of the target block.
5. The image decoding method as described in claim 4, wherein, Based on the template matching costs of the multiple candidates, the motion information is generated according to the candidate with the lowest template matching cost.
6. An image coding method for predicting target blocks included in a target image, comprising: Determine the prediction pattern for the target block from among multiple block vector-based prediction patterns; Determine a block vector resolution for the target block, wherein the block vector resolution is determined from a plurality of block vector resolution candidates, including one or more integer unit resolutions and one or more fractional unit resolutions; Based on the determined prediction pattern and the block vector resolution, at least one block vector for the target block is derived; and The predicted block of the target block is generated from pre-reconstructed samples in the target image using at least one block vector.
7. The image encoding method according to claim 6, wherein, When the determined mode is an intra-block copy (IBC) mode based on sub-blocks, the steps for deriving the at least one block vector include: The target block is divided into multiple sub-blocks; Derive motion information for a specified reference block, which will be referenced to at least one block vector used to derive the target block; Based on the motion information, the reference block is determined in the target image; and A block vector corresponding to a sub-block in the target block is generated by using the block vector of a sub-block in the reference block. The predicted block of the target block is generated on a sub-block basis by using the block vectors of each sub-block in the target block.
8. The image encoding method as described in claim 7, wherein, The motion information is derived based on the block vectors of at least one neighboring block of the target block.
9. The image encoding method as described in claim 7, wherein, The motion information is derived based on at least one of a plurality of candidates included in the block vector candidate list of the target block.
10. The image encoding method as described in claim 9, wherein, Based on the template matching costs of the multiple candidates, the motion information is generated according to the candidate with the lowest template matching cost.
11. A method for transmitting a bitstream comprising encoded video data, the method comprising: A bitstream is generated by predicting and encoding target blocks included in the target frame; and Send the bitstream to the image decoding device. The steps for generating the bitstream include: Determine the prediction pattern for the target block from among multiple block vector-based prediction patterns; Determine a block vector resolution for the target block, wherein the block vector resolution is determined from a plurality of block vector resolution candidates, including one or more integer unit resolutions and one or more fractional unit resolutions; Based on the determined prediction pattern and the block vector resolution, at least one block vector for the target block is derived; and The predicted block of the target block is generated from pre-reconstructed samples in the target image using at least one block vector.
Citation Information
Patent Citations
Branched amino acid surfactants for oil and gas production
KR1020230049631A
Meat processing method using mushroom extract and meat processing products using the same
KR1020230083944A
Method for transmitting multichannel audio signal
KR1020240048361A