Image encoding / decoding method and apparatus
By optimizing the video coding method and using transform type and combination to determine the reduced transform matrix set, the quality limitations of high-resolution and high-quality images are solved, achieving more efficient encoding and decoding, improving image quality and reducing transmission and storage costs.
Patent Information
- Application Number
- CN202511419896.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-13
- Filing Date
- 2020-11-26
- Publication Date
- 2026-01-13
AI Technical Summary
Existing video coding techniques have limitations in both objective and subjective image quality during high-resolution and high-quality image processing, especially in transform/inverse transform methods due to inefficiency caused by single transform type and signaling overhead.
The reduced set of secondary transform/inverse transform matrices and whether to perform the reduced secondary transform/inverse transform are determined based on whether a transform is used, the type of one-dimensional transform, and the combination of two-dimensional transforms. This is combined with intra-frame prediction mode, prediction mode, color components, and size to optimize the image encoding/decoding process.
It improves both the objective and subjective quality of images, reduces the amount of data, and lowers transmission and storage costs.
Smart Images

Figure CN121334366A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an image encoding / decoding method and apparatus and a recording medium for storing a bitstream. More particularly, the present application relates to a method and apparatus for encoding / decoding a video image based on a transform. BACKGROUND
[0002] Recently, the demand for high-resolution and high-quality images such as high-definition (HD) or ultra-high-definition (UHD) images has increased in various applications. As the resolution and quality of images increase, the amount of data increases accordingly. This is one of the reasons for an increase in transmission and storage costs when image data is transmitted through existing transmission media such as wired or wireless broadband channels or when image data is stored. In order to solve these problems of high-resolution and high-quality image data, an efficient image encoding / decoding technique is required.
[0003] There are various video compression techniques such as an inter prediction technique of predicting a value of a pixel in a current picture from values of pixels in a previous picture or a subsequent picture, an intra prediction technique of predicting a value of a pixel in a region of a current picture from values of pixels in another region of the current picture, a transform and quantization technique of compressing energy of a residual signal, and an entropy encoding technique of assigning a shorter code to a frequently occurring pixel value and a longer code to a less frequently occurring pixel value.
[0004] In a conventional transform / inverse transform method, there are limitations in both objective and subjective quality of an image due to the use of a single transform / inverse transform type or overhead for signaling of various transform / inverse transform types. SUMMARY
[0005] TECHNICAL PROBLEM
[0006] In order to improve the objective and subjective quality of an image, the present application provides a video encoding / decoding method and apparatus, in which at least one of a reduced secondary transform / inverse transform matrix set, a reduced secondary transform / inverse transform matrix, and whether to perform a reduced secondary transform / inverse transform is determined based on whether a transform is used, a one-dimensional transform type, and at least one of a two-dimensional transform combination, in which whether the transform is used, the one-dimensional transform type, and the two-dimensional transform combination are further based on an intra prediction mode, a prediction mode, a color component, a size, and a form.
[0007] TECHNICAL SOLUTION
[0008] The present disclosure provides a video decoding method, comprising: obtaining a transform skip mode flag indicating whether transform / inverse transform is skipped in a current block; determining that secondary transform / inverse transform is skipped in the current block when transform / inverse transform is skipped in the current block according to the transform skip mode flag; and obtaining a transform matrix index of secondary transform / inverse transform for the current block when transform / inverse transform is not skipped in the current block according to the transform skip mode flag; and determining whether secondary transform / inverse transform is skipped in the current block based on the transform matrix index.
[0009] According to an embodiment, the step of obtaining the transform skip mode flag comprises: obtaining a transform skip mode flag for a luma component, a transform skip mode flag for a Cb component, and a transform skip mode flag for a Cr component.
[0010] According to an embodiment, the step of determining whether secondary transform / inverse transform is skipped in the current block comprises: when a tree structure of the current block is a single tree type, the transform skip mode flag for the luma component indicates that transform skip mode is applied to the luma component, the transform skip mode flag for the Cb component indicates that transform skip mode is applied to the Cb component, and the transform skip mode flag for the Cr component indicates that transform skip mode is applied to the Cr component, determining that secondary transform / inverse transform is skipped in the current block.
[0011] According to an embodiment, the step of determining that secondary transform / inverse transform is skipped in the current block comprises: when the tree structure of the current block is a dual tree luma type, and the transform skip mode flag for the luma component indicates that transform skip mode is applied to the luma component, determining that secondary transform / inverse transform is skipped in the current block.
[0012] According to an embodiment, the step of determining that secondary transform / inverse transform is skipped in the current block comprises: when the tree structure of the current block is a dual tree chroma type, the transform skip mode flag for the Cb component indicates that transform skip mode is applied to the Cb component, and the transform skip mode flag for the Cr component indicates that transform skip mode is applied to the Cr component, determining that secondary transform / inverse transform is skipped in the current block.
[0013] According to an embodiment, the video decoding method can further comprise: when secondary transform / inverse transform is applied to the current block, determining a secondary transform matrix of the current block according to the transform matrix index, and applying secondary transform / inverse transform to the current block according to the secondary transform matrix.
[0014] According to an embodiment, the step of determining the secondary transform matrix of the current block comprises:
[0015] The secondary transform matrix for the current block is determined according to at least one of the transform matrix index, the transform matrix set index of the current block, and the size of the current block.
[0016] According to an embodiment, the video decoding method can further comprise obtaining information on whether an intra residual DPCM method is used, and determining that a transform / inverse transform is skipped in the current block when the information on whether the intra residual DPCM method is used indicates that the intra residual DPCM method is used for the current block, wherein the step of obtaining the transform skip mode flag comprises obtaining the transform skip mode flag when the information on whether the intra residual DPCM method is used indicates that the intra residual DPCM method is not used for the current block.
[0017] According to an embodiment, the transform matrix index of the secondary transform / inverse transform for the current block is obtained when the current block is predicted according to an intra prediction mode that is not a matrix-based intra prediction mode.
[0018] According to an embodiment, the step of determining whether the secondary transform / inverse transform is skipped in the current block according to the transform matrix index comprises determining whether the secondary transform / inverse transform is skipped in the current block according to at least one of the transform matrix index, the size of the current block, and the transform skip mode flag.
[0019] The present disclosure provides a video encoding method, comprising: encoding a transform skip mode flag indicating whether a transform / inverse transform is skipped in a current block; determining that a secondary transform / inverse transform is skipped in the current block when the transform skip mode is skipped in the current block according to the transform skip mode flag; and determining whether the secondary transform / inverse transform is skipped in the current block when the transform skip mode is not skipped in the current block according to the transform skip mode flag; and encoding a transform matrix index of the secondary transform / inverse transform for the current block according to whether the secondary transform / inverse transform is skipped in the current block.
[0020] According to an embodiment, the step of encoding the transform skip mode flag comprises:
[0021] The transform skip mode flag for a luma component, the transform skip mode flag for a Cb component, and the transform skip mode flag for a Cr component are encoded.
[0022] According to an embodiment, the step of determining that the secondary transform / inverse transform is skipped in the current block includes: when the tree structure of the current block is a single tree type, a transform skip mode flag for the luminance component indicates that the transform skip mode is applied to the luminance component, a transform skip mode flag for the Cb component indicates that the transform skip mode is applied to the Cb component, and a transform skip mode flag for the Cr component indicates that the transform skip mode is applied to the Cr component, the secondary transform / inverse transform is skipped in the current block.
[0023] According to an embodiment, the step of determining that the secondary transform / inverse transform is skipped in the current block includes: determining that the secondary transform / inverse transform is skipped in the current block when the tree structure of the current block is a dual-tree luminance type and the transform skip mode flag for the luminance component indicates that the transform skip mode is applied to the luminance component.
[0024] According to an embodiment, the step of determining that the secondary transform / inverse transform is skipped in the current block includes: when the tree structure of the current block is a dual-tree chroma type, a transform skip mode flag for the Cb component indicates that the transform skip mode is applied to the Cb component, and a transform skip mode flag for the Cr component indicates that the transform skip mode is applied to the Cr component, the secondary transform / inverse transform is skipped in the current block.
[0025] According to an embodiment, the step of encoding the transformation matrix index includes: when a secondary transformation / inverse transformation is applied to the current block, determining the secondary transformation matrix of the current block, and encoding the transformation matrix index according to whether the secondary transformation / inverse transformation is applied to the current block.
[0026] According to an embodiment, the step of determining the secondary transformation matrix of the current block includes: determining the secondary transformation matrix of the current block based on at least one of the transformation matrix index, the transformation matrix set index of the current block, and the size of the current block.
[0027] According to an embodiment, the video coding method may further include: encoding information about whether the intra-frame residual DPCM method is used in the current block, and determining that the transform / inverse transform is skipped in the current block when the information about whether the intra-frame residual DPCM method is used indicates that the intra-frame residual DPCM method is used in the current block, wherein the step of encoding the transform skip mode flag includes: encoding the transform skip mode flag when the information about whether the intra-frame residual DPCM method is used indicates that the intra-frame residual DPCM method is not used in the current block.
[0028] According to an embodiment, the step of encoding the transform matrix index for the secondary transform / inverse transform of the current block includes: encoding the transform matrix index when the current block is predicted according to an intra-prediction mode that is not a matrix-based intra-prediction mode.
[0029] This disclosure provides a computer-readable recording medium for storing a bitstream generated by encoding video using a video encoding method, wherein the video encoding method includes: encoding a transform skip mode flag indicating whether a transform / inverse transform is skipped in a current block; determining that a secondary transform / inverse transform is skipped in the current block when the transform skip mode is skipped according to the transform skip mode flag; determining whether a secondary transform / inverse transform is skipped in the current block when the transform skip mode is not skipped according to the transform skip mode flag; and encoding a transform matrix index for the secondary transform / inverse transform of the current block based on whether the secondary transform / inverse transform is skipped in the current block.
[0030] Beneficial effects
[0031] This invention can improve the objective and subjective quality of images by providing an image encoding / decoding method and apparatus. The image encoding / decoding method and apparatus determine at least one of the following: whether a transform is used, the type of one-dimensional transform, and a combination of two-dimensional transforms. This determines a reduced set of secondary transform / inverse transform matrices, a reduced secondary transform / inverse transform matrix, and whether to perform a reduced secondary transform / inverse transform. The use of transforms, the type of one-dimensional transform, and the combination of two-dimensional transforms are again based on intra-frame prediction mode, prediction mode, color components, size, and form. Attached Figure Description
[0032] FIG. 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0033] FIG. 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.
[0034] FIG. 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded.
[0035] FIG. 4 This is a diagram illustrating intra-frame prediction processing.
[0036] FIG. 5 This is a diagram illustrating an embodiment of inter-screen prediction processing.
[0037] FIG. 6This is a diagram illustrating the transformation and quantization processes.
[0038] FIG. 7 This is a diagram showing reference samples that can be used for intra-frame prediction.
[0039] FIG. 8 This is a diagram illustrating an embodiment of a decoding method using the SDST method according to the present invention.
[0040] FIG. 9 This is a diagram illustrating an embodiment of an encoding method using the SDST method according to the present invention.
[0041] FIG. 10 to FIG. 12 This is a diagram illustrating an embodiment of the first sub-block partitioning mode according to the present invention.
[0042] FIG. 13 This is a diagram illustrating an embodiment of the second sub-block partitioning mode according to the present invention.
[0043] FIG. 14 This is a diagram illustrating an embodiment of diagonal scanning.
[0044] FIG. 15 This is a diagram illustrating an embodiment of horizontal scanning.
[0045] FIG. 16 This is a diagram illustrating an embodiment of vertical scanning.
[0046] FIG. 17 This is a diagram illustrating an embodiment of block-based diagonal scanning.
[0047] FIG. 18 This is a diagram illustrating an embodiment of block-based horizontal scanning.
[0048] FIG. 19 This is a diagram illustrating an embodiment of block-based vertical scanning.
[0049] FIG. 20 These are illustrations showing various embodiments of scanning based on the shape of blocks.
[0050] FIG. 21 This is a diagram illustrating the intra-frame prediction mode.
[0051] FIG. 22 to FIG. 26 This is a diagram illustrating an example of encoding or decoding processing using transformations according to an embodiment of the present invention.
[0052] FIG. 27 Examples of performing secondary transformations and / or inverse secondary transformations in encoders / decoders are shown.
[0053] FIG. 28 An example of a secondary transformation matrix is shown.
[0054] FIG. 29 This illustrates the reduced secondary / inverse transformation process.
[0055] FIG. 30 to FIG. 32 Several embodiments are shown for deriving transformation matrices based on block size, transformation matrix set index, and transformation matrix index.
[0056] FIG. 33 to FIG. 36 The present invention illustrates the syntax of a bitstream applied to encoding / decoding methods and apparatuses using transformations and recording media storing bitstreams, according to embodiments thereof.
[0057] FIG. 37 to FIG. 54 Various embodiments are provided for sending conditions for the transformation matrix index using signals.
[0058] FIG. 55 The present invention illustrates the syntax of a bitstream applied to encoding / decoding methods and apparatuses using transformations and recording media storing bitstreams, according to embodiments thereof.
[0059] FIG. 56 A video decoding method according to an embodiment of the present invention is shown.
[0060] FIG. 57 A video encoding method according to an embodiment of the present invention is shown.
[0061] Best mode
[0062] This disclosure provides a video decoding method, comprising: obtaining a transform skip mode flag indicating whether a transform skip mode is applied to a current block; determining, based on the transform skip mode flag, that a secondary transform / inverse transform is not applied to the current block when the transform skip mode is applied to the current block; and obtaining, based on the transform skip mode flag, a transform matrix index for the secondary transform / inverse transform of the current block when the transform skip mode is not applied to the current block, and determining, based on the transform matrix index, whether the secondary transform / inverse transform is applied to the current block. Detailed Implementation
[0063] Various modifications can be made to this invention, and various embodiments of the invention exist, wherein examples of various embodiments of the invention will now be provided and described in detail with reference to the accompanying drawings. However, the invention is not limited thereto, although exemplary embodiments may be interpreted as including all modifications, equivalents, or substitutions within the technical concept and scope of the invention. In various respects, similar reference numerals refer to the same or similar functions. In the drawings, the shape and size of elements may be exaggerated for clarity. In the following detailed description of the invention, reference is made to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice this disclosure. It should be understood that the various embodiments of this disclosure, though different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in conjunction with one embodiment may be implemented in other embodiments without departing from the spirit and scope of this disclosure. Furthermore, it should be understood that the position or arrangement of various elements within each disclosed embodiment may be modified without departing from the spirit and scope of this disclosure. Therefore, the following detailed description should not be considered limiting, and the scope of this disclosure is defined only by the appended claims (which, where properly interpreted, also include the full scope of the equivalents claimed by the claims).
[0064] The terms "first," "second," etc., used in this specification may be used to describe various components, but the components should not be construed as limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the invention, a "first" component may be named a "second" component, and a "second" component may similarly be named a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple items.
[0065] It will be understood that, in this specification, when an element is simply referred to as "connected to" or "coupled to" another element rather than "directly connected to" or "directly coupled to" another element, the element may be "directly connected to" or "directly coupled to" another element, or may be connected to or coupled to another element where there is another element intervening between the element and the other element. Conversely, it should be understood that when an element is referred to as "directly coupled to" or "directly connected" to another element, there is no intermediate element.
[0066] Furthermore, the constituent parts shown in the embodiments of the present invention are illustrated independently to represent different functional characteristics. Therefore, this does not mean that each constituent part is constructed as a separate hardware or software unit. In other words, for convenience, each constituent part includes each of the listed constituent parts. Thus, at least two constituent parts of each constituent part can be combined to form one constituent part, or a constituent part can be divided into multiple constituent parts to perform each function. Embodiments in which each constituent part is combined and embodiments in which each constituent part is divided are also included within the scope of the present invention without departing from its spirit.
[0067] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Unless the context clearly distinguishes them, expressions used in the singular include those used in the plural. In this specification, it will be understood that terms such as “comprising,” “having,” etc., are intended to indicate the presence of features, numbers, steps, actions, elements, components, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, components, or combinations thereof may be present or added. In other words, when a particular element is referred to as “comprising,” it does not exclude elements other than the corresponding element, but rather includes additional elements within the embodiments of the invention or within the scope of the invention.
[0068] Furthermore, some components may not be essential for performing the basic functions of the invention, but rather optional components that only improve its performance. The invention can be implemented by including only the essential components necessary for achieving the essence of the invention, excluding components used to improve performance. Structures that include only the essential components and exclude optional components used only to improve performance are also included within the scope of the invention.
[0069] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In describing exemplary embodiments of the invention, well-known functions or structures will not be described in detail, as they may unnecessarily obscure the understanding of the invention. Like constituent elements in the drawings are indicated by like reference numerals, and repeated descriptions of like elements will be omitted.
[0070] In the following text, an image may refer to a frame that constitutes a video, or it may refer to the video itself. For example, "encoding or decoding an image or both" may refer to "encoding or decoding a moving image or both," and may also refer to "encoding or decoding an image within an image of a moving image or both."
[0071] In the following text, the terms "moving images" and "video" are used to mean the same thing and are interchangeable.
[0072] In the following text, the target image can be an encoded target image that serves as an encoding target and / or a decoded target image that serves as a decoding target. Furthermore, the target image can be an input image input to an encoding device and an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
[0073] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used as having the same meaning and may be used interchangeably with each other.
[0074] In the following text, a target block can be an encoded target block that serves as the encoding target and / or a decoded target block that serves as the decoding target. Furthermore, a target block can be the current block that serves as the target of the current encoding and / or decoding. For example, the terms "target block" and "current block" can be used to mean the same thing and can be used interchangeably.
[0075] In the following text, the terms "block" and "unit" may be used to mean the same thing and may be used interchangeably. Alternatively, "block" may refer to a specific unit.
[0076] In the following text, the terms “region” and “fragment” are used interchangeably.
[0077] In the following text, a specific signal can be a signal representing a specific block. For example, the original signal can be a signal representing the target block. The prediction signal can be a signal representing the prediction block. The residual signal can be a signal representing the residual block.
[0078] In this embodiment, each of the following can have a value: information, data, flags, indexes, elements, and attributes. A value of "0" for information, data, flags, indexes, elements, and attributes can represent logical false or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value can be interchanged. A value of "1" for information, data, flags, indexes, elements, and attributes can represent logical true or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value can be interchanged.
[0079] When variables i or j are used to represent columns, rows, or indices, the value of i can be an integer equal to or greater than 0, or an integer equal to or greater than 1. That is, columns, rows, indices, etc., can be counted starting from 0, or they can be counted starting from 1.
[0080] Description of terms
[0081] Encoder: Represents the device that performs encoding. In other words, it refers to the encoding device.
[0082] Decoder: Represents the device that performs decoding. In other words, it refers to the decoding device.
[0083] A block is an M×N sample array. Here, M and N can represent positive integers, and a block can represent a two-dimensional sample array. A block can refer to a unit. The current block can represent a coding target block that becomes the target during encoding, or a decoding target block that becomes the target during decoding. Furthermore, the current block can be at least one of a coding block, a prediction block, a residual block, and a transform block.
[0084] Samples are the basic units that make up a block. Based on the bit depth (Bd), samples can be represented as numbers from 0 to 2. Bd The value is -1. In this invention, a sample point can be used to represent a pixel. That is, a sample point, a pel, and a pixel can have the same meaning.
[0085] Unit: Refers to an encoding and decoding unit. When encoding and decoding an image, a unit can be a region generated by partitioning a single image. Furthermore, a unit can represent a sub-partitioning unit when a single image is partitioned into sub-partitioning units during encoding or decoding. That is, an image can be partitioned into multiple units. When encoding and decoding an image, predetermined processing can be performed for each unit. A single unit can be partitioned into sub-units smaller than the unit's size. Depending on the function, a unit can represent a block, macroblock, coding tree unit, coding tree block, coding unit, coding block, prediction unit, prediction block, residual unit, residual block, transform unit, transform block, etc. Furthermore, to distinguish a unit from a block, a unit can include a luma component block, a chroma component block associated with the luma component block, and syntax elements for each chroma component block. Units can have various sizes and shapes; specifically, the shape of a unit can be a two-dimensional geometric shape, such as a square, rectangle, trapezoid, triangle, pentagon, etc. In addition, the cell information may include at least one of the following: cell type indicating coding cell, prediction cell, transform cell, etc., cell size, cell depth, and the order of encoding and decoding of the cell.
[0086] A coding tree unit is a single coding tree block configured with the luminance component Y and two coding tree blocks associated with the chrominance components Cb and Cr. Furthermore, a coding tree unit can represent a block and the syntax elements of each block. Each coding tree unit can be partitioned using at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning methods to configure lower-level units such as coding units, prediction units, transform units, etc. A coding tree unit can be used as a term to specify a sample block that becomes a processing unit when encoding / decoding an image as an input image. Here, a quadtree can represent a quaternion tree.
[0087] When the size of the coded block is within a predetermined range, it is possible to partition using only quadtree partitioning. Here, the predetermined range can be defined as at least one of the maximum and minimum sizes of the coded block that can be partitioned using only quadtree partitioning. Information indicating the maximum / minimum size of the coded block that allows quadtree partitioning can be transmitted via a signal in the bitstream, and this information can be transmitted via a signal in at least one unit of sequence, frame parameters, parallel block groups, or stripes (fragments). Optionally, the maximum / minimum size of the coded block can be a predetermined fixed size in the encoder / decoder. For example, when the size of the coded block corresponds to 256×256 to 64×64, it is possible to partition using only quadtree partitioning. Optionally, when the size of the coded block is greater than the size of the maximum transform block, it is possible to partition using only quadtree partitioning. Here, the block to be partitioned can be at least one of a coded block and a transform block. In this case, the information indicating the partitioning of the coded block (e.g., split_flag) can be a flag indicating whether quadtree partitioning is performed. When the size of the coded block falls within a predetermined range, it is possible to partition using only binary or ternary tree partitioning. In this case, the above description of quadtree partitioning can be applied in the same way to binary tree partitioning or ternary tree partitioning.
[0088] Encoding block: Can be used as a term to specify any one of the Y encoding block, Cb encoding block, and Cr encoding block.
[0089] Neighboring blocks: These can represent blocks adjacent to the current block. A block adjacent to the current block can be a block that touches the boundary of the current block, or a block located within a predetermined distance from the current block. A neighboring block can also represent a block adjacent to a vertex of the current block. Here, a block adjacent to a vertex of the current block can be a block that is vertically adjacent to a block horizontally adjacent to the current block, or a block that is horizontally adjacent to a block vertically adjacent to the current block.
[0090] Reconstructing neighboring blocks: This can represent neighboring blocks that are adjacent to the current block and have already been spatially / temporally encoded or decoded. Here, reconstructing neighboring blocks can represent reconstructing neighboring units. Reconstructing spatial neighboring blocks can be blocks within the current frame that have already been reconstructed through encoding or decoding, or both. Reconstructing temporal neighboring blocks are blocks within a reference image that are located at the position corresponding to the current block in the current frame, or are neighboring blocks of said block.
[0091] Cell depth: Represents the degree of cell partitioning. In a tree structure, the highest node (root node) corresponds to the first unpartitioned cell. Furthermore, the highest node can have the minimum depth value. In this case, the depth of the highest node can be level 0. A node with a depth of level 1 represents a cell generated by partitioning the first cell once. A node with a depth of level 2 represents a cell generated by partitioning the first cell twice. A node with a depth of level n represents a cell generated by partitioning the first cell n times. Leaf nodes can be the lowest-level nodes and cannot be further partitioned. The depth of a leaf node can be the highest level. For example, the predefined value of the highest level can be 3. The root node can have the lowest depth, and the leaf nodes can have the deepest depth. Additionally, when cells are represented as a tree structure, the level in which the cell exists can represent the cell depth.
[0092] Bitstream: A bitstream that can represent encoded image information.
[0093] Parameter set: Corresponds to the header information in the configuration within the bitstream. At least one of the video parameter set, sequence parameter set, frame parameter set, and adaptive parameter set may be included in the parameter set. Furthermore, the parameter set may include slice headers, tile group headers, and tile header information. The term "tile group" refers to a set of parallel tiles and has the same meaning as a slice.
[0094] An adaptive parameter set can represent a set of parameters that can be shared by referencing different frames, subframes, stripes, parallel block groups, parallel blocks, or bricks. Furthermore, information in an adaptive parameter set can be used by referencing different adaptive parameter sets for subframes, stripes, parallel block groups, parallel blocks, or bricks within a frame.
[0095] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets within a frame, such as sub-frames, stripes, parallel block groups, parallel blocks, or chunks.
[0096] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets within a sub-picture, such as stripes, parallel block groups, parallel blocks, or chunks.
[0097] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets for parallel blocks or segments within a strip.
[0098] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets for blocks within a parallel block.
[0099] Information about the adaptive parameter set identifier can be included in the parameter set or header of the sub-screen, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the sub-screen.
[0100] Information about the adaptive parameter set identifier can be included in the parameter set or header of the parallel block, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the parallel block.
[0101] Information about the adaptive parameter set identifier can be included in the header of the block, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used for the block.
[0102] The screen can be divided into one or more parallel block rows and one or more parallel block columns.
[0103] A sub-screen can be divided into one or more parallel block rows and one or more parallel block columns within the screen. A sub-screen can be a rectangular / square region within the screen and may include one or more CTUs. In addition, at least one or more parallel blocks / strips / strips may be included within a sub-screen.
[0104] A parallel block can be a rectangular / square region within the frame and may include one or more CTUs. Furthermore, a parallel block may be divided into one or more sub-blocks.
[0105] A block can represent one or more CTU lines within a parallel block. A parallel block can be partitioned into one or more blocks, and each block can have at least one or more CTU lines. A parallel block that is not partitioned into two or more blocks can be represented as a block.
[0106] A strip may include one or more parallel blocks within a frame, and may include one or more sub-blocks within a parallel block.
[0107] Explanation: This could mean determining the value of a syntax element by performing entropy decoding, or it could mean entropy decoding itself.
[0108] Symbols: can represent at least one of the syntax elements, encoding parameters, and transform coefficient values of the encoding / decoding target unit. Additionally, symbols can represent entropy encoding targets or entropy decoding results.
[0109] Prediction mode: This can be information indicating the mode that is encoded / decoded using intra-frame prediction or the mode that is encoded / decoded using inter-frame prediction.
[0110] Prediction Unit: A basic unit that can be represented when performing prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation). A single prediction unit can be partitioned into multiple partitions with smaller sizes, or it can be partitioned into multiple lower-level prediction units. Multiple partitions can be the basic units when performing prediction or compensation. Partitions generated by dividing prediction units can also be prediction units.
[0111] Prediction cell partitioning: can represent the shape obtained by partitioning prediction cells.
[0112] A reference frame list can refer to a list of one or more reference frames used for inter-frame prediction or motion compensation. There are several types of available reference frame lists, including LC (list combination), L0 (list 0), L1 (list 1), L2 (list 2), and L3 (list 3).
[0113] The inter-frame prediction indicator can indicate the direction of inter-frame prediction for the current block (unidirectional prediction, bidirectional prediction, etc.). Optionally, the inter-frame prediction indicator can indicate the number of reference frames used to generate the prediction blocks for the current block. Optionally, the inter-frame prediction indicator can indicate the number of prediction blocks used when performing inter-frame prediction or motion compensation on the current block.
[0114] The prediction list utilization flag indicates whether at least one reference frame from a specific reference frame list is used to generate a prediction block. The prediction list utilization flag can be used to derive an inter-frame prediction indicator, and conversely, the inter-frame prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag has a first value of zero (0), it indicates that a reference frame from the reference frame list is not used to generate a prediction block. On the other hand, when the prediction list utilization flag has a second value of one (1), it indicates that the reference frame list is used to generate a prediction block.
[0115] The reference screen index can refer to the index of a specific reference screen in the reference screen list.
[0116] A reference frame can refer to a frame referenced by a specific block for the purpose of inter-frame prediction or motion compensation for that specific block. Alternatively, a reference frame can be a frame that includes a reference block referenced by the current block for inter-frame prediction or motion compensation. In the following text, the terms "reference frame" and "reference image" have the same meaning and are interchangeable.
[0117] Motion vectors can be two-dimensional vectors used for inter-frame prediction or motion compensation. A motion vector can represent the offset between the encoded / decoded target block and the reference block. For example, (mvX, mvY) can represent a motion vector. Here, mvX can represent the horizontal component, and mvY can represent the vertical component.
[0118] The search range can be a two-dimensional region searched during inter-frame prediction to retrieve motion vectors. For example, the size of the search range can be M×N. Here, M and N are both integers.
[0119] Motion vector candidates can refer to a block of prediction candidates or the motion vectors within a block of prediction candidates when making predictions about motion vectors. Furthermore, motion vector candidates can be included in a list of motion vector candidates.
[0120] A motion vector candidate list can represent a list consisting of one or more motion vector candidates.
[0121] A motion vector candidate index can represent an indicator that points to a motion vector candidate in the motion vector candidate list. Alternatively, it can be an index of a motion vector predictor.
[0122] Motion information may include at least one of the following: motion vector, reference frame index, inter-frame prediction indicator, prediction list utilization flag, reference frame list information, reference frame, motion vector candidate, motion vector candidate index, merge candidate, and merge index.
[0123] A merge candidate list can represent a list consisting of one or more merge candidates.
[0124] Merge candidates can be spatial merge candidates, temporal merge candidates, combined merge candidates, combined double prediction merge candidates, or zero merge candidates. Merge candidates may include motion information such as inter-frame prediction indicators, reference frame indices for each list, motion vectors, prediction list utilization flags, and inter-frame prediction indicators.
[0125] The merge index can represent an indicator pointing to a merge candidate in the merge candidate list. Optionally, the merge index can indicate a block in a reconstructed block that is spatially / temporally adjacent to the current block, from which a merge candidate has been derived. Optionally, the merge index can indicate at least one piece of motion information for a merge candidate.
[0126] Transform unit: This can represent the basic unit used when encoding / decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding / decoding) a residual signal. A single transform unit can be partitioned into multiple lower-level transform units with smaller sizes. Here, the transform / inverse transform may include at least one of a first transform / first inverse transform and a second transform / second inverse transform.
[0127] Scaling: This refers to the process of multiplying the quantization level by a factor. Transform coefficients can be generated by scaling the quantization level. Scaling can also be called inverse quantization.
[0128] Quantization parameters: These represent values used when transform coefficients are used to generate quantization levels during quantization. Quantization parameters can also represent values used when transform coefficients are generated during dequantization by scaling the quantization levels. Quantization parameters can be values mapped to the quantization step size.
[0129] Incremental quantization parameter: can represent the difference between the predicted quantization parameter and the quantization parameter of the encoding / decoding target unit.
[0130] Scan: This can refer to a method of sorting coefficients within a cell, block, or matrix. For example, changing a two-dimensional matrix of coefficients into a one-dimensional matrix can be called a scan, and changing a one-dimensional matrix of coefficients into a two-dimensional matrix can be called a scan or inverse scan.
[0131] Transform coefficients: These represent the coefficient values generated after a transform is performed in the encoder. Transform coefficients can also represent the coefficient values generated after at least one of entropy decoding and dequantization is performed in the decoder. The quantization level obtained by quantizing the transform coefficients or residual signal, or the quantized transform coefficient level, can also fall within the meaning of transform coefficients.
[0132] Quantization level: This can represent the value generated in the encoder by quantizing the transform coefficients or residual signal. Optionally, the quantization level can represent the value of the dequantization target that undergoes dequantization in the decoder. Similarly, the transform coefficient level, as a result of transform and quantization, can also fall within the meaning of quantization level.
[0133] Non-zero transform coefficients: can represent transform coefficients with values other than zero, or transform coefficient levels or quantization levels with values other than zero.
[0134] Quantization matrix: A matrix used in quantization or dequantization processes performed to improve subjective or objective image quality. The quantization matrix can also be referred to as a scaling list.
[0135] Quantization matrix coefficients: These represent each element within the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.
[0136] Default matrix: can represent a predefined quantization matrix in the encoder or decoder.
[0137] Non-default matrix: can represent a quantization matrix that is not predefined in the encoder or decoder but is sent by the user via signal.
[0138] Statistical value: For at least one of the following variables, coding parameters, constant values, etc., that have a computable specific value, a statistical value can be one or more of the following: mean, summation, weighted average, weighted sum, minimum, maximum, most frequent value, median, interpolation.
[0139] FIG. 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0140] Encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include at least one image. Encoding device 100 may encode at least one image sequentially.
[0141] Reference FIG. 1 The encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.
[0142] Encoding device 100 can encode the input image using intra-frame mode, inter-frame mode, or both. Furthermore, encoding device 100 can generate a bitstream including encoding information by encoding the input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium. When intra-frame mode is used as the prediction mode, switcher 115 can switch to intra-frame mode. Optionally, when inter-frame mode is used as the prediction mode, switcher 115 can switch to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. Encoding device 100 can generate prediction blocks for input blocks of the input image. Furthermore, encoding device 100 can encode residual blocks using the residual between the input block and the prediction block after generating the prediction blocks. The input image can be referred to as the current image as the current encoding target. The input block can be referred to as the current block as the current encoding target, or as the encoding target block.
[0143] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from blocks that have been encoded / decoded and are adjacent to the current block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the current block using the reference samples, or generate prediction samples for the input block by performing spatial prediction. Here, intra-frame prediction can refer to prediction within a frame.
[0144] When the prediction mode is inter-frame mode, the motion prediction unit 111 can retrieve the region that best matches the input block from the reference image during motion prediction and derive the motion vector using the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference frame buffer 190. Here, the reference image can be stored in the reference frame buffer 190 when encoding / decoding the reference image is performed.
[0145] The motion compensation unit 112 can generate a prediction block by performing motion compensation on the current block using motion vectors. Here, inter-frame prediction can refer to prediction or motion compensation between frames.
[0146] When the value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 can generate prediction blocks by applying an interpolation filter to a portion of the reference frame. To perform inter-frame prediction or motion compensation on the coding unit, it can be determined which of the following modes—skip mode, merge mode, Advanced Motion Vector Prediction (AMVP) mode, and current frame reference mode—is used for motion prediction and motion compensation on the prediction unit included in the corresponding coding unit. Then, depending on the determined mode, inter-frame prediction or motion compensation can be performed differently.
[0147] Subtractor 125 generates a residual block by using the difference between the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Furthermore, the residual signal can be a signal generated by transforming or quantizing, or transforming and quantizing, the difference between the original signal and the prediction signal. The residual block can be the residual signal of a block cell.
[0148] Transform unit 130 can generate transform coefficients by performing a transform on the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, transform unit 130 can skip the transform on the residual block.
[0149] The level of quantization can be generated by applying quantization to the transform coefficients or to the residual signal. In the following examples, the level of quantization may also be referred to as the transform coefficients.
[0150] The quantization unit 140 can generate a quantization level by quantizing the transform coefficients or residual signal according to parameters, and output the generated quantization level. Here, the quantization unit 140 can quantize the transform coefficients using a quantization matrix.
[0151] Entropy coding unit 150 can generate a bitstream by performing entropy coding on the values calculated by quantization unit 140 according to a probability distribution or on the coding parameter values calculated during encoding, and output the generated bitstream. Entropy coding unit 150 can perform entropy coding on sample information of the image and information used for decoding the image. For example, the information used for decoding the image may include syntax elements.
[0152] When entropy coding is applied, symbols are represented such that fewer bits are allocated to symbols with high generation probability and more bits are allocated to symbols with low generation probability, thus reducing the size of the bitstream used to encode the symbols. The entropy coding unit 150 can use coding methods for entropy coding such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code (VLC) table. Furthermore, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / bits, and perform arithmetic coding by using the derived binarization method and context model.
[0153] In order to encode the transform coefficient levels (quantization levels), the entropy coding unit 150 can change the coefficients in two-dimensional block form into one-dimensional vector form by using a transform coefficient scanning method.
[0154] Encoding parameters may include information such as syntax elements (flags, indexes, etc.) encoded in the encoder and signaled to the decoder, as well as information derived during encoding or decoding. Encoding parameters can represent the information required when encoding or decoding an image. For example, at least one value or combination of the following may be included in the encoding parameters: cell / block size, cell / block depth, cell / block partitioning information, cell / block shape, cell / block partitioning structure, whether quadtree partitioning is performed, whether binary tree partitioning is performed, binary tree partitioning direction (horizontal or vertical), binary tree partitioning type (symmetric or asymmetric), whether the current encoded unit is partitioned via ternary tree partitioning, the direction of the ternary tree partitioning (horizontal or vertical), the type of ternary tree partitioning (symmetric or asymmetric), whether the current encoded unit is partitioned via multi-type tree partitioning, and the type of multi-type tree partitioning. Direction (horizontal or vertical), type of multi-type tree partition (symmetric or asymmetric), tree structure of multi-type tree partition (binary or ternary), prediction mode (intra-frame prediction or inter-frame prediction), luma intra-frame prediction mode / direction, chroma intra-frame prediction mode / direction, intra-frame partition information, inter-frame partition information, coded block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, prediction block filtering method, prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, intra-frame prediction mode Inter-frame prediction mode, motion information, motion vector, motion vector difference, reference frame index, inter-frame prediction angle, inter-frame prediction indicator, prediction list utilization flag, reference frame list, reference frame, motion vector predictor index, motion vector predictor candidate, motion vector candidate list, whether to use merge mode, merge index, merge candidate, merge candidate list, whether to use skip mode, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector magnitude, motion vector representation accuracy, transform type, transform size, information on whether the primary (first) transform is used, information on whether the secondary transform is used, primary transform index, secondary transform index Information on the presence of residual signals, code block style, code block flag (CBF), quantization parameters, quantization parameter residuals, quantization matrix, whether an intra-loop filter is applied, intra-loop filter coefficients, intra-loop filter taps, intra-loop filter shape / form, whether a deblocking filter is applied, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset type, adaptive sample offset function, whether an adaptive intra-loop filter is applied, adaptive intra-loop filter coefficients, adaptive intra-loop filter taps, adaptive intra-loop filter shape / form.Binarization / debinarization method, context model determination method, context model update method, whether to execute normal mode, whether to execute bypass mode, context binary bits, bypass binary bits, valid coefficient flag, last valid coefficient flag, encoding flag for the unit of the coefficient group, position of the last valid coefficient, flag indicating whether the coefficient value is greater than 1, flag indicating whether the coefficient value is greater than 2, flag indicating whether the coefficient value is greater than 3, information about the remaining coefficient values, positive and negative sign information, reconstructed luminance sample, reconstructed chrominance sample, residual luminance sample, residual chrominance sample, luminance transform coefficient, chrominance transform coefficient, quantized luminance level, quantized chrominance level, transform coefficient level scanning method, motion vector on the decoder side. The information includes the size of the motion vector search area, the shape of the motion vector search area on the decoder side, the number of motion vector searches on the decoder side, information about the CTU size, information about the minimum block size, information about the maximum block size, information about the maximum block depth, information about the minimum block depth, image display / output order, strip identification information, strip type, strip partition information, parallel block identification information, parallel block type, parallel block partition information, parallel block group identification information, parallel block group type, parallel block group partition information, image type, bit depth of input samples, bit depth of reconstructed samples, bit depth of residual samples, bit depth of transform coefficients, bit depth of quantization levels, and information about the luminance signal or the chrominance signal.
[0155] Here, sending a flag or index with a signal can indicate that the encoder entropy-encodes the corresponding flag or index and includes it in the bitstream, and can also indicate that the decoder entropy-decodes the corresponding flag or index from the bitstream.
[0156] When the encoding device 100 performs encoding via inter-frame prediction, the encoded current image can be used as a reference image for another image to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded current image, or store the reconstructed or decoded image as a reference image in the reference frame buffer 190.
[0157] The quantization level can be dequantized in dequantization unit 160 or inverse transformed in inverse transform unit 170. The coefficients that have undergone dequantization or inverse transform, or both, can be added to the prediction block by adder 175. A reconstruction block can be generated by adding the coefficients that have undergone dequantization or inverse transform, or both, to the prediction block. Here, the coefficients that have undergone dequantization or inverse transform, or both, can represent coefficients for which at least one of dequantization and inverse transform has been performed, and can represent the reconstructed residual block.
[0158] The reconstructed block can be passed through filter unit 180. Filter unit 180 can apply at least one of deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF) to the reconstructed sample, reconstructed block, or reconstructed image. Filter unit 180 may be referred to as an in-loop filter.
[0159] Deblocking filters remove block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, the number of samples included in several rows or columns within the block can be used. When a deblocking filter is applied to a block, another filter can be applied based on the desired deblocking intensity.
[0160] To compensate for coding errors, a suitable offset value can be added to the sample value using a sample-adaptive offset. The sample-adaptive offset corrects the offset between the deblocked image and the original image on a sample-by-sample basis. This can be achieved by considering edge information about each sample point when applying the offset, or by dividing the image's samples into a predetermined number of regions, determining the regions where the offset will be applied, and then applying the offset to those regions.
[0161] The adaptive in-loop filter (ALF) performs filtering based on a comparison between the filtered reconstructed image and the original image. Samples included in the image can be partitioned into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information regarding whether to apply the ALF can be transmitted via a signal through the coding unit (CU), and the form and coefficients of the ALF to be applied to each block can vary.
[0162] The reconstructed blocks or reconstructed image that have passed through filter unit 180 can be stored in reference frame buffer 190. The reconstructed blocks processed by filter unit 180 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filter unit 180. The stored reference image can be used later in inter-frame prediction or motion compensation.
[0163] FIG. 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment and to which the present invention is applied.
[0164] Decoding device 200 can be a decoder, video decoding device, or image decoding device.
[0165] Reference FIG. 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference frame buffer 270.
[0166] Decoding device 200 can receive bitstreams output from encoding device 100. Decoding device 200 can receive bitstreams stored on a computer-readable recording medium, or bitstreams streamed via wired / wireless transmission media. Decoding device 200 can decode the bitstreams using intra-frame mode or inter-frame mode. Furthermore, decoding device 200 can generate and output reconstructed or decoded images generated through decoding.
[0167] When the prediction mode used during decoding is intra-frame mode, the switcher can be switched to intra-frame mode. Optionally, when the prediction mode used during decoding is inter-frame mode, the switcher can be switched to inter-frame mode.
[0168] Decoding device 200 obtains a reconstructed residual block and generates a prediction block by decoding the input bitstream. Once the reconstructed residual block and prediction block are obtained, decoding device 200 generates a reconstructed block that becomes the decoding target by adding the reconstructed residual block and the prediction block. The decoding target block can be referred to as the current block.
[0169] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols in quantized hierarchical form. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0170] In order to decode the transform coefficient levels (quantization levels), the entropy decoding unit 210 can change the coefficients in unidirectional vector form into two-dimensional block form by using a transform coefficient scanning method.
[0171] The quantization level can be dequantized in the dequantization unit 220, or the quantization level can be inversely transformed in the inverse transform unit 230. The quantization level can be the result of dequantization, inverse transform, or both, and can be generated as a reconstruction residual block. Here, the dequantization unit 220 can apply the quantization matrix to the quantization level.
[0172] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, wherein the spatial prediction uses sample values of blocks that are adjacent to the target block and have already been decoded.
[0173] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, wherein the motion compensation uses motion vectors and a reference image stored in reference frame buffer 270.
[0174] Adder 255 generates a reconstructed block by adding the reconstructed residual block to the prediction block. Filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or reconstructed image. Filter unit 260 can output a reconstructed image. The reconstructed block or reconstructed image can be stored in a reference frame buffer 270 and used when performing inter-frame prediction. The reconstructed block processed by filter unit 260 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filter unit 260. The stored reference image can be used later in inter-frame prediction or motion compensation.
[0175] FIG. 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded. FIG. 3 An example of dividing a single cell into multiple sub-cells is illustrated schematically.
[0176] To effectively partition an image, coding units (CUs) are used during encoding and decoding. A coding unit can serve as the basic unit when encoding / decoding an image. Furthermore, a coding unit can be used to distinguish between intra-frame prediction modes and inter-frame prediction modes during image encoding / decoding. A coding unit can be the basic unit used for prediction, transform, quantization, inverse transform, inverse quantization, or encoding / decoding processing of transform coefficients.
[0177] Reference FIG. 3 Image 300 is partitioned sequentially according to the Largest Coding Unit (LCU), and the LCU unit is determined as the partitioning structure. Here, LCU can be used with the same meaning as Coding Tree Unit (CTU). Unit partitioning can represent partitioning of the block associated with that unit. The block partitioning information may include information about the unit depth. The depth information may represent the number or degree to which the unit is partitioned, or both the number and degree to which the unit is partitioned. A single unit can be partitioned into multiple lower-level units hierarchically associated with the depth information based on a tree structure. In other words, the unit and the lower-level units generated by partitioning the unit may correspond to a node and the child nodes of that node, respectively. Each of the partitioned lower-level units may have depth information. The depth information may be information representing the size of the CU and may be stored in each CU. The unit depth represents the number and / or degree associated with partitioning the unit. Therefore, the partitioning information of the lower-level units may include information about the size of the lower-level units.
[0178] The partitioning structure represents the distribution of coding units (CUs) within the LCU 310. This distribution can be determined by whether a single CU is partitioned into multiple CUs (including positive integers equal to or greater than 2, such as 2, 4, 8, 16, etc.). The horizontal and vertical dimensions of the CUs generated by partitioning can be half the horizontal and vertical dimensions of the CUs before partitioning, respectively, or they can have dimensions smaller than the horizontal and vertical dimensions before partitioning, depending on the number of partitions performed. CUs can be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the height and width of the CU after partitioning can be reduced compared to at least one of the height and width of the CU before partitioning. CU partitioning can be performed recursively until a predefined depth or a predefined size is reached. For example, the depth of the LCU can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, as mentioned above, the LCU can be a coding unit with the maximum coding unit size, and the SCU can be a coding unit with the minimum coding unit size. Partitioning begins at LCU 310. The CU depth increases by 1 when the horizontal or vertical dimension of a CU, or both, are reduced through partitioning. For example, for each depth, the size of an unpartitioned CU can be 2N×2N. Furthermore, in the case of partitioned CUs, a CU of size 2N×2N can be partitioned into four CUs of size N×N. As the depth increases by 1, the size of N can be halved.
[0179] Furthermore, partition information of a CU can be used to indicate whether a CU is partitioned. Partition information can be 1 bit. All CUs except SCUs can include partition information. For example, when the partition information value is the first value, the CU may not be partitioned; when the partition information value is the second value, the CU may be partitioned.
[0180] Reference FIG. 3 An LCU with depth 0 can be a 64×64 block. 0 can be the minimum depth. An SCU with depth 3 can be an 8×8 block. 3 can be the maximum depth. CUs with 32×32 blocks and 16×16 blocks can be represented as depth 1 and depth 2, respectively.
[0181] For example, when a single coding unit is partitioned into four coding units, the horizontal and vertical dimensions of the four partitioned coding units can be half the size of the CU before partitioning. In one embodiment, when a 32×32 coding unit is partitioned into four coding units, each of the four partitioned coding units can have a size of 16×16. When a single coding unit is partitioned into four coding units, the coding unit can be said to be partitioned into a quadtree form.
[0182] For example, when a coding unit is partitioned into two sub-coding units, the horizontal or vertical dimension (width or height) of each of the two sub-coding units can be half the horizontal or vertical dimension of the original coding unit. For example, when a coding unit of size 32×32 is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16×32. For example, when a coding unit of size 8×32 is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 8×16. When a coding unit is partitioned into two sub-coding units, it can be said that the coding unit is binary partitioned or partitioned according to a binary tree partitioning structure.
[0183] For example, when a coding unit is divided into three sub-coding units, the horizontal or vertical dimensions of the coding unit can be divided in a 1:2:1 ratio, resulting in three sub-coding units with a horizontal or vertical dimension ratio of 1:2:1. For instance, when a 16×32 coding unit is horizontally divided into three sub-coding units, these three sub-coding units, in order from the top to the bottom, can have dimensions of 16×8, 16×16, and 16×8, respectively. Similarly, when a 32×32 coding unit is vertically divided into three sub-coding units, these three sub-coding units, in order from the left to the right, can have dimensions of 8×32, 16×32, and 8×32, respectively. When a coding unit is divided into three sub-coding units, it can be said that the coding unit is tri-partitioned or partitioned according to a ternary tree partitioning structure.
[0184] exist FIG. 3 In the example, the coding tree unit (CTU) 320 is an example of a CTU in which quadtree partitioning, binary tree partitioning, and ternary tree partitioning structures are all applied.
[0185] As described above, to partition the CTU, at least one of a quadtree partitioning structure, a binary tree partitioning structure, and a ternary tree partitioning structure can be applied. Various tree partitioning structures can be applied sequentially to the CTU according to a predetermined priority order. For example, a quadtree partitioning structure can be preferentially applied to the CTU. Encoding units that cannot be further partitioned using a quadtree partitioning structure can correspond to leaf nodes of a quadtree. Encoding units corresponding to leaf nodes of a quadtree can be used as root nodes of binary and / or ternary tree partitioning structures. That is, encoding units corresponding to leaf nodes of a quadtree can be further partitioned according to a binary or ternary tree partitioning structure, or they can be left unpartitioned. Therefore, by preventing encoding units obtained from binary or ternary tree partitioning of encoding units corresponding to leaf nodes of a quadtree from undergoing further quadtree partitioning, block partitioning operations and / or the operation of signaling partitioning information can be effectively performed.
[0186] The fact that a coding unit corresponding to a node in a quadtree is partitioned can be signaled using four-partition information. Four-partition information with a first value (e.g., "1") indicates that the current coding unit is partitioned according to the quadtree partitioning structure. Four-partition information with a second value (e.g., "0") indicates that the current coding unit is not partitioned according to the quadtree partitioning structure. The four-partition information can be a flag with a predetermined length (e.g., one bit).
[0187] There may be no priority between binary tree partitions and ternary tree partitions. That is, the coding unit corresponding to the leaf node of the quadtree can further undergo any partition in either binary tree or ternary tree partitions. Furthermore, the coding unit generated by binary tree partitions or ternary tree partitions may undergo further binary tree partitions or further ternary tree partitions, or it may not be further partitioned.
[0188] A tree structure in which there is no priority between binary tree partitions and ternary tree partitions is called a multi-type tree structure. The coding unit corresponding to the leaf node of a quadtree can be used as the root node of a multi-type tree. At least one of multi-type tree partition indication information, partition direction information, and partition tree information can be used to signal whether to partition the coding unit corresponding to a node in the multi-type tree. To partition the coding unit corresponding to a node in the multi-type tree, the multi-type tree partition indication information, partition direction information, and partition tree information can be signaled sequentially.
[0189] A multi-type tree partitioning indication with a first value (e.g., "1") indicates that the current coding unit will undergo a multi-type tree partition. A multi-type tree partitioning indication with a second value (e.g., "0") indicates that the current coding unit will not undergo a multi-type tree partition.
[0190] When the coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partitioning structure, the coding unit may include partitioning direction information. The partitioning direction information may indicate in which direction the current coding unit will be partitioned for the multi-type tree partition. Partitioning direction information with a first value (e.g., "1") may indicate that the current coding unit will be vertically partitioned. Partitioning direction information with a second value (e.g., "0") may indicate that the current coding unit will be horizontally partitioned.
[0191] When the coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partitioning structure, the current coding unit may include partitioning tree information. The partitioning tree information may indicate the tree partitioning structure that will be used to partition the nodes of the multi-type tree. Partitioning tree information with a first value (e.g., "1") may indicate that the current coding unit will be partitioned according to a binary tree partitioning structure. Partitioning tree information with a second value (e.g., "0") may indicate that the current coding unit will be partitioned according to a ternary tree partitioning structure.
[0192] The partition indication information, partition tree information, and partition direction information can all be flags with a predetermined length (e.g., one bit).
[0193] At least one of the following—quadtree partitioning indication information, multi-type tree partitioning indication information, partitioning direction information, and partitioning tree information—can be entropy encoded / decoded. To entropy encode / decode those types of information, information about neighboring coding units adjacent to the current coding unit can be used. For example, it is highly likely that the partitioning type (partitioned or unpartitioned, partitioning tree, and / or partitioning direction) of the left-hand neighboring coding unit and / or above-hand neighboring coding unit is similar to the partitioning type of the current coding unit. Therefore, contextual information for entropy encoding / decoding of information about the current coding unit can be derived from the information about neighboring coding units. Information about neighboring coding units may include at least one of the following: quadtree partitioning information, multi-type tree partitioning indication information, partitioning direction information, and partitioning tree information.
[0194] As another example, in binary tree partitioning and ternary tree partitioning, binary tree partitioning can be performed first. That is, the current coding unit can first undergo binary tree partitioning, and then the coding unit corresponding to the leaf node of the binary tree can be set as the root node for ternary tree partitioning. In this case, for the coding unit corresponding to the node of the ternary tree, neither quadtree partitioning nor binary tree partitioning can be performed.
[0195] Encoding units that cannot be partitioned according to quadtree, binary tree, and / or ternary tree partitioning structures become the basic units for encoding, prediction, and / or transformation. In other words, these encoding units cannot be further partitioned for prediction and / or transformation. Therefore, partitioning structure information and partitioning information for dividing encoding units into prediction and / or transformation units may not exist in the bitstream.
[0196] However, when the size of the coding unit (i.e., the basic unit used for partitioning) is larger than the size of the maximum transform block, the coding unit can be partitioned recursively until the size of the coding unit is reduced to be equal to or smaller than the size of the maximum transform block. For example, when the size of the coding unit is 64×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into four 32×32 blocks for transformation. For example, when the size of the coding unit is 32×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into two 32×32 blocks for transformation. In this case, the partitioning of the coding unit for transformation is not sent separately with a signal, and the partitioning of the coding unit for transformation can be determined by comparing the horizontal or vertical dimensions of the coding unit with the horizontal or vertical dimensions of the maximum transform block. For example, when the horizontal dimension (width) of the coding unit is greater than the horizontal dimension (width) of the maximum transform block, the coding unit can be vertically bisected. For example, when the vertical dimension (height) of the coding unit is greater than the vertical dimension (height) of the maximum transform block, the coding unit can be horizontally bisected.
[0197] Information regarding the maximum and / or minimum size of the coding unit and the maximum and / or minimum size of the transform block can be transmitted or determined by signals at the higher level of the coding unit. The higher level can be, for example, a sequence level, a frame level, a stripe level, a parallel block group level, a parallel block level, etc. For example, the minimum size of the coding unit can be determined to be 4×4. For example, the maximum size of the transform block can be determined to be 64×64. For example, the minimum size of the transform block can be determined to be 4×4.
[0198] Information regarding the minimum size of the coding unit corresponding to the leaf node of the quadtree (minimum size of the quadtree) and / or the maximum depth from the root node to the leaf node of the multi-type tree (maximum depth of the multi-type tree) can be signaled or determined at the higher level of the coding unit. For example, the higher level can be the sequence level, frame level, stripe level, parallel block group level, parallel block level, etc. Information regarding the minimum size of the quadtree and / or the maximum depth of the multi-type tree can be signaled or determined for each of the intra-frame stripes and inter-frame stripes.
[0199] The difference between the size of the CTU and the maximum size of the transform block can be signaled or determined at the higher level of the coding unit. For example, the higher level can be a sequence level, frame level, stripe level, parallel block group level, parallel block level, etc. The maximum size of the coding unit corresponding to each node of the binary tree (hereinafter referred to as the maximum size of the binary tree) can be determined based on the size of the coding tree unit and the difference information. The maximum size of the coding unit corresponding to each node of the ternary tree (hereinafter referred to as the maximum size of the ternary tree) can vary depending on the type of stripe. For example, for an intra-frame stripe, the maximum size of the ternary tree can be 32×32. For example, for an inter-frame stripe, the maximum size of the ternary tree can be 128×128. For example, the minimum size of the coding unit corresponding to each node of the binary tree (hereinafter referred to as the minimum size of the binary tree) and / or the minimum size of the coding unit corresponding to each node of the ternary tree (hereinafter referred to as the minimum size of the ternary tree) can be set as the minimum size of the coding block.
[0200] As another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the stripe level. Alternatively, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the stripe level.
[0201] Based on the size and depth information of the various blocks mentioned above, four-partition information, multi-type tree partition indication information, partition tree information and / or partition direction information may or may not be included in the bitstream.
[0202] For example, when the size of the coding unit is no greater than the minimum size of the quadtree, the coding unit does not include four-partition information. The four-partition information can be inferred as a second value.
[0203] For example, when the size (horizontal and vertical dimensions) of the coding unit corresponding to a node of a multi-type tree is greater than the maximum size (horizontal and vertical dimensions) of a binary tree and / or a ternary tree, the coding unit may not be divided into two or three partitions. Therefore, multi-type tree partition indication information can be sent without a signal, but the multi-type tree partition indication information can be inferred as a second value.
[0204] Optionally, when the size (horizontal and vertical dimensions) of the coding unit corresponding to a node of a multi-type tree is the same as the maximum size (horizontal and vertical dimensions) of a binary tree and / or twice the maximum size (horizontal and vertical dimensions) of a ternary tree, the coding unit may not be further divided into two or three partitions. Therefore, multi-type tree partitioning indication information does not need to be sent by signal, but the multi-type tree partitioning indication information can be inferred as a second value. This is because when the coding unit is partitioned according to the binary tree partitioning structure and / or the ternary tree partitioning structure, coding units smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree are generated.
[0205] Optionally, binary or ternary partitioning can be limited based on the size of the virtual pipeline data unit (hereinafter, the pipeline buffer size). For example, when a coding unit is divided into sub-coding units that are not suitable for the pipeline buffer size through binary or ternary partitioning, the corresponding binary or ternary partitioning may be limited. The pipeline buffer size can be the size of the largest transform block (e.g., 64×64). For example, when the pipeline buffer size is 64×64, the following partitioning can be limited.
[0206] - N×M (N and / or M are 128) ternary tree partitions for encoding units
[0207] - 128×N (N<=64) binary tree partitioning for the horizontal direction of the encoding unit
[0208] - N×128 (N<=64) binary tree partitions used in the vertical direction of the encoding unit
[0209] Optionally, when the depth of the coding unit corresponding to a node in the multi-type tree is equal to the maximum depth of the multi-type tree, the coding unit may not be further divided into two and / or three partitions. Therefore, multi-type tree partition indication information may not be sent by signal, but the multi-type tree partition indication information can be inferred as a second value.
[0210] Optionally, multi-type tree partitioning indication information may be signaled only if at least one of the vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is possible for the coding unit corresponding to the node of the multi-type tree. Otherwise, the coding unit may not be partitioned into two and / or three partitions. Therefore, multi-type tree partitioning indication information may not be signaled, but the multi-type tree partitioning indication information may be inferred as a second value.
[0211] Optionally, partition direction information may be signaled only if both vertical binary tree partitioning and horizontal binary tree partitioning, or both vertical ternary tree partitioning and horizontal ternary tree partitioning, are possible for the coding units corresponding to nodes of multiple tree types. Otherwise, partition direction information may not be signaled, but the partition direction information may be inferred as a value indicating the possible partition direction.
[0212] Optionally, partition tree information may be signaled only if both vertical binary tree partitions and vertical ternary tree partitions, or both horizontal binary tree partitions and horizontal ternary tree partitions, are possible for the encoded tree corresponding to the nodes of the multi-type tree. Otherwise, partition tree information may not be signaled, but the partition tree information may be inferred as a value indicating a possible partition tree structure.
[0213] FIG. 4 This is a diagram illustrating intra-frame prediction processing.
[0214] FIG. 4 The arrows from the center outwards indicate the prediction direction of the intra-frame prediction mode.
[0215] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the current block. A neighboring block can be a reconstructed neighboring block. For example, intra-frame coding and / or decoding can be performed using values of reference samples or coding parameters included in the reconstructed neighboring block.
[0216] A prediction block can represent a block generated by performing intra-frame prediction. A prediction block can correspond to at least one of CU, PU, and TU. The cells of a prediction block can have the size of one of CU, PU, and TU. A prediction block can be a square block with dimensions such as 2×2, 4×4, 16×16, 32×32, or 64×64, or a rectangular block with dimensions such as 2×8, 4×8, 2×16, 4×16, and 8×16.
[0217] Intra-prediction can be performed based on the intra-prediction mode for the current block. The number of intra-prediction modes that the current block can have can be a fixed value, or it can be a value determined differently depending on the attributes of the predicted block. For example, the attributes of the predicted block can include the size and shape of the predicted block.
[0218] Regardless of the block size, the number of intra-prediction modes can be fixed at N. Alternatively, the number of intra-prediction modes can be 3, 5, 9, 17, 34, 35, 36, 65, or 67, etc. Optionally, the number of intra-prediction modes can vary depending on the block size or the color component type, or both. For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chrominance signal. For example, the number of intra-prediction modes can increase as the block size increases. Optionally, the number of intra-prediction modes for the luma component block can be greater than the number of intra-prediction modes for the chrominance component block.
[0219] Intra-prediction modes can be non-angular or angular. Non-angular modes can be DC or planar modes, and angular modes can be prediction modes with a specific direction or angle. Intra-prediction modes can be represented by at least one of mode number, mode value, mode number, mode angle, and mode direction. The number of intra-prediction modes can be greater than 1M, including both non-angular and angular modes. To perform intra-prediction on the current block, a step can be performed to determine whether a sample included in the reconstructed neighboring block can be used as a reference sample for the current block. When there are samples that cannot be used as reference samples for the current block, the value obtained by copying or interpolating at least one sample value included in the reconstructed neighboring block, or both, can be used to replace the unavailable sample value, and the replaced sample value is used as the reference sample for the current block.
[0220] FIG. 7 This is a diagram showing reference samples that can be used for intra-frame prediction.
[0221] like FIG. 7 As shown, at least one of reference sample lines 0 to 3 can be used for intra-frame prediction of the current block. FIG. 7 In this process, samples from fragments A and F can be filled using samples from the nearest fragments B and E, respectively, instead of being retrieved from reconstructed neighboring blocks. Index information of the reference sample lines to be used for intra-frame prediction of the current block can be transmitted using signals. For example, in... FIG. 7 In this configuration, reference sample line indicators 0, 1, and 2 can be signaled as index information indicating reference sample line 0, reference sample line 1, and reference sample line 2. When the upper boundary of the current block is the boundary of the CTU, only reference sample line 0 can be available. Therefore, in this case, index information does not need to be signaled. When reference sample lines other than reference sample line 0 are used, filtering for the prediction block, which will be described later, is not required.
[0222] When performing intra-frame prediction, filters can be applied to at least one of the reference samples and the prediction samples based on the intra-frame prediction mode and the current block size.
[0223] In planar mode, when generating the prediction block for the current block, the sample value of the target sample is generated by using a weighted sum of the upper and left reference samples, and the upper right and lower left reference samples of the current block, based on the position of the target sample within the prediction block. Furthermore, in DC mode, the average of the upper and left reference samples of the current block can be used when generating the prediction block. Additionally, in angled mode, the prediction block can be generated using the upper, left, upper right, and / or lower left reference samples of the current block. Real-valued interpolation can be performed to generate the prediction sample values.
[0224] In the case of intra-frame prediction between color components, a predicted block for the current block of the second color component can be generated based on the corresponding reconstructed block of the first color component. For example, the first color component can be a luma component, and the second color component can be a chroma component. For intra-frame prediction between color components, the parameters of a linear model between the first and second color components can be derived based on a template. The template may include the upper and / or left neighboring samples of the current block and the upper and / or left neighboring samples of the reconstructed block of the corresponding first color component. For example, the parameters of the linear model can be derived using the sample value of the first color component with the maximum value in the template and its corresponding sample value of the second color component, and the sample value of the first color component with the minimum value in the template and its corresponding sample value of the second color component. When deriving the parameters of the linear model, the corresponding reconstructed block can be applied to the linear model to generate a predicted block for the current block. Depending on the video format, subsampling can be performed on the neighboring samples and the corresponding reconstructed block of the reconstructed block of the first color component. For example, when a sample point of the second color component corresponds to four samples of the first color component, the four samples of the first color component can be subsampled to calculate a corresponding sample point. In this case, parameter derivation of the linear model and intra-frame prediction between color components can be performed based on the corresponding subsampled sample points. Whether to perform intra-frame prediction between color components and / or the range of the template can be sent as an intra-frame prediction mode signal.
[0225] The current block can be partitioned into two or four sub-blocks, either horizontally or vertically. The partitioned sub-blocks can be reconstructed sequentially. That is, intra-prediction can be performed on the sub-blocks to generate sub-prediction blocks. Furthermore, inverse quantization and / or inverse transform can be performed on the sub-blocks to generate sub-residual blocks. Reconstructed sub-blocks can be generated by adding the sub-prediction blocks to the sub-residual blocks. The reconstructed sub-blocks can be used as reference samples for intra-prediction of subsequent sub-blocks. A sub-block can be a block comprising a predetermined number (e.g., 16) or more samples. Therefore, for example, when the current block is an 8×4 or 4×8 block, the current block can be partitioned into two sub-blocks. Furthermore, when the current block is a 4×4 block, the current block may not be partitioned into sub-blocks. When the current block has other sizes, the current block can be partitioned into four sub-blocks. Information regarding whether intra-prediction is performed based on sub-blocks and / or partitioning direction (horizontal or vertical) can be signaled. Intra-prediction based on sub-blocks can be limited to being performed only when using reference sample line 0. When performing sub-block-based intra-frame prediction, filtering for the prediction block, which will be described later, may not be performed.
[0226] A final prediction block can be generated by performing filtering on the prediction block that has been intra-predicted. Filtering can be performed by applying predetermined weights to the target sample, the left reference sample, the top reference sample, and / or the top-left reference sample. The weights and / or reference samples (range, position, etc.) used for filtering can be determined based on at least one of the block size, the intra-prediction mode, and the position of the target sample in the prediction block. Filtering can be performed only in a predetermined intra-prediction mode (e.g., DC, planar, vertical, horizontal, diagonal, and / or adjacent diagonal mode). An adjacent diagonal mode can be a mode that is diagonal mode plus k or subtracted from diagonal mode. For example, k can be a positive integer of 8 or less.
[0227] The intra-prediction mode of the current block can be entropy-coded / decoded by predicting the intra-prediction modes of adjacent blocks. When the intra-prediction modes of the current block and its neighboring blocks are the same, information indicating that the intra-prediction modes of the current block and its neighboring blocks are the same can be signaled using predetermined flag information. Furthermore, an indicator of the intra-prediction mode among multiple neighboring blocks that is the same as the intra-prediction mode of the current block can be signaled. When the intra-prediction modes of the current block and its neighboring blocks are different, the intra-prediction mode information of the current block can be entropy-coded / decoded by performing entropy coding / decoding based on the intra-prediction modes of neighboring blocks.
[0228] FIG. 5 This is a diagram illustrating an embodiment of inter-screen prediction processing.
[0229] exist FIG. 5 In this context, rectangles can represent the image. FIG. 5In the image, the arrow indicates the prediction direction. Based on the encoding type of the frame, frames can be classified into intra-frame frames (I-frames), predictive frames (P-frames), and dual-predictive frames (B-frames).
[0230] I-frames can be encoded via intra-frame prediction without requiring inter-frame prediction. P-frames can be encoded via inter-frame prediction using reference frames present in one direction (i.e., forward or backward) relative to the current block. B-frames can be encoded via inter-frame prediction using reference frames present in both directions (i.e., forward and backward) relative to the current block. When using inter-frame prediction, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform the corresponding motion compensation.
[0231] The following section will describe in detail an embodiment of inter-screen prediction.
[0232] Reference frames and motion information can be used to perform inter-frame prediction or motion compensation.
[0233] Motion information of the current block can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction. The motion information of the current block can be derived using motion information of reconstructed neighboring blocks, motion information of co-located blocks (also called col blocks or co-position blocks), and / or motion information of blocks adjacent to the co-position block. A co-position block can represent a block within a previously reconstructed co-located frame (also called a col frame or co-position frame) that is spatially located at the same position as the current block. A co-position frame can be one of one or more reference frames included in a list of reference frames.
[0234] The methods for deriving motion information can vary depending on the prediction mode of the current block. For example, prediction modes applied to inter-frame prediction include AMVP mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, geometric partitioning mode, combined inter-frame-intra-frame prediction mode, affine mode, etc. Here, the merge mode can be referred to as motion merge mode.
[0235] For example, when AMVP is used as a prediction mode, at least one of the motion vectors of reconstructed neighboring blocks, co-located blocks, blocks adjacent to co-located blocks, and (0,0) motion vectors can be identified as motion vector candidates for the current block, and a motion vector candidate list is generated using these motion vector candidates. Motion vector candidates for the current block can be derived using the generated motion vector candidate list. Motion information for the current block can be determined based on the derived motion vector candidates. The motion vectors of co-located blocks or blocks adjacent to co-located blocks can be referred to as temporal motion vector candidates, and the motion vectors of reconstructed neighboring blocks can be referred to as spatial motion vector candidates.
[0236] Encoding device 100 can calculate the motion vector difference (MVD) between the motion vector of the current block and motion vector candidates, and can perform entropy encoding on the motion vector difference (MVD). Furthermore, encoding device 100 can perform entropy encoding on the motion vector candidate index and generate a bitstream. The motion vector candidate index indicates the best motion vector candidate among the motion vector candidates included in the motion vector candidate list. Decoding device 200 can perform entropy decoding on the motion vector candidate index included in the bitstream, and can select motion vector candidates for the target block to be decoded from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index. Furthermore, decoding device 200 can add the entropy-decoded MVD to the motion vector candidate extracted by entropy decoding, thereby deriving the motion vector of the target block to be decoded.
[0237] Additionally, the encoding device 100 can perform entropy encoding on the calculated MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the resolution of the entropy-decoded MVD.
[0238] Additionally, the encoding device 100 calculates the motion vector difference (MVD) between the motion vectors and motion vector candidates in the current block based on an affine model, and performs entropy encoding on the MVD. The decoding device 200 derives motion vectors based on each sub-block by deriving the affine control motion vector of the decoded target block based on the sum of the entropied MVD and the affine control motion vector candidates.
[0239] The bitstream may include a reference frame index indicating a reference frame. The reference frame index may be entropy encoded by the encoding device 100 and subsequently transmitted as a bitstream to the decoding device 200. The decoding device 200 may generate a predicted block for the decoded target block based on the derived motion vectors and the reference frame index information.
[0240] Another example of a method for deriving motion information for the current block could be a merge pattern. A merge pattern can represent a method for merging the motion of multiple blocks. A merge pattern can represent a pattern for deriving motion information for the current block from the motion information of neighboring blocks. When applying a merge pattern, a list of merge candidates can be generated using reconstructed motion information of neighboring blocks and / or motion information of blocks at the same location. Motion information may include at least one of motion vectors, reference frame indices, and inter-frame prediction indicators. The prediction indicators may indicate unidirectional prediction (L0 prediction or L1 prediction) or bidirectional prediction (L0 prediction and L1 prediction).
[0241] The merge candidate list can be a list of stored motion information. The motion information included in the merge candidate list can be at least one of the following: motion information of neighboring blocks adjacent to the current block (spatial merge candidate), motion information of blocks at the same position in the reference frame of the current block (temporal merge candidate), new motion information generated by combining motion information existing in the merge candidate list, motion information of blocks encoded / decoded before the current block (history-based merge candidate), and zero merge candidate.
[0242] Encoding device 100 can generate a bitstream by performing entropy encoding on at least one of a merge flag and a merge index, and can transmit the bitstream as a signal to decoding device 200. The merge flag may be information indicating whether a merge mode is performed for each block, and the merge index may be information indicating which neighboring block among the current block's neighboring blocks is the target block for merging. For example, the neighboring blocks of the current block may include a left neighboring block located to the left of the current block, an upper neighboring block arranged above the current block, and a time neighboring block that is temporally adjacent to the current block.
[0243] Additionally, the encoding device 100 performs entropy encoding on the correction information for correcting motion vectors in the motion information of the merging candidates and sends it as a signal to the decoding device 200. The decoding device 200 can correct the motion vectors of the merging candidates selected by the merging index based on the correction information. Here, the correction information may include at least one of information on whether correction is performed, correction direction information, and correction size information. As described above, the prediction mode that corrects the motion vectors of the merging candidates based on the correction information sent as a signal can be referred to as a merging mode with motion vector difference.
[0244] Skip mode can be a mode in which motion information of neighboring blocks is applied to the current block as is. When skip mode is applied, encoding device 100 can perform entropy encoding on information about which block's motion information will be used as the current block's motion information to generate a bitstream, and can send the bitstream to decoding device 200 as a signal. Encoding device 100 may not send syntax elements regarding at least one of motion vector difference information, coded block flags, and transform coefficient levels to decoding device 200 as a signal.
[0245] Sub-block merging patterns can represent a pattern for deriving motion information on a sub-block basis within a coded block (CU). When applying a sub-block merging pattern, a list of sub-block merging candidates can be generated using motion information of sub-blocks at the same position as the current sub-block in a reference image (based on sub-block temporal merging candidates) and / or affine control point motion vector merging candidates.
[0246] A geometric partitioning pattern can be represented as follows: motion information is derived by partitioning the current block in a predetermined direction, each of the derived motion information is used to derive each prediction sample, and the prediction sample of the current block is derived by weighting each of the derived prediction samples.
[0247] The inter-frame-intra-frame combined prediction mode can be represented as a mode for deriving the prediction samples of the current block by weighting the prediction samples generated by inter-frame prediction and the prediction samples generated by intra-frame prediction.
[0248] The decoding device 200 can self-correct the derived motion information. The decoding device 200 can search a predetermined region based on a reference block indicated by the derived motion information and derive motion information with minimum SAD as corrected motion information.
[0249] Decoding device 200 can use optical flow to compensate for prediction samples derived via inter-frame prediction.
[0250] FIG. 6 This is a diagram illustrating the transformation and quantization processes.
[0251] like FIG. 6 As shown, transform and / or quantization are performed on the residual signal to generate a quantized level signal. The residual signal is the difference between the original block and the predicted block (i.e., an intra-frame predicted block or an inter-frame predicted block). The predicted block is generated through intra-frame prediction or inter-frame prediction. The transform can be a primary transform, a secondary transform, or both. The primary transform of the residual signal generates transform coefficients, and the secondary transform of the transform coefficients generates secondary transform coefficients.
[0252] At least one scheme selected from a variety of predefined transform schemes is used to perform the primary transform. Examples of the predefined transform schemes include the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and Karhunen-Loève Transform (KLT). The transform coefficients generated by the primary transform may undergo secondary transforms. The transform scheme used for the primary and / or secondary transforms can be determined based on the coding parameters of the current block and / or its neighboring blocks. Optionally, transform information indicating the transform scheme can be transmitted via a signal. DCT-based transforms may include, for example, DCT-2, DCT-8, etc. DST-based transforms may include, for example, DST-7.
[0253] A quantized level signal (quantization coefficients) can be generated by performing quantization on the residual signal or on the result of performing a primary transform and / or a secondary transform. Depending on the intra-frame prediction mode or block size / shape, the quantized level signal can be scanned using at least one of diagonal top-right scan, vertical scan, and horizontal scan. For example, when scanning coefficients according to a diagonal top-right scan, the block-form coefficients change to a one-dimensional vector form. In addition to the diagonal top-right scan, depending on the intra-frame prediction mode and / or the size of the transform block, a horizontal scan that horizontally scans the coefficients in two-dimensional block form or a vertical scan that vertically scans the coefficients in two-dimensional block form can be used. The scanned quantized level coefficients can be entropy-encoded for insertion into the bitstream.
[0254] The decoder performs entropy decoding on the bitstream to obtain quantized level coefficients. The quantized level coefficients can be arranged in a two-dimensional block format via inverse scanning. For inverse scanning, at least one of diagonal top-right scanning, vertical scanning, and horizontal scanning can be used.
[0255] The quantized level coefficients can then be dequantized, then subjected to a secondary inverse transform as needed, and finally subjected to a primary inverse transform as needed to generate the reconstructed residual signal.
[0256] Inverse mapping in the dynamic range can be performed on the luma components reconstructed via intra-frame or inter-frame prediction prior to intra-loop filtering. The dynamic range can be divided into 16 equal segments, and a mapping function for each segment can be signaled. The mapping function can be signaled at the strip level or parallel block group level. The inverse mapping function used to perform the inverse mapping can be derived based on the mapping function. Intra-loop filtering, reference frame storage, and motion compensation are performed in the inverse mapping region, and the prediction blocks generated via inter-frame prediction are transformed to the mapping region via mapping using the mapping function and then used to generate reconstructed blocks. However, since intra-frame prediction is performed in the mapping region, the prediction blocks generated via intra-frame prediction can be used to generate reconstructed blocks without mapping / inverse mapping.
[0257] When the current block is a residual block for the chroma component, the residual block can be converted to the inverse-mapped region by scaling the chroma components of the mapped region. The availability of scaling can be signaled at the stripe level or the parallel block group level. Scaling can only be applied if mapping for the luma component is available and the partitioning of the luma component and the partitioning of the chroma component follow the same tree structure. Scaling can be performed based on the average of the sample values of the luma prediction block corresponding to the chroma block. In this case, when inter-frame prediction is used for the current block, the luma prediction block can represent the mapped luma prediction block. The required scaling value can be derived by using an index-referenced lookup table of the segment to which the average of the sample values of the luma prediction block belongs. Finally, by scaling the residual block using the derived value, the residual block can be converted to the inverse-mapped region. Chroma component block recovery, intra-frame prediction, inter-frame prediction, intra-loop filtering, and reference frame storage can then be performed in the inverse-mapped region.
[0258] Information indicating whether the mapping / inverse mapping of the luminance and chrominance components is available can be sent via a sequence parameter set using signals.
[0259] A predicted block for the current block can be generated based on a block vector indicating the displacement between the current block and a reference block in the current frame. In this way, the prediction mode used to generate a predicted block with reference to the current frame is called Intra-Block Copy (IBC) mode. IBC mode can be applied to M×N (M<=64, N<=64) coding units. IBC modes can include skip mode, merge mode, AMVP mode, etc. In the case of skip mode or merge mode, a merge candidate list is constructed, and a merge index is signaled so that a merge candidate can be specified. The block vector of the specified merge candidate can be used as the block vector of the current block. The merge candidate list can include at least one of spatial candidates, history-based candidates, candidates based on the average of two candidates, and zero merge candidates. In the case of AMVP mode, a difference block vector can be signaled. Furthermore, the predicted block vector can be derived from the left neighboring block and the upper neighboring block of the current block. The index of the neighboring block to be used can be signaled. The predicted block in IBC mode is included in the current CTU or the left CTU and is limited to blocks in the already reconstructed area. For example, the value of the block vector can be restricted such that the predicted block of the current block is located in the region of the three 64×64 blocks preceding the 64×64 block to which the current block belongs, in the encoding / decoding order. By restricting the value of the block vector in this way, memory consumption and device complexity can be reduced in implementations according to the IBC mode.
[0260] The following describes a method for enhancing video compression efficiency by improving the transform method, which is one of the video coding processes. More specifically, coding in conventional video coding schematically includes: an intra-frame / inter-frame prediction step, predicting an original block that is part of the current original image; a transform and quantization step for a residual block, wherein the residual block is the difference between the predicted block and the original block; and an entropy coding step, which is a lossless compression method based on the coefficients of the block that has already undergone transform and quantization and the probability of compression information obtained in the previous steps. Thus, a bitstream as a compressed form of the original image is generated and sent to a decoder or stored in a recording medium. The shuffling and discrete sine transform (hereinafter referred to as "SDST"), which will be described below in this specification, aims to enhance compression efficiency by improving transform efficiency.
[0261] The SDST method according to the present invention uses Discrete Sine Transform Type-7 (hereinafter referred to as "DST-VII" or "DST-7") instead of Discrete Cosine Transform Type-2 (hereinafter referred to as "DCT-II" or "DCT-2"), which is widely used as the transform kernel in video coding, thereby better reflecting the common frequency characteristics of the image.
[0262] According to the transformation method of the present invention, high objective video quality can be obtained even at a relatively low bit rate compared with traditional video coding methods.
[0263] DST-7 can be applied to the data of the residual block. The application of DST-7 to the residual block can be performed based on the prediction mode corresponding to the residual block. For example, it can be applied to residual blocks encoded in inter-frame mode. According to embodiments of the invention, DST-7 can be applied after rearranging or shuffling the data of the residual block. Here, shuffling can refer to the rearrangement of image data and can be equivalent to rearranging or flipping the residual signal. Here, residual block can have the same meaning as residual, remaining block, remaining signal, residual signal, remaining data, or residual data. Furthermore, residual block can have the same meaning as reconstructed residual, reconstructed remaining block, reconstructed remaining signal, reconstructed residual signal, reconstructed remaining data, or reconstructed residual data in the encoder and decoder in the form of reconstructed residual block.
[0264] According to an embodiment of the present invention, SDST can use DST-7 as the transform kernel. Here, the transform kernel of SDST is not limited to DST-7, and at least one of various types of DST and DCT (such as Discrete Sine Transform Type-1 (DST-1), Discrete Sine Transform Type-2 (DST-2), Discrete Sine Transform Type-3 (DST-3), ..., Discrete Sine Transform Type-n (DST-n), Discrete Cosine Transform Type-1 (DCT-1), Discrete Cosine Transform Type-2 (DCT-2), Discrete Cosine Transform Type-3 (DCT-3), ..., Discrete Cosine Transform Type-n (DCT-n) etc.) can be used (here, n can be a positive integer 1 or a larger positive integer).
[0265] Equation 1 below represents a method for performing one-dimensional DCT-2 according to an embodiment of the present invention. Here, N may represent the size of the block, k may represent the position of the frequency component, and x n It can represent the value of the nth coefficient in the spatial domain.
[0266] [Equation 1]
[0267]
[0268] DCT-2 in the two-dimensional domain can be achieved by performing horizontal and vertical transformations on the residual block using Equation 1 above.
[0269] The DCT-2 transform kernel can be defined as shown in Equation 2 below. Here, X k It can represent the basis vectors based on their positions in the frequency domain, and N can represent the size of the frequency domain.
[0270] [Equation 2]
[0271]
[0272] in addition, FIG. 7 This is a diagram showing the basis vectors in the frequency domain of the DCT-2 according to the present invention. FIG. 7 The frequency characteristics of DCT-2 in the frequency domain are shown. Here, the value calculated from the X0 basis vector of DCT-2 represents the DC component.
[0273] DCT-2 can be used in the transformation processing of residual blocks with sizes such as 4×4, 8×8, 16×16, 32×32, etc.
[0274] Additionally, DCT-2 can be selectively used based on at least one of the following: the size of the residual block, the color components of the residual block (e.g., luma and chroma components), and the prediction mode corresponding to the residual block. For example, DCT-2 is used when the residual block has a size of 4×4 and is encoded in intra-frame mode, and the components of the residual block are luma components. For example, a first transform kernel can be used for horizontal transform when the horizontal length (width) of the residual block encoded in intra-frame mode is within a predetermined range (e.g., equal to or greater than four pixels and equal to or less than 16 pixels) and the horizontal length (width) is not longer than the vertical length (height). Otherwise, a second transform kernel can be used for horizontal transform. For example, a first transform kernel can be used for vertical transform when the vertical length (height) of the residual block encoded in intra-frame mode is equal to or greater than four pixels and equal to or less than 16 pixels, and the vertical length (height) is not longer than the horizontal length (width). Otherwise, a second transform kernel can be used for vertical transform. The first transform kernel may be different from the second transform kernel. In other words, the horizontal and vertical transform methods for blocks encoded in intra-frame mode can be implicitly determined based on the block shape under predetermined conditions. For example, the first transform kernel can be DST-7, and the second transform kernel can be DCT-2. Here, the residual block is the transform target, and therefore it can have the same meaning as the transform block. Here, the prediction mode can represent inter-frame prediction or intra-frame prediction. Furthermore, in the case of intra-frame prediction, the prediction mode represents the intra-frame prediction mode or intra-frame prediction direction.
[0275] Transforms using the DCT-2 transform kernel can achieve high compression efficiency for blocks with characteristics where changes between neighboring pixels are as small as the background of an image. However, it may not be suitable as a transform kernel for regions with complex patterns, such as textured images. This is because when blocks with low correlation between neighboring pixels are transformed by DCT-2, a large number of transform coefficients appear in the high-frequency components of the frequency domain. When transform coefficients are frequently generated in the high-frequency domain, the compression efficiency of the image may decrease. To improve compression efficiency, coefficients with large values need to appear near the low-frequency components, and the values of the coefficients need to be close to zero in the high-frequency components.
[0276] Equation 3 below represents a method for performing one-dimensional DST-7 according to an embodiment of the present invention. Here, N may represent the size of the block, k may represent the position of the frequency component, and x n It can represent the value of the nth coefficient in the spatial domain.
[0277] [Equation 3]
[0278]
[0279] DST-7 in the two-dimensional domain can be achieved by performing horizontal and vertical transformations on the residual block using Equation 3 above.
[0280] The DST-7 transform kernel can be defined as shown in Equation 4 below. Here, X k It can represent the k-th basis vector of DST-7, i can represent the position in the frequency domain, and N can represent the size of the frequency domain.
[0281] [Equation 4]
[0282]
[0283] DST-7 can be used in the transformation processing of residual blocks with a size of at least one of 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, 128×128, etc.
[0284] Additionally, DST-7 can be applied to rectangular blocks instead of square blocks. For example, DST-7 can be applied to at least one of the vertical and horizontal transformations of rectangular blocks with different horizontal and vertical dimensions (such as 8×4, 16×8, 32×4, 64×16, etc.). When multiple transformation methods can be selectively applied, DCT-2 is applied to the horizontal and vertical transformations of square blocks. When multiple transformation methods cannot be selectively applied, DST-7 is applied to the horizontal and vertical transformations of square blocks.
[0285] Furthermore, DST-7 can be selectively used based on at least one of the following: the size of the residual block, the color components of the residual block (e.g., luma and chroma components), the prediction mode corresponding to the residual block, the intra-prediction mode (direction), and the shape of the residual block. For example, DST-7 is used when the residual block has a size of 4×4 and is encoded in intra-prediction mode, and the components of the residual block are luma components. Here, the prediction mode can represent inter-frame prediction or intra-frame prediction. Furthermore, in the case of intra-prediction, the prediction mode represents the intra-prediction mode or the intra-prediction direction. For example, for the chroma component, the selection of the transformation method based on the block shape may be unavailable. For example, when the intra-prediction mode is a prediction between color components, the selection of the transformation method based on the block shape is unavailable. For example, the transformation method for the chroma component can be specified by information transmitted by signal via a bitstream. When the current block is partitioned into multiple sub-blocks and intra-prediction is performed on each of the multiple sub-blocks, the transformation method for the current block is determined based on the intra-prediction mode and / or the block size (horizontal size and / or vertical size). For example, when the intra-prediction mode is non-directional (DC or planar) and the horizontal length (width) (or vertical length (height)) is within a predetermined range, a first transform kernel is used for horizontal (vertical) transformation. Otherwise, a second transform kernel is used. The first transform kernel may be different from the second transform kernel. For example, the first transform kernel may be DST-7, and the second transform kernel may be DCT-2. The predetermined range may, for example, be from 4 pixels to 16 pixels. When the block size is not within the predetermined range, the same kernel (e.g., the second transform kernel) is used for both horizontal and vertical transformations. When the block size is within the predetermined range, different transform kernels are used for intra-prediction modes that are adjacent to each other. For example, when the second and first transform kernels are used for horizontal and vertical transformations in mode 27, respectively, the first and second transform kernels are used for horizontal and vertical transformations in modes 26 and 28, which are adjacent to mode 27, respectively.
[0286] The following text describes SDST as one of the transformation methods that uses DST-7 as the transformation kernel.
[0287] In the following text, a block may represent one of CU, PU, and TU.
[0288] The SDST according to the present invention can be performed in two steps. The first step is to perform shuffling on the residual signal within the PU of the predicted CU in inter-frame mode or intra-frame mode. The second step is to apply DST-7 to the residual signal within the block that has already been shuffling.
[0289] The residual signals arranged within the current block (e.g., CU, PU, or TU) can be scanned in a first direction and rearranged in a second direction. That is, the residual signals arranged within the current block can be scanned in the first direction and rearranged in the second direction to perform a commingling. Here, the residual signal can represent a signal indicating the difference between the original signal and the predicted signal. That is, the residual signal can represent the signal before performing at least one of transform and quantization. Optionally, the residual signal can represent the signal form after performing at least one of transform and quantization. Furthermore, the residual signal can represent the reconstructed residual signal. That is, the residual signal can represent the signal after performing at least one of inverse transform and dequantization. Furthermore, the residual signal can represent the signal before performing at least one of inverse transform and dequantization.
[0290] Additionally, the first direction (or scanning direction) can be one of the following: raster scanning sequence, upper right diagonal scanning sequence, horizontal scanning sequence, and vertical scanning sequence. Furthermore, the first direction can be defined as at least one of the following (1) to (10).
[0291] Scan from top to bottom rows, and from left to right within a row.
[0292] Scan from top to bottom rows, and from right to left within a row.
[0293] Scan from bottom row to top row, and from left to right within a row.
[0294] Scan from bottom row to top row, and from right to left within a row.
[0295] Scan from the left column to the right column, and from top to bottom within a column.
[0296] Scan from the left column to the right column, and from bottom to top within a column.
[0297] Scan from the right column to the left column, and from top to bottom within a column.
[0298] Scan from the right column to the left column, and from bottom to top within a column.
[0299] Spiral scan: Scanning from the inside (or outside) of the block towards the outside (or inside) of the block, and scanning in a clockwise / counterclockwise direction.
[0300] Diagonal scan: Starting from a vertex within the block, scan diagonally along the upper left, upper right, lower left, or lower right directions.
[0301] Additionally, regarding the second direction (or rearrangement direction), at least one of the scanning directions (1) to (10) may be used selectively. The first direction and the second direction may be the same or different from each other.
[0302] It can perform scanning and rearrangement of residual signals on a block-by-block basis.
[0303] Here, rearrangement can mean that the residual signals scanned in the first direction within a block are arranged in a second direction within blocks of the same size. The size of the block used for scanning in the first direction may differ from the size of the block used for rearrangement in the second direction.
[0304] Furthermore, scanning and rearrangement are described as being performed separately according to a first direction and a second direction, but scanning and rearrangement can also be performed as a single process for the first direction. For example, for residual signals within a block, scanning can be performed from top row to bottom row and from right to left within a row to store (rearrange) them in the block.
[0305] Additionally, scanning and rearrangement processing for the residual signal can be performed within predetermined units of sub-blocks within the current block. Here, a sub-block can be a block with a size equal to or smaller than the current block. A sub-block can be a block obtained by partitioning the current block using a quadtree, binary tree, or similar method.
[0306] Sub-block units may have fixed sizes and / or shapes (e.g., 4×4, 4×8, 8×8, ..., N×M, where N and M are positive integers). Furthermore, the size and / or shape of sub-block units can be variably derived. For example, the size and / or shape of a sub-block unit can be determined based on the size, shape, and / or prediction mode (inter-frame and intra-frame) of the current block.
[0307] The scan direction and / or rearrangement direction can be adaptively determined based on the position of the sub-blocks. In this case, different scan directions and / or rearrangement directions can be used for sub-blocks, or all or part of the sub-blocks of the current block can use the same scan direction and / or the same rearrangement direction.
[0308] For example, for a block predicted inter-frame, a residual block of the same size as the block can be decoded, or a sub-residual block corresponding to a portion of the block can be decoded. Information for this operation can be signaled for the block, and this information can be, for example, a flag. When a residual block of the same size as the block is decoded, information about the transform kernel is determined by decoding information contained in the bitstream. When a sub-residual block corresponding to a portion of the block is decoded, a transform kernel for the sub-residual block is determined based on information specifying the type of the sub-residual block and / or its position within the block. For example, information about the type of the sub-residual block and / or its position within the block can be included in the bitstream for signaling. Here, when the block is larger than 32×32, the determination of the transform kernel based on the type of the sub-residual block and / or its position within the block is not performed. For example, for blocks larger than 32×32, a predetermined transform kernel (e.g., DCT-2) can be applied, or information about the transform kernel can be explicitly signaled. Optionally, when the width or height of the block is greater than 32, the determination of the transform kernel based on the type of the sub-residual block and / or its position within the block is not performed. For example, for a 64×8 block, a predetermined transform kernel (e.g., DCT-2) may be applied, or information about the transform kernel may be explicitly sent by signaling.
[0309] Information about the type of a sub-residual block can be the block's partitioning information. This partitioning information can be, for example, partitioning direction information indicating one of horizontal or vertical partitioning. Optionally, the block's partitioning information can include partitioning ratio information. For example, partitioning ratios can include 1:1, 1:3, and / or 3:1. Partitioning direction information and partitioning ratio information can be sent as separate syntax elements or as a single syntax element using signals.
[0310] Information about the location of a sub-residual block can indicate its position within that block. For example, when the block is partitioned vertically, the information about its position indicates either the left or right side. Furthermore, when the block is partitioned horizontally, the information about its position indicates either the top or the bottom.
[0311] The transform kernel of a sub-residual block can be determined based on type information and / or location information. The transform kernel can be determined independently for horizontal and vertical transforms. For example, the transform kernel can be determined based on the partitioning direction. For example, in the case of vertical partitioning, a first transform kernel can be applied to the vertical transform. In the case of horizontal partitioning, a first transform kernel can be applied to the horizontal transform. For example, a first or second transform kernel can be applied to both the horizontal and vertical transforms in both cases. For example, in the case of vertical partitioning, a second transform kernel can be applied to the horizontal transform at the left position, and a first transform kernel can be applied to the horizontal transform at the right position. Furthermore, in the case of horizontal partitioning, a second transform kernel can be applied to the vertical transform at the top position, and a first transform kernel can be applied to the vertical transform at the bottom position. For example, the first and second transform kernels can be DST-7 and DCT-8, respectively. For example, the first and second transform kernels can be DST-7 and DCT-2, respectively. However, this is not a limitation; any two different transform kernels can be used as the first and second transform kernels in the various transform kernels described in this specification. Here, the block can represent CU or TU. Furthermore, sub-residual blocks can represent sub-TUs.
[0312] Only in the case of the TU within the PU predicted inter-frame can the transform mode information be entropy encoded / decoded in bypass mode. Furthermore, in the case of at least one of transform skip mode, residual differential PCM (RDPCM) mode, and lossless mode, the entropy encoding / decoding of the transform mode information is omitted, and the transform mode information is not transmitted via signal.
[0313] Furthermore, when the coded block flag is zero, entropy encoding / decoding of the transform mode information is omitted, and the transform mode information is not transmitted via signal. When the coded block flag is zero, the inverse transform processing is omitted in the decoder. Therefore, block reconstruction can be performed even when transform mode information is absent in the decoder.
[0314] However, transformation pattern information is not limited to representing transformation patterns through flags, and can be implemented in the form of predefined tables and indexes. Here, the predefined table can be a table that defines the available transformation patterns for each index.
[0315] Furthermore, DCT-2 or SDST transformations can be performed in the horizontal and vertical directions respectively. The same transformation mode can be used in both the horizontal and vertical directions, or different transformation modes can be used.
[0316] Furthermore, transform mode information related to whether DCT-2 is used in the horizontal and vertical directions, whether SDST is used, and whether DST-7 is used can be entropy encoded / entropy decoded, respectively. Transform mode information can be transmitted as an index, for example, using a signal. The transform kernel indicated by the same index can be the same for blocks predicted intra-frame and blocks predicted inter-frame.
[0317] Furthermore, the transformation mode information can be entropy encoded / entropy decoded in units of CU, PU, TU, and blocks.
[0318] Furthermore, transformation mode information can be transmitted using signals based on the luminance component or the chrominance component. In other words, transformation mode information can be transmitted using signals based on the Y component, Cb component, or Cr component. For example, when transmitting transformation mode information related to whether DCT-2 or SDST is performed for the Y component, the transformation mode information transmitted using signals for the Y component can be used as the transformation mode of the block without transmitting any transformation mode information for at least one of the Cb and Cr components.
[0319] Here, the arithmetic coding method using a context model can be used to entropy encode / decode the transform pattern information. When the transform pattern information is implemented in the form of a predefined table and index, the arithmetic coding method using a context model is used to entropy encode / decode all or part of the binary bits in multiple binary bits.
[0320] Furthermore, entropy encoding / decoding of the transformation mode information can be selectively performed based on the block size. For example, when the size of the current block is equal to or greater than 64×64, entropy encoding / decoding of the transformation mode information is not performed. When the size is equal to or less than 32×32, entropy encoding / decoding of the transformation mode information is performed.
[0321] Furthermore, when there are non-zero transform coefficients or L quantization levels within the current block, entropy encoding / decoding of the transform mode information is not performed, and one of the DCT-2, DST-7, and SDST methods is executed. Here, entropy encoding / decoding of the transform mode information is not performed regardless of the position of the non-zero transform coefficients or quantization levels within the block. Additionally, entropy encoding / decoding of the transform mode information is not performed only when the non-zero transform coefficients or quantization levels are located in the upper left position within the block. Here, L can be a positive integer including zero, and can be, for example, 1.
[0322] Furthermore, when there are non-zero transform coefficients or J or more quantization levels within the current block, entropy encoding / decoding is performed on the transform mode information. Here, J is a positive integer.
[0323] Furthermore, the transform mode information is based on the transform mode of the same block, which restricts the use of certain transform modes or the transform mode of the same block is represented by several bits. The binarization method of the transform method can vary.
[0324] The above-mentioned SDST can be used in a limited manner based on at least one of the prediction mode, intra-frame prediction mode, inter-frame prediction mode, TU depth, size and shape of the current block.
[0325] For example, SDST is used when the current block is encoded in inter-frame mode.
[0326] A minimum / maximum depth for SDST can be defined. In this case, SDST is used when the depth of the current block is equal to or greater than the minimum depth. Alternatively, SDST is used when the depth of the current block is equal to or less than the maximum depth. Here, the minimum / maximum depth can be a fixed value, or it can be variably determined based on information indicating the minimum / maximum depth. The information indicating the minimum / maximum depth can be sent by a signal from the encoder and can be derived from the decoder based on the attributes of the current block / neighboring blocks (e.g., size, depth, and / or shape).
[0327] The minimum / maximum size allowed for SDST can be defined. Similarly, SDST is used when the size of the current block is equal to or greater than the minimum size. Optionally, SDST is used when the size of the current block is equal to or less than the maximum size. Here, the minimum / maximum size can be a fixed value, or it can be variably determined based on information indicating the minimum / maximum size. The information indicating the minimum / maximum size can be sent by a signal from the encoder and can be derived from the decoder based on the attributes of the current block / neighboring blocks (e.g., size, depth, and / or shape). For example, when the current block is 4×4, DCT-2 is used as the transform method, and no entropy encoding / decoding is performed on transform mode information related to whether DCT-2 or SDST is used.
[0328] The shape of a block that allows SDST can be defined. In this case, SDST is used when the current block's shape is the defined block shape. Conversely, the shape of a block that does not allow SDST can be defined. In this case, SDST is not used when the current block's shape is the defined block shape. The shapes of blocks that allow or disallow SDST can be fixed, and information about this can be transmitted from the encoder via signals. Alternatively, this information can be derived from the decoder based on attributes of the current block / neighboring blocks (e.g., size, depth, and / or shape). The shape of a block that allows or disallows SDST can represent, for example, M, N, and / or the ratio of M to N in an M×N block.
[0329] Furthermore, when the TU depth is zero, DCT-2 or DST-7 is used as the transform method, and information about the transform mode of which transform method was used is entropy encoded / entropy decoded. When DST-7 is used as the transform method, rearrangement processing of the residual signal is performed. Additionally, when the TU depth is 1 or greater, DCT-2 or SDST is used as the transform method, and information about the transform mode of which transform method was used is entropy encoded / entropy decoded.
[0330] Furthermore, transformation methods can be selectively applied based on the partition shapes of CU and PU or the shape of the current block.
[0331] According to the embodiment, DCT-2 is used when the partition shape of CU and PU or the shape of the current block is 2N×2N. For the remaining partition shapes and block shapes, DCT-2 or SDST can be used selectively.
[0332] In addition, DCT-2 is used when the partition shape of CU and PU or the shape of the current block is 2N×N or N×2N. For the remaining partition shapes and block shapes, DCT-2 or SDST can be used selectively.
[0333] In addition, DCT-2 is used when the partition shape of CU and PU or the shape of the current block is nR×2N, nL×2N, 2N×nU, or 2N×nD. For the remaining partition shapes and block shapes, DCT-2 or SDST can be used selectively.
[0334] Additionally, when SDST or DST-7 is executed in blocks derived from the partitions of the current block, scans and inverse scans of the transform coefficients (quantization levels) can be performed on a block-by-block basis. Furthermore, when SDST or DST-7 is executed in blocks derived from the partitions of the current block, scans and inverse scans of the transform coefficients (quantization levels) can be performed on an unpartitioned block of the current block.
[0335] In addition, a transform / inverse transform using SDST or DST-7 can be performed based on at least one of the intra-prediction mode (direction) of the current block, the size of the current block, and the components of the current block (luminance component or chrominance component).
[0336] Furthermore, in transform / inverse transforms using SDST or DST-7, DST-1 can be used instead of DST-7. Additionally, in transform / inverse transforms using SDST or DST-7, DCT-4 can be used instead of DST-7.
[0337] Furthermore, in the transform / inverse transform using DCT-2, the rearrangement method used to rearrange the residual signals of SDST or DST-7 can be applied. That is, even when using DCT-2, the rearrangement of the residual signals or the rotation of the residual signals using a predetermined angle is performed.
[0338] The following sections will describe various modifications and embodiments of the mixed-signaling method and signaling method.
[0339] The SDST of this invention aims to enhance image compression efficiency by changing the transformation, shuffling, rearrangement, and / or flipping methods. Performing DST-7 by shuffling the residual signal effectively reflects the distribution characteristics of the residual signal within the PU, thus achieving high compression efficiency.
[0340] The residual signal rearrangement method has been described above in relation to the commingling step. In the following sections, in addition to the commingling method for rearranging residual signals, other implementation methods will be described.
[0341] The rearrangement method described below can be applied to at least one embodiment of the embodiments related to the SDST method described above.
[0342] To minimize the hardware complexity of rearranging residual signals, the rearrangement process can be implemented using horizontal and vertical flipping methods. The residual signal rearrangement method can be implemented using flipping as shown in (1) to (4) below. The rearrangement described below can represent flipping.
[0343] (1) r'(x,y) = r(x,y); No flipping
[0344] (2) r'(x,y)=r(w-1-x,y); horizontal flip
[0345] (3) r'(x,y)=r(x,h-1-y); vertical flip
[0346] (4) r'(x,y)=r(w-1-x,h-1-y); Horizontal and vertical flipping
[0347] The expression r'(x,y) represents the residual signal after rearrangement, and the expression r(x,y) represents the residual signal before rearrangement. The width and height of the block are represented by w and h, respectively. The position of the residual signal within the block is represented by x and y. The inverse rearrangement method using the flipped rearrangement method can be executed with the same processing as the rearrangement method. That is, the residual signal rearranged using horizontal flipping can be reconstructed into the original residual signal arrangement by performing horizontal flipping again. The rearrangement method executed by the encoder and the inverse rearrangement method executed by the decoder can be the same flipping method.
[0348] For example, when performing a horizontal flip on a residual block that has already been horizontally flipped, the residual block before the flip is performed is obtained, as shown below.
[0349] r'(w-1-x,y)=r(w-1-(w-1-x),y)=r(x,y).
[0350] For example, when performing a vertical flip on a residual block that has already been vertically flipped, the residual block before the flip is performed is obtained, as shown below.
[0351] r'(x,h-1-y)=r(x,h-1-(h-1-y))=r(x,y).
[0352] For example, when performing a horizontal and vertical flip on a residual block that has already been horizontally and vertically flipped, the residual block before the flip is performed is obtained, as shown below.
[0353] r'(w-1-x,h-1-y)=r(w-1-(w-1-x),h-1-(h-1-y))=r(x,y).
[0354] A residual signal mixing / rearrangement method based on flipping can be used without partitioning the current block. That is, the SDST method describes partitioning the current block (TU, etc.) into sub-blocks and applying DST-7 to each sub-block. However, when using the residual signal mixing / rearrangement method based on flipping, the current block is not partitioned into sub-blocks, and a flip is performed on all or part of the current block before performing the DST-7 transform. Furthermore, when using the residual signal mixing / rearrangement method based on flipping, the current block is not partitioned into sub-blocks, and a flip is performed on all or part of the current block after performing the inverse DST-7 transform.
[0355] The maximum size (M×N) and / or minimum size (O×P) of a block capable of performing flip-based residual signal mixing / rearrangement can be defined. Here, the size may include at least one of a width as a horizontal dimension (M or O) and a height as a vertical dimension (N or P). M, N, O, and P can be positive integers. The maximum size and / or minimum size of the block can be predefined values in the encoder / decoder, or they can be information sent from the encoder to the decoder via signals.
[0356] For example, when the size of the current block is smaller than the minimum size that allows the flip method to be performed, the flip and DST-7 transformation are not performed, and only the DCT-2 transformation is performed. Here, the SDST flag, which serves as a transformation mode information indicating whether the flip and DST-7 are used as transformation modes, is not required to be transmitted.
[0357] For example, when the block width is less than the minimum width required to perform the flip method and the block height is greater than the minimum height required to perform the flip method, only DCT-2 is used to perform the one-dimensional transformation in the horizontal direction. Regarding the one-dimensional transformation in the vertical direction, DST-7 is used to perform the one-dimensional vertical transformation after a vertical flip, or DST-7 is used to perform the one-dimensional vertical transformation without a flip. Here, only for the one-dimensional transformation in the vertical direction, the SDST flag can be sent as transformation mode information indicating whether a flip is used as the transformation mode.
[0358] For example, when the block height is less than the minimum height for performing the flip method and the block width is greater than the minimum width for performing the flip method, for one-dimensional transformations in the horizontal direction, DST-7 is used to perform the one-dimensional horizontal transformation after a horizontal flip, or DST-7 is used to perform the one-dimensional horizontal transformation without a flip. Only DCT-2 is used to perform one-dimensional transformations in the vertical direction. Here, an SDST flag indicating whether a flip is used as the transformation mode can be sent only for one-dimensional transformations in the horizontal direction.
[0359] For example, when the size of the current block is larger than the maximum size that can be used to perform the flip method, the flip and DST-7 transform are not used, and only the DCT-2 transform is used. Here, the SDST flag, which serves as transform mode information indicating whether the flip and DST-7 transform are used as the transform mode, is not required to be transmitted.
[0360] For example, when the size of the current block is larger than the maximum size that can be used to perform the flipping method, only the DCT-2 transformation or the DST-7 transformation is used.
[0361] For example, when the maximum size for which the flip method can be performed is 32×32 and the minimum size is 4×4, flipping and DST-7 transformations are not used for blocks of size 64×64, and only DCT-2 transformations are used. Here, for blocks of size 64×64, the SDST flag, which indicates whether flipping and DST-7 are used as transformation modes, does not need to be signaled. Furthermore, for blocks of sizes 4×4 to 32×32, the SDST flag, which indicates whether flipping and DST-7 are used as transformation modes, can be signaled. In this case, DST-7 transformations are not used for blocks of size 64×64, thus saving memory space used to store DST-7 transformations for blocks of size 64×64.
[0362] For example, when the maximum size that can be used to perform the flip method is 32×32 and the minimum size is 4×4, not only is the flip method used for blocks of size 64×64, but DCT-2 or DST-7 transformations are also used.
[0363] For example, a square block of size M×N can be partitioned into four sub-blocks according to a quadtree, and a shuffling / rearranging method can be performed on each sub-block using a flipping method, followed by a DST-7 transformation. Here, the flipping method can be explicitly signaled for each sub-block. The flipping method can be signaled as a fixed-length code of two bits, or as a truncated unary code. Furthermore, a binarization method based on the probability of occurrence of the flipping method according to each block obtained from the partitioning can be used. Here, M and N can be positive integers, for example, 64×64.
[0364] The transform mode information (sdst_flag or sdst flag) can be used to entropy encode / decode information regarding the use of a flip-based residual signal mixing / rearranging method. That is, by signaling the transform mode information, the same method executed in the encoder can be performed in the decoder. For example, when the flag indicating the transform mode information has a first value, a flip-based and DST-7-based residual signal mixing / rearranging method is used as the transform / inverse transform method. When the flag has a second value, another transform / inverse transform method is used. Here, the transform mode information can be entropy encoded / decoded for each block. Here, the other transform / inverse transform method can be the DCT-2 transform / inverse transform method. Furthermore, in the case of transform skip mode, residual differential PCM (RDPCM) mode, and lossless mode, the entropy encoding / decoding of the transform mode information is omitted, and the transform mode information is not signaled.
[0365] The transform mode information can be entropy encoded / decoded using at least one of the following: the depth of the current block, the size of the current block, the shape of the current block, transform mode information of neighboring blocks, the coded block flag of the current block, and information about whether a transform skip mode is used for the current block. For example, when the coded block flag of the current block is zero, entropy encoding / decoding of the transform mode information is omitted, and the transform mode information is not transmitted by signal. Furthermore, the transform mode information can be predictively encoded / decoded based on the transform mode information of the reconstructed blocks adjacent to the current block during entropy encoding / decoding. Additionally, the transform mode information can be transmitted by signal based on at least one of the encoding parameters of the current block and neighboring blocks.
[0366] Furthermore, by using flipping method information, at least one of the four flipping methods (no flipping, horizontal flipping, vertical flipping, and horizontal and vertical flipping) can be entropy encoded / decoded in the form of a flag or index (flipping_idx). In other words, by sending the flipping method information as a signal, the same flipping method executed in the encoder can be performed in the decoder. Transformation mode information may include the flipping method information.
[0367] Furthermore, in the case of transform skip mode, residual differential PCM (RDPCM) mode, and lossless mode, entropy encoding / decoding of the flip method information is omitted, and the flip method information is not transmitted by signal. Entropy encoding / decoding of the flip method information can be performed using at least one of the following: the depth of the current block, the size of the current block, the shape of the current block, the flip method information of neighboring blocks, the coded block flag of the current block, and information regarding whether the transform skip mode of the current block is used. For example, when the coded block flag of the current block is zero, entropy encoding / decoding of the flip method information is omitted, and transform mode information is not transmitted by signal. Furthermore, the flip method information can be predictively encoded / decoded based on the flip method information of the reconstructed blocks adjacent to the current block during entropy encoding / decoding. Additionally, the flip method information can be transmitted by signal based on at least one of the encoding parameters of the current block and neighboring blocks.
[0368] Furthermore, the encoder can determine the optimal rearrangement method for a portion of the rearrangement step in the aforementioned residual signal rearrangement method, and can send information about the determined rearrangement method (reversal method information) to the decoder via a signal. For example, when four rearrangement methods are used, the encoder can send up to two bits of information about the residual signal rearrangement method to the decoder via a signal.
[0369] Furthermore, when the rearrangement methods used have different occurrence probabilities, fewer bits are used to encode the rearrangement methods with high occurrence probabilities, and relatively more bits are used to encode the rearrangement methods with low occurrence probabilities. For example, the four rearrangement methods are arranged in descending order of occurrence probability and can be transmitted as truncated unary codes (e.g., (0, 10, 110, 111) or (1, 01, 001, 000)).
[0370] Furthermore, the occurrence probability of a rearrangement method can vary depending on coding parameters (such as the prediction mode of the current CU, the intra-frame prediction mode (direction) of the PU, the motion vectors of neighboring blocks, etc.). Therefore, coding methods that utilize information about the rearrangement method (flipping method information) can be used differently depending on the coding parameters. For example, the occurrence probability of a rearrangement method can vary depending on the intra-frame prediction mode. Therefore, for each intra-frame mode, fewer bits can be allocated to rearrangement methods with high occurrence probabilities, and many bits can be allocated to rearrangement methods with low occurrence probabilities. Optionally, depending on the situation, rearrangement methods with extremely low occurrence probabilities may not be used and may not be allocated any bits.
[0371] A rearrangement set, including at least one residual signal rearrangement method, can be constructed based on at least one of the following: the prediction mode (inter-frame mode or intra-frame mode), the intra-frame prediction mode (including directional and non-directional modes), the inter-frame prediction mode, the block size, the block shape (square or non-square), the luma / chroma signal, and transform mode information. Rearrangement can represent flipping. Furthermore, a rearrangement set, including at least one residual signal rearrangement method, can be constructed based on at least one coding parameter from the coding parameters of the current block and neighboring blocks.
[0372] Furthermore, based on at least one of the prediction mode, intra-frame prediction mode, inter-frame prediction mode, block size, block shape, luma / chroma signal, transform mode information, etc., at least one of the following rearrangement sets can be selected. Additionally, based on at least one of the coding parameters of the current block and neighboring blocks, at least one of the rearrangement sets can be selected.
[0373] Rearranged sets can include at least one of "no flip", "horizontal flip", "vertical flip", and "horizontal and vertical flip". An example of a rearranged set is shown below.
[0374] 1. Do not flip
[0375] 2. Horizontal flip
[0376] 3. Vertical Flip
[0377] 4. Horizontal and vertical flipping
[0378] 5. No flipping, and horizontal flipping.
[0379] 6. No flipping, and vertical flipping
[0380] 7. No flipping, and horizontal and vertical flipping.
[0381] 8. Horizontal flip and vertical flip
[0382] 9. Horizontal flip, and both horizontal and vertical flip.
[0383] 10. Vertical flip, and horizontal and vertical flip.
[0384] 11. No flip, horizontal flip, and vertical flip
[0385] 12. No flip, horizontal flip, and both horizontal and vertical flip.
[0386] 13. No flip, vertical flip, and horizontal and vertical flip.
[0387] 14. Horizontal flip, vertical flip, and both horizontal and vertical flip
[0388] 15. No flip, horizontal flip, vertical flip, and both horizontal and vertical flips
[0389] Based on the rearrangement set, at least one residual signal rearrangement method can be used for the rearrangement of the current block.
[0390] Furthermore, based on at least one of the prediction mode of the current block, intra-frame prediction mode, inter-frame prediction mode, block size, block shape, luma / chroma signal, transform mode information, and flip method information, at least one residual signal rearrangement method can be selected from the rearrangement set. Additionally, based on at least one coding parameter of the coding parameters of the current block and neighboring blocks, at least one residual signal rearrangement method can be selected from the rearrangement set.
[0391] At least one rearrangement set can be constructed based on the prediction mode of the current block. For example, when the prediction mode of the current block is intra-frame prediction, multiple rearrangement sets are constructed. When the prediction mode of the current block is inter-frame prediction, one rearrangement set is constructed.
[0392] Based on the luma / chroma signal of the current block, at least one rearrangement set can be constructed. For example, when the current block is a chroma signal, one rearrangement set is constructed. When the current block is a luma signal, multiple rearrangement sets are constructed.
[0393] Furthermore, based on the rearranged set, the index for the residual signal rearrangement method can be entropy encoded / decoded. Here, the index can be entropy encoded / decoded into variable-length code or fixed-length code.
[0394] Furthermore, based on the rearranged set, binarization and debinarization of the index for the residual signal rearrangement method can be performed. Here, the index can be binarized and debinarized into variable-length code or fixed-length code.
[0395] Furthermore, the rearranged set can be in the form of a table in the encoder and decoder, and can be computed through equations.
[0396] Furthermore, the rearrangement set can be constructed in a symmetric manner. For example, a table for the rearrangement set can be constructed in a symmetric manner. Here, the table can be constructed in a symmetric manner for intra-frame prediction modes.
[0397] In addition, the rearrangement set can be constructed based on at least one of the following: whether the intra-prediction mode is in a specific range, and whether the intra-prediction mode is even or odd.
[0398] The following is an example of a method for encoding / decoding residual signals based on the prediction mode of the current block and the intra-frame prediction mode (direction).
[0399] In addition, the following table may use the flipping method information to indicate the use of at least one residual signal rearrangement method among the residual signal rearrangement methods.
[0400] [Table 1]
[0401]
[0402] In Table 1, columns (1) to (4) specify the residual signal rearrangement methods, such as the index for the scan / rearrangement order of the above residual signal rearrangement, the index for the predetermined angle value, the index for the predetermined flipping method, etc. In Table 1, the * symbol in the column of residual signal rearrangement methods indicates that the corresponding rearrangement method is used implicitly without signaling, and the - symbol indicates that the corresponding rearrangement method is not used in the corresponding case. Implicit use of the rearrangement method means that the rearrangement method is used by utilizing the transformation mode information (sdst_flag or sdst flag) without entropy encoding / entropy decoding of the index for the residual signal rearrangement method. Columns (1) to (4) of the residual signal rearrangement methods can refer to (1) no flipping, (2) horizontal flipping, (3) vertical flipping, and (4) horizontal and vertical flipping, respectively. Furthermore, the numbers 0, 1, 10, 11, 110, 111, etc., can be the result of binarization / debinarization for entropy coding / debinarization of the residual signal rearrangement method. Fixed-length codes, truncated unary codes, unary codes, etc., can be used as binarization / debinarization methods. As shown in Table 1, at least one rearrangement method is used in the encoder and decoder when the current block corresponds to at least one of the prediction mode and the intra-frame prediction mode (direction). Here, the diagonal direction at a 45-degree angle can represent a direction towards the upper left position in the current block or a direction from the upper left position in the current block towards the current block.
[0403] [Table 2]
[0404]
[0405] As another example, as shown in Table 2, when the current block corresponds to at least one prediction mode in the prediction mode and at least one intra prediction mode (direction) in the intra prediction mode (direction), at least one rearrangement method is used in the encoder and decoder.
[0406] [Table 3]
[0407]
[0408] As another example, as shown in Table 3, when the current block corresponds to at least one prediction mode in the prediction modes and at least one intra-prediction mode (direction) in the intra-prediction modes (directions), at least one reordering method is used in the encoder and decoder. Here, the residual signal reordering method can represent a type of transform. For example, when the residual signal reordering method is (1), both the horizontal transform and the vertical transform represent the first transform kernel. As another example, when the residual signal reordering method is (2), the horizontal transform and the vertical transform represent the second transform kernel and the first transform kernel, respectively. As another example, when the residual signal reordering method is (3), the horizontal transform and the vertical transform represent the first transform kernel and the second transform kernel, respectively. As another example, when the residual signal reordering method is (4), the horizontal transform and the vertical transform represent the second transform kernel and the second transform kernel, respectively. For example, the first transform kernel can be DST-7, and the second transform kernel can be DCT-8. When the intra-prediction mode is a planar mode or a DC mode, entropy coding / entropy decoding of information about the four reordering methods (flipping method information) is performed using a truncated unary code based on the occurrence frequency. In the case of inter-frame prediction, the probability of occurrence of rearrangement methods (1) to (4) can be considered equal, and information about the rearrangement method can be entropy encoded / entropy decoded into a two-bit fixed-length code.
[0409] Arithmetic encoding / decoding can be used for the code. Furthermore, arithmetic encoding using a context model specific to the code may not be used, and entropy encoding / decoding can be performed in a bypass mode.
[0410] A DST-7 transform / inverse transform can be performed on a region or CTU within a frame, the entire frame, or the current block within a frame group without flipping. Alternatively, a DCT-2 transform / inverse transform can be performed by selecting one of the two methods. In this case, a 1-bit flag indicating whether DST-7 or DCT-2 is used (transform mode information) can be entropy-encoded / entropy-decoded on a current block basis. This method can be used when the energy of the residual signal increases with the distance from the reference sample, or to reduce the computational complexity in encoding and decoding. Information about the region using this method can be signaled on a CTU, strip, PPS, SPS, or other specific region basis, and the 1-bit flag can be signaled on / off.
[0411] For a region or CTU within a frame, the entire frame, or the current block within a frame group, a transform / inverse transform can be performed by selecting one of five methods: DCT-2 transform / inverse transform, DST-7 transform / inverse transform without flipping, DST-7 transform / inverse transform after horizontal flipping, DST-7 transform / inverse transform after vertical flipping, and DST-7 transform / inverse transform after both horizontal and vertical flipping. Information about which of the five transform methods will be selected can be implicitly selected using proximity information of the current block, and can be explicitly selected by sending an index (transformation mode information or flipping method information) using signals. The index can be represented by a truncated unary code as follows: 0 for DCT-2, 10 for DST-7 without flipping, 110 for DST-7 after horizontal flipping, 1110 for DST-7 after vertical flipping, and 1111 for DST-7 after both horizontal and vertical flipping. Furthermore, the binarization of DCT-2 and DST-7 can be swapped for signal transmission based on the current block size and proximity information. Additionally, the first binary digit in the binary number can be signaled in units of CU, and the remaining binary digits can be signaled in units of TU or PU. Furthermore, information can be represented as a fixed-length code by distinguishing the first, second, and third binary digits within the binary number. For example, transformation mode information or flipping method information can be signaled as follows: DCT-2 is 0, DST-7 without flipping is 000, DST-7 after horizontal flipping is 001, DST-7 after vertical flipping is 010, and DST-7 after both horizontal and vertical flipping is 011. Furthermore, depending on the intra-frame prediction mode, only a portion of the five methods can be used. For example, when the intra-frame prediction mode is close to the horizontal prediction mode, only three transformation methods are used: DCT-2, DST-7 without flipping, and DST-7 after vertical flipping. In this case, the transformation mode information or flipping method information can be sent by signal as follows: DCT-2 is 0, DST-7 without flipping is 10, and DST-7 after vertical flipping is 11.
[0412] FIG. 8 This is a diagram illustrating an embodiment of a decoding method using the SDST method according to the present invention.
[0413] Reference FIG. 8 First, in step 801, the transformation mode of the current block can be determined, and in step 802, the residual data of the current block can be inversely transformed according to the transformation mode of the current block.
[0414] Next, in step 803, the residual data of the current block, which has already undergone an inverse transformation according to the transformation mode of the current block, can be rearranged.
[0415] Here, the transformation mode may include at least one of the following: mixed discrete sine transform (SDST), mixed discrete cosine transform (SDCT), discrete sine transform (DST), and discrete cosine transform (DCT).
[0416] The SDST mode indicates a mode in which the inverse transformation is performed in the DST-7 transformation mode and the residual data that has undergone the inverse transformation is rearranged.
[0417] The SDCT mode indicates the mode in which the inverse transform is performed in the DCT-2 transform mode and the residual data that has already undergone the inverse transform is rearranged.
[0418] The DST mode indicates a mode in which the inverse transformation is performed in the DST-7 transformation mode and the residual data that has already undergone the inverse transformation is not rearranged.
[0419] The DCT mode indicates a mode in which the inverse transformation is performed under the DCT-2 transform mode, and the residual data that has already undergone the inverse transformation is not rearranged.
[0420] Therefore, the rearrangement of the residual data is performed only if the transformation mode of the current block is one of SDST and SDCT.
[0421] Although the inverse transformation is described for the SDST and DST modes as described above in the DST-7 transformation mode, transformation modes based on other DSTs (such as DST-1, DST-2, etc.) can be used.
[0422] Additionally, the step of determining the transformation mode of the current block in step 801 may include: obtaining transformation mode information of the current block from the bit stream; and determining the transformation mode of the current block based on the transformation mode information.
[0423] Furthermore, when determining the transformation mode of the current block in step 801, the transformation mode of the current block can be determined based on at least one of the prediction mode of the current block, the depth information of the current block, the size of the current block, and the shape of the current block.
[0424] Specifically, when the prediction mode of the current block is the inter-frame prediction mode, one of SDST and SDCT is determined as the transformation mode of the current block.
[0425] Additionally, the step of rearranging the residual data of the current block that has undergone inverse transformation in step 803 may include: scanning the residual data arranged in the current block that has undergone inverse transformation in a first direction; and rearranging the residual data scanned in the first direction in the current block that has undergone inverse transformation in a second direction. Here, the first direction order may be one of the raster scan order, the upper right diagonal scan order, the horizontal scan order, and the vertical scan order. Furthermore, the first direction order may be defined as follows.
[0426] (1) Scan from top row to bottom row, and scan from left to right within a row.
[0427] (2) Scan from top row to bottom row, and scan from right to left within a row.
[0428] (3) Scan from bottom row to top row, and scan from left to right within a row.
[0429] (4) Scan from bottom row to top row, and scan from right to left within a row.
[0430] (5) Scan from the left column to the right column, and scan from top to bottom in a column.
[0431] (6) Scan from the left column to the right column, and scan from bottom to top in a column.
[0432] (7) Scan from the right column to the left column, and scan from top to bottom within a column.
[0433] (8) Scan from the right column to the left column, and scan from bottom to top within a column.
[0434] (9) Spiral scan: Scan from the inside (or outside) of the block to the outside (or inside) of the block, and scan in a clockwise / counterclockwise direction.
[0435] Furthermore, for the order of the second direction, one of the directions mentioned above may be used selectively. The first and second directions may be the same or different from each other.
[0436] Furthermore, when rearranging the residual data of the current block that has undergone the inverse transformation in step 803, the rearrangement is performed on a sub-block basis within the current block. In this case, the residual data can be rearranged based on the position of the sub-blocks within the current block.
[0437] Furthermore, when rearranging the residual data of the current block that has already undergone inverse transformation in step 803, the residual data arranged within the current block that has already undergone inverse transformation can be rotated by a predefined angle for rearrangement.
[0438] Furthermore, when rearranging the residual data of the current block that has undergone inverse transformation in step 803, the residual data arranged within the current block that has undergone inverse transformation can be flipped to rearrange it according to the flipping method. In this case, the step of determining the transformation mode of the current block in step 801 may include: obtaining flipping method information from the bit stream; and determining the flipping method for the current block based on the flipping method information.
[0439] FIG. 9 This is a diagram illustrating an embodiment of an encoding method using the SDST method according to the present invention.
[0440] Reference FIG. 9 In step 901, the transformation mode of the current block can be determined.
[0441] Next, in step 902, the residual data of the current block can be rearranged according to the transformation mode of the current block.
[0442] Next, in step 903, the residual data of the current block, which has been rearranged according to the transformation mode of the current block, can be transformed.
[0443] Here, the transform mode may include at least one of the following: Mixed Discrete Sine Transform (SDST), Mixed Discrete Cosine Transform (SDCT), Discrete Sine Transform (DST), and Discrete Cosine Transform (DCT). Since already referenced... FIG. 8 The SDST, SDCT, DST, and DCT modes are described, so repeated descriptions will be omitted.
[0444] Additionally, the residual data is rearranged only if the current block's transformation mode is either SDST or SDCT.
[0445] Furthermore, when determining the transformation mode of the current block in step 901, the transformation mode of the current block can be determined based on at least one of the prediction mode of the current block, the depth information of the current block, the size of the current block, and the shape of the current block.
[0446] Here, when the prediction mode of the current block is the inter-frame prediction mode, one of SDST and SDCT is determined as the transformation mode of the current block.
[0447] Additionally, the step of rearranging the residual data of the current block in step 902 may include: scanning the residual data arranged in the current block in a first direction; and rearranging the residual data scanned in the first direction in the current block in a second direction.
[0448] Furthermore, when rearranging the residual data of the current block in step 902, the rearrangement is performed on a per-sub-block basis within the current block.
[0449] In this case, when rearranging the residual data of the current block in step 902, the residual data can be rearranged based on the position of the sub-blocks within the current block.
[0450] Additionally, when rearranging the residual data of the current block in step 902, the residual data arranged within the current block can be rotated at a predefined angle for rearrangement.
[0451] In addition, when rearranging the residual data of the current block in step 902, the residual data arranged in the current block can be flipped and rearranged according to the flipping method.
[0452] The image decoder using the SDST method according to the present invention may include an inverse transform module, wherein the inverse transform module determines the transform mode of the current block, performs an inverse transform on the residual data of the current block according to the transform mode of the current block, and rearranges the residual data of the current block that has already undergone the inverse transform according to the transform mode of the current block. Here, the transform mode may include at least one of mixed discrete sine transform (SDST), mixed discrete cosine transform (SDCT), discrete sine transform (DST), and discrete cosine transform (DCT).
[0453] The image decoder using the SDST method according to the present invention may include an inverse transform module, wherein the inverse transform module determines a transform mode for the current block, rearranges the residual data of the current block according to the transform mode, and performs an inverse transform on the rearranged residual data of the current block according to the transform mode. Here, the transform mode may include at least one of Mixed Discrete Sine Transform (SDST), Mixed Discrete Cosine Transform (SDCT), Discrete Sine Transform (DST), and Discrete Cosine Transform (DCT).
[0454] The image encoder using the SDST method according to the present invention may include a transform module, wherein the transform module determines a transform mode for a current block, rearranges the residual data of the current block according to the transform mode, and transforms the rearranged residual data of the current block according to the transform mode. Here, the transform mode may include at least one of mixed discrete sine transform (SDST), mixed discrete cosine transform (SDCT), discrete sine transform (DST), and discrete cosine transform (DCT).
[0455] The image encoder using the SDST method according to the present invention may include a transform module, wherein the transform module determines a transform mode for a current block, transforms the residual data of the current block according to the transform mode, and rearranges the residual data of the current block transformed according to the transform mode. Here, the transform mode may include at least one of mixed discrete sine transform (SDST), mixed discrete cosine transform (SDCT), discrete sine transform (DST), and discrete cosine transform (DCT).
[0456] A bitstream generated by an encoding method using the SDST method according to the present invention can be provided, wherein the encoding method includes: determining a transform mode of the current block; rearranging the residual data of the current block according to the transform mode of the current block; and transforming the residual data of the current block rearranged according to the transform mode of the current block, wherein the transform mode may include at least one of mixed discrete sine transform (SDST), mixed discrete cosine transform (SDCT), discrete sine transform (DST), and discrete cosine transform (DCT).
[0457] Furthermore, the transforms used in this specification can be selected from a predefined set of N transform candidates for each block. Here, N can be a positive integer. Each transform candidate in the transform candidates can specify a primary horizontal transform, a primary vertical transform, and a secondary transform (which can be the same as the identity transform). The list of transform candidates can vary depending on the block size and prediction mode. The selected transform can be signaled as follows: When the coded block flag is 1, a flag indicating whether the first transform in the candidate list is used is sent. When the flag indicating whether the first transform in the candidate list is used is zero, the following applies: if the number of non-zero transform coefficient levels is greater than a threshold, the transform index indicating the transform candidate used is sent; otherwise, the second transform in the list is used.
[0458] Furthermore, NSST is used as a secondary transform only when DCT-2, as the primary transform, is used as the default transform. Additionally, regarding horizontal or vertical transforms, DST-7 is selected without signal transmission when the width or height is equal to or less than 4.
[0459] Regarding residual blocks, DST-7, instead of DCT-2, is used for the one-dimensional horizontal transform when the block width is equal to or less than K. When the block height is equal to or less than L, DST-7, instead of DCT-2, is used for the one-dimensional vertical transform. Furthermore, DCT-2 is used even when the block width or height is equal to or less than K, when the intra-frame prediction mode is a linear model (LM) chroma mode. Here, K and L can be positive integers, for example, 4. Furthermore, K and L can be the same or have different values. Additionally, the residual block can be a block encoded in intra-frame mode. Furthermore, the residual block can be a chroma block.
[0460] As an alternative to performing a flipping method on the residual signal, a transform / inverse transform can be performed using a transform kernel or transform matrix that has already been flipped. Here, the transform / inverse transform kernel or transform / inverse transform matrix that has already been flipped can be a kernel or matrix that has been predefined in the encoder / decoder. In this case, since the transform / inverse transform matrix that has already been flipped is used to perform the transform / inverse transform, the same effect as performing a flipping on the residual signal can be obtained. Here, the flipping can be no flipping, horizontal flipping, vertical flipping, and at least one of horizontal and vertical flipping. In this case, information about whether a flipped transform / inverse transform has been used can be sent by signal. Furthermore, information about whether a flipped transform / inverse transform has been used can be sent by signal for each of the horizontal and vertical transform / inverse transforms.
[0461] Furthermore, as an alternative to performing a flipping method on the residual signal, the transform kernel or transform matrix can be flipped during the encoding / decoding process to perform a transform / inverse transform. In this case, since the transform / inverse transform matrix is flipped to perform the transform / inverse transform, the same effect as performing a flip on the residual signal can be obtained. Here, the flipping can be no flipping, horizontal flipping, vertical flipping, or at least one of horizontal and vertical flipping. In this case, information about whether a flipping is performed on the transform / inverse transform matrix can be sent by a signal. Furthermore, information about whether a flipping is performed on the transform / inverse transform matrix can be sent by a signal for each of the horizontal and vertical transform / inverse transforms.
[0462] When the flip method is determined based on the intra-prediction mode and two or more intra-prediction modes of the current block are used, the flip is performed before / after the transform / inverse transform of the current block as the flip method for non-directional modes.
[0463] Furthermore, when the flipping method is determined based on the intra-prediction mode and two or more intra-prediction modes for the current block are used, the flipping method as the primary direction mode is performed before / after the transform / inverse transform of the current block. Here, the primary direction mode can be at least one of the vertical mode, horizontal mode, and diagonal mode.
[0464] When the size of the transform is equal to or greater than M×N, all transform coefficients existing in the region from M / 2 to M and from N / 2 to N during or after the transform are set to the value 0. Here, M and N can be positive integers, for example, 64×64.
[0465] To reduce memory requirements, the transform coefficients generated after the transformation can be right-shifted by K. Additionally, the temporary transform coefficients generated after the horizontal transformation can be right-shifted by K. Furthermore, the temporary transform coefficients generated after the vertical transformation can be right-shifted by K. Here, K is a positive integer.
[0466] To reduce memory requirements, the reconstructed residual signal generated after the inverse transform can be right-shifted by K. Furthermore, the temporary transform coefficients generated after the horizontal inverse transform can be right-shifted by K. Additionally, the temporary transform coefficients generated after the vertical inverse transform can be right-shifted by K. Here, K is a positive integer.
[0467] At least one of the following flipping methods can be performed on at least one of the signals generated before, after, and during a transformation / inverse transformation in the horizontal direction, and before and after a transformation / inverse transformation in the vertical direction. In this case, information about the flipping method used in the transformation / inverse transformation in the horizontal or vertical direction can be transmitted via a signal.
[0468] Furthermore, DCT-4 can be used instead of DST-7. A 2N-1 size DCT-4 transform / inverse transform matrix is extracted from a 2N-size DCT-2 transform / inverse transform matrix for use; therefore, only the DCT-2 transform / inverse transform matrix, instead of the DCT-4 transform / inverse transform matrix, is stored in the encoder / decoder, thus reducing the encoder / decoder's memory requirements. Moreover, the 2N-1 size DCT-4 transform / inverse transform logic is utilized from the 2N-size DCT-2 transform / inverse transform logic, thus reducing the chip area required to implement the encoder / decoder. Here, the above examples are applied not only to DCT-2 and DCT-4, but also when there are transform matrices or transform logic shared between at least one type of DST transform / inverse transform and at least one type of DCT transform / inverse transform. That is, one transform / inverse transform matrix or logic can be extracted from one transform / inverse transform matrix or logic for use. Furthermore, given a specific transform / inverse transform size, one transform / inverse transform matrix or logic can be extracted from one transform / inverse transform matrix or logic for use. Furthermore, one transformation / inverse transformation matrix can be extracted from one transformation / inverse transformation matrix according to at least one of the matrix units, basis vector units, and matrix coefficient units.
[0469] Furthermore, when the current block is smaller than M×N, another transform / inverse transform is used for the current block as an alternative to the specific transform / inverse transform. Conversely, when the current block is larger than M×N, another transform / inverse transform is used for the current block as an alternative to the specific transform / inverse transform. Here, M and N are positive integers. The specific transform / inverse transform and the other transform / inverse transform can be predefined transforms / inverse transforms in the encoder / decoder.
[0470] Furthermore, at least one of the transformations of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc. used in this specification can be replaced by at least one of the transformations calculated based on the transformations of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc. Here, the calculated transformation can be a transformation calculated by modifying the coefficient values within the transformation matrices of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc. Furthermore, the coefficient values within the transformation matrices of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc., can have integer values. That is, the transformations of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc., can be integer transformations. Furthermore, the coefficient values within the calculated transformation matrix can have integer values. That is, the calculated transformation can be an integer transformation. Furthermore, the calculated transformation can be the result of left-shifting the coefficient values within the transformation matrices of DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc., by N. Here, N can be a positive integer.
[0471] The DCT-Q and DST-W transforms can be represented as including the DCT-Q and DST-W transforms and their inverses. Here, Q and W can be positive integers of 1 or greater, and for example, the numbers 1 to 9 can have the same meaning as the Roman numerals I to IX.
[0472] Furthermore, the transformations DCT-4, DCT-8, DCT-2, DST-4, DST-1, DST-7, etc., used in this specification are not limited to these, and at least one of the DCT-Q transformation and DST-W transformation can be used by substituting the transformations of DCT-4, DCT-8, DCT-2, DST-4, DST-1, and DST-7. Here, Q and W can be positive integers of 1 or greater, and for example, the numbers 1 to 9 can have the same meaning as the Roman numerals I to IX.
[0473] Furthermore, in the case of square blocks, the transformations used in this specification can be performed in a square transformation form. In the case of non-square blocks, the transformations can be performed in a non-square transformation form. In the case of a square-shaped region that includes at least one of square blocks and non-square blocks, the transformation of that region can be performed in a square transformation form. In the case of a non-square-shaped region that includes at least one of square blocks and non-square blocks, the transformation of that region can be performed in a non-square transformation form.
[0474] Furthermore, in this specification, information regarding the rearrangement method may be information regarding the flipping method.
[0475] Furthermore, the transformations used in this specification may represent at least one of the transformation and the inverse transformation.
[0476] The encoder can perform a transform on the residual block to generate transform coefficients, quantize the transform coefficients to generate quantized coefficient levels, and entropy encode the quantized coefficient levels to improve the subjective / objective image quality of the image.
[0477] The decoder can perform entropy decoding on the quantized coefficient levels, dequantize the quantized coefficient levels to generate transform coefficients, and perform inverse transform on the transform coefficients to generate reconstructed residual blocks.
[0478] Transform type information regarding which transform is used as the transform and its inverse can be explicitly entropy-encoded / entropy-decoded. Alternatively, transform type information regarding which transform is used as the transform and its inverse can be implicitly determined based on at least one of the encoding parameters, without entropy encoding / entropy decoding of this information.
[0479] Hereinafter, embodiments of the image encoding / decoding method and apparatus for performing at least one of the transformations or inverse transformations, as well as the recording medium for storing bitstreams, will be described.
[0480] Using at least one of the embodiments described below, a block can be partitioned into N sub-blocks, and at least one of prediction, transform / inverse transform, quantization / dequantization, or entropy coding / entropy decoding can be performed. Such a mode may be referred to as a first sub-block partitioning mode (e.g., ISP mode or intra-frame sub-partitioning mode).
[0481] A block can represent a coded block, a prediction block, or a transform block. For example, a block can be a transform block.
[0482] Furthermore, the partitioned sub-blocks can represent at least one of a coded block, a prediction block, or a transform block. For example, the partitioned sub-blocks can be transform blocks.
[0483] Furthermore, a block or sub-block can be at least one of an intra-block, an inter-block, or a copy of an intra-block. For example, a block and a sub-block can be intra-blocks.
[0484] Furthermore, a block or sub-block can be at least one of an intra-prediction block, an inter-prediction block, or an intra-block copy prediction block. For example, a block and a sub-block can be intra-prediction blocks.
[0485] Furthermore, a block or sub-block can be at least one of a luminance signal block or a chrominance signal block. For example, a block and a sub-block can be luminance signal blocks.
[0486] When a block is partitioned into N sub-blocks, the block before partitioning can be a coding block, and the partitioned sub-blocks can be at least one of prediction blocks or transform blocks. That is, the size of the partitioned sub-blocks can be used to perform prediction of transform coefficients, transform / inverse transform, quantization / dequantization, and entropy coding / entropy decoding.
[0487] Furthermore, when a block is partitioned into N sub-blocks, the block preceding the partition can be at least one of a coding block or a prediction block, and the partitioned sub-blocks can be transform blocks. That is, prediction can be performed using the size of the block preceding the partition, and transform / inverse transform, quantization / dequantization, and entropy coding / entropy decoding of the transform coefficients can be performed using the size of the partitioned sub-blocks.
[0488] Whether a block is divided into multiple sub-blocks can be determined based on at least one of the block's area (the product of its width and height, etc.), size (width, height, or a combination of width and height), and shape / form (rectangle (not square), square, etc.).
[0489] For example, when the current block is a 64×64 block, the current block can be partitioned into multiple sub-blocks.
[0490] As another example, when the current block is a 32×32 block, the current block can be partitioned into multiple sub-blocks.
[0491] As another example, when the current block is a 32×16 block, the current block can be partitioned into multiple sub-blocks.
[0492] As another example, when the current block is a 16×32 block, the current block can be partitioned into multiple sub-blocks.
[0493] As another example, when the current block is a 4×4 block, the current block may not be divided into multiple sub-blocks.
[0494] As another example, when the current block is a 2×4 block, the current block may not be divided into multiple sub-blocks.
[0495] As another example, when the area of the current block is equal to or greater than 32, the current block can be divided into multiple sub-blocks.
[0496] As another example, when the area of the current block is less than 32, the current block may not be divided into multiple sub-blocks.
[0497] As another example, when the area of the current block is 256 and the shape of the current block is rectangular, the current block can be divided into multiple sub-blocks.
[0498] As another example, when the area of the current block is 16 and the shape of the current block is a square, the current block may not be divided into multiple sub-blocks.
[0499] When a block is partitioned, it can be divided into multiple sub-blocks in at least one partitioning direction, either vertical or horizontal.
[0500] For example, the current block can be divided into two sub-blocks in the vertical direction.
[0501] As another example, the current block can be divided into two sub-blocks in the horizontal direction.
[0502] As another example, the current block can be divided into four sub-blocks in the horizontal direction.
[0503] As another example, the current block can be divided into four sub-blocks in the vertical direction.
[0504] When a block is divided into N sub-blocks, N can be a positive integer, and can be, for example, 2 or 4. Furthermore, N can be determined using at least one of the block's area, size, shape, or partitioning orientation.
[0505] For example, when the current block is a 4×8 or 8×4 block, the current block can be divided into two sub-blocks in the horizontal direction or in the vertical direction.
[0506] As another example, when the current block is a 16×8 or 16×16 block, the current block can be divided into four sub-blocks in the vertical direction or in the horizontal direction.
[0507] As another example, when the current block is an 8×32 or 32×32 block, the current block can be divided into four sub-blocks in the horizontal direction or in the vertical direction.
[0508] As another example, when the current block is a 16×4, 32×4, or 64×4 block, the current block can be divided into four sub-blocks in the vertical direction. Furthermore, when the current block is a 16×4, 32×4, or 64×4 block, the current block can be divided into two sub-blocks in the horizontal direction.
[0509] As another example, when the current block is a 4×16, 4×32, or 4×64 block, the current block can be divided into four sub-blocks in the horizontal direction. Furthermore, when the current block is a 4×16, 4×32, or 4×64 block, the current block can be divided into two sub-blocks in the vertical direction.
[0510] As another example, when the current block is a J×4 block, the current block can be divided into two sub-blocks in the horizontal direction. Here, J can be a positive integer.
[0511] As another example, when the current block is a 4×K block, the current block can be divided into two sub-blocks in the vertical direction. Here, K can be a positive integer.
[0512] As another example, when the current block is a J×K (K>4) block, the current block can be divided into four sub-blocks in the horizontal direction. Here, J can be a positive integer.
[0513] As another example, when the current block is a J×K (J>4) block, the current block can be divided into four sub-blocks in the vertical direction. Here, J can be a positive integer.
[0514] As another example, when the current block is a J×K (K>4) block, the current block can be divided into four sub-blocks in the vertical direction. Here, J can be a positive integer.
[0515] As another example, when the area of the current block is 64, the current block can be divided into four sub-blocks in either the horizontal or vertical direction.
[0516] As another example, when the current block is a 16×4 block and the shape of the current block is rectangular, the current block can be divided into four sub-blocks in the vertical direction.
[0517] As another example, when the area of the current block is 1024 and the shape of the current block is a square, the current block can be divided into four sub-blocks in the horizontal or vertical direction.
[0518] In addition, a sub-block may have at least one of a minimum area, minimum width, or minimum height.
[0519] For example, a sub-block can have a minimum area of S. Here, S can be a positive integer, and can be, for example, 16.
[0520] As another example, a sub-block can have a minimum width of J. Here, J can be a positive integer, and can be, for example, 4.
[0521] As another example, a sub-block can have a minimum height of K. Here, K can be a positive integer, and can be, for example, 4.
[0522] Within each partitioned sub-block, a reconstructed block can be generated by adding the residual block (or reconstructed residual block) and the prediction block. Here, at least one reconstructed sample from each reconstructed sub-block can later be used as a reference sample in the intra-frame prediction of the encoded / decoded sub-block.
[0523] The encoding / decoding order of sub-blocks partitioned from a block can be determined based on at least one of the partition directions.
[0524] For example, the encoding / decoding order of sub-blocks divided by a horizontal partition can be determined as a top-to-bottom order.
[0525] As another example, the encoding / decoding order of the sub-blocks divided vertically can be determined as a left-to-right order.
[0526] For sub-blocks partitioned by the partition, the intra-prediction mode can be shared and used.
[0527] At this point, information about the intra-prediction mode for each sub-block can be entropy-encoded / entropy-decoded only once in the block before partitioning.
[0528] For sub-blocks partitioned from the frame, the intra-block copy mode can be shared and used.
[0529] At this point, information about the intra-block copying mode for each sub-block can be entropy-encoded / entropy-decoded only once in the block before partitioning.
[0530] To indicate a sub-block partitioning pattern that divides a block into N sub-blocks and performs at least one of prediction, transform / inverse transform, quantization / dequantization, or entropy coding / entropy decoding on a sub-block basis, at least one of the sub-block partitioning pattern information or partitioning direction information can be entropy coded / entropy decoded.
[0531] Here, sub-block partitioning mode information can be used to indicate the sub-block partitioning mode. When the sub-block partitioning mode indicator is used (second value), the block can be partitioned into sub-blocks, and at least one of prediction, transform / inverse transform, quantization / dequantization, or entropy coding / decoding can be performed. When the sub-block partitioning mode indicator is not used (first value), the block may not be partitioned into sub-blocks, and at least one of prediction, transform / inverse transform, quantization / dequantization, or entropy coding / decoding can be performed. Here, the first value can be 0, and the second value can be 1.
[0532] Furthermore, partitioning direction information can be used to indicate whether the sub-block partitioning mode is vertical or horizontal. When the partitioning direction information has a first value, the block can be partitioned into sub-blocks in the horizontal direction, and the first value can be 0. Furthermore, when the partitioning direction information has a second value, the block can be partitioned into sub-blocks in the vertical direction, and the second value can be 1.
[0533] When the current block does not use the closest reference sample line (first reference sample line) as the reference sample line, entropy encoding / decoding of at least one of the sub-block partitioning mode information or partitioning direction information is not required. In this case, it can be inferred from the sub-block partitioning mode information that the current block has not been partitioned into sub-blocks.
[0534] Here, the fact that the closest reference sample line (the first reference sample line) is not used as the reference sample line for the current block means that the second or larger reference sample line can be used as the reconstruction reference line around the current block.
[0535] In other words, entropy encoding / entropy decoding can be performed on at least one of the sub-block partitioning mode information or partitioning direction information only when the current block uses the closest reference sample line as the reference sample line.
[0536] At least one of the area, size, shape, or partitioning orientation of the coefficient set used during entropy encoding / decoding of the transform coefficients can be determined based on at least one of the area, size, shape, or partitioning orientation of the sub-block.
[0537] For example, when the area of a sub-block is 16, the area of the coefficient group can be determined to be 16.
[0538] As another example, when the area of the sub-block is 32, the area of the coefficient group can be determined to be 16.
[0539] As another example, when the size of the sub-block is 1×16 or 16×1, the size of the coefficient group can be determined to be 1×16 or 16×1.
[0540] As another example, when the size of the sub-block is 2×8 or 8×2, the size of the coefficient group can be determined to be 2×8 or 8×2.
[0541] As another example, when the size of the sub-block is 4×4, the size of the coefficient group can be determined to be 4×4.
[0542] As another example, when the width of the sub-block is 2, the width of the coefficient group can be determined to be 2.
[0543] As another example, when the width of the sub-block is 4, the width of the coefficient group can be determined to be 4.
[0544] As another example, when the height of the sub-block is 2, the height of the coefficient group can be determined to be 2.
[0545] As another example, when the height of the sub-block is 4, the height of the coefficient group can be determined to be 4.
[0546] As another example, when the shape of the sub-block is rectangular (not square), the shape of the coefficient group can be determined to be rectangular (not square).
[0547] As another example, when the shape of the sub-block is a square, the shape of the coefficient group can be determined to be a square.
[0548] As another example, when the size of the sub-block is 16×4 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 8×2.
[0549] As another example, when the size of the sub-block is 4×8 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 2×8.
[0550] As another example, when the size of the sub-block is 32×4 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 8×2.
[0551] As another example, when the size of the sub-block is 8×64 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 2×8.
[0552] As another example, when the size of the sub-block is 16×4 and the partitioning direction is vertical, the size of the coefficient group can be determined to be 4×4.
[0553] As another example, when the size of the sub-block is 4×8 and the partitioning direction is horizontal, the size of the coefficient group can be determined to be 4×4.
[0554] As another example, when the size of the sub-block is 32×4 and the partitioning direction is horizontal, the size of the coefficient group can be determined to be 8×2.
[0555] As another example, when the size of the sub-block is 8×64 and the partitioning direction is vertical, the size of the coefficient group can be determined to be 2×8.
[0556] For each sub-block of partition, entropy coding / entropy decoding can be performed on the coded block flag that indicates whether there is at least one transform coefficient with a non-zero value in the sub-block unit.
[0557] For example, a coded block flag can indicate, on a sub-block basis, that at least one transform coefficient with a non-zero value exists in at least one sub-block.
[0558] As another example, when m represents the total number of sub-blocks and the coding block flags of the m-1 sub-blocks preceding the sub-blocks indicate that there are no transform coefficients with non-zero values, it can be inferred that the coding block flag of the m-th sub-block indicates that there is at least one transform coefficient with a non-zero value.
[0559] As another example, when the coded block flag is entropy encoded / entropy decoded in units of sub-blocks, the coded block flag can be entropy encoded / entropy decoded in units of blocks preceding the partition.
[0560] As another example, when the coded block flag is entropy encoded / decoded on a block-by-block basis before partitioning, the coded block flag may not be entropy encoded / decoded on a sub-block basis.
[0561] Furthermore, when the current block is in the first sub-block partitioning mode and the current block size is a predefined size, the size of the sub-block used for intra-frame prediction and the size of the sub-block used for transformation can be different from each other. That is, the sub-block partitions used for intra-frame prediction and the sub-block partitions used for transformation can be different from each other. Here, the predefined size can be 4×N or 8×N (N>4). Here, sub-block partitioning can represent vertical partitioning.
[0562] For example, when the current block is in the first sub-block partitioning mode and the current block size is 8×N (N>4), the current block can be vertically partitioned into sub-blocks of size 4×N for intra-frame prediction, and the current block can be vertically partitioned into sub-blocks of size 1×N for transform. In this case, a one-dimensional transform / inverse transform can be performed to perform a transform of size 1×N. That is, a one-dimensional transform / inverse transform can be performed based on at least one of the current block's partitioning mode or the current block's size.
[0563] As another example, when the current block is in the first sub-block partitioning mode and the current block size is 8×N (N>4), the current block can be vertically partitioned into sub-blocks of size 4×N for intra-frame prediction, and the current block can be vertically partitioned into sub-blocks of size 2×N for transform. In this case, a two-dimensional transform / inverse transform can be performed to perform a transform of size 2×N. That is, a two-dimensional transform / inverse transform can be performed based on at least one of the current block's partitioning mode or the current block's size.
[0564] Here, N can represent a positive integer, and can be a positive integer less than 64 or 128.
[0565] In addition, the size of the current block can represent at least one of the size of the coded block, the size of the predicted block, or the size of the transformed block.
[0566] As in FIG. 10 In the example, according to the embodiment of the first sub-block partitioning mode, the current block can be partitioned into two sub-blocks in the horizontal direction and into two sub-blocks in the vertical direction.
[0567] As in FIG. 11 In the example, according to the embodiment of the first sub-block partitioning mode, the current block can be partitioned into two sub-blocks in the horizontal direction and into two sub-blocks in the vertical direction.
[0568] As in FIG. 12In the example, according to the embodiment of the first sub-block partitioning mode, the current block can be partitioned into four sub-blocks in the horizontal direction and into four sub-blocks in the vertical direction.
[0569] Using at least one of the embodiments described below, a block can be partitioned into N sub-blocks to perform at least one of transform / inverse transform, quantization / dequantization, or entropy coding / entropy decoding on a sub-block basis. Such a mode may be referred to as a second sub-block partitioning mode (e.g., SBT mode or sub-block transform mode).
[0570] A block can represent at least one of a coding block, a prediction block, or a transform block. For example, a block can be a transform block.
[0571] Furthermore, the partitioned sub-blocks can represent at least one of a coded block, a prediction block, or a transform block. For example, the partitioned sub-blocks can be transform blocks.
[0572] Furthermore, a sub-block derived from a block or partition can be at least one of an intra-block, an inter-block, or an intra-block copy block. For example, a sub-block derived from a block or partition can be an inter-block.
[0573] Furthermore, a block or a sub-block partitioned from a frame can be at least one of an intra-prediction block, an inter-prediction block, or an intra-block copy prediction block. For example, a block can be an inter-prediction block.
[0574] Furthermore, a block or sub-block can be at least one of a luminance signal block or a chrominance signal block. For example, a block and a sub-block can be luminance signal blocks.
[0575] When a block is partitioned into N sub-blocks, the block before partitioning can be a coded block, and the partitioned sub-blocks can be at least one of a prediction block or a transform block. That is, the size of the partitioned sub-blocks can be used to perform transform / inverse transform, quantization / dequantization, and entropy coding / entropy decoding on the transform coefficients.
[0576] Furthermore, when a block is partitioned into N sub-blocks, the block preceding the partition can be at least one of a coding block or a prediction block, and the partitioned sub-blocks can be transform blocks. That is, prediction is performed using the size of the block preceding the partition, and transform / inverse transform, quantization / dequantization, and entropy coding / decoding of the transform coefficients can be performed using the size of the partitioned sub-blocks.
[0577] Whether a block is divided into multiple sub-blocks can be determined based on at least one of the block's area (the product of its width and height), size (width or height or a combination of width and height), or shape / form (rectangle, square, etc.).
[0578] For example, when the current block is a 64×64 block, the current block can be partitioned into multiple sub-blocks.
[0579] As another example, when the current block is a 32×32 block, the current block can be partitioned into multiple sub-blocks.
[0580] As another example, when the current block is a 32×16 block, the current block can be partitioned into multiple sub-blocks.
[0581] As another example, when the current block is a 16×32 block, the current block can be partitioned into multiple sub-blocks.
[0582] As another example, when the current block is a 4×4 block, the current block may not be divided into multiple sub-blocks.
[0583] As another example, when the current block is a 2×4 block, the current block may not be divided into multiple sub-blocks.
[0584] As another example, the current block may not be divided into multiple sub-blocks when at least one of its width or height is greater than the maximum size of the transform block. That is, when at least one of its width or height is less than or equal to the maximum size of the transform block, a second sub-block partitioning mode may be applied to the current block.
[0585] As another example, the current block may not be divided into multiple sub-blocks when at least one of its width or height is greater than the maximum size of the transform block. That is, when both the width and height of the current block are less than or equal to the maximum size of the transform block, a second sub-block partitioning mode can be applied to the current block.
[0586] Furthermore, when at least one of the width or height of the current block is less than or equal to the maximum size of the transform block, at least one of the information indicating the partitioning mode of the second sub-block (sub-block partitioning mode information, partitioning direction information, sub-block position information, or sub-block size information) can be entropy encoded / entropy decoded. Here, the maximum size of the transform block can be determined based on the maximum size information of the transform block transmitted by a signal in a higher-level unit. For example, the maximum size of the transform block can be determined to be either 64 or 32 based on the maximum size information of the transform block.
[0587] As another example, when the area of the current block is equal to or greater than 32, the current block can be divided into multiple sub-blocks.
[0588] As another example, when the area of the current block is less than 32, the current block may not be divided into multiple sub-blocks.
[0589] As another example, when the area of the current block is 256 and the shape of the current block is rectangular, the current block can be divided into multiple sub-blocks.
[0590] As another example, when the area of the current block is 16 and the shape of the current block is a square, the current block may not be divided into multiple sub-blocks.
[0591] The current block can be divided into multiple sub-blocks when at least one of its width or height is greater than or equal to a predefined value.
[0592] For example, when the width or height of the current block is equal to or greater than 8, the current block can be divided into multiple sub-blocks.
[0593] Conversely, if both the width and height of the current block are less than the predefined values, the current block may not be divided into multiple sub-blocks.
[0594] When the current block is in GPM (Geometric Partitioning Mode), the transform block of the current block may not be partitioned into multiple sub-blocks. Here, GPM can be a prediction mode that partitions the prediction block of the current block into two sub-blocks to perform prediction. When the current block is in GPM, the prediction block of the current block can be partitioned into two sub-blocks. In this case, entropy coding / entropy decoding can be performed on information related to the partitioning direction used to partition the prediction block of the current block into two sub-blocks. Inter-frame prediction can be performed for the two sub-blocks, thereby generating prediction samples for the two sub-blocks. Furthermore, a weighted summation can be performed on the generated prediction samples for the two sub-blocks to derive the prediction samples for the current block. That is, when the prediction block of the current block is partitioned into at least two sub-blocks, the transform block of the current block may not be partitioned into at least two sub-blocks. Similarly, when the prediction block of the current block is not partitioned into at least two sub-blocks, the transform block of the current block can be partitioned into at least two sub-blocks.
[0595] When a block is partitioned, it can be divided into multiple sub-blocks in at least one of the vertical or horizontal directions.
[0596] For example, the current block can be divided into two sub-blocks in the vertical direction.
[0597] As another example, the current block can be divided into two sub-blocks in the horizontal direction.
[0598] When a block is divided into N sub-blocks, N can be a positive integer, and can be, for example, 2. Furthermore, N can be determined using at least one of the block's area, size, shape, or partitioning orientation.
[0599] For example, when the current block is a 4×8 or 8×4 block, the current block can be divided into two sub-blocks in the horizontal direction or in the vertical direction.
[0600] As another example, when the current block is a 16×8 or 16×16 block, the current block can be divided into two sub-blocks in the vertical direction or in the horizontal direction.
[0601] As another example, when the current block is an 8×32 or 32×32 block, the current block can be divided into two sub-blocks in the horizontal direction or in the vertical direction.
[0602] As another example, when the current block is a J×8 block, the current block can be divided into two sub-blocks in the horizontal direction. Here, J can be a positive integer.
[0603] As another example, when the current block is an 8×K block, the current block can be divided into two sub-blocks in the vertical direction. Here, K can be a positive integer.
[0604] As another example, when the current block is a J×K (K>8) block, the current block can be divided into two sub-blocks in the horizontal direction. Here, J can be a positive integer. In this case, the height of the sub-blocks can have a ratio of 1:3 or 3:1.
[0605] As another example, when the current block is a J (J > 8) × K block, the current block can be divided into two sub-blocks in the vertical direction. Here, J can be a positive integer. In this case, the widths of the sub-blocks can have a ratio of 1:3 or 3:1.
[0606] As another example, when the area of the current block is 64, the current block can be divided into two sub-blocks in either the horizontal or vertical direction.
[0607] As another example, when the current block is a 16×4 block and the current block shape is rectangular, the current block can be divided into two sub-blocks in the vertical direction.
[0608] As another example, when the area of the current block is 1024 and the shape of the current block is a square, the current block can be divided into two sub-blocks in the horizontal or vertical direction.
[0609] In addition, a sub-block may have at least one of a minimum area, minimum width, or minimum height.
[0610] For example, a sub-block can have a minimum area of S. Here, S can be a positive integer, and can be, for example, 16.
[0611] As another example, a sub-block can have a minimum width of J. Here, J can be a positive integer, and can be, for example, 4.
[0612] As another example, a sub-block can have a minimum height of K. Here, K can be a positive integer, and can be, for example, 4.
[0613] To indicate a sub-block partitioning mode that divides a block into N sub-blocks to perform at least one of transform / inverse transform, quantization / dequantization, or entropy encoding / entropy decoding, at least one of sub-block partitioning mode information, partitioning direction information, sub-block position information, or sub-block size information can be entropy encoded / entropy decoded.
[0614] Here, sub-block partitioning mode information can be used to indicate the sub-block partitioning mode. When the sub-block partitioning mode information indicates that the sub-block partitioning mode is used (second value), the block can be partitioned into sub-blocks, and at least one of transform / inverse transform, quantization / dequantization, or entropy encoding / decoding can be performed. When the sub-block partitioning mode information indicates that the sub-block partitioning mode is not used (first value), the block may not be partitioned into sub-blocks, and at least one of transform / inverse transform, quantization / dequantization, or entropy encoding / decoding can be performed. Here, the first value can be 0, and the second value can be 1.
[0615] Furthermore, partitioning direction information can be used to indicate whether the sub-block partitioning mode is vertical or horizontal. When the partitioning direction information has a first value, the sub-block can be partitioned vertically, and the first value can be 0. Furthermore, when the partitioning direction information has a second value, the sub-block can be partitioned horizontally, and the second value can be 1.
[0616] Furthermore, sub-block location information can be used to indicate which sub-block's residual signal is encoded / decoded. When the sub-block location information has a first value, the residual signal of the first sub-block can be encoded / decoded, and the first value can be 0. Furthermore, when the sub-block location information has a second value, the residual signal of the second sub-block can be encoded / decoded, and the second value can be 1. Furthermore, when the sub-block location information has a first value, entropy encoding / decoding can be performed on at least one of the coded block flags for the luminance signal or the coded block flags for the chrominance signal of the first sub-block. Furthermore, when the sub-block location information has a second value, entropy encoding / decoding can be performed on at least one of the coded block flags for the luminance signal or the coded block flags for the chrominance signal of the second sub-block.
[0617] Furthermore, sub-block size information can be used to indicate whether the width or height of the partitioned sub-block is 1 / 2 or 1 / 4 of the width or height of the block. When the sub-block size information has a first value, this indicates that the width or height of the sub-block whose residual signal is encoded / decoded by the sub-block position information is 1 / 2 of the width or height of the block, and the first value can be 0. Furthermore, when the sub-block size information has a second value, this indicates that the width or height of the sub-block whose residual signal is encoded / decoded by the sub-block position information is 1 / 4 of the width or height of the block, and the second value can be 1. For example, when only the size of the partitioned sub-block is 1 / 2 of the width or height of the block exists, entropy encoding / decoding of the sub-block size information is not required.
[0618] Within each partitioned sub-block, a reconstructed block can be generated by adding the residual block (or reconstructed residual block) to the predicted block.
[0619] The encoding / decoding order of sub-blocks partitioned from a block can be determined based on at least one of the partition directions.
[0620] For example, the encoding / decoding order of sub-blocks partitioned in the horizontal direction can be determined as a top-to-bottom order.
[0621] As another example, the encoding / decoding order of sub-blocks partitioned in the vertical direction can be determined as a left-to-right order.
[0622] Within the partitioned sub-blocks, a coded block flag indicating the presence of at least one transform coefficient with a non-zero value can be entropy encoded / entropy decoded on a sub-block basis.
[0623] For example, a coded block flag can indicate, on a sub-block basis, that at least one transform coefficient with a non-zero value exists in at least one sub-block.
[0624] As another example, when the coded block flag is entropy encoded / decoded on a sub-block basis, it is not necessary to entropy encode / decode the coded block flag on a block basis prior to the partition.
[0625] As another example, when the coded block flag is entropy encoded / decoded on a block-by-block basis before partitioning, the coded block flag may not be entropy encoded / decoded on a sub-block basis.
[0626] The residual signal can be entropy encoded / entropy decoded only for sub-blocks indicated by sub-block location information.
[0627] At this point, since the residual signal can always exist in the sub-block indicated by the sub-block position information, it can be inferred that the coded block flag indicates the existence of at least one non-zero transform coefficient.
[0628] Furthermore, since the residual signal may not always exist in sub-blocks not indicated by sub-block location information, it can be inferred that the coded block flag indicates the absence of at least one non-zero transform coefficient.
[0629] At least one of the area, size, shape, or partitioning orientation of the coefficient set used when the transform coefficients are entropy encoded / entropy decoded can be determined based on at least one of the area, size, shape, or partitioning orientation of the sub-block.
[0630] For example, when the area of a sub-block is 16, the area of the coefficient group can be determined to be 16.
[0631] As another example, when the area of the sub-block is 32, the area of the coefficient group can be determined to be 16.
[0632] As another example, when the size of the sub-block is 1×16 or 16×1, the size of the coefficient group can be determined to be 1×16 or 16×1.
[0633] As another example, when the size of the sub-block is 2×8 or 8×2, the size of the coefficient group can be determined to be 2×8 or 8×2.
[0634] As another example, when the size of the sub-block is 4×4, the size of the coefficient group can be determined to be 4×4.
[0635] As another example, when the width of the sub-block is 2, the width of the coefficient group can be determined to be 2.
[0636] As another example, when the width of the sub-block is 4, the width of the coefficient group can be determined to be 4.
[0637] As another example, when the height of the sub-block is 2, the height of the coefficient group can be determined to be 2.
[0638] As another example, when the height of the sub-block is 4, the height of the coefficient group can be determined to be 4.
[0639] As another example, when the shape of the sub-block is rectangular (not square), the shape of the coefficient group can be determined to be rectangular (not square).
[0640] As another example, when the shape of the sub-block is a square, the shape of the coefficient group can be determined to be a square.
[0641] As another example, when the size of the sub-block is 16×4 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 8×2.
[0642] As another example, when the size of the sub-block is 4×8 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 2×8.
[0643] As another example, when the size of the sub-block is 32×4 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 8×2.
[0644] As another example, when the size of the sub-block is 8×64 and its form is rectangular, the size of the coefficient group can be determined to be at least one of 4×4 or 2×8.
[0645] As another example, when the size of the sub-block is 16×4 and the partitioning direction is vertical, the size of the coefficient group can be determined to be 4×4.
[0646] As another example, when the size of the sub-block is 4×8 and the partitioning direction is horizontal, the size of the coefficient group can be determined to be 4×4.
[0647] As another example, when the size of the sub-block is 32×4 and the partitioning direction is horizontal, the size of the coefficient group can be determined to be 8×2.
[0648] As another example, when the size of the sub-block is 8×64 and the partitioning direction is vertical, the size of the coefficient group can be determined to be 2×8.
[0649] Furthermore, the area of the coefficient set used when the transform coefficients are entropy encoded / entropy decoded can be determined as a predefined value. Here, the predefined value can be 4 or 16.
[0650] Furthermore, the area or size of the transformation coefficient group can be determined based on the size of the current block, regardless of the color components of the current block. In this case, the size of the current block can include at least one of the width or height of the current block.
[0651] For example, when the width and height of the current block are 2, the size of the coefficient group can be determined as 2×2.
[0652] As another example, when the size of the current block is 2×4 or 4×2, the size of the coefficient group can be determined to be 2×2.
[0653] As in FIG. 13 In the example, according to the embodiment of the second sub-block partitioning mode, the current block can be partitioned into two sub-blocks horizontally (1 / 2 or 1 / 4 of the height) and vertically (1 / 2 or 1 / 4 of the width). FIG. 13 In the example, the gray shading can represent the block in which the residual signal in the sub-block is encoded / decoded, and the sub-block location information can be used to indicate the block.
[0654] The sub-block partitioning mode usage information may be entropy encoded / entropy decoded in at least one of the parameter set or header, wherein the sub-block partitioning mode usage information indicates whether a mode is used to partition the block into N sub-blocks and perform at least one of prediction, transform / inverse transform, quantization / dequantization or entropy encoding / entropy decoding.
[0655] Here, the sub-block partitioning mode information can represent at least one of the first sub-block partitioning mode or the second sub-block partitioning mode.
[0656] At this time, at least one of the parameter set or header can be a video parameter set, a decoding parameter set, a sequence parameter set, an adaptive parameter set, a picture parameter set, a picture header, a strip header, a parallel block group header, or a parallel block header.
[0657] For example, to indicate whether a sub-block partitioning mode is used within a video, entropy encoding / entropy decoding of the sub-block partitioning mode usage information can be performed in the video parameter set.
[0658] As another example, to indicate whether a sub-block partitioning pattern is used within the decoding process, the sub-block partitioning pattern usage information can be entropy encoded / entropy decoded in the sequence parameter set.
[0659] As another example, in order to indicate whether a sub-block partitioning pattern is used within a sequence, the sub-block partitioning pattern usage information can be entropy encoded / entropy decoded in the sequence parameter set.
[0660] As another example, to indicate whether a sub-block partitioning mode is used within a number of frames, the sub-block partitioning mode usage information can be entropy encoded / entropy decoded in the adaptive parameter set or the adaptive header.
[0661] As another example, to indicate whether a sub-block partitioning mode is used within the frame, the sub-block partitioning mode usage information can be entropy encoded / entropy decoded in the frame parameter set or frame header.
[0662] As another example, to indicate whether a sub-block partitioning mode is used within a stripe, the sub-block partitioning mode usage information can be entropy encoded / entropy decoded in the stripe header.
[0663] As another example, to indicate whether a sub-block partitioning pattern is used within a parallel block group, the sub-block partitioning pattern usage information can be entropy encoded / entropy decoded in the parallel block group header.
[0664] As another example, to indicate whether a sub-block partitioning pattern is used within a parallel block, the sub-block partitioning pattern usage information can be entropy encoded / entropy decoded in the parallel block header.
[0665] The transformation / inverse transformation type of each block or sub-block can be determined using at least one of the embodiments in the following examples.
[0666] The type of one-dimensional transform, combination of two-dimensional transforms, or whether to use at least one of the following can be determined based on at least one of the following: prediction mode for a block or sub-block, intra-frame prediction mode, color components, size, shape (form), sub-block partitioning information, secondary transform execution information, or matrix-based intra-frame prediction execution information. Matrix-based intra-frame prediction can indicate matrix-based intra-frame prediction.
[0667] For example, a one-dimensional transform type indicating at least one of the horizontal or vertical transform types can be determined based on at least one of the following: prediction mode for blocks or sub-blocks, intra-frame prediction mode, color components, size, shape (form), sub-block partitioning information, secondary transform execution information, or matrix-based intra-frame prediction execution information.
[0668] As another example, a two-dimensional transform combination indicating a combination of horizontal and vertical transform types can be determined based on at least one of intra-frame prediction mode, prediction mode, color components, size, shape, or sub-block partitioning information for a block or sub-block.
[0669] As another example, information indicating whether to perform a transformation can be determined based on at least one of intra-prediction mode, prediction mode, color components, size, shape, or sub-block partitioning related information for a block or sub-block.
[0670] At this point, the one-dimensional transformation type, the combination of two-dimensional transformations, or whether at least one of the transformations is used may differ from each other based on at least one of the intra-frame prediction mode, prediction mode, color components, size, shape, or sub-block partitioning information for the block or sub-block.
[0671] Furthermore, when determining at least one of the following for a block or sub-block: intra-prediction mode, prediction mode, color components, size, shape (form), sub-block partitioning information, secondary transform execution information, or matrix-based intra-prediction execution information, the information regarding the one-dimensional transform type, the two-dimensional transform combination, or whether a transform is used may not be entropy-coded / entropy-decoded.
[0672] In other words, the type of one-dimensional transform, combination of two-dimensional transforms, or whether to use at least one of the transforms for a block or sub-block can be implicitly determined according to predetermined rules in the encoder / decoder. These predetermined rules can be set based on encoding parameters in the encoder / decoder.
[0673] Here, matrix-based intra-frame prediction can represent an intra-frame prediction mode that performs at least one of boundary averaging, matrix-vector multiplication, or linear interpolation to generate a prediction block.
[0674] Here, transformation can refer to at least one of transformation or inverse transformation.
[0675] In addition, a block can represent each sub-block partitioned from a block.
[0676] The primary transform can represent at least one integer transform based on DCT-J or DST-K (such as DCT-2, DCT-8, DST-7, DCT-4, or DST-4 performed on the residual block to generate transform coefficients). Here, J and K can be positive integers.
[0677] A primary transformation can be performed using a transformation matrix extracted from the transformation matrix of at least one integer transformation based on DCT-J or DST-K (such as DCT-2, DCT-8, DST-7, DCT-4, or DST-4). That is, a primary transformation can be performed using the extracted transformation matrix. Furthermore, at least one coefficient in the extracted transformation matrix can be equal to at least one coefficient in the transformation matrix of at least one integer transformation based on DCT-J or DST-K (such as DCT-2, DCT-8, DST-7, DCT-4, or DST-4). Additionally, the extracted transformation matrix can be included in the transformation matrix to be extracted. Furthermore, the extracted transformation matrix can be obtained by performing at least one of flipping or sign changing on specific coefficients in the transformation matrix to be extracted.
[0678] For example, at least one integer transform based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4 or DST-4) can be extracted from the transform matrix of DCT-2 and used for the primary transform.
[0679] Here, at least one integer transformation based on DCT-J or DST-K (such as DCT-2, DCT-8, DST-7, DCT-4 or DST-4) may have different coefficients in the transformation matrix than at least one integer transformation based on DCT-J or DST-K (such as DCT-2, DCT-8, DST-7, DCT-4 or DST-4).
[0680] For example, an integer transformation matrix based on DCT-8 can be derived by performing at least one of a horizontal flip of the DST-7-based integer transformation matrix or a sign change of at least one coefficient of the DST-7 transformation matrix. In this case, a vertical flip can be used instead of a horizontal flip.
[0681] As another example, a DST-7-based integer transformation matrix can be derived by performing at least one of a horizontal flip of the DCT-8-based integer transformation matrix or a sign change of at least one coefficient of the DCT-8 transformation matrix. In this case, a vertical flip can be used instead of a horizontal flip.
[0682] As another example, a DST-4-based integer transformation matrix can be derived by performing at least one of a horizontal flip of the DST-4-based integer transformation matrix or a sign change of at least one coefficient of the DST-4 transformation matrix. In this case, a vertical flip can be used instead of a horizontal flip.
[0683] As another example, a DST-4-based integer transformation matrix can be derived by performing at least one of a horizontal flip of the DCT-4-based integer transformation matrix or a sign change of at least one coefficient of the DCT-4 transformation matrix. In this case, a vertical flip can be used instead of a horizontal flip.
[0684] A secondary transformation can represent at least one of a transformation based on at least one of the transformation coefficients in the angular rotation transformation. A secondary transformation can be performed after the primary transformation.
[0685] In an encoder, a secondary transformation can be performed on the coefficients in the low-frequency region to the upper left of the transform coefficients after the primary transformation. The size of the low-frequency region to which the secondary transformation is applied can be determined based on the size of the transform block.
[0686] In the decoder, a secondary inverse transform may be performed before the primary inverse transform. In the following description, the secondary transform may include the secondary inverse transform.
[0687] The secondary transform can be called LFNST (Low Frequency Non-Separable Transform) because it uses a non-separable transform kernel instead of horizontal and vertical separable transform kernels (or types).
[0688] Secondary transforms can be performed only for intra-frame predictive coding / decoding, and the secondary transform kernels can be determined based on the intra-frame prediction mode. Specifically, a transform set including multiple transform kernels can be determined based on the intra-frame prediction mode. Furthermore, the transform kernels to be applied to the secondary transforms can be determined from the transform set determined based on index information. Here, the transform set can include four types of transform sets.
[0689] Furthermore, when the intra-prediction mode of the current block is CCLM mode, the transform set for the chroma block can be determined based on the intra-prediction mode of the luma block corresponding to the chroma block. Here, when the luma block corresponding to the chroma block is in matrix-based intra-prediction mode, this can be considered a planar mode, and the transform set for the chroma block can be determined. Furthermore, when the luma block corresponding to the chroma block is in IBC mode, this can be considered a DC mode, and the transform set for the chroma block can be determined.
[0690] As another example, when the luma block corresponding to the chroma block is in matrix-based intra-prediction mode, IBC mode, or palette mode, this can be considered a planar mode and the transform set for the chroma block can be determined.
[0691] As another example, when the luma block corresponding to the chroma block is in matrix-based intra-prediction mode, IBC mode, or palette mode, this can be regarded as DC mode and the transform set for the chroma block can be determined.
[0692] Furthermore, secondary transformation index information can be used to determine the secondary transformation kernel in the transformation set. Here, the secondary transformation kernel can indicate the secondary transformation matrix.
[0693] Whether to use a transformation indicates whether at least one of a primary transformation or a secondary transformation is used in the residual block. Whether to use a transformation may include whether at least one of a primary transformation or a secondary transformation is used. Whether to use a transformation may indicate whether a transformation skip mode is applied. Additionally, `transform_skip_flag` indicates a transformation skip flag.
[0694] For example, when transform_skip_flag, which serves as information indicating whether at least one of the primary or secondary transformations is used, has a first value (e.g., 0), this can indicate that at least one of the primary or secondary transformations is used.
[0695] As another example, when transform_skip_flag, which serves as information indicating whether at least one of the primary or secondary transformations is used, has a second value (e.g., 1), this can indicate that at least one of the primary or secondary transformations has not been used.
[0696] A one-dimensional transformation type can represent the type of a primary transformation, and can also represent a horizontal transformation type (trTypeHor) or a vertical transformation type (trTypeVer) for at least one of the integer transformation types based on DCT-J or DST-K. Here, J and K can be positive integers.
[0697] As a type of primary transformation, transformations from the first to the Nth can be used. Here, N can be a positive integer of 2 or a larger positive integer.
[0698] For example, the first transformation can represent an integer transformation based on DCT-2.
[0699] As another example, when the first transformation is applied to both horizontal and vertical transformations, trTypeHor, the transformation type for the horizontal transformation, and trTypeVer, the transformation type for the vertical transformation, can have values Q and R, respectively. Here, Q and R can be at least one of negative integers, 0, or positive integers. For example, Q and R can be 0 and 0, respectively.
[0700] For example, when trTypeHor has a first value, this can represent an integer-level transformation based on DCT-2.
[0701] As another example, when trTypeVer has a first value, this can represent an integer vertical transform based on DCT-2.
[0702] The first value can be 0.
[0703] For example, the second transformation may represent at least one integer transformation based on an integer transformation other than DCT-2, such as DCT-J or DST-K (e.g., DCT-8, DST-7, DCT-4, or DST-4). Here, J and K can be positive integers. That is, the second transformation may represent at least one transformation other than the first transformation.
[0704] As another example, when the second transformation is applied to at least one of a horizontal or vertical transformation, trTypeHor, the transformation type for the horizontal transformation, and trTypeVer, the transformation type for the vertical transformation, can have values T and U, respectively. Here, T and U can be at least one of a negative integer, 0, or a positive integer. For example, T and U can be values equal to or greater than 1 and 1, respectively. Furthermore, T and U can be greater than Q and R, respectively.
[0705] For example, when trTypeHor has a second value, this can represent an integer level transformation based on DST-7.
[0706] As another example, when trTypeHor has a third value, this can represent an integer-level transformation based on DCT-8.
[0707] As another example, when trTypeVer has a second value, this can represent an integer vertical transformation based on DST-7.
[0708] As another example, when trTypeVer has a third value, this can represent an integer vertical transform based on DCT-8.
[0709] The second value can be 1. Furthermore, the third value can be 2.
[0710] DST-4 can be used instead of DST-7. Additionally, DCT-4 can be used instead of DCT-8.
[0711] For example, the first transform can be an integer transform based on DCT-2. Furthermore, the second transform can be an integer transform based on DST-7. Furthermore, the third transform can be an integer transform based on DCT-8. Additionally, the second transform can represent at least one of the second or third transforms.
[0712] As another example, the first transformation can be an integer transformation based on DCT-2. Furthermore, the second transformation can be an integer transformation based on DST-4. Furthermore, the third transformation can be an integer transformation based on DCT-4. Additionally, the second transformation can represent at least one of the second or third transformations.
[0713] In other words, the first transformation can be an integer transformation based on DCT-2, and the second to Nth transformations can represent at least one integer transformation based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4, or DST-4) other than DCT-2. Here, N can be a positive integer equal to or greater than 3.
[0714] For example, the first transformation can be a DCT-2 based integer transformation. Furthermore, the second transformation can be a DST-7 based integer transformation extracted from a DCT-2 based integer transformation matrix. Furthermore, the third transformation can be a DCT-8 based integer transformation extracted from a DCT-2 based integer transformation matrix. Additionally, the second transformation can represent at least one of the second or third transformations.
[0715] As another example, the first transformation can be a DCT-2 based integer transformation. Furthermore, the second transformation can be a DCT-4 based integer transformation extracted from a DCT-2 based integer transformation matrix. Furthermore, the third transformation can be a DCT-4 based integer transformation extracted from a DCT-2 based integer transformation matrix. Additionally, the second transformation can represent at least one of the second or third transformations.
[0716] In other words, the first transformation can be an integer transformation based on DCT-2, and the second to Nth transformations can represent at least one integer transformation among those based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4, or DST-4) extracted from the integer transformation matrix based on DCT-2. Here, N can be a positive integer equal to or greater than 3. Furthermore, the second transformation can represent at least one of the second to Nth transformations.
[0717] As an alternative to DCT-2, at least one integer transform based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4 or DST-4) may be used.
[0718] Two-dimensional transform combinations can represent combinations of primary transforms, and can also represent combinations of horizontal transform type trTypeHor and vertical transform type trTypeVer for at least one integer transform type among DCT-J or DST-K based integer transform types. Furthermore, two-dimensional transform combinations can represent mts_idx as a multi-transform selection index.
[0719] For example, when the first transformation is used for both horizontal and vertical transformations, the mts_idx, which serves as the multi-transformation selection index, can have a value P. Here, P can be at least one of a negative integer, 0, or a positive integer. For example, P can be 0.
[0720] For example, when mts_idx is 0, trTypeHor and trTypeVer can each have a first value (e.g., 0). That is, when mts_idx is 0, this can represent an integer horizontal transform based on DCT-2 and an integer vertical transform based on DCT-2.
[0721] As another example, when the second transformation is used for at least one of the horizontal or vertical transformations, the mts_idx, which serves as the multi-transformation selection index, can have a value S or greater. Here, S can be at least one of a negative integer, 0, or a positive integer. For example, S can be 1. Furthermore, S can be greater than P.
[0722] For example, when mts_idx is 1, trTypeHor and trTypeVer can each have a second value (e.g., 1).
[0723] As another example, when mts_idx is 2, trTypeHor and trTypeVer can have a third value (e.g., 2) and a second value (e.g., 1), respectively.
[0724] As another example, when mts_idx is 3, trTypeHor and trTypeVer can have a second value (e.g., 1) and a third value (e.g., 2), respectively.
[0725] As another example, when mts_idx is 4, trTypeHor and trTypeVer can each have a third value (e.g., 2).
[0726] For example, when trTypeHor has a second value, this can represent an integer level transformation based on DST-7.
[0727] For another example, when trTypeHor has a third value, this could represent an integer-level transformation based on DCT-8.
[0728] For another example, when trTypeVer has a second value, this could represent an integer vertical transformation based on DST-7.
[0729] For another example, when trTypeVer has a third value, this could represent an integer vertical transform based on DCT-8.
[0730] The second value can be 1. Furthermore, the third value can be 2.
[0731] In the above embodiments, DST-4 can be used instead of DST-7. Furthermore, DCT-4 can be used instead of DCT-8.
[0732] For example, in the first transformation, the horizontal and vertical transformations can be DCT-2-based integer transformations, respectively. Furthermore, in the second transformation, the horizontal and vertical transformations can be DST-7-based integer transformations and DST-7-based integer transformations, respectively. Furthermore, in the third transformation, the horizontal and vertical transformations can be DCT-8-based integer transformations and DST-7-based integer transformations, respectively. Furthermore, in the fourth transformation, the horizontal and vertical transformations can be DST-7-based integer transformations and DCT-8-based integer transformations, respectively. Furthermore, in the fifth transformation, the horizontal and vertical transformations can be DCT-8-based integer transformations and DCT-8-based integer transformations, respectively. Moreover, the second transformation can represent at least one of the second, third, fourth, or fifth transformations.
[0733] As another example, in the first transformation, the horizontal and vertical transformations can be DCT-2-based integer transformations, respectively. Furthermore, in the second transformation, the horizontal and vertical transformations can be DST-4-based integer transformations and DST-4-based integer transformations, respectively. Furthermore, in the third transformation, the horizontal and vertical transformations can be DCT-4-based integer transformations and DST-4-based integer transformations, respectively. Furthermore, in the fourth transformation, the horizontal and vertical transformations can be DST-4-based integer transformations and DCT-4-based integer transformations, respectively. Furthermore, in the fifth transformation, the horizontal and vertical transformations can be DCT-4-based integer transformations and DCT-4-based integer transformations, respectively. Moreover, the second transformation can represent at least one of the second, third, fourth, or fifth transformations.
[0734] That is, in the first transformation, the horizontal and vertical transformations can each be integer transformations based on DCT-2, and in the second to Nth transformations, the horizontal and vertical transformations can represent at least one integer transformation based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4, or DST-4) other than DCT-2. Here, N can be an integer equal to or greater than 3.
[0735] For example, in the first transformation, the horizontal and vertical transformations can be DCT-2-based integer transformations, respectively. Furthermore, in the second transformation, the horizontal and vertical transformations can be DST-7-based integer transformations and DST-7-based integer transformations extracted from the DCT-2-based integer transformation matrix, respectively. Furthermore, in the third transformation, the horizontal and vertical transformations can be DCT-8-based integer transformations and DST-7-based integer transformations extracted from the DCT-2-based integer transformation matrix, respectively. Furthermore, in the fourth transformation, the horizontal and vertical transformations can be DST-7-based integer transformations and DCT-8-based integer transformations extracted from the DCT-2-based integer transformation matrix, respectively. Furthermore, in the fifth transformation, the horizontal and vertical transformations can be DCT-8-based integer transformations and DCT-8-based integer transformations extracted from the DCT-2-based integer transformation matrix, respectively. Moreover, the second transformation can represent at least one of the second, third, fourth, or fifth transformations.
[0736] As another example, in the first transformation, the horizontal and vertical transformations can be DCT-2 based integer transformations, respectively. Furthermore, in the second transformation, the horizontal and vertical transformations can be DST-4 based integer transformations extracted from the DCT-2 based integer transformation matrix and DST-4 based integer transformations, respectively. Furthermore, in the third transformation, the horizontal and vertical transformations can be DCT-4 based integer transformations extracted from the DCT-2 based integer transformation matrix and DST-4 based integer transformations extracted from the DCT-2 based integer transformation matrix, respectively. Furthermore, in the fourth transformation, the horizontal and vertical transformations can be DST-4 based integer transformations extracted from the DCT-2 based integer transformation matrix and DCT-4 based integer transformations extracted from the DCT-2 based integer transformation matrix, respectively. Furthermore, in the fifth transformation, the horizontal and vertical transformations can be DCT-4 based integer transformations extracted from the DCT-2 based integer transformation matrix and DCT-4 based integer transformations. Moreover, the second transformation can represent at least one of the second, third, fourth, or fifth transformations.
[0737] That is, in the first transformation, the horizontal and vertical transformations can each be integer transformations based on DCT-2, and in the second to Nth transformations, the horizontal and vertical transformations can represent at least one integer transformation among those based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4, or DST-4) extracted from the integer transformation matrix based on DCT-2. Here, N can be a positive integer equal to or greater than 3. In this case, the second transformation can represent at least one of the second to Nth transformations.
[0738] As an alternative to the DCT-2 transform, at least one of the integer transforms based on DCT-J or DST-K (such as DCT-8, DST-7, DCT-4 or DST-4) may be used.
[0739] The prediction mode can represent the prediction mode of a block, and can indicate which of the intra-prediction mode, inter-prediction mode, and IBC (intra-block copy) mode is used to perform encoding / decoding.
[0740] For example, when both intra-frame prediction and inter-frame prediction are performed in a specific mode to generate a prediction block, the specific mode may represent an inter-frame prediction mode.
[0741] For example, when the current image is used as the reference image and vectors are used for prediction in a specific mode, the specific mode can represent an intra-block copy prediction mode. The intra-block copy prediction mode can be an IBC mode. Here, an IBC mode can represent a mode in which a reference region is set within the current image / strip / parallel block / parallel block group / CTU, the position of which is indicated by a block vector, and prediction is performed using the region indicated by the block vector.
[0742] Color components can represent the color components of a block and can also represent the luminance (Y) component or the chromaticity component.
[0743] For example, a chromaticity component can represent at least one of a Cb component or a Cr component. That is, a color component can represent a Y component, a Cb component, or a Cr component.
[0744] As another example, a color component can represent at least one of an R component, a G component, or a B component.
[0745] As another example, when an image is decomposed into multiple components and encoded / decoded, the color components can represent the decomposed components.
[0746] Sub-block partition information can indicate that a block is divided into multiple sub-blocks.
[0747] For example, sub-block partition information may include at least one of sub-block partition mode information or partition direction information.
[0748] As another example, sub-block partitioning related information may include at least one of sub-block partitioning mode information, partitioning direction information, sub-block location information, or sub-block size information.
[0749] A dimension can represent at least one of a block size, a sub-block size, or a transformation size. Here, a dimension can represent at least one of a width, a height, or a combination of width and height.
[0750] Transformation dimensions can represent the transformation dimensions used in the corresponding block. Transformation dimensions can be less than or equal to the block size.
[0751] The size can be M×N, such as 2×2, 4×2, 2×4, 4×4, 8×4, 8×2, 2×8, 8×8, 16×8, 16×4, 16×2, 2×16, 4×16, 8×16, 16×16, 32×16, 32×8, 32×4, 32×2, 2×32, 4×32, 8×32, 16×32, 32×32, 64×32, 64×16, 64×8, 64×4, 64×2, 2×64, 4×64, 8×64, 16×64, 32×64, 64×64, 128×64, 128×32, 32×128, 64×128, or 128×128. Here, M and N can be positive integers and can be the same or different. Furthermore, M can be S*N. N can be S*M. Here, S can be a positive integer.
[0752] Here, M can represent the width, and N can represent the height.
[0753] For example, in the case of a 64×64 block, a 32×32 transformation can be performed in the upper left region of the block. In this case, a 32×32 quantization matrix can be used.
[0754] As another example, in the case of a 64×32 block, a 32×32 transformation can be performed in the upper left region of the block. In this case, a 32×32 quantization matrix can be used.
[0755] As another example, in the case of a 32×64 block, a 16×32 transformation can be performed in the upper left region of the block. In this case, a 16×32 quantization matrix can be used.
[0756] As another example, in the case of a 32×32 block, a 32×32 transformation can be performed within the block. In this case, a 32×32 quantization matrix can be used.
[0757] A form (or shape) can represent at least one of the following: the form of a block, the form of a sub-block, or the form of a transformation.
[0758] The form can be square or non-square.
[0759] A square form can represent the shape of a square.
[0760] A non-square form can be represented as a rectangle.
[0761] The form of a transformation can represent the form of the transformation used in the corresponding block. When the horizontal and vertical transformation dimensions are different from each other, the transformation form can be non-square. Furthermore, when the horizontal and vertical transformation dimensions are the same, the transformation form can be square. The form of the transformation can be equal to or different from the form of the corresponding block.
[0762] The quantization matrix can be represented by the form of the quantization matrix used in the corresponding block. When the horizontal and vertical transform dimensions are different, the quantization matrix can be non-square. Conversely, when the horizontal and vertical transform dimensions are the same, the quantization matrix can be square. The form of the quantization matrix can be equal to or different from the form of the corresponding block. The form of the quantization matrix can be equal to or different from the form of the transform.
[0763] For example, in the case of a 64×64 square block, a 32×32 square transformation can be performed in the upper left region of the block. In this case, a 32×32 square quantization matrix can be used.
[0764] As another example, in the case of a 16×16 square block, a 16×16 square transformation can be performed within the block. In this case, a 16×16 square quantization matrix can be used.
[0765] As another example, in the case of a non-square block of size 16×4, a non-square transformation of size 16×4 can be performed within the block. In this case, a quantization matrix of size 16×4 can be used.
[0766] As another example, in the case of a non-square block of size 2×8, a transformation of size 2×8 can be performed within the block. In this case, a quantization matrix of size 2×8 can be used.
[0767] The following illustrates an embodiment for determining the type of one-dimensional transformation, a combination of two-dimensional transformations, or whether to use at least one of the transformations for a block or sub-block. Here, the type of one-dimensional transformation can be determined for at least one of the horizontal or vertical transformations for a block or sub-block.
[0768] When the width W or height H of a block or sub-block is less than X, the horizontal transform type trTypeHor or the vertical transform type trTypeVer can be determined as the first transform indicating an integer transform based on DCT-2.
[0769] At this point, the first transformation can represent that trTypeHor or trTypeVer has a first value. Here, the first value can be 0.
[0770] Here, X can be a positive integer, and can be, for example, 2 or 4.
[0771] When the width or height of a block or sub-block is greater than Y, the horizontal transform type trTypeHor or the vertical transform type trTypeVer can be determined as the first transform indicating an integer transform based on DCT-2.
[0772] Here, Y can be a positive integer, and can be, for example, 16, 32 or 64.
[0773] When the width or height of a block or sub-block is greater than or equal to X or less than or equal to Y, the horizontal transform type trTypeHor or the vertical transform type trTypeVer can be determined as a second transform indicating an integer transform based on DST-7.
[0774] At this point, the second transformation can indicate that trTypeHor or trTypeVer has a second value.
[0775] Here, the second value can be 1. Here, X can be a positive integer, and can be, for example, 2 or 4. Here, Y can be a positive integer, and can be, for example, 16, 32 or 64.
[0776] In addition, when the width or height of a block or sub-block is Z, horizontal or vertical transformations may not be performed.
[0777] Here, Z can be a positive integer including 0, and can be, for example, 1.
[0778] Here, the same horizontal transformation type or the same vertical transformation type can be used to transform all sub-blocks partitioned from the block partition.
[0779] Furthermore, a one-dimensional transformation type can be determined for at least one of the horizontal or vertical transformations of a block or sub-block, regardless of the intra-frame prediction mode. That is, a transformation type can be determined for at least one of the horizontal or vertical transformations of a block or sub-block, regardless of the intra-frame prediction mode. The transformation type can be selected from at least two transformation types.
[0780] When the current block is in sub-block partitioning mode, the transformation type of at least one of the horizontal or vertical transformation can be determined based on the width or height of the current block, regardless of the intra-prediction mode. Here, the sub-block partitioning mode can be either the first sub-block partitioning mode (ISP mode) or the second sub-block partitioning mode (SBT mode).
[0781] As shown in the example in Table 4, at least one of the horizontal transformation type trTypeHor or the vertical transformation type trTypeVer can be determined for a block or sub-block.
[0782] Table 4
[0783] trTypeHor trTypeVer (W >= 4 && W <= 16)? 1 : 0 (H >= 4 && H <= 16)? 1 : 0
[0784] For example, when the current block is in sub-block partitioning mode, the transformation type for at least one of the horizontal or vertical transformations can be determined based on Table 4, regardless of the intra-frame prediction mode.
[0785] Furthermore, to reduce the implementation complexity of the encoder / decoder, the one-dimensional transformation type for horizontal and vertical transformations can be determined on the same standard based on the width and height of the block.
[0786] The conditions used to determine horizontal and vertical transformations can be the same, regardless of the block's width or height. In other words, the conditions used to determine horizontal and vertical transformations can be identical.
[0787] The above conditions can represent a comparison between the width or height of the current block and a specific positive integer.
[0788] Here, the implementation complexity of the encoder / decoder can be reduced because the logic for determining the conditions for horizontal and vertical transformations can be shared.
[0789] Furthermore, as shown in the examples in Table 4, the transformation type for at least one of the horizontal or vertical transformations used in the sub-block can be determined without performing a comparison between the width and height of the current block, thereby reducing computational complexity.
[0790] Compared to Tables 5 and 6 below, Table 4 reduces computational complexity because the transformation type for at least one of the horizontal or vertical transformations used in the sub-block can be determined without performing a comparison between the width and height of the current block.
[0791] mt can be used as a selection index for multiple transformations s The idx performs entropy encoding / decoding to determine the two-dimensional transform combination of a block or sub-block. The horizontal transform type (trTypehor) and vertical transform type (trTypever) can be determined using a pre-defined two-dimensional transform combination table in the encoder and decoder. In this case, the two-dimensional transform combination can represent each entry in the two-dimensional transform combination table.
[0792] When a block is not partitioned into sub-blocks, entropy encoding / decoding can be performed on mts_idx, which serves as the multi-transformation selection index.
[0793] The multitransform selection index mts_idx can be determined based on lfnst_idx (i.e., secondary transform execution information). In other words, the multitransform selection index mts_idx can be entropy encoded / entropy decoded based on lfnst_idx (i.e., secondary transform execution information).
[0794] When the multitransform selection index mts_idx does not exist in the bitstream, the multitransform selection index mts_idx can be inferred to be the first value (e.g., 0). In other words, when the multitransform selection index mts_idx is not entropy encoded / entropy decoded, the multitransform selection index mts_idx can be inferred to be the first value (e.g., 0).
[0795] According to an embodiment, whether the multitransform selection index mts_idx is sent by signaling can be determined based on at least one of the intra-prediction explicit multitransform selection enable information (sps_explicit_mts_intra_enabled_flag) and the inter-prediction explicit multitransform selection enable information (sps_explicit_mts_inter_enabled_flag). For example, when the current block is intra-predicted and the intra-prediction explicit multitransform selection enable information indicates that multitransformation is enabled for the intra-prediction block, the multitransform selection index for the current block can be sent by signaling. Conversely, when the intra-prediction explicit multitransform selection enable information indicates that multitransformation is not enabled for the intra-prediction block, the multitransform selection index for the current block can be sent without signaling.
[0796] Furthermore, when the current block is inter-predicted and the inter-prediction explicit multi-transform selection enable information indicates that multi-transformation is enabled for the inter-prediction block, a multi-transformation selection index for the current block can be sent using a signal. Conversely, when the inter-prediction explicit multi-transformation selection enable information indicates that multi-transformation is not enabled for the inter-prediction block, a multi-transformation selection index for the current block does not need to be sent using a signal.
[0797] Furthermore, when predicting the current block according to IBC mode, the signaling of the multi-transform selection index can be determined based on the intra-frame prediction explicit multi-transform selection enable information and / or the inter-frame prediction explicit multi-transform selection enable information. Additionally, when predicting the current block according to IBC mode, the multi-transform selection index for the current block can be signaled without signaling, regardless of the intra-frame prediction explicit multi-transform selection enable information and the inter-frame prediction explicit multi-transform selection enable information. Here, when predicting the current block according to IBC mode, secondary transform / inverse transform may not be performed for the current block. Furthermore, when the current block is in intra-frame prediction mode instead of IBC mode and / or inter-frame prediction mode, secondary transform / inverse transform may be performed for the current block.
[0798] As shown in the example in Table 5, at least one of the horizontal transformation type trTypeHor or the vertical transformation type trTypeVer can be determined for a block or sub-block. Based on Table 5, a two-dimensional transformation combination table can be predetermined, and the two-dimensional transformation combination indicated by the multi-transformation selection index can be determined.
[0799] Table 5
[0800] mts_idx[x0][y0] 0 1 2 3 4 trTypeHor 0 1 2 1 2 trTypeVer 0 1 1 2 2
[0801] The following illustrates an embodiment for determining the type of one-dimensional transformation, a combination of two-dimensional transformations, or whether to use at least one of the transformations for a block or sub-block. Here, the type of one-dimensional transformation can be determined for at least one of the horizontal or vertical transformations for a block or sub-block.
[0802] When the width W or height H of a block or sub-block is less than X, the horizontal transform type trTypeHor or the vertical transform type trTypeVer can be determined as the first transform indicating an integer transform based on DCT-2.
[0803] At this point, the first transformation can represent that trTypeHor or trTypeVer has a first value. Here, the first value can be 0.
[0804] Here, X can be a positive integer, and can be, for example, 2 or 4.
[0805] When the width or height of a block or sub-block is greater than Y, the horizontal or vertical transformation type can be determined as the first transformation indicating an integer transformation based on DCT-2.
[0806] Here, Y can be a positive integer, and can be, for example, 16, 32 or 64.
[0807] When the width or height of a block or sub-block is greater than or equal to X or less than or equal to Y, the horizontal or vertical transformation type can be determined based on the following conditions.
[0808] When the partition direction information has a first value (0) and the sub-block position information has a first value (0), the horizontal transformation type can be determined as a third transformation indicating an integer transformation based on DCT-8, and the vertical transformation type can be determined as a second transformation indicating an integer transformation based on DST-7.
[0809] When the partition direction information has a first value (0) and the sub-block position information has a second value (1), the horizontal transformation type can be determined as a second transformation indicating an integer transformation based on DST-7, and the vertical transformation type can be determined as a second transformation indicating an integer transformation based on DST-7.
[0810] When the partition direction information has a second value (1) and the sub-block position information has a first value (0), the horizontal transformation type can be determined as a second transformation indicating an integer transformation based on DST-7, and the vertical transformation type can be determined as a third transformation indicating an integer transformation based on DCT-8.
[0811] When the partition direction information has a second value (1) and the sub-block position information has a second value (1), the horizontal transformation type can be determined as a second transformation indicating an integer transformation based on DST-7, and the vertical transformation type can be determined as a second transformation indicating an integer transformation based on DST-7.
[0812] At this point, the second transformation can represent that trTypeHor or trTypeVer has a second value. Here, the second value can be 1.
[0813] At this point, the third transformation can represent that trTypeHor or trTypeVer has a third value. Here, the third value can be 2.
[0814] Here, X can be a positive integer, and can be, for example, 2 or 4. Here, Y can be a positive integer, and can be, for example, 16, 32 or 64.
[0815] In addition, when the width or height of a block or sub-block is Z, horizontal or vertical transformations may not be performed.
[0816] Here, Z can be a positive integer including 0, and can be, for example, 1.
[0817] As shown in the example in Table 6, at least one of the horizontal transformation type trTypeHor or the vertical transformation type trTypeVer can be determined for a block or sub-block.
[0818] Table 6
[0819] outer zone direction information sub-block position information trTypeHor trTypeVer 0 0 2 1 0 1 1 1 1 0 1 2 1 1 1 1
[0820] When the width W or height H of a block or sub-block is less than X, the horizontal transform type trTypeHor or the vertical transform type trTypeVer can be determined as the first transform indicating an integer transform based on DCT-2.
[0821] At this point, the first transformation can represent that trTypeHor or trTypeVer has a first value. Here, the first value can be 0.
[0822] Here, X can be a positive integer, and can be, for example, 2, 4 or 8.
[0823] When the width or height of a block or sub-block is greater than Y, the horizontal or vertical transformation type can be determined as the first transformation indicating an integer transformation based on DCT-2.
[0824] Here, Y can be a positive integer, and can be, for example, 8, 16, 32 or 64.
[0825] When the width or height of a block or sub-block is greater than or equal to X or less than or equal to Y, the horizontal or vertical transformation type can be determined as a second transformation indicating an integer transformation based on DST-7.
[0826] At this point, the second transformation can represent that trTypeHor or trTypeVer has a second value. Here, the second value can be 1.
[0827] Here, X can be a positive integer, and can be, for example, 2, 4, or 8. Here, Y can be a positive integer, and can be, for example, 8, 16, 32, or 64.
[0828] In addition, when the width or height of a block or sub-block is Z, horizontal or vertical transformations may not be performed.
[0829] Here, Z can be a positive integer including 0, and can be, for example, 1.
[0830] Here, the same horizontal transform type or the same vertical transform type can be used to transform all sub-blocks partitioned from the block. Furthermore, a one-dimensional transform type can be determined for at least one of the horizontal or vertical transforms of a block or sub-block, regardless of the intra-prediction mode.
[0831] As shown in the example in Table 7, at least one of the horizontal transformation type trTypeHor or the vertical transformation type trTypeVer can be determined for a block or sub-block.
[0832] Table 7
[0833] trTypeHor trTypeVer (W >= 4 && W <= 8)? 1 : 0 (H >= 4 && H <= 8)? 1 : 0
[0834] The following illustrates an embodiment for determining the type of one-dimensional transformation, a combination of two-dimensional transformations, or whether to use at least one of the transformations for a block or sub-block. Here, the type of one-dimensional transformation can be determined for at least one of the horizontal or vertical transformations for a block or sub-block.
[0835] When the width W or height H of a block or sub-block is less than X, the horizontal transform type trTypeHor or the vertical transform type trTypeVer can be determined as the first transform indicating an integer transform based on DCT-2.
[0836] At this point, the first transformation can represent that trTypeHor or trTypeVer has a first value. Here, the first value can be 0.
[0837] Here, X can be a positive integer, and can be, for example, 2, 4 or 8.
[0838] When the width or height of a block or sub-block is greater than Y, the horizontal or vertical transformation type can be determined as the first transformation indicating an integer transformation based on DCT-2.
[0839] Here, Y can be a positive integer, and can be, for example, 8, 16, 32 or 64.
[0840] When the width or height of a block or sub-block is greater than or equal to X or less than or equal to Y, the horizontal or vertical transformation type can be determined as a second transformation indicating an integer transformation based on DST-7.
[0841] At this point, the second transformation can represent that trTypeHor or trTypeVer has a second value. Here, the second value can be 1.
[0842] Here, X can be a positive integer, and can be, for example, 2, 4, or 8. Here, Y can be a positive integer, and can be, for example, 8, 16, 32, or 64.
[0843] In addition, when the width or height of a block or sub-block is Z, horizontal or vertical transformations may not be performed.
[0844] Here, Z can be a positive integer including 0, and can be, for example, 1.
[0845] Here, the same horizontal transform type or the same vertical transform type can be used to transform all sub-blocks partitioned from the block. Furthermore, a one-dimensional transform type can be determined for at least one of the horizontal or vertical transforms of a block or sub-block, regardless of the intra-prediction mode.
[0846] As shown in the example in Table 8, at least one of the horizontal transformation type trTypeHor or the vertical transformation type trTypeVer can be determined for a block or sub-block.
[0847] Table 8
[0848] trTypeHor trTypeVer (W >= 4 && W <= 32)? 1 : 0 (H >= 4 && H <= 32)? 1 : 0
[0849] The following illustrates an embodiment for determining the type of one-dimensional transformation, a combination of two-dimensional transformations, or whether to use at least one of the transformations for a block or sub-block. Here, the type of one-dimensional transformation can be determined for at least one of the horizontal or vertical transformations for a block or sub-block.
[0850] When the prediction mode of a block or sub-block is an intra-prediction mode or an intra-block copy prediction mode, a one-dimensional transformation type for at least one of the horizontal or vertical transformations of the block or sub-block can be determined based on at least one embodiment in the embodiments.
[0851] Various transformation / inverse transformation type determination methods can be performed for blocks or sub-blocks, and at least one of the following scanning methods can be performed after a transformation or before an inverse transformation.
[0852] At least one of the following scanning methods can be performed on the quantized coefficient level or quantized level that has undergone at least one transformation or quantization in the encoder / decoder.
[0853] Here, the quantization coefficient level can represent the result generated by performing transformation and quantization on the residual block. Furthermore, the quantization level can represent the result generated by performing quantization on the residual block.
[0854] Furthermore, the quantization coefficient level and the quantization level can have the same meaning, and can also have the same meaning as the transform coefficient. That is, the quantization coefficient level, the quantization level, and the transform coefficient can represent the object when the residual block is entropy encoded / entropy decoded.
[0855] like FIG. 14 As shown in the example, a diagonal scan can be used to arrange the quantized coefficient levels in a two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a diagonal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in a two-dimensional residual block.
[0856] The scanning direction from the lower left to the upper right can be called the upper right diagonal scan. Conversely, the scanning direction from the upper right to the lower left can be called the lower left diagonal scan.
[0857] FIG. 14 The example shows the top-right scan in a diagonal scan.
[0858] As in FIG. 15 In the example, a horizontal scan can be used to arrange the quantized coefficient levels in a two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a horizontal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in a two-dimensional residual block.
[0859] In this case, horizontal scanning can be a method that prioritizes scanning the coefficients corresponding to the first row.
[0860] As in FIG. 16 In the example, a vertical scan can be used to arrange the quantized coefficient levels in a two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a vertical scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in a two-dimensional residual block.
[0861] In this case, vertical scanning can be a method that prioritizes scanning the coefficients corresponding to the first column.
[0862] As in FIG. 17 In the example, a block-based diagonal scan can be used to arrange the quantized coefficient levels in a two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a block-based diagonal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in a two-dimensional residual block.
[0863] In this case, the block size can be M×N. Here, at least one of M or N can be a positive integer and can be 4. Furthermore, the block size can be equal to the size of the coefficient set used in transform coefficient encoding / decoding.
[0864] FIG. 17 The example shows a block-based top-right scan in a block-based diagonal scan.
[0865] In this case, a block can represent a sub-block partitioned from a block of a specific size. If block-based scanning is used, even the same scanning method as that used in blocks can be used to scan sub-blocks within a block of a specific size.
[0866] As in FIG. 17 In the example, when using a block-based diagonal scan, after an 8×8 block is divided into sub-blocks of size 4×4, a diagonal scan can be used to scan the sub-blocks of size 4×4, and a diagonal scan can be used to scan the coefficients in the sub-blocks.
[0867] As in FIG. 18 In the example, a block-based horizontal scan can be used to arrange the quantized coefficient levels in a two-dimensional residual block into a one-dimensional coefficient level array. Alternatively, a block-based horizontal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in a two-dimensional residual block. In this case, the block size can be 4×4, and the blocks corresponding to the first row can be scanned first.
[0868] In this case, the block size can be M×N. Here, at least one of M or N can be a positive integer and can be 4. Furthermore, the block size can be equal to the size of the coefficient set used in transform coefficient encoding / decoding.
[0869] At this point, a block-based horizontal scan can be a method that prioritizes scanning the coefficients corresponding to the first row.
[0870] As in FIG. 19 In the example, a block-based vertical scan can be used to arrange the quantized coefficient levels in a two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a block-based vertical scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in a two-dimensional residual block.
[0871] In this case, the block size can be M×N. Here, at least one of M or N can be a positive integer and can be 4. Furthermore, the block size can be equal to the size of the coefficient set used in transform coefficient encoding / decoding.
[0872] At this point, a block-based vertical scan can be a method that prioritizes scanning the coefficients corresponding to the first column.
[0873] As in FIG. 14 to FIG. 19In the examples, the scan corresponding to (a) can be used for a residual block of size J×K for a block of size J×K, and the scan corresponding to (b) can be used for a residual block of size M×N or larger for a block of size at least one of 8×8 / 16×16 / 32×32 / 64×64, or for a residual block of size M×N. J, K, M, and N can be positive integers. Furthermore, J and K can be less than M and N, respectively. Additionally, J×K can be 4×4, and M×N can be 8×8.
[0874] As in FIG. 14 to FIG. 19 In the example, although only the scanning method corresponding to the maximum size of 8×8 is shown, the scanning method corresponding to the size of 8×8 is applicable to the scanning method corresponding to the size of larger than 8×8, and the above scanning is applicable not only to residual blocks with square form, but also to residual blocks with non-square form.
[0875] To arrange the quantized coefficient levels in a two-dimensional residual block with a square / non-square form in the encoder, the quantized coefficient levels in the residual block can be scanned. Furthermore, to arrange the one-dimensional reconstructed coefficient level array in the decoder as the quantized coefficient levels in a two-dimensional residual block with a square / non-square form, the coefficient levels can be scanned.
[0876] As in FIG. 20 In the example, at least one of the quantized coefficient levels can be scanned.
[0877] For example, as in FIG. 20 In example (a), a diagonal scan can be used to arrange the quantized coefficient levels in the two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a diagonal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in the two-dimensional residual block.
[0878] At this moment, as if in FIG. 20 In example (a), the diagonal scan direction can be from the bottom left to the top right or from the top right to the bottom left.
[0879] The scanning direction from the lower left to the upper right can be called the upper right diagonal scan. Conversely, the scanning direction from the upper right to the lower left can be called the lower left diagonal scan.
[0880] FIG. 20 Example (a) shows the top-right scan in a diagonal scan.
[0881] As another example, such as in FIG. 20In example (b), a vertical scan can be used to arrange the quantized coefficient levels in the two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a vertical scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in the two-dimensional residual block.
[0882] In this case, vertical scanning can be a method that prioritizes scanning the coefficients corresponding to the first column.
[0883] As another example, such as in FIG. 20 In example (c), a horizontal scan can be used to arrange the quantized coefficient levels in the two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a horizontal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in the two-dimensional residual block.
[0884] In this case, horizontal scanning can be a method that prioritizes scanning the coefficients corresponding to the first row.
[0885] As another example, such as in FIG. 20 In example (d), a block-based diagonal scan can be used to arrange the quantized coefficient levels in the two-dimensional residual block into a one-dimensional coefficient level array. Furthermore, a block-based diagonal scan can be used to arrange the one-dimensional reconstructed coefficient level array into the quantized coefficient levels in the two-dimensiona...
Claims
1. A video decoding method, the method comprising: Transform skip mode flag, which indicates whether the inverse transform is skipped, is decoded from the bitstream; Determine whether to decode the transform matrix index of the current block from the bitstream; Determine whether to skip the secondary inverse transformation for the current block based on the value of the transformation matrix index; as well as Based on determining whether to skip the secondary inverse transform, and whether to perform or not perform the secondary inverse transform for the current block, the residual samples of the current block are obtained. Specifically, the transform skip mode flag is decoded for each color component, including the luminance component, Cb component, and Cr component. Specifically, when the tree structure of the current block is a dual-tree type and the current block is a chroma component block, it is determined whether to decode the transform matrix index of the current block from the bitstream based on the transform skip mode flag for the Cb component and the transform skip mode flag for the Cr component, regardless of the transform skip mode flag for the luma component. Specifically, the transform matrix index is decoded when both the transform skip mode flag for the Cb component and the transform skip mode flag for the Cr component indicate that the inverse transform is not skipped.
2. The video decoding method as described in claim 1, in, When the tree structure for the current block is a single tree type, based on all the transform skip mode flags for the luminance component, the transform skip mode flag for the Cb component, and the transform skip mode flag for the Cr component, it is determined whether to decode the transform matrix index of the current block from the bitstream.
3. The video decoding method as described in claim 1, in, When the tree structure of the current block is a dual-tree type and the current block is a luminance component block, it is determined whether to decode the transform matrix index of the current block from the bitstream based on the transform skip mode flag for the luminance component, regardless of the transform skip mode flags for the Cb component and the transform skip mode flags for the Cr component.
4. The video decoding method as described in claim 2, in, When the tree structure for the current block is a single tree type, the transform matrix index of the current block is decoded from the bitstream only if all the transform skip mode flags for the luminance component, the transform skip mode flags for the Cb component, and the transform skip mode flags for the Cr component indicate that the inverse transform is not skipped.
5. The video decoding method as described in claim 1, in, When the value of the transformation matrix index is greater than 0, the secondary inverse transformation is applied to the current block, and the transformation matrix used for the secondary inverse transformation is determined according to the value of the transformation matrix index.
6. The video decoding method as described in claim 1, further comprising: Decoding information indicating whether the intra-frame residual DPCM method was applied to the current block. Specifically, when the information indicates that the intra-frame residual DPCM method is not applied to the current block, the transform matrix index of the current block is decoded from the bitstream, and Specifically, when the information indicating the intra-frame residual DPCM method is applied to the current block, the transform skip mode flag for decoding the current block is omitted, and the inverse transform is skipped for the current block.
7. The video decoding method as described in claim 1, in, Further considering whether the current block is predicted by an intra-prediction mode that is not matrix-based, determine whether to decode the transform matrix index of the current block from the bitstream.
8. A video encoding method, the method comprising: Obtain the residual samples of the current block; The transformation coefficients of the current block are obtained by performing a secondary transformation on the residual samples of the current block or by not performing a secondary transformation on the residual samples of the current block. as well as The transform coefficients of the current block are encoded. The transform skip mode flag, which indicates whether a transform is skipped, is encoded into the bitstream. Specifically, when the transform matrix index of the current block is encoded, the value of the transform matrix index is determined based on whether secondary transforms are skipped for the current block. Specifically, the transform skip mode flag is encoded for each color component, including the luminance component, Cb component, and Cr component. Specifically, when the tree structure of the current block is a dual-tree type and the current block is a chroma component block, the transformation matrix index of the current block is determined based on the transform skip mode flag for the Cb component and the transform skip mode flag for the Cr component, regardless of the transform skip flag for the luminance component. Specifically, the transform matrix index is encoded when both the transform skip mode flag for the Cb component and the transform skip mode flag for the Cr component are encoded to indicate that the inverse transform is not skipped.
9. The video encoding method as described in claim 8, in, When the tree structure of the current block is a single tree type, based on all the transform skip mode flags for the luminance component, the transform skip mode flag for the Cb component, and the transform skip mode flag for the Cr component, it is determined whether to encode the transform matrix index of the current block.
10. The video encoding method as described in claim 9, in, When the tree structure of the current block is a single tree type, the transform matrix index of the current block is encoded only if the transform skip mode flag for the luminance component, the transform skip mode flag for the Cb component, and the transform skip mode flag for the Cr component are all encoded to indicate that the inverse transform is not skipped.
11. The video encoding method as described in claim 8, in, When a secondary transformation is applied to the current block, the value of the transformation matrix index is greater than 0, and the value of the transformation matrix index is set to indicate the secondary transformation matrix used for the secondary transformation.
12. The video encoding method as described in claim 8, further comprising: Information indicating whether the intra-frame residual DPCM method is applied to the current block is encoded. Specifically, when the information is encoded as indicating that the intra-frame residual DPCM method is not applied to the current block, the transform skip mode flag of the current block is encoded, and Specifically, when the information is encoded as indicating that the intra-frame residual DPCM method is applied to the current block, the encoding of the transform skip mode flag for the current block is omitted, and the transform is skipped for the current block.
13. The video encoding method as described in claim 8, in, Further considering whether the current block is predicted by an intra-prediction mode that is not matrix-based, it is determined whether the transform matrix index of the current block should be encoded.
14. An apparatus for transmitting compressed video data, comprising: The processor is configured to acquire the compressed video data; as well as The transmitting unit is configured to transmit the compressed video data. Obtaining the compressed video data includes: Obtain the residual samples of the current block; The transformation coefficients of the current block are obtained by performing a secondary transformation on the residual samples of the current block or by not performing a secondary transformation on the residual samples of the current block; and The transform coefficients of the current block are encoded. The transform skip mode flag, which indicates whether a transform is skipped, is encoded into the bitstream. Specifically, when the transform matrix index of the current block is encoded, the value of the transform matrix index is determined based on whether secondary transforms are skipped for the current block. Specifically, the transform skip mode flag is encoded for each color component, including the luminance component, Cb component, and Cr component. Specifically, when the tree structure of the current block is a dual-tree type and the current block is a chroma component block, the transformation matrix index of the current block is determined based on the transform skip mode flag for the Cb component and the transform skip mode flag for the Cr component, regardless of the transform skip flag for the luminance component. Specifically, the transform matrix index is encoded when both the transform skip mode flag for the Cb component and the transform skip mode flag for the Cr component are encoded to indicate that the inverse transform is not skipped.