Image encoding / decoding methods and devices, and recording media for storing bitstreams.
By optimizing image encoding and decoding through a change in skip mode, the problem of increased data volume in high-resolution images is solved, resulting in a more efficient compression and storage solution.
Patent Information
- Application Number
- CN202080070456.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-10
- Filing Date
- 2020-10-12
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-10-12
AI Technical Summary
As the resolution and quality of high-definition and ultra-high-definition images improve, the amount of image data increases, leading to higher transmission and storage costs, and existing video coding technologies are not efficient enough in compression.
A transform skip mode is adopted, and the image encoding and decoding process is optimized by encoding and decoding the chroma residual joint flag and the transform skip mode flag, including the transformation skip mode judgment and processing of luminance and chroma components.
It improves the compression efficiency of image encoding and decoding, and reduces the cost of data transmission and storage.
Smart Images

Figure CN114503566B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a video encoding / decoding method and apparatus. More specifically, this disclosure relates to an image encoding / decoding method and apparatus using transform skipping. Background Technology
[0002] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) or ultra-high-definition (UHD) images, has increased in various applications. As image resolution and quality improve, the amount of data increases accordingly. This is one reason for the increased transmission and storage costs when transmitting or storing image data via existing transmission media such as wired or wireless broadband channels. To address these issues related to high-resolution and high-quality image data, efficient image encoding / decoding technologies are needed.
[0003] Various video compression techniques exist, such as inter-frame prediction techniques that predict pixel values in the current frame from pixel values in previous or subsequent frames, intra-frame prediction techniques that predict pixel values in a region of the current frame from pixel values in another region of the current frame, energy transformation and quantization techniques for compressing residual signals, and entropy coding techniques that assign shorter codes to frequently occurring pixel values and longer codes to less frequently occurring pixel values. Summary of the Invention
[0004] Technical issues
[0005] This disclosure aims to provide an image encoding / decoding method and apparatus with enhanced compression efficiency, as well as a recording medium for storing bitstreams generated by the video encoding / decoding method and apparatus of this disclosure.
[0006] This disclosure also aims to provide an image encoding / decoding method and apparatus for improving compression efficiency by using a transform skip mode, and a recording medium for storing bitstreams.
[0007] Technical solution
[0008] This disclosure provides a video decoding method, comprising: obtaining a chroma residual joint flag indicating whether a chroma residual joint mode is applied to the current block; obtaining a transform skip mode flag for the luminance component and a transform skip mode flag for a first chroma component from a bitstream; when the chroma residual joint flag indicates that the chroma residual joint mode is applied to the current block, determining whether a transform skip mode is applied to a second chroma component of the current block based on the transform skip mode flag for the first chroma component; and when the chroma residual joint flag indicates that the chroma residual joint mode is not applied to the current block, obtaining a transform skip mode flag for the second chroma component of the current block from the bitstream.
[0009] According to an embodiment, the video decoding method may further include: when the chroma residual joint flag indicates that the chroma residual joint mode is applied to the current block, determining the residual sample points of the second chroma component of the current block based on the residual sample points of the first chroma component of the current block.
[0010] According to an embodiment, the size of the residual sample points of the second chromaticity component can be determined based on the size of the residual sample points of the first chromaticity component, and the sign of the residual sample points of the second chromaticity component can be determined to be opposite to the sign of the residual sample points of the first chromaticity component.
[0011] According to an embodiment, the video decoding method may further include: obtaining chroma residual joint symbol information from the bitstream, which indicates the relationship between the symbols of residual samples of a first chroma component and the symbols of residual samples of a second chroma component. The symbols of the residual samples of the second chroma component can be determined based on the chroma residual joint symbol information and the symbols of the residual samples of the first chroma component.
[0012] According to an embodiment, the step of obtaining a transform skip mode flag for the luminance component of the current block and a transform skip mode flag for the first chrominance component may include: obtaining a transform skip mode flag for the luminance component of the current block according to the size of the current block; and obtaining a transform skip mode flag for the first chrominance component of the current block according to the size of the current block.
[0013] According to an embodiment, the step of obtaining a transition skip mode flag for the luminance component of the current block based on the size of the current block may include: obtaining a transition skip mode flag for the luminance component of the current block when the height of the current block is equal to or less than the maximum block height and the width of the current block is equal to or less than the maximum block width.
[0014] According to an embodiment, the step of obtaining a transform skip mode flag for the first chroma component of the current block based on the size of the current block may include: obtaining a transform skip mode flag for the first chroma component of the current block based on the size and color format of the current block.
[0015] According to an embodiment, the video decoding method may further include: obtaining a chroma residual joint enable flag from the bitstream, indicating whether a chroma residual joint mode is enabled for the parent unit of the current block. The step of obtaining the chroma residual joint flag of the current block may include: when the chroma residual joint enable flag indicates that the chroma residual joint mode is enabled for the parent unit of the current block, obtaining the chroma residual joint flag of the current block.
[0016] According to an embodiment, the upper-level unit can be at least one of video, coded video sequence (CVS), frame, sub-frame, strip, parallel block, and coding tree unit.
[0017] According to an embodiment, the first chromaticity component and the second chromaticity component can be Cb component and Cr component, respectively, or Cr component and Cb component, respectively.
[0018] This disclosure provides a video encoding method, comprising: encoding a chroma residual joint flag of the current block indicating whether a chroma residual joint mode is applied to the current block; encoding a transform skip mode flag for the luminance component of the current block and a transform skip mode flag for a first chroma component of the current block; skipping the encoding of a transform skip mode flag for a second chroma component of the current block when the chroma residual joint mode is applied to the current block; and encoding a transform skip mode flag for the second chroma component of the current block when the chroma residual joint mode is not applied to the current block.
[0019] According to an embodiment, the video coding method may further include: determining whether the chroma residual joint mode is applied to the current block based on the residual sample points of the first chroma component of the current block and the residual sample points of the second chroma component of the current block.
[0020] According to an embodiment, when the chroma residual joint mode is applied to the current block, the video coding method may further include: encoding chroma residual joint symbol information indicating the relationship between the symbols of the residual samples of the first chroma component and the residual samples of the second chroma component of the current block, based on the residual samples of the first chroma component and the residual samples of the second chroma component of the current block.
[0021] According to an embodiment, the step of encoding the transform skip mode flag for the luminance component of the current block and the transform skip mode flag for the first chrominance component of the current block may include: encoding the transform skip mode flag for the luminance component of the current block according to the size of the current block; and encoding the transform skip mode flag for the first chrominance component of the current block according to the size of the current block.
[0022] According to an embodiment, the step of encoding the transform skip mode flag for the luminance component of the current block according to the size of the current block may include: encoding the transform skip mode flag for the luminance component of the current block when the height of the current block is equal to or less than the maximum block height and the width of the current block is equal to or less than the maximum block width.
[0023] According to an embodiment, the step of encoding the transform skip mode flag for the first chroma component of the current block according to the size of the current block may include: encoding the transform skip mode flag for the first chroma component of the current block according to the size and color format of the current block.
[0024] According to an embodiment, the video decoding method may further include: encoding a chroma residual joint enable flag that indicates whether a chroma residual joint mode is enabled for the parent unit of the current block. Encoding the chroma residual joint flag of the current block may include: encoding the chroma residual joint flag of the current block when the chroma residual joint enable flag indicates that a chroma residual joint mode is enabled for the parent unit of the current block.
[0025] According to an embodiment, the upper-level unit can be at least one of video, coded video sequence (CVS), frame, sub-frame, strip, parallel block, and coding tree unit.
[0026] According to an embodiment, the first chromaticity component and the second chromaticity component may be Cb component and Cr component, respectively, or Cr component and Cb component, respectively.
[0027] This disclosure provides a computer-readable recording medium for storing a bitstream comprising image data encoded according to a video coding method. The bitstream includes: a chroma residual joint flag indicating whether a chroma residual joint mode is applied to the current block; a transform skip mode flag for the luma component of the current block; and a transform skip mode flag for a first chroma component of the current block. When the chroma residual joint flag indicates that the chroma residual joint mode is not applied to the current block, the bitstream further includes a transform skip mode flag for a second chroma component of the current block. When the chroma residual joint flag indicates that the chroma residual joint mode is applied to the current block, the transform skip mode flag for the second chroma component of the current block is skipped in the bitstream. Whether a transform skip mode is applied to the second chroma component of the current block is determined based on the transform skip mode flag for the first chroma component.
[0028] Beneficial effects
[0029] According to this disclosure, an image encoding / decoding method and apparatus with enhanced compression efficiency, as well as a recording medium for storing bitstreams generated by the video encoding / decoding method and apparatus of this disclosure, can be provided.
[0030] According to this disclosure, an image encoding / decoding method and apparatus for improving compression efficiency by using a transform skip mode, and a recording medium for storing bit streams are also provided. Attached Figure Description
[0031] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0032] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the invention.
[0033] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded.
[0034] Figure 4 This is a diagram illustrating intra-frame prediction processing.
[0035] Figure 5 This is a diagram illustrating an embodiment of inter-screen prediction processing.
[0036] Figure 6 This is a diagram illustrating the transformation and quantization processes.
[0037] Figure 7 This is a diagram showing reference samples that can be used for intra-frame prediction.
[0038] Figure 8 This is a view illustrating an encoding / decoding method according to an embodiment of the present disclosure.
[0039] Figure 9 A video decoding method according to an embodiment of the present disclosure is shown.
[0040] Figure 10 A video encoding method according to an embodiment of the present disclosure is shown. Detailed Implementation
[0041] This disclosure provides a video decoding method, comprising: obtaining a chroma residual joint flag indicating whether a chroma residual joint mode is applied to the current block; obtaining a transform skip mode flag for the luminance component and a transform skip mode flag for a first chroma component from a bitstream; when the chroma residual joint flag indicates that the chroma residual joint mode is applied to the current block, determining whether a transform skip mode is applied to a second chroma component of the current block based on the transform skip mode flag for the first chroma component; and when the chroma residual joint flag indicates that the chroma residual joint mode is not applied to the current block, obtaining a transform skip mode flag for the second chroma component of the current block from the bitstream.
[0042] Various modifications can be made to this invention, and various embodiments of the invention exist, wherein examples of various embodiments of the invention will now be provided and described in detail with reference to the accompanying drawings. However, the invention is not limited thereto, although exemplary embodiments may be interpreted as including all modifications, equivalents, or substitutions within the technical concept and scope of the invention. In various respects, similar reference numerals refer to the same or similar functions. In the drawings, the shape and size of elements may be exaggerated for clarity. In the following detailed description of the invention, reference is made to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice this disclosure. It should be understood that the various embodiments of this disclosure, though different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in conjunction with one embodiment may be implemented in other embodiments without departing from the spirit and scope of this disclosure. Furthermore, it should be understood that the position or arrangement of various elements within each disclosed embodiment may be modified without departing from the spirit and scope of this disclosure. Therefore, the following detailed description should not be considered limiting, and the scope of this disclosure is defined only by the appended claims (which, where properly interpreted, also include the full scope of the equivalents claimed by the claims).
[0043] The terms "first," "second," etc., used in this specification may be used to describe various components, but the components should not be construed as limited to these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the invention, a "first" component may be named a "second" component, and a "second" component may similarly be named a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple items.
[0044] It will be understood that, in this specification, when an element is simply referred to as "connected to" or "coupled to" another element rather than "directly connected to" or "directly coupled to" another element, the element may be "directly connected to" or "directly coupled to" another element, or may be connected to or coupled to another element where there is another element intervening between the element and the other element. Conversely, it should be understood that when an element is referred to as "directly coupled to" or "directly connected" to another element, there is no intermediate element.
[0045] Furthermore, the constituent parts shown in the embodiments of the present invention are illustrated independently to represent different functional characteristics. Therefore, this does not mean that each constituent part is constructed as a separate hardware or software unit. In other words, for convenience, each constituent part includes each of the listed constituent parts. Thus, at least two constituent parts of each constituent part can be combined to form one constituent part, or a constituent part can be divided into multiple constituent parts to perform each function. Embodiments in which each constituent part is combined and embodiments in which each constituent part is divided are also included within the scope of the present invention without departing from its spirit.
[0046] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Unless the context clearly distinguishes them, expressions used in the singular include those used in the plural. In this specification, it will be understood that terms such as “comprising,” “having,” etc., are intended to indicate the presence of features, numbers, steps, actions, elements, components, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, components, or combinations thereof may be present or added. In other words, when a particular element is referred to as “comprising,” it does not exclude elements other than the corresponding element, but rather includes additional elements within the embodiments of the invention or within the scope of the invention.
[0047] Furthermore, some components may not be essential for performing the basic functions of the invention, but rather optional components that only improve its performance. The invention can be implemented by including only the essential components necessary for achieving the essence of the invention, excluding components used to improve performance. Structures that include only the essential components and exclude optional components used only to improve performance are also included within the scope of the invention.
[0048] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In describing exemplary embodiments of the invention, well-known functions or structures will not be described in detail, as they may unnecessarily obscure the understanding of the invention. Like constituent elements in the drawings are indicated by like reference numerals, and repeated descriptions of like elements will be omitted.
[0049] In the following text, an image may refer to a frame that constitutes a video, or it may refer to the video itself. For example, "encoding or decoding an image or both" may refer to "encoding or decoding a moving image or both," and may also refer to "encoding or decoding an image within an image of a moving image or both."
[0050] In the following text, the terms "moving images" and "video" are used to mean the same thing and are interchangeable.
[0051] In the following text, the target image can be an encoded target image that serves as an encoding target and / or a decoded target image that serves as a decoding target. Furthermore, the target image can be an input image input to an encoding device and an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
[0052] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used as having the same meaning and may be used interchangeably with each other.
[0053] In the following text, a target block can be an encoded target block that serves as the encoding target and / or a decoded target block that serves as the decoding target. Furthermore, a target block can be the current block that serves as the target of the current encoding and / or decoding. For example, the terms "target block" and "current block" can be used to mean the same thing and can be used interchangeably.
[0054] In the following text, the terms "block" and "unit" may be used to mean the same thing and may be used interchangeably. Alternatively, "block" may refer to a specific unit.
[0055] In the following text, the terms “region” and “fragment” are used interchangeably.
[0056] In the following text, a specific signal can be a signal representing a specific block. For example, the original signal can be a signal representing the target block. The prediction signal can be a signal representing the prediction block. The residual signal can be a signal representing the residual block.
[0057] In this embodiment, each of the following can have a value: information, data, flags, indexes, elements, and attributes. A value of "0" for information, data, flags, indexes, elements, and attributes can represent logical false or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value can be interchanged. A value of "1" for information, data, flags, indexes, elements, and attributes can represent logical true or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value can be interchanged.
[0058] When variables i or j are used to represent columns, rows, or indices, the value of i can be an integer equal to or greater than 0, or an integer equal to or greater than 1. That is, columns, rows, indices, etc., can be counted starting from 0, or they can be counted starting from 1.
[0059] Description of terms
[0060] Encoder: Represents the device that performs encoding. In other words, it refers to the encoding device.
[0061] Decoder: Represents the device that performs decoding. In other words, it refers to the decoding device.
[0062] A block is an M×N sample array. Here, M and N can represent positive integers, and a block can represent a two-dimensional sample array. A block can refer to a unit. The current block can represent a coding target block that becomes the target during encoding, or a decoding target block that becomes the target during decoding. Furthermore, the current block can be at least one of a coding block, a prediction block, a residual block, and a transform block.
[0063] Samples: These are the basic units that make up a block. Based on bit depth (B... d Sample points can be represented as numbers from 0 to 2. Bd The value is -1. In this invention, a sample point can be used to represent a pixel. That is, a sample point, a pel, and a pixel can have the same meaning.
[0064] Unit: Refers to an encoding and decoding unit. When encoding and decoding an image, a unit can be a region generated by partitioning a single image. Furthermore, a unit can represent a sub-partitioning unit when a single image is partitioned into sub-partitioning units during encoding or decoding. That is, an image can be partitioned into multiple units. When encoding and decoding an image, predetermined processing can be performed for each unit. A single unit can be partitioned into sub-units smaller than the unit's size. Depending on the function, a unit can represent a block, macroblock, coding tree unit, coding tree block, coding unit, coding block, prediction unit, prediction block, residual unit, residual block, transform unit, transform block, etc. Furthermore, to distinguish a unit from a block, a unit can include a luma component block, a chroma component block associated with the luma component block, and syntax elements for each chroma component block. Units can have various sizes and shapes; specifically, the shape of a unit can be a two-dimensional geometric shape, such as a square, rectangle, trapezoid, triangle, pentagon, etc. In addition, the cell information may include at least one of the following: cell type indicating coding cell, prediction cell, transform cell, etc., cell size, cell depth, and the order of encoding and decoding of the cell.
[0065] A coding tree unit is a single coding tree block configured with the luminance component Y and two coding tree blocks associated with the chrominance components Cb and Cr. Furthermore, a coding tree unit can represent a block and the syntax elements of each block. Each coding tree unit can be partitioned using at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning methods to configure lower-level units such as coding units, prediction units, transform units, etc. A coding tree unit can be used as a term to specify a sample block that becomes a processing unit when encoding / decoding an image as an input image. Here, a quadtree can represent a quaternion tree.
[0066] When the size of the coded block is within a predetermined range, it is possible to partition using only quadtree partitioning. Here, the predetermined range can be defined as at least one of the maximum and minimum sizes of the coded block that can be partitioned using only quadtree partitioning. Information indicating the maximum / minimum size of the coded block that allows quadtree partitioning can be transmitted via a signal in the bitstream, and this information can be transmitted via a signal in at least one unit of sequence, frame parameters, parallel block groups, or stripes (fragments). Optionally, the maximum / minimum size of the coded block can be a predetermined fixed size in the encoder / decoder. For example, when the size of the coded block corresponds to 256×256 to 64×64, it is possible to partition using only quadtree partitioning. Optionally, when the size of the coded block is greater than the size of the maximum transform block, it is possible to partition using only quadtree partitioning. Here, the block to be partitioned can be at least one of a coded block and a transform block. In this case, the information indicating the partitioning of the coded block (e.g., split_flag) can be a flag indicating whether quadtree partitioning is performed. When the size of the coded block falls within a predetermined range, it is possible to partition using only binary or ternary tree partitioning. In this case, the above description of quadtree partitioning can be applied in the same way to binary tree partitioning or ternary tree partitioning.
[0067] Encoding block: Can be used as a term to specify any one of the Y encoding block, Cb encoding block, and Cr encoding block.
[0068] Neighboring blocks: These can represent blocks adjacent to the current block. A block adjacent to the current block can be a block that touches the boundary of the current block, or a block located within a predetermined distance from the current block. A neighboring block can also represent a block adjacent to a vertex of the current block. Here, a block adjacent to a vertex of the current block can be a block that is vertically adjacent to a block horizontally adjacent to the current block, or a block that is horizontally adjacent to a block vertically adjacent to the current block.
[0069] Reconstructing neighboring blocks: This can represent neighboring blocks that are adjacent to the current block and have already been spatially / temporally encoded or decoded. Here, reconstructing neighboring blocks can represent reconstructing neighboring units. Reconstructing spatial neighboring blocks can be blocks within the current frame that have already been reconstructed through encoding or decoding, or both. Reconstructing temporal neighboring blocks are blocks within a reference image that are located at the position corresponding to the current block in the current frame, or are neighboring blocks of said block.
[0070] Cell depth: Represents the degree of cell partitioning. In a tree structure, the highest node (root node) corresponds to the first unpartitioned cell. Furthermore, the highest node can have the minimum depth value. In this case, the depth of the highest node can be level 0. A node with a depth of level 1 represents a cell generated by partitioning the first cell once. A node with a depth of level 2 represents a cell generated by partitioning the first cell twice. A node with a depth of level n represents a cell generated by partitioning the first cell n times. Leaf nodes can be the lowest-level nodes and cannot be further partitioned. The depth of a leaf node can be the highest-level. For example, the predefined value of the highest-level can be 3. The root node can have the lowest depth, and the leaf nodes can have the deepest depth. Additionally, when cells are represented as a tree structure, the level in which the cell exists can represent the cell depth.
[0071] Bitstream: A bitstream that can represent encoded image information.
[0072] Parameter set: Corresponds to the header information in the configuration within the bitstream. At least one of the video parameter set, sequence parameter set, frame parameter set, and adaptive parameter set may be included in the parameter set. Furthermore, the parameter set may include slice headers, tile group headers, and tile header information. The term "tile group" refers to a set of parallel tiles and has the same meaning as a slice.
[0073] An adaptive parameter set can represent a set of parameters that can be shared by referencing different frames, subframes, stripes, parallel block groups, parallel blocks, or bricks. Furthermore, information in an adaptive parameter set can be used by referencing different adaptive parameter sets for subframes, stripes, parallel block groups, parallel blocks, or bricks within a frame.
[0074] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets within a frame, such as sub-frames, stripes, parallel block groups, parallel blocks, or chunks.
[0075] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets within a sub-picture, such as stripes, parallel block groups, parallel blocks, or chunks.
[0076] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets for parallel blocks or segments within a strip.
[0077] Furthermore, regarding adaptive parameter sets, different adaptive parameter sets can be referenced by using identifiers for different adaptive parameter sets for blocks within a parallel block.
[0078] Information about the adaptive parameter set identifier can be included in the header or parameter set of the sub-screen, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the sub-screen.
[0079] Information about the adaptive parameter set identifier can be included in the header or parameter set of the parallel block, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used in the parallel block.
[0080] Information about the adaptive parameter set identifier can be included in the header of the block, and the adaptive parameter set corresponding to the adaptive parameter set identifier can be used for the block.
[0081] The screen can be divided into one or more parallel block rows and one or more parallel block columns.
[0082] A sub-screen can be divided into one or more parallel block rows and one or more parallel block columns within the screen. A sub-screen can be a rectangular / square region within the screen and may include one or more CTUs. In addition, at least one or more parallel blocks / strips / strips may be included within a sub-screen.
[0083] A parallel block can be a rectangular / square region within the frame and may include one or more CTUs. Furthermore, a parallel block may be divided into one or more sub-blocks.
[0084] A block can represent one or more CTU lines within a parallel block. A parallel block can be partitioned into one or more blocks, and each block can have at least one or more CTU lines. A parallel block that is not partitioned into two or more blocks can be represented as a block.
[0085] A strip may include one or more parallel blocks within a frame, and may include one or more sub-blocks within a parallel block.
[0086] Explanation: This could mean determining the value of a syntax element by performing entropy decoding, or it could mean entropy decoding itself.
[0087] Symbols: can represent at least one of the syntax elements, encoding parameters, and transform coefficient values of the encoding / decoding target unit. Additionally, symbols can represent entropy encoding targets or entropy decoding results.
[0088] Prediction mode: This can be information indicating the mode that is encoded / decoded using intra-frame prediction or the mode that is encoded / decoded using inter-frame prediction.
[0089] Prediction Unit: A basic unit that can be represented when performing prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation). A single prediction unit can be partitioned into multiple partitions with smaller sizes, or it can be partitioned into multiple lower-level prediction units. Multiple partitions can be the basic units when performing prediction or compensation. Partitions generated by dividing prediction units can also be prediction units.
[0090] Prediction cell partitioning: can represent the shape obtained by partitioning prediction cells.
[0091] A reference frame list can refer to a list of one or more reference frames used for inter-frame prediction or motion compensation. There are several types of available reference frame lists, including LC (list combination), L0 (list 0), L1 (list 1), L2 (list 2), and L3 (list 3).
[0092] The inter-frame prediction indicator can indicate the direction of inter-frame prediction for the current block (unidirectional prediction, bidirectional prediction, etc.). Optionally, the inter-frame prediction indicator can indicate the number of reference frames used to generate the prediction blocks for the current block. Optionally, the inter-frame prediction indicator can indicate the number of prediction blocks used when performing inter-frame prediction or motion compensation on the current block.
[0093] The prediction list utilization flag indicates whether at least one reference frame from a specific reference frame list is used to generate a prediction block. The prediction list utilization flag can be used to derive an inter-frame prediction indicator, and conversely, the inter-frame prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag has a first value of zero (0), it indicates that a reference frame from the reference frame list is not used to generate a prediction block. On the other hand, when the prediction list utilization flag has a second value of one (1), it indicates that the reference frame list is used to generate a prediction block.
[0094] The reference screen index can refer to the index of a specific reference screen in the reference screen list.
[0095] A reference frame can refer to a frame referenced by a specific block for the purpose of inter-frame prediction or motion compensation for that specific block. Alternatively, a reference frame can be a frame that includes a reference block referenced by the current block for inter-frame prediction or motion compensation. In the following text, the terms "reference frame" and "reference image" have the same meaning and are interchangeable.
[0096] Motion vectors can be two-dimensional vectors used for inter-frame prediction or motion compensation. A motion vector can represent the offset between the encoded / decoded target block and the reference block. For example, (mvX, mvY) can represent a motion vector. Here, mvX can represent the horizontal component, and mvY can represent the vertical component.
[0097] The search range can be a two-dimensional region searched during inter-frame prediction to retrieve motion vectors. For example, the size of the search range can be M×N. Here, M and N are both integers.
[0098] Motion vector candidates can refer to a block of prediction candidates or the motion vectors within a block of prediction candidates when making predictions about motion vectors. Furthermore, motion vector candidates can be included in a list of motion vector candidates.
[0099] A motion vector candidate list can represent a list consisting of one or more motion vector candidates.
[0100] A motion vector candidate index can represent an indicator that points to a motion vector candidate in the motion vector candidate list. Alternatively, it can be an index of a motion vector predictor.
[0101] Motion information may include at least one of the following: motion vector, reference frame index, inter-frame prediction indicator, prediction list utilization flag, reference frame list information, reference frame, motion vector candidate, motion vector candidate index, merge candidate, and merge index.
[0102] A merge candidate list can represent a list consisting of one or more merge candidates.
[0103] Merge candidates can be spatial merge candidates, temporal merge candidates, combined merge candidates, combined double prediction merge candidates, or zero merge candidates. Merge candidates may include motion information such as inter-frame prediction indicators, reference frame indices for each list, motion vectors, prediction list utilization flags, and inter-frame prediction indicators.
[0104] The merge index can represent an indicator pointing to a merge candidate in the merge candidate list. Optionally, the merge index can indicate a block in a reconstructed block that is spatially / temporally adjacent to the current block, from which a merge candidate has been derived. Optionally, the merge index can indicate at least one piece of motion information for a merge candidate.
[0105] Transform unit: This can represent the basic unit used when encoding / decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding / decoding) a residual signal. A single transform unit can be partitioned into multiple lower-level transform units with smaller sizes. Here, the transform / inverse transform may include at least one of a first transform / first inverse transform and a second transform / second inverse transform.
[0106] Scaling: This refers to the process of multiplying the quantization level by a factor. Transform coefficients can be generated by scaling the quantization level. Scaling can also be called inverse quantization.
[0107] Quantization parameters: These represent values used when transform coefficients are used to generate quantization levels during quantization. Quantization parameters can also represent values used when transform coefficients are generated during dequantization by scaling the quantization levels. Quantization parameters can be values mapped to the quantization step size.
[0108] Incremental quantization parameter: can represent the difference between the predicted quantization parameter and the quantization parameter of the encoding / decoding target unit.
[0109] Scan: This can refer to a method of sorting coefficients within a cell, block, or matrix. For example, changing a two-dimensional matrix of coefficients into a one-dimensional matrix can be called a scan, and changing a one-dimensional matrix of coefficients into a two-dimensional matrix can be called a scan or inverse scan.
[0110] Transform coefficients: These represent the coefficient values generated after a transform is performed in the encoder. Transform coefficients can also represent the coefficient values generated after at least one of entropy decoding and dequantization is performed in the decoder. The quantization level obtained by quantizing the transform coefficients or residual signal, or the quantized transform coefficient level, can also fall within the meaning of transform coefficients.
[0111] Quantization level: This can represent the value generated in the encoder by quantizing the transform coefficients or residual signal. Optionally, the quantization level can represent the value of the dequantization target that undergoes dequantization in the decoder. Similarly, the transform coefficient level, as a result of transform and quantization, can also fall within the meaning of quantization level.
[0112] Non-zero transform coefficients: can represent transform coefficients with values other than zero, or transform coefficient levels or quantization levels with values other than zero.
[0113] Quantization matrix: A matrix used in quantization or dequantization processes performed to improve subjective or objective image quality. The quantization matrix can also be referred to as a scaling list.
[0114] Quantization matrix coefficients: These represent each element within the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.
[0115] Default matrix: can represent a predefined quantization matrix in the encoder or decoder.
[0116] Non-default matrix: can represent a quantization matrix that is not predefined in the encoder or decoder but is sent by the user via signal.
[0117] Statistical value: For at least one of the following variables, coding parameters, constant values, etc., that have a computable specific value, a statistical value can be one or more of the following: mean, summation, weighted average, weighted sum, minimum, maximum, most frequent value, median, interpolation.
[0118] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0119] Encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include at least one image. Encoding device 100 may encode at least one image sequentially.
[0120] Reference Figure 1 The encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.
[0121] Encoding device 100 can encode the input image using intra-frame mode, inter-frame mode, or both. Furthermore, encoding device 100 can generate a bitstream including encoding information by encoding the input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium. When intra-frame mode is used as the prediction mode, switcher 115 can switch to intra-frame mode. Optionally, when inter-frame mode is used as the prediction mode, switcher 115 can switch to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. Encoding device 100 can generate prediction blocks for input blocks of the input image. Furthermore, encoding device 100 can encode residual blocks using the residual between the input block and the prediction block after generating the prediction blocks. The input image can be referred to as the current image as the current encoding target. The input block can be referred to as the current block as the current encoding target, or as the encoding target block.
[0122] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from blocks that have been encoded / decoded and are adjacent to the current block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the current block using the reference samples, or generate prediction samples for the input block by performing spatial prediction. Here, intra-frame prediction can refer to prediction within a frame.
[0123] When the prediction mode is inter-frame mode, the motion prediction unit 111 can retrieve the region that best matches the input block from the reference image during motion prediction and derive the motion vector using the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference frame buffer 190. Here, the reference image can be stored in the reference frame buffer 190 when encoding / decoding the reference image is performed.
[0124] The motion compensation unit 112 can generate a prediction block by performing motion compensation on the current block using motion vectors. Here, inter-frame prediction can refer to prediction or motion compensation between frames.
[0125] When the value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 can generate prediction blocks by applying an interpolation filter to a portion of the reference frame. To perform inter-frame prediction or motion compensation on the coding unit, it can be determined which of the following modes—skip mode, merge mode, Advanced Motion Vector Prediction (AMVP) mode, and current frame reference mode—is used for motion prediction and motion compensation on the prediction unit included in the corresponding coding unit. Then, depending on the determined mode, inter-frame prediction or motion compensation can be performed differently.
[0126] Subtractor 125 generates a residual block by using the difference between the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Furthermore, the residual signal can be a signal generated by transforming or quantizing, or transforming and quantizing, the difference between the original signal and the prediction signal. The residual block can be the residual signal of a block cell.
[0127] Transform unit 130 can generate transform coefficients by performing a transform on the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, transform unit 130 can skip the transform on the residual block.
[0128] The level of quantization can be generated by applying quantization to the transform coefficients or to the residual signal. In the following examples, the level of quantization may also be referred to as the transform coefficients.
[0129] The quantization unit 140 can generate a quantization level by quantizing the transform coefficients or residual signal according to parameters, and output the generated quantization level. Here, the quantization unit 140 can quantize the transform coefficients using a quantization matrix.
[0130] Entropy coding unit 150 can generate a bitstream by performing entropy coding on the values calculated by quantization unit 140 according to a probability distribution or on the coding parameter values calculated during encoding, and output the generated bitstream. Entropy coding unit 150 can perform entropy coding on sample information of the image and information used for decoding the image. For example, the information used for decoding the image may include syntax elements.
[0131] When entropy coding is applied, symbols are represented such that fewer bits are allocated to symbols with high generation probability and more bits are allocated to symbols with low generation probability, thus reducing the size of the bitstream used to encode the symbols. The entropy coding unit 150 can use coding methods for entropy coding such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code (VLC) table. Furthermore, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / bits, and perform arithmetic coding by using the derived binarization method and context model.
[0132] In order to encode the transform coefficient levels (quantization levels), the entropy coding unit 150 can change the coefficients in two-dimensional block form into one-dimensional vector form by using a transform coefficient scanning method.
[0133] Encoding parameters may include information such as syntax elements (flags, indexes, etc.) encoded in the encoder and signaled to the decoder, as well as information derived during encoding or decoding. Encoding parameters can represent the information required when encoding or decoding an image. For example, at least one value or combination of the following may be included in the encoding parameters: cell / block size, cell / block depth, cell / block partitioning information, cell / block shape, cell / block partitioning structure, whether quadtree partitioning is performed, whether binary tree partitioning is performed, binary tree partitioning direction (horizontal or vertical), binary tree partitioning type (symmetric or asymmetric), whether the current encoded unit is partitioned via ternary tree partitioning, the direction of the ternary tree partitioning (horizontal or vertical), the type of ternary tree partitioning (symmetric or asymmetric), whether the current encoded unit is partitioned via multi-type tree partitioning, and the type of multi-type tree partitioning. Direction (horizontal or vertical), type of multi-type tree partition (symmetric or asymmetric), tree structure of multi-type tree partition (binary or ternary), prediction mode (intra-frame prediction or inter-frame prediction), luma intra-frame prediction mode / direction, chroma intra-frame prediction mode / direction, intra-frame partition information, inter-frame partition information, coded block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, prediction block filtering method, prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, intra-frame prediction mode Inter-frame prediction mode, motion information, motion vector, motion vector difference, reference frame index, inter-frame prediction angle, inter-frame prediction indicator, prediction list utilization flag, reference frame list, reference frame, motion vector predictor index, motion vector predictor candidate, motion vector candidate list, whether to use merge mode, merge index, merge candidate, merge candidate list, whether to use skip mode, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector magnitude, motion vector representation accuracy, transform type, transform size, information on whether the primary (first) transform is used, information on whether the secondary transform is used, primary transform index, secondary transform index Information on the presence of residual signals, code block style, code block flag (CBF), quantization parameters, quantization parameter residuals, quantization matrix, whether an intra-loop filter is applied, intra-loop filter coefficients, intra-loop filter taps, intra-loop filter shape / form, whether a deblocking filter is applied, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset type, adaptive sample offset function, whether an adaptive intra-loop filter is applied, adaptive intra-loop filter coefficients, adaptive intra-loop filter taps, adaptive intra-loop filter shape / form.Binarization / debinarization method, context model determination method, context model update method, whether to execute normal mode, whether to execute bypass mode, context binary bits, bypass binary bits, valid coefficient flag, last valid coefficient flag, encoding flag for the unit of the coefficient group, position of the last valid coefficient, flag indicating whether the coefficient value is greater than 1, flag indicating whether the coefficient value is greater than 2, flag indicating whether the coefficient value is greater than 3, information about the remaining coefficient values, symbol information, reconstructed luminance samples, reconstructed chrominance samples, residual luminance samples, residual chrominance samples, luminance transformation coefficient, chrominance transformation coefficient, quantized luminance level, quantized chrominance level, transformation coefficient level scanning method, motion vector search area on the decoder side. The data includes the domain size, the shape of the motion vector search area on the decoder side, the number of motion vector searches on the decoder side, information about the CTU size, information about the minimum block size, information about the maximum block size, information about the maximum block depth, information about the minimum block depth, image display / output order, stripe identification information, stripe type, stripe partition information, parallel block identification information, parallel block type, parallel block partition information, parallel block group identification information, parallel block group type, parallel block group partition information, image type, bit depth of input samples, bit depth of reconstructed samples, bit depth of residual samples, bit depth of transform coefficients, bit depth of quantization levels, and information about the luminance signal or the chrominance signal.
[0134] Here, sending a flag or index with a signal can indicate that the encoder entropy-encodes the corresponding flag or index and includes it in the bitstream, and can also indicate that the decoder entropy-decodes the corresponding flag or index from the bitstream.
[0135] When the encoding device 100 performs encoding via inter-frame prediction, the encoded current image can be used as a reference image for another image to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded current image, or store the reconstructed or decoded image as a reference image in the reference frame buffer 190.
[0136] The quantization level can be dequantized in dequantization unit 160 or inverse transformed in inverse transform unit 170. The coefficients that have undergone dequantization or inverse transform, or both, can be added to the prediction block by adder 175. A reconstruction block can be generated by adding the coefficients that have undergone dequantization or inverse transform, or both, to the prediction block. Here, the coefficients that have undergone dequantization or inverse transform, or both, can represent coefficients for which at least one of dequantization and inverse transform has been performed, and can represent the reconstructed residual block.
[0137] The reconstructed block can be passed through filter unit 180. Filter unit 180 can apply at least one of deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF) to the reconstructed sample, reconstructed block, or reconstructed image. Filter unit 180 may be referred to as an in-loop filter.
[0138] Deblocking filters remove block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, the number of samples included in several rows or columns within the block can be used. When a deblocking filter is applied to a block, another filter can be applied based on the desired deblocking intensity.
[0139] To compensate for coding errors, a suitable offset value can be added to the sample value using a sample-adaptive offset. The sample-adaptive offset corrects the offset between the deblocked image and the original image on a sample-by-sample basis. This can be achieved by considering edge information about each sample point when applying the offset, or by dividing the image's samples into a predetermined number of regions, determining the regions where the offset will be applied, and then applying the offset to those regions.
[0140] The adaptive in-loop filter (ALF) performs filtering based on a comparison between the filtered reconstructed image and the original image. Samples included in the image can be partitioned into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information regarding whether to apply the ALF can be transmitted via a signal through the coding unit (CU), and the form and coefficients of the ALF to be applied to each block can vary.
[0141] The reconstructed blocks or reconstructed image that have passed through filter unit 180 can be stored in reference frame buffer 190. The reconstructed blocks processed by filter unit 180 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filter unit 180. The stored reference image can be used later in inter-frame prediction or motion compensation.
[0142] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment and to which the present invention is applied.
[0143] Decoding device 200 can be a decoder, video decoding device, or image decoding device.
[0144] Reference Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference frame buffer 270.
[0145] Decoding device 200 can receive bitstreams output from encoding device 100. Decoding device 200 can receive bitstreams stored on a computer-readable recording medium, or bitstreams streamed via wired / wireless transmission media. Decoding device 200 can decode the bitstreams using intra-frame mode or inter-frame mode. Furthermore, decoding device 200 can generate and output reconstructed or decoded images generated through decoding.
[0146] When the prediction mode used during decoding is intra-frame mode, the switcher can be switched to intra-frame mode. Optionally, when the prediction mode used during decoding is inter-frame mode, the switcher can be switched to inter-frame mode.
[0147] Decoding device 200 obtains a reconstructed residual block and generates a prediction block by decoding the input bitstream. Once the reconstructed residual block and prediction block are obtained, decoding device 200 generates a reconstructed block that becomes the decoding target by adding the reconstructed residual block and the prediction block. The decoding target block can be referred to as the current block.
[0148] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols in quantized hierarchical form. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0149] In order to decode the transform coefficient levels (quantization levels), the entropy decoding unit 210 can change the coefficients in unidirectional vector form into two-dimensional block form by using a transform coefficient scanning method.
[0150] The quantization level can be dequantized in the dequantization unit 220, or the quantization level can be inversely transformed in the inverse transform unit 230. The quantization level can be the result of dequantization, inverse transform, or both, and can be generated as a reconstruction residual block. Here, the dequantization unit 220 can apply the quantization matrix to the quantization level.
[0151] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, wherein the spatial prediction uses sample values of blocks that are adjacent to the target block and have already been decoded.
[0152] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, wherein the motion compensation uses motion vectors and a reference image stored in reference frame buffer 270.
[0153] Adder 255 generates a reconstructed block by adding the reconstructed residual block to the prediction block. Filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive in-loop filter to the reconstructed block or reconstructed image. Filter unit 260 can output a reconstructed image. The reconstructed block or reconstructed image can be stored in a reference frame buffer 270 and used when performing inter-frame prediction. The reconstructed block processed by filter unit 260 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filter unit 260. The stored reference image can be used later in inter-frame prediction or motion compensation.
[0154] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded. Figure 3 An example of dividing a single cell into multiple sub-cells is illustrated schematically.
[0155] To effectively partition an image, coding units (CUs) are used during encoding and decoding. A coding unit can serve as the basic unit when encoding / decoding an image. Furthermore, a coding unit can be used to distinguish between intra-frame prediction modes and inter-frame prediction modes during image encoding / decoding. A coding unit can be the basic unit used for prediction, transform, quantization, inverse transform, inverse quantization, or encoding / decoding processing of transform coefficients.
[0156] Reference Figure 3 Image 300 is partitioned sequentially according to the Largest Coding Unit (LCU), and the LCU unit is determined as the partitioning structure. Here, LCU can be used with the same meaning as Coding Tree Unit (CTU). Unit partitioning can represent partitioning of the block associated with that unit. The block partitioning information may include information about the unit depth. The depth information may represent the number or degree to which the unit is partitioned, or both the number and degree to which the unit is partitioned. A single unit can be partitioned into multiple lower-level units hierarchically associated with the depth information based on a tree structure. In other words, the unit and the lower-level units generated by partitioning the unit may correspond to a node and the child nodes of that node, respectively. Each of the partitioned lower-level units may have depth information. The depth information may be information representing the size of the CU and may be stored in each CU. The unit depth represents the number and / or degree associated with partitioning the unit. Therefore, the partitioning information of the lower-level units may include information about the size of the lower-level units.
[0157] The partitioning structure represents the distribution of coding units (CUs) within the LCU 310. This distribution can be determined by whether a single CU is partitioned into multiple CUs (including positive integers equal to or greater than 2, such as 2, 4, 8, 16, etc.). The horizontal and vertical dimensions of the CUs generated by partitioning can be half the horizontal and vertical dimensions of the CUs before partitioning, respectively, or they can have dimensions smaller than the horizontal and vertical dimensions before partitioning, depending on the number of partitions performed. CUs can be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the height and width of the CU after partitioning can be reduced compared to at least one of the height and width of the CU before partitioning. CU partitioning can be performed recursively until a predefined depth or a predefined size is reached. For example, the depth of the LCU can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, as mentioned above, the LCU can be a coding unit with the maximum coding unit size, and the SCU can be a coding unit with the minimum coding unit size. Partitioning begins at LCU 310. The CU depth increases by 1 when the horizontal or vertical dimension of a CU, or both, are reduced through partitioning. For example, for each depth, the size of an unpartitioned CU can be 2N×2N. Furthermore, in the case of partitioned CUs, a CU of size 2N×2N can be partitioned into four CUs of size N×N. As the depth increases by 1, the size of N can be halved.
[0158] Furthermore, partition information of a CU can be used to indicate whether a CU is partitioned. Partition information can be 1 bit. All CUs except SCUs can include partition information. For example, when the partition information value is the first value, the CU may not be partitioned; when the partition information value is the second value, the CU may be partitioned.
[0159] Reference Figure 3 An LCU with depth 0 can be a 64×64 block. 0 can be the minimum depth. An SCU with depth 3 can be an 8×8 block. 3 can be the maximum depth. CUs with 32×32 blocks and 16×16 blocks can be represented as depth 1 and depth 2, respectively.
[0160] For example, when a single coding unit is partitioned into four coding units, the horizontal and vertical dimensions of the four partitioned coding units can be half the size of the CU before partitioning. In one embodiment, when a 32×32 coding unit is partitioned into four coding units, each of the four partitioned coding units can have a size of 16×16. When a single coding unit is partitioned into four coding units, the coding unit can be said to be partitioned into a quadtree form.
[0161] For example, when a coding unit is partitioned into two sub-coding units, the horizontal or vertical dimension (width or height) of each of the two sub-coding units can be half the horizontal or vertical dimension of the original coding unit. For example, when a coding unit of size 32×32 is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16×32. For example, when a coding unit of size 8×32 is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 8×16. When a coding unit is partitioned into two sub-coding units, it can be said that the coding unit is binary partitioned or partitioned according to a binary tree partitioning structure.
[0162] For example, when a coding unit is divided into three sub-coding units, the horizontal or vertical dimensions of the coding unit can be divided in a 1:2:1 ratio, resulting in three sub-coding units with a horizontal or vertical dimension ratio of 1:2:1. For instance, when a 16×32 coding unit is horizontally divided into three sub-coding units, these three sub-coding units, in order from the top to the bottom, can have dimensions of 16×8, 16×16, and 16×8, respectively. Similarly, when a 32×32 coding unit is vertically divided into three sub-coding units, these three sub-coding units, in order from the left to the right, can have dimensions of 8×32, 16×32, and 8×32, respectively. When a coding unit is divided into three sub-coding units, it can be said that the coding unit is tri-partitioned or partitioned according to a ternary tree partitioning structure.
[0163] exist Figure 3 In the example, the coding tree unit (CTU) 320 is an example of a CTU in which quadtree partitioning, binary tree partitioning, and ternary tree partitioning structures are all applied.
[0164] As described above, to partition the CTU, at least one of a quadtree partitioning structure, a binary tree partitioning structure, and a ternary tree partitioning structure can be applied. Various tree partitioning structures can be applied sequentially to the CTU according to a predetermined priority order. For example, a quadtree partitioning structure can be preferentially applied to the CTU. Encoding units that cannot be further partitioned using a quadtree partitioning structure can correspond to leaf nodes of a quadtree. Encoding units corresponding to leaf nodes of a quadtree can be used as root nodes of binary and / or ternary tree partitioning structures. That is, encoding units corresponding to leaf nodes of a quadtree can be further partitioned according to a binary or ternary tree partitioning structure, or they can be left unpartitioned. Therefore, by preventing encoding units obtained from binary or ternary tree partitioning of encoding units corresponding to leaf nodes of a quadtree from undergoing further quadtree partitioning, block partitioning operations and / or the operation of signaling partitioning information can be effectively performed.
[0165] The fact that a coding unit corresponding to a node in a quadtree is partitioned can be signaled using four-partition information. Four-partition information with a first value (e.g., "1") indicates that the current coding unit is partitioned according to the quadtree partitioning structure. Four-partition information with a second value (e.g., "0") indicates that the current coding unit is not partitioned according to the quadtree partitioning structure. The four-partition information can be a flag with a predetermined length (e.g., one bit).
[0166] There may be no priority between binary tree partitions and ternary tree partitions. That is, the coding unit corresponding to the leaf node of the quadtree can further undergo any partition in either binary tree or ternary tree partitions. Furthermore, the coding unit generated by binary tree partitions or ternary tree partitions may undergo further binary tree partitions or further ternary tree partitions, or it may not be further partitioned.
[0167] A tree structure in which there is no priority between binary tree partitions and ternary tree partitions is called a multi-type tree structure. The coding unit corresponding to the leaf node of a quadtree can be used as the root node of a multi-type tree. At least one of multi-type tree partition indication information, partition direction information, and partition tree information can be used to signal whether to partition the coding unit corresponding to a node in the multi-type tree. To partition the coding unit corresponding to a node in the multi-type tree, the multi-type tree partition indication information, partition direction information, and partition tree information can be signaled sequentially.
[0168] A multi-type tree partitioning indication with a first value (e.g., "1") indicates that the current coding unit will undergo a multi-type tree partition. A multi-type tree partitioning indication with a second value (e.g., "0") indicates that the current coding unit will not undergo a multi-type tree partition.
[0169] When the coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partitioning structure, the coding unit may include partitioning direction information. The partitioning direction information may indicate in which direction the current coding unit will be partitioned for the multi-type tree partition. Partitioning direction information with a first value (e.g., "1") may indicate that the current coding unit will be vertically partitioned. Partitioning direction information with a second value (e.g., "0") may indicate that the current coding unit will be horizontally partitioned.
[0170] When the coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partitioning structure, the current coding unit may include partitioning tree information. The partitioning tree information may indicate the tree partitioning structure that will be used to partition the nodes of the multi-type tree. Partitioning tree information with a first value (e.g., "1") may indicate that the current coding unit will be partitioned according to a binary tree partitioning structure. Partitioning tree information with a second value (e.g., "0") may indicate that the current coding unit will be partitioned according to a ternary tree partitioning structure.
[0171] The partition indication information, partition tree information, and partition direction information can all be flags with a predetermined length (e.g., one bit).
[0172] At least one of the following—quadtree partitioning indication information, multi-type tree partitioning indication information, partitioning direction information, and partitioning tree information—can be entropy encoded / decoded. To entropy encode / decode those types of information, information about neighboring coding units adjacent to the current coding unit can be used. For example, it is highly likely that the partitioning type (partitioned or unpartitioned, partitioning tree, and / or partitioning direction) of the left-hand neighboring coding unit and / or above-hand neighboring coding unit is similar to the partitioning type of the current coding unit. Therefore, contextual information for entropy encoding / decoding of information about the current coding unit can be derived from the information about neighboring coding units. Information about neighboring coding units may include at least one of the following: quadtree partitioning information, multi-type tree partitioning indication information, partitioning direction information, and partitioning tree information.
[0173] As another example, in binary tree partitioning and ternary tree partitioning, binary tree partitioning can be performed first. That is, the current coding unit can first undergo binary tree partitioning, and then the coding unit corresponding to the leaf node of the binary tree can be set as the root node for ternary tree partitioning. In this case, for the coding unit corresponding to the node of the ternary tree, neither quadtree partitioning nor binary tree partitioning can be performed.
[0174] Encoding units that cannot be partitioned according to quadtree, binary tree, and / or ternary tree partitioning structures become the basic units for encoding, prediction, and / or transformation. In other words, these encoding units cannot be further partitioned for prediction and / or transformation. Therefore, partitioning structure information and partitioning information for dividing encoding units into prediction and / or transformation units may not exist in the bitstream.
[0175] However, when the size of the coding unit (i.e., the basic unit used for partitioning) is larger than the size of the maximum transform block, the coding unit can be partitioned recursively until the size of the coding unit is reduced to be equal to or smaller than the size of the maximum transform block. For example, when the size of the coding unit is 64×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into four 32×32 blocks for transformation. For example, when the size of the coding unit is 32×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into two 32×32 blocks for transformation. In this case, the partitioning of the coding unit for transformation is not sent separately with a signal, and the partitioning of the coding unit for transformation can be determined by comparing the horizontal or vertical dimensions of the coding unit with the horizontal or vertical dimensions of the maximum transform block. For example, when the horizontal dimension (width) of the coding unit is greater than the horizontal dimension (width) of the maximum transform block, the coding unit can be vertically bisected. For example, when the vertical dimension (height) of the coding unit is greater than the vertical dimension (height) of the maximum transform block, the coding unit can be horizontally bisected.
[0176] Information regarding the maximum and / or minimum size of the coding unit and the maximum and / or minimum size of the transform block can be transmitted or determined by signals at the higher level of the coding unit. The higher level can be, for example, a sequence level, a frame level, a stripe level, a parallel block group level, a parallel block level, etc. For example, the minimum size of the coding unit can be determined to be 4×4. For example, the maximum size of the transform block can be determined to be 64×64. For example, the minimum size of the transform block can be determined to be 4×4.
[0177] Information regarding the minimum size of the coding unit corresponding to the leaf node of the quadtree (minimum size of the quadtree) and / or the maximum depth from the root node to the leaf node of the multi-type tree (maximum depth of the multi-type tree) can be signaled or determined at the higher level of the coding unit. For example, the higher level can be the sequence level, frame level, stripe level, parallel block group level, parallel block level, etc. Information regarding the minimum size of the quadtree and / or the maximum depth of the multi-type tree can be signaled or determined for each of the intra-frame stripes and inter-frame stripes.
[0178] The difference between the size of the CTU and the maximum size of the transform block can be signaled or determined at the higher level of the coding unit. For example, the higher level can be a sequence level, frame level, stripe level, parallel block group level, parallel block level, etc. The maximum size of the coding unit corresponding to each node of the binary tree (hereinafter referred to as the maximum size of the binary tree) can be determined based on the size of the coding tree unit and the difference information. The maximum size of the coding unit corresponding to each node of the ternary tree (hereinafter referred to as the maximum size of the ternary tree) can vary depending on the type of stripe. For example, for an intra-frame stripe, the maximum size of the ternary tree can be 32×32. For example, for an inter-frame stripe, the maximum size of the ternary tree can be 128×128. For example, the minimum size of the coding unit corresponding to each node of the binary tree (hereinafter referred to as the minimum size of the binary tree) and / or the minimum size of the coding unit corresponding to each node of the ternary tree (hereinafter referred to as the minimum size of the ternary tree) can be set as the minimum size of the coding block.
[0179] As another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the stripe level. Alternatively, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the stripe level.
[0180] Based on the size and depth information of the various blocks mentioned above, four-partition information, multi-type tree partition indication information, partition tree information and / or partition direction information may or may not be included in the bitstream.
[0181] For example, when the size of the coding unit is no greater than the minimum size of the quadtree, the coding unit does not include four-partition information. The four-partition information can be inferred as a second value.
[0182] For example, when the size (horizontal and vertical dimensions) of the coding unit corresponding to a node of a multi-type tree is greater than the maximum size (horizontal and vertical dimensions) of a binary tree and / or a ternary tree, the coding unit may not be divided into two or three partitions. Therefore, multi-type tree partition indication information can be sent without a signal, but the multi-type tree partition indication information can be inferred as a second value.
[0183] Optionally, when the size (horizontal and vertical dimensions) of the coding unit corresponding to a node of a multi-type tree is the same as the maximum size (horizontal and vertical dimensions) of a binary tree and / or twice the maximum size (horizontal and vertical dimensions) of a ternary tree, the coding unit may not be further divided into two or three partitions. Therefore, multi-type tree partitioning indication information does not need to be sent via signal, but the multi-type tree partitioning indication information can be derived from the second value. This is because when the coding unit is partitioned according to the binary tree partitioning structure and / or the ternary tree partitioning structure, coding units smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree are generated.
[0184] Optionally, binary or ternary partitioning can be limited based on the size of the virtual pipeline data unit (hereinafter, the pipeline buffer size). For example, when a coding unit is divided into sub-coding units that are not suitable for the pipeline buffer size through binary or ternary partitioning, the corresponding binary or ternary partitioning may be limited. The pipeline buffer size can be the size of the largest transform block (e.g., 64×64). For example, when the pipeline buffer size is 64×64, the following partitioning can be limited.
[0185] - N×M (N and / or M are 128) ternary tree partitions for encoding units
[0186] - 128×N (N<=64) binary tree partitioning for the horizontal direction of the encoding unit
[0187] - N×128 (N<=64) binary tree partitions used in the vertical direction of the encoding unit
[0188] Optionally, when the depth of the coding unit corresponding to a node in the multi-type tree is equal to the maximum depth of the multi-type tree, the coding unit may not be further divided into two and / or three partitions. Therefore, multi-type tree partition indication information may not be sent by signal, but the multi-type tree partition indication information can be inferred as a second value.
[0189] Optionally, multi-type tree partitioning indication information may be signaled only if at least one of the vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is possible for the coding unit corresponding to the node of the multi-type tree. Otherwise, the coding unit may not be partitioned into two and / or three partitions. Therefore, multi-type tree partitioning indication information may not be signaled, but the multi-type tree partitioning indication information may be inferred as a second value.
[0190] Optionally, partition direction information may be signaled only if both vertical binary tree partitioning and horizontal binary tree partitioning, or both vertical ternary tree partitioning and horizontal ternary tree partitioning, are possible for the coding units corresponding to nodes of multiple tree types. Otherwise, partition direction information may not be signaled, but the partition direction information may be derived from values indicating possible partition directions.
[0191] Optionally, partition tree information may be signaled only if both vertical binary tree partitions and vertical ternary tree partitions, or both horizontal binary tree partitions and horizontal ternary tree partitions, are possible for the encoded tree corresponding to the nodes of the multi-type tree. Otherwise, partition tree information may not be signaled, but the partition tree information may be inferred as a value indicating a possible partition tree structure.
[0192] Figure 4 This is a diagram illustrating intra-frame prediction processing.
[0193] Figure 4 The arrows from the center outwards indicate the prediction direction of the intra-frame prediction mode.
[0194] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the current block. A neighboring block can be a reconstructed neighboring block. For example, intra-frame coding and / or decoding can be performed using values of reference samples or coding parameters included in the reconstructed neighboring block.
[0195] A prediction block can represent a block generated by performing intra-frame prediction. A prediction block can correspond to at least one of CU, PU, and TU. The cells of a prediction block can have the size of one of CU, PU, and TU. A prediction block can be a square block with dimensions such as 2×2, 4×4, 16×16, 32×32, or 64×64, or a rectangular block with dimensions such as 2×8, 4×8, 2×16, 4×16, and 8×16.
[0196] Intra-prediction can be performed based on the intra-prediction mode for the current block. The number of intra-prediction modes that the current block can have can be a fixed value, or it can be a value determined differently depending on the attributes of the predicted block. For example, the attributes of the predicted block can include the size and shape of the predicted block.
[0197] Regardless of the block size, the number of intra-prediction modes can be fixed at N. Alternatively, the number of intra-prediction modes can be 3, 5, 9, 17, 34, 35, 36, 65, or 67, etc. Optionally, the number of intra-prediction modes can vary depending on the block size or the color component type, or both. For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chrominance signal. For example, the number of intra-prediction modes can increase as the block size increases. Optionally, the number of intra-prediction modes for the luma component block can be greater than the number of intra-prediction modes for the chrominance component block.
[0198] Intra-prediction modes can be non-angular or angular. Non-angular modes can be DC or planar modes, and angular modes can be prediction modes with a specific direction or angle. Intra-prediction modes can be represented by at least one of mode number, mode value, mode number, mode angle, and mode direction. The number of intra-prediction modes can be greater than 1M, including both non-angular and angular modes. To perform intra-prediction on the current block, a step can be performed to determine whether a sample included in the reconstructed neighboring block can be used as a reference sample for the current block. When there are samples that cannot be used as reference samples for the current block, the value obtained by copying or interpolating at least one sample value included in the reconstructed neighboring block, or both, can be used to replace the unavailable sample value, and the replaced sample value is used as the reference sample for the current block.
[0199] Figure 7 This is a diagram showing reference samples that can be used for intra-frame prediction.
[0200] like Figure 7 As shown, at least one of reference sample lines 0 to 3 can be used for intra-frame prediction of the current block. Figure 7 In this process, samples from fragments A and F can be filled using samples from the nearest fragments B and E, respectively, instead of being retrieved from reconstructed neighboring blocks. Index information of the reference sample lines to be used for intra-frame prediction of the current block can be transmitted using signals. For example, in... Figure 7 In this configuration, reference sample line indicators 0, 1, and 2 can be signaled as index information indicating reference sample line 0, reference sample line 1, and reference sample line 2. When the upper boundary of the current block is the boundary of the CTU, only reference sample line 0 can be available. Therefore, in this case, index information does not need to be signaled. When reference sample lines other than reference sample line 0 are used, filtering for the prediction block, which will be described later, is not required.
[0201] When performing intra-frame prediction, filters can be applied to at least one of the reference samples and the prediction samples based on the intra-frame prediction mode and the current block size.
[0202] In planar mode, when generating the prediction block for the current block, the sample value of the target sample is generated by using a weighted sum of the upper and left reference samples, and the upper right and lower left reference samples of the current block, based on the position of the target sample within the prediction block. Furthermore, in DC mode, the average of the upper and left reference samples of the current block can be used when generating the prediction block. Additionally, in angled mode, the prediction block can be generated using the upper, left, upper right, and / or lower left reference samples of the current block. Real-valued interpolation can be performed to generate the prediction sample values.
[0203] In the case of intra-frame prediction between color components, a predicted block for the current block of the second color component can be generated based on the corresponding reconstructed block of the first color component. For example, the first color component can be a luma component, and the second color component can be a chroma component. For intra-frame prediction between color components, the parameters of a linear model between the first and second color components can be derived based on a template. The template may include the upper and / or left neighboring samples of the current block and the upper and / or left neighboring samples of the reconstructed block of the corresponding first color component. For example, the parameters of the linear model can be derived using the sample value of the first color component with the maximum value in the template and its corresponding sample value of the second color component, and the sample value of the first color component with the minimum value in the template and its corresponding sample value of the second color component. When deriving the parameters of the linear model, the corresponding reconstructed block can be applied to the linear model to generate a predicted block for the current block. Depending on the video format, subsampling can be performed on the neighboring samples and the corresponding reconstructed block of the reconstructed block of the first color component. For example, when a sample point of the second color component corresponds to four samples of the first color component, the four samples of the first color component can be subsampled to calculate a corresponding sample point. In this case, parameter derivation of the linear model and intra-frame prediction between color components can be performed based on the corresponding subsampled sample points. Whether to perform intra-frame prediction between color components and / or the range of the template can be sent as an intra-frame prediction mode signal.
[0204] The current block can be partitioned into two or four sub-blocks, either horizontally or vertically. The partitioned sub-blocks can be reconstructed sequentially. That is, intra-prediction can be performed on the sub-blocks to generate sub-prediction blocks. Furthermore, inverse quantization and / or inverse transform can be performed on the sub-blocks to generate sub-residual blocks. Reconstructed sub-blocks can be generated by adding the sub-prediction blocks to the sub-residual blocks. The reconstructed sub-blocks can be used as reference samples for intra-prediction of subsequent sub-blocks. A sub-block can be a block comprising a predetermined number (e.g., 16) or more samples. Therefore, for example, when the current block is an 8×4 or 4×8 block, the current block can be partitioned into two sub-blocks. Furthermore, when the current block is a 4×4 block, the current block may not be partitioned into sub-blocks. When the current block has other sizes, the current block can be partitioned into four sub-blocks. Information regarding whether intra-prediction is performed based on sub-blocks and / or partitioning direction (horizontal or vertical) can be signaled. Intra-prediction based on sub-blocks can be limited to being performed only when using reference sample line 0. When performing sub-block-based intra-frame prediction, filtering for the prediction block, which will be described later, may not be performed.
[0205] A final prediction block can be generated by performing filtering on the prediction block that has been intra-predicted. Filtering can be performed by applying predetermined weights to the target sample, the left reference sample, the top reference sample, and / or the top-left reference sample. The weights and / or reference samples (range, position, etc.) used for filtering can be determined based on at least one of the block size, the intra-prediction mode, and the position of the target sample in the prediction block. Filtering can be performed only in a predetermined intra-prediction mode (e.g., DC, planar, vertical, horizontal, diagonal, and / or adjacent diagonal mode). An adjacent diagonal mode can be a mode that is diagonal mode plus k or subtracted from diagonal mode. For example, k can be a positive integer of 8 or less.
[0206] The intra-prediction mode of the current block can be entropy-coded / decoded by predicting the intra-prediction modes of adjacent blocks. When the intra-prediction modes of the current block and its neighboring blocks are the same, information indicating that the intra-prediction modes of the current block and its neighboring blocks are the same can be signaled using predetermined flag information. Furthermore, an indicator of the intra-prediction mode among multiple neighboring blocks that is the same as the intra-prediction mode of the current block can be signaled. When the intra-prediction modes of the current block and its neighboring blocks are different, the intra-prediction mode information of the current block can be entropy-coded / decoded by performing entropy coding / decoding based on the intra-prediction modes of neighboring blocks.
[0207] Figure 5 This is a diagram illustrating an embodiment of inter-screen prediction processing.
[0208] exist Figure 5 In this context, rectangles can represent the image. Figure 5In the image, the arrow indicates the prediction direction. Based on the encoding type of the frame, frames can be classified into intra-frame frames (I-frames), predictive frames (P-frames), and dual-predictive frames (B-frames).
[0209] I-frames can be encoded via intra-frame prediction without requiring inter-frame prediction. P-frames can be encoded via inter-frame prediction using reference frames present in one direction (i.e., forward or backward) relative to the current block. B-frames can be encoded via inter-frame prediction using reference frames present in both directions (i.e., forward and backward) relative to the current block. When using inter-frame prediction, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform the corresponding motion compensation.
[0210] The following section will describe in detail an embodiment of inter-screen prediction.
[0211] Reference frames and motion information can be used to perform inter-frame prediction or motion compensation.
[0212] Motion information of the current block can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction. The motion information of the current block can be derived using motion information of reconstructed neighboring blocks, motion information of co-located blocks (also called col blocks or co-position blocks), and / or motion information of blocks adjacent to the co-position block. A co-position block can represent a block within a previously reconstructed co-located frame (also called a col frame or co-position frame) that is spatially located at the same position as the current block. A co-position frame can be one of one or more reference frames included in a list of reference frames.
[0213] The methods for deriving motion information can vary depending on the prediction mode of the current block. For example, prediction modes applied to inter-frame prediction include AMVP mode, merge mode, skip mode, merge mode with motion vector difference, sub-block merge mode, geometric partitioning mode, combined inter-frame-intra-frame prediction mode, affine mode, etc. Here, the merge mode can be referred to as motion merge mode.
[0214] For example, when AMVP is used as a prediction mode, at least one of the motion vectors of reconstructed neighboring blocks, co-located blocks, blocks adjacent to co-located blocks, and (0,0) motion vectors can be identified as motion vector candidates for the current block, and a motion vector candidate list is generated using these motion vector candidates. Motion vector candidates for the current block can be derived using the generated motion vector candidate list. Motion information for the current block can be determined based on the derived motion vector candidates. The motion vectors of co-located blocks or blocks adjacent to co-located blocks can be referred to as temporal motion vector candidates, and the motion vectors of reconstructed neighboring blocks can be referred to as spatial motion vector candidates.
[0215] Encoding device 100 can calculate the motion vector difference (MVD) between the motion vector of the current block and motion vector candidates, and can perform entropy encoding on the motion vector difference (MVD). Furthermore, encoding device 100 can perform entropy encoding on the motion vector candidate index and generate a bitstream. The motion vector candidate index indicates the best motion vector candidate among the motion vector candidates included in the motion vector candidate list. Decoding device 200 can perform entropy decoding on the motion vector candidate index included in the bitstream, and can select motion vector candidates for the target block to be decoded from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index. Furthermore, decoding device 200 can add the entropy-decoded MVD to the motion vector candidate extracted by entropy decoding, thereby deriving the motion vector of the target block to be decoded.
[0216] Additionally, the encoding device 100 can perform entropy encoding on the calculated MVD resolution information. The decoding device 200 can use the MVD resolution information to adjust the resolution of the entropy-decoded MVD.
[0217] Additionally, the encoding device 100 calculates the motion vector difference (MVD) between the motion vectors and motion vector candidates in the current block based on an affine model, and performs entropy encoding on the MVD. The decoding device 200 derives motion vectors based on each sub-block by deriving the affine control motion vector of the decoded target block based on the sum of the entropied MVD and the affine control motion vector candidates.
[0218] The bitstream may include a reference frame index indicating a reference frame. The reference frame index may be entropy encoded by the encoding device 100 and subsequently transmitted as a bitstream to the decoding device 200. The decoding device 200 may generate a predicted block for the decoded target block based on the derived motion vectors and the reference frame index information.
[0219] Another example of a method for deriving motion information for the current block could be a merge pattern. A merge pattern can represent a method for merging the motion of multiple blocks. A merge pattern can represent a pattern for deriving motion information for the current block from the motion information of neighboring blocks. When applying a merge pattern, a list of merge candidates can be generated using reconstructed motion information of neighboring blocks and / or motion information of blocks at the same location. Motion information may include at least one of motion vectors, reference frame indices, and inter-frame prediction indicators. The prediction indicators may indicate unidirectional prediction (L0 prediction or L1 prediction) or bidirectional prediction (L0 prediction and L1 prediction).
[0220] The merge candidate list can be a list of stored motion information. The motion information included in the merge candidate list can be at least one of the following: motion information of neighboring blocks adjacent to the current block (spatial merge candidate), motion information of blocks at the same position in the reference frame of the current block (temporal merge candidate), new motion information generated by combining motion information existing in the merge candidate list, motion information of blocks encoded / decoded before the current block (history-based merge candidate), and zero merge candidate.
[0221] Encoding device 100 can generate a bitstream by performing entropy encoding on at least one of a merge flag and a merge index, and can transmit the bitstream as a signal to decoding device 200. The merge flag may be information indicating whether a merge mode is performed for each block, and the merge index may be information indicating which neighboring block among the current block's neighboring blocks is the target block for merging. For example, the neighboring blocks of the current block may include a left neighboring block located to the left of the current block, an upper neighboring block arranged above the current block, and a time neighboring block that is temporally adjacent to the current block.
[0222] Additionally, the encoding device 100 performs entropy encoding on the correction information for correcting motion vectors in the motion information of the merging candidates and sends it as a signal to the decoding device 200. The decoding device 200 can correct the motion vectors of the merging candidates selected by the merging index based on the correction information. Here, the correction information may include at least one of information on whether correction is performed, correction direction information, and correction size information. As described above, the prediction mode that corrects the motion vectors of the merging candidates based on the correction information sent as a signal can be referred to as a merging mode with motion vector difference.
[0223] Skip mode can be a mode in which motion information of neighboring blocks is applied to the current block as is. When skip mode is applied, encoding device 100 can perform entropy encoding on information about which block's motion information will be used as the current block's motion information to generate a bitstream, and can send the bitstream to decoding device 200 as a signal. Encoding device 100 may not send syntax elements regarding at least one of motion vector difference information, coded block flags, and transform coefficient levels to decoding device 200 as a signal.
[0224] Sub-block merging patterns can represent a pattern for deriving motion information on a sub-block basis within a coded block (CU). When applying a sub-block merging pattern, a list of sub-block merging candidates can be generated using motion information of sub-blocks at the same position as the current sub-block in a reference image (based on sub-block temporal merging candidates) and / or affine control point motion vector merging candidates.
[0225] A geometric partitioning pattern can be represented as follows: motion information is derived by partitioning the current block in a predetermined direction, each of the derived motion information is used to derive each prediction sample, and the prediction sample of the current block is derived by weighting each of the derived prediction samples.
[0226] The inter-frame-intra-frame combined prediction mode can be represented as a mode for deriving the prediction samples of the current block by weighting the prediction samples generated by inter-frame prediction and the prediction samples generated by intra-frame prediction.
[0227] The decoding device 200 can self-correct the derived motion information. The decoding device 200 can search a predetermined region based on a reference block indicated by the derived motion information and derive motion information with minimum SAD as corrected motion information.
[0228] Decoding device 200 can use optical flow to compensate for prediction samples derived via inter-frame prediction.
[0229] Figure 6 This is a diagram illustrating the transformation and quantization processes.
[0230] like Figure 6 As shown, transform and / or quantization are performed on the residual signal to generate a quantized level signal. The residual signal is the difference between the original block and the predicted block (i.e., an intra-frame predicted block or an inter-frame predicted block). The predicted block is generated through intra-frame prediction or inter-frame prediction. The transform can be a primary transform, a secondary transform, or both. The primary transform of the residual signal generates transform coefficients, and the secondary transform of the transform coefficients generates secondary transform coefficients.
[0231] At least one scheme selected from a variety of predefined transform schemes is used to perform the primary transform. Examples of the predefined transform schemes include the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and Karhunen-Loève Transform (KLT). The transform coefficients generated by the primary transform may undergo secondary transforms. The transform scheme used for the primary and / or secondary transforms can be determined based on the coding parameters of the current block and / or its neighboring blocks. Optionally, transform information indicating the transform scheme can be transmitted via a signal. DCT-based transforms may include, for example, DCT-2, DCT-8, etc. DST-based transforms may include, for example, DST-7.
[0232] A quantized level signal (quantization coefficients) can be generated by performing quantization on the residual signal or on the result of performing a primary transform and / or a secondary transform. Depending on the intra-frame prediction mode or block size / shape, the quantized level signal can be scanned using at least one of diagonal top-right scan, vertical scan, and horizontal scan. For example, when scanning coefficients according to a diagonal top-right scan, the block-form coefficients change to a one-dimensional vector form. In addition to the diagonal top-right scan, depending on the intra-frame prediction mode and / or the size of the transform block, a horizontal scan that horizontally scans the coefficients in two-dimensional block form or a vertical scan that vertically scans the coefficients in two-dimensional block form can be used. The scanned quantized level coefficients can be entropy-encoded for insertion into the bitstream.
[0233] The decoder performs entropy decoding on the bitstream to obtain quantized level coefficients. The quantized level coefficients can be arranged in a two-dimensional block format via inverse scanning. For inverse scanning, at least one of diagonal top-right scanning, vertical scanning, and horizontal scanning can be used.
[0234] The quantized level coefficients can then be dequantized, then subjected to a secondary inverse transform as needed, and finally subjected to a primary inverse transform as needed to generate the reconstructed residual signal.
[0235] Inverse mapping in the dynamic range can be performed on the luma components reconstructed via intra-frame or inter-frame prediction prior to intra-loop filtering. The dynamic range can be divided into 16 equal segments, and a mapping function for each segment can be signaled. The mapping function can be signaled at the strip level or parallel block group level. The inverse mapping function used to perform the inverse mapping can be derived based on the mapping function. Intra-loop filtering, reference frame storage, and motion compensation are performed in the inverse mapping region, and the prediction blocks generated via inter-frame prediction are transformed to the mapping region via mapping using the mapping function and then used to generate reconstructed blocks. However, since intra-frame prediction is performed in the mapping region, the prediction blocks generated via intra-frame prediction can be used to generate reconstructed blocks without mapping / inverse mapping.
[0236] When the current block is a residual block for the chroma component, the residual block can be converted to the inverse-mapped region by scaling the chroma components of the mapped region. The availability of scaling can be signaled at the stripe level or the parallel block group level. Scaling can only be applied if mapping for the luma component is available and the partitioning of the luma component and the partitioning of the chroma component follow the same tree structure. Scaling can be performed based on the average of the sample values of the luma prediction block corresponding to the chroma block. In this case, when inter-frame prediction is used for the current block, the luma prediction block can represent the mapped luma prediction block. The required scaling value can be derived by using an index-referenced lookup table of the segment to which the average of the sample values of the luma prediction block belongs. Finally, by scaling the residual block using the derived value, the residual block can be converted to the inverse-mapped region. Chroma component block recovery, intra-frame prediction, inter-frame prediction, intra-loop filtering, and reference frame storage can then be performed in the inverse-mapped region.
[0237] Information indicating whether the mapping / inverse mapping of the luminance and chrominance components is available can be sent via a sequence parameter set using signals.
[0238] A predicted block for the current block can be generated based on a block vector indicating the displacement between the current block and a reference block in the current frame. In this way, the prediction mode used to generate a predicted block with reference to the current frame is called Intra-Block Copy (IBC) mode. IBC mode can be applied to M×N (M<=64, N<=64) coding units. IBC modes can include skip mode, merge mode, AMVP mode, etc. In the case of skip mode or merge mode, a merge candidate list is constructed, and a merge index is signaled so that a merge candidate can be specified. The block vector of the specified merge candidate can be used as the block vector of the current block. The merge candidate list can include at least one of spatial candidates, history-based candidates, candidates based on the average of two candidates, and zero merge candidates. In the case of AMVP mode, a difference block vector can be signaled. Furthermore, the predicted block vector can be derived from the left neighboring block and the upper neighboring block of the current block. The index of the neighboring block to be used can be signaled. The predicted block in IBC mode is included in the current CTU or the left CTU and is limited to blocks in the already reconstructed area. For example, the value of the block vector can be restricted such that the predicted block of the current block is located in the region of the three 64×64 blocks preceding the 64×64 block to which the current block belongs, in the encoding / decoding order. By restricting the value of the block vector in this way, memory consumption and device complexity can be reduced in implementations according to the IBC mode.
[0239] In the following, an image encoding / decoding method according to an embodiment of the present disclosure will be described.
[0240] When encoding / decoding an image, block transformations are typically performed. However, in some cases, it may be advantageous not to perform block transformations. Here, not performing block transformations is called transform skipping (TS). Additionally, the encoded information indicating whether to skip a transformation for the current block to be encoded (or decoded) is called the transform skip flag. The transform skip flag is sent from the encoder to the decoder by being marked in the bitstream. Furthermore, the decoder parses the transform skip flag in the bitstream to derive its value. When the transform skip flag indicates 0, the block transformation is performed. When the transform skip flag indicates 1, the block transformation is skipped.
[0241] Figure 8 This illustrates a method for constructing a bitstream that sends only the transform skip flag information to the luminance channel.
[0242] Sending a transform skip flag to every YcbCr channel may be detrimental to coding efficiency. Therefore, transform skipping can be applied to the luma channel but not to the chroma channel. Thus, the transform skip flag can be sent only to the luma channel.
[0243] Here, the spatial coordinates (x0, y0) indicate the upper-left spatial position of the current encoded (decoded) block, and the corresponding transform skip flag for the current luma block is indicated as transform_skip_flag[x0][y0]. According to... Figure 8 The method for constructing the bitstream in the code involves the encoder only sending or parsing the transform_skip_flag for the luminance channel. Additionally, according to... Figure 8 The method for constructing a bitstream in the code involves the decoder parsing only the transform skip flags for the luminance channel when parsing transform skip information from the bitstream.
[0244] according to Figure 8 In this embodiment, because a transform skip flag for the chroma channel is not sent, the transform is always performed even when the transform skip mode is more advantageous for the chroma block, which is a problem. Furthermore, transform skipping may be advantageous for both chroma and luma blocks when performing lossless encoding or when creating screen content or other images by computer.
[0245] Typically, image channels exhibit considerable correlation. Therefore, this disclosure provides a method and apparatus for determining whether to skip a transform for a chroma channel using minimal information obtained by sharing or referencing transform skip flag information between multiple channels. According to this disclosure, problems of compression and image quality degradation can be resolved when transform skip flags are sent and parsed for only one or two color channels, rather than for each color channel, and the remaining channels reference the transform skip flags.
[0246] Although a transform skip flag is described below in this disclosure, this technique is not exclusively applicable to transform skip flags. According to this disclosure, the method of sharing or referencing information between color channels can be applied not only to transform skip flags, but also to first transform selection or second transform selection (multiple transform selection, MTS) information, intra-frame smoothing filtering information, position-related intra-prediction combination (PDPC) flag information, information required for deblocking filtering, information required for adaptive loop filtering (ALF), information required for sample adaptive offset (SAO) filtering, information for image quality improvement, and information for improving compression efficiency. However, for ease of description, transforms are used as specific examples to describe the image encoding / decoding operations according to this disclosure in detail. This description can also be applied to other examples of the encoding information listed above.
[0247] To share encoded information among channels in embodiments of this disclosure, one or more predetermined channels may be selected as representative channels among the color channels. Additionally, the encoded information of the representative channels may be shared or referenced in the remaining channels. In other words, the decoder can parse only the encoded information from the representative channels of the encoder via the bitstream. The decoder can parse the encoded information of the remaining channels by referencing or sharing the encoded information of the representative channels without separate signaling.
[0248] For example, when the luminance channel with the highest correlation to the luminance signal in the YCbCr color space is selected as the representative channel, since there is a correlation between the luminance (Y) channel and the chrominance (Cb and / or Cr) channels in the image, the chrominance channel's coded information can be sent and parsed by referring to the luminance channel's coded information, instead of sending (for encoding) and parsing (for decoding) the coded information required to encode and decode the image in each color channel.
[0249] According to this disclosure, a detailed method and apparatus for sending and parsing a transform skip flag and using the flag for decoding can be implemented as follows.
[0250] <Example A-1>
[0251] Table 1 below shows the syntax structure according to Example A-1.
[0252] Table 1
[0253]
[0254] Here, `transform_skip_flag[x0][y0][cIdx]` is the transform skip flag for the block corresponding to the channel indicated by `cIdx` and the spatial coordinates (x0, y0). `cIdx` indicates the channel currently being encoded (or decoded). In other words, the luminance channel corresponds to `cIdx = 0`, and the Cb and Cr channels correspond to `cIdx = 1` and `cIdx = 2`, respectively. Therefore, the transform skip flag information for the current block of the currently encoded (decoded) channel can usually be indicated by `transform_skip_flag[x0][y0][cIdx]`. However, in the following operational description, `transform_skip_flag[x0][y0][cIdx]` will simply be indicated as `transform_skip_flag` or `transform_skip_flag[cIdx]`, where some information is omitted for convenience.
[0255] According to embodiment A-1, for each of the three channels (i.e., cIdx = 0, 1, 2 or Y, Cb, Cr), a transform_skip_flag is transmitted (by the encoder) or parsed (by the decoder). In this case, the transform skip flag value transmitted or parsed can be transform_skip_flag[x0][y0][0], transform_skip_flag[x0][y0][1], and transform_skip_flag[x0][y0][2] corresponding to the Y, Cb, and Cr channels (i.e., cIdx = 0, 1, 2), respectively.
[0256] Alternatively, the tu_joint_cbcr_residual technique can be applied to encode (or decode) only the Cb block and generate the Cr block using the value of the decoded Cb block. Whether to use this technique is determined by the tu_joint_cbcr_residual_flag. The tu_joint_cbcr_residual technique is applied when the tu_joint_cbcr_residual_flag has a value of 1. In this case, the transform_skip_flag is sent or parsed only for the Y and Cb channels (here, a total of 2 bits of transform skip flags are needed), instead of sending or parsing for each of the three channels (i.e., cIdx = 0, 1, 2 or Y, Cb, Cr) (here, a total of 3 bits of transform skip flags are needed). In other words, when the transform skip flag for the Cb channel replaces the transform skip flag for the Cr channel, not only is the encoding efficiency improved, but a transform skip mode for the chroma channel is also implemented.
[0257] Unlike Example A-1 above, as shown in Example A-2 below, the order of operations for parsing the flag value and constructing the residual signal (or prediction error signal) can be modified.
[0258] <Example A-2>
[0259] Table 2 below shows the syntax structure according to embodiment A-2. The operational description according to embodiment A-2 is the same as that according to embodiment A-1.
[0260] Table 2
[0261]
[0262] <Example B>
[0263] According to an embodiment, another simple method can be implemented such that when transform skipping is applied to the Y, Cb, and Cr channels respectively, a transform skip flag is sent (or parsed) for the luma block, and the transform skip flag used for the luma block is shared for the chroma block, instead of sending (or parsing) a separate transform skip flag for the chroma block. This embodiment consists of the following two operations. First (operation 1), the decoder only parses the transform_skip_flag indicating whether transform skipping is applied to the luma channel. However (operation 2), the decoder shares the transform_skip_flag for the luma block (i.e., uses the same value), and does not separately parse the transform_skip_flag for the Cb and Cr blocks. The decoder according to this embodiment can be implemented to perform operations 1 and 2 as described in Table 3.
[0264] Table 3
[0265]
[0266] According to the embodiment, without considering the block partitioning tree structure such as a single tree or dual tree set during encoding, or the prediction mode such as intra-frame prediction or inter-frame prediction, the transform skip flag of the luma block corresponding to the chroma block can be implemented as referenced or shared by the chroma block. According to the embodiment, the parsing operation corresponding to operation 1 of this embodiment can only be implemented if all or some of the conditions for parsing transform_skip_flag[x0][y0][0] of embodiment A-1 or embodiment A-2 are satisfied. Table 4 describes operation 1 with the parsing conditions of embodiment A-1 or embodiment A-2 applied without change.
[0267] Table 4
[0268]
[0269] In another embodiment, it may be more advantageous to share the transform skip flag of the luma block corresponding to the current chroma block only under specific conditions, rather than always sharing the transform skip flag of the luma block as in embodiment B. Various modifications can be described here depending on the conditions used. Such modified configurations can be implemented such that operation 1 of embodiment B-1 is performed, and then operation 2 is modified under various conditions. Therefore, in the following description of the modified embodiments of embodiment B (embodiments B-1, B-2, ...), only operation 2 is described. However, it should be understood that the following embodiments describe additional operations performed after operation 1 described above.
[0270] <Example B-1>
[0271] According to an embodiment, the high-level syntax can define that transform skipping is not allowed for chroma blocks. Since transform skipping is predefined as not being used for chroma blocks, no operation is needed to indicate whether transform skipping is used for chroma blocks. That is, in this embodiment, the transform skipping flag is sent or parsed only for luma blocks.
[0272] In this embodiment, the encoding performance of the chroma block may be slightly reduced. On the other hand, the reduced encoding complexity is an advantage. However, in some cases, as shown in Table 5, the transformation skip operation can be indicated by explicitly setting transform_skip_flag[x0][y0][2] = transform_skip_flag[x0][y0][1] = 0.
[0273] Table 5
[0274]
[0275] <Example B-2>
[0276] According to the embodiment, when the condition that the current chroma block is in derivation mode (DM) (i.e., intra-prediction mode) is met in operation 2, the transform skip flag of the luma block corresponding to the current chroma block can be shared. As the intra-prediction mode of the chroma block, DM refers to the intra-prediction mode of the luma block used through sharing. CuPredMode[chType][x0][y0] indicates the prediction mode of the current coding block corresponding to spatial coordinates (x0, y0). chType = 0 indicates that the current coding block is a luma block. chType = 1 or chType = 2 indicates that the current coding block is a chroma block.
[0277] `intra_chroma_pred_mode` indicates the intra-prediction mode of the current chroma block. `intra_chroma_pred_mode = 4` corresponds to the case where the intra-prediction mode of the current chroma block is DM. The decoder according to this embodiment can then perform operation 2 as shown in Table 6.
[0278] Table 6
[0279]
[0280] <Example B-3>
[0281] According to an embodiment, when the condition that the current chroma block is in inter-frame prediction mode (“MODE_INTER”) is met in operation 2, the transform skip flag of the luma block corresponding to the current chroma block can be shared. The decoder according to this embodiment can then execute operation 2 as shown in Table 7.
[0282] Table 7
[0283]
[0284] <Example B-4>
[0285] Since Examples B-2 and B-3 are applied to the cases where the current chroma block is in intra-frame prediction mode and the cases where the current chroma block is in inter-frame prediction mode, respectively, these two examples can be combined to implement the following operations in Table 8.
[0286] Table 8
[0287]
[0288] <Example B-5>
[0289] To determine whether to skip the transform for blocks Y, Cb, and Cr respectively, the transform skip flag for block Y and the transform skip flag for block Cb can be sent (or parsed) separately, and then block Cr can share the transform skip flag for block Cb. According to this embodiment, only the transform_skip_flag indicating whether to apply the transform skip mode is sent for channels Y and Cb, and the decoder only parses the transform_skip_flag in the bitstream for channels Y and Cb. The decoder according to this embodiment can be implemented to perform operation 2 as shown in Table 9.
[0290] Table 9
[0291]
[0292] According to an embodiment, after the transformation skip flag of the Cr block is sent (or parsed), the Cb block can share the transformation skip flag of the Cr block.
[0293] <Example B-6>
[0294] Typically, the signal characteristics of the Cb and Cr channels are similar, except that they have opposite signs in many cases. Therefore, instead of encoding the prediction error signal blocks (“residual blocks”) for the Cb and Cr channels separately, only the residual blocks of the Cb channels can be encoded, and the residual blocks of the Cr channels can be left unencoded. In this case, it can be implemented such that the opposite sign of the decoded residual block of the Cr channel is used for the residual block of the Cb block. This technique is called the “tu_joint_cbcr_residual” technique.
[0295] The `tu_joint_cbcr_flag` can be used to indicate whether the Cb and Cr blocks are encoded using the techniques described above. When `tu_joint_cbcr_flag` = 1, the current block can be decoded according to the `tu_joint_cbcr_residual` mode. When the conditions for using the `tu_joint_cbcr_residual` mode are met, the residual block of the Cb channel can be decoded, and then the residual block of the Cr channel can be decoded by inverting the sign of the block value of the Cb channel. When `tu_joint_cbcr_flag` = 0, the `tu_joint_cbcr_residual` mode is not applied to the current block. After parsing the compressed residual data for the Cb and Cr blocks in the bitstream separately, each block can be decoded.
[0296] When tu_joint_cbcr_flag = 1, the decoder can parse the transform_skip_flag of the Cb block and apply it to the Cr block, without using the parsing operation for the transform_skip_flag of the Cr block. Alternatively, conversely, when tu_joint_cbcr_flag = 1, the decoder can parse the transform_skip_flag of the Cr block and apply it to the Cb block, without using the parsing operation for the transform_skip_flag of the Cb block.
[0297] When tu_joint_cbcr_flag = 0, the decoder can decode the Cb and Cr blocks after parsing the compressed residual data of the Cb and Cr blocks in the bitstream respectively. As shown in Table 10, the decoder according to this embodiment can be implemented using one (combination) of the operations in Embodiment B or Embodiments B-1 to B-5 (referred to as Operation 2).
[0298] Table 10
[0299]
[0300] <Example B-7>
[0301] According to the embodiment, operation 2 is performed as follows. When tu_joint_cbcr_flag = 1, the decoder can parse the transform_skip_flag of the Cb block and apply the transform_skip_flag of the Cb block to the Cr block as is. Alternatively, when tu_joint_cbcr_flag = 0, as shown in Table 11, the decoder can parse the transform_skip_flag for both the Cb and Cr blocks respectively.
[0302] Table 11
[0303]
[0304] <Example B-8>
[0305] Operation 2, performed in Examples B-1 to B-7, can be implemented as being restricted according to the size or shape of the current block, as shown in Table 12.
[0306] Table 12
[0307]
[0308] The conditions according to the embodiment (i.e., the recommended practice of condition check 1) can be configured according to the application as follows.
[0309] The suggested practice condition check 1 → (tbWidth != tbHeight)
[0310] b. Recommendation for implementation: Condition check 1 → (tbWidth == tbHeight)
[0311] c suggests checking the following condition for implementation: 1 → (tbWidth >= N*tbHeight)
[0312] d suggests checking the following conditions for implementation: 1 → (tbWidth >= N*tbHeight)
[0313] Here, tbWidth and tbHeight refer to the horizontal length and vertical height of the current encoding unit, respectively. Additionally, according to the embodiment, the inequality symbol "=<" indicating "equal to or less than" on the left side can be replaced by "<" indicating "less than" on the right side. Furthermore, N, representing the ratio of width to height (aspect ratio), is calculated as width / height and can be at least one of 2, 4, 8, 16, and 32.
[0314] <Example B-9>
[0315] As shown in Table 13, the size of the block can be applied in a limited manner in all cases of Embodiment B, Embodiment B-1 to Embodiment B-7.
[0316] Table 13
[0317]
[0318] The conditions according to the embodiment (i.e., the recommended practice condition check 2) can be configured according to the application as follows.
[0319] a. Recommendation for practice: Condition check 2 → ((tbWidth) <N1)&&(cbHeight<N2))
[0320] b suggests checking the conditions for implementation 2→((tbWidth) <N1)OR(cbHeight<N2))
[0321] c suggests checking the conditions for implementation 2 → (min(tbWidth, tbHeight) <N3)
[0322] d suggests checking the conditions for implementation 2 → (max(tbwidth,tbheight)). <N4)
[0323] e suggests checking the conditions for implementation 2→(tbWidth*tbHeight) <N5)
[0324] f suggests checking the conditions for implementation 2→(tbWidth+tbHeight) <N6)
[0325] g suggests checking the conditions for implementation 2 → (tbwidth==N7&&tbheight==N6)
[0326] h suggests checking the conditions for implementation 2 → (tbwidth==N7&&tbheight==N7)
[0327] i suggests implementing the following condition check: 2 → (log(tbwidth) + log(tbheight)) <N8)
[0328] In the various embodiments described above, depending on the application, the inequality symbol "<" indicating that the left side is "less than" the right side can be replaced with "=<" indicating that the left side is equal to or less than the right side. Additionally, N1 to N7, which indicate predetermined boundary condition values, can be one of 2, 4, 8, 12, 16, 32, 64, and 128, respectively, and N8 can be one of 0, 1, 2, 3, 4, 5, and 6.
[0329] <Example C>
[0330] According to the embodiment, it may be necessary to enable or disable the use of transform skip mode. For example, when transform skipping is beneficial to part of the image but detrimental to the rest, transform skip mode can be enabled or disabled globally at the parent level of the CU. In this regard, whether transform skipping is possible can be set by setting transform_skip_enabled_flag and transform_skip_chroma_enabled_flag.
[0331] The `transform_skip_enabled_flag`, which indicates whether transformation is enabled, is applied to both the luma and chroma blocks. That is, `transform_skip_enabled_flag = 0` indicates that the transformation skip mode will not be used. Conversely, `transform_skip_enabled_flag = 1` indicates that the transformation skip mode will be used.
[0332] The `transform_skip_chroma_enabled_flag` indicates whether transform skipping can be enabled for chroma blocks. When both `transform_skip_enabled_flag` and `transform_skip_chroma_enabled_flag` are 1, transform skipping mode is enabled for chroma blocks. When both `transform_skip_enabled_flag` and `transform_skip_chroma_enabled_flag` are 0, transform skipping mode is enabled for luma blocks, but may not be enabled for chroma blocks.
[0333] When `transform_skip_enabled_flag = 0`, transform skipping may not be applied to either the luma or chroma blocks. Therefore, when `transform_skip_enabled_flag = 0`, `transform_skip_chroma_enabled_flag` may not be included in the bitstream. Furthermore, `transform_skip_chroma_enabled_flag` can be set to 0.
[0334] The transform_skip_enabled_flag and transform_skip_chroma_enabled_flag can be implemented as shown in Table 14 and are sent by being included in a higher-level unit such as the SPS (Sequence Parameter Set).
[0335] Table 14
[0336]
[0337] However, in the case of embodiments that frequently repeat natural video and computer-generated video, transform_skip_enabled_flag and transform_skip_chroma_enabled_flag can be sent by being included in the PPS (Picture Parameter Set), as shown in Table 15.
[0338] Table 15
[0339]
[0340] In another embodiment, transform_skip_enabled_flag and transform_skip_chroma_enabled_flag may be included in both SPS and PPS.
[0341] The implementation of enabling or disabling the transform skip mode globally by using transform_skip_enabled_flag and transform_skip_chroma_enabled_flag can be applied to all the above implementations.
[0342] As shown in Table 16, <Example A-1> can be implemented by modifying it to <Example C-1> below.
[0343] Table 16
[0344]
[0345] As shown in Table 17, <Example A-2> can be implemented by modifying it to <Example C-2> below.
[0346] <Example C-2>
[0347] Table 17
[0348]
[0349] In the same manner as in implementing <Example C-1> and <Example C-2>, all the embodiments described above can be implemented to globally enable or disable the transformation skip mode by using transform_skip_enabled_flag and transform_skip_chroma_enabled_flag.
[0350] Figure 9 A video decoding method according to an embodiment of the present disclosure is shown.
[0351] In step S902, a chromatic residual joint flag indicating whether the chromatic residual joint mode is applied to the current block can be obtained.
[0352] According to an embodiment, the video decoding method may further include obtaining a chroma residual union enable flag from the bitstream, indicating whether the chroma residual union mode is enabled for the parent unit of the current block. Additionally, when the chroma residual union enable flag indicates that the chroma residual union mode is enabled for the parent unit of the current block, the chroma residual union flag of the current block can be obtained.
[0353] According to an embodiment, the upper-level unit can be at least one of video, coded video sequence (CVS), frame, sub-frame, strip, parallel block, and coding tree unit.
[0354] In step S904, a transform skip mode flag for the luminance component of the current block and a transform skip mode flag for the first chrominance component can be obtained from the bit stream.
[0355] According to an embodiment, a transform skip mode flag for the luminance component of the current block can be obtained based on the size of the current block. Additionally, a transform skip mode flag for the first chrominance component of the current block can be obtained based on the size of the current block. Furthermore, a transform skip mode flag for the second chrominance component of the current block can be obtained based on the size of the current block.
[0356] According to an embodiment, a transform skip mode flag for the first chromaticity component of the current block can be obtained based on the size and color format of the current block.
[0357] According to an embodiment, when the height of the current block is equal to or less than the maximum block height and the width of the current block is equal to or less than the maximum block width, a transition skip mode flag can be obtained for the luminance component of the current block. Additionally, when the height of the current block adjusted according to the color format is equal to or less than the maximum block height and the width of the current block adjusted according to the color format is equal to or less than the maximum block width, a transition skip mode flag can be obtained for the first chrominance component of the current block.
[0358] In step S906, when the chromaticity residual joint flag indicates that the chromaticity residual joint mode is applied to the current block, it can be determined whether the transform skip mode is applied to the second chromaticity component of the current block according to the transform skip mode flag for the first chromaticity component.
[0359] In step S908, when the chroma residual union flag indicates that the chroma residual union mode is not applied to the current block, the transform skip mode flag for the second chroma component of the current block can be obtained from the bit stream.
[0360] According to an embodiment, a transform skip mode flag for the second chroma component of the current block can be obtained based on the size of the current block. Additionally, the transform skip mode flag for the second chroma component of the current block can be obtained based on the size and color format of the current block.
[0361] According to an embodiment, when the height of the current block adjusted according to the color format is equal to or less than the maximum block height and the width of the current block adjusted according to the color format is equal to or less than the maximum block width, a transform skip mode flag can be obtained for the second chromaticity component of the current block.
[0362] According to an embodiment, the video decoding method may further include: when a chroma residual joint flag indicates that a chroma residual joint mode is applied to the current block, determining the residual sample points of the second chroma component of the current block based on the residual sample points of the first chroma component of the current block.
[0363] According to an embodiment, the size of the residual sample point of the second chromaticity component can be determined based on the size of the residual sample point of the first chromaticity component, and the sign of the residual sample point of the second chromaticity component can be determined to be opposite to the sign of the residual sample point of the first chromaticity component.
[0364] According to an embodiment, the video decoding method may further include obtaining chroma residual joint symbol information from the bitstream, which indicates the relationship between the symbols of the residual samples of the first chroma component and the symbols of the residual samples of the second chroma component. Furthermore, the symbols of the residual samples of the second chroma component may be determined based on the chroma residual joint symbol information and the symbols of the residual samples of the first chroma component.
[0365] Figure 10 A video encoding method according to an embodiment of the present disclosure is shown.
[0366] In step S1002, the chroma residual joint flag of the current block, which indicates whether the chroma residual joint mode is applied to the current block, is encoded.
[0367] According to an embodiment, it can be determined whether to apply the chromaticity residual joint mode to the current block based on the residual samples of the first chromaticity component and the second chromaticity component of the current block.
[0368] According to an embodiment, when the chromaticity residual joint mode is applied to the current block, the chromaticity residual joint symbol information indicating the relationship between the symbols of the residual samples of the first chromaticity component and the residual samples of the second chromaticity component of the current block can be encoded based on the residual samples of the first chromaticity component and the residual samples of the second chromaticity component of the current block.
[0369] According to an embodiment, the video decoding method may further include encoding a chroma residual joint enable flag that indicates whether a chroma residual joint mode is enabled for the parent unit of the current block. Additionally, when the chroma residual joint enable flag indicates that a chroma residual joint mode is enabled for the parent unit of the current block, the chroma residual joint flag of the current block may be encoded.
[0370] According to an embodiment, the upper-level unit can be at least one of video, coded video sequence (CVS), frame, sub-frame, strip, parallel block, and coding tree unit.
[0371] In step S1004, the transform skip mode flag for the luminance component of the current block and the transform skip mode flag for the first chrominance component of the current block are encoded.
[0372] According to an embodiment, the transform skip mode flag for the luminance component of the current block can be encoded according to the size of the current block. Additionally, the transform skip mode flag for the first chrominance component of the current block can be encoded according to the size of the current block.
[0373] According to an embodiment, the transform skip mode flag for the first chromaticity component of the current block can be encoded according to the size and color format of the current block.
[0374] According to an embodiment, when the height of the current block is equal to or less than the maximum block height and the width of the current block is equal to or less than the maximum block width, a transform skip mode flag for the luminance component of the current block can be encoded. Additionally, when the height of the current block adjusted according to the color format is equal to or less than the maximum block height and the width of the current block adjusted according to the color format is equal to or less than the maximum block width, a transform skip mode flag for the first chrominance component of the current block can be encoded.
[0375] In step S1006, when the chroma residual joint mode is applied to the current block, the encoding of the transform skip mode flag for the second chroma component used in the current block is skipped.
[0376] In step S1008, when the chroma residual joint mode is not applied to the current block, the transform skip mode flag for the second chroma component used in the current block is encoded.
[0377] According to an embodiment, the transform skip mode flag for the second chroma component of the current block can be encoded according to the size of the current block. Alternatively, the transform skip mode flag for the second chroma component of the current block can be encoded according to the size and color format of the current block.
[0378] According to an embodiment, Figure 9 Video decoding methods and Figure 10In the video coding method, the first chromaticity component and the second chromaticity component can be Cb component and Cr component respectively, or Cr component and Cb component respectively.
[0379] Figure 9 and Figure 10 The embodiments described are illustrative and can be readily modified by those skilled in the art. Figure 9 and Figure 10 Each step. Additionally, Figure 9 and Figure 10 Each configuration can be omitted or replaced by another configuration. Figure 9 Video decoding methods can be used Figure 2 Executed in the decoder. Additionally, Figure 10 Video encoding methods can be used Figure 1 Executed in the encoder. Furthermore, one or more processors can execute the implementation. Figure 9 and Figure 10 The commands for each step. Additionally, implementation... Figure 9 and Figure 10 The program product of each step of the command can be stored in a memory device or circulated online.
[0380] The above embodiments can be performed in the same way in both the encoder and decoder.
[0381] At least one embodiment or combination of the above embodiments can be used to encode / decode video.
[0382] The order in which the above embodiments are applied may differ between the encoder and the decoder, or the order in which the above embodiments are applied may be the same in the encoder and the decoder.
[0383] The above embodiments can be performed on each luminance signal and each chrominance signal, or the above embodiments can be performed on both luminance and chrominance signals in the same way.
[0384] The block form of the above embodiments of the present invention can be square or non-square.
[0385] At least one of the syntax elements (flags, indexes, etc.) entropy encoded by the encoder and entropy decoded by the decoder can be obtained using at least one of the following binarization / debinarization methods and entropy encoding / decoding methods.
[0386] Signed 0th-order Exp_Golomb binarization / inverse binarization method (se(v))
[0387] Signed k-order Exp_Golomb binarization / inverse binarization method (sek(v))
[0388] Unsigned 0th-order Exp_Golomb binarization / inverse binarization method (ue(v))
[0389] Unsigned k-order Exp_Golomb binarization / inverse binarization method (uek(v))
[0390] Fixed-length binarization / inverse binarization method (f(n))
[0391] Truncation of Rice binarization / inverse binarization methods or truncation of univariate binarization / inverse binarization methods (tu(v))
[0392] Truncated binary binarization / inverse binarization method (tb(v))
[0393] Context-adaptive arithmetic coding / decoding method (ae(v))
[0394] Byte unit bit string (b(8))
[0395] Signed integer binarization / inverse binarization methods (i(n))
[0396] Unsigned integer binarization / debinarization method (u(n))
[0397] Univariate binarization / inverse binarization methods
[0398] The above embodiments of the present invention can be applied based on the size of at least one of the coding block, prediction block, transform block, block, current block, coding unit, prediction unit, transform unit, unit, and current unit. Here, the size can be defined as a minimum size or a maximum size, or both, such that the above embodiments are applied, or the size can be defined as a fixed size for applying the above embodiments. Furthermore, in the above embodiments, the first embodiment can be applied to a first size, and the second embodiment can be applied to a second size. In other words, the above embodiments can be applied in combination based on the size. Furthermore, the above embodiments can be applied when the size is equal to or greater than the minimum size and equal to or less than the maximum size. In other words, the above embodiments can be applied when the block size is included within a specific range.
[0399] For example, the above embodiments can be applied when the size of the current block is 8×8 or larger. For example, the above embodiments can be applied when the size of the current block is only 4×4. For example, the above embodiments can be applied when the size of the current block is 16×16 or smaller. For example, the above embodiments can be applied when the size of the current block is equal to or greater than 16×16 and equal to or less than 64×64.
[0400] The above embodiments of the present invention can be applied according to time layers. To identify the time layer to which the above embodiments can be applied, a corresponding identifier can be sent by a signal, and the above embodiments can be applied to the specified time layer identified by the corresponding identifier. Here, the identifier can be defined as the lowest or highest layer to which the above embodiments can be applied, or both, or it can be defined as indicating a specific layer to which the embodiments are applied. Furthermore, a fixed time layer to which the embodiments are applied can be defined.
[0401] For example, the above embodiments can be applied when the time layer of the current image is the lowest layer. For example, the above embodiments can be applied when the time layer identifier of the current image is 1. For example, the above embodiments can be applied when the time layer of the current image is the highest layer.
[0402] The strip type or parallel block group type of the above embodiments of the present invention can be defined, and the above embodiments can be applied according to the corresponding strip type or parallel block group type.
[0403] In the above embodiments, the method is described based on a flowchart having a series of steps or units. However, the present invention is not limited to the order of the steps, but some steps may be performed simultaneously with other steps or in a different order. Furthermore, those skilled in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps may be added to the flowchart or some steps may be deleted from the flowchart without affecting the scope of the present invention.
[0404] The embodiments described include various aspects of the examples. Not all possible combinations of these aspects may be described, but those skilled in the art will recognize different combinations. Therefore, the invention may include all substitutions, modifications, and alterations within the scope of the claims.
[0405] Embodiments of the invention may be implemented in the form of program instructions executable by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include individual program instructions, data files, data structures, etc., or combinations thereof. The program instructions recorded in the computer-readable recording medium may be specifically designed and constructed for the present invention or may be well known to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic recording media (such as hard disks, floppy disks, and magnetic tapes), optical data storage media (such as CD-ROMs or DVD-ROMs), magneto-optical media (such as floppy disks), and hardware devices (such as read-only memory (ROM), random access memory (RAM), flash memory, etc.) specifically configured to store and implement program instructions. Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code that can be implemented by a computer using an interpreter. The hardware device may be configured to operate by one or more software modules to perform processing according to the present invention, or vice versa.
[0406] Although the invention has been described with reference to specific items such as detailed elements and limited embodiments and drawings, these are provided only to aid in a broader understanding of the invention, and the invention is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and changes can be made from the above description.
[0407] Therefore, the spirit of the present invention should not be limited to the above embodiments, and the entire scope of the claims and their equivalents shall fall within the scope and spirit of the present invention.
[0408] Industrial applicability
[0409] This invention can be used to encode or decode images.
Claims
1. A video decoding method, the method comprising: Obtain the chroma residual union enable flag from the bitstream, indicating whether to enable the chroma residual union mode for the parent unit of the current block; When the chroma residual joint enable flag indicates that the chroma residual joint mode is enabled for the parent unit of the current block, a chroma residual joint flag indicating whether the chroma residual joint mode is applied to the current block is obtained. Obtain from the bitstream the transform skip mode flag for the luminance component of the current block and the transform skip mode flag for the first chrominance component; When the chroma residual joint flag indicates that the chroma residual joint mode is applied to the current block, it is determined whether to obtain the transform skip mode flag for the second chroma component of the current block from the bit stream based on the coded block flag for the first chroma component. as well as When the chroma residual union flag indicates that the chroma residual union mode is not applied to the current block, obtain the transform skip mode flag for the second chroma component of the current block from the bit stream.
2. The video decoding method according to claim 1 further includes: When the chromaticity residual joint flag indicates that the chromaticity residual joint mode is applied to the current block, the residual sample points of the second chromaticity component of the current block are determined based on the residual sample points of the first chromaticity component of the current block.
3. The video decoding method according to claim 2, in, The size of the residual sample points of the second chromaticity component is determined based on the size of the residual sample points of the first chromaticity component, and The sign of the residual samples of the second chromaticity component is determined to be opposite to the sign of the residual samples of the first chromaticity component.
4. The video decoding method according to claim 2 further includes: The chromaticity residual joint symbol information, which indicates the relationship between the symbols of the residual samples of the first chromaticity component and the symbols of the residual samples of the second chromaticity component, is obtained from the bitstream. Specifically, the sign of the residual sample of the second chromaticity component is determined based on the joint sign information of the chromaticity residual and the sign of the residual sample of the first chromaticity component.
5. The video decoding method according to claim 1, in, The step of obtaining the transform skip mode flag for the luminance component of the current block and the transform skip mode flag for the first chrominance component includes: Based on the size of the current block, obtain a transform skip mode flag for the luminance component of the current block; and The transform skip mode flag for the first chromaticity component of the current block is obtained based on the size of the current block.
6. The video decoding method according to claim 5, in, The step of obtaining the transform skip mode flag for the luminance component of the current block based on the size of the current block includes: When the height of the current block is equal to or less than the maximum block height and the width of the current block is equal to or less than the maximum block width, a transform skip mode flag for the luminance component of the current block is obtained.
7. The video decoding method according to claim 5, in, The step of obtaining the transform skip mode flag for the first chromaticity component of the current block based on the size of the current block includes: The transform skip mode flag for the first chromaticity component of the current block is obtained based on the size and color format of the current block.
8. The video decoding method according to claim 1, in, The higher-level unit is at least one of video, encoded video sequence (CVS), frame, sub-frame, strip, parallel block, and coding tree unit.
9. The video decoding method according to claim 1, in, The first chromaticity component and the second chromaticity component are Cb component and Cr component, respectively, or Cr component and Cb component, respectively.
10. A video encoding method, the method comprising: The chroma residual joint enable flag, which indicates whether the chroma residual joint mode is enabled for the parent unit of the current block, is encoded. When the chroma residual joint enable flag indicates that the chroma residual joint mode is enabled for the parent unit of the current block, the chroma residual joint flag of the current block that indicates whether the chroma residual joint mode is applied to the current block is encoded. The transform skip mode flag for the luminance component of the current block and the transform skip mode flag for the first chrominance component of the current block are encoded. When the chroma residual joint mode is applied to the current block, the encoding of the transform skip mode flag for the second chroma component used in the current block is skipped; as well as When the chroma residual joint mode is not applied to the current block, the transform skip mode flag for the second chroma component of the current block is encoded.
11. The video encoding method as described in claim 10, further comprising: Based on the residual samples of the first chromaticity component and the second chromaticity component of the current block, determine whether the chromaticity residual joint mode is applied to the current block.
12. The video encoding method according to claim 11, further comprising: Based on the residual samples of the first chromaticity component and the second chromaticity component of the current block, the chromaticity residual joint symbol information indicating the relationship between the symbols of the residual samples of the first chromaticity component and the symbols of the residual samples of the second chromaticity component is encoded.
13. The video encoding method according to claim 10, in, The step of encoding the transform skip mode flag for the luminance component of the current block and the transform skip mode flag for the first chrominance component of the current block includes: The transform skip mode flag for the luminance component of the current block is encoded according to the size of the current block; and The transform skip mode flag for the first chromaticity component of the current block is encoded according to the size of the current block.
14. The video encoding method according to claim 13, in, The step of encoding the transform skip mode flag for the luminance component of the current block according to the size of the current block includes: When the height of the current block is equal to or less than the maximum block height and the width of the current block is equal to or less than the maximum block width, the transform skip mode flag for the luminance component of the current block is encoded.
15. The video encoding method according to claim 13, in, The step of encoding the transform skip mode flag for the first chroma component of the current block according to the current block size includes: The transform skip mode flag for the first chromaticity component of the current block is encoded according to the size and color format of the current block.
16. The video encoding method according to claim 10, in, The higher-level unit is at least one of video, encoded video sequence (CVS), frame, sub-frame, strip, parallel block, and coding tree unit.
17. The video encoding method according to claim 10, in, The first chromaticity component and the second chromaticity component are Cb component and Cr component, respectively, or Cr component and Cb component, respectively.
18. A computer-readable recording medium for storing program instructions for transmitting a bitstream comprising video data encoded according to a video coding method. in, The processor executes program instructions to generate and transmit the bitstream, the bitstream comprising: The chroma residual joint enable flag of the current block indicates whether the chroma residual joint mode is enabled for the parent unit of the current block; The chroma residual joint mode is indicated as to whether the chroma residual joint flag of the current block is applied to the current block; Transform skip mode flag for the luminance component of the current block; and Transform skip mode flag for the first chromaticity component of the current block Wherein, when the chroma residual joint flag indicates that the chroma residual joint mode is not applied to the current block, the bitstream also includes a transform skip mode flag for the second chroma component of the current block. Specifically, when the chroma residual union enable flag indicates that the chroma residual union mode is not enabled for the parent unit of the current block, the chroma residual union flag of the current block is skipped in the bitstream. Specifically, when the chroma residual joint flag indicates that the chroma residual joint mode is applied to the current block, it is determined whether to skip the transform skip mode flag for the second chroma component of the current block in the bitstream based on the coding block flag for the first chroma component.
Citation Information
Patent Citations
Image processing method and device therefor
KR1020190090865A