Method and apparatus for encoding and decoding using selective information sharing between channel
By sharing the encoding decision information of representative channels in video encoding and decoding, and adopting cross-channel prediction and intra prediction modes, the efficient encoding and decoding problems of high-resolution and high-quality images are solved, reducing transmission and storage costs, and improving encoding and decoding efficiency.
Patent Information
- Application Number
- CN202510649793.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-21
- Filing Date
- 2018-12-07
- Publication Date
- 2025-08-08
AI Technical Summary
With the increasing demand for high-definition and ultra-high-definition images, the increase in the amount of image data leads to an increase in transmission and storage costs, and existing image encoding/decoding technologies are difficult to efficiently process high-resolution and high-quality images.
By sharing the encoding decision information of the representative channel of the target block as the encoding decision information of the target channel during the video encoding and decoding process, duplicate information transmission is reduced, and cross-channel prediction and intra prediction modes are adopted to improve encoding and decoding efficiency.
It effectively reduces the amount of image data, reduces transmission and storage costs, and improves encoding and decoding efficiency.
Smart Images

Figure CN120455655A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with an application date of December 7, 2018, application number 201880088927.1, and invention name “Method and device for encoding and decoding using selective information sharing between channels”. Technical Field
[0002] The following embodiments generally relate to a video decoding method and apparatus and a video encoding method and apparatus, and more particularly, to a video decoding method and apparatus and a video encoding method and apparatus using sharing of selective information between channels. Background Art
[0003] With the continuous development of the information and communication industry, broadcast services supporting high-definition (HD) resolution have become popular around the world. With this popularity, a large number of users have become accustomed to high-resolution and high-definition images and / or videos.
[0004] To meet user demand for higher definition, numerous organizations have accelerated the development of next-generation imaging devices. In addition to high-definition television (HDTV) and full-high-definition (FHD) television, interest in ultra-high-definition television (UHDTV), which offers over four times the resolution of full-high-definition (FHD) television, has also grown. This growing interest is driving a growing demand for image encoding and decoding technologies that deliver higher resolution and definition.
[0005] Image encoding / decoding devices and methods may use inter-frame prediction technology, intra-frame prediction technology, entropy coding technology, etc. to perform encoding / decoding on high-resolution and high-definition images. Inter-frame prediction technology may be a technology for predicting the values of pixels included in a target picture using a temporally previous picture and / or a temporally subsequent picture. Intra-frame prediction technology may be a technology for predicting the values of pixels included in a target picture using information about the pixels in the target picture. Entropy coding technology may be a technology for assigning short codewords to frequently occurring symbols and long codewords to rarely occurring symbols.
[0006] Recently, the demand for high-quality images (such as ultra-high-definition (UHD) images) that can provide high resolution, a further broadened color space, and excellent image quality has increased in various application fields. As images tend toward higher resolution and higher quality images, the amount of image data required to provide the images may increase to exceed the existing image data amount. In the case of transmitting image data through a communication medium (such as a wired / wireless broadband line) or various broadcasting media (such as a satellite, terrestrial wave, Internet Protocol (IP) network, wireless network, cable or mobile communication network), or in the case of storing image data in various types of storage media (such as compact discs (CDs), digital versatile discs (DVDs), universal serial bus (USB) media, and high-definition (HD)-DVDs), as the amount of image data increases, the transmission cost and storage cost increase.
[0007] With the use of high-resolution and high-quality images, in order to solve inevitable and more serious problems in image data and provide images with higher resolution and higher image quality, efficient image encoding / decoding technology is required. Summary of the Invention
[0008] Technical issues
[0009] The embodiments are directed to providing an encoding apparatus and method and a decoding apparatus and method using sharing of selective information between channels.
[0010] Technical Solution
[0011] According to one aspect, a decoding method is provided, including: sharing encoding decision information of a representative channel of a target block as encoding decision information of a target channel of the target block; and performing decoding on the target block using the encoding decision information of the target channel.
[0012] The decoding method may further include receiving a bitstream including information about the target block.
[0013] The information about the target block may include encoding decision information of the representative channel.
[0014] The information about the target block may not include encoding decision information of the target channel.
[0015] The encoding decision information of the representative channel may be transform skip information indicating whether transform is to be skipped.
[0016] The encoding decision information of the representative channel may indicate which transform is to be used for the transform block of the channel.
[0017] The encoding decision information of the representative channel may be intra-frame encoding decision information of the representative channel.
[0018] The representative channel and the target channel may be channels in a YCbCr color space.
[0019] The representative channel may be a luminance channel.
[0020] The target channel may be a chroma channel.
[0021] The representative channel may be a color channel having the highest correlation with the luminance signal.
[0022] The representative channel may be determined by an index in the bitstream indicating the selected representative channel.
[0023] The sharing operation may be performed when image properties of a plurality of channels of the target block are similar to each other.
[0024] When the intra prediction mode of the chroma channel of the target block is the direct mode, image properties of the plurality of channels may be determined to be similar to each other.
[0025] The shared operation may be performed when cross-channel prediction is used.
[0026] Whether to use cross-channel prediction can be derived based on information obtained from the bitstream.
[0027] The shared operation may be performed when cross-channel prediction is used.
[0028] Whether to use cross-channel prediction may be determined based on the intra prediction mode of the target block.
[0029] When the intra prediction mode of the target block is one of the INTRA_CCLM mode, the INTRA_MMLM mode, and the INTRA_MFLM mode, cross-channel prediction may be used.
[0030] Whether sharing will be performed may be determined based on the size of the target block.
[0031] The encoding decision information of a representative channel among the multiple channels of the target block may be used for all channels among the multiple channels.
[0032] According to another aspect, an encoding method is provided, comprising: determining encoding decision information of a representative channel of a target block; and performing encoding on the target block using the encoding decision information of the representative channel, wherein the encoding decision information of the representative channel is shared with another channel of the target block.
[0033] The encoding method may further include generating a bitstream including information about the target block.
[0034] The information about the target block may include encoding decision information of the representative channel.
[0035] The information about the target block may not include encoding decision information of the additional channel.
[0036] The representative channel and the further channels may be channels in a YCbCr color space.
[0037] According to another aspect, a computer-readable storage medium storing a bitstream for image decoding is provided, the bitstream including information about a target block, wherein the information about the target block includes encoding decision information of a representative channel of the target block, wherein the encoding decision information of the representative channel of the target block is used and shared as encoding decision information of a target channel of the target block, and wherein decoding of the target block is performed using the encoding decision information of the target channel.
[0038] Beneficial effects
[0039] Provided are an encoding apparatus and method and a decoding apparatus and method using sharing of selective information between channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied;
[0041] Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied;
[0042] Figure 3 is a diagram schematically illustrating a partition structure of an image when the image is encoded and decoded;
[0043] Figure 4 is a diagram illustrating a form of prediction unit (PU) that a coding unit (CU) can include;
[0044] Figure 5 is a diagram illustrating a form of a transform unit (TU) that can be included in a CU;
[0045] Figure 6 shows the division of blocks according to an example;
[0046] Figure 7 is a diagram for explaining an embodiment of an intra prediction process;
[0047] Figure 8 is a diagram for explaining the positions of reference samples used in the intra prediction process;
[0048] Figure 9 is a diagram for explaining an embodiment of an inter-frame prediction process;
[0049] Figure 10shows spatial candidates according to an embodiment;
[0050] Figure 11 shows the order in which motion information of spatial candidates is added to a merge list according to an embodiment;
[0051] Figure 12 shows a transform and quantization process according to an example;
[0052] Figure 13 shows a diagonal scan according to an example;
[0053] Figure 14 shows a horizontal scan according to an example;
[0054] Figure 15 shows vertical scanning according to an example;
[0055] Figure 16 is a configuration diagram of an encoding device according to an embodiment;
[0056] Figure 17 is a configuration diagram of a decoding device according to an embodiment;
[0057] Figure 18 is a flowchart of a method for decoding encoding decision information according to an embodiment;
[0058] Figure 19 is a flowchart of a decoding method for determining whether a transform is to be skipped according to an embodiment;
[0059] Figure 20 is a flowchart of a decoding method for determining whether a transform is to be skipped with reference to an intra mode according to an embodiment;
[0060] Figure 21 is a flowchart of a method for sharing transform selection information according to an embodiment;
[0061] Figure 22 shows a single tree block partition structure according to an example;
[0062] Figure 23 shows a dual-tree block partition structure according to an example;
[0063] Figure 24 A scheme for specifying corresponding blocks based on positions in corresponding areas according to an example is shown;
[0064] Figure 25 A scheme for specifying corresponding blocks based on areas in corresponding regions according to an example is shown;
[0065] Figure 26 Another scheme for specifying corresponding blocks based on areas in corresponding regions according to an example is shown;
[0066] Figure 27 A scheme for specifying corresponding blocks based on the form of the blocks in the corresponding region according to an example is shown;
[0067] Figure 28 Another scheme for specifying corresponding blocks based on the form of the blocks in the corresponding region according to an example is shown;
[0068] Figure 29 shows a scheme for specifying corresponding blocks based on aspect ratios of blocks in a corresponding region according to an example;
[0069] Figure 30 Another scheme for specifying corresponding blocks based on aspect ratios of blocks in a corresponding region according to an example is shown;
[0070] Figure 31 A scheme for specifying corresponding blocks based on coding features of blocks in a corresponding region according to an example is shown;
[0071] Figure 32 is a flowchart of an encoding method according to an embodiment; and
[0072] Figure 33 is a flowchart of a decoding method according to an embodiment. DETAILED DESCRIPTION
[0073] The present invention can be variously modified and can have various embodiments, and specific embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that these embodiments are not intended to limit the present invention to the specific disclosed forms, and they include all changes, equivalent forms or modified forms included in the spirit and scope of the present invention.
[0074] The following exemplary embodiments will be described in detail with reference to the accompanying drawings showing specific embodiments. These embodiments are described so that those of ordinary skill in the art to which the present disclosure pertains can easily put these embodiments into practice. It should be noted that the various embodiments are different from one another, but do not need to be mutually exclusive. For example, the specific shapes, structures, and characteristics described herein can be implemented as other embodiments without departing from the spirit and scope of the multiple embodiments associated with an embodiment. In addition, it should be understood that the position or arrangement of the various components in each disclosed embodiment can be changed without departing from the spirit and scope of the embodiments. Therefore, the attached detailed description is not intended to limit the scope of the present disclosure, and the scope of the exemplary embodiments is limited only by the attached claims and their equivalents (as long as they are appropriately described).
[0075] In the accompanying drawings, like reference numerals are used to designate the same or similar functions in various aspects. The shapes, sizes, etc. of components in the accompanying drawings may be exaggerated to make the description clear.
[0076] Terms such as "first" and "second" may be used to describe various components, but the components are not limited by the terms. The terms are only used to distinguish one component from another. For example, a first component may be referred to as a second component without departing from the scope of this specification. Similarly, a second component may be referred to as a first component. The term "and / or" may include a combination of multiple related description items or any one of the multiple related description items.
[0077] It will be understood that when a component is referred to as being “connected” or “coupled” to another component, the two components may be directly connected or coupled to each other, or intervening components may be present between the two components. It will be understood that when a component is referred to as being “directly connected or coupled,” there are no intervening components between the two components.
[0078] In addition, the components described in the embodiments are shown independently to represent different feature functions, but this does not mean that each component is formed by a separate hardware or software. That is, for the convenience of description, multiple components are arranged and included separately. For example, at least two components of the multiple components can be integrated into a single component. Conversely, a component can be divided into multiple components. As long as it does not deviate from the essence of this specification, embodiments in which multiple components are integrated or embodiments in which some components are separated are included in the scope of this specification.
[0079] Furthermore, it should be noted that, in exemplary embodiments, the expression describing components “including” specific components means that additional components may be included within the scope of practice or technical spirit of the exemplary embodiments, but does not exclude the existence of components other than the specific components.
[0080] The terms used in this specification are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context specifically indicates the contrary. In this specification, it should be understood that terms such as "including" or "having" are only intended to indicate the presence of features, numbers, steps, operations, components, parts or combinations thereof, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.
[0081] The embodiments will be described in detail below with reference to the accompanying drawings so that those skilled in the art can easily practice the embodiments. In the following description of the embodiments, detailed descriptions of well-known functions or configurations that are deemed to obscure the main points of this specification will be omitted. In addition, the same reference numerals are used to designate the same components throughout the drawings, and repeated descriptions of the same components will be omitted.
[0082] Hereinafter, "image" may refer to a single frame constituting a video, or may refer to the video itself. For example, "encoding and / or decoding an image" may refer to "encoding and / or decoding a video" or "encoding and / or decoding any one of a plurality of images constituting a video."
[0083] Hereinafter, the terms "video" and "moving picture" may be used to have the same meaning and may be used interchangeably with each other.
[0084] Hereinafter, the target image may be an encoding target image as a target to be encoded and / or a decoding target image as a target to be decoded. In addition, the target image may be an input image input to an encoding device or an input image input to a decoding device.
[0085] Hereinafter, the terms "image," "picture," "frame," and "screen" may be used to have the same meaning and may be used interchangeably with each other.
[0086] Hereinafter, a target block may be an encoding target block (i.e., a target to be encoded) and / or a decoding target block (i.e., a target to be decoded). In addition, a target block may be a current block, i.e., a target to be currently encoded and / or decoded. Herein, the terms "target block" and "current block" may be used to have the same meaning and may be used interchangeably.
[0087] Hereinafter, the terms "block" and "unit" may be used to have the same meaning and may be used interchangeably with each other. Alternatively, a "block" may refer to a specific unit.
[0088] Hereinafter, the terms "region" and "segment" are used interchangeably with each other.
[0089] Hereinafter, a specific signal may be a signal indicating a specific block. For example, an original signal may be a signal indicating a target block. A prediction signal may be a signal indicating a prediction block. A residual signal may be a signal indicating a residual block.
[0090] In the following embodiments, specific information, data, flags, elements, and attributes may have their own values. The value "0" corresponding to each of the information, data, flags, elements, and attributes may indicate a logical false value or a first predefined value. In other words, the values "0," false, logical false, and the first predefined value may be used interchangeably. The value "1" corresponding to each of the information, data, flags, elements, and attributes may indicate a logical true value or a second predefined value. In other words, the values "1," true, logical true, and the second predefined value may be used interchangeably.
[0091] When a variable such as i or j is used to indicate a row, column, or index, the value i may be an integer 0 or greater than 0, or may be an integer 1 or greater than 1. In other words, in an embodiment, each of the row, column, and index may be counted starting from 0, or may be counted starting from 1.
[0092] Hereinafter, terms to be used in the embodiments will be described.
[0093] Encoder: An encoder refers to a device used to perform encoding.
[0094] Decoder: A decoder refers to a device used to perform decoding.
[0095] Unit: A “unit” may refer to a unit of image encoding and decoding. The terms “unit” and “block” may be used to have the same meaning and may be used interchangeably with each other.
[0096] - A "cell" may be an M x N array of samples. M and N may be positive integers, respectively. The term "cell" may generally refer to a two-dimensional (2D) array of samples.
[0097] During image encoding and decoding, a "unit" may be a region generated by partitioning an image. In other words, a "unit" may be a designated region within an image. A single image may be partitioned into multiple units. Alternatively, an image may be partitioned into sub-parts, and a unit may represent each sub-part when encoding or decoding is performed on the partitioned sub-parts.
[0098] - During encoding and decoding of an image, predefined processing may be performed on each unit according to the type of the unit.
[0099] - According to the function, the unit type can be classified into a macro unit, a coding unit (CU), a prediction unit (PU), a residual unit, a transform unit (TU), etc. Alternatively, according to the function, the unit can refer to a block, a macro block, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc.
[0100] The term "unit" may mean information including a luma component block, a chroma component block corresponding to the luma component block, and syntax elements for the respective blocks such that the unit is designated to be distinguished from the block.
[0101] - The size and shape of the unit can be implemented differently. In addition, the unit can have any of a variety of sizes and shapes. Specifically, the shape of the unit can include not only a square, but also a geometric shape that can be represented in two dimensions (2D), such as a rectangle, a trapezoid, a triangle, and a pentagon.
[0102] In addition, the unit information may include one or more of a unit type, a unit size, a unit depth, a unit encoding order, a unit decoding order, etc. For example, the unit type may indicate one of CU, PU, residual unit, and TU.
[0103] - A cell may be partitioned into sub-cells, each sub-cell having a size smaller than that of the associated cell.
[0104] - Depth: Depth may indicate the degree to which a cell is partitioned. In addition, cell depth may indicate the level at which a corresponding cell exists when the cell is represented in a tree structure.
[0105] - The cell partition information may include a depth indicating the depth of the cell. The depth may indicate the number of times the cell is partitioned and / or the extent to which the cell is partitioned.
[0106] - In a tree structure, the root node can be considered to have the smallest depth and the leaf node can be considered to have the largest depth.
[0107] A single cell can be hierarchically partitioned into multiple sub-cells, with the sub-cell having depth information based on a tree structure. In other words, a cell and the sub-cells generated by partitioning the cell may correspond to a node and the sub-nodes of the node, respectively. Each partitioned sub-cell may have a cell depth. Since the depth indicates the number of times a cell has been partitioned and / or the extent to which a cell has been partitioned, the partition information of a sub-cell may include information about the size of the sub-cell.
[0108] In a tree structure, the top node may correspond to the initial node before partitioning. The top node may be referred to as a "root node." Furthermore, the root node may have the smallest depth value. Here, the depth of the top node may be level "0."
[0109] - Nodes at a depth level of "1" may represent cells generated when the original cell is partitioned once. Nodes at a depth level of "2" may represent cells generated when the original cell is partitioned twice.
[0110] - A leaf node at depth level "n" may represent a cell generated when the initial cell is partitioned n times.
[0111] A leaf node may be a bottom node that cannot be partitioned further. The depth of a leaf node may be a maximum level. For example, a predefined value for the maximum level may be 3.
[0112] -QT depth can represent the depth for four partitions. BT depth can represent the depth for two partitions. TT depth can represent the depth for three partitions.
[0113] - Sample: Sample can be the basic unit of a block. Available from 0 to 2 according to the bit depth (Bd) Bd-The value of 1 represents the sample point.
[0114] - A sample can be a pixel or a pixel value.
[0115] - Hereinafter, the terms "pixel" and "sample" may be used to have the same meaning and may be used interchangeably with each other.
[0116] Coding Tree Unit (CTU): A CTU may consist of a single luma component (Y) coding tree block and two chroma component (Cb, Cr) coding tree blocks associated with the luma component coding tree block. In addition, a CTU may represent information including the above blocks and syntax elements for each block.
[0117] - Each coding tree unit (CTU) can be partitioned using one or more partitioning methods, such as quadtree (QT), binary tree (BT), and ternary tree (TT), to configure sub-units such as coding units, prediction units, and transform units. In addition, each coding tree unit can be partitioned using multiple types of trees using one or more partitioning methods.
[0118] - "CTU" may be used as a term to designate a pixel block as a processing unit in image decoding and encoding processes (such as in the case of partitioning an input image).
[0119] Coding Tree Block (CTB): “CTB” may be used as a term designating any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.
[0120] Neighboring block: A neighboring block (or neighboring block) may refer to a block adjacent to the target block. A neighboring block may refer to a reconstructed neighboring block.
[0121] Hereinafter, the terms “neighboring block” and “adjacent block” may be used to have the same meaning and may be used interchangeably with each other.
[0122] Spatially neighboring blocks: Spatially neighboring blocks may be blocks that are spatially adjacent to the target block. Neighboring blocks may include spatially neighboring blocks.
[0123] -The target block and spatially neighboring blocks may be included in the target picture.
[0124] The spatially neighboring block may mean a block whose boundary contacts the target block or a block located within a predetermined distance from the target block.
[0125] - A spatially adjacent block may refer to a block adjacent to a vertex of the target block. Here, a block adjacent to a vertex of the target block may refer to a block vertically adjacent to a neighboring block horizontally adjacent to the target block or a block horizontally adjacent to a neighboring block vertically adjacent to the target block.
[0126] Temporally neighboring blocks: Temporally neighboring blocks may be blocks that are temporally adjacent to the target block. Neighboring blocks may include temporally neighboring blocks.
[0127] - Temporally neighboring blocks may include co-located blocks (col blocks).
[0128] The col block may be a block in a previously reconstructed co-located picture (col picture). The position of the col block in the col picture may correspond to the position of the target block in the target picture. Alternatively, the position of the col block in the col picture may be equal to the position of the target block in the target picture. The col picture may be a picture included in a reference picture list.
[0129] - The temporally neighboring block may be a block that is temporally adjacent to the spatially neighboring block of the target block.
[0130] Prediction unit: A prediction unit may be a basic unit for prediction such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation.
[0131] A single prediction unit can be divided into multiple partitions or sub-prediction units of smaller size. The multiple partitions can also be the basic units when performing prediction or compensation. Partitions generated by dividing the prediction unit can also be prediction units.
[0132] Prediction unit partition: A prediction unit partition may be a shape into which a prediction unit is divided.
[0133] Reconstructed neighboring cells: Reconstructed neighboring cells may be cells around the target cell that have been decoded and reconstructed.
[0134] The reconstructed neighboring cell may be a cell that is spatially adjacent to the target cell or temporally adjacent to the target cell.
[0135] The reconstructed spatially neighboring unit may be a unit included in the target picture that has been reconstructed through encoding and / or decoding.
[0136] The reconstructed temporally adjacent unit may be a unit included in a reference picture and already reconstructed through encoding and / or decoding. The position of the reconstructed temporally adjacent unit in the reference picture may be the same as the position of the target unit in the target picture, or may correspond to the position of the target unit in the target picture.
[0137] Parameter set: Parameter set can be header information in the structure of bitstream. For example, parameter set can include video parameter set, sequence parameter set, picture parameter set, adaptation parameter set, etc.
[0138] In addition, the parameter set may include slice header information and tile header information.
[0139] Rate-distortion optimization: The encoding device may use rate-distortion optimization to provide high encoding efficiency by utilizing a combination of the following: the size of the coding unit (CU), the prediction mode, the size of the prediction unit (PU), motion information, and the size of the transform unit (TU).
[0140] The rate-distortion optimization scheme can calculate the rate-distortion cost of each combination to select the optimal combination from these combinations. The rate-distortion cost can be calculated using the following equation 1. Generally, the combination that minimizes the rate-distortion cost can be selected as the optimal combination under the rate-distortion optimization scheme.
[0141] [Equation 1]
[0142] D+λ*R
[0143] -D may represent distortion. D may be the average of the squares of the differences between the original transform coefficients and the reconstructed transform coefficients in the transform unit (ie, the mean square error).
[0144] -R may represent the rate, which may use relevant context information to represent the bit rate.
[0145] -λ represents the Lagrange multiplier. R may include not only encoding parameter information such as prediction mode, motion information, and coding block flag, but also bits generated by encoding transform coefficients.
[0146] - The encoding device may perform processes such as inter-frame prediction and / or intra-frame prediction, transformation, quantization, entropy coding, inverse quantization (dequantization), and inverse transformation in order to calculate accurate D and R. These processes may greatly increase the complexity of the encoding device.
[0147] - Bitstream: A bitstream may refer to a stream of bits including encoded image information.
[0148] -Parameter set: Parameter set can be header information in the structure of the bitstream.
[0149] The parameter set may include at least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptation parameter set. In addition, the parameter set may include information about a slice header and information about a tile header.
[0150] Parsing: Parsing can be the decision of the value of the syntax element by performing entropy decoding on the bitstream. Alternatively, the term "parsing" can refer to such entropy decoding itself.
[0151] Symbol: A symbol may be at least one of a syntax element, a coding parameter, and a transform coefficient of a coding target unit and / or a decoding target unit. In addition, a symbol may be a target of entropy coding or a result of entropy decoding.
[0152] Reference picture: A reference picture may be an image referenced by a unit in order to perform inter-frame prediction or motion compensation. Alternatively, a reference picture may be an image including a reference unit referenced by a target unit in order to perform inter-frame prediction or motion compensation.
[0153] Hereinafter, the terms “reference picture” and “reference image” may be used to have the same meaning and may be used interchangeably with each other.
[0154] Reference picture list: A reference picture list may be a list including one or more reference images used for inter prediction or motion compensation.
[0155] - The type of reference picture list may include merged list (LC), list 0 (L0), list 1 (L1), list 2 (L3), list 3 (L3), etc.
[0156] - For inter prediction, one or more reference picture lists may be used.
[0157] Inter-frame prediction indicator: The inter-frame prediction indicator may indicate the inter-frame prediction direction for the target unit. Inter-frame prediction can be either unidirectional or bidirectional. Optionally, the inter-frame prediction indicator may indicate the number of reference pictures used to generate the prediction unit for the target unit. Optionally, the inter-frame prediction indicator may indicate the number of prediction blocks used for inter-frame prediction or motion compensation for the target unit.
[0158] Reference picture index: The reference picture index may be an index indicating a specific reference picture in a reference picture list.
[0159] Motion Vector (MV): A motion vector is a 2D vector used for inter-frame prediction or motion compensation. A motion vector represents the offset between a target image and a reference image.
[0160] - For example, you can use a command such as (mv x , mv y ) to express MV. mv x Can indicate the horizontal component, mv y May indicate the vertical component.
[0161] -Search range: The search range may be a 2D area where a search for an MV is performed during inter prediction. For example, the size of the search range may be M×N. M and N may be positive integers, respectively.
[0162] Motion vector candidate: A motion vector candidate may be a block that is a prediction candidate when a motion vector is predicted or a motion vector of a block that is a prediction candidate.
[0163] - The motion vector candidate may be included in the motion vector candidate list.
[0164] Motion vector candidate list: A motion vector candidate list may be a list configured using one or more motion vector candidates.
[0165] Motion vector candidate index: The motion vector candidate index may be an indicator for indicating a motion vector candidate in the motion vector candidate list. Alternatively, the motion vector candidate index may be an index of a motion vector predictor.
[0166] Motion information: The motion information may be information including at least one of a reference picture list, a reference image, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index, as well as a motion vector, a reference picture index, and an inter prediction indicator.
[0167] Merge candidate list: The merge candidate list may be a list configured using merge candidates.
[0168] Merge candidate: A merge candidate may be a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-predictive merge candidate, a zero merge candidate, etc. The merge candidate may include motion information such as prediction type information, reference picture index for each list, and a motion vector.
[0169] Merge index: The merge index may be an indicator for indicating a merge candidate in a merge candidate list.
[0170] The merge index may indicate a reconstructed unit for deriving a merge candidate between a reconstructed unit spatially adjacent to the target unit and a reconstructed unit temporally adjacent to the target unit.
[0171] - The merge index may indicate at least one of a plurality of pieces of motion information of a merge candidate.
[0172] Transformation unit: A transformation unit can be a basic unit for residual signal encoding and / or residual signal decoding (such as transformation, inverse transformation, quantization, inverse quantization, transformation coefficient encoding, and transformation coefficient decoding). A single transformation unit can be partitioned into multiple transformation units with smaller sizes.
[0173] Scaling: Scaling can be referred to as the process of multiplying the levels of the transform coefficients by a factor.
[0174] - As a result of scaling the transform coefficient levels, transform coefficients may be generated. Scaling may also be referred to as "inverse quantization."
[0175] Quantization Parameter (QP): A quantization parameter may be a value used to generate transform coefficient levels for a transform coefficient during quantization. Alternatively, a quantization parameter may also be a value used to generate a transform coefficient by scaling the transform coefficient levels during inverse quantization. Alternatively, a quantization parameter may be a value mapped to a quantization step size.
[0176] Delta quantization parameter: The delta quantization parameter is the difference between the target unit's quantization parameter and the predicted quantization parameter.
[0177] Scanning: Scanning can refer to a method of arranging the order of coefficients in a cell, block, or matrix. For example, a method for arranging a 2D array in the form of a one-dimensional (1D) array can be referred to as "scanning." Alternatively, a method for arranging a 1D array in the form of a 2D array can also be referred to as "scanning" or "inverse scanning."
[0178] Transform coefficient: The transform coefficient may be a coefficient value generated when the encoding device performs transformation. Alternatively, the transform coefficient may be a coefficient value generated when the decoding device performs at least one of entropy decoding and inverse quantization.
[0179] - A quantized level or a quantized transform coefficient level generated by applying quantization to a transform coefficient or a residual signal may also be included in the meaning of the term "transform coefficient".
[0180] Quantization level: The quantization level may be a value generated when the encoding device performs quantization on the transform coefficient or residual signal. Alternatively, the quantization level may be a value that is a target of inverse quantization when the decoding device performs inverse quantization.
[0181] - A quantized transform coefficient level as a result of transformation and quantization may also be included in the meaning of the quantization level.
[0182] Non-zero transform coefficient: A non-zero transform coefficient may be a transform coefficient having a value other than 0, or may be a transform coefficient level having a value other than 0. Alternatively, a non-zero transform coefficient may be a transform coefficient having a value with a magnitude other than 0, or may be a transform coefficient level having a value with a magnitude other than 0.
[0183] Quantization Matrix: A quantization matrix may be a matrix used in a quantization process or an inverse quantization process to improve the subjective or objective image quality of an image. A quantization matrix may also be referred to as a "scaling list."
[0184] Quantization matrix coefficient: A quantization matrix coefficient can be each element in the quantization matrix. A quantization matrix coefficient can also be called a "matrix coefficient."
[0185] Default matrix: The default matrix may be a quantization matrix predefined by the encoding device and the decoding device.
[0186] Non-default matrix: A non-default matrix may be a quantization matrix that is not pre-defined by the encoding device and the decoding device. The non-default matrix may be signaled by the encoding device to the decoding device.
[0187] Most Probable Mode (MPM): The MPM may represent an intra prediction mode that is highly likely to be used for intra prediction for a target block.
[0188] The encoding apparatus and the decoding apparatus may determine one or more MPMs based on encoding parameters related to the target block and properties of an entity related to the target block.
[0189] The encoding device and the decoding device may determine one or more MPMs based on the intra-frame prediction mode of the reference block. The reference block may include multiple reference blocks. The multiple reference blocks may include a spatially neighboring block adjacent to the left of the target block and a spatially neighboring block adjacent to the top of the target block. In other words, depending on which intra-frame prediction mode has been used for the reference block, one or more different MPMs may be determined.
[0190] One or more MPMs may be determined in the same manner in both the encoding device and the decoding device. That is, the encoding device and the decoding device may share the same MPM list including one or more MPMs.
[0191] MPM list: The MPM list may be a list including one or more MPMs. The number of the one or more MPMs in the MPM list may be predefined.
[0192] MPM indicator: The MPM indicator may indicate an MPM to be used for intra prediction for a target block among one or more MPMs in an MPM list. For example, the MPM indicator may be an index for the MPM list.
[0193] Since the MPM list is determined in the same manner in both the encoding device and the decoding device, there may be no need to transmit the MPM list itself from the encoding device to the decoding device.
[0194] The MPM indicator may be signaled from the encoding apparatus to the decoding apparatus.Since the MPM indicator is signaled, the decoding apparatus may determine an MPM to be used for intra prediction for a target block among the MPMs in the MPM list.
[0195] MPM usage indicator: The MPM usage indicator may indicate whether an MPM usage mode is to be used for prediction for a target block. The MPM usage mode may be a mode of determining an MPM to be used for intra prediction for a target block using an MPM list.
[0196] The MPM usage indicator may be signaled from the encoding device to the decoding device.
[0197] Signaling: "Signaling" may mean that information is sent from an encoding device to a decoding device. Alternatively, "signaling" may mean that information is included in a bitstream or storage medium. Information signaled by an encoding device can be used by a decoding device.
[0198] Figure 1 is a block diagram showing a configuration of an embodiment of an encoding device to which the present disclosure is applied.
[0199] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images (pictures). The encoding device 100 may sequentially encode one or more images of a video.
[0200] Reference Figure 1 , the encoding device 100 includes an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization (dequantization) unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.
[0201] The encoding apparatus 100 may perform encoding on a target image using an intra mode and / or an inter mode.
[0202] In addition, the encoding device 100 can generate a bitstream including information about encoding by encoding the target image, and can output the generated bitstream. The generated bitstream can be stored in a computer-readable storage medium and can be streamed through a wireless / wired transmission medium.
[0203] When the intra mode is used as the prediction mode, the switch 115 may switch to the intra mode. When the inter mode is used as the prediction mode, the switch 115 may switch to the inter mode.
[0204] The encoding apparatus 100 may generate a prediction block of the target block. In addition, after having generated the prediction block, the encoding apparatus 100 may encode a residual between the target block and the prediction block.
[0205] When the prediction mode is intra mode, the intra prediction unit 120 may use pixels of a previously encoded / decoded neighboring block around the target block as reference samples. The intra prediction unit 120 may perform spatial prediction on the target block using the reference samples and may generate prediction samples for the target block through spatial prediction.
[0206] The inter prediction unit 110 may include a motion prediction unit and a motion compensation unit.
[0207] When the prediction mode is the inter mode, the motion prediction unit may search the reference image for a region that best matches the target block during the motion prediction process, and may derive a motion vector for the target block and the found region based on the found region.
[0208] The reference image may be stored in the reference picture buffer 190. More specifically, when encoding and / or decoding of the reference image has been processed, the reference image may be stored in the reference picture buffer 190.
[0209] The motion compensation unit can generate a prediction block for the target block by performing motion compensation using a motion vector. Here, the motion vector can be a two-dimensional (2D) vector used for inter-frame prediction. In addition, the motion vector can represent the offset between the target image and the reference image.
[0210] When the motion vector has a value other than an integer, the motion prediction unit and the motion compensation unit may generate a prediction block by applying an interpolation filter to a partial area of the reference image. In order to perform inter-frame prediction or motion compensation, it may be determined which mode among the skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to a method for predicting and compensating the motion of the PU included in the CU based on the CU, and inter-frame prediction or motion compensation may be performed according to the mode.
[0211] The subtractor 125 may generate a residual block, which is the difference between the target block and the prediction block. The residual block may also be referred to as a "residual signal."
[0212] The residual signal may be the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming or quantizing the difference between the original signal and the predicted signal, or a signal generated by transforming and quantizing the difference. The residual block may be a residual signal for a block unit.
[0213] The transform unit 130 may generate a transform coefficient by transforming the residual block, and may output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by transforming the residual block.
[0214] The transform unit 130 may use one of a plurality of predefined transform methods when performing the transform.
[0215] The plurality of predefined transform methods may include discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loeve transform (KLT), and the like.
[0216] The transform method used to transform the residual block may be determined based on at least one of the encoding parameters for the target block and / or the neighboring blocks. For example, the transform method may be determined based on at least one of the inter-prediction mode for the PU, the intra-prediction mode for the PU, the size of the TU, and the shape of the TU. Alternatively, transform information indicating the transform method may be signaled from the encoding device 100 to the decoding device 200.
[0217] When the transform skip mode is used, the transform unit 130 may omit the operation of transforming the residual block.
[0218] By performing quantization on the transform coefficients, quantized transform coefficient levels or quantized levels may be generated. Hereinafter, in an embodiment, each of the quantized transform coefficient levels and the quantized levels may also be referred to as a 'transform coefficient'.
[0219] The quantization unit 140 may generate a quantized transform coefficient level or a quantized level by quantizing the transform coefficient according to the quantization parameter. The quantization unit 140 may output the generated quantized transform coefficient level or the quantized level. In this case, the quantization unit 140 may quantize the transform coefficient using a quantization matrix.
[0220] The entropy encoding unit 150 may generate a bitstream by performing entropy encoding based on a probability distribution based on a value calculated by the quantization unit 140 and / or an encoding parameter value calculated during encoding. The entropy encoding unit 150 may output the generated bitstream.
[0221] The entropy encoding unit 150 may perform entropy encoding on information about pixels of an image and information required for decoding the image. For example, the information required for decoding the image may include syntax elements and the like.
[0222] When entropy coding is applied, fewer bits are allocated to symbols that appear more frequently, and more bits are allocated to symbols that appear less frequently. Since symbols are represented by this allocation, the size of the bit string used to encode the target symbol can be reduced. Therefore, entropy coding can improve the compression performance of video encoding.
[0223] In addition, to perform entropy coding, the entropy coding unit 150 may use a coding method such as exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 may use a variable length coding / code (VLC) table to perform entropy coding. For example, the entropy coding unit 150 may derive a binarization method for the target symbol. In addition, the entropy coding unit 150 may derive a probability model for the target symbol / bin. The entropy coding unit 150 may use the derived binarization method, probability model, and context model to perform arithmetic coding.
[0224] The entropy encoding unit 150 may transform coefficients in a 2D block form into a 1D vector form through a transform coefficient scanning method in order to encode quantized transform coefficient levels.
[0225] Coding parameters may be information required for encoding and / or decoding. Coding parameters may include information encoded by the encoding device 100 and transmitted from the encoding device 100 to the decoding device, and may also include information that can be derived during the encoding or decoding process. For example, the information transmitted to the decoding device may include syntax elements.
[0226] The coding parameters may include not only information (or flags or indices) such as syntax elements that are encoded by the encoding device and transmitted by the encoding device to the decoding device using a signal, but also information derived in the encoding or decoding process. In addition, the coding parameters may include information required for encoding or decoding an image. For example, the coding parameters may include at least one value of the following items, a combination of the following items, or statistics: the size of the unit / block, the depth of the unit / block, the partition information of the unit / block, the partition structure of the unit / block, information indicating whether the unit / block is partitioned in a quadtree structure, information indicating whether the unit / block is partitioned in a binary tree structure, the partition direction of the binary tree structure (horizontal or vertical), the partition form of the binary tree structure (symmetric partitioning or asymmetric partitioning), information indicating whether the unit / block is partitioned in a ternary tree structure, the partition direction of the ternary tree structure (horizontal or vertical), the partition form of the ternary tree structure (symmetric partitioning or asymmetric partitioning, etc.), an indication of the unit / Information on whether the block is partitioned in a composite tree structure, combination and direction of partitions of the composite tree structure (horizontally or vertically, etc.), prediction scheme (intra-frame prediction or inter-frame prediction), intra-frame prediction mode / direction, reference sample filtering method, prediction block filtering method, prediction block boundary filtering method, filter taps for filtering, filter coefficients for filtering, inter-frame prediction mode, motion information, motion vector, reference picture index, inter-frame prediction direction, inter-frame prediction indicator, reference picture list, reference image, motion vector predictor, motion vector prediction candidate, motion vector candidate list, information indicating whether merge mode is used, merge candidate, merge candidate list, indication Information on whether skip mode is used, type of interpolation filter, taps of interpolation filter, filter coefficients of interpolation filter, size of motion vector, accuracy of motion vector representation, transform type, transform size, information indicating whether primary transform is used, information indicating whether additional (secondary) transform is used, first transform selection information (or first transform index), second transform selection information (or second transform index), information indicating the presence or absence of residual signal, coding block pattern, coding block flag, quantization parameter, quantization matrix, information on in-loop filter, information indicating whether in-loop filter is applied, coefficients of in-loop filter, taps of in-loop filter, shape / form of the in-loop filter, information indicating whether a deblocking filter is applied, coefficients of the deblocking filter, taps of the deblocking filter, deblocking filter strength, shape / form of the deblocking filter, information indicating whether an adaptive sample offset is applied, value of the adaptive sample offset, category of the adaptive sample offset, type of the adaptive sample offset, information indicating whether an adaptive loop filter is applied, coefficients of the adaptive loop filter, taps of the adaptive loop filter, shape / form of the adaptive loop filter, binarization / debinarization method, context model, context model determination method, context model update method, information indicating whether a normal mode is performed,Information indicating whether bypass mode is performed, context binary bit, bypass binary bit, transform coefficient, transform coefficient level, transform coefficient level scanning method, image display / output order, slice identification information, slice type, slice partition information, tile identification information, tile type, tile partition information, picture type, bit depth, information about luma signal, and information about chroma signal. The prediction scheme may indicate one of intra prediction mode and inter prediction mode.
[0227] The first transform selection information may indicate a first transform applied to the target block.
[0228] The secondary transform selection information may indicate a secondary transform applied to the target block.
[0229] The residual signal may represent the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a residual signal for a block.
[0230] Here, signaling a flag or index may indicate that the encoding device 100 includes an entropy-coded flag or an entropy-coded index generated by performing entropy encoding on the flag or index in the bitstream, and may indicate that the decoding device 200 obtains the flag or index by performing entropy decoding on the entropy-coded flag or entropy-coded index extracted from the bitstream.
[0231] Since the encoding apparatus 100 performs encoding via inter-frame prediction, the encoded target image can be used as a reference image for another image to be subsequently processed. Therefore, the encoding apparatus 100 can reconstruct or decode the encoded target image and store the reconstructed or decoded image as a reference image in the reference picture buffer 190. For decoding, the encoded target image can be dequantized and inversely transformed.
[0232] The quantization level may be dequantized by the dequantization unit 160 and inversely transformed by the inverse transform unit 170. The dequantized and / or inversely transformed coefficients may be added to the prediction block by the adder 175. The dequantized and / or inversely transformed coefficients and the prediction block are added to generate a reconstructed block. Here, the dequantized and / or inversely transformed coefficients may represent coefficients on which one or more of dequantization and inverse transformation have been performed, and may also represent a reconstructed residual block.
[0233] The reconstructed block may be filtered by the filter unit 180. The filter unit 180 may apply one or more filters of a deblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), and a non-local filter (NLF) to the reconstructed block or the reconstructed picture. The filter unit 180 may also be referred to as a "loop filter."
[0234] The deblocking filter may remove block distortion occurring at boundaries between blocks. To determine whether to apply the deblocking filter, the number of columns or rows of pixels included in the block and comprising pixels on which to determine whether to apply the deblocking filter to the target block may be determined.
[0235] When a deblocking filter is applied to a target block, the filter applied may differ depending on the desired strength of the deblocking filter. In other words, among different filters, a filter determined in consideration of the strength of the deblocking filter may be applied to the target block. When a deblocking filter is applied to a target block, a filter corresponding to either a strong filter or a weak filter may be applied to the target block depending on the desired strength of the deblocking filter.
[0236] Also, when vertical filtering and horizontal filtering are performed on a target block, the horizontal filtering and the vertical filtering may be performed in parallel.
[0237] SAO can add an appropriate offset to pixel values to compensate for coding errors. SAO can perform pixel-by-pixel correction on an image to which deblocking has been applied, wherein the correction uses an offset that is the difference between the original image and the image to which deblocking has been applied. To perform offset correction on an image, a method can be used for dividing the pixels included in the image into a specific number of regions, determining the regions to which the offset is applied within the divided regions, and applying the offset to the determined regions. A method can also be used for applying the offset while taking into account edge information for each pixel.
[0238] The ALF can perform filtering based on values obtained by comparing the reconstructed image with the original image. After the pixels included in the image have been divided into a predetermined number of groups, the filter to be applied to each group can be determined, and filtering can be performed differently for each group. For luma signals, information related to whether an adaptive loop filter is applied can be signaled for each CU. The shape and filter coefficients of the ALF to be applied to each block can be different for each block. Alternatively, a fixed ALF can be applied to the block regardless of its characteristics.
[0239] A non-local filter can perform filtering based on a reconstructed block similar to a target block. A region similar to the target block can be selected from the reconstructed image, and the target block can be filtered using statistical properties of the selected similar region. Information on whether a non-local filter is applied can be signaled for each coding unit (CU). Furthermore, the shape and filter coefficients of the non-local filter applied to a block can vary depending on the block.
[0240] The reconstructed block or reconstructed image filtered by the filter unit 180 may be stored in the reference picture buffer 190. The reconstructed block filtered by the filter unit 180 may be part of a reference picture. In other words, the reference picture may be a reconstructed picture composed of the reconstructed blocks filtered by the filter unit 180. The stored reference picture may then be used for inter-frame prediction.
[0241] Figure 2 is a block diagram showing a configuration of an embodiment of a decoding device to which the present disclosure is applied.
[0242] The decoding device 200 may be a decoder, a video decoding device, or an image decoding device.
[0243] Reference Figure 2 , the decoding device 200 may include an entropy decoding unit 210, an inverse quantization (dequantization) unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, an inter-frame prediction unit 250, a switch 245, an adder 255, a filter unit 260 and a reference picture buffer 270.
[0244] The decoding apparatus 200 may receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 may receive a bitstream stored in a computer-readable storage medium, and may receive a bitstream streamed through a wired / wireless transmission medium.
[0245] The decoding apparatus 200 may perform decoding on a bitstream in an intra mode and / or an inter mode. In addition, the decoding apparatus 200 may generate a reconstructed image or a decoded image through decoding, and may output the reconstructed image or the decoded image.
[0246] For example, an operation of switching to intra mode or inter mode based on the prediction mode used for decoding may be performed by the switch 245. When the prediction mode used for decoding is intra mode, the switch 245 may be operated to switch to intra mode. When the prediction mode used for decoding is inter mode, the switch 245 may be operated to switch to inter mode.
[0247] The decoding device 200 can obtain a reconstructed residual block by decoding the input bit stream and can generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as a decoding target by adding the reconstructed residual block to the prediction block.
[0248] The entropy decoding unit 210 may generate symbols by performing entropy decoding on the bitstream based on the probability distribution of the bitstream. The generated symbols may include quantized transform coefficient level format symbols. Here, the entropy decoding method may be similar to the entropy encoding method described above. That is, the entropy decoding method may be the inverse process of the entropy encoding method described above.
[0249] The entropy decoding unit 210 may change a coefficient having a one-dimensional (1D) vector form into a 2D block shape through a transform coefficient scanning method in order to decode quantized transform coefficient levels.
[0250] For example, the coefficients of the block can be changed to a 2D block shape by scanning the block coefficients using an upper right diagonal scan. Optionally, which of the upper right diagonal scan, vertical scan, and horizontal scan to use can be determined based on the size of the corresponding block and / or the intra prediction mode.
[0251] The quantized coefficients may be dequantized by the dequantization unit 220. The dequantization unit 220 may generate dequantized coefficients by performing dequantization on the quantized coefficients. Furthermore, the dequantized coefficients may be inversely transformed by the inverse transform unit 230. The inverse transform unit 230 may generate a reconstructed residual block by performing inverse transform on the dequantized coefficients. As a result of performing dequantization and inverse transform on the quantized coefficients, a reconstructed residual block may be generated. Here, when generating the reconstructed residual block, the dequantization unit 220 may apply a quantization matrix to the quantized coefficients.
[0252] When the intra mode is used, the intra prediction unit 240 may generate a prediction block by performing spatial prediction using pixel values of previously decoded neighboring blocks around a target block.
[0253] The inter-frame prediction unit 250 may include a motion compensation unit. Alternatively, the inter-frame prediction unit 250 may be designated as a "motion compensation unit."
[0254] When the inter mode is used, the motion compensation unit 250 may generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer 270 .
[0255] The motion compensation unit may apply an interpolation filter to a partial area of a reference image when a motion vector has a value other than an integer, and may generate a prediction block using the reference image to which the interpolation filter is applied. In order to perform motion compensation, the motion compensation unit may determine which mode of skip mode, merge mode, advanced motion vector prediction (AMVP) mode, and current picture reference mode corresponds to a motion compensation method for a PU included in the CU based on the CU, and may perform motion compensation according to the determined mode.
[0256] The reconstructed residual block and the prediction block may be added to each other by the adder 255. The adder 255 may generate a reconstructed block by adding the reconstructed residual block and the prediction block.
[0257] The reconstructed block may be filtered by the filter unit 260. The filter unit 260 may apply at least one of a deblocking filter, an SAO filter, an ALF, and an NLF to the reconstructed block or the reconstructed image. The reconstructed image may be a picture including the reconstructed block.
[0258] The filtered reconstructed image may be output by the encoding apparatus 100 and may be used by the encoding apparatus.
[0259] The reconstructed image filtered by the filter unit 260 may be stored as a reference picture in the reference picture buffer 270. The reconstructed block filtered by the filter unit 260 may be part of the reference picture. In other words, the reference picture may be an image composed of the reconstructed blocks filtered by the filter unit 260. The stored reference picture may then be used for inter-frame prediction.
[0260] Figure 3 is a diagram schematically showing a partition structure of an image when the image is encoded and decoded.
[0261] Figure 3 An example in which a single unit is partitioned into a plurality of subunits may be schematically shown.
[0262] To efficiently partition an image, coding units (CUs) may be used in encoding and decoding. The term "unit" may be used to collectively designate 1) a block containing image samples and 2) a syntax element. For example, "partition of a unit" may refer to "partition of blocks corresponding to the unit."
[0263] A CU may be used as a basic unit for image encoding / decoding. A CU may be used as a unit to which one mode selected from intra mode and inter mode is applied in image encoding / decoding. In other words, in image encoding / decoding, it may be determined which mode of intra mode and inter mode is to be applied to each CU.
[0264] Also, a CU may be a basic unit for predicting, transforming, quantizing, inversely transforming, dequantizing, and encoding / decoding a transform coefficient.
[0265] Reference Figure 3 , the image 300 may be sequentially partitioned into units corresponding to the largest coding unit (LCU), and a partition structure may be determined for each LCU. Here, LCU may be used to have the same meaning as a coding tree unit (CTU).
[0266] Partitioning a unit may mean partitioning the block corresponding to the unit. Block partition information may include depth information regarding the depth of the unit. The depth information may indicate the number of times the unit is partitioned and / or the degree to which the unit is partitioned. A single unit may be hierarchically partitioned into sub-units, with the single unit having depth information based on a tree structure. Each partitioned sub-unit may have depth information. The depth information may be information indicating the size of a CU. The depth information may be stored for each CU.
[0267] Each CU may have depth information. When a CU is partitioned, the depth of the CU generated from the partition may increase by 1 from the depth of the partitioned CU.
[0268] The partition structure may indicate the distribution of coding units (CUs) in the LCU 310 for efficiently encoding an image. This distribution may be determined based on whether a single CU is to be partitioned into multiple CUs. The number of CUs generated by partitioning may be a positive integer of 2 or greater, including 2, 3, 4, 8, 16, and the like. Depending on the number of CUs generated by partitioning, the horizontal and vertical sizes of each CU generated by partitioning may be smaller than those of the CU before partitioning.
[0269] Each partitioned CU may be recursively partitioned into four CUs in the same manner. Compared to at least one of the horizontal size and the vertical size of the CU before partitioning, at least one of the horizontal size and the vertical size of each partitioned CU may be reduced through recursive partitioning.
[0270] Partitioning of a CU may be recursively performed up to a predefined depth or a predefined size. For example, the depth of a CU may have a value ranging from 0 to 3. The size of a CU may range from 64×64 to 8×8 depending on the depth of the CU.
[0271] For example, the depth of the LCU may be 0, and the depth of the minimum coding unit (SCU) may be a predefined maximum depth. Here, as described above, the LCU may be a CU with a maximum coding unit size, and the SCU may be a CU with a minimum coding unit size.
[0272] Partitioning may begin at the LCU 310, and each time the horizontal and / or vertical dimensions of the CU are reduced by partitioning, the depth of the CU may increase by one.
[0273] For example, for each depth, a non-partitioned CU may have a size of 2N×2N. In addition, when a CU is partitioned, a CU of size 2N×2N may be partitioned into four CUs of size N×N. Whenever the depth increases by 1, the value of N may be halved.
[0274] Reference Figure 3 , an LCU with a depth of 0 may have 64×64 pixels or a 64×64 block. 0 may be the minimum depth. An SCU with a depth of 3 may have 8×8 pixels or an 8×8 block. 3 may be the maximum depth. Here, a CU with a 64×64 block as an LCU may be represented by a depth of 0. A CU with a 32×32 block may be represented by a depth of 1. A CU with a 16×16 block may be represented by a depth of 2. A CU with an 8×8 block as an SCU may be represented by a depth of 3.
[0275] Information about whether a corresponding CU is partitioned can be represented by the CU's partition information. The partition information can be 1-bit information. All CUs except the SCU can include partition information. For example, the partition information value of a non-partitioned CU can be 0. The partition information value of a partitioned CU can be 1.
[0276] For example, when a single CU is partitioned into four CUs, the horizontal size and vertical size of each of the four CUs generated by partitioning can be half the horizontal size and vertical size of the CU before partitioning. When a CU of size 32×32 is partitioned into four CUs, the size of each of the four partitioned CUs can be 16×16. When a single CU is partitioned into four CUs, it can be considered that the CU has been partitioned in a quadtree structure.
[0277] For example, when a single CU is partitioned into two CUs, the horizontal size or vertical size of each of the two CUs generated by partitioning may be half the horizontal size or vertical size of the CU before partitioning. When a CU of size 32×32 is partitioned vertically into two CUs, the size of each of the two partitioned CUs may be 16×32. When a CU of size 32×32 is partitioned horizontally into two CUs, the size of each of the two partitioned CUs may be 32×16. When a single CU is partitioned into two CUs, it can be considered that the CU has been partitioned in a binary tree structure.
[0278] Both quadtree partitioning and binary tree partitioning are applied to Figure 3 LCU310.
[0279] In the encoding apparatus 100, a coding tree unit (CTU) of size 64×64 may be partitioned into multiple smaller CUs using a recursive quadtree structure. A single CU may be partitioned into four CUs of the same size. Each CU may be recursively partitioned and may have a quadtree structure.
[0280] By recursive partitioning of CUs, the optimal partitioning method that incurs the minimum rate-distortion cost can be selected.
[0281] Figure 4 is a diagram illustrating a form of prediction units (PUs) that a coding unit (CU) can include.
[0282] In a CU partitioned from an LCU, the CU that is no longer partitioned can be divided into one or more prediction units (PUs). This division is also called "partitioning."
[0283] PU can be a basic unit for prediction. PU can be encoded and decoded in any one of skip mode, inter mode and intra mode. PU can be partitioned into various shapes according to each mode. For example, Figure 1 The target block described above refers to Figure 2 The target blocks described may all be PUs.
[0284] A CU may not be split into PUs. When a CU is not split into PUs, the size of the CU and the size of the PU may be equal to each other.
[0285] In skip mode, partitioning may not exist in a CU.In skip mode, a 2Nx2N mode 410 may be supported without partitioning, wherein in the 2Nx2N mode 410, the size of the PU and the size of the CU are the same as each other.
[0286] In inter mode, eight types of partition shapes may exist in a CU. For example, in inter mode, 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445 may be supported.
[0287] In intra mode, 2N×2N mode 410 and N×N mode 425 may be supported.
[0288] In 2N×2N mode 410, a PU of size 2N×2N may be encoded. A PU of size 2N×2N may represent a PU of the same size as a CU. For example, a PU of size 2N×2N may have a size of 64×64, 32×32, 16×16, or 8×8.
[0289] In NxN mode 425, PUs of size NxN may be encoded.
[0290] For example, in intra prediction, when the size of a PU is 8×8, four partitioned PUs may be encoded, and the size of each partitioned PU may be 4×4.
[0291] When encoding a PU in intra mode, any one of multiple intra prediction modes may be used to encode the PU. For example, HEVC technology provides 35 intra prediction modes, and a PU may be encoded in any of the 35 intra prediction modes.
[0292] Which mode of the 2Nx2N mode 410 and the NxN mode 425 is to be used to encode the PU may be determined based on the rate-distortion penalty.
[0293] The encoding device 100 may perform an encoding operation on a PU of size 2N×2N. Here, the encoding operation may be an operation of encoding the PU in each of a plurality of intra-prediction modes that can be used by the encoding device 100. Through the encoding operation, an optimal intra-prediction mode for the PU of size 2N×2N may be derived. The optimal intra-prediction mode may be an intra-prediction mode that, among the plurality of intra-prediction modes that can be used by the encoding device 100, results in a minimum rate-distortion cost when encoding the PU of size 2N×2N.
[0294] In addition, the encoding device 100 may sequentially perform encoding operations on each PU obtained by performing N×N partitioning. Here, the encoding operation may be an operation of encoding the PU in each of a plurality of intra-frame prediction modes that can be used by the encoding device 100. Through the encoding operation, the optimal intra-frame prediction mode for the PU of size N×N may be derived. The optimal intra-frame prediction mode may be the intra-frame prediction mode that produces the minimum rate-distortion cost when encoding the PU of size N×N among the plurality of intra-frame prediction modes that can be used by the encoding device 100.
[0295] The encoding apparatus 100 may determine which of the PU of size 2N×2N and the PU of size N×N to be encoded based on a comparison between the rate-distortion cost of the PU of size 2N×2N and the rate-distortion cost of the PU of size N×N.
[0296] A single CU may be partitioned into one or more PUs, and a PU may be partitioned into multiple PUs.
[0297] For example, when a single PU is partitioned into four PUs, the horizontal and vertical sizes of each of the four PUs generated by partitioning can be half the horizontal and vertical sizes of the PU before partitioning. When a PU of size 32×32 is partitioned into four PUs, the size of each of the four partitioned PUs can be 16×16. When a single PU is partitioned into four PUs, it can be considered that the PU has been partitioned in a quadtree structure.
[0298] For example, when a single PU is partitioned into two PUs, the horizontal size or vertical size of each of the two PUs generated by partitioning may be half the horizontal size or vertical size of the PU before partitioning. When a PU of size 32×32 is partitioned vertically into two PUs, the size of each of the two partitioned PUs may be 16×32. When a PU of size 32×32 is partitioned horizontally into two PUs, the size of each of the two partitioned PUs may be 32×16. When a single PU is partitioned into two PUs, the PU may be considered to have been partitioned in a binary tree structure.
[0299] Figure 5 is a diagram illustrating a form of a transform unit (TU) that can be included in a CU.
[0300] A transform unit (TU) may be a basic unit used for processes such as transform, quantization, inverse transform, inverse quantization, entropy encoding, and entropy decoding in a CU.
[0301] A TU may have a square shape or a rectangular shape. The shape of a TU may be determined based on the size and / or shape of a CU.
[0302] Among the CUs partitioned from the LCU, the CUs that are no longer partitioned into CUs may be partitioned into one or more TUs. Here, the partition structure of the TU may be a quadtree structure. For example, Figure 5 As shown in , a single CU 510 can be partitioned one or more times according to a quadtree structure. Through this partitioning, a single CU 510 can be composed of TUs of various sizes.
[0303] It can be considered that a single CU is recursively split when it is split two or more times. Through splitting, a single CU can be composed of transform units (TUs) of various sizes.
[0304] Alternatively, a single CU may be split into one or more TUs based on the number of vertical lines and / or horizontal lines that partition the CU.
[0305] A CU may be divided into symmetric TUs or asymmetric TUs. To divide into asymmetric TUs, information about the size and / or shape of each TU may be signaled from the encoding apparatus 100 to the decoding apparatus 200. Alternatively, the size and / or shape of each TU may be derived from the information about the size and / or shape of the CU.
[0306] A CU may not be divided into TUs. When a CU is not divided into TUs, the size of the CU and the size of the TU may be equal to each other.
[0307] A single CU may be partitioned into one or more TUs, and a TU may be partitioned into multiple TUs.
[0308] For example, when a single TU is partitioned into four TUs, the horizontal size and vertical size of each of the four TUs generated by partitioning can be half the horizontal size and vertical size of the TU before partitioning. When a 32×32 TU is partitioned into four TUs, the size of each of the four partitioned TUs can be 16×16. When a single TU is partitioned into four TUs, the TU can be considered to have been partitioned in a quadtree structure.
[0309] For example, when a single TU is partitioned into two TUs, the horizontal size or vertical size of each of the two TUs generated by partitioning may be half the horizontal size or vertical size of the TU before partitioning. When a 32×32 TU is partitioned vertically into two TUs, the size of each of the two partitioned TUs may be 16×32. When a 32×32 TU is partitioned horizontally into two TUs, the size of each of the two partitioned TUs may be 32×16. When a single TU is partitioned into two TUs, the TU may be considered to have been partitioned in a binary tree structure.
[0310] Can be used with Figure 5 The CU is divided in different ways as shown in FIG.
[0311] For example, a single CU may be split into three CUs, and the horizontal sizes or vertical sizes of the three CUs generated by the splitting may be 1 / 4, 1 / 2, and 1 / 4 of the horizontal size or vertical size of the original CU before the splitting, respectively.
[0312] For example, when a CU of size 32×32 is vertically split into three CUs, the sizes of the three CUs generated by the splitting may be 8×32, 16×32, and 8×32, respectively. In this way, when a single CU is split into three CUs, it can be considered that the CU is split in the form of a ternary tree.
[0313] One of the exemplary partitioning forms (i.e., quadtree partitioning, binary tree partitioning, and ternary tree partitioning) may be applied to the partitioning of the CU, and multiple partitioning schemes may be combined and used together for the partitioning of the CU. Here, the case where multiple partitioning schemes are combined and used together may be referred to as "composite tree format partitioning."
[0314] Figure 6 The division of blocks according to an example is shown.
[0315] In the video encoding and / or decoding process, such as Figure 6 As shown in , the target block can be divided.
[0316] For splitting of the target block, an indicator indicating splitting information may be signaled from the encoding apparatus 100 to the decoding apparatus 200. The splitting information may be information indicating how the target block is split.
[0317] The split information may be one or more of a split flag (hereinafter referred to as "split_flag"), a quad-binary flag (hereinafter referred to as "QB_flag"), a quadtree flag (hereinafter referred to as "quadtree_flag"), a binary tree flag (hereinafter referred to as "binarytree_flag") and a binary type flag (hereinafter referred to as "Btype_flag").
[0318] "split_flag" may be a flag indicating whether a block is split. For example, a split_flag value of 1 may indicate that the corresponding block is split. A split_flag value of 0 may indicate that the corresponding block is not split.
[0319] "QB_flag" may be a flag indicating which of the quadtree and binary tree forms corresponds to the shape of the block partition. For example, a QB_flag value of 0 may indicate that the block is partitioned in a quadtree form. A QB_flag value of 1 may indicate that the block is partitioned in a binary tree form. Alternatively, a QB_flag value of 0 may indicate that the block is partitioned in a binary tree form. A QB_flag value of 1 may indicate that the block is partitioned in a quadtree form.
[0320] "quadtree_flag" may be a flag indicating whether the block is partitioned in a quadtree form. For example, a quadtree_flag value of 1 may indicate that the block is partitioned in a quadtree form. A quadtree_flag value of 0 may indicate that the block is not partitioned in a quadtree form.
[0321] "binarytree_flag" may be a flag indicating whether the block is partitioned in a binary tree form. For example, a binarytree_flag value of 1 may indicate that the block is partitioned in a binary tree form. A binarytree_flag value of 0 may indicate that the block is not partitioned in a binary tree form.
[0322] "Btype_flag" may be a flag indicating which of vertical and horizontal divisions corresponds to the division direction when the block is divided in a binary tree form. For example, a Btype_flag value of 0 may indicate that the block is divided in the horizontal direction. A Btype_flag value of 1 may indicate that the block is divided in the vertical direction. Alternatively, a Btype_flag value of 0 may indicate that the block is divided in the vertical direction. A Btype_flag value of 1 may indicate that the block is divided in the horizontal direction.
[0323] For example, it can be derived by signaling at least one of quadtree_flag, binarytree_flag, and Btype_flag. Figure 6 The partitioning information of the blocks in is shown in Table 1 below.
[0324] Table 1
[0325]
[0326] For example, it can be derived by signaling at least one of split_flag, QB_flag, and Btype_flag. Figure 6 The partitioning information of the blocks in is shown in Table 2 below.
[0327] Table 2
[0328]
[0329] The partitioning method may be limited to a quadtree or a binary tree depending on the size and / or shape of the block. When this limitation is applied, split_flag may be a flag indicating whether the block is partitioned in a quadtree or a flag indicating whether the block is partitioned in a binary tree. The size and shape of the block may be derived based on the depth information of the block, and the depth information may be signaled from the encoding device 100 to the decoding device 200.
[0330] When the block size falls within a specific range, partitioning only in quadtree form is possible. For example, the specific range may be defined by at least one of a maximum block size and a minimum block size that can be partitioned only in quadtree form.
[0331] Information indicating the maximum block size and the minimum block size that can be divided only in the quadtree form may be signaled from the encoding apparatus 100 to the decoding apparatus 200 through a bitstream. In addition, this information may be signaled for at least one of units such as a video, a sequence, a picture, and a slice (or segment).
[0332] Alternatively, the maximum block size and / or the minimum block size may be a fixed size predefined by the encoding device 100 and the decoding device 200. For example, when the size of the block is larger than 64×64 and smaller than 256×256, only quadtree division is possible. In this case, split_flag may be a flag indicating whether quadtree division is performed.
[0333] When the size of the block falls within a specific range, partitioning only in the binary tree form is possible. For example, the specific range may be defined by at least one of a maximum block size and a minimum block size that can be partitioned only in the binary tree form.
[0334] Information indicating a maximum block size and / or a minimum block size that can be divided only in a binary tree form may be signaled from the encoding apparatus 100 to the decoding apparatus 200 through a bitstream. In addition, this information may be signaled for at least one of units such as a sequence, a picture, and a slice (or segment).
[0335] Alternatively, the maximum block size and / or the minimum block size may be a fixed size predefined by the encoding device 100 and the decoding device 200. For example, when the size of the block is greater than 8×8 and less than 16×16, only binary tree division is possible. In this case, split_flag may be a flag indicating whether binary tree division is performed.
[0336] The partitioning of a block may be limited by the previous partitioning. For example, when a block is partitioned in a binary tree form and a plurality of partition blocks are generated, each partition block may be further partitioned only in a binary tree form.
[0337] When the horizontal size or the vertical size of the partition block is a size that cannot be further divided, the above-mentioned indicator may not be signaled.
[0338] Figure 7 is a diagram for explaining an embodiment of an intra prediction process.
[0339] from Figure 7 The arrow extending radially from the center of the diagram in represents the prediction direction of the intra prediction mode. Furthermore, numbers appearing near the arrows indicate examples of mode values assigned to the intra prediction mode or the prediction direction of the intra prediction mode.
[0340] Intra-frame encoding and / or decoding may be performed using reference samples of blocks adjacent to the target block. The adjacent blocks may be adjacent reconstructed blocks. For example, intra-frame encoding and / or decoding may be performed using values of reference samples included in each adjacent reconstructed block or encoding parameters of the adjacent reconstructed blocks.
[0341] The encoding device 100 and / or the decoding device 200 may generate a prediction block by performing intra-frame prediction on the target block based on information about samples in the target image. When intra-frame prediction is performed, the encoding device 100 and / or the decoding device 200 may generate a prediction block for the target block by performing intra-frame prediction based on information about samples in the target image. When intra-frame prediction is performed, the encoding device 100 and / or the decoding device 200 may perform directional prediction and / or non-directional prediction based on at least one reconstructed reference sample.
[0342] The prediction block may be a block generated as a result of performing intra prediction. The prediction block may correspond to at least one of a CU, a PU, and a TU.
[0343] The prediction block may have a size corresponding to at least one of a CU, a PU, and a TU. The prediction block may have a square shape of 2N×2N or N×N. The N×N size may include sizes of 4×4, 8×8, 16×16, 32×32, 64×64, etc.
[0344] Alternatively, the prediction block may be a square block of size 2×2, 4×4, 8×8, 16×16, 32×32, 64×64, etc. or a rectangular block of size 2×8, 4×8, 2×16, 4×16, 8×16, etc.
[0345] Intra-frame prediction may be performed considering the intra-frame prediction mode used for the target block. The number of intra-frame prediction modes that the target block can have may be a predefined fixed value, or may be a value determined differently according to the properties of the prediction block. For example, the properties of the prediction block may include the size of the prediction block, the type of the prediction block, etc.
[0346] For example, regardless of the size of the prediction block, the number of intra prediction modes may be fixed to 35. Alternatively, the number of intra prediction modes may be 3, 5, 9, 17, 34, 35, or 36, for example.
[0347] The intra prediction mode can be a non-directional mode or a directional mode. Figure 7 As shown in , the intra prediction modes may include two non-directional modes and 33 directional modes.
[0348] The two non-directional modes may include a DC mode and a planar mode.
[0349] The direction pattern may be a pattern having a specific direction or a specific angle.
[0350] Each intra-frame prediction mode can be represented by at least one of a mode number, a mode value, and a mode angle. The number of intra-frame prediction modes may be M. The value of M may be 1 or greater. In other words, the number of intra-frame prediction modes may be M, where M includes the number of non-directional modes and the number of directional modes.
[0351] Regardless of the size of the block and / or the color components, the number of intra prediction modes may be fixed to M. For example, the number of intra prediction modes may be fixed to any one of 35 and 67 regardless of the size of the block.
[0352] Alternatively, the number of intra prediction modes may differ according to the size of the block and / or the type of color component.
[0353] For example, the larger the block size, the greater the number of intra-frame prediction modes. Alternatively, the larger the block size, the fewer the number of intra-frame prediction modes. When the block size is 4×4 or 8×8, the number of intra-frame prediction modes may be 67. When the block size is 16×16, the number of intra-frame prediction modes may be 35. When the block size is 32×32, the number of intra-frame prediction modes may be 19. When the block size is 64×64, the number of intra-frame prediction modes may be 7.
[0354] For example, the number of intra prediction modes may differ depending on whether the color component is a luma signal or a chroma signal. Alternatively, the number of intra prediction modes corresponding to a luma component block may be greater than the number of intra prediction modes corresponding to a chroma component block.
[0355] For example, in vertical mode with a mode value of 26, prediction may be performed in a vertical direction based on pixel values of reference samples. For example, in horizontal mode with a mode value of 10, prediction may be performed in a horizontal direction based on pixel values of reference samples.
[0356] Even in a directional mode other than the above-described modes, the encoding apparatus 100 and the decoding apparatus 200 may perform intra prediction on a target unit using reference samples according to an angle corresponding to the directional mode.
[0357] An intra-frame prediction mode located to the right relative to the vertical mode may be referred to as a "vertical-right mode". An intra-frame prediction mode located below the horizontal mode may be referred to as a "horizontal-below mode". For example, in Figure 7 , an intra prediction mode whose mode value is one of 27, 28, 29, 30, 31, 32, 33, and 34 may be a vertical-right mode 613. An intra prediction mode whose mode value is one of 2, 3, 4, 5, 6, 7, 8, and 9 may be a horizontal-bottom mode 616.
[0358] The non-directional mode may include a DC mode and a planar mode. For example, the value of the DC mode may be 1. The value of the planar mode may be 0.
[0359] The directional mode may include an angular mode. Among the plurality of intra prediction modes, the remaining modes except the DC mode and the planar mode may be directional modes.
[0360] When the intra prediction mode is the DC mode, the prediction block may be generated based on the average value of the pixel values of the plurality of reference pixels. For example, the pixel values of the prediction block may be determined based on the average value of the pixel values of the plurality of reference pixels.
[0361] The number of intra prediction modes and the mode values of each intra prediction mode described above are merely exemplary and may be defined differently depending on embodiments, implementations, and / or requirements.
[0362] In order to perform intra prediction on a target block, a step of checking whether samples included in a reconstructed neighboring block can be used as reference samples for the target block may be performed. When a sample that cannot be used as a reference sample for the target block exists among the samples in the neighboring block, a value generated by interpolation and / or replication using at least one sample value among the samples included in the reconstructed neighboring block may replace the sample value of the sample that cannot be used as a reference sample. When the value generated by replication and / or interpolation replaces the sample value of the existing sample, the sample may be used as a reference sample for the target block.
[0363] In intra prediction, a filter may be applied to at least one of a reference sample and a prediction sample based on at least one of an intra prediction mode and a size of a target block.
[0364] The type of filter to be applied to at least one of the reference sample and the prediction sample may differ according to at least one of the intra prediction mode of the target block, the size of the target block, and the shape of the target block. The type of filter may be classified according to one or more of the number of filter taps, the value of the filter coefficient, and the filter strength.
[0365] When the intra prediction mode is the planar mode, the sample value of the predicted target block may be generated by using a weighted sum of the upper reference sample of the target block, the left reference sample of the target block, the upper right reference sample of the target block, and the lower left reference sample of the target block according to the position of the predicted target sample in the prediction block when generating the prediction block of the target block.
[0366] When the intra prediction mode is the DC mode, an average value of reference samples above the target block and reference samples to the left of the target block may be used when generating a prediction block for the target block. Furthermore, filtering using the values of the reference samples may be performed on specific rows or columns in the target block. The specific rows may be one or more rows above the reference sample. The specific columns may be one or more columns to the left of the reference sample.
[0367] When the intra prediction mode is a directional mode, the prediction block may be generated using the upper reference sample, the left reference sample, the upper right reference sample, and / or the lower left reference sample of the target block.
[0368] In order to generate the above-mentioned prediction samples, real number-based interpolation may be performed.
[0369] The intra prediction mode of the target block may be predicted from intra prediction modes of neighboring blocks adjacent to the target block, and information used for the prediction may be entropy encoded / decoded.
[0370] For example, when the intra prediction modes of the target block and the neighboring block are identical to each other, a predefined flag may be used to signal that the intra prediction modes of the target block and the neighboring block are the same.
[0371] For example, an indicator indicating the same intra prediction mode as that of the target block among intra prediction modes of a plurality of neighboring blocks may be signaled.
[0372] When the intra prediction modes of the target block and the neighboring blocks are different from each other, information about the intra prediction mode of the target block may be encoded and / or decoded using entropy encoding and / or entropy decoding.
[0373] Figure 8 is a diagram for explaining positions of reference samples used in an intra prediction process.
[0374] Figure 8 The position of the reference sample used for intra prediction of the target block is shown. Figure 8 , the reconstructed reference samples used for intra-frame prediction of the target block may include a lower left reference sample 831 , a left reference sample 833 , an upper left corner reference sample 835 , an upper reference sample 837 and an upper right reference sample 839 .
[0375] For example, left reference sample 833 may represent a reconstructed reference pixel adjacent to the left side of the target block. Upper reference sample 837 may represent a reconstructed reference pixel adjacent to the top of the target block. Upper left corner reference sample 835 may represent a reconstructed reference pixel located at the upper left corner of the target block. Lower left reference sample 831 may represent a reference sample located below the left sample line among samples located on the same line as the left sample line formed by left reference sample 833. Upper right reference sample 839 may represent a reference sample located to the right of the upper sample line among samples located on the same line as the upper sample line formed by upper reference sample 837.
[0376] When the size of the target block is N×N, the numbers of the lower left reference sample 831 , the left reference sample 833 , the upper reference sample 837 , and the upper right reference sample 839 may all be N.
[0377] By performing intra prediction on the target block, a prediction block can be generated. The process of generating the prediction block may include determining the values of the pixels in the prediction block. The target block and the prediction block may be the same size.
[0378] The reference samples used for intra-frame prediction of the target block may change according to the intra-frame prediction mode of the target block. The direction of the intra-frame prediction mode may indicate the dependency between the reference samples and the pixels of the prediction block. For example, the value of a specified reference sample may be used as the value of one or more specified pixels in the prediction block. In this case, the specified reference sample and the one or more specified pixels in the prediction block may be samples and pixels located on a straight line along the direction of the intra-frame prediction mode. In other words, the value of the specified reference sample may be copied as the value of a pixel located in a direction opposite to the direction of the intra-frame prediction mode. Alternatively, the value of a pixel in the prediction block may be the value of a reference sample located in the direction of the intra-frame prediction mode relative to the position of the pixel.
[0379] In this example, when the intra prediction mode of the target block is a vertical mode with a mode value of 26, the upper reference sample 837 can be used for intra prediction. When the intra prediction mode is a vertical mode, the value of a pixel in the prediction block can be the value of a reference sample located vertically above the position of the pixel. Therefore, the upper reference sample 837 adjacent to the top of the target block can be used for intra prediction. In addition, the value of the pixels in a row of the prediction block can be the same as the value of the pixel of the upper reference sample 837.
[0380] In this example, when the intra prediction mode of the target block is horizontal mode with a mode value of 10, the left reference sample 833 can be used for intra prediction. When the intra prediction mode is horizontal mode, the value of a pixel in the prediction block can be the value of a reference sample located horizontally to the left of the pixel. Therefore, the left reference sample 833 adjacent to the left side of the target block can be used for intra prediction. In addition, the value of the pixel in a column of the prediction block can be the same as the value of the pixel of the left reference sample 833.
[0381] In an example, when the mode value of the intra prediction mode of the current block is 18, at least some of the left reference samples 833, the upper left reference sample 835, and at least some of the upper reference samples 837 may be used for intra prediction. When the mode value of the intra prediction mode is 18, the value of a pixel in the prediction block may be the value of a reference sample diagonally located at the upper left corner of the pixel.
[0382] In addition, in a case where an intra prediction mode with a mode value of 27, 28, 29, 30, 31, 32, 33, or 34 is used, at least a portion of the upper right reference sample 839 may be used for intra prediction.
[0383] In addition, in the case where an intra prediction mode with a mode value of 2, 3, 4, 5, 6, 7, 8, or 9 is used, at least a portion of the lower left reference sample 831 may be used for intra prediction.
[0384] Also, in the case of an intra prediction mode in which the mode value is a value ranging from 11 to 25, the upper left corner reference sample 835 may be used for intra prediction.
[0385] The number of reference samples used to determine the pixel value of one pixel in the prediction block may be 1 or 2 or more.
[0386] As described above, the pixel value of a pixel in the prediction block may be determined based on the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode. When the position of the pixel and the position of the reference sample indicated by the direction of the intra prediction mode are integer positions, the value of one reference sample indicated by the integer position may be used to determine the pixel value of the pixel in the prediction block.
[0387] When the position of a pixel and the position of a reference sample indicated by the direction of the intra-frame prediction mode are not integer positions, an interpolated reference sample based on the two reference samples closest to the position of the reference sample may be generated. The value of the interpolated reference sample may be used to determine the pixel value of a pixel in the prediction block. In other words, when the position of a pixel in the prediction block and the position of a reference sample indicated by the direction of the intra-frame prediction mode indicate a position between two reference samples, an interpolated value based on the values of these two samples may be generated.
[0388] The prediction block generated via prediction may be different from the original target block. In other words, there may be a prediction error, which is the difference between the target block and the prediction block, and there may also be a prediction error between the pixels of the target block and the pixels of the prediction block.
[0389] Hereinafter, the terms “difference”, “error” and “residual” may be used to have the same meaning and may be used interchangeably with each other.
[0390] For example, in the case of directional intra prediction, the longer the distance between the pixels of the prediction block and the reference samples, the greater the prediction error that may occur. This prediction error can lead to discontinuities between the generated prediction block and neighboring blocks.
[0391] To reduce prediction error, a filtering operation may be applied to the prediction block. The filtering operation may be configured to adaptively apply a filter to areas of the prediction block that are considered to have a large prediction error. For example, the areas considered to have a large prediction error may be boundaries of the prediction block. Furthermore, the areas of the prediction block considered to have a large prediction error may vary depending on the intra-frame prediction mode, and the characteristics of the filter may also vary depending on the intra-frame prediction mode.
[0392] Figure 9 is a diagram for explaining an embodiment of an inter-frame prediction process.
[0393] Figure 9 The rectangle shown in can represent an image (or picture). Figure 9 In FIG, an arrow may indicate a prediction direction. That is, each image may be encoded and / or decoded according to the prediction direction.
[0394] Images can be classified into intra-frame pictures (I pictures), uni-predictive pictures or predictive coded pictures (P pictures), and bi-predictive pictures or bi-predictive coded pictures (B pictures) according to the coding type. Each picture can be encoded and / or decoded according to the coding type of each picture.
[0395] When the target image to be encoded is an I picture, the target image can be encoded using data contained in the image itself without performing inter-frame prediction with reference to other images. For example, the I picture can be encoded only through intra-frame prediction.
[0396] When the target image is a P picture, the target image may be encoded through inter-frame prediction using a reference picture existing in one direction. Here, the one direction may be a forward direction or a backward direction.
[0397] When the target image is a B picture, the image may be encoded by inter-frame prediction using reference pictures existing in both directions, or may be encoded by inter-frame prediction using reference pictures existing in one of the forward direction and the backward direction. Here, the two directions may be the forward direction and the backward direction.
[0398] P-pictures and B-pictures encoded and / or decoded using reference pictures may be regarded as images using inter-frame prediction.
[0399] Hereinafter, inter prediction in inter mode according to an embodiment will be described in detail.
[0400] Inter-frame prediction may be performed using motion information.
[0401] In the inter mode, the encoding apparatus 100 may perform inter prediction and / or motion compensation on the target block, and the decoding apparatus 200 may perform inter prediction and / or motion compensation corresponding to the inter prediction and / or motion compensation performed by the encoding apparatus 100 on the target block.
[0402] The motion information of the target block may be separately derived during inter prediction by the encoding apparatus 100 and the decoding apparatus 200. The motion information may be derived using motion information of a reconstructed neighboring block, motion information of a col block, and / or motion information of a block adjacent to the col block.
[0403] For example, the encoding apparatus 100 or the decoding apparatus 200 may perform prediction and / or motion compensation by using the motion information of the spatial candidate and / or the temporal candidate as the motion information of the target block. The target block may represent a PU and / or a PU partition.
[0404] The spatial candidate may be a reconstructed block that is spatially adjacent to the target block.
[0405] The temporal candidate may be a reconstructed block corresponding to the target block in a previously reconstructed co-located picture (col picture).
[0406] In inter-frame prediction, the encoding device 100 and the decoding device 200 can improve encoding efficiency and decoding efficiency by utilizing motion information of spatial candidates and / or temporal candidates. The motion information of the spatial candidate may be referred to as "spatial motion information." The motion information of the temporal candidate may be referred to as "temporal motion information."
[0407] Hereinafter, the motion information of a spatial candidate may be the motion information of a PU including the spatial candidate. The motion information of a temporal candidate may be the motion information of a PU including the temporal candidate. The motion information of a candidate block may be the motion information of a PU including the candidate block.
[0408] Inter prediction may be performed using reference pictures.
[0409] The reference picture may be at least one of a picture before the target picture and a picture after the target picture.The reference picture may be an image used for prediction of the target block.
[0410] In inter prediction, a region in a reference picture may be specified using a reference picture index (or refIdx) indicating a reference picture, a motion vector to be described later, etc. Here, the region specified in the reference picture may indicate a reference block.
[0411] Inter prediction can select a reference picture and can also select a reference block corresponding to the target block from the reference picture. In addition, inter prediction can use the selected reference block to generate a prediction block for the target block.
[0412] Motion information may be derived by each of the encoding apparatus 100 and the decoding apparatus 200 during inter prediction.
[0413] A spatial candidate may be a block that 1) exists in the target picture 2) has been previously reconstructed via encoding and / or decoding and 3) is adjacent to the target block or is located at a corner of the target block. Here, a "block located at a corner of the target block" may be a block vertically adjacent to a neighboring block horizontally adjacent to the target block, or a block horizontally adjacent to a neighboring block vertically adjacent to the target block. In addition, a "block located at a corner of the target block" may have the same meaning as a "block adjacent to a corner of the target block." The meaning of a "block located at a corner of the target block" may be included in the meaning of a "block adjacent to the target block."
[0414] For example, the spatial candidate can be a reconstructed block located to the left of the target block, a reconstructed block located above the target block, a reconstructed block located at the lower left corner of the target block, a reconstructed block located at the upper right corner of the target block, or a target block located at the upper left corner of the target block.
[0415] Each of the encoding apparatus 100 and the decoding apparatus 200 may identify a block existing in the col picture at a position spatially corresponding to the target block. The position of the target block in the target picture and the position of the identified block in the col picture may correspond to each other.
[0416] Each of the encoding apparatus 100 and the decoding apparatus 200 may determine a col block existing at a predefined relative position with respect to the identified block as a temporal candidate. The predefined relative position may be a position existing inside and / or outside the identified block.
[0417] For example, the col block may include a first col block and a second col block. When the coordinates of the identified block are (xP, yP) and the size of the identified block is represented by (nPSW, nPSH), the first col block may be a block located at the coordinates (xP+nPSW, yP+nPSH). The second col block may be a block located at the coordinates (xP+(nPSW>>1), yP+(nPSH>>1)). When the first col block is unavailable, the second col block may be selectively used.
[0418] The motion vector of the target block may be determined based on the motion vector of the col block. Each of the encoding apparatus 100 and the decoding apparatus 200 may scale the motion vector of the col block. The scaled motion vector of the col block may be used as the motion vector of the target block. In addition, the motion vector of the running information of the temporal candidate stored in the list may be the scaled motion vector.
[0419] The ratio of the motion vector of the target block to the motion vector of the col block may be the same as the ratio of the first distance to the second distance. The first distance may be the distance between the reference picture and the target picture of the target block. The second distance may be the distance between the reference picture and the col picture of the col block.
[0420] The scheme for deriving motion information may vary depending on the inter-frame prediction mode of the target block. For example, as inter-frame prediction modes applied to inter-frame prediction, there may be Advanced Motion Vector Predictor (AMVP) mode, merge mode, skip mode, current picture reference mode, etc. Merge mode may also be referred to as "motion merge mode." Each mode will be described in detail below.
[0421] 1) AMVP model
[0422] When using the AMVP mode, the encoding device 100 may search for similar blocks in the vicinity of the target block. The encoding device 100 may obtain a prediction block by performing prediction on the target block using the motion information of the found similar blocks. The encoding device 100 may encode a residual block that is the difference between the target block and the prediction block.
[0423] 1-1) Create a list of predicted motion vector candidates
[0424] When the AMVP mode is used as the prediction mode, each of the encoding device 100 and the decoding device 200 can create a list of prediction motion vector candidates using the motion vector of the spatial candidate, the motion vector of the temporal candidate, and the zero vector. The prediction motion vector candidate list may include one or more prediction motion vector candidates. At least one of the motion vector of the spatial candidate, the motion vector of the temporal candidate, and the zero vector may be determined and used as the prediction motion vector candidate.
[0425] Hereinafter, the terms "prediction motion vector (candidate)" and "motion vector (candidate)" may be used to have the same meaning and may be used interchangeably with each other.
[0426] Hereinafter, the terms “predicted motion vector candidate” and “AMVP candidate” may be used to have the same meaning and may be used interchangeably with each other.
[0427] Hereinafter, the terms “motion vector prediction candidate list” and “AMVP candidate list” may be used to have the same meaning and may be used interchangeably with each other.
[0428] The spatial candidate may include a reconstructed spatial neighboring block. In other words, the motion vector of the reconstructed neighboring block may be referred to as a "spatial prediction motion vector candidate."
[0429] The temporal candidate may include the col block and blocks adjacent to the col block. In other words, the motion vector of the col block or the motion vector of the block adjacent to the col block may be referred to as a "temporal prediction motion vector candidate".
[0430] The zero vector may be the (0,0) motion vector.
[0431] The predicted motion vector candidate may be a motion vector predictor for predicting a motion vector. In addition, in the encoding apparatus 100, each predicted motion vector candidate may be an initial search position for a motion vector.
[0432] 1-2) Searching for motion vector using a list of predicted motion vector candidates
[0433] The encoding device 100 may determine a motion vector to be used for encoding the target block within the search range using the list of predicted motion vector candidates. In addition, the encoding device 100 may determine a predicted motion vector candidate to be used as the predicted motion vector of the target block from among the predicted motion vector candidates in the predicted motion vector candidate list.
[0434] The motion vector to be used for encoding the target block may be a motion vector that can be encoded at a minimum cost.
[0435] Also, the encoding apparatus 100 may determine whether to encode the target block using the AMVP mode.
[0436] 1-3) Transmission of inter-frame prediction information
[0437] The encoding apparatus 100 may generate a bitstream including inter prediction information required for inter prediction, and the decoding apparatus 200 may perform inter prediction on a target block using the inter prediction information of the bitstream.
[0438] The inter prediction information may include 1) mode information indicating whether the AMVP mode is used, 2) a prediction motion vector index, 3) a motion vector difference (MVD), 4) a reference direction, and 5) a reference picture index.
[0439] Hereinafter, the terms “predicted motion vector index” and “AMVP index” may be used to have the same meaning and may be used interchangeably with each other.
[0440] In addition, the inter-frame prediction information may include a residual signal.
[0441] When the mode information indicates that the AMVP mode is used, the decoding apparatus 200 may acquire a predicted motion vector index, an MVD, a reference direction, and a reference picture index from a bitstream through entropy decoding.
[0442] The predicted motion vector index may indicate a predicted motion vector candidate to be used for prediction of the target block among the predicted motion vector candidates included in the predicted motion vector candidate list.
[0443] 1-4) Inter-frame prediction in AMVP mode using inter-frame prediction information
[0444] The decoding apparatus 200 may derive a prediction motion vector candidate using the prediction motion vector candidate list, and may determine motion information of a target block based on the derived prediction motion vector candidate.
[0445] The decoding apparatus 200 may determine a motion vector candidate for the target block from among the predicted motion vector candidates included in the predicted motion vector candidate list using the predicted motion vector index. The decoding apparatus 200 may select the predicted motion vector candidate indicated by the predicted motion vector index from among the predicted motion vector candidates included in the predicted motion vector candidate list as the predicted motion vector of the target block.
[0446] The motion vector actually used for inter-frame prediction of the target block may not match the predicted motion vector. To indicate the difference between the motion vector actually used for inter-frame prediction of the target block and the predicted motion vector, MVD may be used. The encoding apparatus 100 may derive a predicted motion vector similar to the motion vector actually used for inter-frame prediction of the target block in order to use as small an MVD as possible.
[0447] The MVD may be a difference between a motion vector of a target block and a predicted motion vector. The encoding apparatus 100 may calculate the MVD and may entropy encode the MVD.
[0448] The MVD may be transmitted from the encoding device 100 to the decoding device 200 via a bitstream. The decoding device 200 may decode the received MVD. The decoding device 200 may derive a motion vector of the target block by summing the decoded MVD and the predicted motion vector. In other words, the motion vector of the target block derived by the decoding device 200 may be the sum of the entropy-decoded MVD and the motion vector candidate.
[0449] The reference direction may indicate a list of reference pictures to be used for prediction of the target block. For example, the reference direction may indicate one of reference picture list L0 and reference picture list L1.
[0450] The reference direction only indicates the reference picture list to be used for prediction of the target block and may not mean that the direction of the reference picture is limited to the forward direction or the backward direction. In other words, each of the reference picture lists L0 and L1 may include pictures in the forward direction and / or the backward direction.
[0451] A unidirectional reference direction may mean that a single reference picture list is used. A bidirectional reference direction may mean that two reference picture lists are used. In other words, the reference direction may indicate one of the following cases: a case where only reference picture list L0 is used, a case where only reference picture list L1 is used, or a case where two reference picture lists are used.
[0452] The reference picture index may indicate a reference picture to be used for prediction of the target block among the reference pictures in the reference picture list. The reference picture index may be entropy-encoded by the encoding apparatus 100. The entropy-encoded reference picture index may be signaled by the encoding apparatus 100 to the decoding apparatus 200 via a bitstream.
[0453] When two reference picture lists are used to predict a target block, a single reference picture index and a single motion vector can be used for each reference picture list. In addition, when two reference picture lists are used to predict a target block, two prediction blocks can be specified for the target block. For example, the (final) prediction block for the target block can be generated using an average or weighted sum of the two prediction blocks for the target block.
[0454] The motion vector of the target block may be derived by predicting the motion vector index, MVD, reference direction, and reference picture index.
[0455] The decoding apparatus 200 may generate a prediction block for the target block based on the derived motion vector and the reference picture index. For example, the prediction block may be a reference block indicated by the derived motion vector in the reference picture indicated by the reference picture index.
[0456] Since the predicted motion vector index and the MVD are encoded but the motion vector itself of the target block is not encoded, the number of bits transmitted from the encoding apparatus 100 to the decoding apparatus 200 may be reduced and encoding efficiency may be improved.
[0457] The reconstructed motion information of the neighboring blocks can be used for the target block. In certain inter-frame prediction modes, the encoding device 100 may not separately encode the actual motion information of the target block. Instead of encoding the motion information of the target block, additional information may be encoded, where the additional information enables the motion information of the target block to be derived using the reconstructed motion information of the neighboring blocks. Since this additional information is encoded, the number of bits transmitted to the decoding device 200 can be reduced, and encoding efficiency can be improved.
[0458] For example, as an inter-frame prediction mode in which the motion information of the target block is not directly encoded, a skip mode and / or a merge mode may exist. Here, each of the encoding device 100 and the decoding device 200 may use an indicator and / or an index indicating a unit whose motion information is to be used as the motion information of the target unit among the reconstructed neighboring units.
[0459] 2) Merge mode
[0460] As a scheme for deriving motion information of a target block, there is a merge mode. The term "merge" may mean merging the motions of multiple blocks. "Merge" may also mean that the motion information of one block is also applied to other blocks. In other words, the merge mode may be a mode in which the motion information of the target block is derived from the motion information of neighboring blocks.
[0461] When using merge mode, the encoding device 100 may use the motion information of the spatial candidate and / or the motion information of the temporal candidate to predict the motion information of the target block. The spatial candidate may include a reconstructed spatial neighboring block that is spatially adjacent to the target block. The spatial neighboring block may include a left neighboring block and an upper neighboring block. The temporal candidate may include a col block. The terms "spatial candidate" and "spatial merge candidate" may be used to have the same meaning and may be used interchangeably with each other. The terms "temporal candidate" and "temporal merge candidate" may be used to have the same meaning and may be used interchangeably with each other.
[0462] The encoding apparatus 100 may obtain a prediction block through prediction. The encoding apparatus 100 may encode a residual block that is a difference between the target block and the prediction block.
[0463] 2-1) Create a merge candidate list
[0464] When using merge mode, each of the encoding device 100 and the decoding device 200 can use the motion information of the spatial candidate and / or the motion information of the temporal candidate to create a merge candidate list. The motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction can be unidirectional or bidirectional.
[0465] The merge candidate list may include a merge candidate. The merge candidate may be motion information. In other words, the merge candidate list may be a list storing multiple pieces of motion information.
[0466] A merge candidate may be a plurality of pieces of motion information of temporal candidates and / or spatial candidates. In addition, a merge candidate list may include new merge candidates generated by combining merge candidates already in the merge candidate list. In other words, the merge candidate list may include new motion information generated by combining multiple pieces of motion information previously in the merge candidate list.
[0467] A merge candidate may be a specific mode for deriving inter-frame prediction information. A merge candidate may be information indicating a specific mode for deriving inter-frame prediction information. Inter-frame prediction information for a target block may be derived according to the specific mode indicated by the merge candidate. Furthermore, the specific mode may include a process for deriving a series of inter-frame prediction information. This specific mode may be an inter-frame prediction information derivation mode or a motion information derivation mode.
[0468] Inter prediction information of the target block may be derived according to a mode indicated by a merge candidate selected from among merge candidates in the merge candidate list by a merge index.
[0469] For example, the motion information derivation mode in the merge candidate list may be at least one of the following modes: 1) a motion information derivation mode for a sub-block unit; 2) an affine motion information derivation mode.
[0470] In addition, the merge candidate list may include motion information of a zero vector. A zero vector may also be referred to as a "zero merge candidate."
[0471] In other words, the multiple motion information in the merge candidate list can be at least one of the following information: 1) motion information of spatial candidates, 2) motion information of temporal candidates, 3) motion information generated by combining multiple motion information previously existing in the merge candidate list, and 4) zero vector.
[0472] Motion information may include 1) a motion vector, 2) a reference picture index, and 3) a reference direction. The reference direction may also be referred to as an "inter prediction indicator." The reference direction may be unidirectional or bidirectional. A unidirectional reference direction may indicate L0 prediction or L1 prediction.
[0473] A merge candidate list may be created before performing prediction in merge mode.
[0474] The number of merge candidates in the merge candidate list may be predefined. Each of the encoding device 100 and the decoding device 200 may add merge candidates to the merge candidate list according to a predefined scheme and a predefined priority, so that the merge candidate list has a predefined number of merge candidates. The merge candidate list of the encoding device 100 and the merge candidate list of the decoding device 200 may be made identical to each other using the predefined scheme and the predefined priority.
[0475] Merging can be applied on a CU or PU basis. When performing merging on a CU or PU basis, the encoding device 100 may transmit a bitstream including predefined information to the decoding device 200. For example, the predefined information may include 1) information indicating whether merging is performed for each block partition, and 2) information about blocks on which merging is to be performed among blocks that are spatial candidates and / or temporal candidates for the target block.
[0476] 2-2) Searching for motion vectors using the merge candidate list
[0477] The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using a merge candidate in the merge candidate list and may generate a residual block for the merge candidate. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in encoding the prediction and residual block.
[0478] Also, the encoding apparatus 100 may determine whether to encode the target block using the merge mode.
[0479] 2-3) Transmission of inter-frame prediction information
[0480] The encoding apparatus 100 may generate a bitstream including inter-frame prediction information required for inter-frame prediction. The encoding apparatus 100 may generate entropy-coded inter-frame prediction information by performing entropy coding on the inter-frame prediction information, and may transmit the bitstream including the entropy-coded inter-frame prediction information to the decoding apparatus 200. The entropy-coded inter-frame prediction information may be signaled from the encoding apparatus 100 to the decoding apparatus 200 through the bitstream.
[0481] The decoding apparatus 200 may perform inter prediction on a target block using inter prediction information of a bitstream.
[0482] The inter prediction information may include 1) mode information indicating whether the merge mode is used and 2) a merge index.
[0483] In addition, the inter-frame prediction information may include a residual signal.
[0484] The decoding apparatus 200 may acquire a merge index from a bitstream only when the mode information indicates that the merge mode is used.
[0485] The mode information may be a merge flag. The unit of the mode information may be a block. The information about the block may include the mode information, and the mode information may indicate whether the merge mode is applied to the block.
[0486] The merge index may indicate a merge candidate to be used for predicting the target block among the merge candidates included in the merge candidate list. Alternatively, the merge index may indicate a block to be merged with the target block among neighboring blocks that are spatially or temporally adjacent to the target block.
[0487] The encoding apparatus 100 may select a merge candidate having the highest encoding performance among the merge candidates included in the merge candidate list, and set a value of the merge index to indicate the selected merge candidate.
[0488] 2-4) Inter-frame prediction using merge mode of inter-frame prediction information
[0489] The decoding apparatus 200 may perform prediction on the target block using a merge candidate indicated by a merge index among merge candidates included in the merge candidate list.
[0490] The motion vector of the target block may be specified by the motion vector of the merge candidate indicated by the merge index, the reference picture index, and the reference direction.
[0491] 3) Skip mode
[0492] Skip mode can be a mode in which the motion information of a spatial candidate or a temporal candidate is applied to the target block without change. In addition, skip mode can be a mode in which a residual signal is not used. In other words, when skip mode is used, the reconstructed block can be a predicted block.
[0493] The difference between the merge mode and the skip mode is whether the residual signal is transmitted or used. That is, the skip mode may be similar to the merge mode except that the residual signal is not transmitted or used.
[0494] When the skip mode is used, the encoding apparatus 100 may transmit information about a block, among blocks serving as spatial candidates or temporal candidates, whose motion information is to be used as the motion information of the target block, to the decoding apparatus 200 through a bitstream. The encoding apparatus 100 may generate entropy-coded information by performing entropy coding on the information, and may signal the entropy-coded information to the decoding apparatus 200 through a bitstream.
[0495] In addition, when the skip mode is used, the encoding device 100 may not transmit other syntax information (such as MVD) to the decoding device 200. For example, when the skip mode is used, the encoding device 100 may not signal syntax elements related to at least one of MVC, coded block flags, and transform coefficient levels to the decoding device 200.
[0496] 3-1) Create a merge candidate list
[0497] The merge candidate list can also be used in skip mode. In other words, the merge candidate list can be used in both merge mode and skip mode. In this regard, the merge candidate list can also be referred to as a "skip candidate list" or a "merge / skip candidate list."
[0498] Alternatively, the skip mode may use an additional candidate list different from the candidate list of the merge mode. In this case, in the following description, the merge candidate list and the merge candidate may be replaced by the skip candidate list and the skip candidate, respectively.
[0499] A merge candidate list may be created before performing prediction in skip mode.
[0500] 3-2) Searching for motion vectors using the merge candidate list
[0501] The encoding device 100 may determine a merge candidate to be used for encoding the target block. For example, the encoding device 100 may perform prediction on the target block using the merge candidate in the merge candidate list. The encoding device 100 may encode the target block using the merge candidate that generates the minimum cost in prediction.
[0502] Also, the encoding apparatus 100 may determine whether to encode the target block using the skip mode.
[0503] 3-3) Transmission of inter-frame prediction information
[0504] The encoding apparatus 100 may generate a bitstream including inter prediction information required for inter prediction, and the decoding apparatus 200 may perform inter prediction on a target block using the inter prediction information of the bitstream.
[0505] The inter prediction information may include 1) mode information indicating whether the skip mode is used and 2) a skip index.
[0506] The skip index may be the same as the merge index described above.
[0507] When skip mode is used, the target block may be encoded without using a residual signal. The inter prediction information may not include a residual signal. Alternatively, the bitstream may not include a residual signal.
[0508] The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the skip mode is used. As described above, the merge index and the skip index may be the same as each other. The decoding device 200 may obtain the skip index from the bitstream only when the mode information indicates that the merge mode or the skip mode is used.
[0509] The skip index may indicate a merge candidate to be used for prediction of a target block among merge candidates included in the merge candidate list.
[0510] 3-4) Inter-frame prediction in skip mode using inter-frame prediction information
[0511] The decoding apparatus 200 may perform prediction on the target block using the merge candidate indicated by the skip index among the merge candidates included in the merge candidate list.
[0512] The motion vector of the target block may be specified by the motion vector of the merge candidate indicated by the skip index, the reference picture index, and the reference direction.
[0513] 4) Current picture reference mode
[0514] The current picture reference mode may denote a prediction mode that uses a previously reconstructed area in a target picture to which the target block belongs.
[0515] A motion vector for specifying a previously reconstructed area may be used.A reference picture index of a target block may be used to determine whether the target block has been encoded in a current picture reference mode.
[0516] A flag or index indicating whether the target block is a block encoded in the current picture reference mode may be signaled by the encoding apparatus 100 to the decoding apparatus 200. Alternatively, whether the target block is a block encoded in the current picture reference mode may be inferred through the reference picture index of the target block.
[0517] When the target block is encoded in the current picture reference mode, the current picture may exist at a fixed position or an arbitrary position in the reference picture list for the target block.
[0518] For example, the fixed position may be a position where the value of the reference picture index is 0 or a last position.
[0519] When a target picture exists at an arbitrary position in the reference picture list, an additional reference picture index indicating such an arbitrary position may be signaled by the encoding apparatus 100 to the decoding apparatus 200 .
[0520] In the AMVP mode, merge mode, and skip mode described above, an index of a list may be used to specify motion information to be used for prediction of a target block among a plurality of pieces of motion information in the list.
[0521] To improve encoding efficiency, the encoding apparatus 100 may signal only the index of the element generating the minimum cost in inter-frame prediction of the target block among the elements in the list. The encoding apparatus 100 may encode the index and signal the encoded index.
[0522] Therefore, the encoding device 100 and the decoding device 200 must be able to derive the above-described lists (i.e., the prediction motion vector candidate list and the merge candidate list) based on the same data using the same scheme. Here, the same data may include reconstructed pictures and reconstructed blocks. In addition, in order to specify an element using an index, the order of the elements in the list must be fixed.
[0523] Figure 10 Illustrated are spatial candidates according to an embodiment.
[0524] exist Figure 10 , the positions of the spatial candidates are shown.
[0525] The large block in the center of the figure may represent the target block, and the five small blocks may represent spatial candidates.
[0526] The coordinates of the target block may be (xP, yP), and the size of the target block may be represented by (nPSW, nPSH).
[0527] The spatial candidate A0 may be a block adjacent to the lower left corner of the target block. A0 may be a block occupying a pixel located at coordinates (xP-1, yP+nPSH+1).
[0528] Spatial candidate A1 may be a block adjacent to the left side of the target block. A1 may be the lowest block among the blocks adjacent to the left side of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block occupying a pixel located at coordinates (xP-1, yP+nPSH).
[0529] The spatial candidate B0 may be a block adjacent to the upper right corner of the target block. B0 may be a block occupying a pixel located at coordinates (xP+nPSW+1, yP-1).
[0530] Spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Alternatively, B1 may be a block adjacent to the left of B0. B1 may be a block occupying a pixel located at coordinates (xP+nPSW, yP-1).
[0531] The spatial candidate B2 may be a block adjacent to the upper left corner of the target block. B2 may be a block occupying a pixel located at coordinates (xP-1, yP-1).
[0532] Determination of availability of spatial and temporal candidates
[0533] In order to include the motion information of the spatial candidate or the motion information of the temporal candidate in the list, it is necessary to determine whether the motion information of the spatial candidate or the motion information of the temporal candidate is available.
[0534] Hereinafter, candidate blocks may include spatial candidates and temporal candidates.
[0535] For example, the determination may be performed by sequentially applying the following steps 1) to 4) below.
[0536] Step 1) When the PU including the candidate block is located outside the boundary of the picture, the availability of the candidate block may be set to “false.” The expression “availability is set to false” may have the same meaning as “set to unavailable.”
[0537] Step 2) When the PU including the candidate block is located outside the boundary of the slice, the availability of the candidate block may be set to “false.” When the target block and the candidate block are located in different slices, the availability of the candidate block may be set to “false.”
[0538] Step 3) When the PU including the candidate block is located outside the boundary of the tile, the availability of the candidate block may be set to “false.” When the target block and the candidate block are located in different tiles, the availability of the candidate block may be set to “false.”
[0539] Step 4) When the prediction mode of the PU including the candidate block is intra prediction mode, the availability of the candidate block may be set to “false.” When the PU including the candidate block does not use inter prediction, the availability of the candidate block may be set to “false.”
[0540] Figure 11 An order in which motion information of spatial candidates is added to a merge list according to an embodiment is shown.
[0541] like Figure 11 As shown in , when multiple pieces of motion information of spatial candidates are added to the merge list, the order of A1, B1, B0, A0, and B2 may be used. That is, multiple pieces of motion information of available spatial candidates may be added to the merge list in the order of A1, B1, B0, A0, and B2.
[0542] Methods for deriving merge lists in merge mode and skip mode
[0543] As described above, the maximum number of merge candidates in a merge list can be set. The set maximum number can be indicated by "N". The set number can be transmitted from the encoding device 100 to the decoding device 200. The slice header of the slice may include N. In other words, the maximum number of merge candidates in the merge list for the target block of the slice can be set by the slice header. For example, the value of N can basically be 5.
[0544] A plurality of pieces of motion information (ie, merge candidates) may be added to the merge list in the order of the following steps 1) to 4) below.
[0545] Step 1) Among the space candidates, available space candidates can be added to the merge list. Figure 10 The plurality of pieces of motion information of the available spatial candidates are added to the merge list in the order shown in FIG. Here, when the motion information of the available spatial candidate overlaps with other motion information already in the merge list, the motion information of the available spatial candidate may not be added to the merge list. The operation of checking whether the corresponding motion information overlaps with other motion information in the list may be simply referred to as "overlap check."
[0546] The maximum number of pieces of motion information added may be N.
[0547] Step 2) When the number of motion information items in the merge list is less than N and a temporal candidate is available, the motion information of the temporal candidate may be added to the merge list. Here, when the motion information of the available temporal candidate overlaps with other motion information already in the merge list, the motion information of the available temporal candidate may not be added to the merge list.
[0548] Step 3) When the number of pieces of motion information in the merge list is less than N and the type of the target slice is 'B', combined motion information generated by combining bidirectional predictions (bi-prediction) may be added to the merge list.
[0549] The target slice may be a slice including the target block.
[0550] The combined motion information may be a combination of L0 motion information and L1 motion information. The L0 motion information may be motion information referring only to the reference picture list L0. The L1 motion information may be motion information referring only to the reference picture list L1.
[0551] In the merge list, there may be one or more pieces of L0 motion information. In addition, in the merge list, there may be one or more pieces of L1 motion information.
[0552] The combined motion information may include one or more pieces of combined motion information. When generating the combined motion information, the L0 motion information and the L1 motion information to be used in the step of generating the combined motion information may be predefined among the one or more pieces of L0 motion information and the one or more pieces of L1 motion information. The one or more pieces of combined motion information may be generated in a predefined order by combining bidirectional prediction using a pair of different motion information in a merge list. One piece of the pair of different motion information may be L0 motion information, and the other piece of the pair of different motion information may be L1 motion information.
[0553] For example, the combined motion information added with the highest priority may be a combination of L0 motion information with a merge index of 0 and L1 motion information with a merge index of 1. When the motion information with a merge index of 0 is not L0 motion information or when the motion information with a merge index of 1 is not L1 motion information, the combined motion information may be neither generated nor added. Next, the combined motion information added with the next highest priority may be a combination of L0 motion information with a merge index of 1 and L1 motion information with a merge index of 0. The subsequent detailed combinations may conform to other combinations in the field of video encoding / decoding.
[0554] Here, when the combined motion information overlaps with other motion information already present in the merge list, the combined motion information may not be added to the merge list.
[0555] Step 4) When the number of pieces of motion information in the merge list is less than N, motion information of a zero vector may be added to the merge list.
[0556] The zero-vector motion information may be motion information in which a motion vector is a zero vector.
[0557] The number of pieces of zero-vector motion information may be one or more. The reference picture indexes of one or more pieces of zero-vector motion information may be different from each other. For example, the reference picture index value of the first zero-vector motion information may be 0. The reference picture index value of the second zero-vector motion information may be 1.
[0558] The number of pieces of zero-vector motion information may be the same as the number of reference pictures in the reference picture list.
[0559] The reference direction of the zero-vector motion information may be bidirectional. Both motion vectors may be zero vectors. The number of pieces of zero-vector motion information may be the smaller of the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1. Alternatively, when the number of reference pictures in reference picture list L0 and the number of reference pictures in reference picture list L1 are different, a unidirectional reference direction may be used for a reference picture index applicable only to a single reference picture list.
[0560] The encoding apparatus 100 and / or the decoding apparatus 200 may then add zero-vector motion information to the merge list while changing the reference picture index.
[0561] When the zero-vector motion information overlaps with other motion information already in the merge list, the zero-vector motion information may not be added to the merge list.
[0562] The order of steps 1) to 4) is merely exemplary and may be changed. In addition, some of the steps above may be omitted according to predefined conditions.
[0563] Method for deriving a candidate list of motion vector prediction in AMVP mode
[0564] The maximum number of motion vector predictor candidates in the motion vector predictor candidate list may be predefined. The predefined maximum number may be indicated by N. For example, the predefined maximum number may be 2.
[0565] A plurality of pieces of motion information (ie, prediction motion vector candidates) may be added to the prediction motion vector candidate list in the order of steps 1) to 3) below.
[0566] Step 1) Available spatial candidates among the spatial candidates may be added to the motion vector prediction candidate list.The spatial candidates may include a first spatial candidate and a second spatial candidate.
[0567] The first spatial candidate may be one of A0, A1, scaled A0, and scaled A1. The second spatial candidate may be one of B0, B1, B2, scaled B0, scaled B1, and scaled B2.
[0568] Multiple pieces of motion information of available spatial candidates may be added to the motion vector prediction candidate list in the order of the first spatial candidate and the second spatial candidate. In this case, if the motion information of an available spatial candidate overlaps with other motion information already in the motion vector prediction candidate list, the motion information of the available spatial candidate may not be added to the motion vector prediction candidate list. In other words, when the value of N is 2, if the motion information of the second spatial candidate is the same as the motion information of the first spatial candidate, the motion information of the second spatial candidate may not be added to the motion vector prediction candidate list.
[0569] The maximum number of pieces of motion information added may be N.
[0570] Step 2) When the number of pieces of motion information in the motion vector predictor candidate list is less than N and a temporal candidate is available, the motion information of the temporal candidate may be added to the motion vector predictor candidate list. In this case, if the motion information of the available temporal candidate overlaps with other motion information already in the motion vector predictor candidate list, the motion information of the available temporal candidate may not be added to the motion vector predictor candidate list.
[0571] Step 3) When the number of pieces of motion information in the motion vector predictor candidate list is less than N, zero vector motion information may be added to the motion vector predictor candidate list.
[0572] The zero-vector motion information may include one or more pieces of zero-vector motion information. Reference picture indices of the one or more pieces of zero-vector motion information may be different from each other.
[0573] The encoding apparatus 100 and / or the decoding apparatus 200 may sequentially add a plurality of pieces of zero-vector motion information to the prediction motion vector candidate list while changing the reference picture index.
[0574] When the zero-vector motion information overlaps with other motion information already existing in the motion vector predictor candidate list, the zero-vector motion information may not be added to the motion vector predictor candidate list.
[0575] The above description of the zero vector motion information in conjunction with the merge list is also applicable to the zero vector motion information, and its repeated description will be omitted.
[0576] The order of steps 1) to 3) described above is merely exemplary and may be changed. In addition, some of the steps may be omitted according to predefined conditions.
[0577] Figure 12 Transformation and quantization processes according to examples are shown.
[0578] like Figure 12 As shown in , the quantized levels may be generated by performing a transform and / or quantization process on the residual signal.
[0579] The residual signal may be generated as a difference between the original block and the predicted block.Here, the predicted block may be a block generated via intra prediction or inter prediction.
[0580] The residual signal may be transformed into a signal in the frequency domain through a transform process as part of the quantization process.
[0581] The transform kernels used for the transform may include various DCT kernels, such as discrete cosine transform (DCT) type 2 (DCT-II) and discrete sine transform (DST) kernels.
[0582] These transform kernels may perform separable transform or two-dimensional (2D) non-separable transform on the residual signal. A separable transform may be a transform indicating that a one-dimensional (1D) transform is performed on the residual signal in each of the horizontal and vertical directions.
[0583] The DCT type and DST type adaptively used for 1D transform may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II, as shown in Table 3 below.
[0584] Table 3
[0585]
[0586]
[0587] As shown in Table 3, when deriving the DCT type or DST type to be used for transformation, a transform set can be used. Each transform set can include multiple transform candidates. Each transform candidate can be a DCT type or a DST type.
[0588] Table 4 below shows an example of a transform set applied to the horizontal direction according to the intra prediction mode.
[0589] Table 4
[0590]
[0591]
[0592] In Table 4, the number of each transform set in the horizontal direction to be applied to the residual signal is indicated according to the intra prediction mode of the target block.
[0593] Table 5 below shows an example of a transform set applied to the vertical direction of the residual signal according to the intra prediction mode.
[0594] Table 5
[0595]
[0596]
[0597] As illustrated in Tables 4 and 5, the transform sets to be applied in the horizontal and vertical directions may be predefined according to the intra prediction mode of the target block. The encoding apparatus 100 may perform transform and inverse transform on the residual signal using the transform included in the transform set corresponding to the intra prediction mode of the target block. In addition, the decoding apparatus 200 may perform inverse transform on the residual signal using the transform included in the transform set corresponding to the intra prediction mode of the target block.
[0598] In the transformation and inverse transformation, the transform set to be applied to the residual signal may be determined and may not be signaled as illustrated in Tables 3, 4, and 5. Transform indication information may be signaled from the encoding apparatus 100 to the decoding apparatus 200. The transform indication information may be information indicating which one of a plurality of transform candidates to be applied to the residual signal included in the transform set is to be used.
[0599] As described above, methods using various transforms may be applied to a residual signal generated via intra prediction or inter prediction.
[0600] The transformation may include at least one of a primary transformation and a secondary transformation. A transformation coefficient may be generated by performing the primary transformation on the residual signal, and a secondary transformation coefficient may be generated by performing the secondary transformation on the transformation coefficient.
[0601] The first transform may be referred to as a "primary transform." Furthermore, the first transform may also be referred to as an "adaptive multi-transform (AMT) scheme." As described above, AMT may refer to applying different transforms to each 1D direction (i.e., vertical direction and / or horizontal direction) or selected directions.
[0602] Alternatively, AMT may be referred to as Multi-Transform Select (MTS) or Extended Multi-Transform (EMT).
[0603] The secondary transform may be a transform for improving the energy concentration of the transform coefficients generated by the primary transform. Similar to the primary transform, the secondary transform may be a separable transform or a non-separable transform. Such a non-separable transform may be a non-separable secondary transform (NSST).
[0604] The first transform may be performed using at least one of a plurality of predefined transform methods, such as discrete cosine transform (DCT), discrete sine transform (DST), Karhunen-Loeve transform (KLT), and the like.
[0605] Furthermore, the first transform may be a transform of various types depending on the kernel function defining the DCT or DST.
[0606] For example, according to the transform kernel presented in Table 6 below, the first transform may include transforms such as DCT-2, DCT-5, DST-7, DCT-7, DST-8, DST-1, and DCT-8. In Table 6 below, various transform types and transform kernel functions for multi-transform selection (MTS) are exemplified.
[0607] MTS may refer to the selection of a combination of one or more DCT and / or DST kernels to transform the residual signal in the horizontal and / or vertical directions.
[0608] Table 6
[0609]
[0610]
[0611] In Table 6, i and j may be integer values equal to or greater than 0 and less than or equal to N-1.
[0612] A secondary transform may be performed on transform coefficients generated by performing the first transform.
[0613] The primary transform and / or secondary transform may be applied to signal components corresponding to one or more of a luma component and a chroma component. Whether to apply the primary transform and / or secondary transform may be determined based on at least one of the coding parameters for the target block and / or the neighboring blocks. For example, whether to apply the primary transform and / or secondary transform may be determined based on the size and / or shape of the target block.
[0614] The transform method to be applied to the first transform and / or the second transform may be determined based on at least one of the encoding parameters for the target block and / or the neighboring blocks. The determined transform method may also indicate that the first transform and / or the second transform is not used.
[0615] Alternatively, transform information indicating a transform method may be signaled from the encoding apparatus 100 to the decoding apparatus 200. For example, the transform information may include an index of a transform to be used for a primary transform and / or a secondary transform.
[0616] A quantized transform coefficient (ie, a quantization level) may be generated by performing quantization on a result generated by performing the primary transform and / or the secondary transform or performing quantization on a residual signal.
[0617] Figure 13 A diagonal scan according to an example is shown.
[0618] Figure 14 A horizontal scan according to an example is shown.
[0619] Figure 15 A vertical scan according to an example is shown.
[0620] The quantized transform coefficients may be scanned via at least one of (upper right) diagonal scanning, vertical scanning, and horizontal scanning according to at least one of an intra prediction mode, a block size, and a block shape. The block may be a transform unit (TU).
[0621] Each scan may be initiated at a specific start point and may be terminated at a specific end point.
[0622] For example, by using Figure 13 The coefficients of the block are scanned by diagonal scanning to change the quantized transform coefficients into 1D vector form. Optionally, the quantized transform coefficients can be used according to the size of the block and / or the intra prediction mode. Figure 14 Horizontal scan or Figure 15 vertical scanning instead of diagonal scanning.
[0623] Vertical scanning may be an operation of scanning 2D block type coefficients in a column direction, and horizontal scanning may be an operation of scanning 2D block type coefficients in a row direction.
[0624] In other words, which of diagonal scanning, vertical scanning, and horizontal scanning is to be used may be determined according to the size of a block and / or an inter prediction mode.
[0625] like Figure 13 、 Figure 14 and Figure 15 As shown in , the quantized transform coefficients may be scanned along a diagonal direction, a horizontal direction, or a vertical direction.
[0626] The quantized transform coefficients can be represented by a block shape. Each block can include multiple sub-blocks. Each sub-block can be defined according to a minimum block size or a minimum block shape.
[0627] In the scanning, a scanning order according to a type or direction of scanning may be first applied to a subblock. In addition, a scanning order according to a direction of scanning may be applied to quantized transform coefficients in each subblock.
[0628] For example, Figure 13 、 Figure 14 and Figure 15 As shown in , when the size of the target block is 8×8, quantized transform coefficients can be generated by primary transform, secondary transform, and quantization of the residual signal of the target block. Therefore, one of three types of scanning orders can be applied to the four 4×4 sub-blocks, and the quantized transform coefficients can also be scanned for each 4×4 sub-block according to the scanning order.
[0629] The scanned quantized transform coefficients may be entropy encoded, and the bitstream may include the entropy encoded quantized transform coefficients.
[0630] The decoding apparatus 200 may generate quantized transform coefficients by entropy decoding a bitstream. The quantized transform coefficients may be arranged in the form of 2D blocks by inverse scanning. Here, as the inverse scanning method, at least one of upper right diagonal scanning, vertical scanning, and horizontal scanning may be performed.
[0631] Inverse quantization may be performed on the quantized transform coefficients. Depending on whether the secondary inverse transform is to be performed, a secondary inverse transform may be performed on the result generated by performing the inverse quantization. Furthermore, depending on whether the primary inverse transform is to be performed, a primary inverse transform may be performed on the result generated by performing the secondary inverse transform. A reconstructed residual signal may be generated by performing the primary inverse transform on the result generated by performing the secondary inverse transform.
[0632] Figure 16 is a configuration diagram of an encoding device according to an embodiment.
[0633] The encoding apparatus 1600 may correspond to the encoding apparatus 100 described above.
[0634] The encoding apparatus 1600 may include a processing unit 1610, a memory 1630, a user interface (UI) input device 1650, a UI output device 1660, and a storage 1640 communicating with each other through a bus 1690. The encoding apparatus 1600 may further include a communication unit 1620 connected to a network 1699.
[0635] The processing unit 1610 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1630 or the storage 1640. The processing unit 1610 may be at least one hardware processor.
[0636] The processing unit 1610 may generate and process signals, data, or information input to, output from, or used in the encoding device 1600, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in embodiments, the generation and processing of data or information, as well as checks, comparisons, and determinations related to the data or information, may be performed by the processing unit 1610.
[0637] The processing unit 1610 may include an inter-frame prediction unit 110, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180 and a reference picture buffer 190.
[0638] At least some of the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy encoding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180, and the reference picture buffer 190 may be program modules and may communicate with an external device or system. The program modules may be included in the encoding apparatus 1600 in the form of an operating system, an application module, or other program modules.
[0639] The program modules may be physically stored in various types of known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device that can communicate with the encoding device 1200.
[0640] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to embodiments or for implementing abstract data types according to embodiments.
[0641] The program modules may be implemented using instructions or codes executed by at least one processor of the encoding device 1600 .
[0642] The processing unit 1610 can execute instructions or codes in the inter-frame prediction unit 110, the intra-frame prediction unit 120, the switch 115, the subtractor 125, the transform unit 130, the quantization unit 140, the entropy coding unit 150, the inverse quantization unit 160, the inverse transform unit 170, the adder 175, the filter unit 180 and the reference picture buffer 190.
[0643] The storage unit may represent the memory 1630 and / or the storage 1640. Each of the memory 1630 and the storage 1640 may be any of various types of volatile or non-volatile storage media. For example, the memory 1630 may include at least one of a read-only memory (ROM) 1631 and a random access memory (RAM) 1632.
[0644] The storage unit may store data or information used for the operation of the encoding apparatus 1600. In an embodiment, data or information of the encoding apparatus 1600 may be stored in the storage unit.
[0645] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, and the like.
[0646] The encoding device 1600 may be implemented in a computer system including a computer-readable storage medium.
[0647] The storage medium may store at least one module required for the operation of the encoding apparatus 1600. The memory 1630 may store at least one module and may be configured such that the at least one module is executed by the processing unit 1610.
[0648] Functions related to communication of data or information of the encoding apparatus 1600 may be performed through the communication unit 1620 .
[0649] For example, the communication unit 1620 may transmit the bitstream to the decoding apparatus 1600 which will be described later.
[0650] Figure 17 is a configuration diagram of a decoding device according to an embodiment.
[0651] The decoding apparatus 1700 may correspond to the decoding apparatus 200 described above.
[0652] The decoding apparatus 1700 may include a processing unit 1710, a memory 1730, a user interface (UI) input device 1750, a UI output device 1760, and a storage 1740 communicating with each other through a bus 1790. The decoding apparatus 1700 may also include a communication unit 1720 connected to a network 1799.
[0653] The processing unit 1710 may be a central processing unit (CPU) or a semiconductor device for executing processing instructions stored in the memory 1730 or the storage 1740. The processing unit 1710 may be at least one hardware processor.
[0654] The processing unit 1710 may generate and process signals, data, or information input to, output from, or used in the decoding device 1700, and may perform checks, comparisons, determinations, etc. related to the signals, data, or information. In other words, in embodiments, the generation and processing of data or information, as well as checks, comparisons, and determinations related to the data or information, may be performed by the processing unit 1710.
[0655] The processing unit 1710 may include an entropy decoding unit 210 , an inverse quantization unit 220 , an inverse transform unit 230 , an intra prediction unit 240 , an inter prediction unit 250 , a switch 245 , an adder 255 , a filter unit 260 , and a reference picture buffer 270 .
[0656] At least some of the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the adder 255, the switch 245, the filter unit 260, and the reference picture buffer 270 of the decoding device 200 may be program modules and may communicate with an external device or system. The program modules may be included in the decoding device 1700 in the form of an operating system, an application module, or other program modules.
[0657] The program modules may be physically stored in various types of well-known storage devices. In addition, at least some of the program modules may also be stored in a remote storage device that can communicate with the decoding device 1700.
[0658] Program modules may include, but are not limited to, routines, subroutines, programs, objects, components, and data structures for performing functions or operations according to embodiments or for implementing abstract data types according to embodiments.
[0659] The program modules may be implemented using instructions or codes executed by at least one processor of the decoding device 1700 .
[0660] The processing unit 1710 can execute instructions or codes in the entropy decoding unit 210, the inverse quantization unit 220, the inverse transform unit 230, the intra-frame prediction unit 240, the inter-frame prediction unit 250, the switch 245, the adder 255, the filter unit 260 and the reference picture buffer 270.
[0661] The storage unit may represent the memory 1730 and / or the storage 1740. Each of the memory 1730 and the storage 1740 may be any of various types of volatile or non-volatile storage media. For example, the memory 1730 may include at least one of the ROM 1731 and the RAM 1732.
[0662] The storage unit may store data or information used for the operation of the decoding apparatus 1700. In an embodiment, data or information of the decoding apparatus 1700 may be stored in the storage unit.
[0663] For example, the storage unit may store pictures, blocks, lists, motion information, inter-frame prediction information, bitstreams, and the like.
[0664] The decoding device 1700 may be implemented in a computer system including a computer-readable storage medium.
[0665] The storage medium may store at least one module required for the operation of the decoding apparatus 1700. The memory 1730 may store at least one module and may be configured such that the at least one module is executed by the processing unit 1710.
[0666] Functions related to communication of data or information of the decoding apparatus 1700 may be performed through the communication unit 1720 .
[0667] For example, the communication unit 1720 may receive a bitstream from the encoding apparatus 1600 .
[0668] Image processing method using information sharing between channels
[0669] According to the methods and devices of the embodiments, transform coding (transcoding) techniques using prediction and various transformations may be applied to high-resolution images (such as 4K or 8K resolution images), images may be encoded and / or decoded through various types of predefined coding decision information between shared channels, and compressed bit streams or compressed data for encoded images may be decoded through the coding decision information sent between shared channels.
[0670] The multiple channels may represent multiple components of a block. For example, the multiple channels may include a color channel, a depth channel, an alpha channel, etc.
[0671] Hereinafter, the terms "channel" and "color" may have the same meaning and may be used interchangeably with each other. In addition, the term "color" may indicate one of the channels. The term "channel" may be used interchangeably with one or more of the terms "color," "depth," and "alpha."
[0672] The technology of this embodiment can be used to solve the problem of compression rate and image quality degradation that occurs when conventional technology is applied to image encoding and decoding. Specifically, when conventional technology is applied to an image in which changes in pixel values are concentrated in space, the problem of compression rate and image quality degradation may be serious.
[0673] In an example, as various types of encoding decision information shared between channels to perform encoding and decoding according to an embodiment, there are the following pieces of information: In the names of the following information, "flag" may be omitted.
[0674] 1) Transform skip flag (transform_skip_flag) information may indicate whether to selectively skip transform. Alternatively, the transform_skip_flag information may indicate one of using transform and skipping transform.
[0675] 2) Intra smoothing filter information may indicate whether smoothing filtering is applied to reference pixels used in intra prediction.
[0676] 3) Position-dependent intra prediction combination (PDPC) flag (PDPC_flag) information may indicate whether intra prediction will be performed by using neighboring pixels to which smoothing (i.e., filtering) is applied and neighboring pixels to which smoothing (i.e., filtering) is not applied when a specific intra prediction (e.g., planar prediction) is performed.
[0677] 4) Residual differential pulse code modulation (RDPCM) flag (rdpcm_flag) information may indicate whether differential pulse code modulation (DPCM) is additionally performed on a residual signal acquired through primary prediction and RDPCM of the residual signal is acquired again will be performed.
[0678] 5) Multiple transform selection (MTS) flag (mts_flag) information may indicate whether an extended multiple transform (EMT) based encoding method is to be used.
[0679] EMT may be an encoding method that selects and uses a designated transform for a transform block that is a target block from among a plurality of provided transforms.
[0680] EMT may also stand for "Enhanced Multi-Transform" and may also indicate "Multi-Transform Select (MTS)".
[0681] 6) EMT flag information may indicate whether EMT will be used.
[0682] 7) MTS index (mts_idx) information may indicate which transforms will be used in the horizontal and vertical directions when MTS is used.
[0683] A portion of the mts_idx information (eg, one designated bit in mts_idx) may be information indicating transform used in the horizontal direction of the residual signal.
[0684] Another part of the mts_idx information or a part of the rest of the mts_idx information (eg, one other designated bit in mts_idx) may be information indicating a transform used in the vertical direction of the residual signal.
[0685] For example, determination of transformation according to mts_idx information may be configured as shown in Table 7 and Table 8 below.
[0686] [Table 7]
[0687]
[0688]
[0689] [Table 8]
[0690]
[0691] Table 7 illustrates transform in the horizontal direction and transform in the vertical direction used according to the intra prediction mode and the value of mts_idx information.
[0692] According to Table 7, when mts_idx information is acquired in order to perform intra prediction or inter prediction of a target block, transform in a horizontal direction and transform in a vertical direction to be used for transform of the target block may be determined according to the value of the mts_idx information.
[0693] For example, when the intra prediction mode of the target block is 6 and the value of the mts_idx information is 2, DCT-7 may be used as transform in the horizontal direction, and DCT-8 may be used as transform in the vertical direction.
[0694] Table 8 shows an example of a modification of Table 7. In Table 8, MTS_CU_flag may indicate that flag information mts_flag is determined and transmitted based on a CU, wherein the flag information mts_flag indicates whether the multi-transform selection (MTS) method is used. In addition, MTS_Hor_flag and MTS_Ver_flag may indicate the transform used in the horizontal direction and the transform used in the vertical direction, respectively. Table 8 may exemplify the transform used in the horizontal direction and the transform used in the vertical direction through the values of MTS_Hor_flag and MTS_Ver_flag.
[0695] Alternatively, determination of transformation according to the mts_idx information of Table 7 may be configured as shown in Table 9 below.
[0696] [Table 9]
[0697]
[0698]
[0699] Table 9 illustrates horizontal transform types and vertical transform types used according to the value of mts_idx information in intra prediction and inter prediction.
[0700] Each value of the horizontal transform type can indicate a specific transform. For example, a horizontal transform type value of "1" can indicate DST-7. A horizontal transform type value of "2" can indicate DCT-8.
[0701] 8) Non-separable secondary transform (NSST) flag (nsst_flag) information may indicate whether an NSST encoding method for additionally performing a non-separable secondary transform on all or some transform coefficients obtained via a primary transform will be used.
[0702] 9) NSST index (nsst_idx) information may indicate the type of secondary transform to be applied to all or some transform coefficients when the NSST encoding method is used.
[0703] The nsst_idx information may indicate a transform to be used for the non-separable secondary transform.
[0704] 10) CU skip flag information may indicate whether a step of transmitting encoded data about a CU will be skipped.
[0705] 11) CU Local Illumination Compensation (LIC) flag (CU_lic_flag) information may indicate whether to compensate for a difference between luminance values of a block.
[0706] 12) Overlapped Block Motion Compensation (OBMC) flag (obmc_flag) information may indicate whether a plurality of overlapped motion compensation blocks are used to generate a final motion compensation block.
[0707] 13) codeAlfCtuEnable flag (codeAlfCtuEnable_flag) information may indicate whether an adaptive loop filter (ALF) may be applied to pixel values of the current CTU.
[0708] When such encoding decision information is shared between channels, an image with excellent image quality can be obtained while improving the image compression rate.
[0709] When describing the encoding decision information to be shared between channels according to this embodiment, in order to facilitate the overall description and understanding of the embodiment (such as description of operations, drawings, and equations), transform_skip_flag information can be used as an example of the encoding decision information to be shared between channels.
[0710] However, the transform_skip_flag information is only a single example, and encoding decision information to be shared between channels to which the present embodiment is applied does not necessarily represent only the transform_skip_flag information.
[0711] For example, it should be understood that one or more of the above-mentioned multiple encoding decision information required for decoding (such as 1) rdpcm_flag information, 2) multiple transform-related selection information such as mts_flag information, mts_idx information, nsst_flag information and nsst_idx information, 3) obmc_flag information and 4) PDPC_flag information) are included in the encoding decision information to be shared between channels.
[0712] In addition, when describing the channels that will share the encoding decision information required for decoding according to the embodiment, the YCbCr color space may be used as an example. However, the YCbCr color space is only a single detailed example, and the embodiment can be applied to various color spaces such as the YUV color space, the XYZ color space, and the RGB color space.
[0713] The color index cIDX may be a channel index indicating one of channels in a color space.
[0714] For the YCbCr color space and the YUV color space, cIDX may have values such as "0 / 1 / 2" for channels displayed sequentially in the corresponding color space. The values "a / b / c" may indicate that the value of cIDX indicating the first channel is 'a', the value of cIDX indicating the second channel is 'b', and the value of cIDX indicating the third channel is 'c'.
[0715] Alternatively, for the YCbCr color space and the YUV color space, cIDX may have values such as “0 / 2 / 1” for channels sequentially displayed in the corresponding color space.
[0716] For the RGB color space and the XYZ color space, cIDX may have values such as “1 / 0 / 2” or “2 / 0 / 1” for channels sequentially displayed in the corresponding color space.
[0717] As image compression technologies that have been developed or are being developed for the purpose of achieving efficient image encoding / decoding, there are various technologies, such as 1) inter-frame prediction technology that predicts the values of pixels included in a target picture from pictures before or after the target picture, 2) intra-frame prediction technology that predicts the values of pixels included in the current target picture using information of pixels in the target picture, 3) transform and quantization technology that compresses the energy of a residual signal remaining as a prediction error, 4) entropy coding technology that assigns short codes to values that appear more frequently and long codes to values that appear less frequently, and arithmetic coding technology. By utilizing these image compression technologies, image data can be efficiently compressed, transmitted, and stored.
[0718] There are various compression techniques that can be applied to encoding an image. In addition, depending on the properties of the image to be encoded, certain compression techniques may be more advantageous than other compression techniques. Therefore, the encoding device 1600 can perform the most advantageous compression on the target block by adaptively determining whether to use any one of a plurality of compression techniques of various types for the target block.
[0719] Therefore, in order to select the most advantageous compression technique for the target block from various selectable compression techniques, the encoding device 1600 may generally perform rate-distortion optimization (RDO). From a rate-distortion perspective, it may not be known in advance which of the various image coding decisions that can be selected for encoding the image is optimal. Therefore, the encoding device 1600 may calculate rate-distortion values for all combinations of available image coding decisions by performing encoding (or simplified encoding) on each combination of all available image coding decisions, and may determine and use the image coding decision having the minimum rate-distortion value among the calculated rate-distortion values as the final image coding decision for the target block.
[0720] In addition, the encoding device 1600 may record, in a bitstream, an encoding decision derived by performing such RDO or using an additional decision method selected by the encoding device 1600. The decoding device 1700 may read (i.e., parse) the encoding decision recorded in the bitstream and accurately perform decoding on the target block by performing an inverse process corresponding to encoding according to the encoding decision.
[0721] Here, information indicating an encoding decision may be referred to as "encoding decision information" or "encoding information" required for decoding.
[0722] Hereinafter, the terms “coding decision information” and “coding information” may have the same meaning and may be used interchangeably with each other.
[0723] Typically, multiple channels (e.g., YUV, YCbCr, RGB, and XYZ) used for an image may not always have the same or similar properties. Therefore, from the perspective of improving compression rate, making encoding decisions independently for each of the multiple channels can often achieve better performance.
[0724] For example, as one of the above-mentioned encoding decisions, there is a transform_skip_flag which is an encoding decision indicating whether to perform a transform on a target block. That is, it is possible to determine whether to skip the transform for each of the blocks, and the transform_skip_flag information indicating such a decision can be recorded as encoding decision information in the bitstream for each of the multiple channels.
[0725] Generally, in encoding for image compression, it has been considered that a transform is always performed on a target block. However, when the spatial change in the value of a pixel in the target block to be compressed is very large, or particularly when the change in the pixel value is very locally limited, even if a transform is applied, the degree to which the image energy is concentrated in the low frequencies may not be very large, and instead, a large number of transform coefficients with relatively large values for high-frequency areas may appear.
[0726] Therefore, when a large amount of low-frequency signal components are retained and high-frequency signal components are eliminated through transformation and quantization processing, or when a transformation and quantization processing for reducing the amount of data by applying strong quantization is applied, severe degradation of image quality may occur. In particular, when the spatial change of pixel values is very large or the change of pixel values is concentrated in a very locally limited area, this degradation of image quality may be further enhanced.
[0727] To address the above issues, a method for directly encoding pixel values in the spatial domain without requiring a transform can be used, rather than uniformly applying a transform to the target block. This method determines whether to perform a transform on each transform block. By performing or skipping the transform based on this decision, encoding of the transform block can be performed. The bitstream may include transform_skip_flag information, which indicates whether to skip performing the transform.
[0728] For example, when the value of the transform_skip_flag information is 1, the transformation may be skipped. When the value of the transform_skip_flag information is 0, the transformation may be performed. The encoding device 1600 may transmit information on whether the transformation is to be skipped for the target block to the decoding device 1700 through the transform_skip_flag information, and the above-mentioned problem may be solved with the help of such transmission.
[0729] In addition, a plurality of transform_skip_flag information may be set for each of the luma channel (ie, Y channel) and the chroma channel (ie, Cb channel and Cr channel), and then the plurality of transform_skip_flag information may be transmitted. The decoding apparatus 1700 may decode the target block by skipping or performing a transform on the channel of the target block according to the value of the transform_skip_flag information for each channel read (ie, parsed) from the bitstream.
[0730] However, when a plurality of pieces of transform_skip_flag information for channels such as Y, Cb, and Cr are transmitted for all transform blocks, another problem may occur in that overhead may increase due to signaling of the plurality of pieces of transform_skip_flag information and a compression rate of an image may deteriorate.
[0731] To mitigate issues such as compression ratio degradation, flag information indicating whether transforms will be skipped can be limited to being transmitted only when the transform block size is less than or equal to a specific transform block size. However, even with this solution, multiple pieces of flag information indicating whether transforms will be skipped must be transmitted for all channels for each transform block size greater than the specific block size, which can still degrade image compression ratio. Furthermore, this degraded compression ratio inevitably reduces the quality of the compressed image.
[0732] In order to solve the degradation of the compression rate caused by transmitting a plurality of pieces of encoding decision information selected by the encoding device 1600 for all channels, an encoding and / or decoding method using sharing of information between channels is disclosed in an embodiment.
[0733] First, conditions for determining that image attributes of channels are similar to each other may be predefined. When these conditions are met, encoding decision information for an image or block determined by the encoding device 1600 for a representative channel among multiple channels may be transmitted to the decoding device 1700.
[0734] The encoding decision information transmitted for the representative channel can be shared and used by all channels other than the representative channel, or channels selected from the remaining channels. This sharing and use improves the image compression rate. Therefore, even without transmitting individual pieces of encoding decision information for multiple channels, the encoding and / or decoding method according to this embodiment can provide excellent encoding efficiency.
[0735] Here, the encoding decision information to be shared may include one or more of the above-mentioned transform_skip_flag information, intra-frame smoothing filtering information, rdpcm_flag information, mts_flag information, mts_idx information, PDPC_flag information, MTS_CU_flag information, MTS_Hor_flag information, MTS_Ver_flag information, nsst_flag information, nsst_idx information, CU skip flag information, CU_lic_flag information, obmc_flag information, codeAlfCtuEnable_flag information and PDPC_flag information.
[0736] The image properties of each channel are determined to be similar to each other under the condition
[0737] In order to determine that image properties of channels are similar to each other, it may be checked whether cross-channel prediction (inter-channel prediction) has been used for the target block.
[0738] That is, in order to predict the decoding target channel of the target block, it can be checked whether a prediction method for obtaining a prediction value for decoding the target channel by applying a specific model to reconstruction information of another channel (e.g., a luma channel) is used. For example, the reconstruction information can be the pixel value of the reconstructed pixel or the value of the transform coefficient. The specific model can be a linear model.
[0739] A decoding target channel may be a channel currently being decoded among multiple channels. An encoding target channel may be a channel currently being encoded among multiple channels. Hereinafter, the encoding target channel and / or the decoding target channel may also be referred to as simply a "target channel."
[0740] For example, it may be checked whether intra prediction for a target block uses an intra prediction mode that derives a prediction value for a target channel by using reconstruction information of another channel.
[0741] In order to use the reconstruction information of another channel to derive the prediction value for the target channel, a cross-component linear model (CCLM) using a single linear model, a multi-mode linear model (MMLM) using multiple linear models, and a multi-filter linear model using multiple filters can be used. In CCLM, the term "component" can be replaced by "channel".
[0742] The INTRA_CCLM mode may be an intra prediction mode using CCLM, the INTRA_MMLM mode may be an intra prediction mode using MMLM, and the INTRA_MFLM mode may be an intra prediction mode using MFLM.
[0743] Alternatively, determination that image properties of channels are similar to each other may be achieved by checking whether the intra prediction mode of the target block (e.g., intra_chroma_pred_mode information indicating the intra prediction mode for the chroma channel of the target block) is one of the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode.
[0744] Alternatively, the determination that the image properties of the channels are similar to each other can be achieved by checking whether the encoding mode of the target channel of the target block uses the encoding mode of another channel (e.g., the luma channel) without change. For example, the determination that the image properties of the channels are similar to each other can be achieved by checking whether the intra prediction mode (e.g., intra_chroma_pred_mode information) of the target block is a direct mode (DM). The direct mode may also be referred to as a "derivation mode." The DM may be a mode that indicates that the intra prediction mode of the luma channel is used as the intra prediction mode of the chroma channel without change due to the characteristic that the correlation between the luma channel and the chroma channel may be high.
[0745] The characteristics of DM, one of the intra prediction modes, and its detailed operation may be defined in more detail with reference to Tables 10 and 11 below.
[0746] Table 10 shows a method for setting the IntraPredModeC value used for intra prediction of a chroma signal (when a value of the sps_cclm_enabled_flag information is 0).
[0747] Table 11 shows a method for setting the IntraPredModeC value used for intra prediction of a chroma signal (when a value of the sps_cclm_enabled_flag information is 1).
[0748] [Table 10]
[0749]
[0750] [Table 11]
[0751]
[0752] Generally, in intra prediction, it can be determined whether to use an intra cross-component linear model (CCLM) mode, an intra multi-model LM (MMLM) mode, or an intra multi-filter LM (MFLM) mode, in which pixel values of reconstructed pixels for a single channel (e.g., a luma channel, or more generally, a representative channel) are used to calculate prediction values for another channel (e.g., a chroma channel, or more generally, a target channel).
[0753] An indication of a case where the INTRA_CCLM mode, the INTRA_MMLM mode, or the INTRA_MFLM mode is used may be classified into two types and defined in detail according to a value of the sps_cclm_enabled_flag information, as shown in Tables 10 and 11.
[0754] The sps_cclm_enabled_flag information may be information indicating whether the INTRA_CCLM mode, the INTRA_MMLM mode, and the INTRA_MFLM mode are to be enabled. Alternatively, the sps_cclm_enabled_flag information may be information indicating whether the INTRA_CCLM mode, the INTRA_MMLM mode, and the INTRA_MFLM mode are already enabled.
[0755] When the value of sps_cclm_enabled_flag is 0, the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode may not be used, and when the value of sps_cclm_enabled_flag is 1, the INTRA_CCLM mode may be used. Alternatively, when the value of sps_cclm_enabled_flag is 1, at least one of the INTRA_MMLM mode and the INTRA_MFLM mode may be used.
[0756] Whether the intra prediction mode of the target channel (e.g., intra_chroma_pred_mode information) is DM can be determined by checking whether the intra prediction mode of the chroma channel (i.e., the value of intra_chroma_pred_mode) is a specific value (e.g., 4 in Table 10 and 5 in Table 11). In the description of this operation, when the value of sps_cclm_enabled_flag is 0, refer to Table 10, and when the value of sps_cclm_enabled_flag is 1, refer to Table 11.
[0757] When the value of sps_cclm_enabled_flag is 0 and the value of the intra prediction mode of the target channel (e.g., intra_chroma_pred_mode) is 4, DM may be considered to be applied. Alternatively, when the value of sps_cclm_enabled_flag is 1 and the value of the intra prediction mode of the target channel (e.g., intra_chroma_pred_mode) is 5, DM may be considered to be applied. For intra prediction of a target channel (e.g., a chroma channel) of a target block indicated by DM, the value of IntraPredModeY indicating the intra prediction mode of a representative channel (e.g., a luma channel) may be used as the value of IntraPredModeC without change.
[0758] Here, the intra prediction mode intra_chroma_pred_mode of the chroma signal may be index information indicating which type of intra prediction is to be used for the chroma signal.
[0759] With the help of such index information, the final value indicating the intra prediction mode actually used for intra prediction of the chrominance signal may be the value of IntraPredModeC. In other words, IntraPredModeC may indicate the intra prediction mode actually used for intra prediction of the chrominance signal.
[0760] When the value of sps_cclm_enabled_flag is 0 and DM is applied (ie, the value of intra_chroma_pred_mode is 4), if the value of IntraPredModeY is 0, 50, 18, or 1, the value of IntraPredModeC may also be 0, 50, 18, or 1.
[0761] Here, a value of 0 may represent a planar mode (ie, a planar prediction or a planar direction), a value of 1 may represent a DC mode, a value of 18 may represent a horizontal mode, a value of 50 may represent a vertical mode, and a value of 66 may represent a diagonal mode.
[0762] When the value of IntraPredModeY is another value X different from any one of the four values 0, 50, 18, and 1, the value of IntraPredModeC may also be X which is equal to the value of IntraPredModeY.
[0763] In addition, as shown in the first four rows in Table 10, when the value of cclm_enabled_flag is 0, if the value of IntraPredModeY is 0, 50, 18, or 1, the value of IntraPredModeC may be determined according to the value of IntraPredModeY.
[0764] For example, as described in the first row in Table 10, when the value of IntraPredModeY is 0, 50, 18, or 1, the value of IntraPredModeC may be 66, 0, 0, or 0. When the value of IntraPredModeY is a value other than 0, 50, 18, and 1, the value of IntraPredModeC may be 0.
[0765] In addition, when the value of sps_cclm_enabled_flag is 1 and DM is applied (ie, the value of intra_chroma_pred_mode is 5), if the value of IntraPredModeY is 0, 50, 18, or 1, the value of IntraPredModeC may also be 0, 50, 18, or 1.
[0766] Here, a value of 0 may represent a planar mode (ie, a planar prediction or a planar direction), a value of 1 may represent a DC mode, a value of 18 may represent a horizontal mode, a value of 50 may represent a vertical mode, and a value of 66 may represent a diagonal mode.
[0767] When the value of IntraPredModeY is another value X different from any one of the four values 0, 50, 18, and 1, the value of IntraPredModeC may also be X which is equal to the value of IntraPredModeY.
[0768] In addition, as shown in the first five rows in Table 11, when the value of cclm_enabled_flag is 1, if the value of IntraPredModeY is 0, 50, 18, or 1, the value of IntraPredModeC may be determined according to the value of IntraPredModeY.
[0769] For example, as described in the first row of Table 11, when the value of IntraPredModeY is 0, 50, 18, or 1, the value of IntraPredModeC may be 66, 0, 0, or 0. When the value of IntraPredModeY is a value other than 0, 50, 18, and 1, the value of IntraPredModeC may be 0.
[0770] In another embodiment, the determination that the image properties of the channels are similar to each other can be achieved by checking whether a mode indicating that only a specific mode limited by the encoding mode of another channel (e.g., the luma channel) is to be used is used as the encoding mode of the target channel for the target block. For example, the determination that the image properties of the channels are similar to each other can be performed by checking whether the intra prediction mode of the target channel is direct mode (DM).
[0771] Cross-channel prediction using correlations between channels
[0772] Cross-channel prediction may be a technique of using pixel values of pixels in another channel when predicting pixel values of pixels in a target channel instead of using intra prediction or inter prediction.
[0773] The fact that cross-channel prediction performs better than other types of prediction when the target block is encoded may indicate that there is considerable similarity between pixel values of pixels in channels of the target block.
[0774] Therefore, in this case, when the value of the encoding decision information of the representative channel is determined, it may be advantageous to use the determined value of the encoding decision information of the representative channel similarly for the encoding decision information of another channel or to use a specific value indicated by the value of the encoding decision information of the representative channel for the encoding decision information of another channel.
[0775] For example, when the value of the transform_skip_flag information of the representative channel is 0 (indicating that transform is not skipped), there may be a high probability that the determined value of the transform_skip_flag information will be '0' even in other channels.
[0776] Therefore, for an image or block for which cross-channel prediction is effective, there may be a situation where it is not necessary to specify multiple pieces of transform_skip_flag information separately for multiple channels. This is because the similarity between channels is high, and therefore the probability that multiple pieces of transform_skip_flag information for each channel will be the same as each other may be high.
[0777] Despite such image properties, when a plurality of pieces of transform_skip_flag information are respectively transmitted for channels of an image, compression rate and image quality may deteriorate.
[0778] This principle can also be applied to additional coding decision information, namely, mts_flag information, mts_idx information, nsst_flag information, nsst_idx information, intra-frame smoothing filter information, PDPC_flag information and rdpcm_flag information, and the probability that the value of the coding decision information for the representative channel will be the same as the value of the coding decision information for the other channels may be high.
[0779] Therefore, based on the condition that image properties of channels are determined to be similar to each other, it may be determined whether cross-channel prediction using correlation between channels has been determined as the encoding mode of the target block.
[0780] For example, the determination of whether cross-channel prediction has been determined as the encoding mode of the target block may be intended to determine whether the image properties of the channels are similar to each other according to whether a color component linear prediction mode (CCLM) indicating cross-channel prediction is applied to the target block. In order to determine whether CCLM is applied to the target block, it may be checked whether the prediction mode of the target block is one of the INTRA_CCLM mode, the INTRA_MMLM mode, and the INTRA_MFLM mode.
[0781] For example, the determination of whether cross-channel prediction has been determined as the coding mode of the target block may be intended to determine whether the intra-frame mode of the representative channel (e.g., the luma channel) is used for another channel (e.g., the chroma channels Cb and Cr) without change. Alternatively, the determination of whether cross-channel prediction has been determined as the coding mode of the target block may be intended to determine whether a specific intra-frame mode indicated by the intra-frame mode of the representative channel is used for another channel. Alternatively, the determination of whether cross-channel prediction has been determined as the coding mode of the target block may be intended to determine whether a specific intra-frame mode derived from the intra-frame mode of the representative channel is used for another channel.
[0782] For example, the determination of whether cross-channel prediction has been determined as the encoding mode of the target block may be intended to determine whether the encoding device 1600 and the decoding device 1700 use a specific encoding mode (e.g., inter-channel shared mode) in agreement with each other.
[0783] For example, the determination of whether cross-channel prediction has been determined as the encoding mode of the target block may be intended to determine whether the intra prediction mode of the chroma channel is DM.
[0784] Such a DM may be a mode indicating that the intra prediction mode of the chroma channel is used as the intra prediction mode of the luma channel without change due to the characteristic that the correlation between the luma channel and the chroma channel may be high. Therefore, when the intra prediction mode of the chroma channel of the target block is DM, it can be determined that the condition that the image properties of the channels are determined to be similar to each other is satisfied.
[0785] In addition to the conditions described in the above examples, whether cross-channel prediction has been determined as the encoding mode of the target block may be determined based on the size of the block.
[0786] For example, the larger the block size, the higher the probability that pixels with heterogeneous properties will exist in the corresponding block. Therefore, the similarity between channels of a block with a larger size may be smaller than the similarity between channels of a block with a smaller size. In addition, when the block size is too small, the similarity between channels of the block may be unstable.
[0787] For example, information sharing between channels may be performed only for blocks whose size is less than or equal to a specific size. The specific size may be 64×64, 32×32, or 16×16. When the block size is less than or equal to the specific size, information sharing between channels is performed, and thus the condition that the image attributes of the channels are determined to be similar to each other can be more reliably satisfied.
[0788] Alternatively, information sharing between channels may be performed only for blocks larger than a specific size. The specific size may be 4×4. When the block size is larger than the specific size, information sharing between channels is performed, and thus the condition for determining that the image attributes of the channels are similar to each other can be more reliably met.
[0789] Alternatively, information sharing between channels may be performed only when the size of the block is greater than a first specific size (e.g., 4×4) and less than or equal to a second specific size (e.g., 32×32 or 64×64). Information sharing between channels may be performed only when the size of the block falls within a specific range, and thus the condition that the image properties of the channels are determined to be similar to each other may be more reliably satisfied.
[0790] In an embodiment, a method and apparatus for encoding a target block through sharing of information between channels will be described below, and may provide the following functions.
[0791] - The coding decision information of one channel of the target block can be parsed from the compressed bit stream, and the coding decision information of one channel can be used The decoding of all channels or some selected channels of the target block is performed based on the coding decision information of each channel.
[0792] - The bitstream may be configured such that coding decision information is sent only for a representative channel or some selected channels of a target block.
[0793] - A determination may be made as to whether a transform is to be skipped for one channel, and the determination as to whether a transform is to be skipped may be applied to another channel.
[0794] - For one channel of a transform block, transform_skip_flag information may be parsed from the compressed bitstream. Whether transform will be skipped may be determined for one channel or multiple channels of the transform block by utilizing the parsed transform_skip_flag information.
[0795] - For one channel of a transform block, the transform_skip_flag information may be signaled. The transform_skip_flag information for one channel may even be used for another channel.
[0796] - Coding decision information can be efficiently signaled by sharing information between channels. This efficient signaling can improve coding efficiency and subjective image quality.
[0797] In particular, when the spatial changes in pixel values within a block are very large or very dramatic, even if a transform is applied to the target block, the image energy may not be concentrated to a large extent in low frequencies. Furthermore, when transform and quantization are applied to such blocks, low-frequency signal components are largely retained while high-frequency signal components are eliminated, or when strong quantization is applied to such blocks, significant degradation in image quality may occur. In an embodiment, whether to skip a transform for a block can be economically indicated based on a determination by the encoding device 1600 without incurring significant overhead. This economical indication can improve the image compression rate and minimize degradation in image quality.
[0798] When using a cross-channel prediction technique that leverages inter-channel correlation, multiple pieces of encoding decision information may not be used separately for multiple channels. In an embodiment, encoding decision information may be transmitted for one channel, and the transmitted encoding decision information may be shared and used with all or some selected channels from the remaining channels. This sharing can address the issue of compression rate and image quality degradation.
[0799] Determine representative channels based on color space
[0800] As color spaces used for image encoding and decoding, there are YCbCr and YUV spaces used for encoding and decoding of general images, and in addition, there are RGB, XYZ, and YCoCg spaces. When one of the various color spaces is a target color space for encoding and decoding an image, one of the channels of the target color space can be determined as a representative channel of the target color space.
[0801] In an embodiment, the color channel with the highest correlation with the luma signal among the channels may be determined as the representative channel. For example, in the RGB color space, the G channel may have the highest correlation with the luma signal, and thus the G channel may be selected as the representative channel. In the XYZ color space, the Y channel may have the highest correlation with the luma signal, and thus the Y channel may be selected as the representative channel. In the YCoCg color space, the Y channel may have the highest correlation with the luma signal, and thus the Y channel may be selected as the representative channel.
[0802] A channel in a color space may be represented by an index value such as "0 / 1 / 2". SelectedCIDX may be an index of a selected color. Alternatively, SelectedCIDX may be an index value indicating a selected representative channel. The representative channel may be determined by an index SelectedCIDX indicating the selected representative channel among a plurality of pieces of information about the target block in the bitstream.
[0803] For example, in the YCbCr color space, the value of SelectedCIDX may be 0 indicating the Y channel.
[0804] For example, in the YCbCr color space, the Cb channel may be determined as a representative channel. When the Cb channel is determined as the representative channel, the value of SelectedCIDX may be 1 indicating the Cb channel.
[0805] For example, in the YUV color space, the U channel may be determined as the representative channel. When the U channel is determined as the representative channel, the value of SelectedCIDX may be 1 indicating the U channel.
[0806] To share coding decision information between channels, a specific channel in the color space may be selected as a representative channel. During encoding and decoding of an image, the coding decision information of the representative channel may be shared among one or more remaining channels.
[0807] For example, the encoding device 1600 may signal only the coding decision information of the representative channel to the decoding device 1700 via a bitstream. Alternatively, the decoding device 1700 may use the bitstream to derive the coding decision information of the representative channel. Coding decision information for at least some of the remaining channels may not be separately signaled. The decoding device 1700 may use the coding decision information of the representative channel to derive coding decision information for at least some of the remaining channels. In other words, the coding decision information of the representative channel may be shared with at least some of the remaining channels.
[0808] For example, when the Y channel having the highest correlation with the luma signal is selected as the representative channel in the YCbCr color space, there may be a correlation between the luma channel (i.e., Y) and the chroma channels (i.e., Cb and / or Cr). Therefore, when prediction is performed for image compression, encoding decision information for the luma channel as the representative channel can be implicitly shared as multiple pieces of encoding decision information for one or more chroma blocks, rather than applying independent predictions to the three channels in the color space, respectively. The one or more chroma blocks may include one or more of a Cb block and a Cr block.
[0809] For example, when the Cb channel is selected as the representative channel in the YCbCr color space, there may be a correlation between the Cb signal and the Cr signal constituting the chrominance channel. Therefore, when prediction is performed for image compression, encoding decision information for the Cb channel as the representative channel can be implicitly shared as encoding decision information for the Cr channel, rather than applying independent predictions to the two chrominance channels respectively.
[0810] Encoding decision information shared between channels
[0811] Coding decision information that can be shared between channels may be information such as syntax elements, which are encoded by the encoding device 1600 and signaled to the decoding device 1700 as information included in the bitstream. For example, the coding decision information may include flags, indexes, etc. Furthermore, the coding decision information may include information derived during the encoding and / or decoding process. Furthermore, the coding decision information may indicate information required for encoding and / or decoding an image.
[0812] For example, the coding decision information may include the size of the unit / block, the depth of the unit / block, the partition information of the unit / block, the partition structure of the unit / block, partition flag information indicating whether the unit / block is partitioned in the form of a quadtree, partition flag information indicating whether the unit / block is partitioned in the form of a binary tree, a partition direction (horizontal or vertical) in the form of a binary tree, a partition form (symmetric partitioning or asymmetric partitioning) in the form of a binary tree, partition flag information indicating whether the unit / block is partitioned in the form of a ternary tree, a partition direction (horizontal or vertical) in the form of a ternary tree, a partition form (symmetric partitioning or asymmetric partitioning) in the form of a ternary tree, a prediction scheme (intra-frame prediction or inter-frame prediction), an intra-frame prediction mode / direction, a reference sample filtering method, a prediction block filtering method, a prediction block boundary filtering method, a filtering filter tap, a filtering filter coefficient, an inter-frame prediction mode, motion information, a motion vector, a reference picture index, an inter-frame prediction direction, an inter-frame prediction indicator, a reference picture list, a reference image, a motion vector predictor, and a motion vector prediction candidate.Motion vector candidate list, information indicating whether merge mode is used, merge candidate, merge candidate list, information indicating whether skip mode is used, type of interpolation filter, filter taps of interpolation filter, filter coefficients of interpolation filter, size of motion vector, accuracy of representation of motion vector, transform type, transform size, information indicating whether primary transform is used, information indicating whether additional (secondary) transform is used, primary transform selection information (or primary transform index), secondary transform selection information (or secondary transform index), information indicating the presence or absence of a residual signal, coding block pattern, coding block flag, quantization parameter, quantization matrix, information about in-loop filter, information about whether in-loop filter is applied, coefficients of in-loop filter, filter taps of in-loop filter, shape / form of in-loop filter, information indicating whether deblocking filter is applied, coefficients of deblocking filter, taps of deblocking filter, strength of deblocking filter, shape / form of deblocking filter, information indicating whether adaptive sample offset is applied, information indicating whether adaptive sample offset is applied, value of adaptive sample offset, category of adaptive sample offset, At least one or a combination of the type of adaptive sample offset, information indicating whether an adaptive loop filter is applied, coefficients of the adaptive loop filter, taps of the adaptive loop filter, shape / form of the adaptive loop filter, binarization / debinarization method, context model, context model determination method, context model updating method, information indicating whether a normal mode is performed, information indicating whether a bypass mode is performed, context bins, bypass bins, transform coefficients, transform coefficient levels, transform coefficient level scanning method, image display / output sequence, slice identification information, slice type, slice partition information, tile identification information, tile type information, tile partition information, picture type, bit depth, information about a luminance signal and information about a chrominance signal, transform_skip_flag information, primary transform selection information, secondary transform selection information, reference sample filtering information, PDPC_flag information, rdpcm_flag information, EMT flag information, mts_flag information, mts_idx information, nsst_flag information, and nsst_idx information.
[0813] Among the multiple pieces of encoding decision information that can be shared between channels, the primary transform selection information may be transform information required to perform a transform process on the residual signal using a combination of one or more DCT transform kernels and / or DST transform kernels related to the horizontal direction and / or the vertical direction. For example, the primary transform selection information may be information required to use the MTS in the primary transform. The primary transform selection information may include mts_flag information and mts_idx information.
[0814] The primary transform selection information applied to the target block may be explicitly signaled, or alternatively, may be implicitly derived by the encoding apparatus 1600 and the decoding apparatus 1700 using encoding decision information of the target block and encoding decision information of neighboring blocks.
[0815] After the primary transform is completed in the encoding apparatus 1600 , a secondary transform may be performed in order to improve energy concentration of transform coefficients.
[0816] The secondary transform selection information applied to the target block may be explicitly signaled, or alternatively, may be implicitly derived using the encoding decision information of the target block and the encoding decision information of the neighboring blocks by the encoding apparatus 1600 and the decoding apparatus 1700. The decoding apparatus 1700 may perform the secondary inverse transform according to whether the secondary inverse transform is to be performed, and may perform the primary inverse transform on the result of performing the secondary inverse transform according to whether the primary inverse transform is to be performed.
[0817] The encoding apparatus 1600 may generate rdpcm_flag information for a target block and may record the rdpcm_flag information in a bitstream. The decoding apparatus 1700 may acquire the rdpcm_flag information through the bitstream and may perform RDPCM according to information indicated by the rdpcm_flag information or may not perform RDPCM.
[0818] Figure 18 is a flowchart of a method for decoding encoding decision information according to an embodiment.
[0819] According to an embodiment, when a specific channel in a color space is selected as a representative channel to share information among multiple channels, encoding decision information of the representative channel of the target block may be shared by one or more remaining channels of the multiple channels except the representative channel.
[0820] For example, when performing intra-frame prediction for a target block in the YCbCr color space, the Y channel may be set as a representative channel, and thereafter, intra-frame coding decision information of the representative channel may be shared and used to perform decoding of channels other than the representative channel (i.e., the Cb channel and / or the Cr channel), rather than separately transmitting multiple pieces of intra-frame coding decision information for the three channels in the color space. Alternatively, after the Cb channel is assumed to be the representative channel in the YCbCr color space, the intra-frame coding decision information of the representative channel may be implicitly used for decoding the Cr channel, where the Cr channel is a channel other than the representative channel.
[0821] For example, the intra-frame coding decision information of the representative channel that can be shared by the remaining channels may include one or more of the intra-frame prediction mode, the intra-frame prediction direction, the prediction block boundary filtering method, the filter taps for the prediction block boundary filtering, the filter coefficients for the prediction block boundary filtering, the transform_skip_flag information, the primary transform selection information, the secondary transform selection information, the mts_flag information, the mts_idx information, the PDPC_flag information, the rdpcm_flag information, the EMT flag information, the nsst_flag information, the nsst_idx information, the intra-frame smoothing filtering information, the CU skip flag information, the CU_lic flag information, the obmc_flag information, the codeAlfCtuEnable_flag information and the PDPC_flag information.
[0822] The case where cross-channel prediction is selected from various techniques for predicting chrominance signals (including angular prediction, DC prediction, planar prediction, etc.) or the case where cross-channel prediction is more advantageous may mean that the properties of the luminance signal (i.e., the Y signal) are very similar to the properties of the chrominance signals (i.e., the Cb signal and / or the Cr signal).
[0823] In this case, when the channel of the Y signal block is a representative channel, the coding decision information determined in the process of encoding and / or decoding the representative channel can be similarly applied to the chrominance block. By applying this, the number of bits required to transmit the coding decision information can be reduced. Therefore, encoding and decoding can be performed so that a single piece of coding decision information is used for multiple channels through cross-channel prediction.
[0824] For example, after the coding decision information for the luma channel has been determined, the determined coding decision information can be shared among the remaining channels, and encoding can be performed based on the shared coding decision information. Alternatively, multiple pieces of coding decision information can be independently applied to three channels instead of being shared between the channels, and the three channels can be independently encoded and / or decoded. The encoding device 1600 can determine a method that is advantageous from the perspective of rate distortion as the encoding method among the methods for sharing coding decision information between channels and the methods for independently encoding channels. Based on the determination, the encoding device 1600 can explicitly write information about whether the coding decision information is shared between channels into the bitstream, and this information can be sent to the decoding device 1700 through the bitstream.
[0825] For example, when a specific encoding condition is met in an encoding and / or decoding process, information on whether encoding decision information is shared between channels may not be explicitly signaled, and encoding decision information of a representative channel may be shared by the remaining channels.
[0826] For example, the specific encoding condition may be a condition indicating whether cross-component prediction (CCP), DM, or cross-component linear model (CCLM) is used.
[0827] For example, when at least one of the original signal, reconstructed signal, residual signal and prediction signal of the luminance signal that has been encoded or decoded is used to predict the remaining chrominance signals (i.e., Cb signal and / or Cr signal), sharing of encoding decision information between channels can be applied.
[0828] For example, when predicting a luminance signal using at least one of an original signal, a reconstructed signal, a residual signal, and a prediction signal of a chrominance signal that has been completely encoded or decoded, sharing of encoding decision information between channels may be applied.
[0829] For example, when predicting a Cr signal using at least one of an original signal, a reconstructed signal, a residual signal, and a predicted signal of a Cb signal that has been completely encoded or decoded, sharing of encoding decision information between channels may be applied.
[0830] For example, when predicting a Cb signal using at least one of an original signal, a reconstructed signal, a residual signal, and a predicted signal of a Cr signal that has been completely encoded or decoded, sharing of encoding decision information between channels may be applied.
[0831] When the luma channel is set as the representative channel in the YCbCr color space, encoding decision information may be signaled only for the luma signal, and encoding decision information may not be separately transmitted for the remaining chroma channels, rather than signaling encoding decision information for the remaining channels except the luma channel. By selectively transmitting such information, the compressibility can be improved.
[0832] Furthermore, encoding decision information may be transmitted only for the Cb signal, and the transmitted encoding decision information may be shared with the Cr signal, thereby improving the compression ratio. Alternatively, encoding decision information may be transmitted only for the Cr signal, and the transmitted encoding decision information may be shared with the Cb signal, thereby improving the compression ratio.
[0833] In step 1810, the communication unit 1720 may receive a bitstream. The bitstream may include encoding decision information.
[0834] At step 1820 , the processing unit 1710 may determine whether sharing of encoding decision information is to be used for a target channel of the target block.
[0835] When the encoding decision information is not shared, step 1830 may be performed.
[0836] When the encoding decision information is to be shared, step 1840 may be performed.
[0837] In step 1830, the processing unit 1710 may obtain the encoding decision information of the target channel from the bitstream. The processing unit 1710 may parse and read the encoding decision information of the target channel from the bitstream.
[0838] In step 1840 , the processing unit 1710 may set the encoding decision information so that the encoding decision information of the representative channel is used as the encoding decision information of the target channel.
[0839] Step 1820, step 1830, and step 1840 may be represented by the following code 1:
[0840] [Code 1]
[0841]
[0842] cIdx may indicate a target channel of the target block. For example, when the number of channels in the target image is 3, cIdx may be one of predefined specific values and may be one of {0, 1, 2}.
[0843] In an embodiment, the cIdx of the representative channel may be assumed to be 0.
[0844] “cidx != 0” may indicate that the target channel is not a representative channel (eg, a luma channel). In other words, a cIdx value of “0” may indicate that the target channel is a representative channel.
[0845] In other words, in step 1820, when the target channel is not the representative channel and cross-channel prediction is used for the target block, the processing unit 1710 may determine to use sharing of encoding decision information for the target channel. When the target channel is the representative channel or when cross-channel prediction is not used for the target block, the processing unit 1710 may determine not to use sharing of encoding decision information for the target channel.
[0846] Whether cross-channel prediction is used for a target block may be 1) inferred based on information obtained from a bitstream and 2) implicitly inferred according to whether a specific condition is satisfied.
[0847] As described above, whether cross-channel prediction is used may be determined based on the intra prediction mode of the target block. Whether cross-channel prediction is used may be determined based on whether the intra prediction mode of the target block is one of the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode. For example, when the intra prediction mode of the target block is one of the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode, the processing unit 1710 may determine that cross-channel prediction is used.
[0848] As described above, whether cross-channel prediction is used may be determined based on whether the intra-frame prediction mode of the chroma channel of the target block has a specific value. For example, when the intra-frame prediction mode of the chroma channel of the target block has a specific value, the processing unit 1710 may determine that cross-channel prediction is used.
[0849] The intra prediction mode of the chroma channel of the target block may be indicated by intra_chroma_pred_mode information.
[0850] As described above, whether cross-channel prediction is used may be determined based on whether the intra prediction mode of the target channel is DM. For example, when the intra prediction mode of the target channel is DM, the processing unit 1710 may determine that cross-channel prediction is used.
[0851] In step 1830 , the processing unit 1710 may obtain coding decision information of a block of a channel indicated by cIdx from the bitstream.
[0852] The processing unit 1710 may parse and read the coding decision information of the block of the channel indicated by cIdx from the bitstream.
[0853] In step 1840, the processing unit 1710 may share the encoding decision information of the representative channel as the encoding decision information of the block of the channel indicated by cIdx. In other words, the processing unit 1710 may set the encoding decision information of the representative channel as the encoding decision information of the block of the channel indicated by cIdx.
[0854] Depending on the embodiment, an operation corresponding to the condition or execution may be additionally performed before step 1820 or between steps 1820 and 1840 .
[0855] According to an embodiment, for the Cb signal, the coding decision information can be sent, and for the Cr signal, the coding decision information of the Cb signal can be shared without sending the coding decision information. In this case, the above code 1 can be modified to the following code 2:
[0856] [Code 2]
[0857]
[0858] Figure 19 is a flowchart of a decoding method for determining whether a transform is to be skipped, according to an embodiment.
[0859] The encoding apparatus 1600 may determine whether transformation (eg, primary transformation and / or secondary transformation) will be skipped according to the size of the target block.
[0860] The target block may be a transform block.
[0861] log2TrafoSize may represent the size of the target block.
[0862] For example, when the size of the target block is less than or equal to a threshold value indicating a boundary value of the block size, the encoding apparatus 1600 may skip transformation of the target block.
[0863] Log2MaxTransformSkipSize may represent a threshold value indicating a boundary value of a block size.
[0864] When the transformation of the target block is skipped, the encoding apparatus 1600 may set a value of the transform_skip_flag information to 1 without performing the transformation. The transform_skip_flag information may be transmitted to the decoding apparatus 1700 through a bitstream.
[0865] Also, when performing transformation of the target block, the encoding apparatus 1600 may perform the transformation and may set a value of the transform_skip_flag information to 0. The transform_skip_flag information may be transmitted to the decoding apparatus 1700 through a bitstream.
[0866] Here, a plurality of pieces of transform_skip_flag information may be transmitted separately for channels constituting the color space of the image.
[0867] The decoding apparatus 1700 may acquire the value of the transform_skip_flag information from the bitstream. In other words, the decoding apparatus 1700 may parse and read the transform_skip_flag information from the bitstream.
[0868] Here, the decoding apparatus 1700 may acquire the value of the transform_skip_flag information from the bitstream only when the size of the block is less than or equal to a threshold value indicating a boundary value of the block size.
[0869] Also, the decoding apparatus 1700 may acquire a plurality of pieces of transform_skip_flag information for a plurality of channels of an image.
[0870] The acquisition of transform_skip_flag information can be represented by the following code 3:
[0871] [Code 3]
[0872] If(log2TrafoSize<=Log2MaxTransformSkipSize)
[0873] transform_skip_flag[x0][y0][cIdx]
[0874] x0 and y0 may denote spatial coordinates indicating the position of the target block.
[0875] cIdx may indicate a target channel of target block information.
[0876] When there are three image channels, cIdx may have one of predefined values {0, 1, 2}. The value of the representative channel may be 0.
[0877] Code 3 can be modified to the following code 4:
[0878] [Code 4]
[0879] If((log2TbWidth<=Log2MaxTransformSkipSize_W)&&(log2TbHeight<=Log2MaxTransformSkipSize_H))
[0880] transform_skip_flag[x0][y0][cIdx]
[0881] log2TbWidth may be a value based on the following Equation 11. 'width' may be the width of the target block (ie, the horizontal length of the target block).
[0882] [Equation 11]
[0883] log2TbWidth=log2width
[0884] log2TbHeight may have a value based on the following Equation 12. 'Height' may be the height of the target block (ie, the vertical length of the target block).
[0885] [Equation 12]
[0886] log2TbHeight=log2height
[0887] The predefined thresholds Log2MaxTransformSkipSize_W and Log2MaxTransformSkipSize_H may be equal to each other or may be different from each other. For example, the value of Log2MaxTransformSkipSize_W may be 2, and the value of Log2MaxTransformSkipSize_H may be 2.
[0888] Code 3 can be modified to the following code 5:
[0889] [Code 5]
[0890] If((log2TbWidth<=2)&&(log2TbHeight<=2))
[0891] transform_skip_flag[x0][y0][cIdx]
[0892] As described above, instead of separately signaling multiple pieces of transform_skip_flag information for multiple channels, transform_skip_flag information may be signaled only for the luma (Y) signal, and may not be separately signaled for the remaining chroma channels. Alternatively, transform_skip_flag information may be signaled only for the Cb signal, and may not be separately signaled for the Cr signal, and the transform_skip_flag information transmitted for the Cb signal may be shared for the Cr signal.
[0893] Next, an embodiment of sharing transform_skip_flag information will be described.
[0894] In an embodiment, the decoding device 1700 may obtain transform_skip_flag information from the bitstream. This acquisition may be represented by the following code 6:
[0895] [Code 6]
[0896]
[0897]
[0898] x0 and y0 may be spatial coordinates indicating the location of the target block.
[0899] cIdx may indicate a target channel of a target block.
[0900] In Code 6 and other codes including the condition “if (log2TrafoSize <= Log2MaxTransformSkipSize)”, the condition “if ((log2TbWidth <= Log2MaxTransformSkipSize_W) && (log2TbHeight <= Log2MaxTransformSkipSize_H))” or the condition “if ((log2TbWidth <= 2) && (log2TbHeight <= 2))” may be used instead of the condition “if (log2TrafoSize <= Log2MaxTransformSkipSize)”.
[0901] When the number of channels of the image is 3, cIdx may have one of predefined values {0, 1, 2}. The value of the representative channel may be 0. Alternatively, the value of the representative channel may be 1 or 2.
[0902] In step 1910 , the communication unit 1720 may receive a bitstream.
[0903] At step 1920 , the processing unit 1710 may determine whether it is possible to skip transformation for the target block.
[0904] If it is determined that the transformation may be skipped, step 1930 may be performed.
[0905] If it is determined that skipping the transition is not possible, step 1960 may be performed.
[0906] For example, when the size of the target block is less than or equal to a specific size, the processing unit 1710 may determine that skipping transformation is impossible.
[0907] For example, when the size of the target block is greater than a specific size, the processing unit 1710 may determine that skipping transformation is impossible.
[0908] Here, the specific size may be a boundary value of a block size that allows skipping of transformation.
[0909] For example, when the condition in the following Code 7 is satisfied (i.e., when the result of the condition in Code 7 is true), the processing unit 1710 may determine that it is impossible to skip the transformation, and when the condition in the following Code 7 is not satisfied (i.e., when the result of the condition in Code 7 is false), the processing unit 1710 may determine that it is possible to skip the transformation.
[0910] [Code 7]
[0911] if(log2TrafoSize<=Log2MaxTransformSkipSize)
[0912] In step 1930 , the processing unit 1710 may determine whether sharing of transform_skip_flag information with a target channel of a target block is to be used.
[0913] If it is determined that the transform_skip_flag information will not be shared, step 1940 may be performed.
[0914] If it is determined that the transform_skip_flag information is to be shared, step 1950 may be performed.
[0915] For example, when the condition in the following Code 8 is met (i.e., when the result of the condition in Code 8 is true), the processing unit 1710 may determine that the transform_skip_flag information will be shared, and when the condition in the following Code 8 is not met (i.e., when the result of the condition in Code 8 is false), the processing unit 1710 may determine that the transform_skip_flag information will not be shared.
[0916] [Code 8]
[0917] if ((cIdx != 0) && (cross-channel prediction is used))
[0918] In other words, in step 1930, when the target channel is not a representative channel and cross-channel prediction is used for the target block, the processing unit 1710 may determine that the transform_skip_flag information will be shared with the target channel. On the contrary, when the target channel is a representative channel or when cross-channel prediction is not used for the target block, the processing unit 1710 may determine that the transform_skip_flag information will not be shared.
[0919] In step 1940, the processing unit 1710 may obtain transform_skip_flag information of the target channel from the bitstream. The processing unit 1710 may parse and read the transform_skip_flag information of the target channel from the bitstream.
[0920] The transform_skip_flag information may be stored in transform_skip_flag[x0][y0][cIdx].
[0921] In step 1950 , the processing unit 1710 may set transform_skip_flag information so that the transform_skip_flag information of the representative channel is used as the transform_skip_flag information of the target channel.
[0922] The processing unit 1710 may use the transform_skip_flag information of the representative channel as the transform_skip_flag information of the target channel without parsing and reading the transform_skip_flag information of the target channel from the bitstream. In other words, the processing unit 1710 may store the value of transform_skip_flag[x0][y0][0] in transform_skip_flag[x0][y0][cIdx].
[0923] That is, without requiring a process for parsing and reading transform_skip_flag information of a target channel from a bitstream, a value previously stored in transform_skip_flag[x0][y0][0] may be used in transform_skip_flag[x0][y0][cIdx] as well.
[0924] For example, transform_skip_flag information signaled for the luma (Y) channel as a representative channel may be used even for chroma channels (Cb and / or Cr).
[0925] According to an embodiment, for the Cb signal, transform_skip_flag information may be transmitted, and for the Cr signal, the transform_skip_flag information for the Cb signal may be shared without transmitting the transform_skip_flag information. In this case, the above code 6 may be modified to the following code 9:
[0926] [Code 9]
[0927]
[0928] When it is not possible to skip transformation for the target block, step 1960 may be performed.
[0929] Information indicating that transformation will not be skipped for the target block may be set in step 1960. Since skipping transformation for the target block is not allowed, the value of transform_skip_flag[x0][y0][cIdx] may be set to 0.
[0930] Figure 20 is a flowchart of a decoding method for determining whether a transform is to be skipped according to an intra mode according to an embodiment.
[0931] There can be considerable correlation between the luma (i.e., Y) channel and the chroma (i.e., Cb and / or Cr) channels of an image. For example, the luma channel may include a lot of information about the texture of the image, and the Cb and Cr channels, which are chroma channels, may additionally provide color information that will be added to the texture.
[0932] Therefore, when performing prediction required for compression and reconstruction of an image, prediction values for the Cb block and the Cr block for which prediction is performed based on the signal of the luminance channel previously obtained through decoding can be calculated instead of performing independent predictions for each of the three channels of the color space.
[0933] The technique used to calculate these predictions may be referred to as "Cross-Channel Prediction (CCP)" or "CCLM" as described above.
[0934] The decoding apparatus 1700 may determine whether cross-channel prediction has been used by checking whether the intra prediction mode of the target block is one of the INTRA_CCLM mode, the INTRA_MMLM mode, and the INTRA_MFLM mode.
[0935] Since a significant portion of the texture information of the chrominance signal is also included in the luminance signal, such cross-channel prediction may be effective. Similarly, cross-channel prediction can be used to calculate the prediction value for the Cr block as the prediction target based on the signal of the Cb channel.
[0936] A case where cross-channel prediction is selected for prediction of a chroma signal from among various techniques including angular prediction, DC prediction, and planar prediction, or a case where cross-channel prediction is advantageous may indicate that signal characteristics of a channel corresponding to SelectedCIDX are very similar to those of another channel.
[0937] Therefore, when it is advantageous to skip (or perform) transforms for blocks of a channel corresponding to SelectedCIDX, it may be equally advantageous to also skip (or perform) transforms for blocks of the remaining channels.
[0938] Therefore, when cross-channel prediction is used, the bitstream may not be parsed separately for the three channels to obtain transform_skip_flag information. When the transform_skip_flag information of the representative channel is parsed, the transform_skip_flag information of the remaining channels may not be parsed separately. The transform_skip_flag information of the representative channel may be shared and used as the transform_skip_flag information of the remaining channels, and information indicating such sharing may be recorded in the bitstream. For example, for such sharing, the channel corresponding to SelectedCIDX may be used to determine whether to skip transforms.
[0939] Alternatively, the rate-distortion value may be calculated for the case where the transform is skipped equally for the three channels, and the rate-distortion value may be calculated for the case where the transform is performed equally for the three channels. The rate-distortion value calculated when the transform is skipped and the rate-distortion value calculated when the transform is performed may be compared with each other, and based on the result of the comparison, a more favorable scheme between the scheme for skipping the transform and the scheme for performing the transform may be used to encode the channels.
[0940] In an embodiment, instead of signaling a plurality of pieces of transform_skip_flag information for a plurality of channels, transform_skip_flag information may be signaled only for a channel corresponding to SelectedCIDX, and transform_skip_flag information may not be separately signaled for the remaining channels.
[0941] Next, an embodiment in which such transform_skip_flag information is shared will be described.
[0942] In an embodiment, the decoding device 1700 may obtain transform_skip_flag information from the bitstream. Such acquisition may be represented by the following code 10:
[0943] [Code 10]
[0944]
[0945] x0 and y0 may be spatial coordinates indicating the location of the target block.
[0946] cIdx may indicate a target channel of a target block.
[0947] When the number of channels in the image is 3, the value of cIdx in code 10 may be one of values {0, 1, 2}. For example, the value of cIdx may be one of predefined values {0, 1, 2}.
[0948] In step 2010 , the communication unit 1720 may receive a bitstream.
[0949] At step 2020 , the processing unit 1710 may determine whether it is possible to skip transformation for the target block.
[0950] When it is possible to skip the transformation, step 2030 may be performed.
[0951] When skipping the transition is not possible, step 2060 may be performed.
[0952] For example, when the size of the target block is less than or equal to a specific size, the processing unit 1710 may determine that skipping transformation is impossible.
[0953] For example, when the size of the target block is greater than a specific size, the processing unit 1710 may determine that skipping transformation is impossible.
[0954] Here, the specific size may be a boundary value of a block size that allows skipping of transformation.
[0955] For example, when the condition in the following code 11 is satisfied (i.e., when the result of the condition in code 11 is true), the processing unit 1710 may determine that it is impossible to skip the transformation, and when the condition in the following code 11 is not satisfied (i.e., when the result of the condition in code 11 is false), the processing unit 1710 may determine that it is possible to skip the transformation.
[0956] [Code 11]
[0957] if(log2TrafoSize<=Log2MaxTransformSkipSize)
[0958] In step 2030 , the processing unit 1710 may determine whether to use sharing of transform_skip_flag information with a target channel of a target block based on the selected representative channel.
[0959] If it is determined that the transform_skip_flag information will not be shared, step 2040 may be performed.
[0960] If it is determined that the transform_skip_flag information is to be shared, step 2050 may be executed.
[0961] For example, when the condition in the following code 12 is met (i.e., when the result of the condition in code 12 is true), the processing unit 1710 may determine that the transform_skip_flag information will be shared, and when the condition in the following code 12 is not met (i.e., when the result of the condition in code 12 is false), the processing unit 1710 may determine that the transform_skip_flag information will not be shared.
[0962] [Code 12]
[0963] if((cIdx!=SelectedCIDX)&&"cross-channel prediction is used")
[0964] In other words, in step 2030, when the target channel is not the selected representative channel indicated by SelectedCIDX and cross-channel prediction is used for the target block, the processing unit 1710 may determine to share the transform_skip_flag information with the target channel. In addition, when the target channel is the selected representative channel indicated by SelectedCIDX or when cross-channel prediction is not used for the target block, the processing unit 1710 may determine not to share the transform_skip_flag information.
[0965] In step 2040, the processing unit 1710 may obtain the transform_skip_flag information of the target channel from the bitstream. The processing unit 1710 may parse and read the transform_skip_flag information of the target channel from the bitstream.
[0966] The transform_skip_flag information may be stored in transform_skip_flag[x0][y0][cIdx].
[0967] In step 2050 , the processing unit 1710 may set the transform_skip_flag information so that the transform_skip_flag information of the selected representative channel indicated by SelectedCIDX is used as the transform_skip_flag information of the target channel.
[0968] The processing unit 1710 may use the transform_skip_flag information of the selected representative channel indicated by SelectedCIDX as the transform_skip_flag information of the target channel without parsing and reading the transform_skip_flag information of the target channel from the bitstream. In other words, the processing unit 1710 may store the value of transform_skip_flag[x0][y0][SelectedCIDX] in transform_skip_flag[x0][y0][cIdx].
[0969] That is, the value previously stored in transform_skip_flag[x0][y0][SelectedCIDX] can also be used in transform_skip_flag[x0][y0][cIdx] in the same manner without requiring a process for parsing and reading transform_skip_flag information of a target channel from a bitstream.
[0970] Information indicating that the target channel for the target block will not skip transform may be set in step 2060. Since skipping transform for the target block is not allowed, the value of transform_skip_flag[x0][y0][cIdx] may be set to 0.
[0971] In other words, for a target channel of a target block indicated by cIdx, a predefined value of 0 may be set in the transform_skip_flag information so as to indicate that transform will not be skipped for the target channel, without parsing and reading the transform_skip_flag information indicating whether transform will be skipped from the bitstream.
[0972] As described above, in the above embodiment, the encoding decision information of the selected representative channel may be shared with all channels except the selected representative channel.
[0973] The above embodiment may be partially modified. In other words, the encoding decision information of the selected representative channel may be shared as the encoding decision information of another designated channel. For example, the encoding decision information may include transform_skip_flag information.
[0974] For example, when the value of SelectedCIDX is 1, a cIDX value of 1 indicates a Cb signal, and a cIDX value of 2 indicates a Cr signal, encoding decision information for the Cb signal may be shared as encoding decision information for the Cr signal.
[0975] In other words, encoding decision information for the Cb signal may be transmitted from the encoding device 1600 to the decoding device 1700 , and encoding decision information for the Cr signal may be set using the encoding decision information for the Cb signal without separately transmitting the encoding decision information for the Cr signal.
[0976] When the coding decision information for the Cb signal is shared as the coding decision information for the Cr signal, the above code 10 may be modified to the following code 13:
[0977] [Code 13]
[0978]
[0979] In step 2030, when the condition in the following code 14 is met (i.e., when the result of the condition in code 14 is true), the processing unit 1710 may determine that the transform_skip_flag information will be shared, and when the condition in the following code 14 is not met (i.e., when the result of the condition in code 14 is false), the processing unit 1710 may determine that the transform_skip_flag information will not be shared.
[0980] [Code 14]
[0981] if((cIdx!=2)or(!"cross-channel prediction is used"))
[0982] In other words, in step 2030, 1) when the target channel is not a channel that shares the encoding decision information of the selected representative channel indicated by SelectedCIDX, or 2) when cross-channel prediction is not used for the target block, the processing unit 1710 may determine not to share the transform_skip_flag information with the target channel. In addition, 1) when the target channel is a channel that shares the encoding decision information of the selected representative channel indicated by SelectedCIDX, and 2) when cross-channel prediction is used for the target block, the processing unit 1710 may determine to share the transform_skip_flag information.
[0983] The steps in code 13 can be implemented as other steps with the same meaning. For example, code 13 can be modified to the following code 15:
[0984] [Code 15]
[0985]
[0986] Sharing of transformation selection information
[0987] Figure 21 is a flowchart of a method for sharing transform selection information according to an embodiment.
[0988] In the above embodiment, transform_skip_flag information has been described as encoding decision information to be shared. The transform_skip_flag information in the above embodiment may be replaced with another type of encoding decision information. In the following, transform selection information will be described as encoding decision information to be shared.
[0989] The transform selection information may be information indicating which transform is to be used for the transform block of the target channel. The transform selection information may include the primary transform selection information and / or the secondary transform selection information described above.
[0990] There can be considerable correlation between the luma channel (i.e., Y) and the chroma channels (i.e., Cb and / or Cr) of an image. For example, the luma channel may include a lot of information about the texture of the image, and the Cb and Cr channels, which are chroma channels, may additionally provide color information that will be added to the texture.
[0991] Therefore, when performing prediction required for image compression and reconstruction, prediction values for Cb blocks and Cr blocks can be calculated instead of performing independent prediction for each of the three channels of the color space, where prediction is performed for the Cb blocks and Cr blocks based on the signal of the luma channel previously obtained through decoding. Such cross-channel prediction is effective because a considerable amount of texture information of the chroma signal can be included in the luma signal.
[0992] The case where cross-channel prediction is selected from various techniques for predicting chrominance signals (including angular prediction, DC prediction, planar prediction, etc.) or the case where cross-channel prediction is more advantageous may mean that the signal properties of the luminance channel are very similar to the signal properties of the chrominance channels (i.e., Cb and / or Cr).
[0993] Therefore, when it is advantageous to use a particular transform from among multiple transforms for the luma block, it may be advantageous to use the same transform for the chroma blocks (ie, Cb blocks and / or Cr blocks).
[0994] In other words, generally, the same transform may be used for luma blocks and chroma blocks. Alternatively, when a specific transform is used for the luma block, another specific transform corresponding to the specific transform used for the luma block may be used for the chroma blocks.
[0995] Therefore, when cross-channel prediction is used, if one transform is determined, the determined transform may be used equally for the luma channel and the chroma channels (ie, three channels) without separately signaling the transforms for the luma channel and the chroma channels.
[0996] Alternatively, if a transform is determined for the luma channel, a transform corresponding to the transform determined for the luma channel may be used for the chroma channels.
[0997] When the transform to be used for the channel is determined in this manner, the luma channel and the chroma channels can be encoded based on the determination. For such encoding, one of the multiple available transforms can be selected only for the luma channel, and the transform for the chroma channels can be automatically determined by selecting the transform for the luma channel.
[0998] Alternatively, encoding using the same transform can be performed for all three channels. With this encoding, rate-distortion values can be calculated for multiple transforms. Thereafter, by comparing the rate-distortion values of the applicable transforms, the transform with the most favorable rate-distortion value can be selected, and encoding can be performed based on the selected transform.
[0999] Alternatively, encoding using a transform set may be performed on the three channels. The transform set may include a specific transform for the luma channel and transforms for the chroma channels that correspond to the specific transform.
[1000] By means of such encoding, rate-distortion values can be calculated for a plurality of transform sets. Thereafter, by comparing the rate-distortion values of the applicable transform sets, the transform set with the most favorable rate-distortion value can be selected, and encoding can be performed according to the selected transform set.
[1001] In an embodiment, instead of separately signaling multiple pieces of transform selection information for multiple channels, transform selection information may be signaled only for the luma signal (luminance channel), and transform selection information may not be separately signaled for the remaining channels which are chroma channels.
[1002] In an embodiment, the decoding apparatus 1700 may acquire transform selection information from a bitstream.
[1003] In step 2110 , the communication unit 1720 may receive a bitstream.
[1004] In step 2120 , the processing unit 1710 may determine whether sharing of transform selection information with a target channel of a target block is to be used.
[1005] When the transform selection information is not to be shared, step 2130 may be performed.
[1006] When the transformation selection information is to be shared, step 2140 may be performed.
[1007] In step 2130, the processing unit 1710 may obtain the transform selection information of the target channel from the bitstream. The processing unit 1710 may parse and read the transform selection information of the target channel from the bitstream.
[1008] In step 2140 , the processing unit 1710 may set the transform selection information so that the transform selection information of the representative channel is used as the transform selection information of the target channel.
[1009] Step 2120, step 2130, and step 2140 may be represented by the following code 16:
[1010] [Code 16]
[1011]
[1012] x0 and y0 may be spatial coordinates indicating the location of the target block.
[1013] cIdx may indicate a target channel of a target block.
[1014] When the number of channels in the image is 3, the value of cIdx in code 16 may be one of values {0, 1, 2}. For example, the value of cIdx may be one of predefined values {0, 1, 2}.
[1015] The cIdx of the representative channel may be assumed to be 0.
[1016] “cIdx != 0” may indicate that the target channel is not a representative channel (eg, luma channel). A cIdx value of 0 may indicate that the target channel is a representative channel.
[1017] In other words, in step 2120, when the target channel is not a representative channel and cross-channel prediction is used for the target block, the processing unit 1710 may determine that sharing of transform selection information with the target channel is to be used. When the target channel is a representative channel or when cross-channel prediction is not used for the target block, the processing unit 1710 may determine that sharing of transform selection information with the target channel is not to be used.
[1018] In step 2130, the processing unit 1710 may obtain the transform selection information of the target channel from the bitstream. The processing unit 1710 may parse and read the transform selection information of the target channel from the bitstream.
[1019] The transform selection information may be stored in transform selection information[x0][y0][cIdx].
[1020] In step 2140 , the processing unit 1710 may set the transform selection information so that the transform selection information of the representative channel is used as the transform selection information of the target channel.
[1021] The processing unit 1710 may use the transform selection information of the representative channel as the transform selection information of the target channel without parsing and reading the transform selection information of the target channel from the bitstream. That is, the processing unit 1710 may store the value of transformselectioninformation[x0][y0][0] in transformselectioninformation[x0][y0][cIdx].
[1022] That is, without requiring a process for parsing and reading the transform selection information of the target channel from the bitstream, the value previously stored in transform selection information[x0][y0][0] can be similarly used for transform selection information[x0][y0][cIdx].
[1023] Determines whether cross-channel prediction will be used
[1024] Refer to above Figures 18 to 21In the described embodiments, it has been illustrated that whether to use cross-channel prediction is determined by checking whether the intra prediction mode of the target block is one of the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode. In other words, when the intra prediction mode of the target block is one of the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode, cross-channel prediction can be used. When the intra prediction mode of the target block is not one of the INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode, cross-channel prediction may not be used.
[1025] The above determination of whether cross-channel prediction is to be used is merely an example, and whether cross-channel prediction is to be used may be determined by one of the following codes 17 to 23. For example, when the value of the condition in each of the following codes is true, cross-channel prediction may be used, and when the value of the condition in each of the following codes is false, cross-channel prediction may not be used. "intra_chroma_pred_mode" may be an intra prediction mode for the chroma channel.
[1026] [Code 17]
[1027] if(intra_chroma_pred_mode==CCLM mode)
[1028] [Code 18]
[1029] if(intra_chroma_pred_mode==DM mode)
[1030] [Code 19]
[1031] if(intra_chroma_pred_mode==INTRA_CCLM mode)
[1032] [Code 20]
[1033] if(intra_chroma_pred_mode==INTRA_MMLM mode)
[1034] [Code 21]
[1035] if(intra_chroma_pred_mode==INTRA_MFLM mode)
[1036] [Code 22]
[1037] if((intra_chroma_pred_mode==INTRA_CCLM mode)||
[1038] (intra_chroma_pred_mode==INTRA_MMLM mode)||
[1039] (intra_chroma_pred_mode==INTRA_MFLM mode))
[1040] [Code 23]
[1041] if((intra_chroma_pred_mode==DM mode)||
[1042] (intra_chroma_pred_mode==INTRA_CCLM mode)||
[1043] (intra_chroma_pred_mode==INTRA_MMLM mode)||
[1044] (intra_chroma_pred_mode==INTRA_MFLM mode))
[1045] Each of the CCLM mode, DM mode, INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode described in Code 17 to Code 23 may indicate one value of intra_chroma_pred_mode presented in the first column of the above Tables 10 and 11. With regard to the CCLM mode, DM mode, INTRA_CCLM mode, INTRA_MMLM mode, and INTRA_MFLM mode, reference may be made to the aforementioned description regarding Tables 10 and 11.
[1046] When aiming to determine whether cross-channel prediction is to be used, the size of the block may be additionally considered in the schemes in the above-mentioned Code 17 to Code 23.
[1047] Determining whether cross-channel prediction is to be used may also be performed by one of the following codes 24 to 30. For example, when the value of the condition in each of the following codes is true, cross-channel prediction may be used, and when the value of the condition in each of the following codes is false, cross-channel prediction may not be used.
[1048] [Code 24]
[1049] if ((intra_chroma_pred_mode == CCLM mode) && block size condition)
[1050] [Code 25]
[1051] if ((intra_chroma_pred_mode == DM mode) && block_size condition)
[1052] [Code 26]
[1053] if ((intra_chroma_pred_mode == INTRA_CCLM mode) && block_size condition)
[1054] [Code 27]
[1055] if ((intra_chroma_pred_mode == INTRA_MMLM mode) && block_size condition)
[1056] [Code 28]
[1057] if ((intra_chroma_pred_mode == INTRA_MFLM mode) && block_size condition)
[1058] [Code 29]
[1059] if((intra_chroma_pred_mode==INTRA_CCLM mode)||
[1060] (intra_chroma_pred_mode==INTRA_MMLM mode)||
[1061] (intra_chroma_pred_mode == INTRA_MFLM mode) &&block size condition)
[1062] [Code 30]
[1063] if((intra_chroma_pred_mode==DM(direct mode)mode)||
[1064] (intra_chroma_pred_mode==INTRA_CCLM mode)||
[1065] (intra_chroma_pred_mode==INTRA_MMLM mode)||
[1066] (intra_chroma_pred_mode == INTRA_MFLM mode) &&block size condition)
[1067] The block size conditions presented in Code 24 to Code 30 may be replaced by one of the following Codes Code 31 , Code 32 , Code 33 , and Code 34 .
[1068] [Code 31]
[1069] ((log2TbWidth<=Log2MaxSizeWidth)&&(log2TbHeight<=Log2MaxSizeHeight))
[1070] [Code 32]
[1071] ((log2TbWidth>=Log2MinSizeWidth)&&(log2TbHeight>=Log2MinSizeHeight))
[1072] [Code 33]
[1073] ((log2TbWidth>Log2MinSizeWidth)&&(log2TbHeight<=Log2MaxSizeHeight))
[1074] [Code 34]
[1075] ((log2TbWidth>Log2MinSizeWidth)&&(log2TbHeight>Log2MaxSizeHeight))
[1076] log2TbWidth and log2TbHeight have been described above with reference to Equation 11 and Equation 12.
[1077] Log2MaxSizeWidth, Log2MaxSizeHeight, Log2MinSizeWidth, and Log2MinSizeHeight may be predefined values. Log2MaxSizeWidth may be the width of the block with the largest size. Log2MaxSizeHeight may be the height of the block with the largest size. Log2MinSizeWidth may be the width of the block with the smallest size. Log2MinSizeHeight may be the height of the block with the smallest size.
[1078] For example, the value of Log2MaxSizeWidth may be 16, the value of Log2MaxSizeHeight may be 16, the value of Log2MinSizeWidth may be 4, and the value of Log2MinSizeHeight may be 4.
[1079] Alternatively, the value of Log2MaxSizeWidth may be 32, and the value of Log2MaxSizeHeight may be 32.
[1080] When DM is used, the intra prediction mode of the chrominance signal may not be signaled separately.When DM is used, the intra prediction mode signaled for the luma signal may also be used in the chrominance mode without change.
[1081] Encoding and decoding using shared selective information between channels under a block partitioning structure
[1082] Typically, when an image is encoded, appropriate encoding schemes may be applied separately to a plurality of spatial regions taking into account spatial characteristics in the image. For such encoding, the image may be partitioned into CUs, and the CUs generated from the partitions may be encoded separately.
[1083] To perform this encoding, the same block partition structure may be used for both the luma and chroma channels.
[1084] However, the characteristics of the luminance signal and the characteristics of the chrominance signal may be different from each other. Considering the difference between the characteristics, different block partition structures may be used for the luminance channel and the chrominance channel, respectively, in order to achieve more efficient encoding.
[1085] Hereinafter, the case where the block partition structure of an image is the same for a luma signal and a chroma signal (or multiple channels) is referred to as a "single-tree block partition structure" or "single-tree".
[1086] Hereinafter, the case where the block partition structure of an image is different for a luminance signal and a chrominance signal (or multiple channels) is referred to as a "dual-tree block partition structure" or "dual-tree".
[1087] In an embodiment, a block of another channel corresponding to a target block of a target channel may be specified between a luminance channel and a chrominance channel (or multiple channels). The block of another channel corresponding to a target block of a target channel is referred to as a "corresponding block (col-block)".
[1088] In an embodiment, when the luma channel and the chroma channel (or channels) have the same block partition structure (ie, when a single tree is applied), a block of another channel corresponding to a target block of a target channel may be designated.
[1089] In an embodiment, when the luma channel and the chroma channel (or channels) have different block partition structures (ie, when dual-tree is applied), a block of another channel corresponding to a target block of a target channel may be designated.
[1090] When the luma channel and chroma channel (or channels) have different block partition structures (i.e., when dual-tree is applied), the encoding decision information of the corresponding block can be shared with the target block of the target channel. By sharing the encoding decision information, the target block is encoded, thereby improving encoding efficiency.
[1091] Encoding decision information of a corresponding block corresponding to a target block of a target channel may be parsed from the compressed bitstream.
[1092] Encoding decision information of a corresponding block corresponding to the target block of the target channel may be parsed from the compressed bitstream, and the target block of the target channel may be decoded using the encoding decision information of the corresponding block.
[1093] For example, transform_skip_flag information of a corresponding block corresponding to a target block of a target channel may be parsed, and the transform_skip_flag information of the corresponding block may be used to determine whether transformation is to be skipped for the target block.
[1094] The shared information may be encoding decision information shared between the luminance block and the chrominance block.
[1095] When the block partition structure of the first channel is identical to the block partition structure of the second channel, if the spatial position of the first block of the first channel corresponds to the spatial position of the second block of the second channel, the first block and the second block may correspond to each other (i.e., may be co-located). In other words, corresponding blocks of different channels may be blocks in different channels with corresponding (co-located) spatial positions. Shared information of the second block corresponding to the first block may be used for encoding and / or decoding of the first block.
[1096] When the block partition structure of the chroma channel is the same as the block partition structure of the luma channel, the luma block corresponding to a specific chroma block of the chroma channel may be a luma block at a spatial position corresponding to the spatial position of the specific chroma block. In this case, shared information of the luma block corresponding to the chroma block may be used for encoding and / or decoding of the chroma block.
[1097] Figure 22 A single-tree block partition structure is shown.
[1098] Figure 23 A dual-tree block partition structure is shown.
[1099] Under the 4:2:0 color subsampling structure, the luma signal region spatially corresponding to the chroma block can occupy an area four times larger than the chroma block. In other words, the horizontal length (width) and vertical length (height) of the luma signal region can be twice the horizontal length and vertical length of the chroma block.
[1100] like Figure 22As shown in , in corresponding image regions of the chroma channel and the luma channel, the block partition structure of the chroma channel and the block partition structure of the luma channel may be identical to each other. In other words, a single tree may be used for the chroma channel and the luma channel.
[1101] exist Figure 22 and subsequent drawings, image regions corresponding to each other are indicated by “corresponding regions”.
[1102] like Figure 23 As shown in , in the corresponding image areas of the chroma channel and the luma channel, the block partition structure of the chroma channel and the block partition structure of the luma channel may be different from each other. In other words, a dual tree may be used for the chroma channel and the luma channel.
[1103] For example, in Figure 23 In , the area of the luma channel spatially corresponding to one chroma block can be partitioned into eight blocks.
[1104] When the block partition structures in corresponding image areas of the luma channel and the chroma channels are identical to each other, a block of the luma channel that spatially corresponds to a specific chroma block can be explicitly specified.
[1105] On the contrary, when the block partition structure of the luma channel is different from the block partition structure of the chroma channel, for a designated chroma block determined by partitioning according to the block partition structure of the chroma channel, the luma block corresponding to the designated chroma block may not be explicitly designated in the luma channel. Figure 23 As shown in , this ambiguity is attributed to the fact that the block partition structure of the chroma channel and the block partition structure of the luma channel are different from each other.
[1106] In an embodiment, for a case where a block partition structure of a luma channel is different from a block partition structure of a chroma channel, a method for specifying a luma block corresponding to a chroma block will be described.
[1107] For example, multiple luma blocks corresponding to a chroma block may be specified. With this specification, one (or more) pieces of shared information may be obtained from one or more luma blocks corresponding to the chroma block as the target block, and the obtained one (or more) pieces of shared information may be used to perform encoding and / or decoding of the chroma block.
[1108] Below, a method for specifying one or more blocks of the second channel corresponding to the blocks of the first channel will be described for a case where the block partition structure of the first channel is different from the block partition structure of the second channel. In the following description, although the first channel will be described as a chroma channel and the second channel will be described as a luma channel, the chroma channel and the luma channel are merely exemplary, and as described above, the first channel and the second channel may be different types of channels.
[1109] Figure 24A scheme for specifying a corresponding block based on a position in a corresponding area according to an example is shown.
[1110] The corresponding area indicating the luminance block corresponding to the chrominance block may be specified as a rectangular area. The position of the uppermost leftmost pixel in the rectangular area may be (xCb, yCb). The position of the lowermost rightmost pixel in the rectangular area may be (xCb+cbWidth-1, yCb+cbHeight-1).
[1111] The position (xCb, yCb) may indicate the position of a luma pixel corresponding to the position of the uppermost leftmost pixel in a chroma block (ie, a chroma coding block).
[1112] cbWidth and cbHeight may be values indicating the width and height of the target block based on luma pixels, respectively.
[1113] In other words, the corresponding area indicating the luminance block corresponding to the chrominance block can be defined as a rectangular area, where the position of the uppermost and leftmost pixel based on the position of the luminance pixel in the rectangular area is (xCb, yCb), and the rectangular area has a horizontal width cbWidth and a vertical height cbHeight.
[1114] The corresponding area indicating the luminance block corresponding to the above chrominance block can be applied to the above reference Figures 24 to 31 Described embodiment.
[1115] exist Figure 24 In the embodiment, the luminance block corresponding to the chrominance block may be a luminance block existing at a predefined position in a corresponding region of a luminance channel spatially corresponding to the chrominance block.
[1116] In other words, a luma block existing at a predefined position in a corresponding region of a luma channel spatially corresponding to a chroma block may be designated as one or more luma blocks corresponding to the chroma block. Alternatively, a luma block occupying a predefined position in a corresponding region of a luma channel spatially corresponding to a chroma block may be designated as one or more luma blocks corresponding to the chroma block.
[1117] For example, the predefined positions may be a center (CR) position, a top left (TL) position, a top right (TR) position, a bottom left (BL) position, and a bottom right (BR) position in an area of a luma channel spatially corresponding to a chroma block.
[1118] The CR position can be indicated as (xCb+cbWidth / 2, yCb+cbHeight / 2). The TL position can be indicated as (xCb, yCb). The TR position can be indicated as (xCb+cbWidth-1, yCb). The BL position can be indicated as (xCb, yCb+cbHeight-1). The BR position can be indicated as (xCb+cbWidth-1, yCb+cbHeight-1).
[1119] In an embodiment, a block of another channel corresponding to a target block of a target channel may be specified in multiple channels (such as a luminance channel and a chrominance channel). The block of another channel corresponding to the target block of the target channel is referred to as a "corresponding block".
[1120] In an embodiment, in an area of a luma channel spatially corresponding to a chroma block, a luma block including a luma pixel present at a position (xCb+cbWidth / 2, yCb+cbHeight / 2) indicating a center (CR) may be a corresponding block. Therefore, when a target block is partitioned in the form of a dual tree, a specific block of another channel (e.g., a luma channel) corresponding to a target block of a target channel (e.g., a chroma channel) among multiple channels may be explicitly specified. In an embodiment, in encoding and / or decoding of a chroma channel, a luma block including a luma pixel present at a position (xCb+cbWidth / 2, yCb+cbHeight / 2) may be specified as a corresponding block. Information about the specified corresponding block may be used to encode and / or decode the target block.
[1121] For example, the predefined positions may be some of the CR position, the TL position, the TR position, the BL position, and the BR position in the region of the luma channel spatially corresponding to the chroma block.
[1122] For example, the luminance block corresponding to the chrominance block may be a block including at least one of pixels located at the following positions in the area of the luminance channel spatially corresponding to the chrominance block: a center position, an upper left position, an upper right position, a lower left position, and a lower right position.
[1123] For example, the luminance blocks corresponding to the chrominance blocks may be some blocks of blocks including at least one of the pixels located at the following positions in the area of the luminance channel spatially corresponding to the chrominance block: the center position, the upper left position, the upper right position, the lower left position, and the lower right position.
[1124] For example, a luma block corresponding to a chroma block may include a block including at least one of pixels located at the following positions in an area of the luma channel spatially corresponding to the chroma block: a center position, an upper left position, an upper right position, a lower left position, and a lower right position.
[1125] Figure 25A scheme for specifying corresponding blocks based on areas in corresponding regions according to an example is shown.
[1126] Figure 26 Another scheme for specifying a corresponding block based on an area in a corresponding region according to an example is shown.
[1127] The luma block corresponding to the chroma block may indicate a luma block having a largest area in a region of a luma channel spatially corresponding to the chroma block.
[1128] Alternatively, the luminance blocks corresponding to the chrominance blocks may be a predefined number of luminance blocks having the largest area in a region of the luminance channel spatially corresponding to the chrominance blocks.
[1129] Such a designation scheme is attributable to the fact that there is a high probability that the characteristics of the luma block with the largest area or a predefined number of luma blocks with the largest areas in the region of the luma channel corresponding to the chroma blocks will be similar to those of the chroma blocks.
[1130] like Figure 25 As shown in , the predetermined number may be 2. Figure 25 , two blocks (ie, block 1 and block 2) having the largest areas may be selected from eight luminance blocks in a region of a luminance channel corresponding to a chrominance block.
[1131] like Figure 26 As shown in , the predetermined number may be 3. Figure 26 , three blocks (ie, block 1, block 2, and block 3) having the largest areas may be selected from eight luminance blocks in a region of a luminance channel corresponding to a chrominance block.
[1132] Through this designation scheme, even if the block partition structures of the luma channel and the chroma channel are different from each other, encoding efficiency can be improved by utilizing shared information.
[1133] Figure 27 A scheme for specifying a corresponding block based on the form of the block in the corresponding region according to an example is shown.
[1134] Figure 28 Another scheme for specifying a corresponding block based on the form of the block in the corresponding region according to an example is shown.
[1135] The luma block corresponding to the chroma block may indicate a luma block having the same form as the chroma block in a region of a luma channel spatially corresponding to the chroma block.
[1136] For example, the form of the block may include the size of the block.
[1137] According to the dual-tree block partition structure, such as Figure 27 and 28, the block partition structure of the CU of the chroma channel may be different from the block partition structure of the CU in the luma channel region corresponding to the CU of the chroma channel. Even in this case, the region of the luma channel corresponding to the region of the target block in the CU of the chroma channel may exist as a single block.
[1138] In this case, a single luminance block in the region of the luminance channel corresponding to the chrominance block can be accurately matched with the chrominance channel, and shared information of the matched luminance blocks can be shared as encoding decision information of the chrominance block.
[1139] Figure 29 A scheme for specifying corresponding blocks based on aspect ratios of blocks in a corresponding region according to an example is shown.
[1140] Figure 30 Another scheme for specifying corresponding blocks based on aspect ratios of blocks in a corresponding region according to an example is shown.
[1141] The luminance block corresponding to the chrominance block may be a luminance block having the same aspect ratio as the chrominance block in a region of a luminance channel spatially corresponding to the chrominance block.
[1142] Alternatively, the luminance block corresponding to the chrominance block may be a luminance block having an aspect ratio similar to that of the chrominance block in a region of a luminance channel spatially corresponding to the chrominance block.
[1143] For example, in Figure 29 In the area of the luminance channel, luminance block 2 can be selected, and Figure 30 In the area of the luminance channel, luminance block 1 can be selected.
[1144] Here, the aspect ratio of the block may be a ratio of the horizontal length to the vertical length of the corresponding block. In other words, the aspect ratio of the block may be a value obtained by dividing the horizontal length of the block by the vertical length of the block.
[1145] For example, the following Equation 13 may be used to determine whether the aspect ratios of the blocks are equal to each other:
[1146] [Equation 13]
[1147] (log2Width Chroma -log2Height Chroma )==(log2Width Luma -log2Height Luma )
[1148] Width Chroma Can be the width of the chroma block. Chroma Can be the height of the chroma block.
[1149] Width Luma Can be the width of the luma block corresponding to the chroma block. Luma It can be the height of the luma block corresponding to the chroma block.
[1150] The following equation 14 can be used to determine whether the aspect ratios are similar to each other:
[1151] [Equation 14]
[1152] |(log2WidthChroma-log2HeightChroma)-(log2WidthLuma-log2HeightLuma)| <THD
[1153] “|x|” may indicate the absolute value of x.
[1154] THD may be a threshold value. For example, the THD value may be 2.
[1155] According to Equations 13 and 14, one or more luma blocks having the same aspect ratio as the chroma blocks may be designated as corresponding blocks. Alternatively, one or more luma blocks having a similar aspect ratio as the chroma blocks may be designated as corresponding blocks.
[1156] Figure 31 A scheme for specifying corresponding blocks based on encoding properties of blocks in a corresponding region according to an example is shown.
[1157] Although the block partition structure of the chroma channel and the block partition structure of the luma channel are independent of each other, if there is a luma block having the same encoding decision information as the chroma block among the luma blocks in the area of the luma channel spatially corresponding to the chroma block, the shared information of the chroma block and the shared information of the luma block may be the same as each other.
[1158] In order to utilize these characteristics, a luminance block having the same value as a chrominance block for predefined coding decision information may be designated as a luminance block corresponding to the chrominance block. Alternatively, a luminance block having a value similar to that of the chrominance block for predefined coding decision information may be designated as a luminance block corresponding to the chrominance block.
[1159] For example, the predefined encoding decision information may be information on whether intra prediction is used, intra prediction mode, motion prediction information, motion vector, information on whether merge mode is used, derivation mode, transform selection information, and the like.
[1160] For example, intra prediction mode can be used as predefined coding decision information. Figure 31, luma blocks 1, 2, and 3 are shown in the area of the luma channel. Luma blocks 1, 2, and 3 may be luma blocks having the same (or similar) intra-prediction mode as that of the chroma blocks. Luma blocks 1, 2, and 3 may be designated as luma blocks corresponding to the chroma blocks.
[1161] According to the above specific embodiment, a luma block corresponding to a chroma block may include multiple luma blocks. When the multiple pieces of shared information for the multiple luma blocks are identical, there may be no problem when the multiple pieces of shared information are used for the chroma blocks. On the other hand, when the multiple pieces of shared information for the multiple luma blocks are different, the shared information to be used for encoding and / or decoding of the chroma blocks may be unclear.
[1162] The value corresponding to the majority of the values of the multiple pieces of shared information for the multiple corresponding blocks may be used as the value of the shared information for the chrominance block. This decision method may be referred to as a "majority-based shared information decision method." This method allows the information to be shared to be efficiently shared without additional signaling.
[1163] For example, when sharing transform_skip_flag information, values occupying most of the values of multiple pieces of transform_skip_flag information of multiple luma blocks corresponding to the chroma block can be shared as the value of the transform_skip_flag information of the chroma block. With this sharing, encoding and / or decoding of the chroma block can be performed.
[1164] For example, shared information is used only when the number of luma blocks corresponding to a chroma block is only one. Shared information is used only when there is only one luma block that satisfies the above-mentioned specific conditions in the region of the luma channel spatially corresponding to the chroma block. Alternatively, when the region of another channel spatially corresponding to the target block of the target channel is specified by only a single block, coding decision information can be shared between channels.
[1165] Reference Figure 27 When the region of the luma channel spatially corresponding to the target block of the chroma channel is partitioned into only one block, the one block may be the corresponding block of the luma channel. Such shared information of the corresponding block may be used to encode and / or decode the chroma block.
[1166] When a region of the luma channel spatially corresponding to a target block of the chroma channel is not partitioned into one block, encoding decision information may not be shared without separate signaling.
[1167] Figure 32 is a flowchart of an encoding method according to an embodiment.
[1168] The encoding method and the bitstream generating method according to this embodiment may be performed by the encoding apparatus 1600. This embodiment may be a part of a target encoding method or a video encoding method.
[1169] In step 3210 , the processing unit 1610 may determine encoding decision information of a representative channel of the target block.
[1170] In step 3220 , the processing unit 1610 may generate information about the target block by performing encoding o...
Claims
1. A video decoding method, comprising: Determine the coded information; as well as Decoding the current block using the encoding information, wherein: The current block includes at least three blocks.
2. The video decoding method according to claim 1, wherein: The encoding information includes intra prediction directions for intra prediction, and the intra prediction directions are used to perform intra prediction on the at least three blocks, respectively.
3. The video decoding method according to claim 2, wherein: The at least three blocks have the same size.
4. The video decoding method according to claim 1, wherein: Whether to perform the decoding on the at least three blocks using the encoding information is determined based on information acquired from a bitstream.
5. The video decoding method according to claim 1, wherein: The decoding includes a multiple transform selection (MTS) for the block.
6. The video decoding method according to claim 1, wherein: The at least three blocks include a luminance block and two chrominance blocks.
7. The video decoding method according to claim 6, wherein: When a single tree structure is used for the luminance block and the two chrominance blocks, the encoding information is applied to the at least three blocks.
8. The video decoding method according to claim 1, wherein: The decoding includes filtering of reference samples of the at least three blocks.
9. The video decoding method according to claim 1, wherein: The decoding includes filtering of a prediction block of the at least three blocks.
10. The video decoding method according to claim 1, wherein: The decoding includes at least one of loop filtering, transform, cross-component linear model (CCLM) processing, and multi-transform selection (MTS) for the at least three blocks.
11. A video encoding method, comprising: Generate bitstream information for the current block by performing encoding on the current block; wherein the encoding of the current block is performed using encoding information, and The current block includes at least three blocks.
12. The video encoding method according to claim 11, wherein: The encoding information includes intra prediction directions for intra prediction, and the intra prediction directions are used to perform intra prediction on the at least three blocks, respectively.
13. The video encoding method according to claim 11, wherein: The encoding includes a multiple transform selection (MTS) for a block.
14. The video encoding method according to claim 11, wherein: The at least three blocks include a luminance block and two chrominance blocks.
15. The video encoding method according to claim 11, wherein: The encoding includes filtering of a prediction block of the at least three blocks.
16. The video encoding method according to claim 11, wherein: The encoding includes at least one of reference sample filtering, loop filtering, transform, cross-component linear model (CCLM) processing, and multi-transform selection (MTS) for the at least three blocks. 17 . A computer-readable medium storing a bit stream generated by an encoding device through the video encoding method according to claim 11 .
18. A computer-readable recording medium storing a bit stream, the bit stream comprising: Encoded information; The encoding information is information used to perform decoding on the current block, and The current block includes at least three blocks.
19. The computer-readable recording medium according to claim 18, wherein The encoding information includes intra prediction directions for intra prediction, and the intra prediction directions are used to perform intra prediction on the at least three blocks, respectively.
20. The computer-readable recording medium according to claim 18, wherein The decoding includes a multiple transform selection (MTS) for the block.
21. The computer-readable recording medium according to claim 18, wherein The at least three blocks include a luminance block and two chrominance blocks.
22. The computer-readable recording medium according to claim 18, wherein The decoding includes filtering of a prediction block of the at least three blocks.
23. The computer-readable recording medium according to claim 18, wherein The decoding includes at least one of reference sample filtering, loop filtering, transform, cross-component linear model (CCLM) processing, and multi-transform selection (MTS) for the at least three blocks.
24. A computer-readable medium storing a bitstream generated by a video encoding device performing a video encoding method, the video encoding method comprising: Generate bitstream information for the current block by performing encoding on the current block; as well as storing a bitstream including the bitstream information in the computer-readable medium, wherein the encoding of the current block is performed using encoding information, and The current block includes at least three blocks.
25. A method for transmitting a bitstream, the method comprising: sending the bit stream including the encoded information; The encoding information is information used to perform decoding on the current block, and The current block includes at least three blocks.
26. The method of claim 25, wherein: The encoding information includes intra prediction directions for intra prediction, and the intra prediction directions are used to perform intra prediction on the at least three blocks, respectively.
27. The method of claim 25, wherein: The decoding includes a multiple transform selection (MTS) for the block.
28. The method of claim 25, wherein: The at least three blocks include a luminance block and two chrominance blocks.
29. The method of claim 25, wherein: The decoding includes filtering of a prediction block of the at least three blocks.
30. The method of claim 25, wherein: The decoding includes at least one of reference sample filtering, loop filtering, transform, cross-component linear model (CCLM) processing, and multi-transform selection (MTS) for the at least three blocks.
31. A video decoding method, comprising: Determine the coded information; as well as Perform decoding on the current block using the encoding information, The current block includes multiple blocks.
32. A video encoding method, comprising: Generate bitstream information for the current block by performing encoding on the current block; wherein the encoding of the current block is performed using encoding information, and The current block includes a plurality of blocks.
33. A computer-readable medium storing a bit stream generated by an encoding device through the video encoding method according to claim 32.
34. A computer-readable medium storing a bitstream, the bitstream comprising: Encoded information; The encoding information is information used to perform decoding on the current block, and The current block includes multiple blocks.
35. A computer-readable medium storing a bitstream generated by a video encoding device performing a video encoding method, the video encoding method comprising: Generate bitstream information for the current block by performing encoding on the current block; as well as storing a bitstream including the bitstream information in the computer-readable medium, wherein the encoding of the current block is performed using encoding information, and The current block includes a plurality of blocks.
36. A method for transmitting a bitstream, the method comprising: sending the bit stream including the encoded information; The encoding information is information used to perform decoding on the current block, and The current block includes multiple blocks.