Image encoding / decoding method and device, and recording medium storing bitstream
By employing multiple transform selection (MTS) for chroma components with tailored transform types and independent signaling, the method addresses inefficiencies in encoding/decoding high-resolution images, enhancing compression performance and efficiency.
Patent Information
- Application Number
- PCT/KR2025/010397
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Existing video encoding/decoding technologies face challenges in efficiently compressing high-resolution, high-quality images due to limitations in transform techniques for chroma components, leading to suboptimal compression performance and inefficiencies in signaling transformation indices.
The implementation of multiple transform selection (MTS) for chroma components, allowing for separable or non-separable transformations, and independent signaling of transformation indices, tailored to the properties and modes of chroma blocks, enhances transform types based on block characteristics and prediction modes.
This approach improves conversion and transformation performance by selectively utilizing various transform types for chroma components, leading to more efficient encoding and decoding of high-resolution images.
Smart Images

Figure KR2025010397_22012026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.
[0003] There are various technologies for image compression, such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information in the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data.
[0004] The present disclosure seeks to provide a method and device for performing a transformation using a separable transformation or a non-separable transformation.
[0005] The present disclosure provides a method and apparatus for performing transformation based on multiple transform selection (MTS) for chroma components.
[0006] The present disclosure seeks to provide a method and apparatus for signaling a transformation index or a transformation set index for MTS.
[0007] A video decoding method and device according to the present disclosure can derive transform coefficients of a current block from a bitstream, derive residual samples of the current block based on inverse transform of the transform coefficients of the current block, and reconstruct the current block based on the residual samples of the current block. The current block can include at least one of a luma block and a chroma block. A transform type for inverse transform of the chroma block can be determined based on a transform index specifying one of the available transform types.
[0008] In the video decoding method and device according to the present disclosure, the available transformation types may be defined differently for each video unit. The video unit may include at least one of a coding unit, a transformation unit, a slice, a picture, or a sequence.
[0009] In the video decoding method and device according to the present disclosure, the transformation types available to the chroma block may be included in the transformation types available to the luma block.
[0010] In the video decoding method and device according to the present disclosure, the chroma block may have available transform types separate from those of the luma block.
[0011] In the video decoding method and device according to the present disclosure, the number of transform types available to the chroma block may be determined based on at least one of the properties of transform coefficients for the chroma block, the size of the current block, the size of the chroma block, the MTS property for the luma block, the MTS property of a neighboring chroma block, or mode information of the current block.
[0012] In the video decoding method and device according to the present disclosure, the transformation types available to the chroma block can be defined respectively for the horizontal transformation and the vertical transformation of the chroma block.
[0013] In the video decoding method and device according to the present disclosure, the transformation types available to the chroma block can be determined based on the prediction mode of the chroma block.
[0014] In the video decoding method and device according to the present disclosure, if the chroma block is a block encoded in a linear model-based prediction mode, one of a plurality of transformation sets may be selected based on the intra prediction mode of the luma block. The transformation types belonging to the selected transformation set may be determined as transformation types available to the chroma block.
[0015] In the video decoding method and device according to the present disclosure, if the chroma block is not a block encoded based on an intra prediction mode of a directional mode or a non-directional mode, one of a plurality of transformation sets may be selected based on a DIPM (derived intra prediction mode) for the chroma block. The transformation types belonging to the selected transformation set may be determined as transformation types available to the chroma block.
[0016] In the video decoding method and device according to the present disclosure, the transformation types available to the chroma block can be determined based on the transformation set of the luma block and a pre-defined lookup table. The lookup table can define a mapping relationship between the transformation set of the luma block and the transformation set of the chroma block.
[0017] In the image decoding method and device according to the present disclosure, the transformation index can be signaled independently from the transformation index for the luma block.
[0018] In the image decoding method and device according to the present disclosure, the transformation index can be derived based on the transformation index for the luma block.
[0019] A video encoding method and device according to the present disclosure can derive residual samples of a current block, derive transform coefficients of the current block based on a transformation of the residual samples of the current block, and encode residual information regarding the transform coefficients of the current block. The current block can include at least one of a luma block and a chroma block. A transform type for transforming the chroma block can be determined as any one of the available transform types.
[0020] A computer-readable digital storage medium is provided, which stores encoded video / image information that causes a decoding device according to the present disclosure to perform a video decoding method.
[0021] A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided.
[0022] A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided.
[0023] The present disclosure can improve conversion performance by using separable conversion or non-separable conversion.
[0024] The present disclosure can improve transformation performance by applying an MTS that selectively utilizes at least one of a plurality of transformation types for a chroma component.
[0025] The present disclosure can efficiently signal a transformation index or a transformation set index for MTS.
[0026] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0027] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0028] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0029] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0030] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs an image decoding method according to the present disclosure.
[0031] FIG. 6 illustrates an image encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0032] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure.
[0033] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0034] The present disclosure may be subject to various modifications and embodiments. Therefore, specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Similar reference numerals have been used to designate similar components throughout the description of each drawing.
[0035] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0036] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0037] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).
[0039] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0040] In this specification, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0041] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0042] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0043] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0044] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0045] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0046] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0047] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."
[0048] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.
[0049] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0050] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).
[0051] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0052] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0053] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0054] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0055] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0056] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0057] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0058] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0059] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0060] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0061] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0062] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0063] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0064] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0065] The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0066] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0067] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal.
[0068] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.
[0069] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0070] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients.
[0071] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0072] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0073] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0074] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0075] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks within the current picture and transfer them to the intra prediction unit (222).
[0076] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0077] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (331) and an intra-prediction unit (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0078] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0079] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (300) can be reproduced through a reproduction device.
[0080] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).
[0081] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the adding unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).
[0082] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0083] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0084] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.
[0085] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be included and signaled in the video / image information.
[0086] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0087] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0088] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.
[0089] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0090] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0091] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (260) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331).
[0092] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0093] Referring to FIG. 4, transform coefficients of the current block can be derived from the bitstream (S400).
[0094] Residual information of the current block can be obtained from the bitstream, and the obtained residual information can be decoded to derive quantized transform coefficients. Transform coefficients can be derived by performing inverse quantization based on the quantized transform coefficients.
[0095] Referring to FIG. 4, residual samples of the current block can be derived based on inverse transformation of the transform coefficients of the current block (S410).
[0096] The above inverse transformation can be performed by selectively utilizing at least one of a plurality of pre-defined transformation types. Hereinafter, selectively utilizing at least one of a plurality of transformation types is referred to as MTS (multiple transform selection).
[0097] The multiple transform types for MTS may include at least one of a DCT-based transform, a DST-based transform, or an identity transform (IDT). Here, the DCT-based transform may include at least one of DCT2, DCT3, DCT5, or DCT8. The DST-based transform may include at least one of DST1, DST4, or DST7.
[0098] MTS may be based on a separable transform. A separable transform may refer to a method of performing a two-dimensional transformation by separating it into a one-dimensional transformation in the horizontal direction and a one-dimensional transformation in the vertical direction. Alternatively, MTS may be based on a non-separable transform. A non-separable transform may refer to a method of performing a transformation all at once based on a two-dimensional transformation matrix without separating it into horizontal and vertical transformations. Here, the transformation kernels for the non-separable transform may be identically pre-defined in the encoding device and the decoding device. Alternatively, the transformation kernels for the non-separable transform may be signaled through the bitstream.
[0099] The available transformation types for MTS can be defined differently for each image unit. Here, an image unit can include at least one of a coding unit (CU), a transform unit (TU), a slice, a picture, or a sequence.
[0100] For example, seven available transform types can be defined for a sequence. The available transform types for a sequence can be defined in a Sequence Parameter Set (SPS). Five available transform types can be defined for a current slice. The available transform types for the current slice can be defined in a Slice Header (SH). Three available transform types can be defined for a current block (CU or TU). Here, the available transform types can be explicitly signaled through the bitstream. Alternatively, to reduce signaling overhead, the available transform types can be implicitly derived. The available transform types can be implicitly derived based on at least one of a block size, a CU coding type, a block type, or a characteristic of a transform coefficient. Here, the CU coding type can mean whether a subblock transform (SBT) is applied to the corresponding CU. The characteristic of the transform coefficient can include at least one of the number of non-zero transform coefficients or the magnitude of the transform coefficients.
[0101] In addition to the luma component of the current block, MTS can also be applied to the chroma component of the current block. For convenience of explanation, the luma component of the current block will be referred to as a luma block, and the chroma component of the current block will be referred to as a chroma block. That is, the inverse transformation for a chroma block can be performed by selectively using at least one of the transformation types available to the chroma block. Here, the available transformation types can be all or some of the multiple pre-defined transformation types described above.
[0102] The transformation types available to a chroma block may be included in the transformation types available to a luma block. For example, the transformation types available to a chroma block may be the same as all of the transformation types available to a luma block. Alternatively, the transformation types available to a chroma block may be a subset of the transformation types available to a luma block.
[0103] A chroma block may have separate available transformation types from a luma block. For example, available transformation types may be defined separately for chroma blocks and luma blocks. At least one of the available transformation types for a chroma block may not be among the available transformation types for a luma block.
[0104] The number of transformation types available to a chroma block may be a fixed number pre-defined equally for the encoding device and the decoding device.
[0105] Alternatively, the number of transform types available to a chroma block may be determined based on at least one of properties of transform coefficients for the chroma block, the size of the current block (or luma block), the size of the chroma block, an MTS property for the luma block, an MTS property of a neighboring chroma block, or mode information of the current block.
[0106] The properties of the above transform coefficients may include at least one of the number of non-zero transform coefficients, the sum of the absolute values of the transform coefficients, or the position of the last significant coefficient.
[0107] The above size may be defined by at least one of width, height, maximum / minimum of width and height, product of width and height, or ratio of width and height.
[0108] The above-mentioned neighboring chroma block may refer to a chroma component of a neighboring block adjacent to the current block. The neighboring block may include at least one of a left block, an upper block, an upper-left block, an upper-right block, or a lower-left block.
[0109] The above MTS property may include at least one of whether MTS is enabled, whether MTS is applied, the number of available transformation types, or an MTS index.
[0110] The above mode information may relate to any one of the pre-defined prediction modes. The pre-defined prediction modes may include at least one of an intra mode, a linear model-based prediction mode, an intraTMP (intra template matching prediction) mode, a MIP (matrix-based intra prediction) mode, or an inter mode.
[0111] The available transformation types described above may be equally applied to the horizontal and vertical directions of a chroma block. Alternatively, the available transformation types may be defined separately for the horizontal and vertical directions of a chroma block. That is, at least one of the transformation types available for the horizontal direction may not belong to the transformation types available for the vertical direction.
[0112] Depending on the prediction mode of the chroma block, the available transformation types for the chroma block can be determined.
[0113] For example, if a chroma block is a block encoded with a linear model-based prediction mode, the intra prediction mode of the chroma block can be derived based on the intra prediction mode of the luma block. Based on the derived intra prediction mode of the chroma block, one of a plurality of transform sets for the chroma component can be selected. The transform types belonging to the selected transform set can be determined as transform types available to the chroma block. A cross-component linear model (CCLM), a multi-model linear model (MMLM), a convolutional cross-component model (CCCM), etc. can be used as the linear model-based prediction mode.
[0114] For example, if a chroma block is not a block encoded based on an intra prediction mode such as a directional mode or a non-directional mode, a derived intra prediction mode (DIPM) can be derived for the chroma block, and one of a plurality of transform sets for the chroma component can be selected based on the DIPM. The transform types belonging to the selected transform set can be determined as the transform types available to the chroma block.
[0115] The above DIPM can be derived by applying a predetermined filter to a surrounding area of a chroma block. Specifically, a predetermined filter can be applied to each sample belonging to the surrounding area of the chroma block to calculate a horizontal direction sample variation and a vertical direction sample variation for each sample. An intensity value can be derived based on the calculated horizontal / vertical direction sample variation. The corresponding intensity value can be accumulated in an intra prediction mode mapped to an angle according to the horizontal / vertical direction sample variation. Through the above-described process, an intensity value can be accumulated for one or more intra prediction modes, and a DIPM can be derived from one or more intra prediction modes for which intensity values have been accumulated.
[0116] DIPM can also be derived by applying a predetermined filter to all or part of the prediction samples of a chroma block generated through intra prediction. In this case, the horizontal sample variation and the vertical sample variation for each prediction sample can be calculated, and intensity values can be accumulated for one or more intra prediction modes, and the DIPM can be derived from one or more intra prediction modes for which intensity values have been accumulated.
[0117] When the chroma block is not a block encoded based on an intra prediction mode such as a directional mode or a non-directional mode, it may mean that the chroma block is a block encoded based on an IntraTMP or MIP (matrix-based intra prediction) mode. Alternatively, when the chroma block is not a block encoded based on an intra prediction mode such as a directional mode or a non-directional mode, it may mean that the chroma block is a block encoded based on an inter mode.
[0118] For example, if a chroma block is a block encoded based on an intra prediction mode such as a directional mode or a non-directional mode, one of a plurality of transform sets for the chroma component may be selected based on the corresponding intra prediction mode of the chroma block. The transform types belonging to the selected transform set may be determined as the transform types available to the chroma block.
[0119] The transformation types available to a chroma block may also be determined based on the transformation set of the luma block.
[0120] For example, a look-up table defining a mapping relationship between a transformation set of a luma block and a transformation set of a chroma block may be used. From the look-up table, a transformation set of a chroma block corresponding to the transformation set of the luma block may be derived. Transformation types belonging to the derived transformation set may be determined as transformation types available to the chroma block. Here, the transformation set of the luma block may be any one of a plurality of transformation sets for the luma component. The transformation set of the chroma block may be any one of a plurality of transformation sets for the chroma component.
[0121] The plurality of transform sets for the aforementioned chroma component may be identical to the plurality of transform sets for the luma component. That is, the luma component and the chroma component may share the plurality of transform sets pre-defined in the encoding device and the decoding device. Alternatively, the plurality of transform sets for the chroma component may be defined separately from the plurality of transform sets for the luma component. For example, the number of transform sets for the chroma component may be different from the number of transform sets for the luma component. At least one of the plurality of transform sets for the chroma component may not have the same transform types as at least one of the plurality of transform sets for the luma component.
[0122] The transform type of a chroma block can be derived based on at least one of the available transform types. To this end, a transform index specifying at least one of the available transform types can be signaled through the bitstream. Alternatively, the transform index can be derived based on the transform index of the corresponding luma block.
[0123] At this time, a single transformation index can be signaled / derived for the chroma block. In this case, the transformation index can specify a transformation type of the chroma block among available transformation types. Here, the available transformation types can correspond to transformation types for non-separable transformation. Based on the single transformation index, a transformation type for non-separable transformation of the chroma block can be derived.
[0124] Alternatively, a single transformation index may be signaled / derived for a chroma block. In this case, the transformation index may specify a transformation type of the chroma block among available transformation types. Here, each of the available transformation types may include a transformation type for transformation in the horizontal direction (i.e., a horizontal transformation type) and a transformation type for transformation in the vertical direction (i.e., a vertical transformation type). In other words, the available transformation types may correspond to transformation types for separate transformations. Based on the single transformation index, the horizontal transformation type and the vertical transformation type of the chroma block may be derived.
[0125] Alternatively, two transformation indices may be signaled / derived for a chroma block. In this case, one of the two transformation indices may specify a horizontal transformation type among the available transformation types, and the other may specify a vertical transformation type among the available transformation types. Here, the available transformation types may correspond to transformation types for separate transformations. That is, the transformation indices may be signaled / derived for the horizontal direction and the vertical direction of the chroma block, respectively. Based on the two transformation indices, the horizontal transformation type and the vertical transformation type of the chroma block, respectively, may be derived.
[0126] Alternatively, the transformation types available to a chroma block can be defined for the horizontal and vertical directions of the chroma block, respectively. The transformation types available for the horizontal direction are called available horizontal transformation types, and the transformation types available for the vertical direction are called available vertical transformation types. In this case, two transformation indices can be signaled / derived for the chroma block. Either one of the two transformation indices can specify any one of the available horizontal transformation types, and the other one can specify any one of the available vertical transformation types. That is, the transformation indices can be signaled / derived for the horizontal and vertical directions of the chroma block, respectively. Based on the two transformation indices, the horizontal transformation type and the vertical transformation type of the chroma block, respectively, can be derived.
[0127] Through the above-described method, a transformation type of a chroma block can be derived based on a transformation index, and a non-separable transformation or a separable transformation-based inverse transformation can be performed on the chroma block based on the derived transformation type.
[0128] Each of the aforementioned multiple transformation sets may be composed of two transformation types, wherein one of the two transformation types may correspond to a horizontal transformation type and the other may correspond to a vertical transformation type.
[0129] The above horizontal transformation type may be any one of the multiple transformation types for the aforementioned MTS. Similarly, the vertical transformation type may also be any one of the multiple transformation types for the aforementioned MTS. At least one of the multiple transformation sets may be a transformation set in which the horizontal transformation type and the vertical transformation type are the same. Alternatively, the multiple transformation sets may be transformation sets in which the horizontal transformation type and the vertical transformation type are different from each other.
[0130] The horizontal and vertical transform types of a chroma block can be derived based on the horizontal and vertical transform types belonging to any one of a plurality of transform sets, respectively. An inverse transform can be performed on the chroma block based on the derived horizontal and vertical transform types. A transform set index specifying any one of the plurality of transform sets can be signaled through the bitstream. Alternatively, the transform set index can be derived based on the transform set index of the corresponding luma block.
[0131] Below, we will examine the method for signaling the aforementioned transformation index. As previously mentioned, a transformation set index may be defined instead of the transformation index, and we will also examine the method for signaling the transformation set index.
[0132] Example 1
[0133] The luma component, the first chroma component (Cb component), and the second chroma component (Cr component) can each have their own transform type. If MTS is allowed, the transform index can be signaled independently for each component.
[0134] For example, if the tree type of the current block is dual tree chroma, the transformation index for the Cb component and the transformation index for the Cr component can be signaled respectively for the current block, which is a chroma block.
[0135] The luma, Cb, and Cr components can each have their own transform sets. If MTS is enabled, the transform set index can be signaled independently for each component.
[0136] For example, if the tree type of the current block is dual tree chroma, the transform set index for the Cb component and the transform set index for the Cr component can be signaled respectively for the current block, which is a chroma block.
[0137] Example 2
[0138] Chroma components (Cb and Cr components) can have independent transform types from the luma component. On the other hand, Cb and Cr components can share the same transform type. If MTS is allowed, transform indices can be signaled for the luma and chroma components separately. A single transform index can be signaled for the Cb and Cr components, and the transform type specified by the transform index can be applied to the Cb and Cr components.
[0139] For example, if the tree type of the current block is dual tree chroma, one transform index for the Cb component and one for the Cr component can be signaled for the current block.
[0140] Chroma components (Cb and Cr components) can have independent transform sets from the luma component. On the other hand, Cb and Cr components can share the same transform set. If MTS is allowed, transform set indices can be signaled for the luma and chroma components separately. A single transform set index can be signaled for the Cb and Cr components, and the Cb and Cr components can use the same transform set specified by the transform set index.
[0141] For example, if the tree type of the current block is dual tree chroma, one transform set index for the Cb component and one for the Cr component can be signaled for the current block.
[0142] Example 3
[0143] MTS for chroma components may be dependent on MTS for luma components. That is, if MTS is applied to luma components, MTS may be allowed / applied to chroma components. Conversely, if MTS is not applied to luma components, MTS may not be allowed / applied to chroma components.
[0144] When MTS is applied to the luma component (or, when MTS is allowed / applied to the chroma component), a first flag may be signaled through the bitstream. Here, the first flag may relate to whether the transform type for the chroma component is derived based on the transform type for the luma component.
[0145] For example, if the value of the first flag is 1, this may indicate that the transform index for the chroma component is set to the transform index for the corresponding luma component. On the other hand, if the value of the first flag is 0, this may indicate that the transform index for the chroma component is not set to the transform index for the corresponding luma component. In this case, the transform index for the chroma component may be set to the value of the transform index that specifies the default transform type pre-defined in the encoding device and the decoding device (e.g., 0), and MTS may not be applied to the chroma component.
[0146] Alternatively, if the value of the first flag is 1, this may indicate that the chroma component uses the same transform type as the corresponding luma component. Conversely, if the value of the first flag is 0, this may indicate that the chroma component is not restricted to use the same transform type as the corresponding luma component. That is, if the value of the first flag is 0, the chroma component may use the same transform type as the luma component or a different transform type. In this case, the transform type for the chroma component may be set to a default transform type pre-defined in the encoding device and the decoding device, and MTS may not be applied to the chroma component.
[0147] If MTS is not applied to the luma component (or if MTS is not allowed / applied to the chroma component), the first flag may not be signaled via the bitstream. In this case, the value of the first flag may be set to 0.
[0148] If MTS is applied to the luma component (or, if MTS is allowed / applied to the chroma component), a second flag may be signaled via the bitstream. Here, the second flag may indicate whether the transform set for the chroma component is derived based on the transform set for the luma component.
[0149] For example, if the value of the second flag is 1, this may indicate that the transform set index for the chroma component is set to the transform set index for the corresponding luma component. On the other hand, if the value of the second flag is 0, this may indicate that the transform set index for the chroma component is not set to the transform set index for the corresponding luma component. In this case, the transform set index for the chroma component may be set to the value of the transform set index (e.g., 0) that specifies a default transform set pre-defined in the encoding device and the decoding device, and MTS may not be applied to the chroma component.
[0150] Alternatively, if the value of the second flag is 1, this may indicate that the chroma component uses the same transform set as the corresponding luma component. Conversely, if the value of the second flag is 0, this may indicate that the chroma component is not restricted to use the same transform set as the corresponding luma component. That is, if the value of the second flag is 0, the chroma component may use the same transform set as the luma component or a different transform set. In this case, the transform set for the chroma component may be set to a default transform set pre-defined in the encoding device and the decoding device, and MTS may not be applied to the chroma component.
[0151] If MTS is not applied to the luma component (or if MTS is not allowed / applied to the chroma component), the second flag may not be signaled via the bitstream. In this case, the value of the second flag may be set to 0.
[0152] In the case of a single tree, there is no separate coding block for the chroma component. However, in the case of a single tree, each component can have its own transformation type (or transformation set). In the case of a single tree, the chroma component has a transformation type (or transformation set) independent of the luma component, and the Cb component and Cr component can share the same transformation type (or transformation set). In the case of a single tree, the chroma component can use a transformation type (or transformation set) according to the transformation type (or transformation set) of the luma component.
[0153] Any one of embodiments 1 to 3 may be applied based on at least one of the tree type, block size, prediction mode, or intra prediction mode derivation method. For example, if the tree type is not a single tree, the signaling method of embodiments 1 or 2 may be applied. If the tree type is a single tree, the signaling method of embodiment 3 may be applied.
[0154] Whether MTS is allowed / applied to a chroma component can be determined based on at least one of a tree type, a block size, a prediction mode, or a method of deriving an intra prediction mode. For example, if the intra prediction mode for the chroma component is set to the same mode as the intra prediction mode for the luma component, MTS can be determined to be allowed or not for the chroma component. If a linear model-based prediction mode is applied to the chroma component, MTS can be determined to be allowed or not for the chroma component. If a JointCbCr mode utilizing the correlation between the Cb component and the Cr component is applied to the chroma component, MTS can be determined to be allowed or not for the chroma component.
[0155] Referring to FIG. 4, the current block can be restored based on the residual samples of the current block (S420).
[0156] The current block can be reconstructed based on the prediction samples and residual samples of the current block. Here, the prediction samples of the current block may be generated based on a predetermined prediction mode (e.g., intra mode or inter mode).
[0157] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs an image decoding method according to the present disclosure.
[0158] Referring to FIG. 5, a decoding device (300) according to the present disclosure may include a transform coefficient derivation unit (500), a residual sample derivation unit (510), and a restoration block generation unit (520). The transform coefficient derivation unit (500) may be configured in the entropy decoding unit (310) of FIG. 3, the residual sample derivation unit (510) may be configured in the residual processing unit (320) of FIG. 3, and the restoration block generation unit (520) may be configured in the adding unit (340) of FIG. 3.
[0159] The transform coefficient derivation unit (500) can obtain residual information of the current block from the bitstream and decode it to derive quantized transform coefficients of the current block.
[0160] The residual sample derivation unit (510) can derive residual samples of the current block by performing at least one of inverse quantization or inverse transformation on the quantized transform coefficients of the current block. In particular, the residual sample derivation unit (510) can determine transformation types available to the chroma components of the current block, and can derive residual samples of the current block by performing inverse transformation based on at least one of the available transformation types. This has been described with reference to FIG. 4, and a detailed description thereof will be omitted herein.
[0161] The restoration block generation unit (520) can restore the current block based on residual samples of the current block.
[0162] FIG. 6 illustrates an image encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0163] Referring to Fig. 6, residual samples of the current block can be derived (S600).
[0164] The residual samples of the current block can be derived based on the prediction samples of the current block. Here, the prediction samples may be generated based on a predetermined prediction mode (e.g., intra mode or inter mode).
[0165] Referring to FIG. 6, transform coefficients of the current block can be derived by performing at least one of transform and quantization on residual samples of the current block (S610).
[0166] Here, the transformation can be performed based on a separable transformation or a non-separable transformation. The method for determining the transformation type for the transformation of the chroma component is as described with reference to Fig. 4, and a detailed description thereof will be omitted. In addition, the method for signaling the transformation index (or transformation set index) was described with reference to Fig. 4, and this method can be equally applied to the method for encoding the transformation index (or transformation set index).
[0167] Referring to FIG. 6, a bitstream can be generated by encoding residual information regarding the transform coefficients of the current block (S620).
[0168] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an image encoding method according to the present disclosure.
[0169] Referring to FIG. 7, an encoding device (200) according to the present disclosure may include a residual sample derivation unit (700), a transform coefficient derivation unit (710), and a transform coefficient encoding unit (720). The residual sample derivation unit (700) and the transform coefficient derivation unit (710) may be configured in the residual processing unit (230) of FIG. 2, and the transform coefficient encoding unit (720) may be configured in the entropy encoding unit (240) of FIG. 2.
[0170] The residual sample derivation unit (700) can derive residual samples of the current block based on the prediction samples of the current block.
[0171] The transform coefficient derivation unit (710) can derive transform coefficients of the current block by performing at least one of transform and quantization on residual samples of the current block. In particular, the transform coefficient derivation unit (710) can perform the transform coefficient derivation method in step S610, and a duplicate description thereof will be omitted.
[0172] Additionally, the conversion coefficient derivation unit (710) can generate residual information regarding the conversion coefficients.
[0173] The residual information encoding unit (720) can generate a bitstream by encoding residual information of the current block.
[0174] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0175] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0176] When the embodiments in this document are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described function. The module can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuit, and / or data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored on a digital storage medium.
[0177] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0178] In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0179] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0180] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0181] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0182] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0183] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0184] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.
[0185] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0186] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0187] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0188] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
Claims
A step of deriving transform coefficients of the current block from a bitstream; A step of deriving residual samples of the current block based on inverse transformation of the transform coefficients of the current block; and A step of restoring the current block based on residual samples of the current block, The current block contains at least one of a luma block or a chroma block, A method in which a transformation type for inverse transformation of the above chroma block is determined based on a transformation index that specifies one of the available transformation types. In the first paragraph, The above available transformation types are defined differently for each image unit, A method wherein the image unit comprises at least one of a coding unit, a transform unit, a slice, a picture, or a sequence. In the first paragraph, A method in which the transformation types available to the chroma block are included in the transformation types available to the luma block. In the first paragraph, A method wherein the chroma block has available transformation types separate from the luma block. In the first paragraph, A method in which the number of transform types available to the chroma block is determined based on at least one of the properties of transform coefficients for the chroma block, the size of the current block, the size of the chroma block, the MTS property for the luma block, the MTS property of a neighboring chroma block, or mode information of the current block. In the first paragraph, A method in which the transformation types available to the above chroma block are defined respectively for horizontal transformation and vertical transformation of the above chroma block. In the first paragraph, A method in which the transformation types available to the chroma block are determined based on the prediction mode of the chroma block. In paragraph 7, If the above chroma block is a block encoded in a linear model-based prediction mode, one of a plurality of transformation sets is selected based on the intra prediction mode of the luma block, A method in which the transformation types belonging to the above-mentioned selected transformation set are determined as the transformation types available to the chroma block. In paragraph 7, If the chroma block is not a block encoded based on an intra prediction mode of a directional mode or a non-directional mode, one of a plurality of transform sets is selected based on a DIPM (derived intra prediction mode) for the chroma block, A method in which the transformation types belonging to the above-mentioned selected transformation set are determined as the transformation types available to the chroma block. In the first paragraph, The transformation types available to the above chroma block are determined based on the transformation set of the luma block and a pre-defined lookup table, A method wherein the above lookup table defines a mapping relationship between a transformation set of the luma block and a transformation set of the chroma block. In the first paragraph, A method wherein the above transformation index is signaled independently from the transformation index for the luma block. In the first paragraph, A method in which the above transformation index is derived based on the transformation index for the luma block. A step of deriving residual samples of the current block; A step of deriving transform coefficients of the current block based on transforms of residual samples of the current block; and A step of encoding residual information regarding transform coefficients of the current block, The current block contains at least one of a luma block or a chroma block, A method in which the transformation type for transformation of the above chroma block is determined as one of the available transformation types. A non-transitory computer-readable storage medium storing a bitstream generated by the method according to claim 13. A step of obtaining a bitstream for image information; wherein the bitstream is obtained based on deriving residual samples of a current block, deriving transform coefficients of the current block based on a transform of the residual samples of the current block, and encoding residual information about the transform coefficients of the current block, and Including a step of transmitting data including the above bitstream, The current block contains at least one of a luma block or a chroma block, A method in which the transformation type for transformation of the above chroma block is determined as one of the available transformation types.
Citation Information
Patent Citations
Method and apparatus for encoding / decoding image
KR1020160108617A
Table for nail
KR1020220026842A
Plasma thin film deposition apparatus and method of predicting thin film
KR1020250128028A
Method for increasing stability of target protein immobilized in silica nanoparticle by salt treatment
KR102420704B1
KR20200110236A