Transform-based image coding method and device therefor
The transform-based image/video coding method with MTS enhances compression efficiency by using DCT, DST, and KLT kernels, addressing the challenge of high data size in high-resolution media, thereby reducing transmission and storage costs.
Patent Information
- Application Number
- PCT/KR2025/000184
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-03
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-17
AI Technical Summary
The increasing demand for high-resolution, high-quality images and immersive media has led to a surge in data size, resulting in higher transmission and storage costs, necessitating a more efficient image/video compression technology.
A transform-based image/video coding method utilizing Multiple Transform Selection (MTS) with discrete cosine transform (DCT), discrete sine transform (DST), and Karhunen-Loève Transform (KLT) kernels, adapted based on block size and prediction mode, to enhance compression efficiency.
Improves overall video/image compression efficiency by optimizing transformation performance and signaling, reducing data size and costs associated with transmission and storage.
Smart Images

Figure KR2025000184_17072025_PF_FP_ABST
Abstract
Description
Transform-based image coding method and device thereof
[0001] The present disclosure relates to a method for image / video coding and a device therefor.
[0002] Image / video coding is used in various applications such as digital storage media, television broadcasting, video streaming services, and real-time communications, and the demand for high-resolution, high-quality images / videos is increasing in various fields.
[0003] As the image / video becomes higher resolution and higher quality, the data size of the image / video increases, and the amount of information or bits transmitted increases relatively. Therefore, when transmitting image data using media such as existing wired or wireless broadband lines or storing image / video data using existing storage media, the transmission and storage costs increase.
[0004] In addition, interest in and demand for immersive media such as VR (virtual reality), AR (artificial reality), MR (mixed reality) content and holograms have been increasing recently, and attempts to provide immersive experiences using immersive media in games, education, medicine, real estate, marketing, etc. are increasing.
[0005] Accordingly, a highly efficient image / video compression technology is required to effectively compress, transmit, store, and play high-resolution, high-quality image / video information having various characteristics as described above.
[0006] According to one embodiment of the present disclosure, a method and device for improving video / image coding efficiency are provided.
[0007] According to one embodiment of the present disclosure, a transform-based video / image coding method and device are provided.
[0008] According to one embodiment of the present disclosure, a video decoding method performed by a decoding device is provided. The method includes the steps of: receiving video information including prediction-related information and residual information for a current block; deriving a prediction mode for the current block based on the prediction-related information; deriving prediction samples for the current block based on the prediction mode; deriving transform coefficients for the current block based on the residual information; and deriving residual samples for the current block by applying MTS (Multiple Transform Selection) to the transform coefficients, wherein the video information includes MTS index information, the MTS index information indicates a set of transform kernels used in the MTS, and the set of transform kernels includes at least one of a DCT (discrete cosine transform) kernel, a DST (discrete sine transform) kernel, or a KLT (Karhunen-Loe`ve Transform) kernel based on a size of the current block.
[0009] According to one embodiment of the present disclosure, an encoding method performed by an encoding device is provided. The method includes the steps of determining a prediction mode for a current block, deriving prediction samples for the current block based on the prediction mode, deriving residual samples for the current block based on the prediction samples, deriving transform coefficients for the current block by applying MTS (Multiple Transform Selection) to the residual samples, generating residual information for the current block based on the transform coefficients, and encoding image information including prediction-related information for the current block and the residual information, wherein the image information includes MTS index information, the MTS index information indicates a set of transform kernels used for the MTS, and the set of transform kernels includes at least one of a DCT (discrete cosine transform) kernel, a DST (discrete sine transform) kernel, or a KLT (Karhunen-Loe`ve Transform) kernel based on a size of the current block.
[0010] According to one embodiment of the present disclosure, a decoding device for video decoding is provided. The decoding device includes a memory and at least one processor connected to the memory, and the at least one processor is configured to perform the steps of: receiving video information including prediction-related information and residual information for a current block; deriving a prediction mode for the current block based on the prediction-related information; deriving prediction samples for the current block based on the prediction mode; deriving transform coefficients for the current block based on the residual information; and deriving residual samples for the current block by applying MTS (Multiple Transform Selection) to the transform coefficients, wherein the video information includes MTS index information, and the MTS index information indicates a set of transform kernels used in the MTS, and the set of transform kernels includes at least one of a discrete cosine transform (DCT) kernel, a discrete sine transform (DST) kernel, or a Karhunen-Loe've Transform (KLT) kernel based on a size of the current block.
[0011] According to one embodiment of the present disclosure, an encoding device for video encoding is provided. The encoding device comprises a memory and at least one processor connected to the memory, and the at least one processor is configured to perform the steps of determining a prediction mode for a current block, deriving prediction samples for the current block based on the prediction mode, deriving residual samples for the current block based on the prediction samples, deriving transform coefficients for the current block by applying MTS (Multiple Transform Selection) to the residual samples, generating residual information for the current block based on the transform coefficients, and encoding image information including prediction-related information for the current block and the residual information, wherein the image information includes MTS index information, and the MTS index information indicates a transform kernel set used for the MTS, and the transform kernel set includes at least one of a DCT (discrete cosine transform) kernel, a DST (discrete sine transform) kernel, or a KLT (Karhunen-Loe`ve Transform) kernel based on a size of the current block.
[0012] According to one embodiment of the present disclosure, a method is provided for transmitting video / video data including a bitstream generated according to a video / video encoding method according to at least one of the embodiments of the present disclosure.
[0013] According to one embodiment of the present disclosure, a device is provided for transmitting video / video data including a bitstream generated according to a video / video encoding method according to at least one of the embodiments of the present disclosure.
[0014] According to one embodiment of the present disclosure, a computer-readable storage medium storing a program for performing a method according to at least one of the embodiments of the present disclosure may be provided.
[0015] According to one embodiment of the present disclosure, a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method according to at least one of the embodiments of the present disclosure is provided.
[0016] According to one embodiment of the present disclosure, there is provided a computer-readable digital storage medium storing encoded information or encoded video / image information that causes a decoding device to perform a video / image decoding method according to at least one of the embodiments of the present disclosure.
[0017] According to one embodiment of the present disclosure, the overall video / image compression efficiency can be improved.
[0018] According to one embodiment of the present disclosure, the transformation performance for the current block can be improved.
[0019] According to one embodiment of the present disclosure, transformation-related information can be efficiently signaled.
[0020] According to one embodiment of the present disclosure, a set of transformation kernels can be determined based on the size of a current block or a prediction mode to perform transformation efficiently.
[0021] According to one embodiment of the present disclosure, a transformation can be performed efficiently by determining a transformation kernel by signaling information representing a set of transformation kernels including a vertical transformation kernel and a horizontal transformation kernel.
[0022] FIG. 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure may be applied.
[0023] FIG. 2 is a drawing schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure can be applied.
[0024] FIG. 3 is a drawing schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied.
[0025] Figure 4 illustrates an example of an inter prediction procedure.
[0026] Figure 5 shows examples of inter prediction based video / image encoding methods.
[0027] Figure 6 shows examples of video / image decoding methods based on inter prediction.
[0028] Figure 7 shows examples of residual processing-based video / image encoding methods.
[0029] Figure 8 shows examples of residual processing-based video / image decoding methods.
[0030] FIG. 9 illustrates an example of a multiple transformation technique according to the present disclosure.
[0031] Figure 10 is a diagram illustrating an example of LFNST.
[0032] Figure 11 shows an example of an ROI for LFNST 16.
[0033] Figure 12 shows an example of ROI for LFNST 8.
[0034] Figure 13 illustrates an example of a MIP prediction sample for HoG construction.
[0035] Figure 14 shows an example of deriving an MTS set.
[0036] Figure 15 shows an example of how DIMD-based intra modes for MTS and LFNST are derived.
[0037] Figure 16 illustrates the IntraTMP mode as an example.
[0038] FIG. 17 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure.
[0039] FIG. 18 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure.
[0040] This disclosure may have various modifications and embodiments, and thus specific embodiments will be illustrated and described in detail in the drawings. However, this is not intended to limit the embodiments of the present disclosure to the specific embodiments. The terminology used herein is only used to describe specific embodiments and is not intended to limit the technical spirit of the present disclosure. The singular forms used herein are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term "and / or" as used herein includes any one or a combination of two or more of the associated listed items. The terms "comprises," "comprises," and "contains" as used herein specify the presence of stated features, numbers, operations, elements, components, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, elements, components, and / or combinations thereof. The use of the term "can" in connection with an example or embodiment (e.g., what the example or embodiment can include or implement) in this disclosure means that there is at least one example or embodiment that includes or implements such feature, but not all examples are limited thereto and such feature or configuration may be omitted.
[0041] Meanwhile, each component in the drawings described in this disclosure is depicted independently for the convenience of explaining different characteristic functions. This does not imply that each component is implemented with separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of the present disclosure, as long as they do not deviate from the essence of the present disclosure.
[0042] In this disclosure, "A or B" can mean "only A," "only B," or "both A and B." In other words, "A or B" in this disclosure can be interpreted as "A and / or B." For example, "A, B or C" in this disclosure can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0043] As used herein, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."
[0044] In the present disclosure, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Additionally, in the present disclosure, the expressions “at least one of A or B” or “at least one of A and / or B” may be interpreted identically to “at least one of A and B.”
[0045] Additionally, in the present disclosure, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0046] Additionally, parentheses used in the present disclosure may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in the present disclosure is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be suggested as an example of "prediction."
[0047] Technical features individually described in one drawing in this disclosure may be implemented individually or simultaneously.
[0048] The present disclosure relates to video / image coding. For example, the methods / embodiments described in the present disclosure can be applied to methods disclosed in the enhanced compression model (ECM) or H.267 standards. In addition, the methods / embodiments disclosed in the present disclosure can be applied to methods disclosed in the AV2 (AOMedia Video 2) standard or next-generation video / image coding standards (e.g., H.268, H.269, etc.).
[0049] In the present disclosure, coding may include encoding and / or decoding. In the present disclosure, image coding may be used interchangeably with video coding.
[0050] In the present disclosure, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). A picture may be composed of one or more slices / tiles. A tile may represent a rectangular area of CTUs within a specific tile row and a specific tile column within a picture.
[0051] Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0052] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can also represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0053] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0054] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, identical reference numerals may be used for identical components in the drawings, and redundant descriptions of identical components may be omitted.
[0055] FIG. 1 schematically illustrates an example of a video / image coding system to which embodiments of the present disclosure may be applied.
[0056] Referring to FIG. 1, a video / image coding system may include a first device (encoding device) and a second device (decoding device). The first device may transmit encoded video / image information or data to the second device via a digital storage medium or a network in the form of a file or streaming.
[0057] The video / image coding system may further include a video / image acquisition device and a video / image renderer. The video / image acquisition device may be included in the encoding device, or may be configured as a separate device or external component. The video / image renderer may be included in the decoding device, or may be configured as a separate device or external component.
[0058] The first device may include the transmission unit as an internal component, or as a separate device or external component.
[0059] The second device may include the receiver as an internal component, or as a separate device or external component.
[0060] An encoder may be referred to as an encoding device, and a decoder may be referred to as a decoding device. A transmitting unit may be included in an encoding device. A receiving unit may be included in a decoding device. A renderer may include a display unit, and the display unit may be comprised of a separate device or an external component.
[0061] The decoding device and encoding device to which the embodiment(s) of the present disclosure are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (augmented reality) device, a video phone video device, a transportation terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0062] A video / image capture device can capture a video / image source. The video / image capture device can capture the video / image through a process of capturing, synthesizing, or generating the video / image. The video / image capture device can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a camcorder, a computer, a tablet, a smartphone, etc., and can (electronically) generate the video / image. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process of generating related data. The video / image source can also perform a video / image preprocessing process to input optimized video / image to an encoder.
[0063] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0064] The transmission unit can transmit encoded video / image information or data output in bitstream form to the reception unit of the receiving device through the network in the form of a file or streaming. The encoded video / image information or data output in bitstream form can also be transmitted to the reception unit through a streaming server. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file through a predetermined file format and an element for transmission through a broadcasting / communication network. The reception unit can receive / extract the bitstream and transmit it to a decoding device.
[0065] The streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream. The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium that informs the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server serves to control commands / responses between each device within the content streaming system.
[0066] The streaming server can receive content from a media storage device and / or an encoding device. For example, when receiving content from the encoding device, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0067] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0068] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0069] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied. The term "encoding device" hereinafter may include an image encoding device and / or a video encoding device.
[0070] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit and an intra prediction unit. The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processor (230) may further include a subtractor (subtractor) 231. The addition unit (250) may be called a reconstruction unit or a reconstructed block generator. The image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0071] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing units may be referred to as coding units (CUs). In this case, the coding units may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above.The above prediction unit may be a unit for sample prediction, and the above transformation unit may be a unit for deriving a transformation coefficient and / or a unit for deriving a residual signal from a transformation coefficient.
[0072] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0073] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from a prediction unit from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, a unit that subtracts a prediction signal (predicted block, prediction sample array) from an input video signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0074] An intra prediction unit can predict a current block by referring to samples within a current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block depending on the prediction mode. In intra prediction, prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, a DC mode and a planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail in the prediction direction. However, this is only an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0075] An inter prediction unit can derive a predicted block for a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in an inter prediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on an inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. A reference picture including the reference block and a reference picture including the temporal neighboring blocks may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and a reference picture including the temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0076] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for screen content coding (SCC), for example. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block based on a block vector within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure.
[0077] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal. The transformation unit (232) can apply a transformation technique to the residual signal to generate transform coefficients. For example, the transformation technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loe've Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT).
[0078] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (240) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) may encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The above bitstream may be transmitted through a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0079] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit. When there is no residual for the target block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next target block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0080] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0081] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0082] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit. Through this, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device when inter prediction is applied, and can also improve encoding efficiency.
[0083] The memory (270) DPB can store the modified reconstructed picture to be used as a reference picture in the inter prediction unit. The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store reconstructed samples of reconstructed blocks within the current picture and transfer them to the intra prediction unit.
[0084] FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied. The term "decoding device" hereinafter may include an image decoding device and / or a video decoding device.
[0085] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit and an intra-prediction unit. The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., decoder chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0086] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.
[0087] The decoding device (300) can receive a signal output from the encoding device in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of a syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of a bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (330), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present disclosure may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), and the prediction unit (330).
[0088] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0089] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0090] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the information regarding the prediction output from the entropy decoding unit (310), the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra / inter-prediction mode.
[0091] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for screen content coding (SCC), for example. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block based on a block vector within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure.
[0092] The intra prediction unit can predict the current block by referencing samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0093] The inter prediction unit can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the mode of inter prediction for the current block.
[0094] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (330). In cases where there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.
[0095] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture.
[0096] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0097] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0098] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit. The memory (360) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transmitted to the inter prediction unit to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks within the current picture and transmit them to the intra prediction unit.
[0099] In this specification, the embodiments described in the filtering unit (260) and the prediction unit (220) of the encoding device (200) can be applied to the filtering unit (350) and the prediction unit (330) of the decoding device (300) in the same or corresponding manner, respectively.
[0100] As described above, prediction is performed to increase compression efficiency when performing video coding. Through this, a predicted block including prediction samples for a current block, which is a coding target block, can be generated. Here, the predicted block includes prediction samples in a spatial domain (or pixel domain). The predicted block is derived identically from an encoding device and a decoding device, and the encoding device can increase video coding efficiency by signaling information (residual information) about the residual between the original block and the predicted block, rather than the original sample value of the original block itself, to a decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and generate a reconstructed block including reconstructed samples by combining the residual block and the predicted block, and can generate a reconstructed picture including the reconstructed blocks.
[0101] The residual information may be generated through a transformation and quantization procedure. For example, the encoding device may derive a residual block between the original block and the predicted block, perform a transformation procedure on residual samples (a residual sample array) included in the residual block to derive transform coefficients, and perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, thereby signaling the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, a transformation technique, a transformation kernel, and quantization parameters of the quantized transform coefficients. The decoding device may perform an inverse quantization / inverse transformation procedure based on the residual information to derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also inversely quantize / inversely transform the quantized transform coefficients to derive a residual block for reference in inter prediction of a subsequent picture, and generate a restored picture based on the residual block.
[0102] In the present disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. If the quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. If the transformation / inverse transformation is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.
[0103] In addition, in the present disclosure, the quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, the residual information may include information about the transform coefficient(s), and the information about the transform coefficient(s) may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficient(s)), and scaled transform coefficients may be derived through inverse transformation (scaling) of the transform coefficients. Residual samples may be derived based on inverse transformation (transformation) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of the present disclosure.
[0104] Intra prediction may refer to a prediction that generates prediction samples for a current block based on reference samples within a picture to which the current block belongs (hereinafter, referred to as the current picture). When intra prediction is applied to a current block, peripheral reference samples to be used for intra prediction of the current block may be derived. The peripheral reference samples of the current block may include H+W samples located to the left of a current block of size WХH, W+H samples located at the top of the current block, and at least one sample neighboring the top-left of the current block. Alternatively, the peripheral reference samples of the current block may include upper peripheral samples of multiple rows and left peripheral samples of multiple columns.
[0105] Some of the surrounding reference samples of the current block may not yet be decoded or available. In this case, the decoder can pad or replace the unavailable samples with available samples to form the surrounding reference samples used for prediction.
[0106] When peripheral reference samples are derived, prediction samples of the current block can be derived based on the peripheral reference samples and intra prediction mode / type information. Here, the intra prediction mode can indicate one of non-directional prediction modes and directional prediction modes that indicate spatial correlation for intra prediction. Here, the directional prediction mode can be called an angular prediction mode, and the non-directional prediction mode can be called a non-angular prediction mode. The intra prediction type can indicate various prediction types for performing intra prediction. Intra prediction types may include, for example, multi-reference line (MRL), intra sub-partitions (ISP), Position dependent intra prediction (PDPC), matrix weighted intra prediction (MIP) or matrix based intra prediction, cross-component linear model (CCLM), multi-model linear model (MMLM), Decoder side intra mode derivation (DIMD), fusion of chroma intra prediction modes, intra template matching, fusion for template-based intra mode derivation (TIMD), intra prediction fusion, cross-component convolutional model (CCCM), cross-component prediction (CCP), spatial geometric partitioning mode (SGPM), etc. In some cases, intra prediction modes and / or intra prediction types may be used to perform intra prediction.
[0107] Specifically, the intra prediction procedure may include an intra prediction mode / type determination step, a reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. Additionally, a post-processing filtering step may be performed on the derived prediction samples, if necessary.
[0108] Below, we explain inter prediction.
[0109] The prediction unit of the encoding device / decoding device can perform inter prediction on a block-by-block basis to derive prediction samples. Inter prediction can refer to a prediction derived in a manner dependent on data elements (e.g., sample values, or motion information, etc.) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis based on the correlation of the motion information between the surrounding blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, a motion information candidate list may be constructed based on the neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.
[0110] The above motion information may include L0 motion information and / or L1 motion information depending on the inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be called an L0 prediction, prediction based on an L1 motion vector may be called an L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be called a bi-prediction (Bi). Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures preceding the current picture in output order as reference pictures, and the reference picture list L1 may include pictures succeeding the current picture in output order. The preceding pictures may be called forward (reference) pictures, and the succeeding pictures may be called backward (reference) pictures. The reference picture list L0 may further include pictures succeeding the current picture in output order as reference pictures. In this case, the preceding pictures may be indexed first and the succeeding pictures may be indexed next within the reference picture list L0. The reference picture list L1 may further include pictures preceding the current picture in output order as reference pictures. In this case, the succeeding pictures may be indexed first and the succeeding pictures may be indexed next within the reference picture list 1. Here, the output order may correspond to a POC (picture order count) order.
[0111] Figure 4 illustrates an example of an inter prediction procedure.
[0112] Referring to FIG. 4, as described above, the inter prediction mode / type determination step, the motion vector derivation / refinement step, and the inter prediction performance (prediction sample generation) step may be included. The inter prediction procedure may be performed in the encoding device and the decoding device as described above. In the present disclosure, the coding device may include the encoding device and / or the decoding device.
[0113] The coding device determines the inter prediction mode / type (S400). The coding device may include an encoding device and / or a decoding device as described above.
[0114] An encoding device can determine an inter-prediction mode / type applied to the current block from among various inter-prediction modes / types described in the present disclosure, and can generate prediction-related information. The prediction-related information can include inter-prediction mode information indicating an inter-prediction mode applied to the current block and / or inter-prediction type information indicating an inter-prediction type applied to the current block. A decoding device can determine an inter-prediction mode / type applied to the current block based on the prediction-related information.
[0115] The coding device derives / refines the motion vector of the current block (S410). The coding device can derive / refine the motion vector of the current block based on the determined inter prediction mode / type. Here, motion information of surrounding blocks of the current block can be used to derive / refine the motion vector.
[0116] For example, when skip mode or merge mode is applied to the current block, the coding device may configure a merge candidate list and select one of the merge candidates included in the merge candidate list. Information indicating the selected merge candidate (e.g., merge index) may be included in the prediction-related information.
[0117] As another example, when the (A)MVP mode is applied to the current block, the decoding device may construct a list of (A)MVP candidates, and use the motion vector of an MVP candidate selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be indicated based on selection information (mvp flag or mvp index). In this case, information about MVD as well as the selection information may be included in the prediction-related information.
[0118] Meanwhile, as described below, the motion information of the current block can be derived without constructing a candidate list, in which case the motion information of the current block can be derived according to the procedure disclosed in the prediction mode / type described below. In this case, the candidate list construction described above can be omitted.
[0119] The coding device performs inter prediction on the current block based on the derived / refined motion vector (S420). For example, the coding device can predict the current block (generate a prediction sample) based on the derived / refined motion vector. The coding device can derive the prediction sample of the current block using samples of the reference block indicated by the motion vector in the reference picture.
[0120] An encoding procedure based on inter prediction may roughly include, for example:
[0121] Figure 5 shows examples of inter prediction based video / image encoding methods.
[0122] Referring to FIG. 5, S500 may be performed by a prediction unit of an encoding device, S505 may be performed by a residual processing unit of the encoding device, and S510 or S515 may be performed by an entropy encoding unit of the encoding device. Specifically, the prediction-related information may be derived by the prediction unit and encoded by the entropy encoding unit. The residual information may be derived by the residual processing unit and encoded by the entropy encoding unit. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients through a transform unit of the encoding device, and the transform coefficients may be derived as quantized transform coefficients through a quantization unit. The information about the quantized transform coefficients may be encoded in the entropy encoding unit through a residual coding procedure.
[0123] An encoding device performs inter prediction on a current block (S500). The encoding device can derive the inter prediction mode / type and motion information of the current block, and generate prediction samples of the current block. Here, the inter prediction mode / type determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the other. For example, the inter prediction unit of the encoding device can search for a block similar to the current block within a certain area (search area) of reference pictures through motion estimation, and derive a reference block whose difference from the current block is minimal or below a certain standard. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine a mode to be applied to the current block among various prediction modes. The encoding device can compare RD costs for the various prediction modes and determine an optimal prediction mode for the current block.
[0124] For example, when the skip mode or merge mode is applied to the current block, the encoding device may configure a merge candidate list described below, and derive a reference block among the reference blocks indicated by the merge candidates included in the merge candidate list, the difference between the current block and the current block being at least or below a certain standard. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to a decoding device. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0125] As another example, when the (A)MVP mode is applied to the current block, the encoding device may configure an (A)MVP candidate list described below, and use the motion vector of an mvp candidate selected from among mvp (motion vector predictor) candidates included in the (A)MVP candidate list as the mvp of the current block. In this case, for example, a motion vector indicating a reference block derived by the above-described motion estimation may be used as the motion vector of the current block, and an mvp candidate having a motion vector with the smallest difference from the motion vector of the current block among the mvp candidates may become the selected mvp candidate. A motion vector difference (MVD), which is a difference obtained by subtracting the mvp from the motion vector of the current block, may be derived. In this case, information about the MVD may be signaled to the decoding device. In addition, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and signaled separately to the decoding device.
[0126] The encoding device may perform residual processing based on the predicted samples (S505). The encoding device may derive residual samples based on the predicted samples. The encoding device may derive the residual samples by comparing the original samples of the current block with the predicted samples. Residual information may be generated based on the residual samples. The residual information may include information regarding quantized transform coefficients as described above.
[0127] An encoding device encodes image information including prediction-related information and / or residual information (S510 or S515). The encoding device can output the encoded image information in the form of a bitstream. The prediction-related information may include information related to the prediction procedure, such as prediction mode information (e.g., skip flag, merge flag, or mode index) and information about motion information. The information about the motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index), which is information for deriving a motion vector. In addition, the information about the motion information may include information about the above-described MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples.
[0128] The output bitstream can be stored on a (digital) storage medium and transmitted to a decoding device, or can be transmitted to a decoding device via a network.
[0129] Meanwhile, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as that performed by the decoding device from the encoding device, thereby improving coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.
[0130] The decoding device can perform operations corresponding to those performed by the encoding device. A video / image decoding procedure based on intra prediction may include, for example, the following.
[0131] Figure 6 shows examples of video / image decoding methods based on inter prediction.
[0132] Referring to FIG. 6, S600 may be performed by an entropy decoding unit of a decoding device, S610 may be performed by a prediction unit of the decoding device, S615 may be performed by a residual processing unit of the decoding device, and S620 may be performed by an adder or restoration unit of the decoding device.
[0133] Specifically, the decoding device obtains image / video information from the bitstream (S600). The image / video information may include prediction-related information and / or residual information.
[0134] The decoding device performs inter prediction based on prediction-related information (S610). The decoding device may derive an inter prediction mode / type for the current block based on the prediction-related information, derive / refine motion information of the current block, and generate prediction samples within the current block based on the intra prediction mode / type and / or the motion information. In this case, the decoding device may perform a prediction sample filtering procedure. The prediction sample filtering procedure may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted.
[0135] The decoding device performs residual processing based on the residual information (S615). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit of the residual processing unit performs inverse transformation on the transform coefficients to derive residual samples for the current block.
[0136] The decoding device generates a reconstructed block / picture (S620). The decoding device can generate reconstructed samples for the current block based on the prediction samples and / or the residual samples, and derive a reconstructed block including the reconstructed samples. A reconstructed picture for the current picture can be generated based on the reconstructed block. As described above, an in-loop filtering procedure, etc., can be further applied to the reconstructed picture.
[0137] The above prediction-related information can be encoded / decoded using the binarization and coding methods described in the present disclosure. For example, the prediction-related information can be binarized using fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, the prediction-related information can be encoded / decoded using entropy coding (e.g., CABAC, CAVLC) coding.
[0138] Below, residual processing is described. The residual processing can be performed in an encoding device and a decoding device. The residual processing can include coefficient coding, transformation, and / or quantization procedures.
[0139] The residual processing procedure at the encoding stage may include a procedure for generating and / or encoding residual information from residual samples for the derived current block. The residual processing procedure may further include a procedure for deriving the residual samples based on prediction samples. The residual processing procedure at the decoding stage may include a procedure for deriving residual samples from residual information of a received bitstream. For example, the residual processing procedure may include an (inverse) transformation and / or an (inverse) quantization procedure. In addition, the residual processing procedure may include an encoding / decoding procedure of the residual information. The residual information may include residual data and / or transformation / quantization related parameters.
[0140] Specifically, for residual processing, a method is provided for deriving (quantized) transform coefficients within a block and generating and (encoding) residual information based on the derived (quantized) residual coefficients within a block to which transform skip is applied and generating and (encoding) residual information for transform skip based on the derived (quantized) residual coefficients. The encoded information can be output in the form of a bitstream as described above.
[0141] Additionally, in the case of the decoding stage, (quantized) transform coefficients or (quantized) residual coefficients within a block can be derived from residual information (or residual information for transform skip) included in the bitstream, and (if necessary) inverse quantization / inverse transformation can be performed to derive residual samples.
[0142] A residual processing-based encoding procedure may roughly include, for example:
[0143] Figure 7 shows examples of residual processing-based video / image encoding methods.
[0144] Referring to FIG. 7, S700 may be performed by a prediction unit of an encoding device, S710 may be performed by a residual processing unit of the encoding device, and S720 may be performed by an entropy encoding unit of the encoding device. Specifically, the prediction-related information may be derived by the prediction unit and encoded by the entropy encoding unit. The residual information may be derived by the residual processing unit and encoded by the entropy encoding unit. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples. As described above, the residual samples may be derived as transform coefficients through a transform unit of the encoding device, and the transform coefficients may be derived as quantized transform coefficients through a quantization unit. The information about the quantized transform coefficients may be encoded in the entropy encoding unit through a residual coding procedure.
[0145] The encoding device derives prediction samples of the current block (S700). The encoding device can derive the prediction samples of the current block based on the inter prediction and / or intra prediction described above.
[0146] The encoding device can perform residual processing based on the prediction samples (S710). The encoding device can derive residual samples based on the prediction samples. The encoding device can derive the residual samples by comparing the original samples of the current block with the prediction samples. The residual processing includes a transformation and / or quantization process for the residual sample, as described above. The encoding device can generate residual information from the residual samples through the residual processing. The residual information can include information about quantized transform coefficients, as described above.
[0147] An encoding device encodes image information including prediction-related information and / or residual information (S720). The encoding device can output the encoded image information in the form of a bitstream. The prediction-related information may include information related to the prediction procedure. The residual information is information about the residual samples. The residual information may include information about quantized transform coefficients for the residual samples.
[0148] The output bitstream can be stored on a (digital) storage medium and transmitted to a decoding device, or can be transmitted to a decoding device via a network.
[0149] Meanwhile, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as that performed by the decoding device from the encoding device, thereby improving coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.
[0150] The decoding device can perform operations corresponding to those performed by the encoding device. A video / image decoding procedure based on residual processing may include, for example, the following.
[0151] Figure 8 shows examples of residual processing-based video / image decoding methods.
[0152] Referring to FIG. 8, S800 may be performed by an entropy decoding unit of a decoding device, S810 may be performed by a prediction unit of the decoding device, S820 may be performed by a residual processing unit of the decoding device, and S830 may be performed by an adder or restoration unit of the decoding device.
[0153] Specifically, the decoding device obtains image / video information from the bitstream (S800). The image / video information may include prediction-related information and / or residual information.
[0154] The decoding device performs prediction (including inter-prediction and / or intra-prediction) based on prediction-related information (S810). The decoding device may derive a prediction mode / type for the current block based on the prediction-related information, and generate prediction samples within the current block based on the prediction mode / type. In this case, the decoding device may perform a prediction sample filtering procedure. The prediction sample filtering procedure may be referred to as post-filtering. Some or all of the prediction samples may be filtered by the prediction sample filtering procedure. In some cases, the prediction sample filtering procedure may be omitted.
[0155] The decoding device performs residual processing based on residual information (S820). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the inverse quantization unit of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit of the residual processing unit performs inverse transformation on the transform coefficients to derive residual samples for the current block.
[0156] The decoding device generates a reconstructed block / picture (S830). The decoding device can generate reconstructed samples for the current block based on the prediction samples and / or the residual samples, and derive a reconstructed block including the reconstructed samples. A reconstructed picture for the current picture can be generated based on the reconstructed block. As described above, an in-loop filtering procedure, etc., can be further applied to the reconstructed picture.
[0157] The residual information can be encoded / decoded through the binarization and coding methods described in the present disclosure. For example, the residual information can be binarized through fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, the residual information can be encoded / decoded through entropy coding (e.g., CABAC, CAVLC) or bypass coding.
[0158] Below we describe the maximum transformation size and zeroing out of the transformation coefficients.
[0159] For example, the CTU size and maximum transform size (all MTS kernels) can be extended up to 256. In this case, the maximum size of the intra prediction block can be set to 128X128. In UHD sequences, the maximum CTU size can be set to 256, and in other cases, it can be set to 128.
[0160] In the first transformation process, zeroing out of the transformation coefficients may not be applied. However, in the second transformation process where LFNST is applied, the first transformation coefficients outside the ROI area where LFNST is applied may be zeroed out.
[0161] Meanwhile, zeroing out of the transform coefficients may be applied during the first transformation process. As the size of the transform block increases, zeroing out may be required after the first transformation, and whether or not to zero out may be determined depending on whether MTS is applied. For example, when MTS is applied, zeroing out may be performed to any of 16, 20, 24, 32, or 40, and when MTS is not applied, zeroing out may be performed to a specific preset value. The specific value may be 32 or 64. When MTS is applied, the number of transform coefficients to be zeroed out may vary depending on the size of the transform block. For example, if the width or height to which MTS is applied in the transform block is 4, 8, or 16 (if it is a 4-point transform, 8-point transform, or 16-point transform), zeroing out may not be applied, and if the width or height to which MTS is applied in the transform block is 32 (if it is a 32-point transform), zeroing out may be performed to 16. Or, if MTS is applied to a width or height greater than 32, it may be zeroed out to 32.
[0162] Meanwhile, residual processing includes transformation / inverse transformation and / or quantization / inverse quantization processes. According to the present disclosure, a first transformation and / or a second transformation may be applied to a residual block to derive a transform coefficient block (transform coefficients), and an inverse second transformation and / or an inverse first transformation may be applied to the transform coefficient block (transform coefficients) to derive a residual block.
[0163] As described above, the transformation on the residual can be performed through a primary transformation and / or a secondary transformation that is optionally performed after the primary transformation. The primary transformation can be called a primary transformation and can be a Discrete Cosine Transform (DCT) and a Discrete Sine Transform (DST) that are applied to all rows and all columns of the residual block. After the primary transformation, a secondary transformation can be additionally applied to a specific transform coefficient at the upper left of the transform block according to the result of the primary transformation. The inverse transformation performed at the decoding stage can be performed by applying an inverse secondary transformation to a specific residual at the upper left of the residual block corresponding to the (inverse quantized) residual information, and applying an inverse primary transformation to the transform block according to the result of the inverse secondary transformation.
[0164] Below, we describe MTS (Multiple Transform Selection), a method of transformation.
[0165] FIG. 9 illustrates an example of a multiple transformation technique according to the present disclosure.
[0166] Referring to FIG. 9, the conversion unit may correspond to the conversion unit in the encoding device of FIG. 2 described above, and the inverse conversion unit may correspond to the inverse conversion unit in the encoding device of FIG. 2 described above or the inverse conversion unit in the decoding device of FIG. 3.
[0167] The transformation unit can derive (primary) transformation coefficients by performing a primary transformation based on residual samples (residual sample array) within the residual block (S900). This primary transformation may be referred to as a core transformation. Here, the primary transformation may be based on multiple transform selection (MTS), and when multiple transformations are applied as the primary transformation, it may be referred to as a multiple core transformation.
[0168] The transformation unit can perform a secondary transformation based on the (primary) transformation coefficients to derive modified (secondary) transformation coefficients (S910). The primary transformation is a transformation from the spatial domain to the frequency domain, and the secondary transformation can be expressed as a transformation into a more compact representation by utilizing the correlation existing between the (primary) transformation coefficients.
[0169] For example, the secondary transform may include a non-separable transform. In this case, the secondary transform may be called a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may represent a transform that generates modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix. Here, the vertical transform and the horizontal transform may be applied at once to the (primary) transform coefficients based on the non-separable transform matrix, without separately applying the vertical transform and the horizontal transform (or independently applying the horizontal transform and the vertical transform).
[0170] That is, the non-separable secondary transform may refer to a transform method that does not separate the vertical and horizontal components of the (primary) transform coefficients, but rather rearranges, for example, two-dimensional signals (transform coefficients) into one-dimensional signals along a specific predetermined direction, and then generates modified transform coefficients (or secondary transform coefficients) based on the non-separable transform matrix. In other words, the transform method may refer to a transform method that rearranges into one-dimensional signals along a row-first direction or a column-first direction, and then generates modified transform coefficients (or secondary transform coefficients) based on the non-separable transform matrix.
[0171] In addition, the inverse transform unit can perform a series of procedures in the reverse order of the procedures performed in the above-described transform unit. The inverse transform unit can receive (inverse quantized) transform coefficients, perform a secondary (inverse) transform to derive (primary) transform coefficients (S920), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain a residual block (residual samples) (S930). Here, the primary transform coefficients may be referred to as modified transform coefficients on the inverse transform unit side. As described above, the encoding device and / or the decoding device can generate a reconstructed block based on the residual block and the predicted block, and can generate a reconstructed picture based on the same.
[0172] In the present disclosure, the primary transform may be referred to as a core transform. Here, the primary transform may be based on multiple transform selection (MTS), and when a transform kernel selected from among multiple transform kernel types is applied as the primary transform, it may be referred to as a multi-core transform.
[0173] Multi-core transform can indicate a method of transforming using Discrete Cosine Transform (DCT) 2 and Discrete Sine Transform (DST) 7, DCT 8, etc. In other words, multi-core transform can indicate a transform method of transforming a residual signal (or residual block) of a spatial domain into transform coefficients (or first transform coefficients) of a frequency domain based on a plurality of transform kernels selected from among the above DCT 2, DST 7, DCT 8, and DST 1. Here, the first transform coefficients can be called temporary transform coefficients from the transform unit's perspective.
[0174] In other words, when a conventional transform method is applied, a transformation from a spatial domain to a frequency domain can be applied to a residual signal (or a residual block) based on DCT 2, so that transform coefficients can be generated. In contrast, when a multi-core transform is applied, a transformation from a spatial domain to a frequency domain can be applied to a residual signal (or a residual block) based on DCT 2, DST 7, DCT 8, and / or DST 1, so that transform coefficients (or first-order transform coefficients) can be generated. Here, DCT 2, DST 7, DCT 8, and DST 1, etc. may be called a transform type, a transform kernel, or a transform core. These DCT / DST transform types can be defined based on basis functions.
[0175] When a multi-core transform is performed, a vertical transform kernel and a horizontal transform kernel for a target block (a block to be transformed) may be selected from among transform kernels, and a vertical transform may be performed for the target block based on the vertical transform kernel, and a horizontal transform may be performed for the target block based on the horizontal transform kernel. Here, the horizontal transform may represent a transform for horizontal components of the target block, and the vertical transform may represent a transform for vertical components of the target block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on a prediction mode and / or a transform index of a target block (CU or sub-block) including a residual block.
[0176] In addition, according to an example, when performing a first transformation by applying MTS, specific basis functions can be set to predetermined values, and a mapping relationship for the transformation kernel can be set by combining which basis functions are applied when performing a vertical transformation or a horizontal transformation. For example, when a horizontal transformation kernel is represented as trTypeHor and a vertical transformation kernel is represented as trTypeVer, a trTypeHor or trTypeVer value of 0 can be set to DCT2, a trTypeHor or trTypeVer value of 1 can be set to DST7, and a trTypeHor or trTypeVer value of 2 can be set to DCT8.
[0177] In this case, MTS index information may be encoded and signaled to the decoding device to indicate which of a plurality of sets of transform kernels. For example, an MTS index of 0 may indicate that both trTypeHor and trTypeVer values are 0, an MTS index of 1 may indicate that both trTypeHor and trTypeVer values are 1, an MTS index of 2 may indicate that trTypeHor values are 2 and trTypeVer values are 1, an MTS index of 3 may indicate that trTypeHor values are 1 and trTypeVer values are 2, and an MTS index of 4 may indicate that both trTypeHor and trTypeVer values are 2.
[0178] These MTSs can be explicitly applied by signaling as described above, or implicitly applied according to specific conditions. Implicit MTSs can independently derive transform kernels for each direction depending on the width or height of the transform block. For example, when a sub-block transform is applied that applies the transform to only one of the sub-blocks, an implicit MTS can be applied if the width or height satisfies a specific condition, or if the conditions for a specific intra mode are satisfied. For example, an implicit MTS can be applied to a block to which SBT is applied and the larger of the width or height is 32 or 64 or less. Alternatively, an implicit MTS can be applied to a block to which ISP is applied, or to a block to which intra prediction is applied but LFNST and MIP are not applied.
[0179] Additionally, implicit MTS can infer primary transform pairs based on the intra prediction mode and TU size of the current block using a lookup table (LUT). For example, the same intra prediction mode as explicit MTC can be considered, and the maximum TU size can be 32x32. In this case, the LUT can include transform pairs based on separable primary transforms, and no additional primary transforms can be added. For example, DCT 2, DCT 5, DCT 8, DST 1, DST 4, and DST 7 can be used.
[0180] For example, for an ISP block, the sizeIdx used as input to the LUT can be based on the location of the current ISP subpartition within the ISP block. Furthermore, as an example, the implicit MTS can only be applied to the ISP block depending on the CTC configuration settings. As another example, the implicit MTS can be applied to angular intra prediction modes. For example, the implicit MTS can be used for the ISP block by using implicit signaling for all modes (non-CTC). For example, the default implicit MTS (i.e., based on the shape of the TU) can be maintained in TIMD and DIMD modes. Furthermore, DCT 2 can be used in MIP, EIP, SGPM, and IntraTMP modes. Furthermore, the default implicit MTS can be used for the ISP block.
[0181] Below, we describe LFNST (Low Frequency Non Separable Transform), which is one of the transformation methods.
[0182] Meanwhile, the transform unit of the encoding device can perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients. Here, the primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform means a transform to a more compact representation by utilizing the correlation that exists between the (primary) transform coefficients. The secondary transform may include a non-separable transform. In this case, the secondary transform may be called a non-separable secondary transform (NSST). The non-separable secondary transform may represent a transform that generates modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix. Here, based on the non-separable transform matrix, the vertical transform and the horizontal transform can be applied at once to the (primary) transform coefficients without separately applying them (or independently applying the horizontal-vertical transform). In other words, the non-separable secondary transform can refer to a transform method of rearranging two-dimensional signals (transform coefficients) into one-dimensional signals in a specific predetermined direction (e.g., row-first direction or column-first direction) without separating the vertical and horizontal components of the (primary) transform coefficients, and then generating modified transform coefficients (or secondary transform coefficients) based on the non-separable transform matrix. For example, the row-first order is to arrange them in a row in the order of the 1st row, the 2nd row, ..., the Nth row for an MxN block, and the column-first order is to arrange them in a row in the order of the 1st column, the 2nd column, ..., the Mth column for an MxN block.The above non-separable secondary transform can be applied to the top-left region of a block composed of (primary) transform coefficients (hereinafter, referred to as a transform coefficient block). The non-separable secondary transform can be selected based on a mode, in which a transform kernel (or transform core, transform type) can be selected. Here, the mode can include an intra-prediction mode and / or an inter-prediction mode.
[0183] A non-separable second-order transform can be performed based on an 8X8 transform or a 4X4 transform determined based on the width (W) and height (H) of a transform coefficient block. An 8x8 transform refers to a transform that can be applied to an 8x8 region contained within a transform coefficient block when both W and H are greater than or equal to 8, and the 8x8 region can be the upper left 8x8 region within the transform coefficient block. Similarly, a 4x4 transform refers to a transform that can be applied to a 4x4 region contained within a transform coefficient block when both W and H are greater than or equal to 4, and the 4x4 region can be the upper left 4x4 region within the transform coefficient block. For example, an 8x8 transform kernel matrix can be a 64x64 / 16x64 matrix, and a 4x4 transform kernel matrix can be a 16x16 / 8x16 matrix.
[0184] Meanwhile, for mode-based transformation kernel selection, a transformation set for non-separable secondary transformations can be set, and k non-separable secondary transformation kernels can be configured for each transformation set. For example, the transformation sets can be 4, 35, etc. Selection of a specific set among the transformation sets can be performed based on, for example, the intra prediction mode of the target block (CU or sub-block).
[0185] For example, if it is determined that a specific set is to be used for a non-separable transform, one of the k transform kernels within the specific set can be selected through a non-separable secondary transform index. The encoding device can derive the non-separable secondary transform index pointing to the specific transform kernel based on a rate-distortion (RD) check, and signal the non-separable secondary transform index to the decoding device. The decoding device can select one of the k transform kernels within the specific set based on the non-separable secondary transform index.
[0186] The inverse transform unit of the encoding device and the decoding device can perform a series of procedures in the reverse order of the procedures performed in the above-described transform unit. The inverse transform unit can receive (inverse quantized) transform coefficients and perform a second (inverse) transform to derive (first) transform coefficients. The first transform coefficients can be called modified transform coefficients from the inverse transform unit's perspective.
[0187] In the present disclosure, in order to reduce the amount of computation and memory required for a non-separable secondary transform, a reduced secondary transform (RST) with a reduced transform matrix (kernel) size can be applied in the concept of NSST. RST can be referred to by various terms such as reduced transform, reduced transform, reduced secondary transform, reduction transform, simplified transform, simple transform, etc., and the names by which RST can be referred to are not limited to the listed examples. Alternatively, since RST is mainly performed in a low-frequency region containing non-zero coefficients in a transform block, it can also be referred to as LFNST (Low-Frequency Non-Separable Transform).
[0188] In LFNST, an N-dimensional vector can be mapped to an R-dimensional vector located in another space to determine a reduced transformation matrix, where R is less than N. N can mean the square of the length of one side of the block to which the transformation is applied or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor can mean the R / N value.
[0189] For example, the size of the LFNST matrix is RxN, which is smaller than the size NxN of a typical transformation matrix, and can be defined as in mathematical expression 1 below.
[0190]
[0191] When the LFNST matrix TRxN is multiplied by the residual samples of the target block for transformation, the transformation coefficients for the target block can be derived. When the size of the block to which the transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), LFNST can be expressed as a matrix operation as in the mathematical expression 2 below.
[0192]
[0193] In mathematical expression 2, r1 to r64 may represent residual samples for the target block, and more specifically, may be transform coefficients generated by applying a primary transform. As a result of the operation of mathematical expression 2, transform coefficients ci for the target block may be derived. ci may be derived from c1 to cR. That is, when R = 16, transform coefficients c1 to c16 for the target block may be derived.
[0194] If a regular transformation instead of LFNST were applied and a transformation matrix of size 64x64 (NxN) were multiplied to residual samples of size 64x1 (Nx1), 64 (N) transformation coefficients for the target block would have been derived. However, since LFNST was applied, only 16 (R) transformation coefficients for the target block are derived. Since the total number of transformation coefficients for the target block is reduced from N to R, the amount of data transmitted from the encoding device to the decoding device is reduced, so the transmission efficiency between the encoding device and the decoding device can be increased.
[0195] The size of the inverse LFNST matrix TNxR is NxR, which is smaller than the size NxN of a normal inverse transform matrix, and is in a transpose relationship with the LFNST matrix TRxN shown in mathematical expression 1. Tt may mean the inverse LFNST matrix TRxNT (the superscript T means transpose). When the inverse LFNST matrix TRxNT is multiplied by the transform coefficients for the target block, modified transform coefficients for the target block or residual samples for the target block can be derived. The inverse LFNST matrix TRxNT may also be expressed as (TRxN)TNxR.
[0196] For example, if the size of the block to which the inverse transformation is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the inverse LFNST can be expressed as a matrix operation as in Equation 3 below.
[0197]
[0198] c1 to c16 may represent transform coefficients for the target block. As a result of the operation in mathematical expression 3, rj representing modified transform coefficients for the target block or residual samples for the target block may be derived. rj may be derived from r1 to rN. That is, when N = 16, transform coefficients r1 to r64 for the target block may be derived.
[0199] Figure 10 is a diagram illustrating an example of LFNST.
[0200] Referring to Fig. 10, 4x4 LFNST can be applied to blocks with min (width, height) < 8, and 8x8 LFNST can be applied to blocks with min (width, height) > 4.
[0201] For example, 16 (first-order) transform coefficients can be input for a 4x4 forward LFNST, and 64 (first-order) transform coefficients can be input for an 8x8 forward LFNST. Since 8 or 16 transform coefficients can be derived from these forward LFNSTs, 8 or 16 transform coefficients can be input for the input of the inverse LFNST, respectively. When a 4x4 inverse LFNST is performed, 16 modified transform coefficients can be output from 8 transform coefficients, and when an 8x8 inverse LFNST is performed, 64 modified transform coefficients can be output from 16 transform coefficients.
[0202] Meanwhile, for example, in the transformation of the encoding process, instead of the 16 x 64 transformation kernel matrix for the 64 data constituting the 8 x 8 region, a maximum of 16 x 48 transformation kernel matrix can be applied to only 48 data. Here, "maximum" means that the maximum value of m is 16 for the m x 48 transformation kernel matrix that can generate m coefficients. That is, when LFNST is performed by applying the m x 48 transformation kernel matrix (m ≤ 16) to the 8 x 8 region, 48 data can be input and m coefficients can be generated. When m is 16, 48 data are input and 16 coefficients are generated. That is, when 48 data form a 48 x 1 vector, a 16 x 1 vector can be generated by sequentially multiplying the 16 x 48 matrix and the 48 x 1 vector. At this time, 48 data forming an 8 x 8 area can be appropriately arranged to form a 48 x 1 vector. At this time, if a matrix operation is performed by applying a maximum 16 x 48 transformation kernel matrix, 16 modified transformation coefficients are generated. The 16 modified transformation coefficients can be arranged in the upper left 4 x 4 area according to the scanning order, and the upper right 4 x 4 area and the lower left 4 x 4 area can be filled with 0.
[0203] The inverse transform of the decoding process can use the transposed matrix of the transform kernel matrix described above. That is, when the inverse LFNST is performed as the inverse transform process performed in the decoding device, the input coefficient data to which the inverse LFNST is to be applied is configured as a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the corresponding inverse LFNST matrix from the left can be arranged in a two-dimensional block according to a predetermined arrangement order. The matrix operation in this case can be expressed as (48 x 16 matrix) * (16x1 transform coefficient vector) = (48 x 1 modified transform coefficient vector). Here, since the nx1 vector can be interpreted as having the same meaning as an nx1 matrix, it can also be expressed as an nx1 column vector. * indicates a matrix multiplication operation. When these matrix operations are performed, 48 modified transformation coefficients can be derived, and the 48 modified transformation coefficients can be arranged in the upper left, upper right, and lower left regions excluding the lower right region of the 8x8 region.
[0204] Meanwhile, LFNST can be applied to sub-blocks. For example, when a coding block is divided into sub-blocks, the size of the sub-blocks can be "width / 2 x height / 2" or "width / 4 x height / 4". In this case, the sub-blocks can include sub-blocks located at the corners or third blocks located at the center. For example, LFNST can be applied to sub-blocks located at the corners. In this case, the division mode and the location of non-zero sub-blocks can be signaled similar to the existing SBT.
[0205] Below we describe an extension of LFSNT for transformation.
[0206] For example, the LFNST described above can be extended with a set of transformations and transformation kernels. For example, the LFNST transformation set can be 35, and each transformation set can be composed of three transformation kernels, i.e., transformation candidates. The transformation set (lfnstTrSetIdx) for the intra prediction mode can have a mapping relationship as shown in the table below.
[0207]
[0208] Referring to Table 1 above, if the intra prediction mode (predModeIntra) is less than 0, lfnstTrSetIdx is 2, if the intra prediction mode is 0 to 34, lfnstTrSetIdx is also mapped to 0 to 34, and if the intra prediction mode is 35 to 66, lfnstTrSetIdx can be mapped to (68- predModeIntra). Also, if the intra prediction mode is greater than 66, lfnstTrSetIdx can be mapped to 2. At this time, lfnstTrSetIdx can be called an LFNST set index.
[0209] Meanwhile, the LFNST kernel can include the LFNST 4 kernel, the LFNST 8 kernel, and the LFNST 16 kernel. For example, the LFNST 16 kernel can be applied in addition to the LFNST 4 kernel and the LFNST 8 kernel. For example, the LFNST 4 kernel can be applied to a block that is 4xN / Nx4 (N≥4), the LFNST 8 kernel can be applied to a block that is 8xN / Nx8 (N≥8), and the LFNST16 kernel can be applied to a block that is 16xN / Nx16 (N≥16).
[0210] Meanwhile, Region-Of-Interest (ROI) and zero-out can be applied in LFNST. Here, ROI can refer to a specific region of interest in an image, or it can refer to a sample within a data set identified for a specific purpose. For example, Forward LFNST is applied to a specific region of interest (ROI) in the upper left corner of the target block. Therefore, when LFNST is applied, the first-order transform coefficients existing in areas outside the ROI can be zeroed out.
[0211] Figure 11 shows an example of an ROI for LFNST 16.
[0212] Referring to Fig. 11, the ROI for LFNST 8 can be composed of six 4x4 sub-blocks that are arranged sequentially in the scan direction from the upper left of the target block. Since a total of 96 transform coefficients are input to the Forward LFNST, the dimension of the Forward LFNST matrix can be Rx96. Here, R can be 32, 48, or 64, which are less than 96. For example, R can be 32, and in this case, LFNST 16 is a 32x96 matrix. In addition, for example, when a 32x96 matrix is used for LFNST 16, the primary transform coefficients existing in the area other than the ROI can be zeroed out.
[0213] Figure 12 shows an example of ROI for LFNST 8.
[0214] Referring to Fig. 12, the ROI for LFNST 8 can be composed of four 4x4 sub-blocks located at the upper left of the target block, i.e., an 8x8 area at the upper left of the target block. Since a total of 64 transform coefficients are input to the Forward LFNST, the dimension of the Forward LFNST matrix can be Rx64. Here, R can be 32 or 48, which are smaller than 64. For example, R can be 32, and in this case, LFNST 8 is a 32x64 matrix. When a 32x64 matrix is used for LFNST 8, the primary transform coefficients existing in areas other than the ROI can be zeroed out. In addition, for example, in the case of an 8X8 block, since 64 transform coefficients are input to the Forward LFNST, the entire block is an ROI, so zeroing out may not be applied.
[0215] Figure 13 illustrates an example of a MIP prediction sample for HoG construction.
[0216] Referring to FIG. 13, a HoG can be constructed using DIMD based on MIP-based prediction samples. For example, for a target block predicted with MIP or IntraTMP, DIMD can be used to derive an intra prediction mode of the target block based on the MIP- or IntraTMP-based prediction samples. For example, for a target block to which MIP is applied, DIMD can be applied based on the MIP prediction samples before upsampling. For example, to construct a HoG, horizontal slopes and vertical slopes are calculated for each prediction sample. Then, an LFNST transform set and an LFNST transpose flag (LFNST Transpose flag) can be determined using the intra prediction mode with the largest histogram amplitude value. In addition, the LFNST transpose flag can be signaled after the LFNST index signaling as information indicating whether the LFNST kernel is transposed. Alternatively, for example, the MIP transpose flag (mip_transposed_flag) can be used as the LFNST transpose flag. In this case, signaling of the LFNST prefix flag may be omitted.
[0217] Meanwhile, if the intra prediction mode is IntraTMP, the intra prediction mode derived through DIMD can be applied as the intra prediction mode for determining the LFNST transformation set. In this case, the intra prediction mode with the largest histogram width can be used as the intra prediction mode for determining the LFNST set.
[0218] Alternatively, for example, if the intra prediction mode is one of IBC, SGPM (Spatial Geometric partitioning mode), TIMD (template-based intra mode derivation), and Palette mode, DIMD can be applied based on the prediction samples derived through the mode to derive the intra prediction mode for determining the LFNST transformation set.
[0219] Alternatively, as an example, LFNST or NSPT can be applied to inter-prediction blocks as well as intra-prediction blocks. For example, a set of transformations can be mapped for each inter-prediction mode, and LFNST or NSPT can be performed by applying one of multiple transformation kernels to the mapped set of transformations. Alternatively, as an example, DIMD can be applied based on inter-prediction samples, and horizontal and vertical gradients can be calculated for each prediction sample to build a HoG. Subsequently, a set of LFNST transformations can be derived using the prediction mode with the largest histogram amplitude value.
[0220] Meanwhile, a modified context model that uses the previous five coefficients in the coding order instead of the neighboring 2D coefficients for coding LFNST / NSPT coefficients can be used. That is, when LFNST / NSPT is applied, the five previous transform coefficients in the coding (scan) order can be used instead of the neighboring coefficients at the 2D position for context modeling or deriving context information about the currently parsed transform coefficient (e.g., it can include at least one of sig_coeff_flag, gt1_flag, or gt2_flag).
[0221] Below we describe Enhanced MTS for intra.
[0222] For example, when MTS is applied to a block predicted intra, the transform kernel may be DCT 5, DST 4 in addition to DCT 2, DST 7, and DCT 8. Also, when MTS is applied to a block predicted intra, the transform kernel may be DCT 5, DST 4, DST 1, and identity transform (IDT) in addition to DCT 2, DST 7, and DCT 8.
[0223] Additionally, intra MTS candidates, i.e., MTS sets, can be derived based on TU sizes and intra prediction mode information. For example, 16 TU sizes can be applied in total, and each TU size can be classified into five groups (classes) based on its intra prediction mode. In this case, if five groups are applied to each of the 16 TU sizes, a total of 80 groups can be considered. However, since the transform set can be shared, fewer than 80 groups, for example, 58 groups, can be considered. Furthermore, the number of groups is not limited to 58, and any positive integer N can be considered.
[0224] Additionally, if the intra prediction mode is a directional mode, the symmetry of the TU shape and the intra prediction direction can be considered. For example, mode i with AХB and mode j with BХA (i > 34) can be mapped to the same group (j = (68 - i)), and the transform pairs for the vertical and horizontal kernels can be swapped. For example, a 16x4 block with an intra prediction mode number of 18 and a 4x16 block with an intra prediction mode number of 50 can be mapped to one group. If the intra prediction mode is a wide-angle intra prediction mode, the closest existing directional mode can be used to determine the transform set. For example, to derive a group for determining the transform set, a wide-angle intra prediction mode between -2 and -14 can use mode 2, and a wide-angle intra prediction mode between 67 and 80 can use mode 66.
[0225] Additionally, multiple transformation candidates (transformation pairs) can be applied to each group. For example, one, four, or six transformation sets can be applied. The number of transformation sets to be applied can be determined based on the position of the last transformation coefficient or based on the absolute value of the transformation coefficient. For example, the position of the last transformation coefficient can be compared with two thresholds, and one of one, four, or six transformation sets to be applied can be selected based on the result of the comparison. Alternatively, the sum of the absolute values of the transformation coefficients can be compared with two thresholds, and one of one, four, or six transformation sets can be selected.
[0226] In addition, the number of transformation sets can be determined based on the sum of the absolute values of the transformation coefficients. For example, if the sum of the absolute values of the transformation coefficients is less than th0, one transformation set can be applied. If the sum of the absolute values of the transformation coefficients is greater than th0 and less than or equal to th1, four transformation sets can be applied. If the sum of the absolute values of the transformation coefficients is greater than th1, six transformation sets can be applied (1 candidate: sum <= th0, 4 candidates: th0 < sum <= th1, 6 candidates: sum > th1). Here, sum can represent the sum of the absolute values of the transformation coefficients. In addition, th0 can be set to 6 and th1 can be set to 32.
[0227] Figure 14 shows an example of deriving an MTS set.
[0228] Referring to Fig. 14, four transform sets can be determined based on the TU size and intra prediction mode information.
[0229] For example, if the TU size is 16, the intra prediction modes can be applied from 0 to 34, and the MIP mode can be added, for a total of 36 intra modes. As explained above, the intra prediction mode of the transform block can be mapped to any one of the intra prediction modes 1 to 34 according to the intra prediction mode information. A total of five groups can be mapped for each TU size. For example, intra prediction modes 0 and 1, intra prediction modes 2 to 12, intra prediction modes 13 to 23, intra prediction modes 24 to 34, and the MIP mode can each be mapped to one group.
[0230] Additionally, each transform pair can be applied with five transform kernels, DCT 8, DST 7, DCT 5, DST 4, and DST 1, in pairs, and each transform pair can be indexed with any one of 0 to 24.
[0231] As described above, if four transformation sets are applied to a single group, a transformation set consisting of four transformation pairs can be applied to each of the 80 groups. In this case, some of the 80 groups may share a transformation set, and the 80 groups can be reduced to 58 groups.
[0232] Figure 15 shows an example of how DIMD-based intra modes for MTS and LFNST are derived.
[0233] Referring to Fig. 15, a DIMD-based intra mode can be derived for determining MTS and LFNST sets.
[0234] Meanwhile, if the intra prediction mode is IntraTMP, the intra prediction mode derived through DIMD can be applied to determine the MTS transform set. For example, the intra prediction mode with the largest histogram width can be used as the intra prediction mode for determining the MTS and LFNST sets.
[0235] Meanwhile, as another example, the intra mode for IntraTMP can also be derived from a reference block or a block surrounding the reference block.
[0236] In addition, if the intra prediction mode is one of IBC, SGPM (Spatial Geometric partitioning mode), TIMD (template-based intra mode derivation), and Palette mode, the intra prediction mode for determining the MTS transformation set can be derived by applying DIMD based on the prediction sample derived through the mode.
[0237] Meanwhile, for example, MTS can be applied even when the size of the transform block is greater than 32. Alternatively, considering complexity, intra prediction modes can be grouped into three or two groups instead of five groups for each transform block size. Alternatively, in addition to the transform pairs in Figure 4.3-2, IDT (identity transform) can be applied. For example, one of the vertical or horizontal transforms can be any one of DST 7, DCT 8, DCT 5, DST 4, or DST 1, as shown in Figure 4.3-2, and the other can be the identity transform. In this case, the number of cases for the transform pair can increase.
[0238] Additionally, regarding the signaling of the MTS index (mts_idx), the following is taken:
[0239] For example, mts_enabled_flag may be signaled at a higher level, and mts_flag and mts_idx may be signaled at the CU or residual coding level. In this case, if mts_flag is 1, mts_idx may be signaled to determine the transform set, and if mts_flag is 0, DCT 2 may be applied to the first transform.
[0240] Also, if there are 4 transformation sets, any one transformation pair among the 4 can be selected through mts_idx, and if there are 6 transformation sets, any one transformation pair among the 6 can be selected through mts_idx. In this case, if there is 1 transformation set, the first transformation pair of the 4 transformation sets or the first transformation pair of the 6 transformation sets can be selected without separate signaling. That is, if there is 1 transformation set, mts_flag can be signaled with a value of 1, and mts_idx can not be signaled. Also, if mts_idx is not signaled, a specific transformation pair can be used or mts_idx can be derived as 0 or 1. Alternatively, if there is one transformation set, mts_flag may be signaled with a value of 1, indicating that the first transformation pair of the four transformation sets is applied if mts_idx is 0, and indicating that the first transformation pair of the six transformation sets is applied if mts_idx is 1.
[0241] For example, binarization according to the value of the MTS index (mts_idx) can be as shown in the table below.
[0242]
[0243] Additionally, according to another example, without signaling mts_flag, if mts_idx is 0, it can indicate that MTS is not applied, i.e., DCT 2 is applied to the first transform. In the above case, the binarization of the MTS index can be as shown in the table below.
[0244]
[0245] The MTS index can be binarized as truncated rice (or truncated unary) as shown in the above tables, or can be binarized based on a fixed-length coding method.
[0246] Alternatively, as another example, if there is only one transformation set, a specific transformation pair may be used rather than a four-transform set or a six-transform set, and mts_idx may point to a specific transformation pair. In this case, mts_idx pointing to a single transformation set may be signaled as 6 in Table 2 or 7 in Table 3. Alternatively, mts_idx may be signaled as an intermediate value, for example, 3 or 4, in Table 2 or Table 3.
[0247] Alternatively, if IDT (Identity Transformation) is applied and any one of the six transformation kernels is applied, information indicating whether IDT (Identity Transformation) is applied in the horizontal direction or the vertical direction may be further signaled after mts_idx. For example, idt_flag indicating whether IDT (Identity Transformation) is applied may be signaled, and if idt_flag is 1, the direction in which the identity transform is applied may be further signaled. Alternatively, flag information indicating whether the identity transform is applied in the horizontal direction and the vertical direction may be signaled, respectively.
[0248] Alternatively, when IDT (Identity Transform) is applied, the upper level can signal mts_inter_enabled_flag or mts_intra_enabled_flag, and the lower level can signal kernel indices that indicate the six kernels individually for each direction. The kernel index in the horizontal direction can be signaled as mts_horizontal_idx, for example, and the kernel index in the vertical direction as mts_vertical_idx.
[0249] Below, we describe Inter MTS, in which MTS is applied to inter prediction blocks.
[0250] For example, when MTS is applied to an inter-predicted block, the transform kernel can be DST 7, DCT 8, and four transform pairs ((DCT 8, DCT 8), (DCT 8, DST 7), (DST 7, DCT 8), (DST 7, DST 7)) can be applied to each CU. For example, for 4-point, 8-point, and 16-point transforms (i.e., when the width or width of the transform block is 16 or less), a separable KLT core kernel can be used instead of DST 7, DCT 8. In addition, for large resolution sequences such as width > 1080, the maximum CU size for using Inter-MTS can be limited to 32X32 or less, and set to 16 for the remaining sequences.
[0251] Figure 16 illustrates the IntraTMP mode as an example.
[0252] Referring to Fig. 16, the template most similar to the current template (a surrounding reference sample (L-shape) template) can be found within a predefined search range of a restored part within the current picture, and prediction of the corresponding block can be performed. In this case, a block vector indicating the location of a matching block (reference block) within the current picture can be derived or stored based on the current block location. At this time, the predefined search range can be located within the current CTU, the left CTU, the upper left CTU, the upper CTU, and the upper right CTU.
[0253] For example, in the case of IntraTMP mode, it is an intra mode, but in the case of the IntraTMP mode, an Inter-MTS kernel instead of an Intra-MTS kernel may be applied. In this case, the Inter-MTS kernel may include four candidates.
[0254] Additionally, although it is an intra mode, it is also possible to apply an inter-MTS kernel instead of an intra-MTS kernel in IBC mode. That is, the inter-MTS kernel can be applied in IBC mode, the inter-MTS kernel can be applied in IntraTMP mode, or the inter-MTS kernel can be applied in both IBC mode and IntraTMP mode.
[0255] For example, even in the case of inter MTS, mts_idx can be signaled and binarized. For example, when DCT 8 or DST 7 is applied and a transform pair to which KLT0 or KLT1 is applied can be binarized as shown in the table below. That is, when DCT 8 or DST 7 is applied and a transform pair to which KLT0 or KLT1 is applied, mts_idx can be binarized as shown in the table below.
[0256]
[0257] Meanwhile, for example, DCT 8 or DST 7 may be applied in the horizontal direction, and KLT0 or KLT1 may be applied in the vertical direction. In this case, the transformation kernel can be derived using the mapping relationship between DCT 8 or DST 7 and KLT0 or KLT1. In this case, mts_idx can be binarized as shown in the table below.
[0258]
[0259] Also, since whether DCT 8 or DST 7 or KLT0 or KLT1 is used as the transform kernel can be derived depending on the size of the transform block, it is possible to indicate through mts_idx whether DCT 8 or DST 7, KLT0 or KLT1 is applied. For example, in the case of an inter block of size 8X32, 8-point KLT0 or KLT1 can be applied in the horizontal direction, and either DCT 8 or DST 7 can be applied in the vertical direction. In the case where KLT0 is applied in the horizontal direction and DST 7 is applied in the vertical direction, mts_idx can be signaled with a value of 1.
[0260] Fig. 17 schematically illustrates a video / image encoding method according to an embodiment(s) of the present disclosure. The method disclosed in Fig. 17 may be performed by the encoding device disclosed in Fig. 2. Specifically, for example, S1700 to S1710 of Fig. 17 may be performed by the prediction unit (220) of the encoding device (200), S1720 to S1740 may be performed by the residual processing unit (230) of the encoding device (200), and S1750 may be performed by the entropy encoding unit (240) of the encoding device (200). The method disclosed in Fig. 17 may include the embodiments described above in the present disclosure.
[0261] Referring to FIG. 17, the encoding device determines a prediction mode for the current block (S1700). For example, the encoding device may determine a prediction mode for the current block. For example, the prediction mode for the current block may be an intra-prediction related mode or an inter-prediction related mode.
[0262] For example, based on the prediction mode being an inter prediction mode, either DST 7 or DCT 8 may be applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 may be applied to the horizontal transformation of the current block.
[0263] For example, if the size of an inter block coded in the inter prediction mode is 8x32, either DST 7 or DCT 8 may be applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 may be applied to the horizontal transformation of the current block. In this case, 8-point KLT 0 or KLT 1 may be applied to the horizontal transformation.
[0264] Additionally, for example, based on the fact that the value of the MTS index information is 1 and the prediction mode is an inter prediction mode, DST 7 may be applied to the vertical transformation of the current block and KLT 0 may be applied to the horizontal transformation of the current block.
[0265] The encoding device derives prediction samples for the current block (S1710). For example, the encoding device may derive prediction samples for the current block based on the prediction mode.
[0266] The encoding device derives residual samples for the current block (S1720). For example, the encoding device may derive residual samples for the current block based on the predicted samples.
[0267] The encoding device derives transform coefficients for the current block (S1730). For example, the encoding device may derive transform coefficients for the current block by applying MTS (Multiple Transform Selection) to the residual samples.
[0268] For example, in applying the MTS to the residual samples, a set of transform kernels including a vertical transform kernel and a horizontal transform kernel can be used. That is, a transform can be applied to the residual samples based on the set of transform kernels including at least one or more of candidates including DCT 2, DST 7, or DCT 8, or the set of transform kernels including at least one or more of candidates including DCT 2, KLT 0, or KLT 1, or the set of transform kernels including at least one or more of candidates including DCT 2, DST 7, DCT 8, KLT 0, and KLT 1. At this time, the transform may be a first transform.
[0269] The encoding device generates residual information for the current block (S1740). For example, the encoding device may generate residual information for the current block based on the transform coefficients. For example, the residual information may include information regarding quantized transform coefficients for the residual samples.
[0270] The encoding device can encode image information including prediction-related information and residual information (S1750). For example, the encoding device can encode image information including prediction-related information and residual information for the current block.
[0271] For example, the image information may include MTS index information. At this time, the MTS index information may indicate a set of transform kernels in the MTS. In addition, the set of transform kernels may include at least one of a discrete cosine transform (DCT) kernel, a discrete sine transform (DST) kernel, or a Karhunen-Loe've Transform (KLT) kernel based on the size of the current block. In addition, the set of transform kernels may include a vertical transform kernel and a horizontal transform kernel.
[0272] Additionally, the set of transform kernels may be composed of at least one of candidates including discrete cosine transform (DCT) 2, discrete sine transform (DST) 7, or DCT 8.
[0273] For example, based on the size of the current block, the set of transform kernels may be composed of at least one of the candidates including DCT 2, DST 7 or DCT 8.
[0274] For example, based on the case where the size of the current block is the first size, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. For example, when the width or height of the current block is greater than 16, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. In addition, for example, when the height of the current block is greater than 16, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. In addition, for example, when the width or height of the current block is greater than 16 or greater than 16, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. However, this is not limited to the case where the width or height of the current block is greater than 16.
[0275] Additionally, for example, based on at least one of the size of the current block or the prediction mode, the set of transform kernels may be composed of at least one of candidates including DCT 2, DST 7, or DCT 8. For example, the prediction mode may be an intra prediction mode. Additionally, the prediction mode may be an inter prediction mode. Additionally, the prediction mode may include a DIMD mode or a TIMD mode, etc. Additionally, the prediction mode may include an IntraTMP or an SGPM, etc.
[0276] For example, a first index value of the MTS index information may indicate that the transform kernel set includes the DCT 2, a second index value of the MTS index information may indicate that the transform kernel set includes the DST 7, a third index value of the MTS index information may indicate that the transform kernel set includes the DCT 8 and the DST 7, a fourth index value of the MTS index information may indicate that the transform kernel set includes the DST 7 and the DCT 8, and a fifth index value of the MTS index information may indicate that the transform kernel set includes the DCT 8.
[0277] Also, for example, a first index value of the MTS index information may indicate that both the vertical transform kernel and the horizontal transform kernel in the transform kernel set are the DCT 2, a second index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the DST 7, a third index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the DST 7, a fourth index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the DCT 8, and a fifth index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the DCT 8.
[0278] Additionally, the set of transform kernels may be composed of at least one of candidates including DCT 2, KLT (Karhunen-Loe`ve Transform) 0, or KLT 1.
[0279] For example, based on the size of the current block, the set of transform kernels may be composed of at least one of candidates including DCT 2, KLT 0 or KLT 1.
[0280] For example, based on the case where the size of the current block is the second size, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. For example, when the width or height of the current block is 16 or less, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. In addition, for example, when the height of the current block is large, 16 or less, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. In addition, for example, when the width or height of the current block is 16 or less and 16 or less, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. However, this is not limited to the case where the width or height of the current block is 16 or less.
[0281] Additionally, for example, based on at least one of the size of the current block or the prediction mode, the set of transform kernels may be composed of at least one of candidates including DCT 2, KLT 0, or KLT 1. For example, the prediction mode may be an intra prediction mode. Additionally, the prediction mode may be an inter prediction mode. Additionally, the prediction mode may include a DIMD mode or a TIMD mode, etc. Additionally, the prediction mode may include an IntraTMP or an SGPM, etc.
[0282] For example, a first index value of the MTS index information may indicate that the transform kernel set includes the DCT 2, a second index value of the MTS index information may indicate that the transform kernel set includes the KLT 0, a third index value of the MTS index information may indicate that the transform kernel set includes the KLO 0 and the KLT 1, a fourth index value of the MTS index information may indicate that the transform kernel set includes the KLO 0 and the KLT 1, and a fifth index value of the MTS index information may indicate that the transform kernel set includes the KLT 1.
[0283] Also, for example, a first index value of the MTS index information may indicate that both the vertical transform kernel and the horizontal transform kernel in the transform kernel set are DCT 2, a second index value of the MTS index information may indicate that the vertical transform kernel is KLT 0 and the horizontal transform kernel is KLT 0, a third index value of the MTS index information may indicate that the vertical transform kernel is KLT 1 and the horizontal transform kernel is KLT 0, a fourth index value of the MTS index information may indicate that the vertical transform kernel is KLT 0 and the horizontal transform kernel is KLT 1, and a fifth index value of the MTS index information may indicate that the vertical transform kernel is KLT 1 and the horizontal transform kernel is KLT 1.
[0284] For example, in the examples described above, the first index value may be 0, the second index value may be 1, the third index value may be 2, the fourth index value may be 3, and the fifth index value may be 4. In addition, the first index value may be binarized to 0, the second index value may be binarized to 10, the third index value may be binarized to 110, the fourth index value may be binarized to 1110, and the fifth index value may be binarized to 1111.
[0285] Additionally, the set of transform kernels may be composed of at least one of candidates including DCT 2, DST 7, DCT 8, KLT 0, and KLT 1.
[0286] For example, the vertical transform kernel may be one of the DCT 2, the DST 7 or the DCT 8, and the horizontal transform kernel may be one of the DCT 2, the KLT 0 or the KLT 1.
[0287] Also, for example, a first index value of the MTS index information may indicate that both the vertical transform kernel and the horizontal transform kernel in the transform kernel set are DCT 2, a second index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the KLT 0, a third index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the KLT 1, a fourth index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the KLT 0, and a fifth index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the KLT 1.
[0288] For example, in the above example, the first index value may be 0, the second index value may be 1, the third index value may be 2, the fourth index value may be 3, and the fifth index value may be 4. In addition, the first index value may be binarized to 0, the second index value may be binarized to 10, the third index value may be binarized to 110, the fourth index value may be binarized to 1110, and the fifth index value may be binarized to 1111.
[0289] Additionally, the set of transform kernels may be determined based on the size of the current block. For example, based on the size of the current block, it may be determined whether the set of transform kernels consists of DST 7 or DCT 8, or KLT 0 or KLT 1.
[0290] Additionally, for example, based on the MTS index information, either DST 7 or DCT 8 may be applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 may be applied to the horizontal transformation of the current block.
[0291] Additionally, the value of the MTS index information indicating the set of transformation kernels used in the MTS can be derived based on TU (truncated unary) binarization.
[0292] Additionally, the image information may include various information according to embodiments of the present disclosure. For example, the image information may include information disclosed in at least one of the tables described above.
[0293] Additionally, encoded image information can be output in the form of a bitstream. The bitstream can be transmitted to a decoding device via a network or storage medium.
[0294] In addition, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and a reconstructed block) based on the reference samples and the residual samples. This is to derive the same prediction result as that performed by the decoding device from the encoding device, thereby increasing coding efficiency. Accordingly, the encoding device can store the reconstructed picture (or reconstructed samples, reconstructed block) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure, etc. can be further applied to the reconstructed picture.
[0295] According to the above-described embodiment(s), even in the case of MTS according to the inter prediction mode, the transformation kernel can be adaptively determined by signaling MTS index information indicating the transformation loss word to which DST 7 or DCT 8 is applied and the transformation loss word to which KLT 0 or KLT 1 is applied, thereby efficiently performing the transformation.
[0296] In addition, according to the above-described embodiment(s), it is possible to efficiently perform transformation by adaptively determining the transformation kernel according to the size of the block by determining whether to use DST 7 or DCT 8 as the transformation kernel, or KLT 0 or KLT 1 as the transformation kernel, depending on the size of the transformation block.
[0297] In addition, according to the above-described embodiment(s), it is possible to efficiently perform transformation by adaptively determining a transformation kernel according to the size of a block by determining whether to use DST 7 or DCT 8 as a transformation kernel, or KLT 0 or KLT 1 as a transformation kernel, depending on the prediction mode of the current block.
[0298] FIG. 18 schematically illustrates a video / image decoding method according to an embodiment(s) of the present disclosure. The method disclosed in FIG. 17 may be performed by the decoding device disclosed in FIG. 3. Specifically, for example, S1800 of FIG. 18 may be performed by the entropy decoding unit (310) of the decoding device (300), S1810 to S1820 may be performed by the prediction unit (330) of the decoding device (300), and S1830 to S1840 may be performed by the residual processing unit (320) of the decoding device (300). The method disclosed in FIG. 18 may include the embodiments described above in the present disclosure.
[0299] Referring to FIG. 18, the decoding device receives image information including prediction-related information and residual information for the current block (S1800). For example, the decoding device may receive image information including the prediction-related information and residual information for the current block via a bitstream.
[0300] For example, the image information may include MTS index information. At this time, the MTS index information may indicate a set of transform kernels used in the MTS. In addition, the set of transform kernels may include at least one of a discrete cosine transform (DCT) kernel, a discrete sine transform (DST) kernel, or a Karhunen-Loe've Transform (KLT) kernel based on the size of the current block. In addition, the set of transform kernels may include a vertical transform kernel and a horizontal transform kernel.
[0301] Additionally, the set of transform kernels may be composed of at least one of candidates including discrete cosine transform (DCT) 2, discrete sine transform (DST) 7, or DCT 8.
[0302] For example, based on the size of the current block, the set of transform kernels may be composed of at least one of the candidates including DCT 2, DST 7 or DCT 8.
[0303] For example, based on the case where the size of the current block is the first size, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. For example, when the width or height of the current block is greater than 16, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. In addition, for example, when the height of the current block is greater than 16, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. In addition, for example, when the width or height of the current block is greater than 16 or greater than 16, the transform kernel set may be configured with at least one or more of the candidates including DCT 2, DST 7, or DCT 8. However, this is not limited to the case where the width or height of the current block is greater than 16.
[0304] Additionally, for example, based on at least one of the size of the current block or the prediction mode, the set of transform kernels may be composed of at least one of candidates including DCT 2, DST 7, or DCT 8. For example, the prediction mode may be an intra prediction mode. Additionally, the prediction mode may be an inter prediction mode. Additionally, the prediction mode may include a DIMD mode or a TIMD mode, etc. Additionally, the prediction mode may include an IntraTMP or an SGPM, etc.
[0305] For example, a first index value of the MTS index information may indicate that the transform kernel set includes the DCT 2, a second index value of the MTS index information may indicate that the transform kernel set includes the DST 7, a third index value of the MTS index information may indicate that the transform kernel set includes the DCT 8 and the DST 7, a fourth index value of the MTS index information may indicate that the transform kernel set includes the DST 7 and the DCT 8, and a fifth index value of the MTS index information may indicate that the transform kernel set includes the DCT 8.
[0306] Also, for example, a first index value of the MTS index information may indicate that both the vertical transform kernel and the horizontal transform kernel in the transform kernel set are the DCT 2, a second index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the DST 7, a third index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the DST 7, a fourth index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the DCT 8, and a fifth index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the DCT 8.
[0307] Additionally, the set of transform kernels may be composed of at least one of candidates including DCT 2, KLT (Karhunen-Loe`ve Transform) 0, or KLT 1.
[0308] For example, based on the size of the current block, the set of transform kernels may be composed of at least one of candidates including DCT 2, KLT 0 or KLT 1.
[0309] For example, based on the case where the size of the current block is the second size, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. For example, when the width or height of the current block is 16 or less, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. In addition, for example, when the height of the current block is large, 16 or less, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. In addition, for example, when the width or height of the current block is 16 or less and 16 or less, the transform kernel set may be configured with at least one or more of candidates including DCT 2, KLT 0, or KLT 1. However, this is not limited to the case where the width or height of the current block is 16 or less.
[0310] Additionally, for example, based on at least one of the size of the current block or the prediction mode, the set of transform kernels may be composed of at least one of candidates including DCT 2, KLT 0, or KLT 1. For example, the prediction mode may be an intra prediction mode. Additionally, the prediction mode may be an inter prediction mode. Additionally, the prediction mode may include a DIMD mode or a TIMD mode, etc. Additionally, the prediction mode may include an IntraTMP or an SGPM, etc.
[0311] For example, a first index value of the MTS index information may indicate that the transform kernel set includes the DCT 2, a second index value of the MTS index information may indicate that the transform kernel set includes the KLT 0, a third index value of the MTS index information may indicate that the transform kernel set includes the KLO 0 and the KLT 1, a fourth index value of the MTS index information may indicate that the transform kernel set includes the KLO 0 and the KLT 1, and a fifth index value of the MTS index information may indicate that the transform kernel set includes the KLT 1.
[0312] Also, for example, a first index value of the MTS index information may indicate that both the vertical transform kernel and the horizontal transform kernel in the transform kernel set are DCT 2, a second index value of the MTS index information may indicate that the vertical transform kernel is KLT 0 and the horizontal transform kernel is KLT 0, a third index value of the MTS index information may indicate that the vertical transform kernel is KLT 1 and the horizontal transform kernel is KLT 0, a fourth index value of the MTS index information may indicate that the vertical transform kernel is KLT 0 and the horizontal transform kernel is KLT 1, and a fifth index value of the MTS index information may indicate that the vertical transform kernel is KLT 1 and the horizontal transform kernel is KLT 1.
[0313] For example, in the examples described above, the first index value may be 0, the second index value may be 1, the third index value may be 2, the fourth index value may be 3, and the fifth index value may be 4. In addition, the first index value may be binarized to 0, the second index value may be binarized to 10, the third index value may be binarized to 110, the fourth index value may be binarized to 1110, and the fifth index value may be binarized to 1111.
[0314] Additionally, the set of transform kernels may be composed of at least one of candidates including DCT 2, DST 7, DCT 8, KLT 0, and KLT 1.
[0315] For example, the vertical transform kernel may be one of the DCT 2, the DST 7 or the DCT 8, and the horizontal transform kernel may be one of the DCT 2, the KLT 0 or the KLT 1.
[0316] Also, for example, a first index value of the MTS index information may indicate that both the vertical transform kernel and the horizontal transform kernel in the transform kernel set are DCT 2, a second index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the KLT 0, a third index value of the MTS index information may indicate that the vertical transform kernel is the DST 7 and the horizontal transform kernel is the KLT 1, a fourth index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the KLT 0, and a fifth index value of the MTS index information may indicate that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the KLT 1.
[0317] For example, in the above example, the first index value may be 0, the second index value may be 1, the third index value may be 2, the fourth index value may be 3, and the fifth index value may be 4. In addition, the first index value may be binarized to 0, the second index value may be binarized to 10, the third index value may be binarized to 110, the fourth index value may be binarized to 1110, and the fifth index value may be binarized to 1111.
[0318] Additionally, the set of transform kernels may be determined based on the size of the current block. For example, based on the size of the current block, it may be determined whether the set of transform kernels consists of DST 7 or DCT 8, or KLT 0 or KLT 1.
[0319] Additionally, for example, based on the MTS index information, either DST 7 or DCT 8 may be applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 may be applied to the horizontal transformation of the current block.
[0320] Additionally, the value of the MTS index information indicating the set of transformation kernels used in the MTS can be derived based on TU (truncated unary) binarization.
[0321] Additionally, the image information may include various information according to embodiments of the present disclosure. For example, the image information may include information disclosed in at least one of the tables described above.
[0322] The decoding device derives a prediction mode for the current block (S1810). For example, the decoding device may derive a prediction mode for the current block based on the prediction-related information.
[0323] Based on the above prediction mode being an inter prediction mode, either DST 7 or DCT 8 may be applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 may be applied to the horizontal transformation of the current block.
[0324] For example, if the size of an inter block coded in the inter prediction mode is 8x32, either DST 7 or DCT 8 may be applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 may be applied to the horizontal transformation of the current block. In this case, 8-point KLT 0 or KLT 1 may be applied to the horizontal transformation.
[0325] Additionally, for example, based on the fact that the value of the MTS index information is 1 and the prediction mode is an inter prediction mode, DST 7 may be applied to the vertical transformation of the current block and KLT 0 may be applied to the horizontal transformation of the current block.
[0326] The decoding device derives prediction samples for the current block (S1820). For example, the decoding device may derive prediction samples for the current block based on the prediction mode.
[0327] The decoding device derives transform coefficients for the current block (S1830). For example, the decoding device may derive transform coefficients for the current block based on the residual information.
[0328] The decoding device derives residual samples for the current block (S1840). For example, the decoding device may derive residual samples for the current block by applying MTS (Multiple Transform Selection) to the transform coefficients.
[0329] For example, in applying the MTS to the transform coefficients, a transform kernel pair including the vertical transform kernel and the horizontal transform kernel described above may be used. That is, an inverse transform may be applied to the transform coefficients based on the transform kernel pair consisting of at least one or more of the candidates including the DCT 2, the DST 7, or the DCT 8, or the transform kernel pair consisting of at least one or more of the candidates including the DCT 2, the KLT 0, or the KLT 1, or the transform kernel set consisting of at least one or more of the candidates including the DCT 2, the DST 7, the DCT 8, the KLT 0, and the KLT 1. At this time, the inverse transform may be an inverse first-order transform.
[0330] The decoding device can generate reconstructed samples based on prediction samples of the current block. For example, the decoding device can generate the reconstructed samples for the current block based on residual samples for the current block and the prediction samples. The residual samples for the current block can be generated based on received residual information. In addition, the decoding device can generate a reconstructed picture including the reconstructed samples, for example. As described above, the decoding device can then apply an in-loop filtering procedure, such as a deblocking filtering and / or an SAO procedure, to the reconstructed picture to improve subjective / objective image quality, as needed.
[0331] According to the above-described embodiment(s), even in the case of MTS according to the inter prediction mode, the transformation kernel can be adaptively determined by signaling MTS index information indicating the transformation loss word to which DST 7 or DCT 8 is applied and the transformation loss word to which KLT 0 or KLT 1 is applied, thereby efficiently performing the transformation.
[0332] In addition, according to the above-described embodiment(s), it is possible to efficiently perform transformation by adaptively determining the transformation kernel according to the size of the block by determining whether to use DST 7 or DCT 8 as the transformation kernel, or KLT 0 or KLT 1 as the transformation kernel, depending on the size of the transformation block.
[0333] In addition, according to the above-described embodiment(s), it is possible to efficiently perform transformation by adaptively determining a transformation kernel according to the size of a block by determining whether to use DST 7 or DCT 8 as a transformation kernel, or KLT 0 or KLT 1 as a transformation kernel, depending on the prediction mode of the current block.
[0334] Although the methods described in the above-described embodiments are described based on a flowchart as a series of steps or blocks, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will appreciate that the steps depicted in the flowchart are not exclusive, and that other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of the present disclosure.
[0335] The method according to the embodiments of the present disclosure described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0336] The embodiments of the present disclosure described above may also be implemented in the form of a recording medium containing computer-executable (program) instructions, such as program modules, executed by a computer. The modules may be stored in a memory and executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. Computer-readable media may be any available media that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include both computer storage media and communication media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any information delivery media.
[0337] In addition, the embodiments of the present disclosure described above may be implemented as a computer program (or computer program product) including computer-executable instructions. The computer program includes programmable machine instructions processed by the processor, and may be implemented in a high-level programming language, an object-oriented programming language, assembly language, or machine language. In addition, the computer program may be recorded on a tangible computer-readable recording medium (e.g., memory, a hard disk, a magnetic / optical medium, or a solid-state drive (SSD), etc.).
[0338] Accordingly, the embodiments of the present disclosure described above can be implemented by executing the computer program described above on a computing device. The computing device may include a processor, memory, a storage device, a high-speed interface connecting the memory and a high-speed expansion port, and at least some of a low-speed interface connecting the low-speed bus and the storage device. Each of these components is connected to one another using various buses and may be mounted on a common motherboard or in another suitable manner.
[0339] Here, the processor can process instructions within the computing device, such as instructions stored in a memory or storage device to display graphical information for providing a graphical user interface (GUI) on an external input / output device, such as a display connected to a high-speed interface. In another embodiment, multiple processors and / or multiple buses may be utilized, as appropriate, together with multiple memories and memory types. The processor may also be implemented as a chipset comprising multiple independent analog and / or digital processors.
[0340] Memory also stores information within a computing device. For example, memory may consist of volatile memory units or a collection of volatile memory units. For another example, memory may consist of nonvolatile memory units or a collection of nonvolatile memory units. Memory may also be another form of computer-readable media, such as magnetic or optical disks.
[0341] A storage device can provide a large amount of storage space to a computing device. The storage device can be a computer-readable medium or a configuration including such a medium, and can include, for example, devices within a storage area network (SAN) or other configurations, and can be a floppy disk device, a hard disk device, an optical disk device, a tape device, flash memory, or other similar semiconductor memory device or device array.
[0342] Additionally, the network may be implemented as a wired network such as a Local Area Network (LAN), a Wide Area Network (WAN), or a Value Added Network (VAN), or as a wireless network of various types such as a mobile radio communication network or a satellite communication network.
[0343] Although the present disclosure has been described above with reference to the embodiments illustrated in the drawings, these are merely exemplary, and those skilled in the art will understand that various modifications and variations of the embodiments are possible from the above-described embodiments. In other words, the scope of the present disclosure is not limited to the above-described embodiments, and various modifications and improvements made by those skilled in the art using the basic concepts of the embodiments defined in the following claims also fall within the scope of the embodiments. Therefore, the true technical protection scope of the present disclosure should be determined by the technical spirit of the appended claims.
Claims
1. In a video decoding method performed by a decoding device, A step of receiving image information including prediction-related information and residual information for a current block; A step of deriving a prediction mode for the current block based on the above prediction-related information; A step of deriving prediction samples for the current block based on the above prediction mode; A step of deriving transformation coefficients for the current block based on the residual information; and A step of applying MTS (Multiple Transform Selection) to the above transformation coefficients to derive residual samples for the current block, The above image information includes MTS index information, The above MTS index information indicates a set of transformation kernels used in the above MTS, An image decoding method, characterized in that the set of transform kernels includes at least one of a DCT (discrete cosine transform) kernel, a DST (discrete sine transform) kernel, or a KLT (Karhunen-Loe've Transform) kernel based on the size of the current block.
2. In paragraph 1, An image decoding method, characterized in that the set of transform kernels is composed of at least one of candidates including DCT 2, DST 7 or DCT 8, based on the case where the size of the current block is the first size.
3. In paragraph 2, The first index value of the above MTS index information indicates that the transform kernel set includes the DCT 2, The second index value of the above MTS index information indicates that the set of transformation kernels includes the DST 7, The third index value of the above MTS index information indicates that the transform kernel set includes the DCT 8 and the DST 7, The fourth index value of the above MTS index information indicates that the transform kernel set includes the DST 7 and the DCT 8, An image decoding method, characterized in that the fifth index value of the above MTS index information indicates that the transform kernel set includes the DCT 8.
4. In paragraph 2, The first index value of the above MTS index information indicates that both the vertical transform kernel and the horizontal transform kernel within the above transform kernel set are DCT 2, The second index value of the above MTS index information indicates that the vertical transformation kernel is the DST 7 and the horizontal transformation kernel is the DST 7. The third index value of the above MTS index information indicates that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the DST 7. The fourth index value of the above MTS index information indicates that the vertical transformation kernel is the DST 7 and the horizontal transformation kernel is the DCT 8. An image decoding method, characterized in that the fifth index value of the MTS index information indicates that the vertical transform kernel is DCT 8 and the horizontal transform kernel is DCT 8.
5. In paragraph 1, An image decoding method, characterized in that the set of transform kernels is composed of at least one of candidates including DCT 2, KLT 0 or KLT 1, based on the case where the size of the current block is the second size.
6. In paragraph 5, The first index value of the above MTS index information indicates that the transform kernel set includes the DCT 2, The second index value of the above MTS index information indicates that the set of transformation kernels includes the KLT 0, The third index value of the above MTS index information indicates that the set of transformation kernels includes the KLO 0 and the KLT 1, The fourth index value of the above MTS index information indicates that the transformation kernel set includes the KLO 0 and the KLT 1, An image decoding method, characterized in that the fifth index value of the above MTS index information indicates that the set of transform kernels includes the KLT 1.
7. In paragraph 5, The first index value of the above MTS index information indicates that both the vertical transform kernel and the horizontal transform kernel within the above transform kernel set are DCT 2, The second index value of the above MTS index information indicates that the vertical transformation kernel is KLT 0 and the horizontal transformation kernel is KLT 0, The third index value of the above MTS index information indicates that the vertical transformation kernel is KLT 1 and the horizontal transformation kernel is KLT 0. The fourth index value of the above MTS index information indicates that the vertical transformation kernel is KLT 0 and the horizontal transformation kernel is KLT 1. An image decoding method, characterized in that the fifth index value of the MTS index information indicates that the vertical transform kernel is KLT 1 and the horizontal transform kernel is KLT 1.
8. In paragraph 1, An image decoding method, characterized in that the set of transform kernels comprises at least one of candidates including DCT 2, DST 7, DCT 8, KLT 0 and KLT 1.
9. In paragraph 8, The vertical transform kernel within the above transform kernel set is one of the DCT 2, the DST 7 or the DCT 8, An image decoding method, characterized in that the horizontal transform kernel in the above transform kernel set is one of the DCT 2, the KLT 0 or the KLT 1.
10. In paragraph 8, The first index value of the above MTS index information indicates that both the vertical transformation kernel and the horizontal transformation kernel within the above converted kernel set are DCT 2. The second index value of the above MTS index information indicates that the vertical transformation kernel is the DST 7 and the horizontal transformation kernel is the KLT 0. The third index value of the above MTS index information indicates that the vertical transformation kernel is the DST 7 and the horizontal transformation kernel is the KLT 1. The fourth index value of the above MTS index information indicates that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the KLT 0. An image decoding method, characterized in that the fifth index value of the MTS index information indicates that the vertical transform kernel is the DCT 8 and the horizontal transform kernel is the KLT 1.
11. In paragraph 10, An image decoding method, characterized in that the first index value is 0, the second index value is 1, the third index value is 2, the fourth index value is 3, and the fifth index value is 4.
12. In paragraph 1, The above set of transformation kernels is determined based on the size of the current block, An image decoding method, characterized in that based on the size of the current block, it is determined whether the transform kernel set consists of DST 7 or DCT 8, or KLT 0 or KLT 1.
13. In paragraph 1, An image decoding method, characterized in that, based on the MTS index information, either DST 7 or DCT 8 is applied to the vertical transformation of the current block, and either KLT 0 or KLT 1 is applied to the horizontal transformation of the current block.
14. In paragraph 1, An image decoding method, characterized in that, based on the above prediction mode being an inter prediction mode, either DST 7 or DCT 8 is applied to a vertical transformation of the current block, and either KLT 0 or KLT 1 is applied to a horizontal transformation of the current block.
15. In paragraph 1, An image decoding method, characterized in that DST 7 is applied to vertical transformation of the current block and KLT 0 is applied to horizontal transformation of the current block based on the value of the above MTS index information being 1 and the above prediction mode being inter prediction mode.
16. In paragraph 1, An image decoding method, characterized in that the value of the MTS index information indicating the set of transformation kernels used in the MTS is derived based on TU (truncated unary) binarization.
17. In a video encoding method performed by an encoding device, A step for determining the prediction mode for the current block; A step of deriving prediction samples for the current block based on the above prediction mode; A step of deriving residual samples for the current block based on the above prediction samples; A step of applying MTS (Multiple Transform Selection) to the above residual samples to derive transform coefficients for the current block; A step of generating residual information for the current block based on the above transformation coefficients; and Comprising a step of encoding image information including prediction-related information for the current block and the residual information, The above image information includes MTS index information, The above MTS index information indicates a set of transformation kernels used in the above MTS, An image encoding method, characterized in that the set of transform kernels includes at least one of a DCT (discrete cosine transform) kernel, a DST (discrete sine transform) kernel, or a KLT (Karhunen-Loe've Transform) kernel based on the size of the current block.
18. In paragraph 17, An image encoding method, characterized in that the set of transform kernels is composed of at least one of candidates including DCT 2, DST 7 or DCT 8, based on the case where the size of the current block is the first size.
19. In Article 17, An image encoding method, characterized in that the set of transform kernels is composed of at least one of candidates including DCT 2, KLT 0 or KLT 1, based on the case where the size of the current block is the second size.
20. A method for transmitting data for an image, wherein a bitstream for the image is obtained, the bitstream being generated based on a step of determining a prediction mode for a current block, a step of deriving prediction samples for the current block based on the prediction mode, a step of deriving residual samples for the current block based on the prediction samples, a step of deriving transform coefficients for the current block based on the residual samples, a step of generating residual information based on the transform coefficients, and a step of encoding image information including prediction-related information for the current block and the residual information; and Comprising a step of transmitting the data including the bitstream, The above image information includes MTS (Multiple Transform Selection) index information, The above MTS index information represents a set of transformation kernels in the above MTS, A transmission method, characterized in that the set of transform kernels includes at least one of a DCT (discrete cosine transform) kernel, a DST (discrete sine transform) kernel, or a KLT (Karhunen-Loe`ve Transform) kernel based on the size of the current block.
Citation Information
Patent Citations
Methods and apparatus for constrained transforms for video coding and decoding having transform selection
KR101753273B1
Urine Splatter Protection Device
KR1020240009864A
Method and apparatus for automatic information classification of criminal investigation material
KR1020250057179A
Method and device for harmonizing between transformation skip mode and multiple transformation selection
KR102591265B1
KR20230169959A