Video encoding / decoding method and apparatus, and recording medium storing bit stream
By deriving the transformation coefficients of the image block from the bitstream and reconstructing the residual samples, determining the transformation kernel for the image block, the problem of configuring the transformation set and selecting the transformation kernel in the prior art is solved, and more efficient image encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380072083.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively configure a predetermined transform set for the current block and determine the transform core candidate, affecting the efficiency of image encoding and decoding.
By deriving the transform coefficients of the current block from the bitstream, performing inverse quantization or inverse transformation, reconstructing the residual sample, and determining the transform kernel for the inverse transformation based on the residual sample, thereby configuring a predetermined transform set and selecting the transform kernel candidate.
Improve the transformation performance, improve the efficiency of image encoding and decoding, and improve the encoding efficiency by effectively sending indexes related to the transformation core of the current block.
Smart Images

Figure CN119999215A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus and a recording medium storing a bit stream. Background Art
[0002] Recently, demands for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images have been increasing in various application fields, and therefore, efficient image compression technology is being discussed.
[0003] There are various technologies, such as inter-frame prediction technology that uses video compression technology to predict pixel values included in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values included in the current picture by using pixel information in the current picture, entropy coding technology that assigns short symbols to values with high occurrence frequency and assigns long symbols to values with low occurrence frequency, etc., and these image compression technologies can be used to effectively compress image data and transmit or store it. Summary of the invention
[0004] Technical issues
[0005] The present disclosure seeks to provide a method and apparatus for configuring a predetermined transform set for an (inverse) transform of a current block.
[0006] The present disclosure seeks to provide a method and apparatus for determining one or more transform kernel candidates for a current block.
[0007] The present disclosure provides a method and apparatus for signaling an index related to a transform kernel of a current block.
[0008] Technical Solution
[0009] According to the image decoding method and apparatus of the present disclosure, transform coefficients of a current block can be derived from a bit stream, at least one of dequantization or inverse transformation can be performed on the transform coefficients of the current block to derive residual samples of the current block, and the current block can be reconstructed based on the residual samples of the current block.
[0010] In the image decoding method and apparatus according to the present disclosure, a transform core of a current block for inverse transformation may be determined from a transform set including a plurality of transform core candidates, and the plurality of transform core candidates may include a reference transform core for the current block.
[0011] In the image decoding method and apparatus according to the present disclosure, the plurality of transform kernels may further include transform kernel candidates derived by pairing a reference transform kernel with a predetermined trigonometric function-based transform kernel.
[0012] In the image decoding method and apparatus according to the present disclosure, the reference transformation kernel may be a transformation kernel based on a non-trigonometric function.
[0013] In the image decoding method and apparatus according to the present disclosure, the reference transform kernel may be one of a plurality of transform kernel candidates belonging to a multiple transform selection (MTS) set.
[0014] In the image decoding method and apparatus according to the present disclosure, the reference transform kernel may be a transform kernel candidate having a minimum index among a plurality of transform kernel candidates belonging to a multiple transform selection (MTS) set.
[0015] In the image decoding method and apparatus according to the present disclosure, the reference transform kernel may be a transform kernel in which the horizontal transform kernel and the vertical transform kernel are DST-7.
[0016] In the image decoding method and apparatus according to the present disclosure, the plurality of transform kernel candidates may include transform kernel candidates derived based on at least one of a width or a height of a current block.
[0017] In the image decoding method and apparatus according to the present disclosure, the number of transform kernel candidates included in the transform set may be adaptively determined based on the sum of absolute values of all or some transform coefficients in the current block.
[0018] In the image decoding method and apparatus according to the present disclosure, the transform set may further include a transform core candidate derived based on a combination of two transform core candidates pre-added to the transform set.
[0019] According to the video encoding method and apparatus of the present disclosure, residual samples of a current block can be derived, transform coefficients of the current block can be derived by performing at least one of transformation or quantization on the residual samples of the current block, and the transform coefficients of the current block can be encoded. Here, a transform kernel of the current block for transformation is determined from a transform set including a plurality of transform kernel candidates, and the plurality of transform kernel candidates may include a reference transform kernel for the current block.
[0020] A computer-readable digital storage medium is provided, which stores encoded video / image information, thereby causing a decoding device according to the present disclosure to perform an image decoding method.
[0021] A computer-readable digital storage medium storing video / image information generated according to an image encoding method according to the present disclosure is provided.
[0022] A method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure are provided.
[0023] Beneficial Effects
[0024] According to the present disclosure, by configuring a transform set including various transform kernel candidates, the performance of the transform can be improved.
[0025] In addition to the transformation kernel based on trigonometric functions, the present disclosure can also improve the performance of transformation by additionally using a transformation kernel based on non-trigonometric functions.
[0026] The present disclosure can improve the encoding efficiency of transform-related information by effectively signaling an index related to a transform kernel of a current block. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A video / image compilation system according to the present disclosure is shown.
[0028] Figure 2 A schematic block diagram showing an encoding device to which an embodiment of the present disclosure is applicable and which performs encoding of a video / image signal.
[0029] Figure 3 A schematic block diagram showing a decoding device to which an embodiment of the present disclosure is applicable and which performs decoding of a video / image signal.
[0030] Figure 4 The diagram illustrates an image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure.
[0031] Figure 5 The intra prediction mode and its prediction direction according to the present disclosure are exemplarily shown.
[0032] Figure 6 The diagram shows a schematic configuration of a decoding device (300) for performing an image decoding method according to the present disclosure.
[0033] Figure 7 The diagram illustrates an image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure.
[0034] Figure 8 The diagram shows a schematic configuration of an encoding device (200) for performing an image encoding method according to the present disclosure.
[0035] Fig. 9 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0036] Because the present disclosure can make various changes and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the present disclosure to specific embodiments, and it should be understood to include all changes, equivalents and substitutes included in the spirit and technical scope of the present disclosure. When describing each of the drawings, similar reference numerals are used for similar components.
[0037] Terms such as first, second, etc. can be used to describe various components, but components should not be limited by these terms. These terms are only used to distinguish one component from other components. For example, without departing from the scope of the present disclosure, a first component can be referred to as a second component, and similarly, a second component can also be referred to as a first component. Terms and / or combinations of any one or more related statement items in a plurality of related statement items are included.
[0038] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but another component may also exist in between. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that another component does not exist in between.
[0039] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to designate the existence of features, numbers, steps, operations, components, parts or combinations thereof described in the specification, but do not exclude the possibility of the existence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof in advance.
[0040] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the Universal Video Coding (VVC) standard. In addition, the methods / embodiments disclosed herein may be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or the next generation video / image coding standard (e.g., H.267 or H.268, etc.).
[0041] This specification proposes various embodiments of video / image coding, and unless otherwise stated, these embodiments may be performed in combination with each other.
[0042] Here, video may refer to a collection of a series of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms a part of a picture in coding. A slice / tile may include at least one coding tree unit (CTU). A picture may be composed of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of a CTU having the same height as the picture and a width assigned by the syntax requirements of the picture parameter set. A tile row is a rectangular area of a CTU having a height assigned by the picture parameter set and the same width as the width of the picture. The CTU within a tile may be arranged continuously according to a CTU raster scan, and the tiles within a picture may be arranged continuously according to a raster scan of the tile. A slice may include an integer number of complete tiles or an integer number of continuous complete CTU rows within a tile of a picture that may be exclusively included in a single NAL unit. At the same time, a picture may be divided into at least two sub-pictures. A sub-picture may be a rectangular area of at least one slice within a picture.
[0043] Pixel, pixel or picture element may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component, or only a pixel / pixel value of a chrominance component.
[0044] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with the corresponding region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0045] Here, "A or B" may refer to "only A", "only B", or "both A and B". In other words, herein, "A or B" may be interpreted as "A and / or B". For example, herein, "A, B or C" may refer to "only A", "only B", "only C", or "any combination of A, B, and C".
[0046] As used herein, a slash ( / ) or a comma may mean "and / or". For example, "A / B" may mean "A and / or B". Thus, "A / B" may mean "only A", "only B", or "both A and B". For example, "A, B, C" may mean "A, B, or C".
[0047] Here, "at least one of A and B" may refer to "only A", "only B", or "both A and B". In addition, herein, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same manner as "at least one of A and B".
[0048] In addition, herein, “at least one of A, B, and C” may refer to “only A”, “only B”, “only C”, or “any combination of A, B, and C”. In addition, “at least one of A, B, or C” or “at least one of A, B and / or C” may refer to “at least one of A, B, and C”.
[0049] In addition, the brackets used herein may refer to "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction". In other words, the "prediction" here is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". In addition, even when the indication is "prediction (ie, intra-frame prediction)", "intra-frame prediction" may be proposed as an example of "prediction".
[0050] Here, technical features described individually in one drawing may be implemented individually or simultaneously.
[0051] Figure 1 A video / image compilation system according to the present disclosure is shown.
[0052] refer to Figure 1 , a video / image coding system may include a first device (source device) and a second device (receiving device).
[0053] The source device may send the encoded video / image information or data to the receiving device in the form of a file or stream transmission through a digital storage medium or a network. The source device may include a video source, an encoding device, and a sending unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0054] The video source may obtain the video / image through the process of capturing, synthesizing or generating the video / image. The video source may include a device for capturing the video / image and a device for generating the video / image. The device for capturing the video / image may include at least one camera, a video / image archive including previously captured videos / images, etc. The device for generating the video / image may include a computer, a tablet computer, a smart phone, etc., and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc., and in this case, the process of capturing the video / image may be replaced by the process of generating the relevant data.
[0055] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bit stream.
[0056] The sending unit may send the encoded video / image information or data output in the form of a bit stream to the receiving unit of the receiving device in the form of a file or stream transmission through a digital storage medium or a network. The digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The sending unit may include an element for generating a media file in a predetermined file format and may include an element for transmission through a broadcast / communication network. The receiving unit may receive / extract a bit stream and send it to a decoding device.
[0057] The decoding device may decode the video / image by performing a series of processes such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operations of the encoding device.
[0058] The renderer may render the decoded video / image. The rendered video / image may be displayed through a display unit.
[0059] Figure 2 A rough block diagram of an encoding device to which an embodiment of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0060] refer to Figure 2, the encoding device 200 may be composed of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the above-mentioned image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include a memory 270 as an internal / external component.
[0061] The image divider 210 may divide the input image (or picture, frame) input to the encoding device 200 into at least one processing unit. As an example, the processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure.
[0062] For example, one coding unit may be segmented into a plurality of coding units having a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding process according to this specification may be performed based on a final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit may be directly used as the final coding unit, or if necessary, the coding unit may be recursively segmented into coding units of a deeper depth, and the coding unit with the best size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction described later.
[0063] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the above-mentioned final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0064] In some cases, a unit may be used interchangeably with terms such as a block or region. In general, an MxN block may represent a set of transform coefficients or samples consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample may be used as a term to make a picture (or image) correspond to a pixel or a picture element.
[0065] The encoding device 200 may subtract the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be referred to as a subtractor 231.
[0066] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a predicted block including a prediction sample for the current block. The predictor 220 may determine whether intra prediction or inter prediction is applied in units of a current block or CU. The predictor 220 may generate various information about the prediction, such as prediction mode information, and send it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0067] The intra-frame predictor 222 can predict the current block by referring to the samples in the current picture. Depending on the prediction mode, the referenced sample can be located near the current block or can be located at a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The non-directional mode may include at least one of the DC mode or the plane mode. Depending on the detail level of the prediction direction, the directional mode may include 33 directional modes or 65 directional modes. However, this is only an example, and more or less directional modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0068] The inter-frame predictor 221 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 221 may configure a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter predictor 221 may use motion information of a neighboring block as motion information of a current block. For skip mode, unlike merge mode, a residual signal may not be transmitted. For motion vector prediction (MVP) mode, a motion vector of a neighboring block is used as a motion vector predictor, and a motion vector difference is signaled to indicate a motion vector of the current block.
[0069] The predictor 220 may generate a prediction signal based on various prediction methods described later. For example, the predictor may not only apply intra prediction or inter prediction to predict a block, but may also apply intra prediction and inter prediction at the same time. It may be referred to as a combined inter and intra prediction (CIIP) mode. In addition, the predictor may be based on an intra block copy (IBC) prediction mode or may be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode may be used for content image / video coding of games, such as screen content coding (SCC), etc. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture may be signaled based on information about the palette table and the palette index. The prediction signal generated by the predictor 220 may be used to generate a reconstructed signal or a residual signal.
[0070] The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size or may be applied to non-square blocks of variable size.
[0071] The quantizer 233 may quantize the transform coefficients and send them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in the block form into a one-dimensional vector form based on the coefficient scanning order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0072] The entropy encoder 240 may perform various encoding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 240 may encode information necessary for video / video image reconstruction (e.g., values of syntax elements, etc.) in addition to transform coefficients quantized together or individually.
[0073] The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of network abstraction layer (NAL) units in the form of a bitstream. The video / image information may further include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. Here, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded and included in the bitstream through the above-mentioned encoding process. The bitstream may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmission and / or a storage unit (not shown) for storing a signal output from the entropy encoder 240 may be configured as an internal / external element of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.
[0074] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the inverse quantizer 234 and the inverse transformer 235. The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the block to be processed, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can also be used for inter-frame prediction of the next picture through filtering described later. At the same time, luminance mapping with chroma scaling (LMCS) can be applied in the picture encoding and / or reconstruction process.
[0075] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, specifically in the DPB of the memory 270. Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information about filtering and send it to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bit stream.
[0076] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied therethrough, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve encoding efficiency.
[0077] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter-frame predictor 221. The memory 270 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the pre-reconstructed picture. The stored motion information may be sent to the inter-frame predictor 221 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 270 may store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 222.
[0078] Figure 3 A rough block diagram of a decoding device to which an embodiment of the present disclosure can be applied and which performs decoding of a video / image signal is shown.
[0079] refer to Figure 3 , the decoding apparatus 300 may be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321.
[0080] According to an embodiment, the above-mentioned entropy decoder 310, residual processor 320, predictor 330, adder 340 and filter 350 may be configured by one hardware component (e.g., decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0081] When a bit stream including video / image information is input, the decoding apparatus 300 may generate a decoded signal in response to the bit stream received in the decoded signal. Figure 2 The image is reconstructed by the process of processing video / image information in the encoding device of the decoding device. For example, the decoding device 300 can derive the unit / block based on the relevant information of the block segmentation obtained from the bit stream. The decoding device 300 can perform decoding by using the processing unit applied in the encoding device. Therefore, the processing unit of decoding can be a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding device 300 can be played by a playback device.
[0082] The decoding device 300 may receive the bit stream from Figure 2 The received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device may further decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or the syntax elements described later in this document may be decoded and obtained from the bitstream by a decoding process. For example, the entropy decoder 310 may decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients of the residual. In more detail, the CABAC entropy decoding method can receive a bin corresponding to each syntax element from a bitstream, determine a context model by using information of syntax elements to be decoded, decoded information of neighboring blocks and blocks to be decoded, or information of symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the probability of occurrence of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins of the context model for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value, that is, the quantized transform coefficient and related parameter information, which is entropy decoded in the entropy decoder 310, can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300 or the receiving unit may be a component of the entropy decoder 310 .
[0083] Meanwhile, the decoding device according to this specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0084] The inverse quantizer 321 may inverse quantize the quantized transform coefficient and output the transform coefficient. The inverse quantizer 321 may rearrange the quantized transform coefficient into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantizer 321 may perform inverse quantization on the quantized transform coefficient by using a quantization parameter (e.g., quantization step size information) and obtain the transform coefficient.
[0085] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0086] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0087] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. It can be called a combined inter and intra prediction (CIIP) mode. In addition, the predictor can be based on an intra block copy (IBC) prediction mode or can be based on a palette mode for prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding of games, such as screen content coding (SCC), etc. IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction because it derives a reference block within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered as an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and the palette index can be included in the video / image information and sent with a signal.
[0088] The intra-frame predictor 331 can predict the current block by referring to samples within the current picture. Depending on the prediction mode, the referenced sample can be located near the current block or can be located at a certain distance away from the current block. In intra-frame prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0089] The inter-frame predictor 332 may derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information sent in the inter-frame prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter-frame predictor 332 may configure a motion information candidate list based on the neighboring blocks, and derive a motion vector and / or a reference picture index for the current block based on the received candidate selection information. Inter-frame prediction may be performed based on various prediction modes, and information about the prediction may include information indicating an inter-frame prediction mode for the current block.
[0090] The adder 340 may add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual of the block to be processed, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0091] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering described later, or may be used for inter prediction of the next picture. At the same time, luminance mapping with chroma scaling (LMCS) may be applied during picture decoding.
[0092] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 360, specifically the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0093] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can be derived from the motion information in its current picture (or decoded) The motion information of the block and / or the motion information of the block in the pre-reconstructed picture. The stored motion information can be sent to the inter-frame predictor 260 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0094] Here, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may also be equally or correspondingly applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.
[0095] Figure 4 The diagram illustrates an image decoding method performed by a decoding device (300) according to an embodiment of the present disclosure.
[0096] refer to Figure 4 , the transform coefficient of the current block may be derived from the bitstream (S400). That is, the bitstream may include residual information of the current block, and the transform coefficient of the current block may be derived by decoding the residual information.
[0097] refer to Figure 4 , a residual sample of the current block may be derived by performing at least one of dequantization and inverse transformation on a transformation coefficient of the current block ( S410 ).
[0098] When adaptive multi-transform selection (MTS) is applied, inverse transform may be performed based on at least one of DCT-2, DST-7, or DCT-8. Here, DCT-2, DST-7, DCT-8, etc. may be referred to as transform types, transform kernels, or transform cores.
[0099] In the present disclosure, the inverse transform may mean a separable transform. However, it is not limited to this, and the inverse transform may mean an inseparable transform, or may be a concept including a separable transform and an inseparable transform. In addition, the inverse transform in the present disclosure means a main transform, but is not limited to this, and may be applied to an auxiliary transform by being modified into the same / similar form.
[0100] For example, as a method for inverse transformation, only DCT-2 and an inseparable transform may be used, or an inseparable transform may be used in addition to at least one of DCT-2, DST-7, or DCT-8, or an inseparable transform may replace the transform kernel of one or more of DCT-2, DST-7, or DCT-8.
[0101] As a more specific embodiment, when there are (DCT-2, DCT-2), (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), (DCT-8, DCT-8) as transform core candidates for separable transforms, an inseparable transform may replace or be added to one or more of the five transform core candidates. Here, the notation (transform 1, transform 2) indicates that transform 1 is applied in the horizontal direction and transform 2 is applied in the vertical direction. When an inseparable transform replaces part of the transform core candidates, the remaining transform core candidates except (DCT-2, DCT-2) and (DST-7, DST-7) may be replaced with inseparable transforms. However, the above-mentioned transform core candidates are only examples, and other types of DCT and / or DST may be included, and transform skipping may be included as transform core candidates.
[0102] The non-separable transform may mean a transform or inverse transform based on a non-separable transform matrix. That is, unlike a separable transform that independently performs horizontal and vertical transforms by separating vertical and horizontal transforms, a non-separable transform may perform horizontal and vertical transforms at the same time.
[0103] For example, when an inseparable transform is performed on a 4×4 block, input data X to the inseparable transform is as shown in Formula 1 below.
[0104] [Formula 1]
[0105] When input data X is expressed in a vector form, vector X' can be expressed as follows.
[0106] [Formula 2]
[0107] In this case, the non-separable transformation can be performed as in the following Formula 3.
[0108] [Formula 3]
[0109] In Formula 3, F represents a transform coefficient vector, T represents a 16x16 non-separable transform matrix, and Represents the product of a matrix and a vector.
[0110] The 16x1 transform coefficient vector F may be derived by Formula 3, and F may be reconfigured into a 4x4 block according to a predetermined scanning order. The scanning order may be horizontal scanning, vertical scanning, diagonal scanning, z scanning, raster scanning, or predefined scanning.
[0111] The inseparable transform set and / or the transform kernel for the inseparable transform may be configured differently based on a prediction mode (e.g., intra mode, inter mode, etc.), a width, a height, or a number of pixels of a current block, a position of a sub-block within the current block, an explicitly signaled syntax element, statistical characteristics of neighboring samples, whether to use an auxiliary transform, or at least one of a quantization parameter (QP).
[0112] Specifically, for intra mode, the predefined intra prediction modes can be grouped to correspond to n inseparable transform sets, and each inseparable transform set can include k transform kernel candidates. Here, n and k can be arbitrary constants according to the same rules (conditions) defined for the encoding device and the decoding device.
[0113] The number of inseparable transform sets and / or the number of transform core candidates included in the inseparable transform set may be configured differently depending on the width and / or height of the current block. For example, for a 4x4 block, n1 inseparable transform sets and k1 transform core candidates may be configured. For a 4x8 block, n2 inseparable transform sets and k2 transform core candidates may be configured. In addition, the number of inseparable transform sets and the number of transform core candidates included in each inseparable transform set may be configured differently depending on the product of the width and height of the current block. For example, when the product of the width and height of the current block is equal to or greater than 256, n3 inseparable transform sets and k3 transform core candidates may be configured, and otherwise, n4 inseparable transform sets and k4 transform core candidates may be configured. That is, because the degree of change in the statistical characteristics of the residual signal varies depending on the block size, the number of inseparable transform sets and transform core candidates may be configured differently to reflect this.
[0114] When the current block is divided into multiple sub-blocks, the statistical characteristics of the residual signal may be different for each sub-block, and therefore the number of inseparable transform sets and transform core candidates may be configured differently. For example, when a 4x8 or 8x4 block is divided into two 4x4 sub-blocks and an inseparable transform is applied to each sub-block, n5 inseparable transform sets and k5 transform core candidates may be configured for the 4x4 sub-block in the upper left corner, and n6 inseparable transform sets and k6 transform core candidates may be configured for other 4x4 sub-blocks.
[0115] Based on a syntax element explicitly signaled, the number of inseparable transform sets and transform core candidates may be configured differently. As a syntax element, information indicating one of a plurality of inseparable transform configurations may be used. For example, when three inseparable transform configurations (i.e., n7 inseparable transform sets and k7 transform core candidates, n8 inseparable transform sets and k8 transform core candidates, n9 inseparable transform sets and k9 transform core candidates) are supported, the syntax element may have values of 0, 1, and 2, and the inseparable transform configuration applied to the current block may be determined based on the value of the syntax element signaled.
[0116] Based on whether to apply the auxiliary transformation and / or which auxiliary transformation is applied, the number of inseparable transformation sets and transformation core candidates can be configured differently. For example, when the auxiliary transformation is not applied, a set including n 10 A set of inseparable transformations and k 10 When applying the auxiliary transformation, you can apply n 11 A set of inseparable transformations and k 11 The inseparable transformation configuration of the transformation kernel candidates.
[0117] Based on the quantization parameter (QP) and / or the range to which the QP value belongs, different inseparable transform configurations may be applied. For example, when the QP value has a small value, a configuration including n 12 A set of inseparable transformations and k 12 On the other hand, when the QP value has a large value, a non-separable transform configuration including n transform kernel candidates can be applied. 13 A set of inseparable transformations and k 13 The non-separable transform configuration of the transform core candidates is selected. When the QP value is less than or equal to a threshold value (e.g., 32), the case is classified as having a smaller QP value, and otherwise, the case is classified as having a larger QP value. Alternatively, the range of QP values may be divided into three or more, and a different non-separable transform configuration may be applied to each range.
[0118] For larger blocks, instead of using an inseparable transform corresponding to the width and height of the block, the block may be divided into multiple sub-blocks and an inseparable transform corresponding to the width and height of the sub-block may be used. For example, when an inseparable transform is performed on a 4x8 block, the 4x8 block may be divided into two 4x4 sub-blocks, and an inseparable transform based on the 4x4 block may be used for each of the 4x4 sub-blocks. Alternatively, an 8x16 block may be divided into two 8x8 sub-blocks, and an inseparable transform based on the 8x8 block may be used.
[0119] The inseparable transform set can be determined based on the intra prediction mode of the current block and the mapping table. The mapping table can define the mapping relationship between the predefined intra prediction mode and the inseparable transform set. The predefined intra prediction mode may include two non-directional modes and 65 directional modes. In general, the inseparable transform has a larger transform kernel size than the separable transform. This means that the computational complexity required for the transform process is high and the memory required to store the transform kernel is large. At the same time, when the separable transform may only consider the statistical characteristics existing in the horizontal and / or vertical directions, the inseparable transform can simultaneously consider the statistical characteristics in the two-dimensional space including the horizontal and vertical directions, thereby providing better compression efficiency. Because the statistical characteristics and diversity of the residual are different depending on the directionality of the intra prediction mode, there may be a situation where the inseparable transform is absolutely necessary, and there may be an intra prediction mode in which the residual characteristics can be fully identified only by the separable transform. Therefore, by predefining which transform to use based on the intra prediction mode in the encoding device and the decoding device, a transform process with optimized complexity and memory requirements can be designed. The non-directional mode may include a planar mode numbered 0 and a DC mode numbered 1, and the directional mode may include intra prediction modes numbered 2 to 66. However, this is merely an example, and the present disclosure may also be applied to a case where the numbers of the predefined intra prediction modes are different.
[0120] The predefined intra prediction modes may further include an intra prediction mode from -14 to -1 and an intra prediction mode from 67 to 80 due to application of wide angle intra prediction (WAIP).
[0121] Figure 5 The intra-frame prediction mode and its prediction direction according to the present disclosure are exemplarily shown. Figure 5 , modes -14 to -1 and 2 to 33 and modes 35 to 80 are symmetrical with respect to mode 34 in terms of the prediction direction. For example, modes 10 and 58 are symmetrical with respect to the direction corresponding to mode 34, and mode -1 is symmetrical with mode 67. Therefore, for the vertical directivity mode symmetrical with respect to the horizontal directivity mode with respect to mode 34, the input data can be transposed and used. Transposing the input data means that the rows and columns in the input data MxN of the two-dimensional block are changed to columns and rows, respectively, to form NxM data.
[0122] For example, when a 4x4 block is used, the 16 data forming the 4x4 block can be appropriately arranged to form a 16x1 1-dimensional vector for an inseparable transform. In this case, the 1-dimensional vector can be formed in a row-first order or in a column-first order. The residual samples caused by the inseparable transform can be arranged in the above order to form a 2-dimensional block.
[0123] For modes -14 to -1 and 2 to 33, when the data arrangement order for forming a 16x1 input vector is row-major order, for modes 35 to 80, the input vector can be formed in column-major order.
[0124] Mode 34 can be regarded as neither a horizontal directivity mode nor a vertical directivity mode, but in the present disclosure, it is classified as belonging to a horizontal directivity mode. That is, for modes -14 to -1 and 2 to 33, the input data arrangement method for the horizontal directivity mode, i.e., the row priority order, is used, and for the vertical directivity mode symmetrical with respect to mode 34, the input data can be transposed and used.
[0125] For non-square blocks, the symmetry in square blocks (i.e., the symmetry between mode P and mode (68-P) in NxN blocks (2<=P<=33) or the symmetry between mode Q and mode (66-Q) (-14<=Q<=-1)) cannot be utilized. Therefore, in addition to the symmetry based only on the intra-frame prediction mode, the symmetry between block shapes that are in a transposed relationship with each other, that is, the symmetry between KxL blocks and LxK blocks, can also be utilized. Specifically, there is a symmetrical relationship between the KxL block predicted by mode P and the LxK block predicted by mode (68-P). Alternatively, there is a symmetrical relationship between the KxL block predicted by mode Q and the LxK block predicted by mode (66-Q).
[0126] Since the KxL block with mode 2 and the LxK block with mode 66 can be regarded as symmetrical to each other, the same transform kernel can be applied to the KxL block and the LxK block. If the inseparable transform set of the intra prediction mode for the KxL block is mapped, in order to apply the inseparable transform to the LxK block, the inseparable transform set can be derived through a mapping table corresponding to the KxL block based on mode (68-P) instead of mode P applied to the LxK block. Alternatively, the inseparable transform set can be derived through a mapping table corresponding to the KxL block based on mode (66-Q) instead of mode Q applied to the LxK block.
[0127] For example, to apply an inseparable transform to an LxK block, an inseparable transform set may be selected based on mode 2 instead of mode 66. In addition, for a KxL block, input data may be read in a predetermined order (e.g., row-first order or column-first order) to form a one-dimensional vector, and then the corresponding inseparable transform may be applied. For an LxK block, input data may be read in a transposed order to form a one-dimensional vector and then the corresponding inseparable transform may be applied. That is, when a KxL block is read in a row-first order, an LxK block may be read in a column-first order. Conversely, when a KxL block is read in a column-first order, an LxK block may be read in a row-first order.
[0128] In addition, when mode 34 is applied to a KxL block, an inseparable transform set may be determined based on mode 34, and input data may be read in a predetermined order to form a one-dimensional vector and a corresponding inseparable transform may be performed. When mode 34 is applied to an LxK block, an inseparable transform set may be determined based on mode 34, but input data may be read in a transposed order to form a one-dimensional vector and a corresponding inseparable transform may be performed.
[0129] In the present disclosure, a method for determining an inseparable transform set and a method for forming input data are described based on a KxL block. However, an inseparable transform can be performed based on an LxK block by utilizing the above-mentioned symmetry for the KxL block. Alternatively, a block having a width greater than a height can be restricted to be used as a reference block. Alternatively, the symmetry can be restricted from being utilized in the case of a non-square block. In this case, a non-square block can use a different number of inseparable transform sets and / or transform kernel candidates than a square block, and a different mapping table than a square block can be used to select an inseparable transform set.
[0130] An example of a mapping table for selecting a set of inseparable transforms is as follows: [Table 1]
[0131] Table 1 shows an example of assigning an inseparable transform set to each intra prediction mode when there are five inseparable transform sets. The value of predModeIntra means the value of the intra prediction mode considering WAIP, and TrSetIdx is an index indicating a specific inseparable transform set. In Table 1, it can be confirmed that the same inseparable transform set is applied to the mode located in the symmetric direction according to the intra prediction mode. Table 1 is only an example of using five inseparable transform sets, and does not limit the total number of inseparable transform sets used for inseparable transforms.
[0132] Alternatively, as shown in Table 2, for compression performance, the non-separable transform may not be applied to WAIP.
[0133] [Table 2]
[0134] Alternatively, as shown in Table 3, instead of configuring a separate inseparable transform set for WAIP, inseparable transform sets corresponding to adjacent intra prediction modes may be shared.
[0135] [Table 3]
[0136] The inseparable transform set may include multiple transform core candidates, and one of the multiple transform core candidates may be selectively used. To this end, an index signaled by a bitstream may be used. Alternatively, one of the multiple transform core candidates may be implicitly determined based on context information of the current block. Here, the context information may mean the size of the current block or whether an inseparable transform is applied to a neighboring block. Here, the size of the current block may be defined as the width, height, maximum / minimum of the width and height, the sum of the width and height, or the product of the width and height.
[0137] Hereinafter, a method of determining a transform kernel for inverse transform of a current block will be described in detail.
[0138] Example 1
[0139] As described above, inverse transforms may be divided into separable transforms and non-separable transforms. Separable transforms mean that transforms in the horizontal direction and the vertical direction are respectively performed on a two-dimensional block, and non-separable transforms may mean that a single transform is performed on samples constituting the entire or a part of the two-dimensional block. When expressing a separable transform, it may be expressed as a pair of a horizontal transform kernel and a vertical transform kernel, and in the present disclosure, it is expressed as (horizontal transform kernel, vertical transform kernel).
[0140] Multiple transform sets may be defined for inverse transform of the current block. Each transform set may include one or more transform kernel candidates.
[0141] For example, one of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), or (DCT-8, DCT-8) may be applied as a separable transform, and the above four transform core candidates may be considered as one transform set. In addition, (DCT-2, DCT-2) may be considered as one transform set. A transform skip that does not apply a transform may also be considered as one transform set, while (DCT-2, DCT-2) and a transform skip may be considered as one transform set. In the present disclosure, a transform core may refer to one transform (e.g., DCT-2, DST-7) or may refer to two transform pairs (e.g., (DCT-2, DCT-2)).
[0142] As another example of a transform set, there may be the aforementioned inseparable transform set. In the present disclosure, an inseparable transform applied as a main transform may be represented as an inseparable main transform (NSPT). In NSPT, multiple inseparable transform sets may be configured, and each inseparable transform set may include one or more transform cores as transform core candidates. In the case of NSPT, one of multiple inseparable transform sets is selected based on the intra prediction mode, and multiple inseparable transform sets for NSPT may be represented as an NSPT set list. This is as described above, and a detailed description thereof will be omitted here.
[0143] A group of one or more transform sets available for the current block may be configured from a plurality of predefined transform sets. The group of one or more transform sets may be configured in a predetermined area unit to which the current block belongs, and is hereinafter referred to as a set. Here, the predetermined area unit may be at least one of a picture, a slice, a coding tree unit row (CTU row), or a coding tree unit (CTU).
[0144] For example, the transform set consisting of (DCT-2, DCT-2) is called S1, and the transform set consisting of (DST-7, DST-7), (DCT-8, DST-7), (DST-7, DCT-8), and (DCT-8, DCT-8) is called S2. In addition, the above NSPT set list may include N inseparable transform sets, and the N inseparable transform sets are respectively called S 3,1 , S 3,2 ,...,S 3,N Here, N may be 35, but is not limited thereto.
[0145] When S3, 13 is selected as the inseparable transform set of NSPT based on the intra prediction mode of the current block, the transform kernel applicable to the current block can belong to S1, S2 or S 3,13 In this case, the set available for the current block can be represented as {S1, S2, S 3,13}.
[0146] As described above, because the set according to the present disclosure is a group of one or more transform sets that can be used for the current block, the set can be configured differently based on the context of the current block. Here, the context may include at least one of a shape, a size, or an intra-prediction mode. If a total of K contexts are defined, K sets can be generated, and each set can be represented as Ci (i=1, 2, ..., N). For example, when the size of the block to which NSPT is applicable is 4x4, 8x8, 16x16, and 32x32 and one of a total of 35 inseparable transform sets is selected based on the intra-prediction mode, if a different transform kernel is applied to each block size, a total of 4x35=140 contexts can be defined.
[0147] The set may be configured based on the context of the current block, and in this case, a process of selecting one of a plurality of transform sets belonging to the set and selecting one of a plurality of transform core candidates belonging to the selected transform set may be performed. Here, the selection of the transform set and the transform core candidate may be implicitly performed based on the context of the current block, or may be performed based on an index explicitly signaled.
[0148] Alternatively, the process of selecting one of the plurality of transform sets belonging to the set and the process of selecting one of the plurality of transform core candidates belonging to the selected transform set may be performed separately. For example, an index for selecting a transform set may be first signaled, and one of the plurality of transform sets belonging to the set may be selected based on the index. Then, an index indicating one of the plurality of transform core candidates belonging to the transform set may be signaled, and one of the transform core candidates may be selected from the transform set based on the signaled index. The transform core of the current block may be determined based on the selected transform core candidate. Alternatively, selecting a transform set from the set may be performed implicitly based on the context of the current block, and selecting a transform core candidate from the selected transform set may be performed based on the signaled index. Alternatively, selecting a transform set from the set may be performed based on the signaled index, and selecting a transform core candidate from the selected transform set may be performed implicitly based on the context of the current block. Alternatively, selecting a transform set from the set may be performed implicitly based on the context of the current block, and selecting a transform core candidate from the selected transform set may be performed implicitly based on the context of the current block.
[0149] Of course, when the number of transform sets belonging to the set is 1, the index for selecting the transform set may not be signaled. Similarly, when the number of transform core candidates belonging to the selected transform set is 1, the index for indicating the transform core candidate may not be signaled.
[0150] Alternatively, an index indicating one of all transform core candidates belonging to the current set may be sent by a signal. In this case, the process of selecting a transform set from the set may be omitted. In this case, all transform sets belonging to the set may be shuffled considering the priority. For example, in the case where a small-length binary code is assigned to a small-value index such as a truncated unary code, it may be advantageous to assign the small-value index to a transform core candidate that is more conducive to improving compilation performance. When all transform core candidates belonging to a set are shuffled according to priority, different shuffling may be applied to each set. In addition, instead of shuffling all transform core candidates belonging to a set, only some of them may be selectively shuffled.
[0151] Example 2
[0152] A transform kernel for inverse transform of the current block may be determined based on MTS (Multiple Transform Select).
[0153] The MTS according to the present disclosure may use at least one of DST-7, DCT-8, DCT-5, DST-4, DST-1, or IDT (identity transform) as a transform core. In addition, the MTS according to the present disclosure may further include a transform core of DCT-2.
[0154] In the present disclosure, multiple MTS sets for MTS may be defined. Based on the size of the current block and / or the intra prediction mode, one of the multiple MTS sets may be determined. For example, when determining an MTS set, 16 transform block sizes may be considered, and for the directional mode, the symmetry between the shape of the transform block and the intra prediction mode may be considered. For WAIP (Wide Angle Intra Prediction) modes (i.e., -1 to -14 (or -15), 67 to 80 (or 81)), the MTS set corresponding to mode 2 may be applied to modes -1 to -14 (or -15), and the MTS set corresponding to mode 66 may be applied to modes 67 to 80 (or 81). A separate MTS set may be assigned to the MIP (Matrix-based Intra Prediction) mode.
[0155] For example, an MTS set according to a transform block size and an intra prediction mode may be assigned / defined as shown in Table 4 below.
[0156] [Table 4]
[0157] Table 4 shows assignment of MTS sets according to 16 transform block sizes and intra prediction modes. The number of predefined MTS sets is 80, and an index indicating one of the 80 MTS sets may have a value from 0 to 79, as shown in Table 4.
[0158] [Table 5]
[0159]
[0160] Table 5 shows the transform core candidates included in each MTS set described in Table 4. Each MTS set may consist of six transform core candidates. The transform core candidate index has a value from 0 to 5 and may indicate one of the six transform core candidates. Here, each transform core candidate may be a combination of a horizontal transform core and a vertical transform core for a separable transform, and 25 transform core candidates with indexes from 0 to 24 may be defined.
[0161] [Table 6]
[0162]
[0163] Table 6 is an example of 25 transform core candidates described in Table 5. Specifically, the horizontal transform and vertical transform of the transform core candidate are expressed as (horizontal transform, vertical transform). For each transform core candidate index, the horizontal / vertical transform when the intra prediction mode is less than 35 can be opposite to the horizontal / vertical transform when the intra prediction mode is greater than or equal to 35. When the value of the intra prediction mode is greater than or equal to 35, a mode symmetric with respect to mode 34 can be derived, and the MTS set can be selected from Table 4 based on the mode. In addition, the symmetry of the block shape can also be additionally considered. When the original transform block has a WxH size, it can be regarded as having an HxW size by symmetrizing the original transform block, and the MTS set can be selected from Table 4. Here, the value of the intra prediction mode can be the value of the modified intra prediction mode. That is, as the mode value of WAIP, for from -14 (or -15) to -1, it is modified to mode 2, for from 67 to 80 (or 81), it is modified to mode 66, and for the remaining modes, the value of the original intra prediction mode can be set to the value of the modified intra prediction mode. In this case, since the extended mode of WAIP is also configured symmetrically with respect to mode 34, the symmetry with respect to mode 34 can be used for all directional modes except the planar mode and the DC mode.
[0164] For example, when predicting a 16x32 block based on mode 54, mode 14 (=68-54) can be derived as a mode symmetric to mode 54, and the block size can be considered as 32x16. In this case, the MTS set with index 72 can be selected as defined in Table 4.
[0165] When the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block shape. Alternatively, when the MIP mode is applied, the MTS set assigned to the MIP mode may be selected based on the symmetric block size considering the symmetry of the block shape. For example, when the MIP mode is applied to an 8x16 block, the 8x16 block may be regarded as a 16x8 block symmetrical thereto, and an MTS set with an index of 49 may be selected, as defined in Table 4. Alternatively, when the MIP mode is applied, the intra-frame prediction mode may be regarded as a planar mode. In this case, the MTS set assigned to the MIP mode may be selected based on the size of the current block without considering the symmetry of the block shape. Alternatively, the MTS set assigned to the MIP mode may be selected based on the symmetric block size considering the symmetry of the block shape.
[0166] For the MIP mode, a flag may be used to indicate whether the MIP mode is applied in the transposed mode. When the MIP mode is applied to the current block of MxN and the flag indicates that the transposed mode is applied, the intra prediction mode may be regarded as a plane mode, and the current block of MxN may be regarded as an NxM block. That is, from Table 4, an MTS set corresponding to a block size of NxM and a plane mode may be selected. As described in Table 6, when the value of the intra prediction mode is greater than or equal to 35, the horizontal transform and the vertical transform are exchanged, but because the intra prediction mode of the current block is regarded as a plane mode, the horizontal transform and the vertical transform of the transform core candidate may not be exchanged. Alternatively, when the MIP mode is applied to the current block of MxN and the flag indicates that the transposed mode is applied, the intra prediction mode may not be regarded as a plane mode, and the current block of MxN may be regarded as an NxM block. That is, from Table 4, an MTS set corresponding to a block size of NxM and a MIP mode may be selected.
[0167] In Table 5, the transform core candidate selected by the transform core candidate index may be set as the transform core of the current block. Alternatively, based on the size of the current block, at least one of the horizontal transform or the vertical transform of the selected transform core candidate may be changed to another transform core. For example, when the transform core candidate index is 3 and both the width and the height of the current block are less than or equal to 16, at least one of the horizontal transform or the vertical transform of the transform core candidate corresponding to the transform core candidate index of 3 may be changed to another transform core. In this case, the horizontal transform and the vertical transform may be changed independently of each other. When the difference (or the absolute value of the difference) between the value of the intra-frame prediction mode of the current block and the value of the horizontal mode is less than or equal to a predetermined threshold, the vertical transform of the selected transform core candidate may be changed to IDT (identical transform). When the difference (or the absolute value of the difference) between the value of the intra-frame prediction mode of the current block and the value of the vertical mode is less than or equal to a predetermined threshold, the horizontal transform of the selected transform core candidate may be changed to IDT (identical transform). Here, the threshold may be determined based on the width and height of the current block, as shown in Table 7 below.
[0168] [Table 7]
[0169] Table 7 is used to change the horizontal transform and / or vertical transform of the transform core candidate selected by the transform core candidate index to another transform core, and defines a threshold value according to the size of the transform block.
[0170] The six transform core candidates constituting an MTS set can be distinguished by transform core candidate indexes from 0 to 5, as defined in Table 5. The transform core candidate index can be signaled via the bitstream. A flag (MTS enable flag or MTS flag) indicating whether the MTS set is available / applied can be signaled, and when the flag indicates that the MTS set is available / applied, the transform core candidate index can be signaled. The MTS flag can consist of a bin, and one or more contexts for CABAC-based entropy coding (hereinafter referred to as CABAC contexts) can be assigned to the bin. For example, different CABAC contexts can be assigned to non-MIP mode and MIP mode, respectively.
[0171] Based on the context of the current block described above, the number of transform core candidates available for the current block can be set differently. For example, as the context of the current block, the sum of the absolute values of all or part of the transform coefficients in the current block can be considered. The sum of the absolute values of the transform coefficients is called AbsSum. When AbsSum is less than or equal to T1, only one transform core candidate corresponding to the transform core candidate index of 0 is available. When AbsSum is greater than T1 and less than or equal to T2, transform core candidates corresponding to transform core candidate indexes of 0 to 3 may be available. When AbsSum is greater than T2, six transform core candidates corresponding to transform core candidate indexes of 0 to 5 may be available. Here, T1 can be 6 and T2 can be 32, but this is just an example.
[0172] When AbsSum is less than or equal to T1, because the number of transform core candidates available for the current block is 1, the transform core candidate corresponding to the transform core candidate index of 0 can be set as the transform core of the current block without signaling the transform core candidate index. When AbsSum is greater than T1 and less than or equal to T2, because four transform core candidates are available, one of the four transform core candidates can be selected based on the transform core candidate index with two bins. That is, the transform core candidate indexes 0 to 3 can be signaled as 00, 01, 10, and 11, respectively. For these two bins, the MSB (most significant bit) can be signaled first, and the LSB (least significant bit) can be signaled later. Different CABAC contexts can be assigned to each bin. For example, a CABAC context other than the CABAC context assigned for the MTS flag can be assigned to each bin for two bins. Alternatively, bypass coding can be applied without assigning a CABAC context to two bins. When AbsSum is greater than T2, the transform core candidate index has a value from 0 to 5, so the transform core candidate index cannot be expressed using only two bins. In this case, the transform core candidate index can be expressed by assigning two or more bins, such as truncated binary coding. For each bin assigned by the truncated binary coding method, a CABAC context can be assigned, or bypass coding can be applied without assigning a CABAC context. Alternatively, the CABAC context can be assigned to some of the multiple bins (e.g., the first bin, or the first bin and the second bin), and bypass coding can be applied to the remaining bins.
[0173] Example 3
[0174] The transform core of the current block may be determined based on a transform set including one or more transform core candidates. The transform core of the current block may be derived as one of the one or more transform core candidates belonging to the transform set.
[0175] The process of determining the transform core of the current block may include at least one of the following: 1) a process of determining a transform set of the current block, or 2) a process of selecting a transform core candidate from the transform set of the current block. The process of determining the transform set may be a process of selecting one of the same predefined multiple transform sets in the encoding device and the decoding device. Alternatively, the process of determining the transform set may be a process of configuring one or more transform sets available for the current block from the same predefined multiple transform sets in the encoding device and the decoding device, and selecting one of the configured transform sets. Alternatively, the process of determining the transform set may be a process of configuring a transform set based on transform core candidates available for the current block from the same predefined multiple transform core candidates in the encoding device and the decoding device.
[0176] When the transform set of the current block includes a plurality of transform core candidates, a process of selecting one of the plurality of transform core candidates for the current block may be performed. However, when the transform set of the current block includes one transform core candidate (i.e., the number of transform core candidates available for the current block is 1), the transform core of the current block may be set to the corresponding transform core candidate.
[0177] The transform set according to the present disclosure may mean the (inseparable) transform set in the above-mentioned embodiment 1, or may mean the MTS set in embodiment 2. Alternatively, the transform set may be defined separately from the (inseparable) transform set in embodiment 1 or the MTS set in embodiment 2. In this case, the transform set may include one or more specific transform cores as transform core candidates. One specific transform core may be defined as a pair of a transform core for horizontal transform and a transform core for vertical transform, or may be defined as one transform core that is equally applied to horizontal and vertical transforms. In the following, the specific transform cores will be described in detail.
[0178] The specific transform core according to the present disclosure may be a transform core predefined identically in the encoding device and the decoding device. Alternatively, the specific transform core may further include a transform core derived based on the above-mentioned predefined transform core. Alternatively, the specific transform core may mean a transform core with a predetermined index in the (inseparable) transform set of embodiment 1 or the MTS set of embodiment 2.
[0179] A specific transform kernel may be defined as a combination of transform kernels based on trigonometric functions (e.g., DCT-2, DST-7, DCT-8, DCT-5, DST-4, DST-1). Alternatively, a specific transform kernel may be defined as a combination of transform kernels based on non-trigonometric functions. Here, examples of non-trigonometric function-based transform kernels may include KLT, SOT, orthogonal transform kernels, non-orthogonal transform kernels, and the like. KLT may be represented as a transform kernel trained with training feature data, i.e., a training-based transform kernel. Alternatively, a specific transform kernel may be defined as a combination of a transform kernel based on trigonometric functions and a transform kernel based on non-trigonometric functions.
[0180] For example, a specific transform kernel is represented as (T_h, T_v). T_h is a horizontal transform kernel, which may be a training-based transform kernel such as KLT. T_v is a vertical transform kernel, which may be a trigonometric function-based transform kernel such as DCT-2. Alternatively, T_h may be KLT, and T_v may be DST-7. Alternatively, when DST-7 and DCT-2 are allowed for T_h and KLT1 and KLT2 corresponding to KLT are allowed for T_v, four specific transform kernels may be defined by a combination of the allowed transform kernels.
[0181] A specific transform core or a transform set based on a specific transform core may be used in a manner that replaces the (inseparable) transform set of Embodiment 1 and / or the MTS set of Embodiment 2. Alternatively, it may be added as a transform set independent of the (inseparable) transform set of Embodiment 1 and / or the MTS set of Embodiment 2.
[0182] A flag may be defined to indicate whether a transform set based on a specific transform core is applied. When the flag is a first value, a transform set based on a specific transform core may be applied, and when the flag is a second value, the (inseparable) transform set of Example 1 or the MTS set of Example 2 may be applied. The flag may be signaled before a syntax element that specifies at least one transform core candidate in the transform set. For example, when the flag is a first value, a transform set based on a specific transform core may be applied, and an index indicating one of multiple specific transform cores belonging to the transform set may be additionally signaled. On the other hand, when the flag is a second value, the (inseparable) transform set of Example 1 or the MTS set of Example 2 may be applied, and an index indicating one of multiple transform core candidates (i.e., a transform core candidate index) may be additionally signaled.
[0183] A specific transform kernel may be applied only to the luma component of the current block, or may be applied to both the luma component and the chroma components of the current block.
[0184] A specific transform kernel according to the present disclosure may be defined as a transform kernel based on length. That is, the specific transform kernel may be one or more transform kernels having a predetermined length. In the following, for ease of explanation, a transform kernel having a length of K is represented as a length-K transform kernel. Here, K may be at least one of 4, 8, 16, 32, 64 or a greater integer. Here, the length may mean the length (width and / or height) of a side that a transform block may have. In the present disclosure, the width that a transform block may have may be referred to as an allowable width, and the height that a transform block may have may be referred to as an allowable height. When a separable transform is applied to a current block of size MxN, a length-M transform kernel having a length equal to the width of the current block may be applied in the horizontal direction, and a length-N transform kernel having a length equal to the height of the current block may be applied in the vertical direction.
[0185] By assigning transform kernels by length, separable transforms can be applied to MxN blocks of all sizes. For example, when the current block is 4xN, a transform kernel with a length equal to the width of the current block can be applied in the horizontal direction. When the current block is Nx4, a transform kernel with a length equal to the height of the current block can be applied in the vertical direction. Here, N can be 4, 8, 16, 32, 64 or a larger integer.
[0186] The length-based transform cores may be configured differently for the horizontal and vertical directions. That is, the length-4 horizontal transform core and the length-4 vertical transform core may be different from each other. When the number of allowable widths of the transform block is P and the number of allowable widths of the transform block is Q, a transform set may be configured with (P+Q) length-based transform cores. For example, when the allowable width and the allowable height of the transform block are 4, 8, 16, and 32, respectively, the length-4 transform core, the length-8 transform core, the length-16 transform core, and the length-32 transform core may be used for the horizontal and vertical directions. In this case, the value of P and the value of Q may both be 4, and a single transform set may be configured based on a total of eight length-based transform cores. In this case, the eight length-based transform cores may be a length-4 horizontal transform core, a length-8 horizontal transform core, a length-16 horizontal transform core, a length-32 horizontal transform core, a length-4 vertical transform core, a length-8 vertical transform core, a length-16 vertical transform core, and a length-32 vertical transform core.
[0187] Even if the transform kernel is applied in the same direction and has the same length, the transform kernel applied to the current block may also be different based on the size of the current block. Here, the size of the current block can be defined as one of the width, height, sum of width and height, product of width and height, or maximum / minimum value of width and height of the current block. A length-M horizontal transform kernel can be assigned for each allowable height of the transform block. Similarly, a length-N vertical transform kernel can be assigned for each allowable width of the transform block. Here, when both M and N can have values of 4, 8, 16, 32, different length-M horizontal transform kernels can be assigned to Mx4, Mx8, Mx16, and Mx32 blocks, respectively, and different length-N vertical transform kernels can be assigned to 4xN, 8xN, 16xN, and 32xN blocks, respectively. For example, a length-4 horizontal transform kernel applied to a 4x8 block may be different from a length-4 horizontal transform kernel applied to a 4x16 block.
[0188] Alternatively, the allowable width and the allowable height of the transform block may be divided into a plurality of groups, and the same length-based transform kernel may be assigned to each group.
[0189] For example, the allowable width and allowable height of the transform block can be divided into two groups based on a predetermined threshold. In an MxN block, when the value of N is less than or equal to a first threshold, a first length-M horizontal transform core can be assigned, and when the value of N is greater than the first threshold, a second length-M horizontal transform core can be assigned. Here, the first length-M horizontal transform core and the second length-M horizontal transform core can have the same length, but can be different transform cores. The first threshold can be an integer of 8, 16 or more. Similarly, in an MxN block, when the value of M is less than or equal to a second threshold, a first length-N vertical transform core can be assigned, and when the value of M is greater than the second threshold, a second length-N vertical transform core can be assigned. Here, the first length-N vertical transform core and the second length-N vertical transform core can have the same length, but can be different transform cores. The second threshold can be an integer of 8, 16 or more.
[0190] Alternatively, the allowable width and allowable height of the transform block can be divided into three or more groups based on at least two thresholds. Here, assuming that the allowable width and allowable height of the transform block are 4, 8, 16 and 32, they are divided into three groups, that is, {4,8}, {16} and {32}. When the value of N in the current block of MxN size is 4 and 8, the first length-M horizontal transform core can be assigned. When the value of N is 16, the second length-M horizontal transform core can be assigned. When the value of N is 32, the third length-M horizontal transform core can be assigned. Similarly, when the value of M is 4 and 8, the first length-N vertical transform core can be assigned. When the value of M is 16, the second length-N vertical transform core can be assigned. When the value of M is 32, the third length-N vertical transform core can be assigned. In this way, one of the first to third length-M horizontal transform cores can be applied based on the height of the current block, and one of the first to third length-N vertical transform cores can be applied based on the width of the current block. Nine specific transform cores may be configured by a combination of the above-mentioned first to third length-M horizontal transform cores and first to third length-N vertical transform cores, and one transform set for the current block may be configured with the nine specific transform cores.
[0191] Alternatively, the length-based transform kernel may not be configured separately for the horizontal and vertical directions. That is, a transform kernel having a length corresponding to the width or height may be applied regardless of the horizontal and vertical directions. For example, a length-4 transform kernel applied in the horizontal direction of a 4x8 block and a length-4 transform kernel applied in the vertical direction of an 8x4 block may be the same.
[0192] Regardless of the horizontal and vertical directions, when a length-M transform kernel having the same length as a side having a length of M is applied, only one transform kernel may be required for each length of the allowable width and / or the allowable height of the transform block. For example, when the allowable width or the allowable height of the transform block is 4, 8, 16, and 32, the transform set may be configured based on four length-based transform kernels. Here, the four length-based transform kernels may be a length-4 transform kernel, a length-8 transform kernel, a length-16 transform kernel, and a length-32 transform kernel. When the current block is an MxN block, a length-M transform kernel may be applied to the horizontal direction and a length-N transform kernel may be applied to the vertical direction.
[0193] Based on the context of the current block, the application of the length-based transform kernel can be limited, or the number of length-based transform kernels available for the current block can be different. Hereinafter, for ease of explanation, it is assumed that the length-based transform kernel is a non-trigonometric function-based transform kernel such as the above-mentioned KLT, but is not limited thereto.
[0194] A transformation kernel based on a non-trigonometric function may be applied to the direction of an edge having a specific length, while a transformation kernel based on a trigonometric function may be applied to another length. Here, the specific length may be determined based on a length predefined identically for the encoding device and the decoding device. A transformation kernel based on a non-trigonometric function may be applied to one of the horizontal direction or the vertical direction, and a transformation kernel based on a trigonometric function may be applied to the other direction.
[0195] For example, when the length of the edge of the current block is greater than or equal to 8, a transformation kernel based on a non-trigonometric function may be applied to the edge, otherwise, a transformation kernel based on a trigonometric function may be applied to the edge. Alternatively, when the length of the edge of the current block is less than 8, a transformation kernel based on a non-trigonometric function may be applied to the edge, otherwise, a transformation kernel based on a trigonometric function may be applied to the edge. Alternatively, when the length of the edge of the current block is less than or equal to 16, a transformation kernel based on a non-trigonometric function may be applied to the edge, otherwise, a transformation kernel based on a trigonometric function may be applied to the edge. Alternatively, when the length of the edge of the current block is greater than 16, a transformation kernel based on a non-trigonometric function may be applied to the edge, otherwise, a transformation kernel based on a trigonometric function may be applied to the edge. In this way, in the case of limiting the application of a transformation kernel based on a non-trigonometric function to an edge of length p, for a block whose one of width or height is p and the other is not p, the following two methods may be considered.
[0196] (1) A first method in which a transformation kernel based on a non-trigonometric function is not applied to a direction having a length p among width or height, and a transformation kernel based on a non-trigonometric function may be applied to a direction not having a length p.
[0197] (2) A second method, in which a non-trigonometric function based transformation kernel is not applied in the horizontal and vertical directions when the width or height has a length p.
[0198] For example, one may restrict the application of non-trigonometric function based transformation kernels of length -4.
[0199] According to the first method, in the case of a block whose width or height is 4, a non-trigonometric function-based transformation kernel of length -4 is not applied to the direction of its side whose length is 4, but a trigonometric function-based transformation kernel of length -4 may be applied. Meanwhile, for the direction of the side whose length is not 4, a non-trigonometric function-based transformation kernel corresponding to the corresponding length may be applied. That is, a non-trigonometric function-based transformation kernel of length -4 may not be applied to a 4x4 block. For a block of 4x8, 4x16, 4x32, 8x4, 16x4, or 32x4, a non-trigonometric function-based transformation kernel corresponding to the length may be applied only to the side whose length is not 4. For blocks of other sizes (e.g., 8x8, 8x16, 8x32, 16x8, 16x16, 16x32, 32x8, 32x16, 32x32, etc.), a non-trigonometric function-based transformation kernel corresponding to the length may be applied in the horizontal and vertical directions.
[0200] Alternatively, according to the second method, a non-trigonometric function-based transform kernel corresponding to the length may be applied only when both the width and the height of the current block are greater than 4. In other words, when the width or the height has a length of 4, a non-trigonometric function-based transform kernel of length -4 may not be applied for the horizontal and vertical directions.
[0201] Alternatively, application of non-trigonometric function based transformation kernels of length-4 and length-8 may be restricted.
[0202] According to the first method, for 4x4, 4x8, 8x4, and 8x8 blocks, non-trigonometric function-based transform kernels of length -4 and length -8 may not be applied, but trigonometric function-based transform kernels of length -4 and length -8 may be applied. For 4x16, 4x32, 8x16, 8x32, 16x4, 16x8, 32x4, and 32x8 blocks, non-trigonometric function-based transform kernels corresponding to the length may be applied only to the edges whose lengths are not 4 and 8. For blocks of other sizes (e.g., 16x16, 16x32, 32x16, 32x32, etc.), non-trigonometric function-based transform kernels corresponding to the length may be applied in the horizontal and vertical directions.
[0203] According to the second method, only when both the width and height of the current block are greater than 8 (for example, a 16x16, 16x32, 32x16, or 32x32 block, etc.), a non-triangular function-based transform kernel corresponding to the length may be applied. In other words, when the length of the width or height is 4 or 8, non-triangular function-based transform kernels of length-4 and length-8 may not be applied for the horizontal and vertical directions.
[0204] Alternatively, the application may be restricted to non-trigonometric function based transformation kernels of length -32.
[0205] According to the first method, for a 32x32 block, a non-trigonometric function-based transform kernel of length -32 may not be applied, but a trigonometric function-based transform kernel of length -32 may be applied. For 4x32, 8x32, 16x32, 32x4, 32x8, 32x16 blocks, a non-trigonometric function-based transform kernel corresponding to the length may be applied only to the side whose length is not 32. For blocks of other sizes (e.g., 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, 16x16), a non-trigonometric function-based transform kernel corresponding to the length may be applied to the horizontal and vertical directions.
[0206] According to the second method, only when both the width and height of the current block are less than 32, the non-trigonometric function-based transform kernel corresponding to the length can be applied. That is, for 4x32, 8x32, 16x32, 32x4, 32x8, 32x16, or 32x32 blocks, the non-trigonometric function-based transform kernel of length -32 may not be applied in the horizontal and vertical directions. For blocks of other sizes (e.g., 4x4, 4x8, 4x16, 8x4, 8x8, 8x16, 16x4, 16x8, 16x16), the non-trigonometric function-based transform kernel corresponding to the length may be applied in the horizontal and vertical directions.
[0207] The above embodiment may be adaptively performed based on the slice type of the current block. For example, when the slice type is an I slice, the above example may be applied, otherwise it may not be applied.
[0208] In addition, in the above-mentioned embodiment, the case where the application of the transformation kernel based on non-trigonometric functions is limited to a specific length is described, but this is only an example. For example, the application of the first type of transformation kernel based on trigonometric functions can be limited to a specific length. In this case, the second type of transformation kernel based on trigonometric functions can be applied. That is, in the above-mentioned embodiment, the transformation kernel based on non-trigonometric functions and the transformation kernel based on trigonometric functions can be understood as being replaced by the first type of transformation kernel based on trigonometric functions and the second type of transformation kernel based on trigonometric functions. The second type can be a heterogeneous transformation type different from the first type.
[0209] Example 4
[0210] The transform set may be determined based on the context of the current block. The transform set may be determined based on the intra prediction mode of the current block and may be composed of one or more transform core candidates. For example, the transform set of the present disclosure may be the (inseparable) transform set of embodiment 1 or the MTS set of embodiment 2. The transform set of the present disclosure may include one or more length-based transform cores with a predetermined length as transform core candidates, such as embodiment 3.
[0211] In the present disclosure, a mapping relationship between predefined intra prediction modes and transform sets may be defined.
[0212] The predefined intra prediction modes may include at least one of a planar mode, a DC mode, a directional mode, a WAIP mode (wide angle mode), a matrix-based intra prediction (MIP) mode, a decoder-side intra mode derivation (DIMD) mode, or a template-based intra mode derivation (TIMD) mode.
[0213] Each mode is assigned a mode value to identify the mode. 0 and 1 are assigned to the planar mode and the DC mode, respectively, and 2 to 66 are assigned to the directional modes, respectively. The wide angle mode is assigned values less than 0 and greater than 66, respectively.
[0214] A separate transform set may be mapped to each predefined intra prediction mode (or each mode value). Alternatively, the predefined intra prediction modes may be divided into at least two groups by a predetermined rule, and a transform set may be mapped to each group. Here, the predefined rule may be one of the rules (1) to (7) described below, or a combination of at least two of them. Hereinafter, a predetermined rule for grouping predefined intra prediction modes will be described.
[0215] (1) Planar mode and DC mode can be grouped into one group.
[0216] (2) Directional patterns can be divided into multiple groups by grouping adjacent patterns.
[0217] (3) Wide-angle modes may be grouped into one group. Alternatively, modes having values less than 0 may be grouped into one group, and modes having values greater than 66 may be grouped into another group. Alternatively, wide-angle modes may be divided into three or more groups by grouping adjacent modes, like directional modes. For example, because a wide-angle mode less than 0 is angularly adjacent to mode 2, a wide-angle mode less than 0 and mode 2 (or including mode 2 and at least one directional mode adjacent to mode 2) may be grouped into one group (Group A). Because a wide-angle mode greater than 66 is angularly adjacent to mode 66, a wide-angle mode greater than 66 and mode 66 (or including mode 66 and at least one directional mode adjacent to mode 66) may be grouped into one group (Group B). Group A and Group B may be merged into one group.
[0218] (4) In the case of MIP mode, it can be grouped together with planar mode, or a separate group can consist of MIP mode only. Alternatively, MIP mode can be considered as planar mode when mapping between intra prediction modes and transform sets.
[0219] (5) In the case of the DIMD mode, the separate group may consist of only the DIMD mode, or may be included in the group to which the intra-frame prediction mode derived based on the DIMD mode belongs. Alternatively, when multiple intra-frame prediction modes are derived based on the DIMD mode, it may also be included in the group to which one of the multiple intra-frame prediction modes (e.g., the first mode, the mode with the minimum value) belongs.
[0220] (6) In the case of TIMD mode, a separate group may be composed only of the TIMD mode, or may be included in the group to which the intra-frame prediction mode derived based on the TIMD mode belongs. Alternatively, when multiple intra-frame prediction modes are derived based on the TIMD mode, it may also be included in the group to which one of the multiple intra-frame prediction modes (e.g., the first mode, the mode with the smallest value) belongs. The value of the intra-frame prediction mode derived based on the TIMD mode may be given based on 131 modes. In this case, the mode value based on 131 modes may be converted into a mode value based on 67 modes. Any value from 0 to 130 may be mapped / converted to any value from 0 to 66 by a mapping function (MAP131TO67) according to the following formula 4. The TIMD mode may be included in the group to which the converted mode value belongs.
[0221] [Formula 4]
[0222] MAP131T067( model ) =(mode < 2? mode:((mode >> 1) + 1))
[0223] (7) In the case of intra template matching mode, it can be included in the group to which the planar mode belongs. In the mapping between intra prediction modes and transform sets, the intra template matching mode can be regarded as the planar mode.
[0224] The intra prediction modes predefined by the above-mentioned predetermined rule may be divided into a plurality of groups, and a transform set may be assigned to each group. The following is an example of a mapping table defining a transform set assigned to each intra prediction mode (or group).
[0225] [Table 8]
[0226] (1) 7-episodes
[0227] [Table 9]
[0228] (2) 4-episode
[0229] [Table 10]
[0230] (3) 3-episode
[0231] [Table 11]
[0232] (4) 2-episode
[0233] [Table 12]
[0234] (5) 1-episode
[0235] The transform set may be determined based on the symmetry between intra prediction modes and / or block shapes. Figure 5 As described above, the intra prediction mode, in particular the directional mode, may be symmetrical with respect to mode 34. According to the symmetry of the intra prediction mode with respect to mode 34, the mode symmetrical with mode p for the MxN block may be mode (68-p), and the mode symmetrical with mode q which is the wide-angle mode may be mode (66-q). In addition, due to the symmetry between block shapes, the block symmetrical with the MxN block may be an NxM block.
[0236] If the mode corresponding to mode r is mode s, the case where intra prediction is performed on an MxN block in mode r and the case where intra prediction is performed on an NxM block in mode s may be symmetrical to each other. Alternatively, the case where a horizontal transform kernel of length-M and a vertical transform kernel of length-N are applied to an MxN block may be symmetrical to the case where a vertical transform kernel of length-M and a horizontal transform kernel of length-N are applied to an NxM block. When it is desired to apply the same specific transform kernel to the above two cases, the horizontal transform kernel of length-M of the MxN block and the vertical transform kernel of length-M of the NxM block may be set to be the same for the above two cases, and the vertical transform kernel of length-N of the MxN block and the horizontal transform kernel of length-N of the NxM block may be set to be the same.
[0237] For example, when the intra prediction mode for an MxN block is a directional mode and mode p is greater than 34, a transform set mapped to mode (68-p) for an NxM block may be used. In this case, a vertical transform kernel of length-M and a horizontal transform kernel of length-N, which are transform kernel candidates belonging to the transform set, may be used as a horizontal transform kernel of length-M and a vertical transform kernel of length-N for the MxN block, respectively. When the above-mentioned symmetry is used, the same transform set may be used between modes that are symmetrical to each other, and the mapping between the intra prediction mode and the transform set may be symmetrically set as follows.
[0238] [Table 13]
[0239] (6) 4-sets: When applying the symmetry to (1) 7-sets
[0240] [Table 14]
[0241] (7) 2-sets: In the case of applying the symmetry of (3) to 3-sets
[0242] [Table 15]
[0243] (8) 1-set: When applying the symmetry to (5) 1-set
[0244] In the present disclosure, a transform set may include one or more transform core candidates, and the selected candidates among them may include a horizontal transform core and a vertical transform core. The transform core candidates may be assigned the length that one side of the transform block can have, respectively. In this case, the transform core of the MxN block predicted by mode r may be derived based on the symmetry between the above-mentioned intra prediction modes and / or block shapes. That is, the transform set may be determined based on a mode s symmetric to mode r, and the transform core of the current block may be derived from the transform core candidates belonging to the transform set. In this case, a horizontal transform core of length-N from the selected transform core candidate for the NxM block symmetric to the MxN block may be applied as a vertical transform core of length-N for the MxN block. Similarly, a horizontal transform core of length-M from the selected transform core candidate for the NxM block symmetric to the MxN block may be applied as a vertical transform core of length-M for the MxN block. However, as described in Example 3, when the transform set is composed of length-based transform cores that are also applied to the horizontal and vertical directions, the horizontal transform core is changed to a vertical transform core, but the transform core is not changed.
[0245] For the MIP mode, a flag may be defined to indicate whether prediction is performed in a transposed form, and the case where the flag is 0 may be symmetrical with the case where the flag is 1. A horizontal transform kernel of length -M in the horizontal direction applied to an MxN block when the flag is 0 may be set to be the same as a vertical transform kernel of length -M in the vertical direction applied to an NxM block when the flag is 1.
[0246] Alternatively, when the flag is 1, the horizontal transform kernel and the vertical transform kernel to be applied to the MxN block can be set to the horizontal transform kernel and the vertical transform kernel of the NxM block, respectively. For example, when DST-7 and DCT-5 are applied as the horizontal transform kernel and the vertical transform kernel of the NxM block, respectively, the horizontal transform kernel and the vertical transform kernel to be applied to the MxN block can be set to DST-7 and DCT-5, respectively. Only the transform type is the same, and the length of the transform kernel can be different. That is, while DST-7 of length-N can be applied as the horizontal transform kernel of the NxM block, DST-7 of length-M can be applied as the horizontal transform kernel of the MxN block.
[0247] For MIP mode, when a non-trigonometric function-based transform kernel (e.g., a trained transform kernel) is applied instead of a trigonometric function-based transform kernel, the non-trigonometric function-based transform kernel is applied in the same manner, but a transform kernel of a different length is applied, so that different non-trigonometric function-based transform kernels can be applied in each direction for MxN blocks and NxM blocks. For MIP mode, the horizontal transform kernel and the vertical transform kernel specified in the transform set can be used as is, regardless of whether transposition is applied.
[0248] Example 5
[0249] One or more specific transform cores may be configured for the current block. One or more transform sets may be configured for the current block based on the one or more specific transform cores. The transform core of the current block may be determined based on at least one of the multiple specific transform cores.
[0250] As specific transform kernels, a horizontal transform kernel (T_h) and a vertical transform kernel (T_v), a transform kernel based on a non-trigonometric function such as KLT, or a transform kernel based on a trigonometric function such as DCT-2 or DST-7 can be used. In this case, T_h is a horizontal transform kernel of length-p, and T_v is a vertical transform kernel of length-q. When p and q are the same, T_h and T_v can be the same transform kernel. Alternatively, even if p and q are the same, the horizontal transform kernel can be a different kernel from the vertical transform kernel.
[0251] The specific transform kernel may include a predetermined reference transform kernel. In addition, the specific transform kernel may include a transform kernel derived based on the reference transform kernel (hereinafter, referred to as a derived transform kernel). Here, the derived transform kernel may be generated by pairing the reference transform kernel with a transform kernel based on a trigonometric function. The reference transform kernel may be used to derive the derived transform kernel and may not be included in the specific transform kernel. Such a specific transform kernel may be used as a transform kernel candidate for the current block. Hereinafter, a method for determining a specific transform kernel according to the present disclosure will be described.
[0252] (Method 1) The reference transform kernel may be a transform kernel based on a non-trigonometric function. In the following, for ease of explanation, the reference transform kernel is assumed to be KLT and is represented as (KLT_h, KLT_v). For example, the reference transform kernel (KLT_h, KLT_v) may be paired with DCT-2. In this case, the specific transform kernel may include at least one of (KLT_h, KLT_v), (KLT_h, DCT-2), or (DCT-2, KLT_v). Alternatively, the reference transform kernel (KLT_h, KLT_v) may be paired with DST-7. In this case, the specific transform kernel may include at least one of (KLT_h, KLT_v), (KLT_h, DST-7), or (DST-7, KLT_v).
[0253] As described above, one transform set may be configured with three transform core candidates from one transform core. Alternatively, one transform set including one or more transform core candidates generated by pairing with DCT-2 may be configured, and another transform set including one or more transform core candidates generated by pairing with DST-7 may be configured. In this case, as in Embodiment 4, one of the two transform sets may be selected based on the intra prediction mode of the current block, and a transform core candidate may be selected from the corresponding transform set. To this end, an index indicating the corresponding transform core candidate may be sent by a signal, or one of the transform core candidates may be implicitly selected based on the context of the current block.
[0254] (Method 2) The reference transform core may be the first transform core candidate in the MTS set of the aforementioned embodiment 2. Here, the first transform core candidate in the MTS set is represented as (AMTS_h, AMTS_v). For example, the reference transform core (AMTS_h, AMTS_v) may be paired with DCT-2. In this case, the specific transform core may include at least one of (AMTS_h, AMTS_v), (AMTS_h, DCT-2), or (DCT-2, AMTS_v). Alternatively, the reference transform core (AMTS_h, AMTS_v) may be paired with DST-7. In this case, the specific transform core may include at least one of (AMTS_h, AMTS_v), (AMTS_h, DST-7), or (DST-7, AMTS_v). Alternatively, the reference transform core (AMTS_h, AMTS_v) may be paired with DST-4. In this case, the specific transform core may include at least one of (AMTS_h, AMTS_v), (AMTS_h, DST-4), or (DST-4, AMTS_v).
[0255] In the MTS set according to Embodiment 2, there is no case where AMTS_h and AMTS_v are DCT-2. Therefore, the three specific transform cores determined by pairing with DCT-2 may be different transform core candidates. However, when the transform set based on the specific transform core according to the present disclosure and the MTS set according to Embodiment 2 coexist in one codec system, (AMTS_h, AMTS_v) overlaps with the first transform core candidate of the MTS set. In this case, a flag may be introduced to distinguish between the transform based on the specific transform core and the transform based on the MTS set. In addition, if the codec system operates in a manner where the signaling cost for specifying (AMTS_h, AMTS_v) when the flag is 0 is different from the signaling cost for specifying (AMTS_h, AMTS_v) when the flag is 1, and thus a more favorable case is selected by comparing the RD (rate distortion) cost, then (AMTS_h, AMTS_v) may be included in the specific transform core. Alternatively, (AMTS_h, AMTS_v) may be excluded from a specific transform core because (AMTS_h, AMTS_v) is included as the first transform core candidate in the MTS set. In this case, (AMTS_h, DCT-2) or (DCT-2, AMTS_v) may be selected, and an index consisting of one bin may be used for this selection.
[0256] When AMTS_h or AMTS_v of the reference transform core is DST-7 (or DST-4), the transform core derived by pairing with DST-7 (or DST-4) can overlap with the reference transform core (AMTS_h, AMTS_v). In addition, when the transform set based on the specific transform core according to the present disclosure and the MTS set according to Embodiment 2 coexist in one codec system, (AMTS_h, AMTS_v) overlaps with the first transform core candidate of the MTS set. In this case, as described above, (AMTS_h, AMTS_v) can be included in the specific transform core, and (AMTS_h, AMTS_v) can be excluded from the specific transform core. When (AMTS_h, DST-7 (or DST-4)) and / or (DST-7 (or DST-4), AMTS_v) overlap with (AMTS_h, AMTS_v), the specific transform core can be configured based on the remaining candidates excluding the overlapping candidates. Finally, when only one transform core candidate is available, the transform core of the current block may be determined based on the transform core candidate without signaling a separate index.
[0257] (Method 3) For a block predicted in an intra prediction mode, (DST-7, DST-7) may be a valid transform kernel. Therefore, (DST-7, DST-7) may be used as a reference transform kernel, and at least one derived transform kernel may be generated by pairing with DCT-2. For example, at least one of (DST-7, DST-7) or (DCT-2, DST-7) may be generated by pairing the reference transform kernel (DST-7, DST-7) with DCT-2. In this case, the specific transform kernel may include at least one of (DST-7, DST-7), (DST-7, DCT-2), or (DCT-2, DST-7).
[0258] Similarly, (DST-7, DST-7) may overlap with the first transform core candidate in the MTS set. In this case, (DST-7, DST-7) may be included in or excluded from the specific transform core. When only (DST-7, DCT-2) and (DCT-2, DST-7) are determined as specific transform cores, a bin may be additionally signaled to select one of the two specific transform cores.
[0259] (Method 4) Based on the length of each side of the current block, a transform kernel for the direction of the corresponding side can be determined, and a specific transform kernel can be generated based on the determined transform kernel. For example, for the width and height of the current block, when the length is greater than or equal to 4 and less than or equal to 16, DST-7 can be selected, and otherwise DCT-2 can be selected. When the current block is a 32x8 block, DCT2 can be selected for the horizontal direction of the current block, DST-7 can be selected for the vertical direction of the current block, and a specific transform kernel of (DCT-2, DST-7) can be generated.
[0260] (Method 5) In addition to the specific transformation kernel according to the above-mentioned method 4, the specific transformation kernel according to the present disclosure may further include a predefined default transformation kernel. Here, the predefined default transformation kernel may be a reference transformation kernel in at least one of methods 1 to 3. That is, the default transformation kernel may include at least one of a transformation kernel based on a non-trigonometric function, the first transformation kernel candidate in the MTS set, or (DST-7, DST-7). When the default transformation kernel overlaps with the first transformation kernel candidate in the MTS set, the default transformation kernel may be included in the specific transformation kernel or excluded from the specific transformation kernel.
[0261] Meanwhile, in the transform set including the specific transform core configured according to method 5, the specific transform core according to method 4 may be added after the default transform core is added to the transform set. When the specific transform core according to method 4 overlaps with the default transform core, the specific transform core according to method 4 may not be added to the transform set. Alternatively, the default transform core may be added after the specific transform core according to method 4 is added to the transform set. When the specific transform core according to method 4 overlaps with the default transform core, the default transform core may not be added to the transform set.
[0262] In the above embodiment, it has been described that up to three specific transform cores can be determined from one reference transform core (T_h, T_v). However, this is merely an example, and multiple reference transform cores may be defined. Here, the multiple reference transform cores may include reference transform cores according to at least two of the above methods 1 to 3. When the number of reference transform cores available for the current block is N, the corresponding reference transform cores may be represented as (T_h1, T_v1), (T_h2, T_v2), ..., (T_hN, T_vN). As described above, for each reference transform core, up to three specific transform cores (e.g., (T_hi, T_vi), (T_hi, DCT-2), (DCT-2, T_vi)) may be generated by pairing with DCT-2, DST-7, or DST-4, etc., and some redundant transform cores may be removed. Each transform set may be configured with (T_hi, T_vi)s (i = 1, 2, ..., N), or may be configured with up to three specific transform cores determined based on (T_hi, T_vi). For example, up to three specific transform kernels determined based on (T_hi, T_vi) may be S_i_1 = {(T_hi, T_vi)}, S_i_2 = {(T_hi, T_vi), (T_hi, DCT-2), (DCT-2, T_vi)}, S_i_3 = {(T_hi, T_vi), (T_hi, DCT-2)}, S_i_4 = {(T_hi, T_vi), (DCT-2, T_vi)}, S_i_5 = {(T_hi, DCT-2), (DCT-2, T_vi)}, S_i_6 = {(T_hi, DCT-2)}, S_i_7 = {(DCT-2, T_vi)}. This example is based on pairing with DCT-2, but can also be applied to pairing with other trigonometric function-based transform kernels such as DST-7 or DST-4.
[0263] The entire transform core candidate can be configured by configuring a transform set based on the corresponding S_i_pi (where pi is one of 1, 2, 3, 4, 5, 6, and 7) for each (T_hi, T_vi). For example, when (KLT_h1, KLT_v1) and (KLT_h2, KLT_v2) are used as (T_h1, T_v1) and (T_h2, T_v2), a transform set can be generated, each of which includes a specific transform core of the following (1) to (4).
[0264] (1) (KLT_h1, KLT_v1), (KLT_h2, KLT_v2)
[0265] (2) (KLT_h1, KLT_v1), (KLT_h1, DCT-2), (DCT-2, KLT_v1), (KLT_h2, KLT_v2), (KLT_h2, DCT-2), (DCT-2, KLT_v2)
[0266] (3) (KLT_h1, KLT_v1), (KLT_h2, KLT_v2), (KLT_h2, DCT-2), (DCT-2, KLT_v2)
[0267] (4) (KLT_h1, KLT_v1), (KLT_h1, DCT-2), (DCT-2, KLT_v1), (KLT_h2, KLT_v2)
[0268] As seen in the above example, one or more specific transform kernels can be generated by pairing with a transform kernel based on a trigonometric function such as DCT-2 or DST-7. Similarly, when there are N reference transform kernels, one or more specific transform kernels can be generated based on each reference transform kernel. In addition, depending on the transform set, the reference transform kernel may be included or not included.
[0269] Depending on the context of the current block, the number of transform core candidates available for the current block may vary. For example, depending on the absolute value sum (AbsSum) of all or part of the transform coefficients in the current block, the number of transform core candidates (or the size of the transform set) available for the current block may be set differently. For ease of explanation, this example assumes that the range of values that AbsSum can have is divided into three ranges and the maximum number of available transform core candidates is three. Here, each transform core candidate is represented as c1, c2, and c3, respectively.
[0270] The range of values that AbsSum can have can be divided into a first range (Range_1) in which the sum of the absolute values of the transform coefficients is greater than or equal to 0 and less than or equal to 6, a second range (Range_2) in which the sum of the absolute values of the transform coefficients is greater than 6 and less than or equal to 32, and a third range (Range_3) in which the sum of the absolute values of the transform coefficients is greater than 32. In this case, the number / position of available transform kernel candidates for each of the first to third ranges may be as shown in Table 16 below.
[0271] [Table 16]
[0272] Referring to Table 16, when the sum of the absolute values of the transform coefficients falls within Range_1, the number of transform core candidates available for the current block may be 1, otherwise (ie, when falling within Range_2 or Range_3), the number of transform core candidates available for the current block may be 3 (Case 1).
[0273] Alternatively, when the sum of the absolute values of the transform coefficients falls within Range_1 or Range_2, the number of transform core candidates available for the current block may be 1, otherwise (ie, when falling within Range_3), the number of transform core candidates available for the current block may be 3 (Case 2).
[0274] Alternatively, when the sum of the absolute values of the transform coefficients falls within Range_1, there may be no transform core candidates available for the current block, otherwise (ie, when falling within Range_2 or Range_3), the number of transform core candidates available for the current block may be 3 (Case 3).
[0275] Alternatively, when the sum of the absolute values of the transform coefficients falls within Range_1, there may be no transform core candidates available for the current block. When the sum of the absolute values of the transform coefficients falls within Range_2, the number of transform core candidates available for the current block may be 1. Otherwise (ie, when falling within Range_3), the number of transform core candidates available for the current block may be 3 (Case 4).
[0276] When the number of transform core candidates available for the current block is 1, the transform core candidate may be a candidate with a minimum index among the three transform core candidates. When there is no transform core candidate available for the current block, a transform core other than a specific transform core may be applied, or (inverse) transform may be skipped.
[0277] When (T_h, T_v) is (KLT_h, KLT_v), the specific transform core as the transform core candidate may be (KLT_h, KLT_v), (KLT_h, DCT-2), (DCT-2, KLT_v). As shown in Table 16, when only one transform core candidate is available for a specific range and the specific transform core is applied, the signaling of the index indicating the one transform core candidate may be omitted. In addition, as shown in Table 16, the specific transform core may not be applied to the specific range. In this case, the signaling of the flag indicating whether the specific transform core is applied may also be omitted.
[0278] When the range is distinguished based on the sum of the absolute values of the transform coefficients, the larger the sum of the absolute values of the transform coefficients, the more diverse the statistical characteristics included in the transform coefficient information, so that a larger number of specific transform cores can be assigned to improve coding performance. On the other hand, the smaller the sum of the absolute values of the transform coefficients, the simpler the statistical characteristics, so from the perspective of the trade-off between the signaling cost of the index and the performance improvement, reducing the number of specific transform cores may be advantageous.
[0279] (Method 6) Four specific transform cores can be determined from a combination of two specific transform cores. Here, the two specific transform cores can be reference transform cores of at least one of the above methods 1 to 3. Alternatively, the two specific transform cores can be specific transform cores pre-added to the transform set for the current block. For example, in the specific transform cores (T_h, T_v) for the current block of size MxN, transform cores H1 and H2 of length-M can be used as T_h, and transform cores V1 and V2 of length-N can be used as T_v. In this case, four specific transform cores can be configured by a combination of these, such as (H1, V1), (H1, V2), (H2, V1), and (H2, V2).
[0280] The transform core of the current block can be determined based on one of four specific transform cores. In this case, the index indicating one of the four specific transform cores can be signaled through the bitstream or implicitly derived based on the context of the current block. Here, a truncated unary code or a fixed length code with two bins can be used as a binarization method for indexing. When the index is assigned based on a fixed length code, it becomes 00, 01, 10, and 11. At least one of bypass coding or CABAC context-based coding can be applied to the bin.
[0281] The above example is based on the case where two transform cores are used as T_h and T_v, respectively, but is not limited thereto. When p and q transform cores are available as T_h and T_v, respectively, it is possible to configure by combining them until (p q) specific transformation kernels. These (p q) to configure a transform set by all or part of a specific transform core. The index used to select a specific transform core from a given transform set can be signaled via the bitstream or implicitly derived based on the context of the current block. Here, a truncated unary code or a fixed-length code with two bins can be used as a binarization method for the index. When the index is assigned based on the fixed-length code, it becomes 00, 01, 10, and 11. At least one of bypass coding or CABAC context-based coding can be applied to the bin.
[0282] The transform core of the current block may be determined based on any one of the above-mentioned embodiments 1 to 5. Alternatively, within the scope that the inventions of the above-mentioned embodiments 1 to 5 do not conflict with each other, the transform core of the current block may be determined based on a combination of at least two of the embodiments 1 to 5. In embodiment 5, a specific transform core may be configured based on one of methods 1 to 6, or may be configured based on a combination of at least two of methods 1 to 6.
[0283] refer to Figure 4 , the current block may be reconstructed based on the residual samples of the current block (S420).
[0284] The prediction samples of the current block may be derived based on the intra prediction mode of the current block. The reconstructed samples of the current block may be generated based on the prediction samples and the residual samples of the current block.
[0285] Figure 6 The diagram shows a schematic configuration of a decoding device (300) for performing an image decoding method according to the present disclosure.
[0286] refer to Figure 6 According to the present disclosure, the decoding device (300) may include a transform coefficient deriver (600), a residual sample deriver (610) and a reconstructed block generator (620). The transform coefficient deriver (600) may be configured in Figure 3 In the entropy decoder (310) of FIG. 1 , the residual sample deriver (610) may be configured in Figure 3 The residual processor (320) and the reconstructed block generator (620) may be configured in Figure 3 The adder (340) is used as the input device.
[0287] The transform coefficient deriver (600) can obtain residual information of the current block from the bitstream and decode it to derive the transform coefficient of the current block.
[0288] The residual sample deriver (610) may derive the residual sample of the current block by performing at least one of dequantization or inverse transformation on the transform coefficients of the current block.
[0289] The residual sample deriver (610) can determine the transform kernel for inverse transform of the current block by a predetermined transform kernel determination method, and derive the residual sample of the current block based on the transform kernel determination method. Figure 4 As described above, and its detailed description will be omitted here.
[0290] The reconstructed block generator (620) may reconstruct the current block based on the residual samples of the current block.
[0291] Figure 7 The diagram illustrates an image encoding method performed by an encoding device (200) according to an embodiment of the present disclosure.
[0292] refer to Figure 7 , the residual samples of the current block can be derived (S700).
[0293] The residual samples of the current block may be derived by subtracting the predicted samples from the original samples of the current block. Here, the predicted samples may be derived based on a predetermined intra prediction mode.
[0294] refer to Figure 7 , a transform coefficient of the current block may be derived by performing at least one of transform or quantization on the residual samples of the current block ( S710 ).
[0295] The method for determining the transformation kernel used for the transformation is as shown in reference Figure 4 That is, the transformation kernel used for transformation may be determined based on at least one of the above-mentioned embodiments 1 to 5.
[0296] For example, as described in Embodiment 1, one or more transform sets for the transform of the current block may be defined / configured, and each transform set may include one or more transform core candidates. In this case, one from the multiple transform sets may be selected as the transform set of the current block. One of the multiple transform core candidates belonging to the transform set of the current block may be selected. The selection may be performed implicitly based on the context of the current block. Alternatively, the best transform set and / or transform core candidate for the current block may be selected, and an index indicating it may be signaled.
[0297] Alternatively, as in Embodiment 2, the transform core of the current block may be determined based on an MTS set. One of multiple MTS sets may be selected based on at least one of the size of the current block or the intra prediction mode. The selected MTS set may include one or more transform core candidates. One of the one or more transform core candidates may be selected, and the transform core of the current block may be determined based on the selected transform core candidate. The selection of the transform core candidate may be performed using a transform core candidate index derived based on the context of the current block. Alternatively, the best transform core candidate for the current block may be selected, and a transform core candidate index indicating the selected transform core candidate may be sent by a signal.
[0298] Alternatively, as in Embodiment 3, the transform kernel of the current block may be determined based on a transform set consisting of one or more specific transform kernels.
[0299] Alternatively, as in Embodiment 4, the transform kernel of the current block may be determined based on a mapping table defining a mapping relationship between intra prediction modes and transform sets. Here, the mapping table may be defined by considering symmetry between intra prediction modes and / or block shapes.
[0300] Alternatively, as in Embodiment 5, a specific transform core may be configured based on a combination of one or at least two of Methods 1 to 6, and a transform core of the current block may be determined from among the configured specific transform cores.
[0301] Alternatively, the transform kernel of the current block may be determined based on a combination of at least two of Embodiments 1 to 5.
[0302] refer to Figure 7 , a bitstream may be generated by encoding the transform coefficient of the current block ( S720 ).
[0303] Residual information about the transformation coefficient may be generated based on the transformation coefficient of the current block, and a bitstream may be generated by encoding the residual information.
[0304] Figure 8 The diagram shows a schematic configuration of an encoding device (200) for performing an image encoding method according to the present disclosure.
[0305] refer to Figure 8 According to the present disclosure, the encoding device (200) may include a residual sample deriver (800), a transform coefficient deriver (810) and a transform coefficient encoder (820). The residual sample deriver (800) and the transform coefficient deriver (810) may be configured in Figure 2 The residual processor (230) and the transform coefficient encoder (820) may be configured in Figure 2 In the entropy encoder (240).
[0306] The residual sample deriver (800) may derive the residual sample of the current block by subtracting the prediction sample from the original sample of the current block. Here, the prediction sample may be derived based on a predetermined intra prediction mode.
[0307] The transform coefficient deriver (810) may derive the transform coefficient of the current block by performing at least one of transform or quantization on the residual samples of the current block. The transform coefficient deriver (810) may determine the transform kernel of the current block based on one or a combination of at least two of the above-mentioned embodiments 1 to 5, and may derive the transform coefficient by applying the transform kernel to the residual samples of the current block.
[0308] The transform coefficient encoder (820) may encode the transform coefficient of the current block to generate a bitstream.
[0309] In the above embodiments, the method is described as a series of steps or boxes based on the flowchart, but the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in a different order than other steps described above. In addition, those skilled in the art will appreciate that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of the present disclosure.
[0310] The above-mentioned method according to the embodiment of the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0311] In the present disclosure, when the embodiment is implemented as software, the above method can be implemented as a module (process, function, etc.) that performs the above functions. The module can be stored in a memory and can be executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, a logic circuit, and / or a data processing device. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. In other words, the embodiments described herein can be implemented on a processor, a microprocessor, a controller, or a chip. For example, the functional unit shown in each of the figures can be implemented on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information (e.g., information about instructions) or an algorithm for implementation can be stored in a digital storage medium.
[0312] In addition, the decoding device and the encoding device of the embodiment of the present disclosure can be included in multimedia broadcast sending and receiving devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, devices for providing video on demand (VoD) services, over-the-top video (OTT) devices, devices for providing Internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, videophone video devices, transportation terminals (e.g., vehicle (including autonomous driving vehicle) terminals, aircraft terminals, ship terminals, etc.) and medical video devices, etc., and can be used to process video signals or data signals. For example, over-the-top video (OTT) devices may include game consoles, Blu-ray players, networked TVs, home theater systems, smart phones, tablet computers, digital video recorders (DVRs), etc.
[0313] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to an embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BD), universal serial buses (USB), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, tapes, floppy disks, and optical media storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmitted via the Internet). In addition, the bit stream generated by the encoding method may be stored in a computer-readable recording medium or may be sent via a wired or wireless communication network.
[0314] In addition, the embodiments of the present disclosure may be implemented by a computer program product through a program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0315] Fig. 9 An example of a content streaming system to which an embodiment of the present disclosure can be applied is shown.
[0316] refer to Fig. 9 The content streaming transmission system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming transmission server, a web server, a media storage, a user device, and a multimedia input device.
[0317] The encoding server generates a bitstream by compressing content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0318] A bitstream may be generated by applying the encoding method or the bitstream generating method of the embodiment of the present disclosure, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0319] The streaming server sends multimedia data to the user device through the web server based on the user's request, and the web server serves as a medium to inform the user what services are available. When the user requests the required service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the command / response between each device in the content streaming system.
[0320] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a certain period of time.
[0321] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays (HMDs), digital televisions, desktops, digital signage, etc.).
[0322] Each server in the content streaming system may be operated as a distributed server, and in this case, data received from each server may be distributed and processed.
[0323] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined and implemented as a device, and the technical features of the device claims of the present disclosure can be combined and implemented as a method. In addition, the technical features of the method claims of the present disclosure and the technical features of the device claims can be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims can be combined and implemented as a method.
Claims
1. An image decoding method, comprising: deriving transform coefficients of the current block from the bitstream; performing at least one of inverse quantization or inverse transformation on the transform coefficients of the current block to derive residual samples of the current block, wherein a transform kernel of the current block used for the inverse transformation is determined from a transform set including a plurality of transform kernel candidates, and the plurality of transform kernel candidates include a reference transform kernel for the current block; as well as The current block is reconstructed based on the residual samples of the current block.
2. The image decoding method according to claim 1, wherein: The plurality of transform kernels further include transform kernel candidates derived by pairing the reference transform kernel with a predetermined trigonometric function-based transform kernel.
3. The image decoding method according to claim 2, wherein: The reference transformation kernel is a transformation kernel based on a non-trigonometric function.
4. The image decoding method according to claim 2, wherein: The reference transform kernel is one of a plurality of transform kernel candidates belonging to a multiple transform selection (MTS) set.
5. The image decoding method according to claim 2, wherein: The reference transform kernel is a transform kernel in which the horizontal transform kernel and the vertical transform kernel are DST-7.
6. The image decoding method according to claim 1, wherein: The plurality of transform kernel candidates include transform kernel candidates derived based on at least one of a width or a height of the current block.
7. The image decoding method according to claim 1, wherein: The number of transform kernel candidates included in the transform set is adaptively determined based on the sum of absolute values of all or some transform coefficients in the current block.
8. The image decoding method according to claim 1, wherein: The transform set further includes a transform core candidate derived based on a combination of two transform core candidates pre-added to the transform set.
9. An image encoding method, comprising: Export the residual samples of the current block; deriving a transform coefficient of the current block by performing at least one of transform or quantization on the residual samples of the current block, wherein a transform kernel of the current block used for the transform is determined from a transform set including a plurality of transform kernel candidates, and the plurality of transform kernel candidates include a reference transform kernel for the current block; and The transform coefficients of the current block are encoded.
10. A computer-readable storage medium storing a bit stream generated by the image encoding method according to claim 9.
11. A method for sending data, comprising: Obtaining a bitstream for image information, wherein the bitstream is generated by: deriving residual samples of a current block, deriving transform coefficients of the current block by performing at least one of transform or quantization on the residual samples of the current block, and encoding the transform coefficients of the current block; and sending said data comprising said bit stream, wherein a transform kernel of the current block used for the transform is determined from a transform set including a plurality of transform kernel candidates, and The multiple transform kernel candidates include a reference transform kernel for the current block.