Transform-based image encoding method and apparatus therefor
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LX SEMICON CO LTD
- Filing Date
- 2025-01-03
- Publication Date
- 2026-08-04
AI Technical Summary
因此,当图像数据通过介质诸如常规有线或无线宽带网络被发送时或者当图像/视频数据使用常规存储介质被存储时,发送成本和存储成本增加
[0019] According to the embodiments of this disclosure, the overall video/image compression efficiency can be improved.
Smart Images

Figure CN122514964A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image / video encoding method and an image / video encoding device. Background Technology
[0002] Image / video encoding is used in a variety of applications, such as digital storage media, television broadcasting, video streaming services, and real-time communications, and in recent years, the demand for high-resolution and high-quality images / videos has increased across various fields.
[0003] As images / videos become higher resolution and higher quality, their data size increases, and the amount of information, or bits, to be transmitted increases accordingly. Therefore, transmission and storage costs increase when image data is transmitted via media such as conventional wired or wireless broadband networks, or when image / video data is stored using conventional storage media.
[0004] In addition, interest and demand for immersive media such as VR (virtual reality), AR (augmented reality), MR (mixed reality) content and holograms have been increasing recently, and such immersive media are being used more and more in fields such as gaming, education, healthcare, real estate and marketing to provide immersive experiences.
[0005] Therefore, there is a need for efficient image / video compression technology that can efficiently compress, transmit, store, and reproduce high-resolution, high-quality image / video information with the aforementioned characteristics. Summary of the Invention
[0006] Solution
[0007] According to embodiments of this disclosure, methods and apparatus for improving video / image coding efficiency are provided.
[0008] According to embodiments of this disclosure, a transformation-based video / image coding method and apparatus are provided.
[0009] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided. The image decoding method includes: receiving image information including residual information for a current block; deriving transform coefficients for the current block based on the residual information; performing an inverse transform based on the transform coefficients to derive residual samples for the current block; and generating reconstructed samples for the current block based on the residual samples, wherein the image information includes MTS (Multiple Transform Selection) information; determining a transform pair from a set of transform pair candidates for the current block based on the MTS information; and the number of transform pair candidates in the transform set is at least one.
[0010] According to embodiments of this disclosure, an image encoding method performed by an encoding device is provided. The image encoding method includes: deriving a prediction sample for a current block; deriving a residual sample for the current block based on the prediction sample; performing a transform based on the residual sample to derive transform coefficients for the current block; generating residual information for the current block based on the transform coefficients; and encoding image information including the residual information, wherein the image information includes MTS (Multiple Transform Selection) information; determining a transform pair from a set of transform pair candidates for the current block based on the MTS information; and the number of transform pair candidates in the transform set is at least one.
[0011] According to embodiments of this disclosure, a decoding apparatus for image decoding is provided. The decoding apparatus includes a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform: receiving image information including residual information for a current block; deriving transform coefficients for the current block based on the residual information; performing an inverse transform based on the transform coefficients to derive residual samples for the current block; and generating reconstructed samples for the current block based on the residual samples, wherein the image information includes MTS (Multiple Transform Selection) information; determining a transform pair from a set of transform pair candidates for the current block based on the MTS information; and the number of transform pair candidates in the transform set is at least one.
[0012] According to embodiments of this disclosure, an encoding apparatus for image encoding is provided. The encoding apparatus includes a memory and at least one processor connected to the memory, wherein the at least one processor is configured to perform: deriving a prediction sample for a current block; deriving a residual sample for the current block based on the prediction sample; performing a transform based on the residual sample to derive transform coefficients for the current block; generating residual information for the current block based on the transform coefficients; and encoding image information including the residual information, wherein the image information includes MTS (Multiple Transform Selection) information; determining a transform pair from a set of transform pair candidates for the current block based on the MTS information; and the number of transform pair candidates in the transform set is at least one.
[0013] According to embodiments of this disclosure, a method is provided for transmitting video / image data including bitstreams generated by at least one of the video / image encoding methods disclosed herein.
[0014] According to embodiments of this disclosure, an apparatus is provided for transmitting video / image data including bitstreams generated by at least one of the video / image encoding methods disclosed herein.
[0015] According to embodiments of the present disclosure, a computer-readable storage medium is provided for storing a program for performing at least one of the methods according to embodiments of the present disclosure.
[0016] According to embodiments of the present disclosure, a computer-readable digital storage medium is provided for storing encoded video / image information generated by at least one of the embodiments of the video / image encoding methods disclosed herein.
[0017] According to embodiments of this disclosure, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information, said encoded information or encoded video / image information causing a decoding device to perform at least one of the embodiments of the video / image decoding method disclosed herein.
[0018] Beneficial effects
[0019] According to the embodiments of this disclosure, the overall video / image compression efficiency can be improved.
[0020] According to embodiments of this disclosure, the transformation performance of the current block can be improved.
[0021] According to embodiments of this disclosure, relevant information can be efficiently communicated using signals.
[0022] According to embodiments of this disclosure, transformation efficiency can be improved by determining the MTS transform set based on various intra-frame prediction modes.
[0023] According to embodiments of this disclosure, transformation efficiency can be improved by determining the transformation set for the current block based on an intra-prediction mode derived from a template.
[0024] According to embodiments of this disclosure, transformation efficiency can be improved by efficiently determining transformation pairs by signaling transformation-related information based on the number of transformation sets. Attached Figure Description
[0025] Figure 1 Examples of video / image coding systems to which embodiments of the present disclosure may be applied are illustrated schematically.
[0026] Figure 2 This is a diagram that schematically illustrates the configuration of a video / image encoding device to which embodiments of the present disclosure may be applied.
[0027] Figure 3This is a diagram that schematically illustrates the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied.
[0028] Figure 4 An example of the intra-frame prediction process is illustrated.
[0029] Figure 5 Examples of video / image coding methods based on intra-frame prediction are shown.
[0030] Figure 6 An example of a video / image decoding method based on intra-frame prediction is shown.
[0031] Figure 7 An example of a video / image coding method based on residual processing is shown.
[0032] Figure 8 An example of a video / image decoding method based on residual processing is shown.
[0033] Figure 9 Multiple transformation techniques according to this disclosure are illustrated by way of example.
[0034] Figure 10 This is a diagram illustrating an example of LFNST.
[0035] Figure 11 An NSPT core based on block size is illustrated as an example.
[0036] Figure 12 An ROI for LFNST 16 is illustrated as an example.
[0037] Figure 13 An ROI for LFNST 8 is illustrated as an example.
[0038] Figure 14 An exemplary MIP prediction sample for constructing HoG is illustrated.
[0039] Figure 15 Context modeling for the LFNST and NSPT transform coefficients is illustrated exemplarily.
[0040] Figure 16 An example of deriving the MTS set is shown.
[0041] Figure 17 An example is given of deriving DIMD-based intra-frame modes for MTS and LFNST.
[0042] Figure 18 A video / image coding method according to one or more embodiments of the present disclosure is illustrated schematically.
[0043] Figure 19A video / image decoding method according to one or more embodiments of the present disclosure is illustrated schematically. Detailed Implementation
[0044] Because this disclosure can have various modifications and implementations, specific implementations are illustrated in the accompanying drawings and will be described in detail. However, it should be understood that there is no intention to limit the implementations of this disclosure to that specific implementation. The terminology used herein is for the purpose of describing specific implementations only and is not intended to limit the technical spirit of this disclosure. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise. As used herein, the term "and / or" includes any and all combinations of the associated listed items. As used herein, the terms "comprising," "including," and "having" specify the presence of the stated features, quantities, operations, elements, components, and / or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, elements, components, and / or combinations thereof. In this disclosure, the use of the term "may" (e.g., regarding what the example or implementation may include or implement) in conjunction with an example or implementation indicates the existence of at least one example or implementation that includes or implements that feature, but all examples are not limited thereto, and the corresponding feature or configuration may be omitted.
[0045] Each component in the accompanying drawings described in this disclosure is shown independently to facilitate the explanation of its different features and functions, which does not imply that each component is implemented as separate hardware or separate software. For example, two or more of these components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which components are integrated and / or separated are also included within the scope of this disclosure without departing from its spirit.
[0046] In this disclosure, “A or B” can mean “A only”, “B only”, or “both A and B”. In other words, “A or B” can be interpreted in this disclosure as “A and / or B”. For example, “A, B or C” in this disclosure can mean “A only”, “B only”, “C only”, or “any and all combinations of A, B and C”.
[0047] As used in this article, a forward slash ( / ) or a comma can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0048] In this disclosure, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, the expressions "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0049] In this disclosure, "at least one of A, B, and C" can mean "A only", "B only", "C only", or "any and all combinations of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0050] The brackets used in this disclosure may mean "for example". Specifically, when indicated as "prediction (intra-frame prediction)", "intra-frame prediction" can be presented as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" can be presented as an example of "prediction". Furthermore, even when indicated as "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be presented as an example of "prediction".
[0051] In this disclosure, the technical features described in a single figure may be implemented independently or simultaneously.
[0052] This disclosure relates to video / image coding. For example, the methods / implementations described in this disclosure can be applied to methods disclosed in ECM (Enhanced Compression Model) or the H.267 standard. Furthermore, the methods / implementations disclosed herein can be applied to methods disclosed in the AV2 (AOMedia Video 2) standard or next-generation video / image coding standards (e.g., H.268, H.269, etc.).
[0053] In this disclosure, encoding may include encoding and / or decoding. In this disclosure, image encoding may be used interchangeably with video encoding.
[0054] In this disclosure, video can refer to a collection of images over time. An image typically refers to a unit representing a single image at a specific point in time, and a slice / tile is a unit that constitutes part of an image in encoding. A slice / tile may include one or more CTUs (Coding Tree Units). An image may consist of one or more slices / tiles. A tile may represent a rectangular area of a CTU within a specific tile row and column of an image.
[0055] Furthermore, an image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices within the image.
[0056] A pixel or cell can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the corresponding term to a pixel. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0057] A unit can represent the basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance blocks (e.g., a Cb block and a Cr block). Depending on the context, the term unit may be used interchangeably with terms such as block or region. Typically, an M×N block may include a set (or array) of samples or transform coefficients arranged in M columns and N rows.
[0058] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. In the following description, the same reference numerals may be used for the same components in the drawings, and redundant descriptions of the same components may be omitted.
[0059] Figure 1 Examples of video / image coding systems to which embodiments of the present disclosure may be applied are illustrated schematically.
[0060] refer to Figure 1 A video / image encoding system may include a first device (encoding device) and a second device (decoding device). The first device may deliver encoded video / image information or data to the second device in the form of a file or stream via a digital storage medium or network.
[0061] A video / image encoding system may also include a video / image acquisition device and a video / image renderer. The video / image acquisition device may be included in the encoding device, or it may be configured as a separate device or an external component. The video / image renderer may be included in the decoding device, or it may be configured as a separate device or an external component.
[0062] The first device may include a transmitter as an internal component, or a transmitter as a separate device or an external component.
[0063] The second device may include a receiver as an internal component, or as a receiver as a separate device or an external component.
[0064] An encoding device can be called an encoder, and a decoding device can be called a decoder. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. A renderer can include a display, and the display can be configured as a separate device or an external component.
[0065] The decoding and encoding devices used in one or more embodiments of this disclosure can be included in multimedia broadcasting transmitters / receivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicle terminals), aircraft terminals, and ship terminals), and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0066] A video / image acquisition device can acquire video / image sources. The video / image acquisition device can acquire video / images through processes of capturing, compositing, or generating video / images. The video / image acquisition device may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras and a video / image archive including previously captured video / images. The video / image generation device may include, for example, a camera, a computer, a tablet PC, and a smartphone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer, in which case the video / image capture process can be replaced by a process of generating relevant data. The video / image source can perform a video / image preprocessing process to input the optimized video / image into the encoder.
[0067] Encoding devices can encode input video / images. They can perform a series of processes, such as prediction, transformation, and quantization, to achieve compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0068] The transmitter can deliver encoded images / image information or data, output in bitstream form, to the receiver of the receiving device via digital storage media or a network, either as a file or a stream. The encoded images / image information or data output in bitstream form can be delivered to the receiver via a streaming server. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and deliver the received bitstream to a decoding device.
[0069] A streaming server can temporarily store bitstreams during the sending or receiving of bitstreams. Based on user requests via a web server, the streaming server sends multimedia data to the user's device, with the web server acting as an intermediary to notify the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, and the streaming server sends the multimedia data to the user. The content streaming system may include a separate control server, in which case the control server manages the commands / responses between devices within the content streaming system.
[0070] A streaming server can receive content from media storage devices and / or encoding devices. For example, when receiving content from an encoding device, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0071] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0072] The renderer can render the decoded video / images. The rendered video / images can then be displayed on the monitor.
[0073] Figure 2 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied. Hereinafter, the encoding apparatus may include image encoding apparatus and / or video encoding apparatus.
[0074] refer to Figure 2 The encoding device 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor and an intra-frame predictor. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured as at least one hardware component (e.g., an encoder chipset or processor). The memory 270 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0075] Image partitioner 210 can partition an input image (or picture or frame) input to encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be partitioned into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer partitioned. In this case, based on the encoding efficiency according to the image characteristics, the maximum coding unit can be directly used as the final coding unit; alternatively, if necessary, the coding unit can be recursively partitioned into deeper coding units, such that the coding unit with the optimal size is used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be split or partitioned from the aforementioned final encoding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving the transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0076] A unit can be used interchangeably with terms such as block or region in some cases. Typically, an M×N block can represent a set of samples or transform coefficients arranged in M columns and N rows. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. The term sample can be used as a term corresponding to the pixels or cells that make up a picture (or image).
[0077] Encoding device 200 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the predictor from the input image signal (original block, original sample array), and the generated residual signal can be sent to converter 232. In this case, as illustrated, the unit in encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and can generate a prediction block including prediction samples for the current block. The predictor can determine whether intra-frame prediction or inter-frame prediction is applied based on the current block or CU. The predictor can generate various types of prediction-related information, such as prediction mode information, as described below in conjunction with each prediction mode, and can send the generated information to entropy encoder 240. The prediction-related information can be encoded by entropy encoder 240 and output as a bitstream.
[0078] The intra-frame predictor can refer to samples within the current image to predict the current block. The referenced samples can be located adjacent to the current block or located far from it, depending on the prediction pattern. Prediction patterns in intra-frame prediction can include multiple non-directional patterns and multiple directional patterns. Non-directional patterns can include, for example, DC patterns and planar patterns. Directional patterns can include, for example, 33 or 65 directional prediction patterns based on the granularity of the prediction direction. However, this is merely an example, and more or fewer directional prediction patterns can be used depending on the configuration. The intra-frame predictor can determine the prediction pattern to apply to the current block by using prediction patterns applied to neighboring blocks.
[0079] Inter-frame predictors can deduce the predicted block for the current block based on reference blocks (reference sample arrays) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between the motion information of neighboring blocks and the current block. Motion information can include motion vectors and reference image indices. Motion information can also include inter-frame prediction direction information (e.g., L0 prediction, L1 prediction, bidirectional prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks existing in the current image and temporal neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporal neighboring block can be the same or different. Temporal neighboring blocks can be referred to as co-location reference blocks, co-location CUs (colCUs), etc., and the reference image including temporal neighboring blocks can be referred to as co-location images (colPics). For example, the inter-frame predictor can configure a motion information candidate list based on neighboring blocks and can generate information indicating which candidate is used to deduce the motion vector and / or reference image index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, the residual signal may not be sent.
[0080] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Furthermore, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can, for example, be used for screen content coding (SCC). Although IBC essentially performs prediction within the current frame, it can be performed similarly to inter-frame prediction in that it derives a reference block based on a block vector within the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0081] The predicted signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loe've Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT).
[0082] Quantizer 233 can quantize the transform coefficients and send the quantized transform coefficients to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on the one-dimensional vector form. Entropy encoder 240 can perform various encoding methods, such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode the information required for video / image reconstruction (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be sent or stored in bitstream form per NAL (Network Abstraction Layer) unit. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). The video / image information may also include general constraint information. In this disclosure, information and / or syntax elements sent / signed from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded and included in a bitstream through the encoding process described above. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting signals output from the entropy encoder 240 and / or a storage unit (not shown) for storing signals can be configured as internal or external components of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0083] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transform 235. Adder 250 can generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the predictor. When there is no residual for the target block to be processed, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current image, and can also be used for inter-frame prediction of the next image after filtering, as described below.
[0084] Meanwhile, LMCS (Luminance Mapping with Chroma Scaling) can be applied during the image encoding and / or reconstruction process.
[0085] Filter 260 can apply filtering to the reconstructed signal to improve subjective and / or objective image quality. For example, filter 260 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter (ALF), and bilateral filter. Filter 260 can generate filtering-related information and send the generated information to entropy encoder 240. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.
[0086] The modified reconstructed image sent to memory 270 can be used as a reference image in the inter-frame predictor. In this way, when inter-frame prediction is applied, the encoding device 200 can avoid prediction mismatch between the encoding device 200 and the decoding device, and the encoding efficiency can also be improved.
[0087] The DPB of memory 270 can store modified reconstructed images for use as reference images in the inter-frame predictor. Memory 270 can store motion information of blocks in the current image with derived (or encoded) motion information and / or motion information of blocks in the reconstructed image. The stored motion information can be delivered to the inter-frame predictor to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can deliver these reconstructed samples to the intra-frame predictor.
[0088] Figure 3This diagram schematically illustrates the configuration of a video / image decoding device to which embodiments of the present disclosure can be applied. Hereinafter, the decoding device may include an image decoding device and / or a video decoding device.
[0089] refer to Figure 3 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor and an intra-frame predictor. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). The memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.
[0090] When a bitstream including video / image information is input, the decoding device 300 can reconstruct the data related to the video / image information. Figure 2 The image corresponds to the process processed in the encoding device. For example, the decoding device 300 can deduce units / blocks based on information related to block partitions obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied in the encoding device. Therefore, the processing unit used for decoding can be, for example, an encoding unit, and the encoding unit can be partitioned from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transformation units can be derived from the encoding units. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be played by a playback device.
[0091] Decoding device 300 can receive signals output from encoding device in the form of a bitstream, and the received signals can be decoded by entropy decoder 310. For example, entropy decoder 310 can parse the bitstream and derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). Video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Video / image information may also include general constraint information. Decoding device can further decode pictures based on information about parameter sets and / or general constraint information. In this disclosure, information and / or syntax elements that are signaled / received as described below can be decoded and obtained from the bitstream through a decoding process. For example, entropy decoder 310 can decode information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and can output the values of syntax elements required for image reconstruction and quantized values of transform coefficients associated with the residuals. More specifically, in the CABAC entropy decoding method, bins corresponding to each syntax element can be received from the bitstream. A context model can be determined using the information of the syntax elements to be decoded, the decoding information of neighboring blocks and the target block, or the information of symbols / bins decoded in previous stages. The occurrence probability of a bin can be predicted based on the determined context model, arithmetic decoding can be performed on the bins, and symbols corresponding to the value of each syntax element can be generated. In this case, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the decoded symbols / bins for use in the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoder 310 can be provided to the predictor 330, and the residual values from the entropy decoding performed by the entropy decoder 310, i.e., the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive the residual signal (residual block, residual sample, residual sample array). Furthermore, filtering-related information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Meanwhile, a receiver (not shown) for receiving signals output from the encoding device can be further configured as an internal or external element of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Furthermore, the decoding device according to this disclosure can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, and a predictor 330.
[0092] Dequantizer 321 can dequantize the quantized transform coefficients to output transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0093] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block or residual sample array).
[0094] The predictor can perform predictions on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on prediction-related information output from the entropy decoder 310, and determine the specific intra-frame / inter-frame prediction mode.
[0095] Predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined inter-frame and intra-frame prediction (CIIP). Furthermore, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used, for example, for screen content coding (SCC). Although IBC essentially performs prediction within the current frame, it can be performed similarly to inter-frame prediction in that it derives a reference block based on a block vector within the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0096] An intra-frame predictor can refer to samples within the current image to predict the current block. These referenced samples can be located adjacent to the current block or located far from it, depending on the prediction pattern. Prediction patterns in intra-frame prediction can include multiple non-directional patterns and multiple directional patterns. The intra-frame predictor can determine the prediction pattern to apply to the current block by using prediction patterns applied to neighboring blocks.
[0097] An inter-frame predictor can deduce the predicted block for the current block based on reference blocks (reference sample arrays) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between the motion information of neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction information (e.g., L0 prediction, L1 prediction, bidirectional prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks existing in the current image and temporal neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can construct a motion information candidate list based on neighboring blocks and can deduce the motion vector and / or reference image index for the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information related to the prediction may include information indicating the inter-frame prediction mode for the current block.
[0098] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from predictor 330. When there is no residual for the target block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.
[0099] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current image and can be output after filtering as described below, or it can be used for inter-frame prediction of the next image.
[0100] Meanwhile, LMCS (Luminance Mapping with Chroma Scaling) can be applied during image decoding.
[0101] Filter 350 can apply filtering to the reconstructed signal to improve subjective and / or objective image quality. For example, filter 350 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and the modified reconstructed image can be sent to memory 360, specifically to the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter (ALF), and bilateral filter.
[0102] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in the inter-frame predictor. Memory 360 can store motion information of blocks in the current image with derived (or decoded) motion information and / or motion information of blocks in the reconstructed image. The stored motion information can be delivered to the inter-frame predictor to be used as motion information of spatial neighbor blocks or temporal neighbor blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and can deliver the reconstructed samples to the intra-frame predictor.
[0103] In this disclosure, the embodiments described for the filter 260 and predictor 220 of the encoding device 200 can be applied equally or correspondingly to the filter 350 and predictor 330 of the decoding device 300.
[0104] As described above, prediction is performed in video coding to improve compression efficiency. Through such prediction, a prediction block can be generated, comprising prediction samples for the current block, which is the block to be encoded. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived identically in both the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device with information about the residual between the original block and the prediction block (residual information) instead of signaling the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by combining the residual block with the prediction block, and generate a reconstructed image including the reconstructed block.
[0105] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block representing the residual between the original block and the predicted block, perform a transformation process on the residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization process on the transform coefficients to derive quantized transform coefficients, and signal the relevant residual information to the decoding device (via a bitstream). Here, residual information may include information such as the values and locations of the quantized transform coefficients, the transform technique, the transform kernel, and the quantization parameters. The decoding device can perform dequantization and inverse transform processes based on the residual information and can derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the predicted block and the residual block. The encoding device can also derive residual blocks by dequantizing and inverse transforming the quantized transform coefficients to use as a reference for inter-frame prediction of subsequent images and generate a reconstructed image based on this.
[0106] In this disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transform coefficients.
[0107] Furthermore, in this disclosure, quantized transform coefficients and transform coefficients can be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about one or more transform coefficients, and this information about one or more transform coefficients can be signaled using residual coding syntax. Transform coefficients can be derived based on residual information (or information about one or more transform coefficients), and scaled transform coefficients can be derived by inverse transforming (scaling) the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaled transform coefficients. This can be equivalently applied / expressed in other parts of this disclosure.
[0108] Intra-frame prediction refers to generating a prediction sample for the current block based on reference samples within the image to which the current block belongs (hereinafter referred to as the current image). When intra-frame prediction is applied to the current block, neighbor reference samples to be used for intra-frame prediction of the current block can be derived. The neighbor reference samples of the current block may include H+W samples located to the left of the current block of size W×H, W+H samples located above the current block, and at least one sample adjacent to the top-left corner of the current block. Alternatively, the neighbor reference samples of the current block may include multiple rows of top neighbor samples and multiple columns of left neighbor samples.
[0109] Some of the neighbor reference samples in the current block may not have been decoded or may be unavailable. In this case, the decoder can construct neighbor reference samples to be used for prediction by filling in or replacing unavailable samples with available samples.
[0110] Once neighboring reference samples have been derived, the prediction samples for the current block can be derived based on the neighboring reference samples and intra-prediction mode / type information. Here, the intra-prediction mode can indicate one of a non-directional prediction mode or a directional prediction mode that represents the spatial correlation of intra-prediction. The directional prediction mode can be referred to as the angular prediction mode, and the non-directional prediction mode can be referred to as the non-angular prediction mode. The intra-prediction type can indicate various prediction types used to perform intra-prediction. Intra-prediction types can include, for example, MRL (Multiple Reference Lines), ISP (Intra-Segmentation), PDPC (Position-Related Intra-Prediction Combination), MIP (Matrix-Weighted Intra-Prediction or Matrix-Based Intra-Prediction), CCLM (Cross-Component Linear Model), MMLM (Multi-Model Linear Model), DIMD (Decoder-Side Intra-Mode Derivation), fusion of chroma intra-prediction modes, intra-template matching, TIMD (Template-Based Intra-Mode Derivation Fusion), intra-prediction fusion, CCCM (Cross-Component Convolutional Model), CCP (Cross-Component Prediction), and SGPM (Spatial Geometric Partitioning Model). Depending on the circumstances, intra-prediction can be performed using intra-prediction modes and / or intra-prediction types.
[0111] Specifically, the intra-frame prediction process may include an intra-frame prediction mode / type determination step, a reference sample derivation step, and a prediction sample derivation step based on the intra-frame prediction mode / type. Furthermore, a post-filtering step may be performed on the derived prediction samples as needed.
[0112] Figure 4 An example of the intra-frame prediction process is illustrated.
[0113] refer to Figure 4 As described above, the intra-frame prediction process may include an intra-frame prediction mode / type determination step, a reference sample derivation step, and an intra-frame prediction execution (prediction sample generation) step. As mentioned above, the intra-frame prediction process can be performed in both the encoding and decoding devices.
[0114] The coding apparatus determines the intra-frame prediction mode / type (S400). As described above, the coding apparatus may include an encoding device and / or a decoding device.
[0115] The encoding device can determine the intra prediction mode / type to be applied to the current block from the various intra prediction modes / types described in this disclosure, and can generate prediction-related information. The prediction-related information may include intra prediction mode information indicating the intra prediction mode applied to the current block and / or intra prediction type information indicating the intra prediction type applied to the current block. The decoding device can determine the intra prediction mode / type to be applied to the current block based on the prediction-related information.
[0116] For example, when applying intra-prediction, the intra-prediction modes of neighboring blocks can be used to determine the intra-prediction mode to be applied to the current block. For instance, the coding device can select an MPM (Most Probable Mode) candidate from a list of MPM candidates derived from the intra-prediction modes and / or additional candidate modes of the current block's neighboring blocks (e.g., left and / or above neighboring blocks) based on received index information, or it can select an MPM candidate from remaining intra-prediction modes not included in the MPM candidates based on remaining MPM information (remaining intra-prediction mode information). The MPM list can be configured to include planar modes as candidates or exclude planar modes as candidates.
[0117] The encoding device can configure a list of most probable modes (MPMs) for the current block. The MPM list can also be called an MPM candidate list. Here, MPM can refer to a mode used to improve coding efficiency when encoding intra-prediction modes by considering the similarity between the current block and its neighboring blocks.
[0118] The encoding device can perform prediction based on various intra-prediction modes and can determine the optimal intra-prediction mode based on rate-distortion optimization (RDO). In this case, the encoding device can determine the optimal intra-prediction mode using MPM candidates included in the MPM list, or it can determine the optimal intra-prediction mode using the remaining intra-prediction modes other than those included in the MPM list. Specifically, for example, when the intra-prediction type of the current block is a specific type other than the normal intra-prediction type (e.g., DIMD, TIMD, MRL, or ISP), the encoding device can determine the optimal intra-prediction mode by only considering MPM candidates as intra-prediction mode candidates for the current block. That is, in this case, the intra-prediction mode for the current block can be determined only from the MPM candidates, and the MPM flag may not be encoded / signaled. In this case, the decoding device can assume that the MPM flag is equal to 1 without receiving signaling for the MPM flag separately.
[0119] Simultaneously, typically, when the intra-prediction mode of the current block is not a planar mode and is one of the MPM candidates included in the MPM list, the encoding device generates an MPM index (mpm_idx) indicating one of the MPM candidates. If the intra-prediction mode of the current block is not included in the MPM list, the encoding device generates MPM remainder information (remaining intra-prediction mode information) indicating the same mode as the intra-prediction mode of the current block, which comes from a remaining intra-prediction mode not included in the MPM list (and planar modes). The MPM remainder information may include, for example, the intra_luma_mpm_remainder syntax element.
[0120] The decoding device obtains intra-prediction mode information from the bitstream. As described above, the intra-prediction mode information may include at least one of the following: MPM flag, MPM index, and remaining MPM information (remaining intra-prediction mode information). The decoding device may construct an MPM list. The MPM list is constructed in the same way as the MPM list constructed by the encoding device. That is, the MPM list may include the intra-prediction modes of neighboring blocks, and may also include specific intra-prediction modes according to a predetermined method.
[0121] The decoding device can determine the intra-prediction mode of the current block based on the MPM list and intra-prediction mode information. In one example, when the MPM flag is equal to 1, the decoding device can deduce the intra-prediction mode of the current block from the candidate indicated by the MPM index in the MPM list.
[0122] In another example, when the value of the MPM flag is equal to 0, the decoding device can deduce the intra-prediction mode indicated by the remaining intra-prediction mode information (which may be referred to as the MPM remaining information) in the remaining intra-prediction mode as the intra-prediction mode of the current block.
[0123] The encoding device derives reference samples for the current block (S410). The reference samples may include neighboring reference samples of the current block. The neighboring reference samples of the current block may include H+W samples located to the left of the current block (size W×H), W+H samples located above the current block, and at least one sample adjacent to the upper left corner of the current block. Alternatively, the neighboring reference samples of the current block may include multiple rows of upper neighbor samples and multiple columns of left neighbor samples.
[0124] The coding device performs intra-frame prediction on the current block to derive prediction samples (S1220). The coding device can derive prediction samples based on the intra-frame prediction mode / type and reference samples. The coding device can derive reference samples corresponding to the intra-frame prediction mode of the current block from the reference samples of the current block, and can derive prediction samples of the current block based on the reference samples.
[0125] Intra-prediction-based coding processes can typically include, for example, the following operations.
[0126] Figure 5 Examples of video / image coding methods based on intra-frame prediction are shown.
[0127] refer to Figure 5 S500 can be executed by the predictor of the encoding device, S505 can be executed by the residual processor of the encoding device, and S510 or S515 can be executed by the entropy encoder of the encoding device. Specifically, prediction-related information can be derived by the predictor and encoded by the entropy encoder. Residual information can be derived by the residual processor and encoded by the entropy encoder. Residual information is information about residual samples. Residual information may include information about the quantized transform coefficients for the residual samples. As mentioned above, residual samples can be derived into transform coefficients by the transformer of the encoding device, and transform coefficients can be derived into quantized transform coefficients by the quantizer. Information about the quantized transform coefficients can be encoded by the entropy encoder via the residual encoding process.
[0128] The encoding device performs intra-prediction on the current block (S500). The encoding device can deduce the intra-prediction mode / type for the current block, deduce reference samples for the current block, and generate prediction samples for the current block based on the intra-prediction mode / type and the reference samples. Here, the process of determining the intra-prediction mode / type, the process of deduce neighboring reference samples, and the process of generating prediction samples can be performed simultaneously, or one process can be performed before the other. The encoding device can determine the mode / type to be applied to the current block from multiple intra-prediction modes / types. The encoding device can compare the RD costs of the intra-prediction modes / types and determine the optimal intra-prediction mode / type for the current block.
[0129] Simultaneously, the encoding device can perform a prediction sample filtering process. Prediction sample filtering can also be called post-filtering. Based on the prediction sample filtering process, some or all of the prediction samples can be filtered. In some cases, the prediction sample filtering process can be omitted.
[0130] The encoding device generates residual samples for the current block based on the predicted samples (S505). The encoding device can derive the residual samples by comparing the predicted samples with the original samples of the current block one by one.
[0131] The encoding device can encode video / image information, including information about intra-frame prediction (prediction-related information) and / or information about residual samples (residual information) (S510 or S515). Prediction-related information may include intra-frame prediction mode information and intra-frame prediction type information. The encoding device can output the encoded video / image information as a bitstream. The output bitstream can be sent to the decoding device via a storage medium or network.
[0132] Residual information can include residual coding syntax elements. The coding device can derive quantized transform coefficients by transforming and quantizing the residual samples. The residual information can include information about the quantized transform coefficients.
[0133] Simultaneously, as described above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks). To this end, the encoding device can derive (modified) residual samples by performing dequantization and inverse transform on the quantized transform coefficients. The reason for performing dequantization and inverse transform after transforming and quantizing the residual samples is to derive residual samples identical to those derived in the decoding device, as described above. The encoding device can generate reconstructed blocks, including reconstructed samples for the current block, based on the predicted samples and the (modified) residual samples. A reconstructed image for the current image can be generated based on the reconstructed blocks. As described above, loop filtering processes, etc., can be further applied to the reconstructed image.
[0134] Decoding devices can perform operations corresponding to those performed by encoding devices. A video / image decoding process based on intra-frame prediction may include, for example, the following operations.
[0135] Figure 6 An example of a video / image decoding method based on intra-frame prediction is shown.
[0136] refer to Figure 6 S600 can be executed by the entropy decoder of the decoding device, S610 can be executed by the predictor of the decoding device, S615 can be executed by the residual processor of the decoding device, and S620 can be executed by the adder or reconstructor of the decoding device.
[0137] Specifically, the decoding device obtains video / image information from the bitstream (S600). The video / image information may include prediction-related information and / or residual information.
[0138] The decoding device performs intra-frame prediction based on prediction-related information (S610). The decoding device can deduce the intra-frame prediction mode / type for the current block, deduce the reference samples for the current block, and generate prediction samples for the current block based on the intra-frame prediction mode / type and the reference samples. In this case, the decoding device can perform a prediction sample filtering process. Prediction sample filtering can be referred to as post-filtering. Depending on the prediction sample filtering process, some or all of the prediction samples can be filtered. In some cases, the prediction sample filtering process can be omitted.
[0139] The decoding device performs residual processing based on the residual information (S615). The decoding device can derive residual samples for the current block based on the residual information. Specifically, the dequantizer of the residual processor can derive the transform coefficients by performing dequantization based on the quantized transform coefficients derived from the residual information, and the inverse transformer of the residual processor can derive residual samples for the current block by performing an inverse transform on the transform coefficients.
[0140] The decoding device generates a reconstructed block / image (S620). The decoding device can generate reconstructed samples for the current block based on predicted samples and / or residual samples, and derive a reconstructed block including the reconstructed samples. A reconstructed image for the current image can be generated based on the reconstructed block. As described above, a loop filtering process, etc., can be further applied to the reconstructed image.
[0141] Intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using the binarization and encoding methods described in this disclosure. For example, intra-prediction mode information and / or intra-prediction type information can be binarized using fixed-length binarization, truncated Ricean binarization, truncated unary binarization, etc. For example, intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using entropy coding (e.g., CABAC, CAVLC).
[0142] The following will describe DIMD (decoder-side intra-mode derivation) as one of the intra-prediction methods.
[0143] According to this disclosure, intra-frame modes can be derived based on the DIMD technique. In DIMD, a histogram of oriented gradients (HoG) can be computed using a template that includes samples from the current block's neighbors. For example, the HoG can be computed using horizontal and vertical Sobel filters based on a template that includes samples from three rows of neighbors. In this case, both horizontal and vertical Sobel filters can be applied to compute the HoG. If the template is located in a different CTU, the Sobel filter can be applied without crossing the upper CTU boundary, or the DIMD technique can be omitted from the current block.
[0144] For example, for candidate intra-frame modes, the n intra-frame modes with the highest histogram values can be selected / extracted through HoG computation. In this case, the n intra-frame modes can be used to derive the prediction block for the current block. Furthermore, the prediction block can be derived using n intra-frame modes and non-directional modes (e.g., planar or DC modes). In this case, a predictor corresponding to each intra-frame mode can be derived, and the prediction block can be derived by weighted summation / weighted averaging of the predictors. In other words, through prediction fusion, the predictors corresponding to the five selected / extracted intra-frame modes and the predictor corresponding to the planar mode can be fused.
[0145] The following section describes TIMD (Template-based Intra-mode Derivation) as one of the intra-prediction methods.
[0146] According to this disclosure, intra-frame modes can be derived based on the fusion of Template-Based Intra-Frame Mode Derivation (TIMD) techniques. TIMD techniques can be referred to as TIMD types. In TIMD, a template including neighbor samples of the current block can be derived, a predictor for the template corresponding to a candidate prediction mode can be derived based on reference samples of the template (i.e., the predicted samples of the template can be derived), and the optimal intra-frame prediction mode can be derived by comparing the predictor with the reconstructed template (i.e., the reconstructed samples of the template).
[0147] The template can include a top template and a left template. Here, the size of the template can be L. However, this is just an example, and the height L2 of the top template and the width L1 of the left template can vary depending on the size of the current block, etc. For example, when the current block is a non-square block with a width greater than its height, L2 can be greater than L1. For example, when the current block is a non-square block with a height greater than its width, L1 can be greater than L2.
[0148] The TIMD candidate intra-prediction mode can be determined based on various criteria. In one example, the TIMD candidate intra-prediction mode can be limited to MPM. For instance, the TIMD candidate intra-prediction mode can be limited to candidate modes included in the MPM (or PMPM) list. In this case, for example, the TIMD flag can be signaled after the MPM flag (or PMPM flag) is signaled.
[0149] In another example, a predetermined set of candidate intra-prediction modes can be defined. The predetermined candidate intra-prediction modes can include some or all of the DIMD derivation modes described above. The predetermined candidate intra-prediction modes can include a default mode. The default mode can include at least one of, for example, a vertical mode, a horizontal mode, a bottom-left diagonal mode, a top-left diagonal mode, or a top-right diagonal mode.
[0150] Furthermore, under certain conditions, only a single TIMD prediction pattern can be derived. For example, when the SATD of an intra-prediction pattern with minimum SAD or SATD calculated based on TIMD is less than a threshold, only one intra-prediction pattern can be derived. As another example, when the difference between the first SATD of a first intra-prediction pattern with minimum SATD calculated based on TIMD and the second SATD of a second intra-prediction pattern with second minimum SATD is greater than a threshold, only one intra-prediction pattern can be derived.
[0151] The following section describes intra-template matching (intraTMP) as one of the intra-prediction methods.
[0152] According to this disclosure, intraTMP can be applied. IntraTMP can be considered a prediction technique or prediction type. According to intraTMP, a template most similar to the current template (neighbor reference sample (L-shaped) template) can be searched within a predefined search range of the reconstructed portion of the current image, and prediction of the corresponding block can be performed based on this. In this case, a block vector indicating the position of the matching block (reference block) relative to the current block in the current image can be derived / stored. Furthermore, the predefined search range can be located within the left CTU, the current CTU, the upper left CTU, the upper CTU, and the upper right CTU.
[0153] For example, intraTMP can be performed in the same way by both the encoder and the decoder, and the encoder can signal whether intraTMP is used.
[0154] Additionally, to perform intraTMP efficiently, a candidate list can be configured for intraTMP. The candidate list can include up to n (e.g., 19) template matching block vectors, and in this case, the candidate list can be sorted in ascending order based on the template's SAD cost.
[0155] For example, intra-template matching may be available when the height and width of the CU are less than or equal to 64. The maximum available size can be predefined or signaled at a higher level. Furthermore, whether intra-template matching (intraTMP) is applied can be signaled via CU-level flags. For example, this signaling can be performed when the DIMD mode is not applied to the current block. That is, when dim_flag equals 0, intratmp_flag can be signaled. Additionally, the block vector (BV) derived via intraTMP can be stored and subsequently used in the IBC candidate list for later blocks.
[0156] The following section will describe SGPM (Spatial Geometric Partitioning Pattern), one of the intra-frame prediction methods.
[0157] According to this disclosure, SGPM can be applied. SGPM can be considered as a prediction technique or prediction type. According to SGPM, the current block can be partitioned into two partitions (geometric partitions), and different intra-prediction patterns can be applied to these two partitions. For example, a first intra-prediction pattern can be applied to the first partition, and a second intra-prediction pattern can be applied to the second partition. The two predictors derived therefrom can be merged or mixed to generate a single prediction block. In this case, n partition patterns can be used for SGPM, and for example, 26 partition patterns can be used. Furthermore, n partition patterns and m intra-prediction patterns can be combined to configure an SGPM candidate list of length k.
[0158] For example, the length k of the SGPM candidate list can be 16. For example, n can be 26 and m can be 3. The value of k can be a predefined value, or it can be set differently based on the size of the current block.
[0159] The following describes residual processing. Residual processing can be performed in both encoding and decoding devices. Residual processing may include coefficient encoding, transformation, and / or quantization processes.
[0160] The residual processing on the encoding side may include the process of generating and / or encoding residual information from the derived residual samples of the current block. The residual processing may also include the process of deriving residual samples based on predicted samples. The residual processing on the decoding side may include the process of deriving residual samples from the residual information of the received bitstream. For example, the residual processing may include (inverse) transform and / or (inverse) quantization processes. Furthermore, the residual processing may include the encoding / decoding process of residual information. The residual information may include residual data and / or transform / quantization related parameters.
[0161] Specifically, for residual processing, one method can be used whereby (quantized) transform coefficients within a block are derived, and residual information is generated and encoded / decoded based on these coefficients. Alternatively, (quantized) residual coefficients within a block to which transform skipping has been applied can be derived, and residual information for transform skipping can be generated and encoded / decoded based on this. The encoded information can be output as a bitstream as described above.
[0162] Furthermore, on the decoding side, the (quantized) transform coefficients or (quantized) residual coefficients within a block can be derived from the residual information (or residual information used for transform skipping) included in the bitstream, and residual samples can be derived by performing dequantization and / or inverse transform processes as needed.
[0163] The encoding process based on residual processing can typically include, for example, the following operations.
[0164] Figure 7An example of a video / image coding method based on residual processing is shown.
[0165] refer to Figure 7 S700 can be executed by the predictor of the encoding device, S710 by the residual processor of the encoding device, and S720 by the entropy encoder of the encoding device. Specifically, prediction-related information can be derived by the predictor and encoded by the entropy encoder. Residual information can be derived by the residual processor and encoded by the entropy encoder. Residual information is information about residual samples. Residual information may include information about the quantized transform coefficients for the residual samples. As mentioned above, residual samples can be transformed into transform coefficients by the transformer of the encoding device, and transform coefficients can be quantized into quantized transform coefficients by the quantizer. Information about the quantized transform coefficients can be encoded by the entropy encoder via the residual encoding process.
[0166] The encoding device derives the prediction sample for the current block (S700). The encoding device can derive the prediction sample for the current block based on the inter-frame prediction and / or intra-frame prediction described above.
[0167] The encoding device can perform residual processing based on the predicted samples (S710). The encoding device can derive residual samples based on the predicted samples. The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples. As described above, residual processing includes a transformation process and / or a quantization process for the residual samples. The encoding device can generate residual information from the residual samples through residual processing. As described above, the residual information can include information about the quantized transformation coefficients.
[0168] The encoding device encodes video / image information, including prediction-related information and / or residual information (S720). The encoding device can output the encoded video / image information as a bitstream. Prediction-related information may include information related to the prediction process. Residual information is information about residual samples. Residual information may include information about the quantized transform coefficients for the residual samples.
[0169] The output bitstream can be stored in a (digital) storage medium and delivered to a decoding device, or it can be delivered to a decoding device via a network.
[0170] Simultaneously, as mentioned above, the encoding device can generate reconstructed images (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is because the encoding device derives the same prediction result as the decoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed images (or reconstructed samples or reconstructed blocks) in memory and use them as reference images for inter-frame prediction. As mentioned above, loop filtering processes, etc., can be further applied to the reconstructed images.
[0171] Decoding devices can perform operations corresponding to those performed by encoding devices. A video / image decoding process based on residual processing may include, for example, the following operations.
[0172] Figure 8 An example of a video / image decoding method based on residual processing is shown.
[0173] refer to Figure 8 S800 can be executed by the entropy decoder of the decoding device, S810 can be executed by the predictor of the decoding device, S820 can be executed by the residual processor of the decoding device, and S830 can be executed by the adder or reconstructor of the decoding device.
[0174] Specifically, the decoding device acquires video / image information from the bitstream (S800). The video / image information may include prediction-related information and / or residual information.
[0175] The decoding device performs prediction (including inter-frame prediction and / or intra-frame prediction) based on prediction-related information (S810). The decoding device can deduce the prediction mode / type of the current block based on the prediction-related information and can generate prediction samples within the current block based on the prediction mode / type. In this case, the decoding device can perform a prediction sample filtering process. Prediction sample filtering can be referred to as post-filtering. Through the prediction sample filtering process, some or all prediction samples can be filtered. In some cases, the prediction sample filtering process can be omitted.
[0176] The decoding device performs residual processing based on the residual information (S820). The decoding device can derive the residual samples of the current block based on the residual information. Specifically, the dequantizer of the residual processor can perform a dequantization process to derive the transform coefficients based on the quantized transform coefficients derived from the residual information, and the inverse transformer of the residual processor can perform an inverse transform process on the transform coefficients to derive the residual samples of the current block.
[0177] The decoding device generates a reconstructed block / image (S830). The decoding device can generate reconstructed samples for the current block based on predicted samples and / or residual samples, and can derive a reconstructed block including the reconstructed samples. A reconstructed image of the current image can be generated based on the reconstructed block. As described above, a loop filtering process, etc., can be further applied to the reconstructed image.
[0178] Residual information can be encoded / decoded using the binarization and encoding methods described in this disclosure. For example, residual information can be binarized using fixed-length binarization, truncated Rice binarization, truncated unary binarization, etc. For example, residual information can be encoded / decoded using entropy coding (e.g., CABAC or CAVLC) or bypass coding.
[0179] The following section describes the setting of the maximum transform size and the transform coefficients to zero.
[0180] For example, the CTU size and maximum transform size (for all MTS cores) can be extended to 256. In this case, the maximum size of the intra-prediction block can be set to 128×128. For UHD sequences, the maximum CTU size can be set to 256, and for other cases, it can be set to 128.
[0181] In the first transformation process, the transformation coefficients may not be zeroed out. However, in the second transformation process applying LFNST, the first transformation coefficients outside the ROI region to which LFNST is applied may be zeroed out.
[0182] Simultaneously, zeroing the transform coefficients can be applied during the first transform process. As the size of the transform block increases, zeroing after the first transform may be necessary, and whether or not to apply zeroing can be determined based on whether MTS is applied. For example, when MTS is applied, zeroing can be performed to one of 16, 20, 24, 32, or 40, and when MTS is not applied, zeroing can be performed to a predetermined specific value. The specific value can be 32 or 64. When MTS is applied, the number of transform coefficients to be zeroed can vary depending on the size of the transform block. For example, when the width or height of the transform block to which MTS is applied is 4, 8, or 16 (i.e., 4-point transform, 8-point transform, or 16-point transform), zeroing may not be applied, and when the width or height of the transform block to which MTS is applied is 32 (i.e., 32-point transform), zeroing can be performed to 16. Alternatively, when MTS is applied to a width or height greater than 32, zeroing can be performed to 32.
[0183] Furthermore, residual processing may include a transform / inverse transform process and / or a quantization / dequantization process. According to this disclosure, a master transform and / or a quadratic transform can be applied to the residual block to derive a transform coefficient block (transform coefficients), and an inverse quadratic transform and / or an inverse master transform can be applied to the transform coefficient block (transform coefficients) to derive the residual block.
[0184] As described above, the transformation of the residuals can be performed through a primary transformation and / or a secondary transformation selectively performed after the primary transformation. The primary transformation can be referred to as the master transformation and can be a Discrete Cosine Transform (DCT) and / or Discrete Sine Transform (DST) applied to all rows and columns of the residual block. Following the primary transformation, a secondary transformation can be additionally applied to specific transform coefficients in the upper-left region of the transform block produced by the primary transformation. The inverse transformation performed during decoding can include applying an inverse secondary transformation to specific residuals in the upper-left region of the residual block corresponding to the (inverse-quantized) residual information, and applying an inverse primary transformation to the transform block produced by the inverse secondary transformation.
[0185] Furthermore, as described below, the NSPT transform can be used in certain situations. The NSPT transform can be a transform that integrates and replaces the main transform and the secondary transform.
[0186] The following will describe MTS (Multiple Transformation Selection) as one of the transformation methods.
[0187] Figure 9 Multiple transformation techniques according to this disclosure are illustrated by way of example.
[0188] refer to Figure 9 The converter can correspond to the above. Figure 2 The converter in the encoding device, and the inverse converter can correspond to the above. Figure 2 Inverse converter in encoding device or the above Figure 3 The inverse converter in the decoding device.
[0189] The transformer can perform a master transform based on residual samples (residual sample array) in the residual block to derive the (master) transform coefficients (S900). Such a master transform can be called a core transform. Here, the master transform can be based on multiple transform selection (MTS), and when multiple transforms are applied as master transforms, the master transform can be called a multi-core transform.
[0190] The converter can perform a quadratic transform based on the (primary) transform coefficients to derive modified (quadratic) transform coefficients (S910). The primary transform is a transform from the spatial domain to the frequency domain, and the quadratic transform can be represented by utilizing the correlation between the (primary) transform coefficients to achieve a more compact representation.
[0191] For example, a quadratic transform can include a non-separable transform. In this case, the quadratic transform can be called a non-separable quadratic transform (NSST) or a mode-dependent non-separable quadratic transform (MDNSST). A non-separable quadratic transform can represent a transform that generates modified transform coefficients (or quadratic transform coefficients) for the residual signal by performing a quadratic transform on the (master) transform coefficients derived via the master transform based on a non-separable transform matrix. Here, based on the non-separable transform matrix, the transform can be applied to the (master) transform coefficients all at once without applying the vertical and horizontal transforms separately (or without applying the horizontal and vertical transforms independently).
[0192] In other words, a non-separable quadratic transform can represent a transform method that does not separate the vertical and horizontal components of the (master) transform coefficients. For example, a two-dimensional signal (transform coefficients) can first be rearranged into a one-dimensional signal according to a predetermined direction, and then modified transform coefficients (or quadratic transform coefficients) can be generated based on the non-separable transform matrix. In other words, the transform method may include rearranging the transform coefficients into a one-dimensional signal in a row-major or column-major direction, and then generating modified transform coefficients (or quadratic transform coefficients) based on the non-separable transform matrix.
[0193] Furthermore, the inverse transformer can perform a series of processes in the reverse order of the processes performed by the transformer described above. The inverse transformer can receive (inversely quantized) transform coefficients, perform an inverse quadratic transform to derive the (main) transform coefficients (S920), and perform an inverse main transform on the (main) transform coefficients to obtain residual blocks (residual samples) (S930). Here, on the inverse transformer side, the main transform coefficients can be referred to as modified transform coefficients. As described above, the encoding and / or decoding devices can generate reconstructed blocks based on the residual blocks and prediction blocks, and generate reconstructed images based on these blocks.
[0194] In this disclosure, the master transform may be referred to as the core transform. Here, the master transform may be based on multiple transform selection (MTS), and when a transform kernel selected from multiple transform kernel types is applied as the master transform, the transform may be referred to as a multi-core transform.
[0195] Multi-core transform can refer to transform schemes using Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, DCT-VIII, etc. That is, multi-core transform can represent a transform method that uses multiple transform kernels selected from DCT-II, DST-VII, DCT-VIII, and DST-I to transform the residual signal (or residual block) in the spatial domain into transform coefficients (or master transform coefficients) in the frequency domain. Here, from the perspective of the transformer, the master transform coefficients can be referred to as temporary transform coefficients.
[0196] In other words, when applying conventional transform methods, the residual signal (or residual block) can be transformed from the spatial domain to the frequency domain based on DCT-II, thereby generating transform coefficients. In contrast, when applying multi-core transforms, the residual signal (or residual block) can be transformed from the spatial domain to the frequency domain based on DCT-II, DST-VII, DCT-VIII, and / or DST-I, thereby generating transform coefficients (or master transform coefficients). Here, DCT-II, DST-VII, DCT-VIII, DST-I, etc., can be referred to as transform types, transform kernels, or transform cores. Such DCT / DST transform types can be defined based on basis functions.
[0197] When performing a multi-core transform, a vertical transform kernel and a horizontal transform kernel can be selected from the transform kernels for the target block (the block to be transformed). A vertical transform can be performed on the target block based on the vertical transform kernel, and a horizontal transform can be performed on the target block based on the horizontal transform kernel. Here, the horizontal transform can represent the transformation of the horizontal components of the target block, and the vertical transform can represent the transformation of the vertical components of the target block. The vertical transform kernel and / or the horizontal transform kernel can be adaptively determined based on the prediction mode and / or transform index of the target block (CU or sub-block), including the residual block.
[0198] Furthermore, according to the example, when performing the master transform by applying the MTS, specific basis functions can be assigned predetermined values, and the mapping relationship of the transform kernels can be established by combining whether specific basis functions are applied to the vertical or horizontal transform. For example, when the horizontal transform kernel is represented by trTypeHor and the vertical transform kernel is represented by trTypeVer, a value of 0 for trTypeHor or trTypeVer corresponds to DCT-II, a value of 1 for trTypeHor or trTypeVer corresponds to DST-VII, and a value of 2 for trTypeHor or trTypeVer corresponds to DCT-VIII.
[0199] In this case, the MTS index information can be encoded and signaled to the decoding device to indicate one of multiple transform kernel sets. For example, an MTS index of 0 can indicate that the values of trTypeHor and trTypeVer are both 0; an MTS index of 1 can indicate that the values of trTypeHor and trTypeVer are both 1; an MTS index of 2 can indicate that the value of trTypeHor is 2 and the value of trTypeVer is 1; an MTS index of 3 can indicate that the value of trTypeHor is 1 and the value of trTypeVer is 2; and an MTS index of 4 can indicate that the values of trTypeHor and trTypeVer are both 2.
[0200] Such an MTS can be applied via explicit signaling as described above, or implicitly based on specific conditions. In an implicit MTS, the transform kernel can be derived independently for each direction based on the width or height of the transform block. For example, when applying a sub-block transform (SBT) that applies the transform to only one of the sub-blocks, an implicit MTS can be applied when the width or height meets specific conditions, or when conditions related to a specific intra-frame mode are met. For instance, an implicit MTS can be applied to a block that has applied SBT and where the larger of its width and height is less than or equal to 32 or 64. Alternatively, an implicit MTS can be applied to a block that has applied ISP, or to a block that has applied intra-frame prediction but has neither applied LFNST nor MIP.
[0201] Furthermore, implicit MTS can use lookup tables (LUTs) to infer the primary transform pair based on the intra-prediction mode of the current block and the size of the TU. For example, the same intra-prediction mode as explicit MTS can be considered, and the maximum TU size can be 32×32. In this case, the LUT can include transform pairs based on separable primary transforms, and no additional primary transforms need to be added. For example, DCT-II, DCT-V, DCT-VIII, DST-I, DST-IV, and DST-VII can be used.
[0202] For example, in the case of ISP blocks, the sizeIdx used as LUT input can be based on the position of the current ISP sub-partition within the ISP block. Furthermore, according to an example, implicit MTS can be applied only to ISP blocks based on CTC configuration settings. According to another example, implicit MTS can be applied to angular intra-prediction modes. For example, implicit MTS for ISP blocks can be used by using implicit signaling for all modes (non-CTC). For example, for TIMD and DIMD modes, the default implicit MTS (i.e., based on the shape of the TU) can be maintained. Furthermore, DCT-II can be used for MIP, EIP, SGPM, and IntraTMP modes. Additionally, for ISP blocks, the default implicit MTS can be used.
[0203] The following will describe LFNST (Low Frequency Non-Separable Transform) as one of the transformation methods.
[0204] Simultaneously, the encoder's transformer can perform a secondary transformation based on the (primary) transform coefficients to derive modified (secondary) transform coefficients. Here, the primary transform is a transformation from the spatial domain to the frequency domain, and the secondary transform refers to converting to a more compact representation by utilizing the correlation between the (primary) transform coefficients. The secondary transform can include a non-separable transform. In this case, the secondary transform can be called a non-separable secondary transform (NSST). A non-separable secondary transform can represent a transformation that generates modified transform coefficients (or secondary transform coefficients) for the residual signal by performing a secondary transform on the (primary) transform coefficients derived via the primary transform based on a non-separable transform matrix. Here, based on the non-separable transform matrix, the (primary) transform coefficients can be transformed all at once without separating the vertical and horizontal transforms (or independently applying the horizontal and vertical transforms). In other words, a non-separable secondary transform can represent a transformation method that does not separate the vertical and horizontal components of the (primary) transform coefficients. For example, a two-dimensional signal (transform coefficients) can first be rearranged into a one-dimensional signal according to a predetermined direction (e.g., row-first or column-first), and then modified transform coefficients (or quadratic transform coefficients) can be generated based on a non-separable transform matrix. For example, row-first order refers to arranging M×N blocks in the order of row 1, row 2, ..., row N, while column-first order refers to arranging M×N blocks in the order of column 1, column 2, ..., column M. A non-separable quadratic transform can be applied to the upper left region of a block consisting of (primary) transform coefficients (hereinafter referred to as a transform coefficient block). In a non-separable quadratic transform, the transform kernel (or transform type) can be selected in a mode-dependent manner. Here, the mode can include intra-frame prediction modes and / or inter-frame prediction modes.
[0205] Non-separable quadratic transformations can be performed based on either an 8×8 transformation or a 4×4 transformation, determined by the width (W) and height (H) of the transform coefficient block. An 8×8 transformation refers to a transformation that can be applied to an 8×8 region within the transform coefficient block when both W and H are greater than or equal to 8, and this 8×8 region can be the top-left 8×8 region within the transform coefficient block. Similarly, a 4×4 transformation refers to a transformation that can be applied to a 4×4 region within the transform coefficient block when both W and H are greater than or equal to 4, and this 4×4 region can be the top-left 4×4 region within the transform coefficient block. For example, an 8×8 transform kernel matrix can be a 64×64 / 16×64 matrix, and a 4×4 transform kernel matrix can be a 16×16 / 8×16 matrix.
[0206] Furthermore, for pattern-based transform kernel selection, transform sets can be defined for non-separable quadratic transforms, and k non-separable quadratic transform kernels can be configured for each transform set. For example, the number of transform sets can be 4, 35, etc. The selection of a specific transform set within the transform set can be performed, for example, based on the intra-prediction pattern of the target block (CU or sub-block).
[0207] For example, when determining that a specific set should be used for a non-separable transform, one of the k transform kernels within that set can be selected using the non-separable quadratic transform index. The encoding device can derive the non-separable quadratic transform index indicating the specific transform kernel based on an RD (rate-distortion) check, and can signal this index to the decoding device. The decoding device can then select one of the k transform kernels within the specific set based on the non-separable quadratic transform index.
[0208] The inverse transformers of encoding and decoding devices can perform a series of processes in the reverse order of the processes performed by the transformers described above. The inverse transformer can receive (inversely quantized) transform coefficients and perform an inverse quadratic transform to derive the (primary) transform coefficients. From the perspective of the inverse transformer, the primary transform coefficients can be referred to as the modified transform coefficients.
[0209] In this disclosure, to reduce the computational complexity and memory requirements associated with non-separable quadratic transforms, a reduced quadratic transform (RST) can be applied, where the size of the transform matrix (kernel) is reduced from the NSST concept. RSTs can be referred to by various terms, such as reduced transform, reduced quadratic transform, simplified transform, or simple transform, and the name of RST is not limited to these examples. Alternatively, because RSTs are primarily performed in the low-frequency region of transform blocks containing non-zero coefficients, it can also be called LFNST (low-frequency non-separable transform).
[0210] In LFNST, an N-dimensional vector can be mapped to an R-dimensional vector in another space to determine the reduced transformation matrix, where R is less than N. N can represent the square of the length of one side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied. The reduction factor can be represented by the R / N value.
[0211] According to the example, the size of the LFNST matrix can be R×N, which is smaller than the size of the regular transformation matrix N×N, and can be defined as Equation 1 below.
[0212] [Formula 1]
[0213] When the LFNST matrix When multiplied with the residual sample of the target block, the transformation coefficients of the target block can be derived. When the size of the block to which the transformation is applied is 8×8 and R=16 (i.e., R / N=16 / 64=1 / 4), LFNST can be represented by matrix operations as shown in Equation 2 below.
[0214] [Equation 2]
[0215] In Equation 2, r1 to r64 can represent the residual samples of the target block, and more specifically, can be the transform coefficients generated by applying the master transform. As a result of Equation 2, the transform coefficients ci of the target block can be derived, where ci can include c1 to cR. That is, when R=16, the transform coefficients c1 to c16 of the target block can be derived.
[0216] If a conventional (standard) transform is applied instead of LFNST, and a transform matrix of size 64×64 (N×N) is multiplied by residual samples of size 64×1 (N×1), 64 (N) transform coefficients of the target block are derived. However, due to the application of LFNST, only 16 (R) transform coefficients of the target block are derived. Since the total number of transform coefficients of the target block is reduced from N to R, the amount of data sent from the encoding device to the decoding device is reduced, thereby improving the transmission efficiency between the encoding and decoding devices.
[0217] Inverse LFNST matrix The size is N×R, which is smaller than the size of the conventional inverse transformation matrix N×N, and is similar to the LFNST matrix shown in Equation 1. It has a transpose relationship. It can represent the inverse LFNST matrix. (Where the superscript T denotes transpose). When the inverse LFNST matrix Multiplying by the transform coefficients of the target block allows for the derivation of either the modified transform coefficients of the target block or the residual samples of the target block. The inverse LFNST matrix... It can also be expressed as .
[0218] According to the example, when the size of the block to which the inverse transformation is applied is 8×8 and R=16 (i.e., R / N=16 / 64=1 / 4), the inverse LFNST can be represented by matrix operations as shown in Equation 3 below.
[0219] [Formula 3]
[0220] c1 to c16 can represent the transformation coefficients of the target block. As a result of Equation 12, rj representing the modified transformation coefficients of the target block or the residual samples of the target block can be derived. rj can include r1 to rN. That is, when N=64, the transformation coefficients r1 to r64 of the target block can be derived.
[0221] Figure 10 An example of LFNST is shown.
[0222] refer to Figure 10 A 4×4 LFNST can be applied to blocks with min(width, height) < 8, and an 8×8 LFNST can be applied to blocks with min(width, height) > 4.
[0223] For example, a 4×4 forward LFNST can input 16 (primary) transform coefficients, and an 8×8 forward LFNST can input 64 (primary) transform coefficients. Since 8 or 16 transform coefficients can be derived through such a forward LFNST, 8 or 16 transform coefficients can be input for the inverse LFNST, respectively. When performing a 4×4 inverse LFNST, 16 modified transform coefficients can be output from 8 transform coefficients, and when performing an 8×8 inverse LFNST, 64 modified transform coefficients can be output from 16 transform coefficients.
[0224] Furthermore, according to the example, in the transformation of the encoding process, instead of applying a 16×64 transformation kernel matrix to the 64 data elements constituting the 8×8 region, only 48 data elements can be selected and a transformation kernel matrix of up to 16×48 can be applied. Here, "up to" means that for an m×48 transformation kernel matrix that can generate m coefficients, the maximum value of m is 16. That is, when performing LFNST by applying an m×48 transformation kernel matrix (m≤16) to an 8×8 region, 48 data elements can be used as input to generate m coefficients. When m is 16, 16 coefficients are generated from the 48 input data elements. That is, when the 48 data elements form a 48×1 vector, a 16×1 vector can be generated by sequentially multiplying the 16×48 matrix and the 48×1 vector. In this case, the 48 data elements constituting the 8×8 region can be appropriately arranged to form a 48×1 vector. When matrix operations are performed by applying a transformation kernel matrix of up to 16×48, 16 modified transformation coefficients are generated, and these 16 modified transformation coefficients can be arranged in the upper left 4×4 region according to the scan order, while the upper right 4×4 region and the lower left 4×4 region can be filled with zeros.
[0225] In the inverse transform of the decoding process, the transpose of the aforementioned transform kernel matrix can be used. That is, when performing the inverse LFNST as an inverse transform process executed by the decoding device, the input coefficient data for which the inverse LFNST is applied can be configured as a one-dimensional vector according to a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the inverse LFNST matrix by the one-dimensional vector from the left can be arranged in a two-dimensional block according to a predetermined arrangement order. In this case, the matrix operation can be represented as (48×16 matrix) × (16×1 transform coefficient vector) = (48×1 modified transform coefficient vector). Here, an n×1 vector can be interpreted as having the same meaning as an n×1 matrix, and therefore can be represented as an n×1 column vector. "×" indicates matrix multiplication. When performing such a matrix operation, 48 modified transform coefficients can be derived, and these 48 modified transform coefficients can be arranged in the upper left, upper right, and lower left regions of an 8×8 region, excluding the lower right region.
[0226] Simultaneously, LFNST can be applied to sub-blocks. For example, when splitting a coded block into sub-blocks, the size of each sub-block can be "width / 2 × height / 2" or "width / 4 × height / 4". In this case, sub-blocks can include corner sub-blocks or center sub-blocks. For example, LFNST can be applied to corner sub-blocks. In this case, similar to existing SBTs, the splitting pattern and the location of non-zero sub-blocks can be indicated by signals.
[0227] The following will describe the Non-Separable Master Transform (NSPT) used for the transformation.
[0228] For example, as mentioned above, separable transforms in which the transform kernel is applied independently in the vertical and horizontal directions can be used as the principal transform or the inverse principal transform. Furthermore, the LFNST, as a non-separable transform, can be used as a quadratic transform or the inverse quadratic transform. Additionally, when the MTS is used as both the principal and inverse principal transforms, the LFNST can be omitted.
[0229] Simultaneously, when MTS is not applied and (DCT-II, DCT-II) is applied as the primary transform along with LFNST, the Non-Separable Primary Transform (NSPT) can be applied as the primary transform. That is, the non-separable transform can be performed as the primary transform, and no additional secondary transforms are required. Such a transform can be applied only to the luma component, to both the luma and chroma components, or, considering the fact that the non-separable transform is applied to the intra-prediction block, only to the intra-prediction block. Furthermore, NSPT can be conditionally applied based on a specific tree structure. For example, in the case of a two-tree chroma structure, NSPT can be performed when the coded_flag for the chroma component is not equal to 0 and transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2] are both equal to 0. In the case of a two-tree luma structure, NSPT can be performed when the coded_flag for the luma component is not equal to 0 and transform_skip_flag[x0][y0][0] is equal to 0. Furthermore, in the case of a single-tree structure, NSPT can be performed when the coded_flag for the luma component is not equal to 0 and transform_skip_flag[x0][y0][0] is equal to 0, and when the coded_flag for the chroma component is not equal to 0 and transform_skip_flag[x0][y0][1] and transform_skip_flag[x0][y0][2] are both equal to 0. That is, in the case of a two-tree chroma structure, NSPT can be applied only when the chroma block is skipped without transformation. In the case of a two-tree luma structure, NSPT can be applied only when the luma block is skipped without transformation. Furthermore, in the case of a single-tree structure, LFNST can be applied only to the luma component, and NSPT can be applied only when none of the blocks are skipped without transformation.
[0230] Furthermore, similar to LFNST, NSPT can apply 35 or 67 transform sets, and can provide three transform candidates for each transform set. The transform kernel can be derived based on the size of the block to which NSPT is applied.
[0231] Figure 11 An NSPT core based on block size is illustrated as an example.
[0232] refer to Figure 11The NSPT core can include NSPT4×4 for 4×4 blocks, NSPT4×8 for 4×8 blocks, NSPT8×4 for 8×4 blocks, NSPT8×8 for 8×8 blocks, NSPT4×16 for 4×16 blocks, NSPT16×4 for 16×4 blocks, NSPT8×16 for 8×16 blocks, NSPT16×8 for 16×8 blocks, NSPT4×32 for 4×32 blocks, NSPT32×4 for 32×4 blocks, NSPT8×32 for 8×32 blocks, and NSPT32×8 for 32×8 blocks.
[0233] Specifically, Figure 11 (a) can be used to illustrate NSPT applied to 4×N / N×4 blocks. Figure 11 (b) can be used to illustrate NSPT applied to 8×N / N×8 blocks, and Figure 11 (c) can be used to illustrate the transformation applied to 16×N / N×16 blocks. As mentioned above, for 4×N / N×4 blocks, the dimensions of the NSPT kernel can be as follows: - NSPT4×4: 16×16 - NSPT4×8 / NSPT8×4: 32×20 - NSPT8×8: 64×32 - NSPT4×16 / NSPT16×4: 64×24 - NSPT8×16 / NSPT16×8: 128×40 - NSPT4×32 / NSPT32×4: 128×20 - NSPT8×32 / NSPT32×8: 256×24 For example, when applying NSPT4×8 / NSPT8×4, 12 transform coefficients can be zeroed; when applying NSPT8×8, 32 transform coefficients can be zeroed; when applying NSPT4×16 / NSPT16×4, 40 transform coefficients can be zeroed; when applying NSPT8×16 / NSPT16×8, 88 transform coefficients can be zeroed; when applying NSPT4×32 / NSPT32×4, 108 transform coefficients can be zeroed; and when applying NSPT8×32 / NSPT32×8, 232 transform coefficients can be zeroed.
[0234] like Figure 11 As illustrated and described above, NSPT can be applied only to transform blocks of a specific size, and LFNST can be omitted when applying NSPT.
[0235] Additionally, the NSPT index (nspt_idx) of the NSPT core can be signaled only when the transform block size condition is met. In this case, when the NSPT index is 0, it can indicate that NSPT is not applied to the transform block, and when the NSPT index is not signaled, it can be inferred to be 0.
[0236] According to the example, since LFNST is not applied when NSPT is applied, the NSPT index can be signaled before lfnst_idx. For example, when signaling the NSPT index, lfnst_idx can be left unsigned. On the other hand, when NSPT index is left unsigned, or when the NSPT index is 0, lfnst_idx can be signaled to perform LFNST as an inverse quadratic transform, and DCT-II can be applied as an inverse principal transform. In this case, mts_idx can be left unsigned.
[0237] According to another example, `lfnst_idx` can be signaled first, and the NSPT index and `mts_idx` can be signaled after `lfnst_idx`. Furthermore, the NSPT index and `mts_idx` can be signaled when `lfnst_idx` is not signaled or when `lfnst_idx` equals 0. In this case, the NSPT index and `mts_idx` can be signaled sequentially or non-sequentially. Additionally, since the NSPT index should be signaled before resolving the transform coefficients, it can be signaled after the `last_sig_coeff_pos` syntax element, similar to how `lfnst_idx` is signaled. Furthermore, the NSPT index can be signaled at the residual level (i.e., in the residual coding syntax), while `mts_idx` can be signaled at the CU level, TU tree level, or TU level.
[0238] Simultaneously, the NSPT index can be replaced by signaling via the LFNST index. That is, the NSPT index can be notified without signaling, and the NSPT core can be selected based on the value of the LFNST index when a specific transform block size condition is met. For example, as... Figure 8As shown, when a specific block size condition is met, the NSPT (non-separable transform) can be applied to the transform. In this case, the NSPT kernel can be selected based on the value of the LFNST index (or index information indicating the non-separable transform kernel) signaled by the signal. On the other hand, when the transform block size corresponds to a transform block for which NSPT is not applied, LFNST can be applied. In this case, the LFNST kernel can be selected based on the value of the LFNST index (or index information indicating the non-separable transform kernel) signaled by the signal. Furthermore, when LFNST is not applied, a separable transform can be applied. In this case, either the DCT-II transform or the MTS can be applied.
[0239] Simultaneously, the NSPT kernel set and LFNST kernel set can be selected based on the transform block size and intra-prediction mode. For example, NSPT can be used for block shapes of 4×4, 4×8, 4×16, 8×8, 8×16, and 8×32, as well as the corresponding transposed blocks (e.g., 4×4, 8×4, 16×4, 8×8, 16×8, and 32×8 blocks), while LFNST can be used for all other block shapes.
[0240] Furthermore, NSPT can be used with most intra-prediction tools such as regular intra-prediction, DIMD, TIMD, SGPM, MIP, EIP, and IntraTMP. NSPT can also be applied to inter-frame prediction (CU). In this case, the existing NSPT kernel set can be replaced with three kernel sets by additionally introducing two NSPT kernel sets. For example, a first NSPT kernel set can be applied to a regular intra-prediction block, a second NSPT kernel set can be applied to a block using TIMD, DIMD, EIP, MIP, or SGPM, and a third NSPT kernel set can be applied to a block using IntraTMP and inter-frame prediction (CU). That is, three kernel sets can be used depending on the intra-prediction mode, and additional signaling is not required. Furthermore, in cases other than regular intra-prediction modes, such as DIMD, TIMD, SGPM, MIP, EIP, IntraTMP, or inter-frame prediction modes, additional kernel sets can be used without signaling additional information. Simultaneously, information indicating the three kernel sets can be signaled. In addition, a flag can be used to indicate whether an additional NSPT transform kernel set is used for DIMD, TIMD, EIP, MIP, SGPM, etc. For example, when the flag value is 1, the additional transform kernel set can be used, and when the flag value is 0, the additional transform kernel set can be not used.
[0241] Simultaneously, multiple transform set selection (MTSS) for intra-frame LFNST / NSPT can be used. For example, multiple LFNST / NSPT transform sets can be used for DIMD, TIMD, OBIC, SGPM, MIP, EIP, and IntraTMP modes. In this case, the DIMD method used to derive the intra-frame prediction mode for transform set selection can operate on subsampled neighbor samples or prediction block samples. For example, the second intra-frame prediction mode used to select the LFNST / NSPT transform kernel for MIP, EIP, SGPM, and IntraTMP modes can be derived from the second high HoG (Histogram of Oriented Gradients) of the prediction block or neighboring blocks. In this case, the first transform set and the second transform set can be different transform sets or different transform types.
[0242] Furthermore, restrictions can be imposed on the block size and the number of transformation candidates for the second set. For example, MTSS can be used based on the block size. Specifically, MTSS can be applied to CUs with a width × height equal to or greater than 128. Furthermore, for CUs with a width × height less than 256, only the first two set candidates can be used, while for larger CUs, all three set candidates can be used. Here, the block size to which MTSS can be applied is not limited to 128. Additionally, whether to apply MTSS can be determined based on the shape of the current block.
[0243] Furthermore, the MTSS used for LFNST / NSPT may include a first transform set and a second transform set. In this case, the intra-prediction mode used for the first transform set can vary depending on the prediction mode. For example, in the cases of DIMD, OBIC, and TIMD, the first intra-prediction mode used for the first transform set may be the prediction mode used in the current prediction mode, while in the cases of SGPM, MIP, EIP, and IntraTMP, the first intra-prediction mode used for the first transform set may be a prediction mode derived based on HoG. Furthermore, for example, in the cases of DIMD, OBIC, and TIMD, the second intra-prediction mode used for the second transform set may be the prediction mode used in the current prediction mode, while in the cases of SGPM, MIP, EIP, and IntraTMP, the second intra-prediction mode used for the second transform set may be a prediction mode derived based on HoG. Additionally, a flag indicating whether MTSS is used can be signaled. For example, when the flag value is 1, MTSS can be used, and when the flag value is 0, MTSS can not be used. Furthermore, a flag or index indicating the first or second transform set can be signaled. That is, the value of the flag or index can indicate the first transformation set or the second transformation set.
[0244] Simultaneously, the transform set of intra-chroma blocks encoded using LFNST / NSPT and CCP modes can be derived. For example, the intra-prediction mode of the current block can be derived using DIMD based on CCP prediction samples. Specifically, DIMD can be applied to CCP prediction samples, and in this case, the LFNST / NSPT transform set can be determined using a first DIMD mode. For example, the horizontal and vertical gradients can be computed for each prediction sample to derive the HoG (Histogram of Oriented Gradients), and the LFNST / NSPT transform set can be determined using the intra-prediction mode corresponding to the maximum histogram magnitude.
[0245] In the SGPM mode, the partition orientation can be used to determine the transform kernel set. As another example, the VIPM (Virtual Intra-Prediction Mode) can be calculated by applying the DIMD procedure to the SGPM prediction signal, and the LFNST / NSPT transform kernel set can be determined based on the VIPM. In this case, a threshold can be used to compare the SGPM intra-prediction modes with the VIPM. For example, when the VIPM value is close to at least one of the intra-prediction modes in the SGPM, the VIPM can be used, and the LFNST / NSPT transform kernel set can be derived using the VIPM. Otherwise, the orientation of the partition mode can be used to derive the LFNST / NSPT transform kernel set. A threshold can be used to determine whether the VIPM value is close to at least one of the intra-prediction modes in the SGPM. For example, when the difference between the VIPM value and one of the intra-prediction modes in the SGPM is less than N, the VIPM can be used to derive the LFNST / NSPT transform kernel set. Here, N can be an integer greater than or equal to 0.
[0246] The following will describe the extension of LFNST for transformation.
[0247] For example, the LFNST described above can be extended in terms of transform sets and transform kernels. For instance, the number of LFNST transform sets can be 35, and each transform set can consist of three transform kernels, i.e., three transform candidates. The transform set (lfnstTrSetIdx) of the intra-prediction mode can have the mapping relationship shown in the table below.
[0248] [Table 1]
[0249] Referring to Table 1, when the intra-prediction mode (predModeIntra) is less than 0, lfnstTrSetIdx can be 2. When the intra-prediction mode is between 0 and 34, lfnstTrSetIdx can be mapped to values between 0 and 34 respectively. When the intra-prediction mode is between 35 and 66, lfnstTrSetIdx can be mapped to (68 - predModeIntra). Furthermore, when the intra-prediction mode is greater than 66, lfnstTrSetIdx can be mapped to 2. Here, lfnstTrSetIdx can be referred to as the LFNST set index.
[0250] Meanwhile, LFNST cores can include LFNST 4 cores, LFNST 8 cores, and LFNST 16 cores. For example, in addition to LFNST 4 cores and LFNST 8 cores, LFNST 16 cores can also be applied. For instance, LFNST 4 cores can be applied to blocks of size 4×N / N×4 (N≥4), LFNST 8 cores can be applied to blocks of size 8×N / N×8 (N≥8), and LFNST 16 cores can be applied to blocks of size 16×N / N×16 (N≥16).
[0251] Simultaneously, the Region of Interest (ROI) and zeroing can be applied to LFNST. Here, ROI can refer to a specific part of the image that is focused and processed, or it can refer to samples within a dataset identified for a specific purpose. For example, forward LFNST can be applied to an ROI, which is a specific region of interest located in the upper left region of the target transform block. Therefore, when applying LFNST, the principal transform coefficients located outside the ROI can be zeroed.
[0252] Figure 12 An ROI for LFNST 16 is illustrated as an example.
[0253] refer to Figure 12 The ROI for LFNST 16 can consist of six 4×4 sub-blocks arranged consecutively in the scan direction, starting from the top left corner of the target block. Since a total of 96 transform coefficients are input into the forward LFNST, the dimension of the forward LFNST matrix can be R×96. Here, R can be 32, 48, or 64, each less than 96. In one example, R can be 32, and in this case, LFNST 16 can be a 32×96 matrix. Furthermore, for example, when using a 32×96 matrix for LFNST 16, the main transform coefficients located outside the ROI can be set to zero.
[0254] Figure 13 An ROI for LFNST 8 is illustrated as an example.
[0255] refer to Figure 13 The ROI for LFNST 8 can consist of four 4×4 sub-blocks located at the top left corner of the target block, i.e., an 8×8 region at the top left corner of the target block. Since a total of 64 transform coefficients are input into the forward LFNST, the dimension of the forward LFNST matrix can be R×64. Here, R can be 32 or 48, each less than 64. In one example, R can be 32, and in this case, LFNST 8 can be a 32×64 matrix. When using a 32×64 matrix for LFNST 8, the main transform coefficients located outside the ROI can be zeroed out. Furthermore, for example, in the case of an 8×8 block, since the entire block corresponds to the ROI because all 64 transform coefficients are input into the forward LFNST, zeroing out is not required.
[0256] Figure 14 An example is provided of MIP prediction samples used to construct a HoG (Histogram of Oriented Gradients).
[0257] refer to Figure 14 HoG can be constructed using DIMD based on MIP prediction samples. For example, for a target block predicted using MIP or IntraTMP, DIMD can be used to derive the intra-prediction mode of the target block based on the MIP or IntraTMP prediction samples. For example, in the case of a target block applying MIP, DIMD can be applied based on the MIP prediction samples before upsampling. For example, to construct HoG, the horizontal and vertical gradients can be calculated for each prediction sample. Subsequently, the intra-prediction mode corresponding to the maximum histogram magnitude can be used to determine the LFNST transform set and the LFNST transpose flag. Furthermore, the LFNST transpose flag, indicating whether the LFNST kernel is transposed, can be signaled after the LFNST index signaling. Alternatively, for example, the MIP transpose flag (mip_transposed_flag) can be used as the LFNST transpose flag. In this case, signaling the LFNST transpose flag can be omitted.
[0258] Meanwhile, when the intra-prediction mode is IntraTMP, the intra-prediction mode derived through DIMD can be used as the intra-prediction mode for determining the LFNST transform set. In this case, the intra-prediction mode corresponding to the maximum histogram magnitude can be used as the intra-prediction mode for determining the LFNST set.
[0259] Alternatively, according to the example, when the intra-prediction mode is one of IBC, SGPM (Spatial Geometric Partitioning Mode), TIMD (Template-based Intra-mode Derivation), or Palette Mode, DIMD can be applied based on prediction samples derived through the corresponding mode to derive the intra-prediction mode used to determine the LFNST transform set.
[0260] Alternatively, according to the example, LFNST or NSPT can also be applied to inter-frame prediction blocks rather than just intra-frame prediction blocks. For example, a set of transforms can be mapped to corresponding inter-frame prediction modes, and LFNST or NSPT can be performed by applying one of multiple transform kernels associated with the mapped set of transforms. Alternatively, according to the example, DIMD can be applied based on inter-frame prediction samples, and the horizontal and vertical gradients can be computed for each prediction sample to construct a HoG (Histogram of Oriented Gradients). The set of LFNST transforms can then be derived using the prediction modes corresponding to the maximum histogram magnitude.
[0261] Furthermore, for the encoding of LFNST / NSPT coefficients, a modified context model can be used, which uses the first five coefficients in the encoding order instead of adjacent two-dimensional coefficients. That is, when applying LFNST / NSPT, for context modeling or context information derivation related to the currently parsed transform coefficients (e.g., including at least one of sig_coeff_flag, gt1_flag, or gt2_flag), the first five transform coefficients in the encoding (scan) order can be used instead of adjacent coefficients at two-dimensional positions.
[0262] Figure 15 Context modeling for the LFNST and NSPT transform coefficients is illustrated exemplarily.
[0263] refer to Figure 15 In (a), context modeling can be performed on the currently parsed transform coefficient (the 5th coefficient) using adjacent two-dimensional coefficients (the 6th, 7th, 9th, 10th, and 13th). In this case, the transform coefficient to be parsed can be the 5th coefficient, and the adjacent two-dimensional coefficients can be the 6th, 7th, 9th, 10th, and 13th coefficients. (See reference...) Figure 15 (b) Context modeling can be performed on the transform coefficient to be parsed (the 5th coefficient) using the first five coefficients (the 2nd, 3rd, 6th, 9th, and 12th). In this case, the transform coefficient to be parsed can be the 5th coefficient, and the coefficients used for context modeling can be the 2nd, 3rd, 6th, 9th, and 12th coefficients. In this case, the values or absolute values of the first five coefficients can be used. For example, contextual information (e.g., context index or context increment) related to the transform coefficient can be derived based on whether the sum of the values or absolute values of the five coefficients is greater than a threshold. Alternatively, contextual information (e.g., context index or context increment) related to the transform coefficient can be derived based on whether the average of the values or absolute values of the five coefficients is greater than a threshold.
[0264] Simultaneously, since parsing the transform coefficients requires lfnstIdx (lfnst_idx), lfnstIdx can be signaled after all last_sig_coeff_pos syntax elements in the CU. For example, when applying LFNST or NSPT, diagonal reordering can be used to arrange the DCT-II transform coefficients within the coefficient block. Alternatively, when applying LFNST or NSPT, the scan order of the transform coefficients can be changed or derived due to zeroing. According to the example, lfnstIdx and / or the NSPT index can be signaled after last_sig_coeff_pos at the residual coding level.
[0265] Meanwhile, in the case of non-separable transforms such as LFNST or NSPT, one-dimensional directionality, such as the horizontal or vertical direction, may have little or no impact on the encoding of transform coefficients. Therefore, when performing transforms using LFNST or NSPT, the transform coefficients can be transformed after being reordered along the diagonal direction, and reflecting this, contextual modeling of the transform coefficients can be performed using modeling information from coefficients encoded earlier according to the diagonal scan order. Thus, by signaling the LFNST index before transform coefficient encoding, contextual modeling of the transform coefficients can be performed separately depending on whether LFNST (or NSPT) is applied.
[0266] In this case, mts_idx can be signaled at the same level as lfnstIdx and / or the NSPT index (e.g., at the residual coding level), or at the CU level, TU tree level, or TU level. Furthermore, mts_idx can be signaled immediately after lfnstIdx, or immediately after the NSPT index.
[0267] In one example, the transformation-related syntax elements can be shown in the table below.
[0268] [Table 2]
[0269] Referring to Table 2, signals can be used to notify lfnst_idx, nspt_idx, mts_flag and / or mts_idx within the same syntax structure, or signals can be used to notify some of them within different syntax structures.
[0270] For example, after signaling lfnst_idx, you can signal nspt_idx, mts_flag, and / or mts_idx. Furthermore, you can signal nspt_idx or mts_flag and / or mts_idx based on the value of lfnst_idx. For example, lfnst_idx can be signaled when the LFNST application condition is met. Similarly, nspt_idx can be signaled when the value of lfnst_idx is 0 and the NSPT application condition is met; otherwise, mts_flag and / or mts_idx can be signaled.
[0271] As another example, the transformation-related syntax elements can be shown in the table below.
[0272] [Table 3]
[0273] Referring to Table 3, nspt_idx, lfnst_idx, mts_flag and / or mts_idx can be notified by signals within the same syntax structure, or some of them can be notified by signals within different syntax structures.
[0274] For example, after signaling nspt_idx, you can signal lfnst_idx, mts_flag, and / or mts_idx. For example, you can signal nspt_idx when the NSPT application condition is met. Furthermore, for example, you can signal lfnst_idx when the value of nspt_idx is 0 and the LFNST application condition is met. Furthermore, for example, you can signal mts_flag and / or mts_idx when the value of lfnst_idx is 0.
[0275] As another example, the transformation-related syntax elements can be shown in the table below.
[0276] [Table 4]
[0277] Referring to Table 4, signals can be used to notify nspt_idx, lfnst_idx, mts_flag and / or mts_idx within the same syntax structure, or signals can be used to notify some of them within different syntax structures.
[0278] For example, when the NSPT application condition is met, nspt_idx can be notified using a signal. Furthermore, when the NSPT application condition is not met but the LFNST application condition is met, lfnst_idx can be notified using a signal. Additionally, for example, when the value of lfnst_idx is 0, mts_flag and / or mts_idx can be notified using a signal.
[0279] Additionally, lfnst_idx can indicate the LFNST transform kernel, and under certain conditions, it can indicate the NSPT transform kernel. For example, when the NSPT conditions described in this disclosure are met (e.g., the block size condition for NSPT application), the value of lfnst_idx can indicate one of a plurality of NSPT candidates.
[0280] As another example, the transformation-related syntax elements can be shown in the table below.
[0281] [Table 5]
[0282] Referring to Table 5, signals can be used to notify lfnst_idx, nspt_idx, mts_flag and / or mts_idx within the same syntax structure, or signals can be used to notify some of them within different syntax structures.
[0283] For example, when the LFNST application conditions are met, lfnst_idx can be notified using a signal. Otherwise, mts_flag and / or mts_idx can be notified using signals.
[0284] As another example, the transformation-related syntax elements can be shown in the table below.
[0285] [Table 6]
[0286] Referring to Table 6, when the value of lfnst_idx is 0 and the value of nspt_idx is 0, mts_idx can be notified by a signal at the coding unit (CU) level.
[0287] In addition, the signal notifications for the transform_tree syntax are shown in the table below.
[0288] [Table 7]
[0289] Referring to Table 7, signals can be used to notify the transform_unit syntax at the transform_tree syntax level.
[0290] In addition, the signal notification for the residual_coding syntax can be shown in the table below.
[0291] [Table 8]
[0292] Referring to Table 8, the residual_coding syntax can be signaled at the transform_unit syntax level.
[0293] Furthermore, as another example, the transformation syntax can be as follows.
[0294] [Table 9]
[0295] Referring to Table 9, you can notify nspt_idx with a signal after notifying lfnst_idx with a signal.
[0296] For example, when the last valid coefficient is located within the LFNST zeroing region, LfnstZeroOutSigCoeffFlag can be set to 0. Furthermore, when the last valid coefficient is located within the LFNST coefficient region (i.e., outside the LFNST zeroing region), LfnstZeroOutSigCoeffFlag can be set to 1. In this case, the last valid coefficient position can be derived based on information about the last valid coefficient position (e.g., including at least one of last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, or last_sig_coeff_y_suffix).
[0297] Specifically, `last_sig_coeff_x_prefix` represents the column position prefix of the last valid coefficient in the scan order within the transform block, `last_sig_coeff_y_prefix` represents the row position prefix of the last valid coefficient in the scan order within the transform block, `last_sig_coeff_x_suffix` represents the column position suffix of the last valid coefficient in the scan order within the transform block, and `last_sig_coeff_y_suffix` represents the row position suffix of the last valid coefficient in the scan order within the transform block. Here, a valid coefficient can refer to a non-zero coefficient. Although the description is based on a transform block (TB), this is only an example, and TB can be used interchangeably with a code block (CB).
[0298] For example, when the value of LfnstZeroOutSigCoeffFlag is 1, lfnst_idx can be notified by a signal. Furthermore, when the value of lfnst_idx is 0 and the NSPT application conditions are met, nspt_idx can be notified by a signal.
[0299] Furthermore, for example, when applying LFNST and / or NSPT, the signaling for sb_coded_flag[xS][yS] of the last sub-block including the last valid coefficient and the sub-block within the DC sub-block can be omitted, and its value can be derived as 1.
[0300] In addition, for example, when the value of lfnst_idx or nspt_idx is greater than 0, some or all of the values in sb_coded_flag[xS][yS] can be omitted.
[0301] Meanwhile, transformation-related information can be encoded based on context information, and the relevant context information can be represented as follows.
[0302] In one example, even when the NSPT condition (e.g., the block size condition for NSPT application) is met and the value of lfnst_idx can indicate one of multiple NSPT candidates (i.e., when lfnst_idx is substituted for or interchanged with nspt_idx), it may be necessary to use a different context model (context information) depending on whether NSPT or LFNST is applied. For this purpose, a structure such as the following can be used. For example, when the NSPT condition (e.g., the block size condition for NSPT application) is met, ApplyNsptFlag can be set to 1.
[0303] For example, examples of using context-encoded bin to assign a context model of ctxInc to a syntax element can be shown in Tables 10 to 13 below.
[0304] [Table 10]
[0305] [Table 11]
[0306] [Table 12]
[0307] [Table 13]
[0308] Meanwhile, when nspt_idx and lfnst_idx are used separately, examples of the context model for assigning ctxInc to syntax elements using context-encoded bin can be shown in Tables 14 to 17 below.
[0309] [Table 14]
[0310] [Table 15]
[0311] [Table 16]
[0312] [Table 17]
[0313] The following describes the enhanced MTS (Multi-Transform Selection) for intra-frame prediction.
[0314] For example, when applying MTS to an intra-prediction block, the transform kernel can include DCT-II, DST-VII, and DCT-VIII, and additionally DCT-V and DST-IV can be applied. Furthermore, when applying MTS to an intra-prediction block, the transform kernel can include DCT-II, DST-VII, and DCT-VIII, and additionally DCT-V, DST-IV, DST-I, and the identity transform (IDT) can be applied.
[0315] Furthermore, the intra-MTS candidate, or MTS set, can be derived based on the TU size and intra-prediction mode information. For example, a total of 16 TU sizes can be supported, and each TU size can be classified into five groups (classes) according to the intra-prediction mode. In this case, if the five groups are applied to each of the 16 TU sizes, a total of 80 groups can be considered. However, since the transform set can be shared between groups, fewer than 80 groups can be considered, such as 58 groups. Moreover, the number of groups is not limited to 58, and N groups can be considered, where N is a positive integer.
[0316] Furthermore, when the intra-prediction mode is an angular mode, the symmetry between the TU shape and the intra-prediction direction can be considered. For example, mode i associated with block A×B and mode j associated with block B×A (where i>34 and j=68-i) can be mapped to the same group. In this case, the transform pairs of the vertical and horizontal kernels can be swapped. For example, a 16×4 block with intra-prediction mode number 18 and a 4×16 block with intra-prediction mode number 50 can be mapped to the same group. When the intra-prediction mode is a wide-angle intra-prediction mode, the nearest regular angular mode can be used for transform set determination. For example, to derive the group used for transform set determination, wide-angle intra-prediction modes between -2 and -14 can use mode 2, while wide-angle intra-prediction modes between 67 and 80 can use mode 66.
[0317] Furthermore, multiple transform candidates (transformation pairs) can be applied to each group. For example, one, four, or six transform sets can be applied. The number of transform sets to be applied can be determined based on the position of the last transform coefficient or based on the absolute value of the transform coefficient. For example, the position of the last transform coefficient can be compared with two thresholds, and based on the comparison result, one, four, or six transform sets can be selected. Alternatively, the sum of the absolute values of the transform coefficients can be compared with two thresholds, and one, four, or six transform sets can be selected.
[0318] Furthermore, the number of transformation sets can be determined based on the sum of the absolute values of the transformation coefficients. For example, when the sum of the absolute values of the transformation coefficients is less than th0, one transformation set can be applied; when the sum is greater than th0 and less than or equal to th1, four transformation sets can be applied; and when the sum is greater than th1, six transformation sets can be applied (one candidate: sum ≤ th0, four candidates: th0 < sum ≤ th1, six candidates: sum > th1). Here, sum can represent the sum of the absolute values of the transformation coefficients. Additionally, th0 can be set to 6, and th1 can be set to 32.
[0319] Figure 16 An example of deriving the MTS set is shown.
[0320] refer to Figure 16 The four transform sets can be determined based on the TU size and intra-frame prediction mode information.
[0321] For example, there can be a total of 16 TU sizes and a total of 36 intra-frame modes available, including intra-prediction modes 0 to 34 and an additional MIP mode. As mentioned above, based on the intra-prediction mode information, the intra-prediction mode of the transform block can be mapped to one of the intra-prediction modes 1 to 34. A total of five groups can be mapped for each TU size. For example, intra-prediction modes 0 and 1, intra-prediction modes 2 to 12, intra-prediction modes 13 to 23, intra-prediction modes 24 to 34, and the MIP mode can each be mapped to their respective groups.
[0322] For example, when the TU size is 4×8 and the intra-prediction mode is 5, four transform sets can be determined. In this case, when the MTS index information is 1, number 12 can be selected, and a transform set consisting of DCT-V and DCT-V can be selected.
[0323] Furthermore, transform pairs can be formed by pairing transform kernels selected from DCT-VIII, DST-VII, DCT-V, DST-IV, and DST-I, and each transform pair can be indexed by a value from 0 to 24.
[0324] As described above, when four sets of transformations are applied to a single group, the set of transformations consisting of four pairs of transformations can be applied to each of the 80 groups. In this case, some of the 80 groups can share the transformation set, and the number of groups can be reduced from 80 to 58.
[0325] Figure 17 An example is given of deriving a DIMD-based intra-frame mode for determining the MTS and LFNST sets.
[0326] refer to Figure 17 This allows us to derive DIMD-based intra-frame modes for determining the MTS and LFNST sets.
[0327] Meanwhile, when the intra-prediction mode is IntraTMP, the intra-prediction mode derived through DIMD can be used as the intra-prediction mode for determining the MTS transform set. For example, the intra-prediction mode corresponding to the maximum histogram magnitude can be used as the intra-prediction mode for determining the MTS and LFNST sets.
[0328] Meanwhile, according to another example, the intra-frame mode for IntraTMP can be derived from the reference block or from the neighboring blocks of the reference block.
[0329] Furthermore, when the intra-prediction mode is one of IBC, SGPM (Spatial Geometric Partitioning Mode), TIMD (Template-based Intra-mode Derivation), or Palette Mode, DIMD can be applied based on the prediction samples derived through the corresponding mode to derive the intra-prediction mode used to determine the MTS transform set.
[0330] Furthermore, according to the example, MTS can be applied even when the transform block size is greater than 32. Alternatively, considering complexity, intra-prediction modes can be grouped into three or two groups for each transform block size instead of five groups. Alternatively, in addition to Figure 16 In addition to the transformation pairs shown, identity transformations (IDTs) can also be applied. For example, one of the vertical and horizontal transformations can be as follows: Figure 16 One of DST-VII, DCT-VIII, DCT-V, DST-IV, and DST-I is shown, while the other can be an identity transformation. In this case, the number of possible transformation pairs can increase.
[0331] In addition, the signaling for the MTS index (mts_idx) can be executed as follows.
[0332] Based on the example, `mts_enabled_flag` can be signaled at a higher level, and both `mts_flag` and `mts_idx` can be signaled at the CU level or the residual coding level. In this case, when `mts_flag` equals 1, `mts_idx` can be signaled and the transform set can be determined. When `mts_flag` equals 0, DCT-II can be applied as the master transform.
[0333] Furthermore, when four transform sets exist, one of the four transform pairs can be selected via `mts_idx`, and when six transform sets exist, one of the six transform pairs can be selected via `mts_idx`. In this case, when only one transform set exists, the first transform pair configured with four transform sets or the first transform pair configured with six transform sets can be selected without additional signaling. That is, when one transform set exists, `mts_flag` with a value of 1 can be signaled, and `mts_idx` can be omitted. Furthermore, when `mts_idx` is not signaled, a specific transform pair can be used, or `mts_idx` can be deduced to be 0 or 1. Alternatively, when one transform set exists, `mts_flag` with a value of 1 can be signaled, and `mts_idx=0` can indicate the application of the first transform pair configured with four transform sets, while `mts_idx=1` can indicate the application of the first transform pair configured with six transform sets.
[0334] For example, the binarization of the MTS index (mts_idx) value can be shown in the table below.
[0335] [Table 18]
[0336] Furthermore, according to another example, when mts_idx equals 0, it can indicate that MTS is not applied; that is, DCT-II is applied as the main transform without signaling mts_flag. In this case, the binarization of the MTS index can be shown in the table below.
[0337] [Table 19]
[0338] MTS indexes can be binarized using truncated Rice (or truncated unary) coding, as shown in the table above, or they can be binarized based on a fixed-length coding scheme.
[0339] Alternatively, according to another example, when a transformation set exists, a specific transformation pair can be used instead of one of the transformation pairs belonging to a configuration of four or six transformation sets, and in this case, mts_idx can indicate that specific transformation pair. In this case, mts_idx indicating a single transformation set can be signaled as 6 in Table 2 and as 7 in Table 3. Alternatively, in Table 2 or Table 3, mts_idx can be signaled with an intermediate value, such as 3 or 4.
[0340] Alternatively, when applying the identity transformation (IDT) to select one of the six transformation kernels, information indicating whether the IDT is applied in the horizontal or vertical direction can be appended after mts_idx. For example, idt_flag can be used to indicate whether the IDT is applied, and when idt_flag equals 1, information indicating the direction of application of the identity transformation can be further indicated by a signal. Alternatively, flag information indicating whether the identity transformation is applied in the horizontal and vertical directions can be used separately.
[0341] Alternatively, when applying the identity transformation (IDT), the `mts_inter_enabled_flag` or `mts_intra_enabled_flag` can be signaled at a higher level, and at a lower level, the kernel index indicating one of the six kernels can be signaled separately for each direction. For example, the kernel index for the horizontal direction can be signaled as `mts_horizontal_idx`, and the kernel index for the vertical direction can be signaled as `mts_vertical_idx`.
[0342] Figure 18 A video / image coding method according to one or more embodiments of the present disclosure is illustrated schematically. Figure 18 The method disclosed in the document can be used by [the party in question]. Figure 2 The encoding device disclosed in the document executes the commands. Specifically, for example, Figure 18 S1800 can be executed by the predictor 220 of the encoding device 200, S1810 to S1830 can be executed by the residual processor 230 of the encoding device 200, and S1840 can be executed by the entropy encoder 240 of the encoding device 200. Figure 18 The methods disclosed herein may include the embodiments described above.
[0343] refer to Figure 18 The encoding device derives the predicted sample for the current block (S1800). For example, the encoding device can derive the predicted sample for the current block.
[0344] The encoding device derives the residual sample of the current block (S1810). For example, the encoding device can derive the residual sample of the current block based on the prediction sample.
[0345] The encoding device derives the transform coefficients of the current block (S1820). The encoding device can derive the transform coefficients of the current block by performing a transform based on the residual samples. For example, the transform can be the master transform.
[0346] For example, a transform pair can be determined from the transform pair candidates in the transform set for the current block based on MTS-related information. In this case, the number of transform pair candidates in the transform set can be at least one. For example, the number of transform pair candidates in the transform set can be 1, 4, or 6. Furthermore, the selected transform pair can be configured based on the kernel associated with DCT-II, DCT-V, DCT-VIII, DST-IV, DST-VII, or Identity Transformation (IDT).
[0347] Furthermore, transform pairs can be determined based on intra-prediction modes. For example, prediction samples can be derived based on whether the prediction mode of the current block is IBC (Intra-Block Copy), SGPM (Spatial Geometric Partitioning), TIMD (Template-Based Intra-Mode Derivation), or Palette Mode. In this case, the intra-prediction mode can be derived using the DIMD mode based on the prediction samples, and transform pairs can be determined based on the derived intra-prediction mode.
[0348] Furthermore, MTS can be applied even when the transform block size is greater than 32. In this case, the transform block size is not limited to greater than 32. Additionally, for each transform block size, intra-prediction modes can be grouped into two or three groups instead of five. Furthermore, in addition to the transform pairs shown in Table 19, identity transform (IDT) can also be applied. In this case, one of the vertical transform and the horizontal transform can be one of DST-VII, DCT-VIII, DCT-V, DST-IV, and DST-I as shown in Table 19, while the other can be the identity transform. Therefore, the number of possible transform pairs can be increased.
[0349] The encoding device generates residual information for the current block (S1830). For example, the encoding device may generate residual information for the current block based on the transform coefficients.
[0350] The encoding device encodes image information including residual information (S1840). For example, the encoding device can encode image information including residual information.
[0351] For example, image information may include MTS (Multiple Transform Selection) related information. In this case, MTS related information may include at least one of MTS enable flag information, MTS flag information, or MTS index information. The MTS flag information and MTS index information can be signaled at the coding unit level or the residual coding level, while the MTS enable flag information can be signaled at a level higher than the MTS flag information and MTS index information. That is, the MTS index information can be signaled in the coding unit syntax or the residual coding syntax, and the MTS enable flag information can be signaled at a syntax level higher than the MTS flag information and MTS index information. Furthermore, the MTS enable flag information may correspond to the syntax element "mts_enabled_flag", the MTS flag information may correspond to the syntax element "mts_flag", and the MTS index information may correspond to the syntax element "mts_idx". In this case, truncated Rice coding or truncated unary coding can be used to binarize the MTS index information. Alternatively, a fixed-length coding scheme can be used to binarize the MTS index information.
[0352] Furthermore, a transform pair can be determined from the transform pair candidates in the transform set for the current block based on MTS-related information. In this case, the number of transform pair candidates in the transform set can be at least one. For example, the number of transform pair candidates in the transform set can be 1, 4, or 6. Moreover, the selected transform pair can be configured based on the kernel associated with DCT-II, DCT-V, DCT-VIII, DST-IV, DST-VII, or the identity transform (IDT).
[0353] Furthermore, the signaling related to MTS can be determined based on the number of candidate transformation sets.
[0354] For example, given that the number of candidate transforms is 1, the value of the MTS flag information can be signaled as 1, and the MTS index information can be signaled without signaling. In this case, based on not signaling the MTS index information, a specific transform pair can be used, or the value of the MTS index information can be derived. When not signaling the MTS index information, the value of the MTS index information can be derived as 0 or 1.
[0355] Furthermore, for example, based on the number of transformation set candidates being 1, the first transformation pair candidate among four transformation set transformation pair candidates or the first transformation pair candidate among six transformation set transformation pair candidates can be selected.
[0356] For example, if the number of transform set candidates is 1, the MTS flag information with a value of 1 can be used to notify the user. In this case, if the MTS index information is 0, the first transform pair candidate among the four transform set transform pair candidates can be selected. Furthermore, if the MTS index information is 1, the first transform pair candidate among the six transform set transform pair candidates can be selected.
[0357] Furthermore, since the number of candidate transform sets is 1, a specific transform pair can be used instead of four or six transform sets. In this case, the MTS index information can indicate that specific transform pair. For example, the MTS index information indicating a single transform set can be signaled as 6 in Table 18 above, and signaled as 7 in Table 19 above. Alternatively, the MTS index information can be signaled as an intermediate value in Table 18 or Table 19. In this case, the intermediate value can be 3 or 4. Here, a single transform set can be a specific transform set.
[0358] Furthermore, since no MTS flag information is required and the MTS index information value is 0, MTS can be omitted from the transform. In this case, the transform coefficients can be transformed based on DCT-II.
[0359] Furthermore, when applying the identity transformation (IDT) to a transformation, the kernel associated with the identity transformation can be selected as one of the six transformation kernels. In this case, information indicating whether the identity transformation is applied in the horizontal or vertical direction can be signaled after the MTS index information.
[0360] In addition, the image information may include flag information indicating whether an identity transformation (IDT) is applied to the transform coefficients. Based on a flag value of 1, information related to the direction of application of the identity transformation can be signaled. Furthermore, the image information may include information related to whether the identity transformation is applied to the transform coefficients in the horizontal direction and information related to whether the identity transformation is applied to the current block in the vertical direction. Here, the information indicating whether the identity transformation is applied may correspond to the syntax element "idt_flag".
[0361] Furthermore, based on applying the identity transform (IDT) to the transform, the image information may include MTS inter-frame enable flag information or MTS intra-frame enable flag information. In this case, the kernel index information indicating one of the six transform kernels can be individually signaled according to the direction of the identity transform application. For example, the kernel index information can be signaled at a syntax level lower than that of the MTS inter-frame enable flag information or the MTS intra-frame enable flag information. In this case, the kernel index information may include horizontal kernel index information and / or vertical kernel index information. Here, the MTS inter-frame enable flag information may correspond to the syntax element "mts_inter_enabled_flag", the MTS intra-frame enable flag information may correspond to the syntax element "mts_intra_enabled_flag", the horizontal kernel index information may correspond to "mts_horizontal_idx", and the vertical kernel index information may correspond to "mts_vertical_idx".
[0362] Furthermore, MTS can be applied even when the transform block size is greater than 32. In this case, the transform block size is not limited to greater than 32. Additionally, for each transform block size, intra-prediction modes can be grouped into two or three groups instead of five. Furthermore, in addition to the transform pairs shown in Table 19, identity transform (IDT) can also be applied. In this case, one of the vertical transform and the horizontal transform can be one of DST-VII, DCT-VIII, DCT-V, DST-IV, and DST-I as shown in Table 19, while the other can be the identity transform. Therefore, the number of possible transform pairs can be increased.
[0363] Furthermore, the image information may include various types of information according to embodiments of this disclosure. For example, the image information may include information disclosed in at least one of the tables above.
[0364] Furthermore, the encoded image information can be output as a bitstream. The bitstream can be sent to a decoding device via a network or storage medium.
[0365] Furthermore, as mentioned above, the encoding device can generate reconstructed images (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is because the encoding device derives the same prediction results as the decoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed images (or reconstructed samples or reconstructed blocks) in memory and use them as reference images for inter-frame prediction. As mentioned above, loop filtering processes, etc., can be further applied to the reconstructed images.
[0366] According to one or more of the above embodiments, in the cases of IBC, SGPM, TIMD, palette mode, etc., the MTS transform set can be determined based on the intra-prediction mode derived by applying DIMD. Therefore, the MTS transform set can be adaptively determined for various intra-prediction modes, thereby improving transform efficiency.
[0367] Furthermore, according to one or more of the above embodiments, transformation pairs can be efficiently determined by signaling MTS index information, signaling MTS related information, and / or whether to apply identity transformation (IDT) based on the number of transformation sets, thereby achieving efficient transformation processing.
[0368] Figure 19 A video / image decoding method according to one or more embodiments of the present disclosure is illustrated schematically. Figure 19 The method disclosed in the document can be used by [the party in question]. Figure 3 The decoding device disclosed in the document performs the operation. Specifically, for example, Figure 19 S1900 can be executed by the entropy decoder 310 of the decoding device 300, S1910 to S1920 can be executed by the residual processor 320 of the decoding device 300, and S1930 can be executed by the adder 340 of the decoding device 300. Figure 19 The methods disclosed herein may include the embodiments described above.
[0369] refer to Figure 19 The decoding device receives image information including residual information for the current block (S1900). For example, the decoding device can receive image information including residual information for the current block via a bitstream. As mentioned above, the image information may also include prediction-related information.
[0370] For example, image information may include MTS (Multiple Transform Selection) related information. In this case, the MTS related information may include at least one of MTS enable flag information, MTS flag information, or MTS index information. MTS flag information and MTS index information may be included at the coding unit level or the residual coding level, while MTS enable flag information may be included at a level higher than the levels of MTS flag information and MTS index information. That is, MTS index information can be signaled in the coding unit syntax or the residual coding syntax, and MTS enable flag information can be signaled at a syntax level higher than the syntax levels of MTS flag information and MTS index information. Furthermore, MTS enable flag information may correspond to the syntax element "mts_enabled_flag", MTS flag information may correspond to the syntax element "mts_flag", and MTS index information may correspond to the syntax element "mts_idx". In this case, truncated Rice coding or truncated unary coding can be used to binarize the MTS index information. Alternatively, a fixed-length coding scheme can be used to binarize the MTS index information.
[0371] Furthermore, a transform pair can be determined from the transform pair candidates in the transform set for the current block based on MTS-related information. In this case, the number of transform pair candidates in the transform set can be at least one. For example, the number of transform pair candidates in the transform set can be 1, 4, or 6. Moreover, the selected transform pair can be configured based on the kernel associated with DCT-II, DCT-V, DCT-VIII, DST-IV, DST-VII, or the identity transform (IDT).
[0372] Furthermore, the signal notification for MTS-related information can be determined based on the number of candidates in the transformation set.
[0373] For example, if the number of candidate transforms is 1, the value of the MTS flag information can be 1, and the MTS index information may not be included in the image information. In this case, since the MTS index information is not included in the image information, a specific transform pair can be used, or the value of the MTS index information can be derived. When the MTS index information is not notified by a signal, the value of the MTS index information can be derived as 0 or 1.
[0374] Furthermore, for example, based on the number of transformation set candidates being 1, the first transformation pair candidate among four transformation set transformation pair candidates or the first transformation pair candidate among six transformation set transformation pair candidates can be selected.
[0375] For example, if the number of transform set candidates is 1, the value of the MTS flag information can be 1. In this case, if the MTS index information is 0, the first transform pair candidate among the four transform set transform pair candidates can be selected. Furthermore, if the MTS index information is 1, the first transform pair candidate among the six transform set transform pair candidates can be selected.
[0376] Furthermore, since the number of candidate transform sets is 1, a specific transform pair can be used instead of four or six transform sets. In this case, the MTS index information can indicate that specific transform pair. For example, the MTS index information indicating a single transform set can be signaled as 6 in Table 18 above, and signaled as 7 in Table 19 above. Alternatively, the MTS index information can be signaled as an intermediate value in Table 18 or Table 19. In this case, the intermediate value can be 3 or 4. Here, a single transform set can be a specific transform set.
[0377] The decoding device derives the transform coefficients of the current block (S1910). For example, the decoding device can derive the transform coefficients of the current block based on residual information.
[0378] The decoding device derives the residual samples of the current block (S1920). For example, the decoding device may perform an inverse transform based on the transform coefficients to derive the residual samples of the current block. For example, the inverse transform may be the inverse principal transform.
[0379] For example, the number of candidate transform pairs in the transform set can be at least one. For example, the number of candidate transform pairs in the transform set can be 1, 4, or 6. Furthermore, the selected transform pairs can be configured based on the kernels associated with DCT-II, DCT-V, DCT-VIII, DST-IV, DST-VII, or identity transforms (IDT).
[0380] Furthermore, transform pairs can be determined based on intra-prediction modes. For example, prediction samples can be derived based on whether the prediction mode of the current block is IBC (Intra-Block Copy), SGPM (Spatial Geometric Partitioning), TIMD (Template-Based Intra-Mode Derivation), or Palette Mode. In this case, the intra-prediction mode can be derived using the DIMD mode based on the prediction samples, and transform pairs can be determined based on the derived intra-prediction mode.
[0381] Furthermore, MTS can be applied even when the transform block size is greater than 32. In this case, the transform block size is not limited to greater than 32. Additionally, for each transform block size, intra-prediction modes can be grouped into two or three groups instead of five. Furthermore, in addition to the transform pairs shown in Table 19, identity transform (IDT) can also be applied. In this case, one of the vertical transform and the horizontal transform can be one of DST-VII, DCT-VIII, DCT-V, DST-IV, and DST-I as shown in Table 19, while the other can be the identity transform. Therefore, the number of possible transform pairs can be increased.
[0382] Furthermore, since the MTS flag information is not included in the image information and the MTS index information value is 0, the MTS can be omitted from the inverse transform. In this case, the inverse transform can be performed on the transform coefficients based on DCT-II.
[0383] Furthermore, when applying the identity transformation (IDT) to the inverse transform, the kernel associated with the identity transformation can be selected as one of the six transformation kernels. In this case, information indicating whether the identity transformation is applied in the horizontal or vertical direction can be signaled after the MTS index information.
[0384] Furthermore, the image information may include flag information indicating whether the identity transformation is applied to the transform coefficients. Based on a flag value of 1, information related to the direction of application of the identity transformation can be signaled. Additionally, the image information may include information related to whether the identity transformation is applied to the transform coefficients in the horizontal direction and information related to whether the identity transformation is applied to the current block in the vertical direction. Here, the information indicating whether the identity transformation is applied may correspond to the syntax element "idt_flag".
[0385] Furthermore, based on the application of the identity transform (IDT) to the inverse transform, the image information may include MTS inter-frame enable flag information or MTS intra-frame enable flag information. In this case, kernel index information indicating one of the six transform kernels can be derived individually based on the direction in which the identity transform is applied. For example, the kernel index information may be included at a level lower than the MTS inter-frame enable flag information or the MTS intra-frame enable flag information. In this case, the kernel index information may include horizontal kernel index information and / or vertical kernel index information. Here, the MTS inter-frame enable flag information may correspond to the syntax element "mts_inter_enabled_flag", the MTS intra-frame enable flag information may correspond to the syntax element "mts_intra_enabled_flag", the horizontal kernel index information may correspond to "mts_horizontal_idx", and the vertical kernel index information may correspond to "mts_vertical_idx".
[0386] Furthermore, MTS can be applied even when the transform block size is greater than 32. In this case, the transform block size is not limited to greater than 32. Additionally, for each transform block size, intra-prediction modes can be grouped into two or three groups instead of five. Furthermore, in addition to the transform pairs shown in Table 19, identity transform (IDT) can also be applied. In this case, one of the vertical transform and the horizontal transform can be one of DST-VII, DCT-VIII, DCT-V, DST-IV, and DST-I as shown in Table 19, while the other can be the identity transform. Therefore, the number of possible transform pairs can be increased.
[0387] The decoding device generates a reconstruction sample for the current block (S1930). For example, the decoding device may generate the reconstruction sample for the current block based on the residual samples. In this case, the decoding device may generate the reconstruction sample for the current block based on the residual samples and the predicted samples of the current block. Furthermore, according to the example, the decoding device may generate a reconstructed image including the reconstruction samples. Thereafter, as described above, the decoding device may apply a loop filtering process, such as a deblocking filtering process and / or a SAO (Sample Adaptive Shift) process, to the reconstructed image as needed to improve subjective and / or objective image quality.
[0388] According to one or more of the above embodiments, in the cases of IBC, SGPM, TIMD, palette mode, etc., the MTS transform set can be determined based on the intra-prediction mode derived by applying DIMD. Therefore, the MTS transform set can be adaptively determined for various intra-prediction modes, thereby improving transform efficiency.
[0389] Furthermore, according to one or more of the above embodiments, transformation pairs can be efficiently determined by signaling MTS index information, signaling MTS related information, and / or whether to apply identity transformation (IDT) based on the number of transformation sets, thereby achieving efficient transformation processing.
[0390] Although the method is described as a series of steps or blocks based on the flowchart in the above embodiments, the embodiments are not limited to the order of the steps, and a step may occur in a different order than another step described above or simultaneously with another step described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this disclosure.
[0391] The methods described above according to embodiments of this disclosure can be implemented in software, and the encoding and / or decoding devices according to this disclosure can be included in devices for performing image processing, such as televisions, computers, smartphones, set-top boxes, and display devices.
[0392] The embodiments described above can be implemented in the form of a recording medium including computer-executable (program) instructions, such as a program module executed by a computer. The module can be stored in memory and executed by a processor. The memory can be located inside or outside the processor and can be connected to the processor by various known means. The computer-readable medium can be any available medium accessible to a computer and can include both volatile and non-volatile media, as well as removable and non-removable media. Furthermore, the computer-readable medium can include both computer storage media and communication media. Computer storage media can include both volatile and non-volatile media, as well as removable and non-removable media, implemented using any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Communication media typically include computer-readable instructions, data structures, program modules, other data in modulated data signals (such as carrier waves), or other transmission mechanisms, and include any information transmission medium.
[0393] Furthermore, the embodiments described above in this disclosure can be implemented as a computer program (or computer program product) including computer-executable instructions. The computer program may include programmable machine instructions processed by a processor and may be implemented in a high-level programming language, an object-oriented programming language, assembly language, or machine language. Additionally, the computer program may be recorded on a tangible computer-readable recording medium (e.g., memory, hard disk, magnetic / optical media, or solid-state drive (SSD)).
[0394] Therefore, embodiments of this disclosure can be implemented by executing the computer program described above using a computing device. The computing device may include at least some of a processor, memory, storage devices, high-speed interfaces connected to the memory and high-speed expansion ports, and low-speed interfaces connected to low-speed buses and storage devices. These components may be interconnected via various buses and may be mounted on a common motherboard or otherwise suitably configured.
[0395] A processor can process instructions within a computing device. These instructions may include those stored in memory or storage devices to display graphical information on an external input / output device (such as a display) connected to a high-speed interface, providing a graphical user interface (GUI). In another embodiment, multiple processors and / or multiple buses may be appropriately utilized along with multiple memories and memory types. Furthermore, the processor may be implemented as a chipset comprising multiple independent analog and / or digital processors.
[0396] Memory stores information within a computing device. For example, memory may include volatile memory cells or a collection of volatile memory cells. In another example, memory may include non-volatile memory cells or a collection of non-volatile memory cells. Memory may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0397] Storage devices can provide large-capacity storage space for computing devices. Storage devices can be computer-readable media or components that include computer-readable media. For example, storage devices can include devices or other components within a storage area network (SAN) and can be floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory, other similar semiconductor storage devices, or device arrays.
[0398] The network can be implemented as a wired network, such as a local area network (LAN), a wide area network (WAN), or a value-added network (VAN), or various types of wireless networks, such as mobile radio communication networks or satellite communication networks.
[0399] Although this disclosure has been described with reference to embodiments illustrated in the accompanying drawings, these embodiments are merely exemplary. Those skilled in the art will understand that various modifications and variations are possible. That is, the scope of this disclosure is not limited to the described embodiments, and various modifications and alterations made by those skilled in the art based on the basic concepts defined in the appended claims also fall within the scope of the claims. Therefore, the true technical scope of this disclosure should be determined by the technical spirit of the appended claims.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Receive image information including residual information for the current block; Based on the residual information, the transformation coefficients for the current block are derived; The residual sample for the current block is derived by performing an inverse transformation based on the transformation coefficients; as well as Based on the residual samples, a reconstruction sample is generated for the current block. The image information includes Multi-Transform Selection (MTS) related information. Specifically, based on the MTS-related information, a transformation pair is determined from the candidate transformation pairs in the transformation set for the current block, and The number of candidate transformation pairs in the transformation set is at least one.
2. The image decoding method according to claim 1, in, The transformation applies to one of the following quantities: 1, 4, and 6. The transformation pairs are based on kernel configurations associated with DCT-II, DCT-V, DCT-VIII, DST-IV, DST-VII, or the identity transformation IDT.
3. The image decoding method according to claim 1, in, The prediction sample is derived based on whether the prediction mode for the current block is Intra-Block Copy (IBC), Spatial Geometric Partitioning (SGPM), Template-Based Intra-Mode Derivation (TIMD), or Palette Mode. Specifically, based on the predicted samples, the decoder-side intra-frame mode is used to derive the DIMD mode to derive the intra-frame prediction mode, and... The transform pair is determined based on the intra-frame prediction mode.
4. The image decoding method according to claim 1, in, The MTS-related information includes at least one of MTS activation flag information, MTS flag information, and MTS index information. The MTS flag information and the MTS index information are included in the coding unit syntax or residual coding syntax, and The MTS enable flag information is included at a syntax level that is higher than the syntax level of the MTS flag information and the MTS index information.
5. The image decoding method according to claim 4, in, Since the number of transform set candidates is 1, the value of the MTS flag information is 1, and the MTS index information is not included in the image information.
6. The image decoding method according to claim 5, in, Based on the fact that the MTS index information is not included in the image information, the value of the MTS index information is derived using specific transform pairs, and The value of the MTS index information is deduced to be 0 or 1.
7. The image decoding method according to claim 4, in, Since the number of transformation set candidates is 1, the first transformation pair candidate among the four transformation set transformation pair candidates or the first transformation pair candidate among the six transformation set transformation pair candidates is selected.
8. The image decoding method according to claim 4, in, Since the number of candidate transform sets is 1, the value of the MTS flag information is 1. Specifically, based on the MTS index information being 0, the first transformation pair candidate is selected from the four transformation pair candidate sets, and Specifically, based on the MTS index information being 1, the first transformation pair candidate is selected from the six transformation pair candidates.
9. The image decoding method according to claim 4, in, Since the MTS flag information is not included in the image information and the value of the MTS index information is 0, MTS is not applied to the inverse transform, and Specifically, the inverse transform is performed on the transform coefficients based on DCT-II.
10. The image decoding method according to claim 4, in, Based on applying the identity transformation IDT to the inverse transformation, a kernel associated with the identity transformation is selected as one of the six transformation kernels, and The MTS index information is followed by a signal indicating whether the identity transformation is applied in the horizontal or vertical direction.
11. The image decoding method according to claim 10, in, The image information includes flag information indicating whether the identity transformation is applied to the transform coefficients, and Wherein, based on the value of the flag information being 1, information related to the direction of applying the identity transformation is signaled.
12. The image decoding method according to claim 11, in, The image information includes information related to whether the identity transformation is applied to the transformation coefficients in the horizontal direction and information related to whether the identity transformation is applied to the transformation coefficients in the vertical direction.
13. The image decoding method according to claim 1, in, Based on applying the identity transformation IDT to the inverse transform, the image information includes MTS inter-frame enable flag information or MTS intra-frame enable flag information. Specifically, the kernel index information indicating the six transformation kernels is derived individually based on the direction of the applied identity transformation, and The kernel index information includes MTS horizontal index information and MTS vertical index information.
14. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Derive the prediction samples for the current block; Based on the predicted samples, derive the residual samples for the current block; The transformation coefficients for the current block are derived by performing a transformation based on the residual samples; Residual information for the current block is generated based on the transformation coefficients; as well as The image information, including the residual information, is encoded. The image information includes Multi-Transform Selection (MTS) related information. Specifically, based on the MTS-related information, a transformation pair is determined from the candidate transformation pairs in the transformation set for the current block, and The number of candidate transformation pairs in the transformation set is at least one.
15. The image encoding method according to claim 14, in, The transformation applies to one of the following quantities: 1, 4, and 6. The transformation pairs are based on kernel configurations associated with DCT-II, DCT-V, DCT-VIII, DST-IV, DST-VII, or the identity transformation IDT.
16. The image encoding method according to claim 14, in, The prediction sample is derived based on whether the prediction mode for the current block is Intra-Block Copy (IBC), Spatial Geometric Partitioning (SGPM), Template-Based Intra-Mode Derivation (TIMD), or Palette Mode. Specifically, based on the predicted samples, the decoder-side intra-frame mode is used to derive the DIMD mode to derive the intra-frame prediction mode, and... The transform pair is determined based on the intra-frame prediction mode.
17. The image encoding method according to claim 14, in, The MTS-related information includes at least one of MTS activation flag information, MTS flag information, and MTS index information. Specifically, the MTS flag information and the MTS index information are communicated via signals in the coding unit syntax or residual coding syntax, and Specifically, at a syntax level higher than the MTS flag information and the MTS index information, the MTS enable flag information is signaled.
18. The image encoding method according to claim 17, in, Since the number of candidates in the transformation set is 1, the value of the MTS flag information is signaled as 1, and the MTS index information is not signaled.
19. The image encoding method according to claim 17, in, Since the number of transformation set candidates is 1, the first transformation pair candidate among the four transformation set transformation pair candidates or the first transformation pair candidate among the six transformation set transformation pair candidates is selected.
20. A method for transmitting data for an image, the method comprising the steps of: A bitstream for the image is obtained, the bitstream being generated based on the following operations: deriving a prediction sample for the current block, deriving a residual sample for the current block based on the prediction sample, deriving transform coefficients for the current block by performing a transform based on the residual sample, generating residual information for the current block based on the transform coefficients, and encoding image information including the residual information. as well as Send the data including the bit stream. The image information includes Multi-Transform Selection (MTS) related information. Specifically, based on the MTS-related information, a transformation pair is determined from the candidate transformation pairs in the transformation set for the current block, and The number of candidate transformation pairs in the transformation set is at least one.