BDPCM-based video coding method and device
The BDPCM-based video coding method enhances compression efficiency by applying BDPCM separately to luma and chroma components and optimizing transform processes, addressing inefficiencies in high-resolution video coding.
Patent Information
- Application Number
- JP2024048161
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-20
- Filing Date
- 2024-03-25
- Publication Date
- 2025-12-03
- Estimated Expiration
- 2040-04-20
AI Technical Summary
Existing video coding technologies face inefficiencies in compressing high-resolution, high-quality images and videos, particularly in handling immersive media like VR and AR content, leading to increased transmission and storage costs.
A video coding method and apparatus based on block differential pulse coded modulation (BDPCM) that improves coding efficiency by applying BDPCM separately to luma and chroma components, utilizing directional information for transform coefficients, and omitting non-separable transforms for certain blocks.
Enhances overall image/video compression efficiency, specifically improving coding efficiency of transform indexes and transform skip flags in BDPCM-based video coding.
Smart Images

Figure 0007779943000031 
Figure 0007779943000032 
Figure 0007779943000033
Abstract
Description
[Technical Field]
[0001] This document relates to video coding technology, and more particularly to a video coding method and apparatus based on block differential pulse coded modulation (BDPCM) in a video coding system. [Background technology]
[0002] In recent years, the demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher UHD (Ultra High Definition) images / videos, has been increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits transmitted increases relatively compared to existing image / video data. Therefore, when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media, the transmission and storage costs increase.
[0003] In addition, interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content and holograms has been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, is increasing.
[0004] Accordingly, there is a demand for highly efficient image / video compression technology to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information having the above-mentioned various characteristics. Summary of the Invention [Problem to be solved by the invention]
[0005] The technical problem of this document is to provide a method and apparatus for improving video coding efficiency.
[0006] Another technical problem of this document is to provide a method and apparatus for improving coding efficiency of transform indexes in BDPCM-based video coding.
[0007] Another technical problem of this document is to provide a method and apparatus for improving coding efficiency of transform skip flags in BDPCM-based video coding.
[0008] Another technical problem of this document is to provide a method and apparatus capable of performing BDPCM coding separately for luma components or chroma components. [Means for solving the problem]
[0009] According to one embodiment of the present document, there is provided a video decoding method performed by a decoding device, the method including the steps of: deriving quantized transform coefficients for a current block based on a BDPCM; deriving transform coefficients by performing inverse quantization on the quantized transform coefficients; and deriving residual samples based on the transform coefficients, wherein if the BDPCM is applied to the current block, an inverse non-separable transform may not be applied to the transform coefficients.
[0010] If the BDPCM is applied to the current block, the value of the transform index for the inverse non-separable transform that may be applied to the current block may be assumed to be 0.
[0011] When the BDPCM is applied to the current block, the value of a transform skip flag indicating whether a transform is skipped for the current block may be assumed to be 0.
[0012] The BDPCM may be applied individually to the luma block of the current block or the chroma block of the current block, and when the BDPCM is applied to the luma block, the transformation index for the luma block may not be received, and when the BDPCM is applied to the chroma block, the transformation index for the chroma block may not be received.
[0013] If the width of the current block is less than or equal to a first threshold and the height of the current block is less than or equal to a second threshold, the BDPCM may be applied to the current block.
[0014] Based on directional information for the direction in which the BDPCM is performed, quantized transform coefficients may be derived.
[0015] The method may further include performing intra prediction on the current block based on a direction in which the BDPCM is performed.
[0016] The direction information may indicate a horizontal or vertical direction.
[0017] According to one embodiment of the present document, there is provided a video encoding method performed by an encoding device, the method including the steps of: deriving prediction samples for a current block based on a BDPCM, deriving residual samples for the current block based on the prediction samples, quantizing the residual samples, deriving quantized residual information based on the BDPCM, and encoding the quantized residual information and coding information for the current block, wherein when the BDPCM is applied to the current block, a non-separable transform may not be applied to the current block.
[0018] According to another embodiment of the present document, a digital storage medium can be provided on which video data including encoded video information and a bitstream generated by a video encoding method performed by an encoding device is stored.
[0019] According to another embodiment of the present document, a digital storage medium may be provided that stores video data including encoded video information and a bitstream that causes a decoding device to perform the video decoding method. [Effects of the Invention]
[0020] This document can improve the overall image / video compression efficiency.
[0021] According to the present disclosure, coding of transform indices can improve overall image / video compression efficiency.
[0022] According to the present disclosure, the efficiency of transform index coding in BDPCM-based video coding can be improved.
[0023] According to the present disclosure, it is possible to improve the coding efficiency of transform skip flags in BDPCM-based video coding.
[0024] Another technical object of this document is to provide a method and apparatus capable of performing BDPCM coding separately for luma components or chroma components.
[0025] The effects obtained through a specific example of the present specification are not limited to the effects listed above. For example, there may be various technical effects that a person having ordinary skill in the related art can understand or derive from the present specification. Therefore, the specific effects of the present specification are not limited to those explicitly described in the present specification, but may include various effects that can be understood or derive from the technical features of the present specification. [Brief explanation of the drawings]
[0026] [Figure 1] 1 illustrates, in simplified form, an example of a video / image coding system to which this document may be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which this document can be applied. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which this document can be applied. [Figure 4] 1 illustrates a schematic diagram of a multiple conversion technique according to one embodiment of the present document; [Figure 5] 65 intra-directional modes of prediction directions are shown exemplarily. [Figure 6] FIG. 1 is a diagram for explaining an RST according to one embodiment of this document. [Figure 7] 1 is a flowchart illustrating the operation of a video decoding device according to one embodiment of the present document. [Figure 8] 1 is a control flow chart illustrating a video decoding method according to one embodiment of the present document. [Figure 9] 1 is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present document. [Figure 10] 1 is a control flowchart illustrating a video encoding method according to one embodiment of the present document; [Figure 11] 1 illustrates an exemplary structure of a content streaming system to which this document applies. DETAILED DESCRIPTION OF THE INVENTION
[0027] Although this document may be modified in various ways and may have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to the specific embodiments. Common terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. A singular expression includes a plural expression unless the context clearly indicates otherwise. In this specification, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more different features, numbers, steps, operations, components, parts, or combinations thereof.
[0028] Meanwhile, each component in the drawings described in this document is shown independently for the convenience of explaining the different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Implementations in which each component is integrated and / or separated are also within the scope of this document as long as they do not deviate from the essence of this document.
[0029] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and redundant description of the same components will be omitted.
[0030] This document relates to video / image coding. For example, methods / embodiments disclosed in this document may be related to the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), next-generation video / image coding standards beyond VVC, or other video coding-related standards (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), the essential video coding (EVC) standard, the AVS2 standard, etc.).
[0031] This document presents various embodiments relating to video / image coding, which may be implemented in combination with one another unless otherwise specified.
[0032] In this document, video can refer to a collection of a series of images over time. A picture generally refers to a unit that shows one image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile can contain one or more coding tree units (CTUs). One picture can be composed of one or more slices / tiles. One picture can be composed of one or more tile groups. One tile group can contain one or more tiles.
[0033] A pixel or a pel may refer to the smallest unit constituting one picture (or video). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or may refer to a transform coefficient in the frequency domain when such a pixel value is transformed into the frequency domain.
[0034] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area, depending on the situation. In a general case, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0035] In this document, the terms " / " and "," should be interpreted to mean "and / or." For example, "A / B" means "A and / or B," and "A, B" means "A and / or B." Furthermore, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." (In this document, the terms " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A, B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")
[0036] Furthermore, in this document, "or" should be interpreted as "and / or." For example, "A or B" may mean 1) only "A," 2) only "B," or 3) "A and B." In other words, "or" in this document may mean "additionally or alternatively." (Further, in the document, the term "or" should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively.")
[0037] As used herein, "at least one of A and B" can mean "only A," "only B," or "both A and B." Furthermore, as used herein, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B."
[0038] Furthermore, in this specification, "at least one of A, B, and C" can mean "only A," "only B," "only C," or "any combination of A, B, and C." Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0039] Furthermore, parentheses used herein may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" may be suggested as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction."
[0040] Technical features individually described in one drawing in this specification may be embodied individually or simultaneously.
[0041] FIG. 1 shows a schematic diagram of an example video / image coding system to which this document may be applied.
[0042] 1, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data to the receiving device via a digital storage medium or a network in the form of a file or streaming.
[0043] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may also include a display unit, which may be a separate device or an external component.
[0044] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.
[0045] An encoding device can encode input video / images. The encoding device can perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0046] The transmitter may transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver may receive / extract the bitstream and transmit it to a decoding device.
[0047] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.
[0048] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.
[0049] 2 is a diagram illustrating the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the term "video encoding device" may include a video encoding device.
[0050] Referring to FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0051] The image division unit 210 may divide an input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure described herein may be performed based on the final coding unit that is not further divided. In this case, the largest coding unit may be immediately used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and the coding unit of the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0052] The term "unit" can be used interchangeably with terms such as "block" or "area" depending on the situation. In a general case, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally refer to a pixel or pixel value, or can refer to only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component. A sample can be used as a term corresponding to one pixel or pel of a picture (or image).
[0053] The subtraction unit 231 may subtract a prediction signal (predicted block, prediction sample, or prediction sample array) output from the prediction unit 220 from an input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 may perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit may generate various information related to prediction, such as prediction mode information, and transmit the information to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The prediction information may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0054] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away, depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0055] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (col CU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of a neighboring block as a motion vector predictor and signaling the motion vector difference.
[0056] The predictor 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may not only apply intra prediction or inter prediction for prediction of a block, but also apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) for prediction of a block. The intra block copy may be used, for example, for coding content images / videos such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein.
[0057] The prediction signal generated by the inter prediction unit 221 and / or the intra prediction unit 222 may be used to generate a reconstructed signal or a residual signal. The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing relationship information between pixels. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0058] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 may perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 240 may encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. Signaled / transmitted information and / or syntax elements described later in this document may be encoded through the encoding procedures described above and included in the bitstream.The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoding unit 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding apparatus 200, or the transmitter may be included in the entropy encoding unit 240.
[0059] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) may be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the predictor 220. When there is no residual for the current block, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra-prediction of the next block to be processed in the current picture, and may also be used for inter-prediction of the next picture after filtering, as described below.
[0060] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0061] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 290, as will be described later in connection with each filtering method. The filtering information may be encoded by the entropy encoding unit 290 and output in the form of a bitstream.
[0062] The modified reconstructed picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. When inter prediction is applied through this, the encoding apparatus can avoid a mismatch in prediction between the encoding apparatus 200 and the decoding apparatus, and can also improve coding efficiency.
[0063] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0064] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which this document can be applied.
[0065] Referring to FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0066] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding apparatus of FIG. 2. For example, the decoding apparatus 300 can derive units / blocks based on information about block division obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processing unit applied by the encoding apparatus. Therefore, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be played back via a playback device.
[0067] The decoding apparatus 300 may receive a signal output from the encoding apparatus of FIG. 2 in the form of a bitstream, and the received signal may be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding apparatus may further decode pictures based on the information on the parameter sets and / or the general constraint information. Signaling / received information and / or syntax elements, which will be described later in this document, may be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential Golomb Coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded, decoding information on neighboring and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element.In this case, after determining a context model, the CABAC entropy decoding method may update the context model using information about the decoded symbol / bin for the context model of the next symbol / bin. Prediction-related information from the information decoded by the entropy decoding unit 310 may be provided to the prediction unit 330, and information about the residual on which entropy decoding is performed by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 321. In addition, filtering-related information from the information decoded by the entropy decoding unit 310 may be provided to the filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding apparatus may be further configured as an internal / external element of the decoding apparatus 300, or the receiving unit may be a component of the entropy decoding unit 310. Meanwhile, the decoding device according to this document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.
[0068] The inverse quantization unit 321 can inverse quantize the quantized transform coefficients and output transform coefficients. The inverse quantization unit 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0069] The inverse transform unit 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0070] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0071] The predictor 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the predictor may not only apply intra prediction or inter prediction for prediction of a block, but also apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) for prediction of a block. The intra block copy may be used for content image / video coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described herein.
[0072] The intra prediction unit 332 can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. Prediction modes in intra prediction can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 332 can also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0073] The inter prediction unit 331 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 331 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0074] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit 330. When there is no residual for the current block, such as when skip mode is applied, the predicted block may be used as the reconstructed block.
[0075] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.
[0076] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during picture decoding.
[0077] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0078] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 331 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 332.
[0079] In this specification, the embodiments described for the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, and filtering unit 350 of the decoding device 300 can also be applied identically or correspondingly to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, and filtering unit 260 of the decoding device 200, respectively.
[0080] As described above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block including predicted samples for a current block, which is a block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived in the same way by an encoding device and a decoding device. The encoding device can improve video coding efficiency by signaling to the decoding device information regarding the residual between the original block and the predicted block (residual information) rather than the original sample values of the original block. The decoding device can derive a residual block including residual samples based on the residual information, combine the residual block with the predicted block to generate a reconstructed block including reconstructed samples, and generate a reconstructed picture including the reconstructed block.
[0081] The residual information may be generated through a transform and quantization procedure. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transform procedure on residual samples (residual sample array) included in the residual block to derive transform coefficients, perform a quantization procedure on the transform coefficients to derive quantized transform coefficients, and signal the related residual information (via a bitstream) to a decoding device. Here, the residual information may include information such as value information, position information, transform technique, transform kernel, and quantization parameter of the quantized transform coefficients. The decoding device may derive residual samples (or residual blocks) by performing an inverse quantization / inverse transform procedure based on the residual information. The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device may also derive a residual block by inverse quantizing / inverse transforming quantized transform coefficients for reference for inter-prediction of a subsequent picture, and generate a reconstructed picture based on the residual block.
[0082] FIG. 4 shows a schematic diagram of the multiple transformation technique according to this document.
[0083] Referring to Figure 4, the transform unit may correspond to the transform unit in the encoding device of Figure 2 described above, and the inverse transform unit may correspond to the inverse transform unit in the encoding device of Figure 2 described above or the inverse transform unit in the decoding device of Figure 3.
[0084] The transform unit may perform a primary transform based on the residual samples (residual sample array) in the residual block to derive (primary) transform coefficients (S410). Such a primary transform may be referred to as a core transform. Here, the primary transform may be based on Multiple Transform Selection (MTS), and when multiple transforms are applied as the primary transform, it may be referred to as a multiple core transform.
[0085] The multi-kernel transform may refer to a transform method further using a Discrete Cosine Transform (DCT) type 2, a Discrete Sine Transform (DST) type 7, a DCT type 8, and / or a DST type 1. That is, the multi-kernel transform may refer to a transform method for transforming a spatial domain residual signal (or a residual block) into frequency domain transform coefficients (or first-order transform coefficients) based on a plurality of transform kernels selected from the DCT type 2, the DST type 7, the DCT type 8, and the DST type 1. Here, the first-order transform coefficients may be referred to as tentative transform coefficients from the perspective of a transform unit.
[0086] In other words, when an existing transform method is applied, a spatial-domain to frequency-domain transform is applied to a residual signal (or residual block) based on DCT type 2 to generate transform coefficients. In contrast, when the multi-kernel transform is applied, a spatial-domain to frequency-domain transform is applied to a residual signal (or residual block) based on DCT type 2, DST type 7, DCT type 8, and / or DST type 1, etc. to generate transform coefficients (or primary transform coefficients). Here, DCT type 2, DST type 7, DCT type 8, DST type 1, etc. may be referred to as transform types, transform kernels, or transform cores. Such DCT / DST transform types may be defined based on basis functions.
[0087] When the multi-kernel transform is performed, a vertical transform kernel and a horizontal transform kernel for a current block may be selected from the transform kernels, and a vertical transform for the current block may be performed based on the vertical transform kernel, and a horizontal transform for the current block may be performed based on the horizontal transform kernel. Here, the horizontal transform may indicate a transform for a horizontal component of the current block, and the vertical transform may indicate a transform for a vertical component of the current block. The vertical transform kernel / horizontal transform kernel may be adaptively determined based on a prediction mode and / or a transform index of a current block (CU or sub-block) including a residual block.
[0088] Also, according to one example, when a linear transform is performed using MTS, a specific basis function is set to a predetermined value, and when a vertical transform or horizontal transform is performed, a mapping relationship for the transform kernel can be set by combining which basis function is applied. For example, if the horizontal transform kernel is represented by trTypeHor and the vertical transform kernel is represented by trTypeVer, a value of 0 for trTypeHor or trTypeVer can be set to DCT2, a value of 1 for trTypeHor or trTypeVer can be set to DCT7, and a value of 2 for trTypeHor or trTypeVer can be set to DCT8.
[0089] In this case, index information of the MTS can be encoded and signaled to a decoding device to indicate one of a number of transform kernel sets. For example, an MTS index of 0 indicates that the values of trTypeHor and trTypeVer are all 0, an MTS index of 1 indicates that the values of trTypeHor and trTypeVer are all 1, an MTS index of 2 indicates that the value of trTypeHor is 2 and the value of trTypeVer is 1, an MTS index of 3 indicates that the value of trTypeHor is 1 and the value of trTypeVer is 2, and an MTS index of 4 indicates that the values of trTypeHor and trTypeVer are all 2.
[0090] As an example, the conversion kernel set according to the index information of the MTS is shown in the table below.
[0091] [Table 1]
[0092] The transform unit may perform a secondary transform based on the (primary) transform coefficients to derive modified (secondary) transform coefficients (S420). The primary transform is a transform from the spatial domain to the frequency domain, and the secondary transform refers to a transform using correlations existing between the (primary) transform coefficients to a more compressed representation. The secondary transform may include a non-separable transform. In this case, the secondary transform may be referred to as a non-separable secondary transform (NSST) or a mode-dependent non-separable secondary transform (MDNSST). The non-separable secondary transform may refer to a transform that performs a secondary transform on the (primary) transform coefficients derived through the primary transform based on a non-separable transform matrix to generate modified transform coefficients (or secondary transform numbers) for the residual signal. Here, based on the non-separable transform matrix, a transform can be applied to the (first-order) transform coefficients at one time, without separately applying a vertical transform and a horizontal transform (or independently applying a horizontal-vertical transform). In other words, the non-separable second-order transform can refer to a transform method in which, for example, a two-dimensional signal (transform coefficient) is rearranged into a one-dimensional signal in a specific direction (e.g., row-first direction or column-first direction) without separating the vertical and horizontal components of the (first-order) transform coefficients, and then a modified transform coefficient (or second-order transform coefficient) is generated based on the non-separable transform matrix. For example, the row-major order is an arrangement of a first row, a second row, ..., an Nth row for an MxN block, and the column-major order is an arrangement of a first column, a second column, ..., an Mth column for an MxN block. The non-separable quadratic transform may be applied to the top-left region of a block composed of (first-order) transform coefficients (hereinafter, referred to as a transform coefficient block).For example, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 8, an 8x8 non-separable quadratic transform may be applied to the 8x8 region at the upper left of the transform coefficient block. Also, if the width (W) and height (H) of the transform coefficient block are both equal to or greater than 4 but the width (W) or height (H) of the transform coefficient block is less than 8, a 4x4 non-separable quadratic transform may be applied to the min(8,W) x min(8,H) region at the upper left of the transform coefficient block. However, embodiments are not limited thereto. For example, even if the width (W) or height (H) of the transform coefficient block only satisfies the condition that both are equal to or greater than 4, a 4x4 non-separable quadratic transform may also be applied to the min(8,W) x min(8,H) region at the upper left of the transform coefficient block.
[0093] Specifically, for example, if a 4x4 input block is used, the non-separable quadratic transform can be performed as follows:
[0094] The 4x4 input block X can be expressed as:
[0095]
number
[0096] When X is expressed in the form of a vector, the vector JPEG0007779943000003.jpg10155 can be represented as:
[0097]
number
[0098] As shown in Equation 2, the vector JPEG0007779943000005.jpg10165 rearranges the two-dimensional block of X in Equation 1 into a one-dimensional vector in row-first order.
[0099] In this case, the second-order non-separable transform can be calculated as follows:
[0100]
number
[0101] where: JPEG0007779943000007.jpg9166 denotes the transform coefficient vector, and T denotes the 16x16 (non-separable) transform matrix.
[0102] Through Equation 3, a 16×1 transform coefficient vector JPEG0007779943000008.jpg10155 can be derived, and JPEG0007779943000009.jpg10155 can be re-organized into 4x4 blocks via the scan order (horizontal, vertical, diagonal, etc.). However, the above calculation is an example, and in order to reduce the calculation complexity of the non-separable quadratic transform, a Hypercube-Givens Transform (HyGT) or the like can also be used for the calculation of the non-separable quadratic transform.
[0103] Meanwhile, the non-separable quadratic transform may be mode-dependent, with the transform kernel (or transform core, transform type) being selectable, where the mode may include an intra-prediction mode and / or an inter-prediction mode.
[0104] As described above, the non-separable quadratic transform may be performed based on an 8x8 transform or a 4x4 transform determined based on the width (W) and height (H) of the transform coefficient block. The 8x8 transform refers to a transform that can be applied to an 8x8 region contained within the corresponding transform coefficient block when W and H are both equal to or greater than 8, and the corresponding 8x8 region may be the upper-left 8x8 region within the corresponding transform coefficient block. Similarly, the 4x4 transform refers to a transform that can be applied to a 4x4 region contained within the corresponding transform coefficient block when W and H are both equal to or greater than 4, and the corresponding 4x4 region may be the upper-left 4x4 region within the corresponding transform coefficient block. For example, an 8x8 transform kernel matrix may be a 64x64 / 16x64 matrix, and a 4x4 transform kernel matrix may be a 16x16 / 8x16 matrix.
[0105] In this case, for mode-based transform kernel selection, two non-separable quadratic transform kernels may be configured per transform set for the non-separable quadratic transform for both the 8×8 transform and the 4×4 transform, and the number of transform sets may be four. That is, four transform sets may be configured for the 8×8 transform and four transform sets may be configured for the 4×4 transform. In this case, each of the four transform sets for the 8×8 transform may include two 8×8 transform kernels, and each of the four transform sets for the 4×4 transform may include two 4×4 transform kernels.
[0106] However, the size of the transform, the number of sets, and the number of transform kernels in a set are examples, and sizes other than 8x8 or 4x4 may be used, or n sets may be constructed and each set may contain k transform kernels.
[0107] The transform set may be referred to as an NSST set, and the transform kernels in the NSST set may be referred to as NSST kernels. Selection of a particular transform set may be based on, for example, the intra prediction mode of the current block (CU or sub-block).
[0108] For reference, for example, the intra prediction modes may include two non-directional (or non-angular) intra prediction modes and 65 directional (or angular) intra prediction modes. The non-directional intra prediction modes may include a planar intra prediction mode numbered 0 and a DC intra prediction mode numbered 1, and the directional intra prediction modes may include 65 intra prediction modes numbered 2 to 66. However, this is merely an example, and this document may also be applied to cases where the number of intra prediction modes is different. Meanwhile, an intra prediction mode numbered 67 may also be used depending on the case, and the 67th intra prediction mode may indicate a linear model (LM) mode.
[0109] FIG. 5 exemplarily shows the intra-directional modes of 65 prediction directions.
[0110] Referring to FIG. 5, intra prediction modes may be classified into those with horizontal directionality and those with vertical directionality, centered on the 34th intra prediction mode, which has a prediction direction from the upper left diagonal. In FIG. 5, H and V represent horizontal and vertical directionality, respectively, and the numbers -32 to 32 indicate displacements in 1 / 32 units on the sample grid position. This may indicate an offset to the mode index value. The 2nd to 33rd intra prediction modes have horizontal directionality, while the 34th to 66th intra prediction modes have vertical directionality. Meanwhile, the 34th intra prediction mode can be considered neither horizontally nor vertically oriented in the strict sense, but can be classified as horizontally oriented from the perspective of determining a transform set for a secondary transform. This is because input data is transposed for vertical modes symmetrical with respect to the 34th intra prediction mode, and the 34th intra prediction mode uses the same input data alignment method as the horizontal mode. Transposing the input data means that for an MxN two-dimensional block of data, rows become columns and columns become rows, creating NxM data. The 18th and 50th intra prediction modes indicate a horizontal intra prediction mode and a vertical intra prediction mode, respectively. The 2nd intra prediction mode predicts in the upper right direction using a reference pixel on the left, so it can be called an upper right diagonal intra prediction mode. In the same context, the 34th intra prediction mode can be called a lower right diagonal intra prediction mode, and the 66th intra prediction mode can be called a lower left diagonal intra prediction mode.
[0111] For example, depending on the intra prediction mode, the mapping of four transform sets may be shown as in the following table.
[0112] [Table 2]
[0113] As shown in Table 2, one of four transform sets, ie, stTrSetIdx, can be mapped to any of 0 to 3, ie, 4, depending on the intra prediction mode.
[0114] On the other hand, if it is determined that a specific set is to be used for a non-separable transform, one of k transform kernels in the specific set can be selected through a non-separable quadratic transform index. The encoding device can derive a non-separable quadratic transform index that points to a specific transform kernel based on a rate-distortion (RD) check and signal the non-separable quadratic transform index to a decoding device. The decoding device can select one of k transform kernels in the specific set based on the non-separable quadratic transform index. For example, an NSST index value of 0 can indicate the first non-separable quadratic transform kernel, an NSST index value of 1 can indicate the second non-separable quadratic transform kernel, and an NSST index value of 2 can indicate the third non-separable quadratic transform kernel. Alternatively, an NSST index value of 0 can indicate that the first non-separable quadratic transform is not applied to the current block, and NSST index values of 1 to 3 can indicate the three transform kernels.
[0115] The transform unit may perform the non-separable quadratic transform based on the selected transform kernel to obtain modified (quadratic) transform coefficients. The modified transform coefficients may be derived as quantized transform coefficients via the quantizer unit as described above, encoded, and signaled to a decoding device and transmitted to an inverse quantization / inverse transform unit in the encoding device.
[0116] On the other hand, when the secondary transform is omitted as described above, the (primary) transform coefficients, which are the output of the primary (separate) transform, can be derived as quantized transform coefficients through the quantization unit as described above, encoded, signaled to the decoding device, and transmitted to the inverse quantization / inverse transform unit within the encoding device.
[0117] The inverse transform unit may perform a series of steps in the reverse order of the steps performed by the transform unit described above. The inverse transform unit may receive (dequantized) transform coefficients, perform a secondary (inverse) transform on them to derive (primary) transform coefficients (S450), and perform a primary (inverse) transform on the (primary) transform coefficients to obtain residual blocks (residual samples) (S460). Here, the primary transform coefficients may be referred to as modified transform coefficients from the perspective of the inverse transform unit. As described above, the encoding device and the decoding device may generate reconstructed blocks based on the residual blocks and predicted blocks, and generate reconstructed pictures based on the reconstructed blocks.
[0118] Meanwhile, the decoding apparatus may further include a secondary inverse transform application determining unit (or an element determining whether to apply the secondary inverse transform) and a secondary inverse transform determining unit (or an element determining the secondary inverse transform). The secondary inverse transform application determining unit may determine whether to apply the secondary inverse transform. For example, the secondary inverse transform may be NSST or RST, and the secondary inverse transform application determining unit may determine whether to apply the secondary inverse transform based on a secondary transform flag parsed from the bitstream. As another example, the secondary inverse transform application determining unit may determine whether to apply the secondary inverse transform based on transform coefficients of a residual block.
[0119] The secondary inverse transform decision unit may determine a secondary inverse transform. In this case, the secondary inverse transform decision unit may determine a secondary inverse transform to be applied to a current block based on an NSST (or RST) transform set specified by an intra prediction mode. In addition, as an embodiment, the secondary transform decision method may be determined depending on the primary transform decision method. Various combinations of primary transform and secondary transform may be determined depending on the intra prediction mode. In addition, as an example, the secondary inverse transform decision unit may determine an area to which the secondary inverse transform is applied based on the size of the current block.
[0120] On the other hand, as described above, if the second-order (inverse) transform is omitted, the (dequantized) transform coefficients can be received and the first-order (separate) inverse transform can be performed to obtain a residual block (residual sample). As described above, the encoding device and the decoding device can generate a reconstructed block based on the residual block and a predicted block, and generate a reconstructed picture based on the reconstructed block.
[0121] On the other hand, in this paper, in order to reduce the computational complexity and memory requirements associated with non-separable secondary transforms, the RST (reduced secondary transform) can be applied, in which the size of the transformation matrix (kernel) is reduced using the concept of NSST.
[0122] Meanwhile, the coefficients constituting the transform kernel, transform matrix, and transform kernel matrix described herein, i.e., kernel coefficients or matrix coefficients, may be expressed in 8 bits. This may be one condition for implementation in a decoding device and an encoding device, and may reduce the memory requirements for storing the transform kernel with a reasonably acceptable performance degradation compared to the existing 9-bit or 10-bit representation. In addition, expressing the kernel matrix in 8 bits may allow the use of a small multiplier and may be more suitable for SIMD (Single Instruction Multiple Data) instructions used for optimal software implementation.
[0123] In this specification, RST may refer to a transformation performed on residual samples of a target block based on a transform matrix whose size is reduced by a simplification factor. When a simplified transformation is performed, the amount of calculation required during the transformation may be reduced due to the reduction in the size of the transform matrix. In other words, RST can be used to resolve issues of computational complexity that arise during the transformation of large blocks or non-separable transformations.
[0124] The RST may be referred to by various terms such as a reduced transform, a reduced transform, a reduced secondary transform, a reduction transform, a simplified transform, a simple transform, etc., and the names by which the RST may be referred to are not limited to the examples given. Alternatively, the RST may be referred to as an LFNST (Low-Frequency Non-Separable Transform) because it is mainly performed in the low-frequency domain including non-zero coefficients in the transform block. The transform index may be called an LFNST index.
[0125] On the other hand, when the second-order inverse transform is performed based on an RST, the inverse transform unit 235 of the encoding apparatus 200 and the inverse transform unit 322 of the decoding apparatus 300 may include an inverse RST unit that derives modified transform coefficients based on the inverse RST for the transform coefficients, and an inverse linear transform unit that derives residual samples for the current block based on an inverse linear transform for the modified transform coefficients. The inverse linear transform refers to the inverse transform of the linear transform applied to the residual. In this document, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying the corresponding transform.
[0126] FIG. 6 is a diagram for explaining an RST according to an embodiment of the present document.
[0127] In this specification, the term "target block" may refer to a current block, a residual block, or a transform block on which coding is performed.
[0128] In an RST according to one embodiment, an N-dimensional vector is mapped to an R-dimensional vector located in a different space, and a reduced transformation matrix can be determined, where R is smaller than N. N may represent the square of the length of one side of the block to which the transformation is applied, or the total number of transformation coefficients corresponding to the block to which the transformation is applied, and the simplification factor may represent an R / N value. The simplification factor may be referred to by various terms such as a reduced factor, reduction factor, simplified factor, or simple factor. Meanwhile, R may be referred to as a reduced coefficient, but in some cases the simplification factor may also represent R. In other cases, the simplification factor may also represent an N / R value.
[0129] In one embodiment, the simplification factors or simplification coefficients may be signaled via the bitstream, but the embodiment is not limited thereto. For example, predefined values for the simplification factors or simplification coefficients may be stored in each encoding device 200 and decoding device 300, in which case the simplification factors or simplification coefficients may not be separately signaled.
[0130] The size of the simplified transformation matrix according to one embodiment is RxN, which is smaller than the size NxN of the normal transformation matrix, and can be defined as Equation 4 below.
[0131]
number
[0132] The matrix T in the Reduced Transform block shown in (a) of FIG. 6 is the matrix T in Equation 4. RxN As shown in FIG. 6(a), the simplified transformation matrix T RxN When multiplied by , the transform coefficients for the current block can be derived.
[0133] In one embodiment, if the size of the block to which the transform is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST according to (a) of Figure 6 can be expressed by the matrix operation shown in Equation 5 below. In this case, the memory and multiplication operations can be reduced by approximately 1 / 4 due to the simplification factor.
[0134] In this document, a matrix operation can be understood as an operation in which a matrix is placed to the left of a column vector and multiplied by the column vector to obtain the column vector.
[0135]
number
[0136] In Equation 5, r1 to r 64 may represent a residual sample for the current block, and more specifically, may be a transform coefficient generated by applying a linear transform. As a result of the calculation of Equation 5, the transform coefficient c for the current block is i can be derived, and c i The derivation process is as shown in Equation 6.
[0137]
number
[0138] The calculation result of Equation 6 is the transform coefficients c1 to c2 for the target block. R That is, when R=16, the transform coefficients c1 to c2 for the target block can be derived. 16can be derived. If a regular transform, rather than RST, were applied and a transform matrix of size 64x64 (NxN) were multiplied by residual samples of size 64x1 (Nx1), 64 (N) transform coefficients for the current block would be derived. However, because RST is applied, only 16 (R) transform coefficients for the current block are derived. Since the total number of transform coefficients for the current block is reduced from N to R, the amount of data transmitted from encoding apparatus 200 to decoding apparatus 300 is reduced, and therefore, transmission efficiency between encoding apparatus 200 and decoding apparatus 300 can be improved.
[0139] Considering the size of the transformation matrix, the size of a normal transformation matrix is 64x64 (NxN), but the size of a simplified transformation matrix is reduced to 16x64 (RxN), so compared to performing normal transformation, memory usage when performing RST can be reduced by a ratio of R / N. Also, compared to the number of multiplication operations when using a normal transformation matrix (NxN), the number of multiplication operations can be reduced by a ratio of R / N when using a simplified transformation matrix (RxN).
[0140] In one embodiment, the transform unit 232 of the encoding apparatus 200 may derive transform coefficients for the current block by performing a primary transform and an RST-based secondary transform on residual samples for the current block. These transform coefficients may be transmitted to an inverse transform unit 322 of the decoding apparatus 300, and the inverse transform unit 322 of the decoding apparatus 300 may derive modified transform coefficients based on an inverse reduced secondary transform (RST) of the transform coefficients and derive residual samples for the current block based on an inverse primary transform of the modified transform coefficients.
[0141] Inverse RST matrix T according to one embodiment NxRThe size of the simplified transformation matrix T shown in Equation 4 is NxR, which is smaller than the size of the normal inverse transformation matrix NxN. RxN and are in a transpose relationship.
[0142] The matrix T in the Reduced Inverse Transform block shown in Figure 6(b) t is the inverse RST matrix T RxN T (The superscript T means transpose.) As shown in FIG. 6(b), the inverse RST matrix T RxN T When the inverse RST matrix T is multiplied by the inverse RST matrix T, the modified transform coefficients for the current block or the residual samples for the current block can be derived. RxN T is (T RxN ) T NxR It is also sometimes expressed as:
[0143] More specifically, when the inverse RST is applied to the secondary inverse transform, the inverse RST matrix T RxN T On the other hand, an inverse RST can be applied to the inverse linear transform, in which case the inverse RST matrix T RxN T When multiplied by , the residual sample for the target block can be derived.
[0144] In one embodiment, when the size of the block to which the inverse transform is applied is 8x8 and R=16 (i.e., R / N=16 / 64=1 / 4), the RST according to (b) of FIG. 6 can be expressed by a matrix operation as shown in Equation 7 below.
[0145]
number
[0146] In Equation 7, c1 to c 16 The result of the calculation of Equation 7 is r, which indicates the modified transform coefficients for the current block or the residual samples for the current block. j can be derived, and r j The derivation process is as shown in Equation 8.
[0147]
number
[0148] The calculation result of Equation 8 is r1 to r2, which indicate the modified transform coefficients for the target block or the residual samples for the target block. N can be derived. Considering the size of the inverse transformation matrix, the size of a normal inverse transformation matrix is 64x64 (NxN), but the size of the simplified inverse transformation matrix is reduced to 64x16 (NxR). Therefore, compared to performing a normal inverse transformation, memory usage when performing inverse RST can be reduced by a ratio of R / N. Also, compared to the number of multiplication operations when using a normal inverse transformation matrix (NxN), when using a simplified inverse transformation matrix, the number of multiplication operations can be reduced by a ratio of R / N (NxR).
[0149] Meanwhile, the transform set configuration shown in Table 2 can also be applied to an 8x8 RST. That is, the corresponding 8x8 RST can be applied according to the transform set in Table 2. Since one transform set is composed of two or three transforms (kernels) depending on the intra-frame prediction mode, it can be configured to select one of up to four transforms, including cases where a secondary transform is not applied. When a secondary transform is not applied, the transform can be considered to be applied with an identity matrix. If the four transforms are assigned indices 0, 1, 2, and 3 (for example, index 0 can be assigned to the identity matrix, i.e., when a secondary transform is not applied), a syntax element called an NSST index can be signaled for each transform coefficient block to specify the transform to be applied. That is, an 8x8 NSST can be specified for the upper left block of an 8x8 RST via the NSST index, and an 8x8 RST can be specified in the RST configuration. The 8x8 NSST and 8x8 RST indicate transformations that can be applied to an 8x8 region contained within a corresponding transform coefficient block when W and H of the target block are both equal to or greater than 8, and the corresponding 8x8 region may be the upper left 8x8 region within the corresponding transform coefficient block. Similarly, the 4x4 NSST and 4x4 RST indicate transformations that can be applied to a 4x4 region contained within a corresponding transform coefficient block when W and H of the target block are both equal to or greater than 4, and the corresponding 4x4 region may be the upper left 4x4 region within the corresponding transform coefficient block.
[0150] Meanwhile, according to one embodiment of this document, during the encoding process, only 48 pieces of data constituting an 8x8 region can be selected, instead of a 16x64 transformation kernel matrix, to apply a maximum 16x48 transformation kernel matrix. Here, "maximum" means that for an mx48 transformation kernel matrix that can generate m coefficients, the maximum value of m is 16. That is, when RST is performed by applying an mx48 transformation kernel matrix (m≦16) to an 8x8 region, 48 pieces of data can be input and m coefficients can be generated. When m is 16, 48 pieces of data can be input and 16 coefficients can be generated. That is, when 48 pieces of data form a 48x1 vector, the 16x48 matrix and the 48x1 vector can be multiplied in order to generate the 16x1 vector. In this case, the 48 pieces of data constituting the 8x8 region can be properly arranged to form a 48x1 vector. In this case, when a matrix operation is performed by applying a maximum 16x48 transform kernel matrix, 16 modified transform coefficients are generated, and the 16 modified transform coefficients may be arranged in the upper left 4x4 area according to the scanning order, and the upper right 4x4 area and the lower left 4x4 area may be filled with 0s.
[0151] A transposed matrix of the above-described transform kernel matrix can be used for the inverse transform in the decoding process. That is, when an inverse RST or an LFNST is performed in the inverse transform process performed in the decoding device, input coefficient data to which the inverse RST is applied is configured as a one-dimensional vector in a predetermined arrangement order, and the modified coefficient vector obtained by multiplying the one-dimensional vector by the matrix of the corresponding inverse RST on the left side can be arranged in a two-dimensional block in a predetermined arrangement order.
[0152] To summarize, when RST or LFNST is applied to an 8x8 region during the transform process, a matrix operation is performed on 48 transform coefficients in the upper left, upper right, and lower left regions of the 8x8 region, excluding the lower right region, of the transform coefficients of the 8x8 region, and a 16x48 transform kernel matrix. For the matrix operation, the 48 transform coefficients are input into a one-dimensional array. After this matrix operation, 16 modified transform coefficients are derived, and the modified transform coefficients may be arranged in the upper left region of the 8x8 region.
[0153] Conversely, when the inverse RST or LFNST is applied to an 8x8 region during the inverse transform process, 16 transform coefficients corresponding to the upper left side of the 8x8 region among the transform coefficients of the 8x8 region are input in a one-dimensional array form according to the scanning order and may be subjected to a matrix operation with a 48x16 transform kernel matrix. That is, the matrix operation in this case may be expressed as (48x16 matrix) * (16x1 transform coefficient vector) = (48x1 modified transform coefficient vector). Here, an nx1 vector may be interpreted as an nx1 matrix and therefore may be expressed as an nx1 column vector. Also, * indicates a matrix multiplication operation. When this matrix operation is performed, 48 modified transform coefficients may be derived, and the 48 modified transform coefficients may be arranged in the upper left, upper right, and lower left regions of the 8x8 region, excluding the lower right region.
[0154] On the other hand, when the second-order inverse transform is performed based on an RST, the inverse transform unit 235 of the encoding apparatus 200 and the inverse transform unit 322 of the decoding apparatus 300 may include an inverse RST unit that derives modified transform coefficients based on the inverse RST for the transform coefficients, and an inverse linear transform unit that derives residual samples for the current block based on an inverse linear transform for the modified transform coefficients. The inverse linear transform refers to the inverse transform of the linear transform applied to the residual. In this document, deriving transform coefficients based on a transform may refer to deriving transform coefficients by applying the corresponding transform.
[0155] Meanwhile, in one embodiment, a block differential pulse coded modulation (BDPCM) technique can be used. BDPCM can also be called quantized residual block-based delta pulse code modulation (RDPCM).
[0156] When predicting a block using BDPCM, reconstructed samples are used to predict rows or columns of the block line-by-line. The reference pixels used may be unfiltered samples. The direction of BDPCM may indicate whether vertical or horizontal prediction is used. A prediction error is quantized in the spatial domain, and pixels are reconstructed by adding the dequantized prediction error to the prediction. As an alternative to BDPCM, quantized residual domain BDPCM may be proposed, and the prediction direction and signaling may be the same as those of BDPCM applied to the spatial domain. That is, through quantized residual domain BDPCM, the quantized coefficients themselves can be superimposed like DPCM (Delta Pulse Code Modulation), and then the residual can be reconstructed through dequantization. Therefore, quantized residual domain BDPCM can be used to mean applying DPCM at the residual coding end. As used below, quantized residual domain refers to the domain of quantized residual samples, in which residuals derived based on prediction are quantized without transformation.
[0157] For a block of size M (rows) x N (columns), the predicted residual is calculated by using the unfiltered samples from the left or upper boundary samples, and performing intra prediction horizontally (copying the left surrounding pixel line to the predicted block line-by-line) or intra prediction vertically (copying the upper surrounding pixel line to the predicted block line-by-line).(i,j) Assume that (0≦i≦M-1, 0≦j≦N-1). Then, the residual r (i,j) Let us denote the quantized version of Q(r (i,j) ) (0≦i≦M-1, 0≦j≦N-1), where residual means the difference between the original block and the predicted block.
[0158] Then, when BDPCM is applied to the quantized residual samples, An M×N transformed array consisting of JPEG0007779943000016.jpg7157 JPEG0007779943000017.jpg8157 is derived.
[0159] When vertical BDPCM is signaled, JPEG0007779943000018.jpg9157 is expressed as follows:
[0160]
number
[0161] Applying the same to horizontal prediction, the residual quantized samples are:
[0162]
number
[0163] Quantized residual samples JPEG0007779943000021.jpg11153 is sent to the decoding device.
[0164] In the decoding device, Q(r (i,j) ) (0≦i≦M-1, 0≦j≦N-1.) The above operations are reversed to derive i<M-1, 0<j<N-1.
[0165] For vertical prediction, the following formula can be applied:
[0166]
number
[0167] For horizontal prediction, the following formula can be applied:
[0168]
number
[0169] Dequantized quantized residual JPEG0007779943000024.jpg9154 is combined with the intra-block prediction values to derive the reconstructed sample values.
[0170] The main advantage of such a technique is that the inverse BDPCM can be performed by simply adding a predictor during the parsing of the coefficients, or even after parsing the coefficients.
[0171] As described above, BDPCM can be applied to the quantized residual domain, which can include quantized residuals (or quantized residual coefficients), and transform skip can be applied to the residuals. That is, transform can be skipped and quantization can be applied to the residual samples. Alternatively, the quantized residual domain can include quantized transform coefficients. A flag indicating whether BDPCM is applicable can be signaled at the sequence level (SPS), and this flag can be signaled only if transform skip mode is signaled as possible in the SPS.
[0172] When BDPCM is applied, intra prediction for the quantized residual domain is performed on the entire block by sample copying in a prediction direction similar to the intra prediction direction (e.g., vertical prediction or horizontal prediction). The residual is quantized, and a delta value, i.e., a difference value, between the quantized residual and a predictor for the horizontal or vertical direction (i.e., the quantized residual for the horizontal or vertical direction) is calculated. JPEG0007779943000025.jpg8154 is coded.
[0173] If BDPCM is applicable, flag information can be transmitted at the CU level when the CU size is smaller than or equal to MaxTsSize (maximum transform skip size) for luma samples and the CU is coded using intra prediction. Here, MaxTsSize refers to the maximum block size for which transform skip mode is allowed. This flag information indicates whether normal intra coding or BDPCM is applied. If BDPCM is applied, a BDPCM prediction direction flag can be transmitted, indicating whether the prediction direction is horizontal or vertical. Then, the block is predicted through normal horizontal or vertical intra prediction using unfiltered reference samples. The residuals are quantized, and the difference value between each quantized residual and its predictor, e.g., the already quantized residual at a neighboring position in the horizontal or vertical direction depending on the BDPCM prediction direction, is coded.
[0174] The syntax elements and their semantics for the above content are shown in the table below.
[0175] [Table 3]
[0176] Table 3 shows the 'sps_bdpcm_enabled_flag' signaled in the SPS (Sequence Parameter Set). When the syntax element 'sps_bdpcm_enabled_flag' is 1, it indicates that flag information indicating whether BDPCM is applied to a coding unit where intra prediction is performed, i.e., 'intra_bdpcm_luma_flag' and 'intra_bdpcm_chroma_flag', exists in the coding unit.
[0177] If the syntax element "sps_bdpcm_enabled_flag" is not present, its value is assumed to be 0.
[0178] [Table 4-1] [Table 4-2]
[0179] The syntax elements 'intra_bdpcm_luma_flag' and 'intra_bdpcm_chroma_flag' in Table 4 indicate whether BDPCM is applied to the current luma coding block or the current chroma coding block, as described in Table 3. If the value of 'intra_bdpcm_luma_flag' or 'intra_bdpcm_chroma_flag' is 1, the transform for the corresponding coding block is skipped, and the prediction mode for the coding block can be set to horizontal or vertical depending on the 'intra_bdpcm_luma_dir_flag' or 'intra_bdpcm_chroma_dir_flag' that indicates the prediction direction. If 'intra_bdpcm_luma_flag' or 'intra_bdpcm_chroma_flag' is not present, this value is considered to be 0.
[0180] When "intra_bdpcm_luma_dir_flag" or "intra_bdpcm_chroma_dir_flag", which indicates the prediction direction, is 0, it indicates that the prediction direction of BDPCM is horizontal, and when the value is 1, it indicates that the prediction direction of BDPCM is vertical.
[0181] The process of intra prediction based on the flag information is shown in the table below.
[0182] [Table 5]
[0183] Table 5 shows the process of deriving the intra prediction mode. When intra_luma_not_planar_flag[xCb][yCb] is 0, the intra prediction mode (IntraPredModeY[xCb][yCb]) is set to INTRA_PLANAR according to "Table 19." When intra_luma_not_planar_flag[xCb][yCb] is 1, the intra prediction mode (IntraPredModeY[xCb][yCb]) can be set to vertical mode (INTRA_ANGULAR50) or horizontal mode (INTRA_ANGULAR18) according to the variable BdpcmDir[xCb][yCb][0].
[0184] The variable BdpcmDir[xCb][yCb][0] is set to the same value as intra_bdpcm_luma_dir_flag or intra_bdpcm_chroma_dir_flag, as shown in Table 4. Therefore, the intra prediction mode can be set to horizontal mode when the variable BdpcmDir[xCb][yCb][0] is 0, and to vertical mode when the variable is 1.
[0185] Also, when BDPCM is applied, the inverse quantization process can be shown as in Table 6.
[0186] [Table 6]
[0187] Table 6 shows the dequantization process for transform coefficients (8.4.3 Scaling process for transform coefficients). When BdpcmFlag[xTbY][yYbY][cIdx] is 1, the dequantized residual value (d[x][y]) can be derived based on the intermediate variable dz[x][y]. When BdpcmDir[xTbY][yYbY][cIdx] is 0, i.e., when intra prediction is performed in horizontal mode, the variable dz[x][y] is derived based on "dz[x-1][y] + dz[x][y]". When BdpcmDir[xTbY][yYbY][cIdx] is 1, i.e., when intra prediction is performed in vertical mode, the variable dz[x][y] is derived based on "dz[x][y-1] + dz[x][y]". That is, the residual at a specific position can be derived based on the sum of the residual at the previous position in the horizontal or vertical direction and the value received as residual information at the specific position. When BDPCM is applied, the difference value between the residual sample value at a specific position (x, y) (x is the horizontal coordinate increasing from left to right, and y is the vertical coordinate increasing from top to bottom, and a position within a two-dimensional block is expressed as (x, y). Also, it indicates the position of (x, y) when the upper left position of the corresponding transform block is placed at (0, 0)) and the residual sample value at the previous position in the horizontal or vertical direction ((x-1, y) or (x, y-1)) is signaled as residual information.
[0188] Meanwhile, according to one example, when BDPCM is applied, an inverse quadratic transform, which is a non-separable transform, such as LFNST, may not be applied. Therefore, when BDPCM is applied, transmission of an LFNST index (transform index) may be omitted. As described above, whether LFNST is applied and which transform kernel matrix for LFNST to apply may be indicated via the LFNST index. For example, an LFNST index value of 0 indicates that LFNST is not applied, and an LFNST index value of 1 or 2 indicates one of two transform kernel matrices constituting an LFNST transform set selected based on the intra prediction mode. More specific embodiments related to BDPCM and LFNST may be applied as follows.
[0189] [First Example]
[0190] BDPCM can be applied only to either the luma component or the chroma component. When the CTU partition tree for the luma component and the CTU partition tree for the chroma component are coded separately (for example, in the dual-tree structure of the VVC standard), if BDPCM is applied only to the luma component, the LFNST index can be transmitted only when BDPCM is not applied to the luma component, and the LFNST index can be transmitted for all blocks to which LFNST can be applied to the chroma component. Conversely, in the dual-tree structure, if BDPCM is applied only to the chroma component, the LFNST index can be transmitted only when BDPCM is not applied to the chroma component, and the LFNST index can be transmitted for all blocks to which LFNST can be applied to the luma component.
[0191] [Second Example]
[0192] When the luma component and the chroma component are coded using the same CTU partition tree, i.e., when they share the same partitioning form (e.g., a single-tree structure in the VVC standard), LFNST may not be applied to both the luma component and the chroma component in a block to which BDPCM is applied. Alternatively, LFNST may be applied to only one component (e.g., the luma component or the chroma component) in a block to which BDPCM is applied, and in this case, only the LFNST index for the corresponding component may be coded and signaled.
[0193] [Third Example]
[0194] When BDPCM is applied only to a specific type of image or partial image (e.g., intra-predicted image, intra-slice, etc.), BDPCM can be configured to be applied only to the corresponding type of image or partial image. For images or partial images to which BDPCM is applied, an LFNST index can be transmitted for each block to which BDPCM is not applied, and for images or partial images to which BDPCM is not applied, an LFNST index can be transmitted for all blocks to which LFNST can be applied. Here, the block may be a coding block or a transform block.
[0195] [Fourth Example]
[0196] BDPCM can be applied only to blocks that are less than a certain size. For example, BDPCM can be configured to be applied only when the width of a block is less than W and the height is less than H, where W and H can both be set to 32. If a block's width is less than W and its height is less than H and BDPCM is applicable, the LFNST index can be transmitted only when the flag indicating whether BDPCM is applicable is coded as 0 (when BDPCM is not applied).
[0197] On the other hand, if the width of a block is greater than W or the height is greater than H, BDPCM is not applied, so there is no need to signal a flag indicating whether BDPCM is applicable, and an LFNST index can be transmitted for all blocks to which LFNST can be applied.
[0198] [Fifth Example]
[0199] Combinations of the first to fourth embodiments can be applied. For example, 1) BDPCM can be applied only to the luma component, 2) BDPCM can be applied only to intra slices, or 3) BDPCM can be applied only when both the width and height are 32 or less, and an LFNST index can be coded and signaled for blocks to which BDPCM is applied.
[0200] The following drawings are created to explain a specific example of the present specification. The names of specific devices and names of specific signals / messages / fields shown in the drawings are provided for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.
[0201] FIG. 7 is a flowchart illustrating the operation of a video decoding device according to one embodiment of the present document.
[0202] Each step disclosed in Fig. 7 may be performed by the decoding apparatus 300 disclosed in Fig. 3. More specifically, S710 may be performed by the entropy decoding unit 310 disclosed in Fig. 3, S720 may be performed by the inverse quantization unit 321 disclosed in Fig. 3, S730 and S740 may be performed by the inverse transform unit 322 disclosed in Fig. 3, and S750 may be performed by the addition unit 340 disclosed in Fig. 3. In addition, the operations of S710 to S750 are based on some of the contents described above with reference to Figs. 4 to 6. Therefore, detailed descriptions that overlap with the contents described above with reference to Figs. 3 to 6 will be omitted or simplified.
[0203] According to one embodiment, a decoding apparatus 300 may derive quantized transform coefficients for a current block from a bitstream (S710). More specifically, the decoding apparatus 300 may decode information on the quantized transform coefficients for the current block from the bitstream and derive quantized transform coefficients for the current block based on the information on the quantized transform coefficients for the current block. The information on the quantized transform coefficients for the current block may be included in a Sequence Parameter Set (SPS) or a slice header, and may include at least one of information on whether a simplified transform (RST) is applied, information on a simplification factor, information on a minimum transform size to which the simplified transform is applied, information on a maximum transform size to which the simplified transform is applied, a simplified inverse transform size, and information on a transform index indicating one of the transform kernel matrices included in the transform set.
[0204] The decoding apparatus 300 according to an embodiment may derive transform coefficients by performing inverse quantization on the quantized transform coefficients for the current block (S720).
[0205] The decoding apparatus 300 according to an embodiment may derive modified transform coefficients based on an inverse non-separable transform or an inverse reduced secondary transform (RST) on the transform coefficients (S730).
[0206] In one example, an inverse non-separable transform or inverse RST can be performed based on an inverse RST matrix, which can be a non-square matrix with fewer columns than rows.
[0207] In one embodiment, S730 may include the steps of: decoding a transform index; determining whether a condition for applying an inverse RST is met based on the transform index; selecting a transform kernel matrix; and, if the condition for applying the inverse RST is met, applying the inverse RST to the transform coefficients based on the selected transform kernel matrix and / or a simplification factor. In this case, the size of the simplified inverse transform matrix may be determined based on the simplification factor.
[0208] The decoding apparatus 300 according to an embodiment may derive residual samples for the current block based on the inverse transform of the modified transform coefficients (S740).
[0209] The decoding device 300 may perform an inverse linear transform on the modified transform coefficients for the current block, in which case the inverse linear transform may be a simplified inverse transform or a conventional separation transform may be used.
[0210] The decoding apparatus 300 according to an embodiment may generate reconstructed samples based on residual samples for a current block and predicted samples for the current block (S750).
[0211] Referring to S730, it can be seen that residual samples for a current block are derived based on the inverse RST for the transform coefficients of the current block. Considering the size of the inverse transform matrix, the size of a typical inverse transform matrix is NxN, but the size of the inverse RST matrix is reduced to NxR. Therefore, compared to performing a conventional transform, memory usage during inverse RST can be reduced by a ratio of R / N. Furthermore, compared to the number of multiplication operations required when using a typical inverse transform matrix (NxN), the number of multiplication operations can be reduced by a ratio of R / N (NxR) when using the inverse RST matrix. Furthermore, since only R transform coefficients need to be decoded when applying the inverse RST, the total number of transform coefficients for the current block is reduced from N to R, compared to the number of transform coefficients required when applying a conventional inverse transform (N), thereby improving decoding efficiency. In summary, S730 can improve the (inverse) transform efficiency and decoding efficiency of the decoding device 300 through the inverse RST.
[0212] FIG. 8 is a control flowchart illustrating a video decoding method according to an embodiment of the present document.
[0213] The decoding apparatus 300 receives coding information such as BDPCM information from a bitstream (S810). The decoding apparatus 300 may also receive transform skip flag information indicating whether a transform skip is applied to the current block, and transform index information for an inverse quadratic transform, i.e., an inverse non-separable transform, i.e., an LFNST index or MTS index information indicating a transform kernel for an inverse linear transform.
[0214] The BDPCM information may include BDPCM flag information indicating whether or not BDPCM is applied to the current block, and direction information indicating the direction in which BDPCM is performed.
[0215] If a BDPCM is applied to the current block, the flag value of the BDPCM may be 1, and if a BDPCM is not applied to the current block, the flag value of the BDPCM may be 0.
[0216] If BDPCM is applied to the current block, the transform skip flag value may be considered to be 1, and if the transform skip flag value is 1, the LFNST index value may be considered to be 0 or may not be received. In other words, if BDPCM is applied to the current block, no transform may be applied to the current block.
[0217] Meanwhile, the tree type of the current block can be classified as a single tree (SINGLE_TREE) or a dual tree (DUAL_TREE) depending on whether the luma block and the corresponding chroma block have separate partition structures. If the chroma block has the same partition structure as the luma block, it can be represented as a single tree, and if the chroma component block has a different partition structure from the luma component block, it can be represented as a dual tree. For example, BDPCM can be applied separately to the luma block or chroma blocks of the current block. If BDPCM is applied to the luma block, no transform index for the luma block may be received, and if BDPCM is applied to the chroma block, no transform index for the chroma block may be received.
[0218] If the tree structure of the current block is a dual tree, BDPCM can be applied to only one of the component blocks, and if the current block has a single tree structure, BDPCM can be applied to only one of the component blocks. In this case, LFNST indexes can be received only for component blocks to which BDPCM is not applied.
[0219] Alternatively, for example, BDPCM can be applied only if the width of the current block is less than or equal to a first threshold and the height of the current block is less than or equal to a second threshold, where the first and second thresholds may be 32 and may be set to the maximum height or width of the transform block on which the transform is performed.
[0220] Meanwhile, the direction information for BDPCM can indicate the horizontal or vertical direction, and quantization information is derived based on the direction information, allowing prediction samples to be derived.
[0221] The decoding apparatus 300 may derive quantized transform coefficients for the current block based on the BDPCM (S820), where the transform coefficients may be untransformed residual sample values.
[0222] When BDPCM is applied to the current block, the residual information received by the decoding apparatus 300 may be a difference value of the quantized residual. Depending on the direction of BDPCM, a difference value of the quantized residual between a previous vertical or horizontal line and a specific line may be received, and the decoding apparatus 300 may derive the quantized residual of the specific line by adding the quantized residual value of the previous vertical or horizontal line to the received difference value of the quantized residual. The quantized residual may be derived based on Equation 11 or Equation 12.
[0223] The decoding apparatus 300 may derive transform coefficients by performing inverse quantization on the quantized transform coefficients (S830), and derive residual samples based on the transform coefficients (S840).
[0224] As described above, when BDPCM is applied to the current block, the dequantized transform coefficients can be derived as residual samples without undergoing a transform process.
[0225] The intra prediction unit 331 may perform intra prediction on the current block based on the direction in which BDPCM is performed (S850).
[0226] When BDPCM is applied to the current block, intra prediction can be performed using it, which may mean that BDPCM can only be applied to intra slices or intra-coded blocks predicted in intra mode.
[0227] Intra prediction is performed based on the direction information for the BDPCM, and the intra prediction mode of the current block can be either a horizontal mode or a vertical mode.
[0228] The decoding apparatus 300 may generate a reconstructed picture based on the residual samples and the predicted samples derived in S750 of FIG. 7 (S860).
[0229] The following drawings are created to explain a specific example of the present specification. The names of specific devices and names of specific signals / messages / fields shown in the drawings are provided for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.
[0230] FIG. 9 is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present document.
[0231] The steps disclosed in Fig. 9 may be performed by the encoding apparatus 200 disclosed in Fig. 2. More specifically, S910 may be performed by the prediction unit 220 disclosed in Fig. 2, S920 may be performed by the subtraction unit 231 disclosed in Fig. 2, S930 and S940 may be performed by the transformation unit 232 disclosed in Fig. 2, and S950 may be performed by the quantization unit 233 and entropy encoding unit 240 disclosed in Fig. 2. In addition, the operations of S910 to S950 are based on some of the content described above with reference to Figs. 4 to 6. Therefore, detailed descriptions that overlap with the content described above with reference to Figs. 2 and 4 to 6 will be omitted or simplified.
[0232] The encoding apparatus 200 according to an embodiment may derive prediction samples based on an intra prediction mode applied to a current block (S910).
[0233] The encoding apparatus 200 according to an embodiment may derive residual samples for the current block (S920).
[0234] According to an embodiment, the encoding apparatus 200 may derive transform coefficients for the current block based on a linear transform of the residual samples (S930). The linear transform may be performed via multiple transform kernels, and in this case, the transform kernels may be selected based on the intra prediction mode.
[0235] The decoding apparatus 300 may perform a quadratic or non-separable transform, specifically, an NSST, on transform coefficients for a current block, where the NSST may be performed based on a simplified transform (RST) or may not be performed based on the RST. When the NSST is performed based on the RST, it may correspond to the operation of S940.
[0236] The encoding apparatus 200 according to one embodiment may derive modified transform coefficients for the current block based on the RST for the transform coefficients (S940). In one example, the RST may be performed based on a simplified transform matrix or a transform kernel matrix, and the simplified transform matrix may be a non-square matrix with the number of rows being less than the number of columns.
[0237] In one embodiment, S940 may include determining whether a condition for applying RST is met, generating and encoding a transform index based on the determination, selecting a transform kernel matrix, and if the condition for applying RST is met, applying RST to the residual samples based on the selected transform kernel matrix and / or a simplification factor. In this case, the size of the simplified transform kernel matrix may be determined based on the simplification factor.
[0238] The encoding device 200 according to one embodiment may perform quantization based on the modified transform coefficients for the current block to derive quantized transform coefficients, and encode information about the quantized transform coefficients (S950).
[0239] More specifically, the encoding apparatus 200 may generate information about the quantized transform coefficients and encode the generated information about the quantized transform coefficients.
[0240] In one example, the information about the quantized transform coefficients may include at least one of information on whether RST is applied, information on a simplification factor, information on the minimum transform size to apply RST to, and information on the maximum transform size to apply RST to.
[0241] Referring to S940, it can be seen that transform coefficients for a current block are derived based on the RST of the residual samples. Considering the size of the transform kernel matrix, the size of a conventional transform kernel matrix is NxN, while the size of a simplified transform matrix is reduced to RxN. Therefore, compared to performing a conventional transform, memory usage during RST can be reduced by a ratio of R / N. Furthermore, compared to the number of multiplication operations (NxN) required when using a conventional transform kernel matrix, the number of multiplication operations can be reduced by a ratio of R / N (RxN) when using a simplified transform kernel matrix. Furthermore, since only R transform coefficients are derived when RST is applied, the total number of transform coefficients for a current block is reduced from N to R, compared to the number of transform coefficients derived when applying a conventional transform (N). This reduces the amount of data transmitted from the encoding apparatus 200 to the decoding apparatus 300. In summary, performing S940 can increase the transform efficiency and coding efficiency of the encoding apparatus 200 via the RST.
[0242] FIG. 10 is a control flowchart illustrating a video encoding method according to an embodiment of the present document.
[0243] The encoding apparatus 200 may derive a predicted sample for the current block based on the BDPCM (S1010).
[0244] The encoding apparatus 200 may derive intra-prediction samples for the current block based on a specific direction in which BDPCM is performed. The specific direction may be a vertical direction or a horizontal direction, and prediction samples for the current block may be generated according to the corresponding intra-prediction mode.
[0245] Meanwhile, the tree type of the current block can be classified as a single tree (SINGLE_TREE) or a dual tree (DUAL_TREE) depending on whether the luma block and the corresponding chroma block have separate partition structures. If the chroma block has the same partition structure as the luma block, it can be represented as a single tree, and if the chroma component block has a different partition structure from the luma component block, it can be represented as a dual tree. For example, BDPCM can be applied separately to the luma block or chroma block of the current block.
[0246] If the tree structure of the current block is a dual tree, BDPCM can be applied to only one of the component blocks, and if the current block has a single tree structure, BDPCM can be applied to only one of the component blocks.
[0247] Alternatively, for example, BDPCM can be applied only if the width of the current block is less than or equal to a first threshold and the height of the current block is less than or equal to a second threshold, where the first and second thresholds may be 32 and may be set to the maximum height or width of the transform block on which the transform is performed.
[0248] The encoding apparatus 200 may derive residual samples for the current block based on the predicted samples (S1020) and perform quantization on the residual samples (S1030).
[0249] Thereafter, the encoding apparatus 200 may derive quantized residual information based on the BDPCM (S1040).
[0250] The encoding apparatus 200 may derive a difference value between a quantized residual sample of a specific line and a quantized residual sample of a previous vertical or horizontal line and the specific line as quantized residual information. That is, a difference value of the quantized residual, rather than a normal residual, is generated as residual information and may be derived based on Equation 9 or 10.
[0251] The encoding apparatus 200 may encode the quantized residual information and coding information for the current block (S1050).
[0252] The encoding apparatus 200 may encode BDPCM information, transform skip flag information indicating whether a transform skip is applied to the current block, and transform index information for an inverse quadratic transform, i.e., an inverse non-separable transform, i.e., an index of an LFNST or an MTS index information indicating a transform kernel of an inverse linear transform.
[0253] The BDPCM information may include flag information of the BDPCM indicating whether the BDPCM is applied to the current block and direction information indicating the direction in which the BDPCM is executed.
[0254] If BDPCM is applied to the current block, the flag value of the BDPCM may be encoded as 1, and if BDPCM is not applied to the current block, the flag value of the BDPCM may be encoded as 0.
[0255] If BDPCM is applied to the current block, the transform skip flag value may be considered to be 1 or may be encoded as 1. Also, if the transform skip flag value is 1, the LFNST index value may be considered to be 0 or may not be encoded. In other words, if BDPCM is applied to the current block, no transform may be applied to the current block.
[0256] Also, as mentioned above, if the tree structure of the current block is a dual tree, BDPCM can be applied to only one of the component blocks, and if the current block is a single tree structure, BDPCM can be applied to only one of the component blocks. In this case, LFNST indices can be encoded only for component blocks to which BDPCM is not applied.
[0257] The direction information for the BDPCM can indicate either a horizontal or vertical direction.
[0258] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may also be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for uniformity of representation.
[0259] Also, in this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about transform coefficients, and the information about the transform coefficients may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0260] In the above-described embodiments, the method is described based on a flowchart as a series of steps or blocks, but this document is not limited to the order of steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Also, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and that different steps may be included, or one or more steps in the flowcharts may be deleted without affecting the scope of this document.
[0261] The method according to the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0262] In this document, when an embodiment is implemented in software, the methods described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be implemented and performed on a computer, processor, microprocessor, controller, or chip.
[0263] In addition, the decoding device and encoding device to which this document is applied may be included in, and may be used to process video signals or data signals, a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a customized video (VoD) service providing device, an over-the-top (OTT) video device, an internet streaming service providing device, a three-dimensional (3D) video device, an image telephone video device, a medical video device, etc. For example, over-the-top (OTT) video devices may include a game console, a Blu-ray player, an internet access TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0264] Furthermore, the processing method to which this document is applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media embodied in the form of a carrier wave (e.g., transmission via the Internet). A bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network. The embodiments of this document may be embodied in a computer program product represented by program code, which can be executed by a computer according to the embodiments of this document. The program code may be stored on a computer-readable carrier.
[0265] FIG. 11 shows an exemplary structure of a content streaming system to which this document is applied.
[0266] Furthermore, the content streaming system to which this document applies can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0267] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted. The bitstream may be generated by an encoding method or a bitstream generation method to which this document applies, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0268] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. The content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0269] The streaming server may receive content from a media storage and / or an encoding server. For example, if content is received from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.
[0270] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can be operated as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
[0271] The claims described herein may be combined in various ways. For example, technical features of method claims herein may be combined to be embodied as an apparatus, and technical features of apparatus claims herein may be combined to be embodied as a method. Furthermore, technical features of method claims herein and technical features of apparatus claims herein may be combined to be embodied as an apparatus, and technical features of method claims herein and technical features of apparatus claims herein may be combined to be embodied as a method.
Claims
1. A video decoding method performed by a decoding device, comprising: determining whether block-based delta pulse code modulation (BDPCM) is applied to the current block; deriving quantized transform coefficients for the current block based on the BDPCM; performing inverse quantization on the quantized transform coefficients to derive transform coefficients; deriving residual samples based on the transform coefficients; determining whether BDPCM is applied to the current block; determining that the width of the current block is less than or equal to a first critical value and that the height of the current block is less than or equal to a second critical value; applying the BDPCM to the current block based on the width of the current block being less than or equal to the first threshold and the height of the current block being less than or equal to the second threshold; Based on the BDPCM being applied to the current block, no inverse non-separable transform is applied to the transform coefficients; A method for decoding video, wherein a value of a transform index for the inverse non-separable transform is assumed to be 0 based on the BDPCM being applied to the current block.
2. the BDPCM is applied to a luma block of the current block or a chroma block of the current block separately; Based on the BDPCM being applied to the luma block, no transform index for the luma block is received; The video decoding method of claim 1 , wherein a transform index for the chroma block is not received based on the BDPCM being applied to the chroma block.
3. The video decoding method of claim 1 , wherein a value of a transform skip flag indicating whether a transform is skipped for the current block is regarded as 1 when the BDPCM is applied to the current block.
4. 2. The image decoding method of claim 1, wherein the first and second threshold values are 32.
5. The video decoding method of claim 1 , wherein the quantized transform coefficients are derived based on direction information regarding a direction in which the BDPCM is performed.
6. The video decoding method of claim 5 , further comprising: performing intra prediction on the current block based on the direction in which the BDPCM is performed.
7. The video decoding method of claim 6 , wherein the direction information indicates a horizontal direction or a vertical direction.
8. A video encoding method performed by a video encoding apparatus, comprising: deriving predicted samples for a current block based on block-based delta pulse code modulation (BDPCM); deriving a residual sample for the current block based on the predicted sample; performing quantization on the residual samples; deriving quantized residual information based on the BDPCM; encoding the quantized residual information and flag information related to the BDPCM for the current block; the flag information is encoded based on the width of the current block being less than or equal to a first critical value and the height of the current block being less than or equal to a second critical value; a non-separable transform is not applied to the current block based on the BDPCM being applied to the current block; and a transform index for the non-separable transform is not encoded based on the BDPCM being applied to the current block.
9. the BDPCM is applied to a luma block of the current block or a chroma block of the current block separately; Based on the fact that the BDPCM is applied to the luma block, a transform index for the luma block is not encoded; The video encoding method of claim 8 , wherein a transform index for the chroma block is not encoded based on the application of the BDPCM to the chroma block.
10. The video encoding method of claim 8 , wherein the first and second threshold values are 32.
11. Intra-predicted samples for the current block are derived based on a particular direction in which the BDPCM is performed; The video encoding method of claim 8 , wherein the quantization is performed on the residual samples based on the particular direction.
12. The video encoding method of claim 11 , wherein the specific direction includes a horizontal direction or a vertical direction.
13. A non-transitory computer-readable digital storage medium for storing a computer program and a bitstream, comprising: A non-transitory computer-readable digital storage medium, the computer program, when executed by one or more processors, causing the one or more processors to perform the video encoding method of any one of claims 8 to 12 to generate the bitstream.
14. A decoding device comprising an entropy decoding unit and a residual processing unit, The entropy decoding unit Determine whether BDPCM (Block-based Delta Pulse Code Modulation) is applied to the current block; deriving quantized transform coefficients for the current block based on the BDPCM; The residual processing unit includes: performing inverse quantization on the quantized transform coefficients to derive transform coefficients; deriving residual samples based on the transform coefficients; The entropy decoding unit configured to determine whether the BDPCM is applied to the current block includes: determining that the width of the current block is equal to or less than a first critical value and that the height of the current block is equal to or less than a second critical value; applying the BDPCM to the current block based on the width of the current block being less than or equal to the first threshold and the height of the current block being less than or equal to the second threshold; Based on the BDPCM being applied to the current block, no inverse non-separable transform is applied to the transform coefficients; A decoding apparatus, wherein a value of a transform index for the inverse non-separable transform applied to the current block is assumed to be 0 based on the fact that the BDPCM is applied to the current block.
15. the BDPCM is applied to a luma block of the current block or a chroma block of the current block separately; Based on the BDPCM being applied to the luma block, no transform index for the luma block is received; The decoding apparatus of claim 14 , wherein a transform index for the chroma block is not received based on the BDPCM being applied to the chroma block.
16. The decoding apparatus of claim 14, wherein a value of a transform skip flag indicating whether a transform is skipped for the current block is regarded as 1 when the BDPCM is applied to the current block.
17. 15. The decoding apparatus of claim 14, wherein the first and second critical values are 32.
18. 15. The decoding apparatus of claim 14, wherein the quantized transform coefficients are derived based on direction information regarding the direction in which the BDPCM is performed.
19. The decoding device further comprises a prediction unit, The decoding apparatus of claim 18, wherein the predictor is configured to perform intra prediction on the current block based on the direction in which the BDPCM is performed.
20. 20. The decoding device of claim 19, wherein the direction information indicates a horizontal direction or a vertical direction.
21. A video encoding device comprising: a prediction unit, a residual processing unit, and an entropy encoding unit; The prediction unit is configured to derive a prediction sample for a current block based on a block-based delta pulse code modulation (BDPCM), The residual processing unit includes: deriving a residual sample for the current block based on the predicted sample; performing quantization on the residual samples; configured to derive quantized residual information based on the BDPCM; the entropy encoding unit is configured to encode the quantized residual information and flag information related to the BDPCM for the current block; the flag information is encoded based on the width of the current block being less than or equal to a first critical value and the height of the current block being less than or equal to a second critical value; a non-separable transform is not applied to the current block based on the BDPCM being applied to the current block; and a transform index for the non-separable transform is not encoded based on the BDPCM being applied to the current block.
22. the BDPCM is applied to a luma block of the current block or a chroma block of the current block separately; Based on the fact that the BDPCM is applied to the luma block, a transform index for the luma block is not encoded; The video encoding apparatus of claim 21 , wherein a transform index for the chroma block is not encoded based on the application of the BDPCM to the chroma block.
23. The video encoding apparatus of claim 21 , wherein the first and second threshold values are 32.
24. Intra-predicted samples for the current block are derived based on a particular direction in which the BDPCM is performed; 22. The video encoding apparatus of claim 21, wherein the quantization is performed on the residual samples based on the particular direction.
25. The video encoding apparatus of claim 24 , wherein the specific direction includes a horizontal direction or a vertical direction.