Method, apparatus, and program for decoding and encoding a symbolization unit

By employing separate transform skip flags and scan patterns for luma and chroma channels, the method optimizes decoding efficiency for high-resolution video formats, addressing the limitations of existing standards in encoding and decoding high-frame-rate video.

JP7712997B2Active Publication Date: 2025-07-24CANON KK

Patent Information

Application Number
JP2023207167
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-03
Filing Date
2023-12-07
Publication Date
2025-07-24
Estimated Expiration
2040-11-04

AI Technical Summary

Technical Problem

The existing video compression standards, such as HEVC, struggle to efficiently encode and decode high-resolution, high-frame-rate video formats like 8K Cube Map Projection, requiring improved compression performance and efficient implementation on modern silicon processes while maintaining a balance between performance and cost.

Method used

The method involves decoding and encoding units with separate transform skip flags for luma and chroma channels, determining a secondary transform index based on these flags, and applying appropriate transforms to residual samples, along with specific scan patterns for residual coefficients to optimize decoding efficiency.

Benefits of technology

This approach enhances decoding efficiency by adapting transform processes based on channel-specific flags, reducing computational overhead and improving performance in high-resolution video decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712997000002
    Figure 0007712997000002
  • Figure 0007712997000003
    Figure 0007712997000003
  • Figure 0007712997000004
    Figure 0007712997000004
Patent Text Reader

Abstract

To provide a decoding method that improves an orthogonal transformation process.SOLUTION: A method is for decoding an encoded unit from a bitstream. In the bitstream, the encoded unit is split from an encoded tree unit of the image using a tree structure, and can have a luma component and a chroma component. The chroma component includes a Cb component and a Cr component. The method includes a first decoding step for decoding a luma conversion skip flag for the luma component from the bitstream when the encoded unit has the luma component, a second decoding step for decoding a first chroma transformation skip flag for the Cb component and a second chroma transformation skip flag for the Cr component from the bitstream when the encoded unit has the chroma component, and a determination step for determining an LFNST index for LFNST processing.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Reference to Related Applications This application claims the benefit of the filing date under 35 U.S.C. § 119 of Australian Patent Application No. 2019275553, filed on Dec. 3, 2019, which is hereby incorporated by reference in its entirety as if fully set forth herein.

[0002] The present invention generally relates to digital video signal processing, and more particularly, to methods, apparatuses, and systems for encoding and decoding blocks of video samples. The present invention also relates to a computer program product including a computer-readable medium having recorded thereon a computer program for encoding and decoding blocks of video samples.

Background Art

[0003] There are currently many applications for video encoding, including applications for the transmission and storage of video data. Also, many video encoding standards have been developed, and others are currently under development. Recent progress in the standardization of video encoding has led to the formation of a group called the “Joint Video Expert Team” (JVET). The Joint Video Expert Team (JVET) includes members of Study Group 16, Question 6 (SG16 / Q6) of the Telecommunication Standardization Sector (ITU-T) of the International Telecommunication Union (ITU), also known as the “Video Coding Experts Group” (VCEG), and members of Joint Technical Committee 1 / Subcommittee 29 / Working Group 11 (ISO / IEC JTC1 / SC29 / WG11) of the International Organization for Standardization / International Electrotechnical Commission, also known as the “Moving Picture Experts Group” (MPEG).

[0004] The Joint Video Expert Team (JVET) issued a Call for Proposals (CfP) and analyzed the responses at its 10th meeting held in San Diego, USA. The submitted proposals demonstrated video compression capabilities significantly exceeding those of the current state-of-the-art video compression standard, "High Efficiency Video Coding" (HEVC). In response to this result, it was decided to initiate a development project for a new video compression standard, "Versatile Video Coding" (VVC). VVC is required to have even higher compression performance than before in response to the increasing sophistication of video formats (higher resolutions and higher frame rates) and the growing market demand for services over WANs where the bandwidth cost is relatively high. In use cases such as immersive video, real-time encoding and decoding of such high-order formats are required. For example, in Cube Map Projection (CMP), an 8K format may be used even if the ultimately rendered "viewport" has a low resolution. VVC must be implementable on modern silicon processes and provide an acceptable trade-off between the achieved performance and the implementation cost. The implementation cost can be considered from one or more perspectives such as silicon area, CPU processor load, memory utilization, and bandwidth. High-order video formats can be processed by dividing the frame area into multiple sections and processing each section in parallel. A bitstream constructed from multiple sections of a compressed frame, suitable for decoding by a "single-core" decoder (i.e., frame-level constraints including bitrate), is allocated to each section according to the needs of the application.

[0005] Video data comprises a sequence of multiple frames of image data where each frame contains one or more color channels. Generally, one primary color channel and two secondary color channels are required. The primary color channel is generally called the "luma" channel, and the secondary color channel(s) are generally called the "chroma" channel(s). Video data is usually displayed in the RGB (Red-Green-Blue) color space, but this color space has a high degree of correlation among its three components. The representation of video data as seen by an encoder or decoder often uses a color space such as YCbCr. YCbCr aggregates the luminance mapped to "luma" by a transfer function into the Y (primary) channel, and chroma into the Cb and Cr (secondary) channels. To use an uncorrelated YCbCr signal, the statistics of the luma channel are significantly different from those of the chroma channels. The main difference is that after quantization, the chroma channels have relatively fewer significant coefficients for a given block compared to the coefficients of the corresponding luma channel block. Further, the Cb and Cr channels may be spatially subsampled (downsampled) at a lower rate than the luma channel, for example, half horizontally and half vertically, which is known as the "4:2:0 chroma format". The 4:2:0 chroma format is used in consumer-oriented applications such as Internet video streaming, television broadcasting, and storage on Blu-ray discs. The method of horizontally subsampling the Cb and Cr channels at half rate and not subsampling vertically is known as the "4:2:2 chroma format". The 4:2:2 chroma format is often used in professional applications such as shooting video for movie production. The 4:2:2 chroma format results in a video that is strong for editing operations such as color grading due to its high sampling rate. Material in the 4:2:2 chroma format is often converted to the 4:2:0 chroma format and then encoded for distribution to consumers. In addition to chroma format, video is also characterized by resolution and frame rate.The resolutions include ultra-high definition (UHD) such as 3840x2160 and "8K" such as 7680x4320, and the frame rates include 60Hz, 120Hz, etc. The sample rate of luma ranges from about 500 megasamples per second to several gigasamples per second. In the case of the 4:2:0 chroma format, the sample rate of each chroma channel is 1 / 4 of the luma sample rate, and in the case of the 4:2:2 chroma format, the sample rate of each chroma channel is 1 / 2 of the luma sample rate.

[0006] The VVC standard is a "block-based" codec, and a frame is first divided into a sequence of square regions known as "Coding Tree Units" (CTUs). If a frame cannot be evenly divided into multiple CTUs, some of the CTUs along the left and bottom edges may be truncated to fit the frame size. CTUs generally occupy relatively large regions, such as 128×128 luma samples. However, CTUs at the right or bottom edge of a frame may be smaller in area. The "coding tree" associated with each CTU may be a single (single) tree ("shared tree") for both the luma and chroma channels, or may include a "fork" into separate trees ("dual trees") for each of the luma and chroma channels. The coding tree defines the decomposition of the CTU region into a set of blocks called "Coding Units" (CUs). CUs are processed in a specific order for encoding or decoding. Separate coding trees for luma and chroma generally start at a luma sample granularity of 64×64, and above that there is a shared tree. Since the 4:2:0 chroma format is adopted, a chroma coding tree with a 32×32 chroma sample region is arranged in the individual coding tree structure starting at a luma sample granularity of 64×64. "Unit" indicates that it applies to all color channels of the coding tree from which the block is derived. A single coding tree results in a coding unit with one luma coding block and two chroma coding blocks. The luma branches of another coding tree result in coding units each with one luma coding block, and the chroma branches of another coding tree result in coding units each with a pair of chroma blocks. The CUs described above are also associated with "Prediction Units" (PUs) and "Transformation Units" (TUs), each of which applies to all color channels of the coding tree from which the CU is derived. Similarly, coding blocks are associated with prediction blocks (PBs) and transformation blocks (TBs), each of which applies to a single color channel. A single tree with CUs spanning color channels in 4:2:0 chroma format video data results in chroma coding blocks having half the width and height of the corresponding luma coding blocks.

[0007] Regardless of the above distinction between "units" and "blocks", the term "block" can be used as a general term for an area or region of a frame to which an operation is applied to all color channels.

[0008] For each CU, a prediction unit (PU) of the content (sample value) of the corresponding area of the frame data is generated. Further, an expression of the difference between the predicted value and the content of the area seen at the input to the encoder (or "spatial area" residual) is formed. The difference for each color channel is converted and encoded as a sequence of residual coefficients and can form one or more transform units (TUs) for a given CU. The conversion applied may be a discrete cosine transform (DCT) or other transform applied to each block of the residual values. This conversion is applied separately and the two-dimensional conversion is performed in two passes. First, a one-dimensional conversion is applied to each row of samples in the block to transform the block. Next, a one-dimensional conversion is applied to each column of the partial results to transform the partial results and generate a final block of transform coefficients that substantially decorrelates the residual samples. In the VVC standard, conversions of various sizes are supported, including conversions of rectangular blocks where each side is a power of two. The transform coefficients are quantized for entropy coding into the bitstream. Further, a non-separable transform stage may be applied. Finally, the application of the transform may be bypassed.

[0009] The features of VVC are intra-frame prediction and inter-frame prediction. Intra-frame prediction generates a prediction of the current sample block within a frame using samples previously processed within the frame. Inter-frame prediction involves generating a prediction of the current block of samples within a frame using a block of samples obtained from a previously decoded frame. The block of samples obtained from the previously decoded frame is offset from the spatial position of the current block according to a motion vector, and in many cases, filtering is applied. Intra-frame prediction blocks can be (i) a uniform sample value ("DC intra prediction"), (ii) a plane with an offset and horizontal and vertical gradients ("plane intra prediction"), (iii) a group of neighboring samples and blocks applied in a specific direction ("angular intra prediction"), or (iv) the result of a matrix product using neighboring samples and selected matrix coefficients. Further discrepancies between the predicted block and the corresponding input samples can be corrected to some extent by encoding the "residual" in the bitstream. The residual is generally transformed from the spatial domain to the frequency domain (in the "primary transform" domain) to form residual coefficients, and may be further transformed by the application of a "secondary transform" (to generate residual coefficients in the "secondary transform" domain). The residual coefficients are quantized according to a quantization parameter, resulting in a loss of the accuracy of the sample reconstruction generated by the decoder, but a reduction in the bitrate of the bitstream.

[0010] Quantization parameters may vary between frames and within each frame. Varying the quantization parameter within a frame is a typical example of a "rate control" encoder. A rate control encoder attempts to generate a bitstream at a substantially constant bitrate regardless of the statistics of the received input samples, such as noise characteristics and degree of motion. Since the bitstream is usually transmitted over a network with limited bandwidth, rate control is a widely used technique to ensure reliable performance on the network regardless of the variability of the original frames input to the encoder. When frames are encoded in parallel sections, flexibility is required in the use of rate control because the required fidelity varies by section.

[0011] Also, implementation costs such as memory usage, high accuracy, and communication efficiency are important.

Summary of the Invention

[0012] An object of the present invention is to substantially overcome or at least improve one or more drawbacks of existing devices.

[0013] One aspect of the present disclosure is a method for decoding an encoding unit of an encoding tree from an encoding tree unit of an image frame from a video bitstream, the encoding unit having one luma color channel and at least one chroma color channel, the method comprising: decoding a luma transform skip flag from the video bitstream for a luma transform block of the encoding unit; decoding at least one chroma transform skip flag from the video bitstream, each of the decoded chroma transform skip flags corresponding to one of at least one chroma transform block of the encoding unit; determining a secondary transform index, the determining comprising: When all of the luma conversion skip flag and the chroma conversion skip flag indicate that the conversion of each conversion block is skipped, determining the secondary conversion index so as to indicate that the secondary conversion is not applied; When all of the luma conversion skip flag and the chroma conversion skip flag indicate that the conversion of each conversion block is not skipped, decoding the secondary conversion index from the video bitstream; including the determining; In order to provide the residual samples of each conversion block of the encoding unit, converting the luma conversion block and the at least one chroma conversion block according to the decoded luma conversion skip flag, the at least one chroma conversion skip flag, and the determined secondary conversion index; Decoding the encoding unit by combining the residual samples of each conversion block of the encoding unit and the prediction block of each block of the encoding unit, wherein each prediction block is generated according to the prediction mode of the encoding unit; providing a method including.

[0014] Another aspect of the present disclosure is a non-transitory computer-readable medium storing a computer program for implementing a method of decoding an encoding unit from an encoding tree unit of an image frame from a video bitstream, wherein the encoding unit has one luma color channel and at least one chroma color channel, and the method includes: decoding a luma conversion skip flag from the video bitstream for a luma conversion block of the encoding unit; decoding at least one chroma conversion skip flag from the video bitstream, wherein each of the decoded chroma conversion skip flags corresponds to one of at least one chroma conversion block of the encoding unit; determining a secondary conversion index, and the determining is: When all of the luma conversion skip flag and the chroma conversion skip flag indicate that the conversion of each conversion block is skipped, determining the secondary conversion index so as to indicate that the secondary conversion is not applied; When all of the luma conversion skip flag and the chroma conversion skip flag indicate that the conversion of each conversion block is not skipped, decoding the secondary conversion index from the video bit stream; including the determining; In order to provide residual samples of each conversion block of the encoding unit, converting the luma conversion block and the at least one chroma conversion block according to the decoded luma conversion skip flag, the at least one chroma conversion skip flag, and the determined secondary conversion index; Decoding the encoding unit by combining the residual samples of each conversion block of the encoding unit and the prediction block of each block of the encoding unit, wherein each prediction block is generated according to the prediction mode of the encoding unit; including providing a non-transitory computer-readable medium.

[0015] Another aspect of the present disclosure provides a system including a memory,

[0016] a processor configured to execute code stored in the memory to implement a method for decoding an encoding unit from an encoding tree unit of an image frame from a video bit stream, the encoding unit having at least one chroma color channel; including, the method comprising: decoding a luma conversion skip flag from the video bit stream for a luma conversion block of the encoding unit; Decoding at least one chroma conversion skip flag from the video bitstream, wherein each decoded chroma conversion skip flag corresponds to one of at least one chroma conversion block of the encoding unit, the decoding, Determining a secondary conversion index, the determining comprising: Determining the secondary conversion index such that when all of the luma conversion skip flag and the chroma conversion skip flag indicate that the conversion of each respective conversion block is skipped, the secondary conversion is not applied; Decoding the secondary conversion index from the video bitstream when all of the luma conversion skip flag and the chroma conversion skip flag indicate that the conversion of each respective conversion block is not skipped; Including, the determining; Converting the luma conversion block and the at least one chroma conversion block according to the decoded luma conversion skip flag, the at least one chroma conversion skip flag, and the determined secondary conversion index to provide residual samples of each conversion block of the encoding unit; Decoding the encoding unit by combining the residual samples of each conversion block of the encoding unit and the prediction block of each block of the encoding unit, wherein each prediction block is generated according to the prediction mode of the encoding unit, the decoding; Including, providing a system.

[0017] Another aspect of the present disclosure is a video decoder, comprising: Receiving an image frame from a video bitstream; Determining an encoding unit of an encoding tree from an encoding tree unit of the image frame, the encoding unit having one luma color channel and at least one chroma color channel; Decoding a luma conversion skip flag from the video bitstream for a luma conversion block of the encoding unit; Decoding at least one chroma transform skip flag from the video bitstream, each of the decoded chroma transform skip flags corresponding to one of at least one chroma transform block of the encoding unit, Determining a secondary transform index, the determination comprising: Determining the secondary transform index such that when all of the luma transform skip flag and the chroma transform skip flag indicate that the transform of each respective transform block is skipped, no secondary transform is applied; Decoding the secondary transform index from the video bitstream when all of the luma transform skip flag and the chroma transform skip flag indicate that the transform of each respective transform block is not skipped; comprising: Transforming the luma transform block and the at least one chroma transform block according to the decoded luma transform skip flag, the at least one chroma transform skip flag, and the determined secondary transform index to provide residual samples of each transform block of the encoding unit; Decoding the encoding unit by combining the residual samples of each transform block of the encoding unit with a prediction block of each block of the encoding unit, each prediction block being generated according to a prediction mode of the encoding unit A video decoder configured as described above is provided.

[0018] One aspect of the present invention is a method for decoding an encoding unit from an encoding tree unit of an image frame from a video bitstream, the method comprising: Determining a scan pattern for a transform block of the encoding unit, wherein the scan pattern traverses the transform block by proceeding through a plurality of non-overlapping collections of sub-blocks of residual coefficients, and the scan pattern proceeds from the current collection of the plurality of collections to the next collection after completing the scan of the current collection, said determining; Decoding residual coefficients from the video bitstream according to the determined scan pattern; Determining a plurality of transform selection indices for the encoding unit, said determining comprising: When the last significant coefficient encountered along the scan pattern is at or within a threshold orthogonal position of the transform block, decoding the plurality of transform selection indices from the video bitstream; When the position of the last significant residual coefficient of the transform block along the scan pattern is outside the threshold orthogonal position, determining the plurality of transform selection indices to indicate that multiple transform selections are not used; Said determining, including; Transforming the decoded residual coefficients by applying a transform according to the plurality of transform selection indices to decode the encoding unit; Providing a method including.

[0019] According to another aspect, the selected scan pattern scans a plurality of residual coefficients of each sub-block in a backward diagonal manner.

[0020] According to another aspect, the selected scan pattern scans a plurality of sub-blocks of each collection in a backward diagonal manner.

[0021] According to another aspect, the selected scan pattern scans the plurality of collections in a backward diagonal manner.

[0022] According to another aspect, the selected scan pattern scans the plurality of collections in a backward raster manner.

[0023] According to another aspect, the plurality of transform selection indexes being zero indicates the application of the inverse DCT-2 in the horizontal and vertical directions.

[0024] According to another aspect, the plurality of transform selection indexes being greater than zero indicates that one of the inverse DST-7 or inverse DCT-8 is applied in the horizontal direction and one of the inverse DST-7 or inverse DCT-8 is applied in the vertical direction.

[0025] According to another aspect, each collection is a two-dimensional array of a plurality of sub-blocks having a width and height of up to 4 sub-blocks.

[0026] Another aspect of the present invention is a non-transitory computer-readable medium storing a computer program for implementing a method of decoding an encoding unit from an encoding tree unit of an image frame from a video bitstream, the method comprising: determining a scan pattern for a transform block of the encoding unit, wherein the scan pattern traverses the transform block by proceeding through a plurality of non-overlapping collections of residual coefficients, and the scan pattern proceeds from the current collection to the next collection of the plurality of collections after completing the scan of the current collection; decoding residual coefficients from the video bitstream according to the determined scan pattern; determining a plurality of transform selection indexes for the encoding unit, the determining comprising: decoding the plurality of transform selection indexes from the video bitstream when the last significant coefficient encountered along the scan pattern is at or within a threshold orthogonal position of the transform block; When the position of the last significant residual coefficient of the transform block along the scan pattern is outside the threshold orthogonal position, determining the multiple transform selection index so as to indicate that multiple transform selections are not being used, including said determining, converting the decoded residual coefficients by applying a transform according to the multiple transform selection index in order to decode the encoding unit, providing a non-transitory computer-readable medium including.

[0027] Another aspect of the present invention is a system, a memory, a processor configured to execute code stored in the memory to implement a method of decoding an encoding unit from an encoding tree unit of an image frame from a video bit stream, including, the method comprising: determining a scan pattern for a transform block of the encoding unit, wherein the scan pattern traverses the transform block by proceeding through a plurality of non-overlapping collections of sub-blocks of residual coefficients, and the scan pattern proceeds from the current collection of the plurality of collections to the next collection after completing the scan of the current collection, said determining; decoding residual coefficients from the video bit stream according to the determined scan pattern; determining a multiple transform selection index for the encoding unit, said determining comprising: decoding the multiple transform selection index from the video bit stream when the last significant coefficient encountered along the scan pattern is at or within the threshold orthogonal position of the transform block; When the position of the last significant residual coefficient of the transform block along the scan pattern is outside the threshold orthogonal position, determining the multiple transform selection index so as to indicate that multiple transform selections are not being used, including said determining, converting the decoded residual coefficients by applying a transform according to the multiple transform selection index to decode the encoding unit, including providing a system.

[0028] Another aspect of the present invention is a video decoder, receiving an image frame from a video bit stream, determining an encoding unit of an encoding tree from an encoding tree unit of the image frame, determining a scan pattern for a transform block of the encoding unit, wherein the scan pattern traverses the transform block by proceeding through a plurality of non-overlapping collections of sub-blocks of residual coefficients, and the scan pattern proceeds from the current collection of the plurality of collections to the next collection after completing the scan of the current collection, decoding residual coefficients from the video bit stream according to the determined scan pattern, determining a multiple transform selection index for the encoding unit, the determination comprising: when the last significant coefficient encountered along the scan pattern is at or within the threshold orthogonal position of the transform block, decoding the multiple transform selection index from the video bit stream; when the position of the last significant residual coefficient of the transform block along the scan pattern is outside the threshold orthogonal position, determining the multiple transform selection index so as to indicate that multiple transform selections are not being used; including, To decode the encoding unit, the decoded residual coefficients are transformed by applying a transformation according to the plurality of transformation selection indices. A video decoder configured as described above is provided.

[0029] Another aspect of the present invention is a method for decoding an encoding unit of an encoding tree of an image frame from a video bit stream, the encoding unit having one luma color channel and at least one chroma color channel, the method comprising: decoding a luma transform skip flag from the video bit stream for the luma transform block of the encoding unit; decoding at least one chroma transform skip flag from the video bit stream, each of the decoded chroma transform skip flags corresponding to one of at least one chroma transform block of the encoding unit; determining a secondary transform index, the determining comprising: when at least one of the luma transform skip flag and the at least one chroma transform skip flag indicates that the transform of each respective transform block is not skipped, decoding a secondary transform index from the video bit stream; when all of the luma transform skip flag and the at least one chroma transform skip flag indicate that the transform of each respective transform block is skipped, determining the secondary transform index to indicate that no secondary transform is applied; including the determining; transforming the luma transform block and the at least one chroma transform block according to the decoded luma transform skip flag, the at least one chroma transform skip flag, and the determined secondary transform index to decode the encoding unit; A method is provided that includes.

[0030] Another aspect of the present invention is a method for decoding an encoding unit of an encoding tree from an encoding tree unit of an image frame from a video bit stream, wherein the encoding unit has at least one chroma color channel, and the method comprises: decoding at least one chroma transform skip flag from the video bit stream, each of the chroma transform skip flags corresponding to one of at least one chroma transform block of the encoding unit; determining a secondary transform index for the at least one chroma transform block of the encoding unit, the determining comprising: when any of the at least one chroma transform skip flag indicates that a transform is applied to each chroma transform block, decoding the secondary transform index from the video bit stream; when all of the chroma transform skip flags indicate that the transform of each transform block is skipped, determining the secondary transform index to indicate that no secondary transform is applied; including the determining; for decoding the encoding unit, transforming each of the at least one chroma transform block according to each chroma transform skip flag and the determined secondary transform index; providing a method.

[0031] Another aspect of the present invention is a non-transitory computer-readable medium storing a computer program for implementing a method for decoding an encoding unit of an encoding tree from an encoding tree unit of an image frame from a video bit stream, wherein the encoding unit has one luma color channel and at least one chroma color channel, and the method comprises: decoding a luma transform skip flag from the video bit stream for a luma transform block of the encoding unit; Decoding at least one chroma transform skip flag from the video bitstream, each of the decoded chroma transform skip flags corresponding to one of at least one chroma transform block of the encoding unit, the decoding, and Determining a secondary transform index, the determining comprising: When at least one of the luma transform skip flag and the at least one chroma transform skip flag indicates that the transform of each respective transform block is not skipped, decoding a secondary transform index from the video bitstream; When all of the luma transform skip flag and the at least one chroma transform skip flag indicate that the transform of each respective transform block is skipped, determining the secondary transform index to indicate that no secondary transform is applied; Including the determining; Converting the luma transform block and the at least one chroma transform block according to the decoded luma transform skip flag, the at least one chroma transform skip flag, and the determined secondary transform index to decode the encoding unit; Providing a non-transitory computer-readable medium including.

[0032] Another aspect of the present invention is a system, A memory, A processor configured to execute code stored in the memory to implement a method of decoding an encoding unit of an encoding tree from an encoding tree unit of an image frame from a video bitstream, the encoding unit having at least one chroma color channel, the processor, and Including, the method comprising: Decoding at least one chroma transform skip flag from the video bitstream, each of the chroma transform skip flags corresponding to one of at least one chroma transform block of the encoding unit, the decoding, and Determining a secondary conversion index for the at least one chroma conversion block of the encoding unit, the determining comprising: When any of the at least one chroma conversion skip flags indicates that conversion is to be applied to each chroma conversion block, decoding the secondary conversion index from the video bitstream; When all of the at least one chroma conversion skip flags indicate that conversion of each conversion block is to be skipped, determining the secondary conversion index so as to indicate that no secondary conversion is to be applied; The determining as described above; For decoding the encoding unit, converting each of the at least one chroma conversion blocks according to the respective chroma conversion skip flag and the determined secondary conversion index; Providing a system as described above.

[0033] Another aspect of the present invention is a video decoder, comprising: Receiving an image frame from a video bitstream; Determining an encoding unit of an encoding tree from an encoding tree unit of the image frame, the encoding unit having one luma color channel and at least one chroma color channel; Decoding a luma conversion skip flag from the video bitstream for a luma conversion block of the encoding unit; Decoding at least one chroma conversion skip flag from the video bitstream, each of the decoded chroma conversion skip flags corresponding to one of the at least one chroma conversion blocks of the encoding unit; Determining a secondary conversion index, the determining comprising: When at least one of the luma conversion skip flag and the at least one chroma conversion skip flag indicates that conversion of each conversion block is not to be skipped, decoding the secondary conversion index from the video bitstream; When all of the luma conversion skip flag and the at least one chroma conversion skip flag indicate that the conversion of each conversion block is skipped, determining the secondary conversion index so as to indicate that the secondary conversion is not applied, including converting the luma conversion block and the at least one chroma conversion block according to the decoded luma conversion skip flag, the at least one chroma conversion skip flag, and the determined secondary conversion index in order to decode the encoding unit A video decoder configured as described above is provided.

[0034] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Hereinafter, at least one embodiment of the present invention will be described with reference to the following drawings and appendices.

[0036]

Figure 1

[0037]

Figure 2A

Figure 2B

[0038]

Figure 3

[0039]

Figure 4

[0040]

Figure 5

[0041]

Figure 6

[0042]

Figure 7A

Figure 7B

[0043]

Figure 8A

Figure 8B

Figure 8C

Figure 8D

[0044]

Figure 9

[0045]

Figure 10

[0046]

Figure 11

[0047]

Figure 12

[0048]

Figure 13

[0049]

Figure 14

[0050]

Figure 15

[0051]

Figure 16

[0052]

Figure 17

[0053]

Figure 18

[0054]

Figure 19

[0055]

Figure 20

DETAILED DESCRIPTION OF THE INVENTION

[0056] In any one or more of the accompanying drawings, when steps and / or features having the same reference numerals are referred to, those steps and / or features have the same function(s) and / or operation(s) for the purposes of this specification, unless a contrary intention appears.

[0057] The syntax of the bitstream format of the video compression standard is defined as a hierarchical structure of "syntax structure". Each syntax structure defines a set of syntax elements, some of which may be conditional on other elements. In the syntax, the compression efficiency is improved by allowing only combinations of syntax elements that correspond to useful combinations of tools. Further, the complexity is reduced by prohibiting combinations of syntax elements that, even though implementable, are considered to have insufficient compression advantages over the resulting implementation cost.

[0058] FIG. 1 is a schematic block diagram showing the functional modules of a video encoding and decoding system 100. The system 100 signals primary and secondary conversion parameters so that a compression efficiency gain is achieved.

[0059] System 100 includes a source device 110 and a destination device 130. Communication channel 120 is used to communicate encoded video information from source device 110 to destination device 130. In some arrangements, either or both of source device 110 and destination device 130 may comprise respective cellular phone handsets or “smartphones,” in which case communication channel 120 is a wireless channel. In other configurations, source device 110 and destination device 130 may be comprised of video conferencing devices, in which case communication channel 120 is typically a wired channel such as an Internet connection. Further, source device 110 and destination device 130 may comprise any of a wide range of devices including over-the-air television broadcasts, cable television applications, Internet video applications (including streaming), and devices that support applications in which encoded video data is captured onto some computer-readable storage medium such as a hard disk drive within a file server.

[0060] As shown in FIG. 1, source device 110 includes a video source 112, a video encoder 114, and a transmitter 116. Video source 112 typically comprises a source (shown as 113) of captured video frame data such as an image capture sensor, a previously captured video sequence stored on a non-transitory storage medium, or a video feed from a remote image capture sensor. Also, video source 112 may be the output of a computer graphics card, e.g., one that displays video output of an operating system and various applications running on a computing device such as a tablet computer. Examples of source device 110 that can include an image capture sensor as video source 112 include smartphones, video camcorders, professional video cameras, and network video cameras.

[0061] As further described with reference to FIG. 3, the video encoder 114 converts (or “encodes”) the captured frame data (indicated by arrow 113) from the video source 112 into a bitstream (indicated by arrow 115). The bitstream 115 is transmitted by the transmitter 116 as encoded video data (or “encoded video information”) via the communication channel 120. Also, the bitstream 115 can be stored in a non-transitory storage device 122, such as a “flash” memory or a hard disk drive, until it is later transmitted via the communication channel 120 or in place of transmission via the communication channel 120. For example, the encoded video data may be provided to a customer on demand via a wide area network (WAN) for a video streaming application.

[0062] The destination device 130 includes a receiver 132, a video decoder 134, and a display device 136. The receiver 132 receives the encoded video data from the communication path 120 and passes the received video data as a bitstream (indicated by arrow 133) to the video decoder 134. Then, the video decoder 134 outputs the decoded frame data (indicated by arrow 135) to the display device 136. The decoded frame data 135 has the same chroma format as the frame data 113. Examples of the display device 136 include a cathode ray tube, a smartphone or a tablet computer, a computer monitor, or a liquid crystal display such as those mounted on a standalone television. Also, the functions of each of the source device 110 and the destination device 130 can be embodied in a single device, examples of which include a mobile phone terminal and a tablet computer. The decoded frame data may be further converted before presentation to the user. For example, a “viewport” having specific latitude and longitude can be rendered from the decoded frame data using a projection format to represent a 360 o view of the scene.

[0063] Regardless of the exemplary devices described above, each of the source device 110 and the destination device 130 may typically be configured within a general-purpose computing system by a combination of hardware and software components. FIG. 2A shows such a computer system 200, which includes a computer module 201, input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 which may be configured as a video source 112, and a microphone 280, output devices such as a printer 215, a display device 214 which may be configured as a display 136, and a loudspeaker 217. The external modulator-demodulator (modem) transceiver device 216 may be used by the computer module 201 to communicate with the communication network 220 via the connection 221. The communication network 220, which may represent the communication channel 120, may be a WAN such as the Internet, a cellular communication network, or a private WAN. If the connection 221 is a telephone line, the modem 216 may be a conventional "dial-up" modem. Alternatively, if the connection 221 is a high-capacity (e.g., cable or optical) connection, the modem 216 may be a broadband modem. Also, a wireless modem may be used for a wireless connection to the communication network 220. The transceiver device 216 may provide the functions of the transmitter 116 and the receiver 132, and the communication channel 120 may be embodied in the connection 221.

[0064] Computer module 201 typically includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 may have a semiconductor random access memory (RAM) and a semiconductor read only memory (ROM). The computer module 201 also includes an audio-video interface 207 coupled to a video display 214, a loudspeaker 217, and a microphone 280, an I / O interface 213 coupled to a keyboard 202, a mouse 203, a scanner 226, a camera 227, and optionally a joystick or other human interface device (not shown), and an interface 208 for an external modem 216 and a printer 215, including a number of input / output (I / O) interfaces. The signal from the audio-video interface 207 to the computer monitor 214 is generally the output of a computer graphics card. In some implementations, the modem 216 may be incorporated within the computer module 201, for example, within the interface 208. The computer module 201 also has a local network interface 211, which enables the coupling of the computer system 200 via a connection 223 to a local area communication network 222 known as a local area network (LAN). As shown in Figure 2A, the local communication network 222 can also be coupled to a wide area network 220 via a connection 224, which will typically include a so-called "firewall" device or a device with a similar function. The local network interface 211 may be composed of an Ethernet (trademark) circuit card, a Bluetooth (trademark) wireless device, or an IEEE802.11 wireless device, but for the interface 211, many other types of interfaces may be implemented. Also, the local network interface 211 may provide the functions of a transmitter 116 and a receiver 132, and the communication channel 120 may also be embodied in the local communication network 222.

[0065] The I / O interfaces 208 and 213 may provide either or both serial and parallel connections, the former typically being implemented in accordance with the Universal Serial Bus (USB) standard and having a corresponding USB connector (not shown). The storage device 209 typically includes a hard disk drive (HDD) 210. Also, other storage devices such as a floppy disk drive or a magnetic tape drive (not shown) can be used. The optical disk drive 212 is typically provided to function as a non-volatile source of data. As a suitable data source to the computer system 200, for example, portable memory devices such as optical disks (e.g., CD-ROM, DVD, Blu-ray Disc (trademark)), USB-RAM, portable, external hard drives, and floppy disks may be used. Typically, any of the HDD 210, optical drive 212, networks 220 and 222 may also be configured to operate as a video source 112 or as a destination for decoded video data stored for playback via the display 214. The source device 110 and the destination device 130 of the system 100 may be embodied in the computer system 200.

[0066] The components 205 - 213 of the computer module 201 typically communicate in a manner that provides a conventional mode of operation of the computer system 200 known to those of skill in the art via an interconnected bus 204. For example, the processor 205 is coupled to the system bus 204 using connection 218. Similarly, the memory 206 and the optical disk drive 212 are coupled to the system bus 204 by connection 219. Examples of computers that can implement the described apparatus include IBM-PCs and compatibles, Sun SPARC stations, Apple Macs (trademark) or similar computer systems.

[0067] When appropriate or desired, video encoder 114 and video decoder 134, and the methods described below can be implemented using computer system 200. In particular, video encoder 114, video decoder 134, and the methods described can be implemented as one or more software application programs 233 executable within computer system 200. In particular, the steps of video encoder 114, video decoder 134, and the methods described are effected by instructions 231 (see FIG. 2B) within software 233 executed within computer system 200. The software instructions 231 may each be formed as one or more code modules for performing one or more particular tasks. Also, the software may be divided into two separate parts, in which case the first part and corresponding code modules perform the methods described, and the second part and corresponding code modules manage the user interface between the first part and the user.

[0068] The software may be stored, for example, on a computer-readable medium including a storage device as described below. The software is loaded from the computer-readable medium into computer system 200 and then executed by computer system 200. A computer-readable medium having such software or a computer program recorded thereon is a computer program product. Use of the computer program product in computer system 200 preferably provides an advantageous apparatus for implementing video encoder 114, video decoder 134 and the methods described.

[0069] Software 233 is typically stored in HDD 210 or memory 206. The software is loaded from the computer-readable medium into computer system 200 and executed by computer system 200. Thus, for example, software 233 may be stored on an optically readable disk storage medium (e.g., CD-ROM) 225 read by optical disk drive 212.

[0070] In some examples, the application program 233 may be encoded on one or more CD-ROMs 225 and supplied to the user, and read via the corresponding drive 212, or alternatively, may be read by the user from the network 220 or 222. Further, the software can also be loaded into the computer system 200 from other computer-readable media. A computer-readable storage medium refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray Discs (trademarks), hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical disks, or computer-readable cards such as PCMCIA cards, regardless of whether such devices are inside or outside the computer module 201. Examples of transitory or non-tangible computer-readable transmission media that can also participate in providing software, application programs, instructions and / or video data or encoded video data to the computer module 401 include wireless or infrared transmission paths, network connections with other computers or network devices, and the Internet or intranets including information recorded in emails and websites.

[0071] The second portion of the application program 233 described above and the corresponding code module may be executed to implement one or more graphical user interfaces (GUIs) that are rendered on the display 214 or otherwise presented. Typically, through the operation of the keyboard 202 and mouse 203, the user of the computer system 200 and the application can operate the interface in a functionally adaptable manner to provide control commands and / or input to the application(s) associated with the GUI(s). Also, other forms of functionally adaptable user interfaces can be implemented, such as audio prompts output via the loudspeaker 217 and an audio interface that utilizes the user's voice commands input via the microphone 280.

[0072] FIG. 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents the logical aggregation of all memory modules (including the HDD 209 and the semiconductor memory 206) accessible by the computer module 201 of FIG. 2A.

[0073] When the computer module 201 is initially powered on, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in the ROM 249 of the semiconductor memory 206 in FIG. 2A. Note that a hardware device that stores software such as the ROM 249 may be referred to as firmware. The POST program 250 inspects the hardware within the computer module 201 to ensure correct functionality and typically checks the processor 205, the memory 234 (209, 206), and also the basic input / output system software (BIOS) module 251, which is typically also stored in the ROM 249. When the POST program 250 is executed successfully, the BIOS 251 boots the hard disk drive 210 in FIG. 2A. The activation of the hard disk drive 210 causes the bootstrap loader program 252 resident on the hard disk drive 210 to be executed via the processor 205. Thereby, the operating system 253 is loaded into the RAM memory 206, and the operation of the operating system 253 is thereby started. The operating system 253 is a system-level application executable by the processor 205 and performs various high-level functions including processor management, memory management, device management, storage management, software application interface, and general-purpose user interface.

[0074] The operating system 253 manages the memory 234 (209, 206) such that each process or application running on the computer module 201 has sufficient memory to execute without colliding with the memory allocated to another process. Further, different types of memory available in the computer system 200 of FIG. 2A must be appropriately used so that each process can be effectively executed. Thus, the aggregated memory 234 is not intended to describe how particular segments of the memory are allocated (unless otherwise specified), but rather to provide a general view of the memory accessible by the computer system 200 and how such is used.

[0075] As shown in FIG. 2B, the processor 205 includes a number of functional modules including a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a number of storage registers 244-246 in a register section. One or more internal buses 241 functionally interconnect these functional modules. Also, the processor 205 typically has one or more interfaces 242 for communicating with external devices via the system bus 204 using the connection 218. The memory 234 is coupled to the bus 204 using the connection 219.

[0076] The application program 233 includes a series of instructions 231 that can include conditional branch instructions and loop instructions. The program 233 may also include data 232 used in the execution of the program 233. The instructions 231 and the data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237 respectively. Depending on the relative sizes of the instructions 231 and the memory locations 228 - 230, a particular instruction may be stored in a single memory location as depicted by the instruction shown at memory location 230. Alternatively, the instructions may be segmented into several parts each stored in a different memory location as depicted by the instruction segments shown at memory locations 228 and 229.

[0077] Generally, the processor 205 is provided with a set of instructions to be executed therein. The processor 205 waits for subsequent inputs, to which the processor 205 reacts by executing another set of instructions. Each input may be provided from one or more of a number of sources including data generated by one or more of the input devices 202, 203, data received from an external source via one of the networks 220, 202, data obtained from one of the storage devices 206, 209, or data obtained from a storage medium 225 inserted into the corresponding reader 212, all of which are depicted in FIG. 2A. The execution of a series of instructions may, in some cases, involve an output of data. Also, the execution may include storing data or variables in the memory 234.

[0078] The video encoder 114, the video decoder 134, and the methods described may use input variables 254 stored in the memory 234 at corresponding memory locations 255, 256, 257. The video encoder 114, the video decoder 134, and the methods described generate output variables 261, which are stored in the memory 234 at corresponding memory locations 262, 263, 264. Intermediate variables 258 may be stored at memory locations 259, 260, 266, 267.

[0079] Referring to the processor 205 of FIG. 2B, the registers 244, 245, 246, the arithmetic logic unit (ALU) 240, and the control unit 239 cooperate to execute a sequence of micro-operations necessary to perform "fetch, decode, and execute" cycles for all instructions within the instruction set that makes up the program 233. Each fetch, decode, execute cycle includes the following: A fetch operation to fetch or read an instruction 231 from memory locations 228, 229, 230, A decode operation where the control unit 239 determines which instruction was fetched, and An execute operation where the control unit 239 and / or the ALU 240 execute the instruction.

[0080] Subsequent fetch, decode, and execute cycles for the next instruction may then be performed. Similarly, a store cycle where the control unit 239 stores or writes a value to the memory location 232 may be performed.

[0081] Each step or sub-process in the methods of FIGS. 13 - 16, which will be described hereinafter, is associated with one or more segments of the program 233 and typically, the register section 244, 245, 247, the ALU 240, and the control unit 239 within the processor 205 cooperate to perform fetch cycles, decode cycles, and execute cycles for all instructions within the instruction set for the indicated segments of the program 233.

[0082] FIG. 3 is a schematic block diagram showing the functional modules of the video encoder 114. FIG. 4 is a schematic block diagram showing the functional modules of the video decoder 134. Generally, data passes between the functional modules in the video encoder 114 and the functional modules in the video decoder 134 in groups of samples or coefficients, or as an array, as if the blocks were divided into sub-blocks of a fixed size. The video encoder 114 and the video decoder 134 may be implemented using the general-purpose computer system 200 as shown in FIGS. 2A and 2B, where the various functional modules are implemented by dedicated hardware in the computer system 200 or by software executable within the computer system 200, such as one or more software code modules of the software application program 233 resident on the hard disk drive 205, and the execution thereof may be controlled by the processor 205. Alternatively, the video encoder 114 and the video decoder 134 may be implemented by a combination of dedicated hardware and software executable within the computer system 200. The video encoder 114, the video decoder 134, and the described methods may alternatively be implemented in dedicated hardware, such as one or more integrated circuits that perform the functions or sub-functions of the described methods. Such dedicated hardware may include a graphics processing unit (GPU), a digital signal processor (DSP), an application specific standard product (ASSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or one or more microprocessors and associated memories. In particular, the video encoder 114 includes modules 310 to 386, and the video decoder 134 includes modules 420 to 496, each of which may be implemented as one or more software code modules of the software application program 233.

[0083] The video encoder 114 of FIG. 3 is an example of a video encoding pipeline for Versatile Video Coding (VVC), but other video codecs can also be used to perform the processing stages described herein. The video encoder 114 receives captured frame data 113 such as a series of frames, and each frame includes one or more color channels. The frame data 113 includes a two-dimensional array of samples of luma ("luma channel") and chroma ("chroma channel") arranged in a "chroma format" such as, for example, a 4:0:0, 4:2:0, 4:2:2, or 4:4:4 chroma format. The block partitioner 310 first divides the frame data 113 into Coding Tree Units (CTUs) that are generally square in shape and configured with a specific size for the CTUs. The size of the CTUs may be, for example, 64×64, 128×128, or 256×256 luma samples.

[0084] The block partitioner 310 further divides each CTU into one or more Coding Units (CUs) according to either a shared coding tree or a luma coding tree and a chroma coding tree at the point where the shared coding tree branches into a luma branch and a chroma branch. The luma channel may sometimes also be called the primary color channel. Each chroma channel may sometimes also be called the secondary color channel. The CUs have various sizes and may include both square and non-square aspect ratios. The operation of the block partitioner 310 will be further described with reference to FIGS. 13 and 14. However, in the VVC standard, the CU / CB, PU / PB, and TU / TB always have side lengths that are powers of two. Thus, the current CU represented as 312 proceeds according to iterations for one or more blocks of the CTU according to the shared coding tree or the luma coding tree and the chroma coding tree of the CTU and is output from the block partitioner 310. Options for partitioning the CTU into CBs are further described below with reference to FIGS. 5 and 6.

[0085] The CTUs obtained from the first split of frame data 113 may be scanned in raster scan order or grouped into one or more "slices". The slice may be an "intra" (or "I") slice. An intra slice (I slice) does not contain inter-predicted CUs and, for example, only intra prediction is used. Alternatively, the slice may be a uni-directional prediction or a bi-directional prediction (a "P" or "B" slice, respectively), indicating the additional availability of one or two reference blocks for predicting CUs, known as "uni-directional prediction" and "bi-directional prediction", respectively.

[0086] In an I slice, the coding tree of each CTU may branch into two coding trees, one for luma and one for chroma, at levels 64×64 and below. By using separate trees, different block structures can exist for luma and chroma within the 64×64 luma region of the CTU. For example, a large number of small luma CBs may be arranged in a large chroma CB, or vice versa. In a P or B slice, a single coding tree for the CTU defines a common block structure for luma and chroma. The blocks resulting from the single tree may be either intra-predicted or inter-predicted.

[0087] For each CTU, the video encoder 114 operates in two stages. In the first stage (referred to as the "search" stage), the block partitioner 310 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CU. In the first stage, various candidate CUs are tested to select a CU that provides relatively low distortion and relatively high compression efficiency. This test generally includes Lagrangian optimization, whereby candidate CUs are evaluated based on a weighted combination of rate (coding cost) and distortion (error with respect to the input frame data 113). The "best" candidate CU (the CU with the smallest evaluated rate / distortion) is selected for encoding into the subsequent bitstream 115. Included in the evaluation of candidate CUs is the option of using a CU for a given region, or further dividing the region according to various partitioning options and encoding each of the resulting smaller regions with a further CU, or further dividing the region. As a result, both the coding tree and the CU itself are selected in the search stage.

[0088] Video encoder 114 generates a prediction block (PU) indicated by arrow 320 for each CU, e.g., CU312. PU 320 predicts the content of the associated CU 312. Subtractor module 322 generates a difference between PB 320 and CB 312, indicated as 324 (or "residual" to indicate that the difference is in the spatial domain). The difference 324 is the block-size difference between corresponding samples of PU 320 and CU 312. The difference 324 is an array of block-sized differences between corresponding samples of PU 320 and CU 312, generated for each color channel of CU 312. If primary and (optionally) secondary transforms are performed, the difference 324 is transformed in modules 326 and 330 and passed to quantization module 334 for quantization via multiplexing 333. If the transform is skipped, the difference 324 is passed directly to quantization module 334 for quantization via multiplexing 333. The selection between transform and transform skip is made independently for each TB associated with CU 312. The resulting quantized residual coefficients are represented as TBs (for each color channel of CU 312) indicated by arrow 336. PU 320 and the associated TB 336 are typically selected from among many possible candidate CUs, e.g., based on an evaluated cost or distortion.

[0089] A candidate CU is a CU obtained from one of the prediction modes available to video encoder 114 for the associated PB and its resulting residual. When combined with the PB predicted in video decoder 114, the addition of TB 336 after inverse transformation to the spatial domain reduces the difference between the decoded CU and the original CU 312 at the expense of additional signaling in the bitstream.

[0090] Each candidate coding unit (CU), i.e., a prediction unit (PU) combined with one transform block (TB), thus has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of the CU is estimated as the difference of sample values, e.g., the sum of absolute differences (SAD) or the sum of squared differences (SSD). The estimated values obtained from each candidate PU may be such that the mode selector 386 determines the intra prediction mode using the difference 324. The prediction mode 387 indicates the decision to use a specific prediction mode, e.g., intra-frame prediction or inter-frame prediction, for the current CU. For intra prediction CUs belonging to the shared coding tree, independent intra prediction modes are specified for the luma PB and the chroma PB. For intra prediction CUs belonging to the luma or chroma branch of the dual coding tree, one intra prediction mode is applied to the luma PB or the chroma PB, respectively. The estimation of the coding cost associated with each candidate prediction mode and the corresponding residual coding can be performed at a much lower cost than the entropy coding of the residuals. Thus, even in a real-time video coder, a large number of candidate modes can be evaluated to determine the optimal mode in terms of rate-distortion.

[0091] The Lagrange or similar optimization process can be employed for both the selection of the optimal partitioning of the CTU into CBs (by the block partitioner 310) and the selection of the best prediction mode from a plurality of possible prediction modes. By applying the Lagrange optimization process for candidate modes in the mode selector module 386, the intra prediction mode 387 that performs the minimum cost measurement, the secondary transform index 388, and the primary transform type 389, and the transform skip flag 390 (one for each TB) are selected.

[0092] In the second stage of the operation of the video encoder 114 (referred to as the "encoding" stage), in the video encoder 114, an iteration for each determined encoding tree (s) of each CTU is executed. For CTUs using separate trees, for each 64×64 luma region of the CTU, first the luma encoding tree is encoded, and then the chroma encoding tree is encoded. Only the luma CBs are encoded within the luma encoding tree, and only the chroma CBs are encoded within the chroma encoding tree. For CTUs using a shared tree, according to the common block structure of the shared tree, the CUs, that is, the luma CBs and the chroma CBs, are described in one tree.

[0093] Entropy encoder 338 supports both variable-length coding of syntax elements and arithmetic coding of syntax elements. Parts of the bitstream such as "parameter sets", e.g., sequence parameter set (SPS), picture parameter set (PPS), picture header (PH), etc., use a combination of fixed-length codewords and variable-length codewords. A slice (also called a continuous part) has a slice header that uses variable-length coding and slice data that uses arithmetic coding. The picture header defines parameters specific to the current slice, such as the quantization parameter offset at the picture level. The slice data includes the syntax elements of each CTU within the slice. When using variable-length coding and arithmetic coding, sequential syntax analysis is required within each part of the bitstream. These parts may be delimited by start codes to form "network abstraction layer units" or "NAL units". Arithmetic coding is supported using a context-adaptive binary arithmetic coding process. The arithmetic-coded syntax elements are composed of a sequence of one or more "bins". A bin, like a bit, has a value of "0" or "1". However, a bin is not coded as an individual bit within bitstream 115. A bin has an associated predicted value (or "likely" or "most likely") and an associated probability, called a "context". If the actual bin to be coded matches the predicted value, the "most probable symbol" (MPS) is coded. Coding of the most probable symbol is relatively inexpensive in terms of the bits consumed in bitstream 115, including the cost corresponding to less than one discrete bit. If the actual bin to be coded does not match the likely value, the "least probable symbol" (LPS) is coded. Coding of the least probable symbol is relatively costly in terms of the number of bits consumed. Bin coding techniques can efficiently code bins where the probabilities of "0" and "1" are skewed. For a syntax element with two possible values (i.e., a "flag"), one bin is sufficient. For a syntax element with many possible values, a sequence of bins is required.

[0094] The presence of subsequent bins within the sequence may be determined based on the values of previous bins within the sequence. Further, each bin may be associated with two or more contexts. The selection of a particular context can depend on previous bins of the syntax element, bin values of adjacent syntax elements (i.e., those from adjacent blocks), and the like. Each time a context-encoded bin is encoded, the context (if any) selected for that bin is updated in a way that reflects the new bin value. Thus, the binary arithmetic coding scheme is said to be adaptive.

[0095] Also supported by video encoder 114 are bins without contexts ("bypass bins"). Bypass bins are encoded assuming an equiprobable distribution between "0" and "1". Thus, each bin encodes the cost of 1 bit of the bitstream 115. The absence of a context can save memory and reduce complexity, and thus bypass bins are used where the distribution of the values of a particular bin is not skewed. An example of an entropy coder that employs context and adaptation is known in the art as CABAC (Context Adaptive Binary Arithmetic Coder), and many variations of this coder have been adopted for video coding.

[0096] Entropy coder 338 encodes secondary transform index 388 using primary transform type 389, one transform skip flag (i.e., 390) for each TB of the current CU, and a combination of context-encoded and bypass-encoded bins, and intra prediction mode 387 if applicable to the current CU. Secondary transform index 388 is signaled when the residual associated with the transform block contains significant residual coefficients only at those coefficient positions that are the target of the transform to primary coefficients by the application of the secondary transform.

[0097] The multiplexing module 384 outputs the PB320 from the intra-frame prediction module 364 according to the determined best intra prediction mode selected from the tested prediction modes of each candidate CB. The candidate prediction modes do not have to include all possible prediction modes supported by the video encoder 114. Intra prediction is classified into three types. "DC intra prediction" includes shifting the PB with a single value representing the average of neighboring reconstructed samples. "Planar intra prediction" includes shifting samples to the PB according to a plane where the DC offset and the vertical and horizontal gradients are derived from neighboring reconstructed neighboring samples. The neighboring reconstructed samples typically include a row of reconstructed samples that is above the current PB and extends to some extent to the right of the PB, and a column of reconstructed samples that is to the left of the current PB and extends downward beyond the PB. "Angular intra prediction" includes shifting the PB with the reconstructed neighboring samples that are filtered and propagated across the PB in a specific direction (or "angle"). In VVC, 65 angles are supported, and additional angles that are not available for square blocks can be utilized for rectangular blocks, and a total of 87 angles can be generated. For chroma PB, as a fourth intra prediction, it is possible to generate the PB from the collocation of luma reconstructed samples by the cross-component linear model (CCLM) mode. There are three different CCLM modes, and each mode uses a different model derived from adjacent luma and chroma samples. This model is used to generate a block of chroma PB samples from the placed luma samples.

[0098] When previously reconstructed samples are not available, such as at the edge of the frame, half of the sample range is used as the default half-tone value. For example, in the case of 10-bit video, the value 512 is used. Since there are no previously available samples for the CB at the upper left position of the frame, the angular and planar intra prediction modes generate a plane of samples with the same output as the DC prediction mode, i.e., the half-tone value as the magnitude.

[0099] In inter-frame prediction, the prediction block 382 is generated by the motion compensation module 380 using samples from one or two frames preceding the current frame in the coded order of frames in the bitstream, and output as PB320 by the multiplexing module 384. Further, in inter-frame prediction, usually, a single coding tree is used for both the luma channel and the chroma channel. The order of the coded frames in the bitstream may be different from the order of the frames during shooting or display. When one frame is used for prediction, the block is called "unidirectional prediction" and one motion vector is associated with it. When two frames are used for prediction, the block is called "bidirectional prediction" and has two associated motion vectors. In the case of a P slice, each CU is intra-predicted or unidirectionally predicted. In the case of a B slice, each CU is either intra-predicted, unidirectionally predicted, or bidirectionally predicted. Frames are usually coded using a "picture group" structure, which enables temporal hierarchicalization of the frames. A plurality of frames may be divided into a plurality of slices, and each slice codes a part of the frame. Due to the temporal hierarchicalization of the frames, the frames can refer to the preceding and subsequent images in the order in which the frames are to be displayed. The images are coded in an order necessary to ensure that the dependencies for decoding each frame are satisfied.

[0100] Samples are selected according to motion vector 378 and reference picture index. The motion vector 378 and reference picture index are applied to all color channels, and thus, inter prediction is explained from the perspective of operations on PUs rather than PBs mainly, i.e., the decomposition of each CTU into one or more inter prediction blocks is described using a single coding tree. Inter prediction may have different numbers and accuracies of motion parameters. Motion parameters typically consist of a reference frame index indicating which reference frame to use from a list of reference frames and a spatial transformation for each of the reference frames, but may include more frames, special frames, or complex affine parameters such as scaling and rotation. Further, a predetermined motion refinement process may be applied to generate a dense motion estimation value based on the referenced sample block.

[0101] PU320 is determined and selected, and when PU320 is subtracted from the original sample block by subtractor 322, the residual with the lowest encoding cost represented as 324 is obtained and subjected to irreversible compression. The irreversible compression process consists of steps of transformation, quantization, and entropy encoding. The forward primary transformation module 326 applies a forward transformation to the difference 324, transforms the difference 324 from the time domain to the frequency domain, and generates primary transformation coefficients represented by arrow 328 according to the primary transformation type 389. The maximum primary transformation size in one dimension is either a 32-point DCT-2 transformation or a 64-point DCT-2 transformation. When the CB to be encoded is larger than the maximum supported primary transformation size represented as the block size, i.e., larger than 64×64 or 32×32, the primary transformation 326 is applied in a tiled manner to transform all samples of the difference 324. When each application of the transformation operates on a TB of difference 324 larger than 32×32, for example, 64×64, all resulting primary transformation coefficients 328 outside the upper left 32×32 region of the TB are set to zero, i.e., discarded. For a TB sized up to 32×32, the primary transformation type 389 can indicate the application of a combination of DST-7 and DCT-8 transformations in the horizontal and vertical directions. The remaining primary transformation coefficients 328 are passed to the forward secondary transformation module 330.

[0102] The secondary transformation module 330 generates secondary transformation coefficients 332 according to the secondary transformation index 388. The secondary transformation coefficients 332 are quantized by module 334 according to the quantization parameters related to the CB, generating residual coefficients 336. When the transform skip flag 390 indicates that transform skip is valid for the TB, the difference 324 is passed to the quantizer 334 via multiplexing 333.

[0103] The forward first transform of module 326 is typically separable and transforms the set of rows of each TB and then the set of columns. The forward first transform module 326 uses either a type II discrete cosine transform (DCT-2) in either the horizontal or vertical direction, according to the first transform type 389, or, for the luma TBS, a combination of a type VII discrete sine transform (DST-7) and a type VIII discrete cosine transform (DCT-8) in either the horizontal or vertical direction. The use of the combination of DST-7 and DCT-8 is called "multi-transform selection set" (MTS) in the VVC standard. When using DCT-2, the maximum TB size is 32×32 or 64×64, which is configurable in the video encoder 114 and signaled in the bitstream 115. Regardless of the configured maximum DCT-2 transform size, only the coefficients in the top-left 32×32 region of the TB are encoded in the bitstream 115. The significant coefficients outside the top-left 32×32 region of the TB are discarded (or "zeroed out") and not encoded in the bitstream 115. MTS is only available for CUs of size up to 32×32, and only the coefficients in the top-left 16×16 region of the associated luma TB are encoded. The individual TBs of the CU are either transformed or bypassed according to the corresponding transform skip flag 390.

[0104] The forward second transform of module 330 is generally an non-separable transform and is only applied to the residuals of intra-predicted CUs and may still be bypassed. The forward second transform operates on either 16 samples (arranged as the top-left 4×4 sub-block of the first transform coefficients 328) or 48 samples (arranged as three 4×4 sub-blocks of the top-left 8×8 coefficients of the first transform coefficients 328) to generate a set of second transform coefficients. The set of second transform coefficients may be fewer in number than the set of first transform coefficients from which they are derived. Due to the application of the second transform only to sets of coefficients that are adjacent to each other and include the DC coefficient, the second transform is referred to as a "low frequency non-separable transform" (LFNST).

[0105] The residual coefficients 336 are supplied to an entropy encoder 338 for encoding within the bitstream 115. Typically, the residual coefficients of each TB having at least one significant residual coefficient of a TU are scanned according to a scan pattern to generate a list of values in sorted order. The scan pattern generally scans the TB as a sequence of 4×4 “sub-blocks”, provides a regular scan operation at the granularity of 4×4 sets of residual coefficients, and the arrangement of the sub-blocks depends on the size of the TB. The scan within each sub-block and the progression from one sub-block to the next typically follow a backward diagonal scan pattern.

[0106] As described above, the video encoder 114 requires access to a frame representation corresponding to the decoded frame representation seen at the video decoder 134. Accordingly, the residual coefficients 336 are passed to an inverse quantizer 340 to generate inverse quantized residual coefficients 342. The inverse quantized residual coefficients 342 are passed to an inverse second transform module 344 operating according to a second transform index 388 to generate intermediate inverse transform coefficients represented by arrow 346. The intermediate inverse transform coefficients 346 are passed to an inverse first transform module 348 to generate residual samples of the TU represented by arrow 399. The quantized residual coefficients 342 are output by multiplexing 349 as residual samples 350 if transform skip 390 indicates that a transform bypass is being performed. Otherwise, multiplexing 349 outputs the residual samples 399 as residual samples 350.

[0107] The type of inverse transform performed by the inverse second transform module 344 corresponds to the type of forward transform performed by the forward second transform module 330. The type of inverse transform performed by the inverse first transform module 348 corresponds to the type of forward transform performed by the first transform module 326. An adder module 352 adds the residual samples 350 and the PU 320 to generate reconstructed samples of the CU (indicated by arrow 354).

[0108] The reconstructed sample 354 is passed to a reference sample cache 356 and an in-loop filter module 368. The reference sample cache 356 is typically implemented using static RAM on an ASIC (thus avoiding costly off-chip memory accesses) and provides the minimum sample storage necessary to satisfy dependencies for generating intra-PBs for subsequent CUs within a frame. The minimum dependencies include a "line buffer" of samples along the bottom of a row of the CTU for use by the next row of the CTU, and a column buffer set by the height of the CTU. The reference sample cache 356 supplies reference samples (represented by arrow 358) to a reference sample filter 360. The sample filter 360 applies a smoothing operation to generate filtered reference samples (indicated by arrow 362). The filtered reference samples 362 are used by an intra-frame prediction module 364 to generate an intra prediction block of samples represented by arrow 366. For each candidate intra prediction mode, the intra-frame prediction module 364 generates a block of samples 366. The block of samples 366 is generated by the module 364 using techniques such as DC, planar, or angular intra prediction according to the intra prediction mode 387.

[0109] The in-loop filter module 368 applies several filtering stages to the reconstructed sample 354. The filtering stages include a "deblocking filter" (DBF) that applies smoothing aligned to the CU boundary to reduce artifacts resulting from discontinuities. Another filtering stage present in the in-loop filter module 368 is the "adaptive loop filter" (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. Another filtering stage present in the in-loop filter module 368 is the "sample adaptive offset" (SAO) filter. The SAO filter first classifies the reconstructed samples into one or more categories and operates by applying an offset at the sample level according to the assigned category.

[0110] The filtered sample represented by arrow 370 is output from the in-loop filter module 368. The filtered sample 370 is stored in the frame buffer 372. The frame buffer 372 typically has the capacity to store a plurality (e.g., up to 16) of pictures and is thus stored in the memory 206. Since the frame buffer 372 requires a large amount of memory consumption, it is typically not stored using on-chip memory. Therefore, accessing the frame buffer 372 is costly in terms of memory bandwidth. The frame buffer 372 provides a reference frame (represented by arrow 374) to the motion estimation module 376 and the motion compensation module 380.

[0111] The motion estimation module 376 estimates a number of "motion vectors" (shown as 378), each of which is a Cartesian space offset from the position of the present CB, by referring to one block of the reference frames within the frame buffer 372. A filtered block of the reference samples (shown as 382) is generated for each motion vector. The filtered reference samples 382 form additional candidate modes available for potential selection by the mode selector 386. Further, for a given CU, the PU 320 may be formed using one reference block ("unidirectional prediction") or two reference blocks ("bidirectional prediction"). For the selected motion vectors, the motion compensation module 380 generates the PB 320 according to a filtering process that supports sub-pixel accuracy of the motion vectors. Thus, the motion estimation module 376 (which operates with many candidate motion vectors) can perform a simplified filtering process compared to the filtering process of the motion compensation module 380 (which operates only on the selected candidates) to achieve a reduction in computational complexity. When the video encoder 114 selects inter prediction for a CU, the motion vectors 378 are encoded into the bitstream 115.

[0112] The video encoder 114 of FIG. 3 is described with reference to Versatile Video Coding (VVC), but other video coding standards or implementations may also employ the processing stages of modules 310 to 386. Also, the frame data 113 (and the bitstream 115) may be read from (or written to) the memory 206, the hard disk drive 210, the CD-ROM, the Blu-ray Disc (trademark), or other computer-readable storage media. Further, the frame data 113 (and the bitstream 115) may be received from (or transmitted to) an external source such as a server or a radio frequency receiver connected to the communication network 220.

[0113] The video decoder 134 is shown in FIG. 4. The video decoder 134 in FIG. 4 is an example of a versatile video coding (VVC) video decoding pipeline, but other video codecs can also be used to execute the processing stages described herein. As shown in FIG. 4, a bitstream 133 is input to the video decoder 134. The bitstream 133 may be read from the memory 206, the hard disk drive 210, the CD-ROM, the Blu-ray Disc (trademark), or other non-transitory computer-readable storage media. Alternatively, the bitstream 133 may be received from an external source such as a server or a radio frequency receiver connected to the communication network 220. The bitstream 133 includes encoded syntax elements representing the captured frame data to be decoded.

[0114] The bitstream 133 is input to the entropy decoder module 420. The entropy decoder module 420 extracts syntax elements from the bitstream 133 by decoding a sequence of "bins" and passes the values of the syntax elements to other modules of the video decoder 134. The entropy decoder module 420 uses variable-length and fixed-length decoding to decode the SPS, PPS, or slice header arithmetic decoding engine and decode the syntax elements of the slice data as a sequence of one or more bins. Each bin can use one or more "contexts", and the context describes the probability level used to code the "1" and "0" values of the bin. When multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the contexts available for decoding the bin.

[0115] The entropy decoder module 420 applies an arithmetic coding algorithm, such as "context adaptive binary arithmetic coding" (CABAC), for example, to decode syntax elements from the bitstream 133. The decoded syntax elements are used to reconstruct the parameters within the video decoder 134. The parameters include residual coefficients (represented by arrow 424), quantization parameters (not shown), secondary transform index 474, and mode selection information such as intra prediction mode (represented by arrow 458). Also, the mode selection information includes information such as motion vectors and the division of each CTU into one or more CUs. The parameters are typically used, in combination with sample data from previously decoded CBs, to generate PB.

[0116] The residual coefficient 424 is passed to the inverse quantization module 428. The inverse quantization module 428 performs inverse quantization (or "scaling") on the residual coefficient 424 (i.e., the primary transform coefficient region) to create a reconstructed transform coefficient represented by arrow 432 according to the quantization parameter. The reconstructed transform coefficient 432 is passed to the inverse secondary transform module 436. The inverse secondary transform module 436 performs either applying the secondary transform or bypassing the operation (bypassing) according to the secondary transform type 474 decoded from the bitstream 113 by the entropy decoder 420 according to the method described with reference to FIGS. 15 and 16. The inverse secondary transform module 436 generates a reconstructed transform coefficient 440 (i.e., the primary transform region coefficient).

[0117] The reconfigured transform coefficient 440 is passed to the inverse primary transform module 444. Module 444 inversely transforms the coefficient 440 from the frequency domain to the spatial domain according to the primary transform type 476 (or "mts_idx") decoded from the bitstream 133 by the entropy decoder 420. The result of the operation of module 444 is a block of residual samples represented by arrow 499. If the transform skip flag 478 for a given TB of the CU indicates bypassing of the transform, multiplexing 449 outputs the reconfigured transform coefficient 432 as residual samples 488 to the summation module 450. Otherwise, multiplexing 449 outputs the residual samples 499 as residual samples 488. The residual samples 448 are of the same size as the corresponding CB. The residual samples 448 are supplied to the summation module 450. In the summation module 450, the residual samples 448 are added to the decoded PB (represented as 452) to generate a block of reconstructed samples represented by arrow 456. The reconstructed samples 456 are supplied to the reconstructed sample cache 460 and the in-loop filtering module 488. The in-loop filtering module 488 generates a reconstructed block of frame samples represented by 492. The frame samples 492 are written into the frame buffer 496, from which the frame data 135 is output later.

[0118] The reconstructed sample cache 460 operates in a similar manner to the reconstructed sample cache 356 of the video encoder 114. The reconstructed sample cache 460 provides storage for a plurality of reconstructed samples necessary for intra prediction of subsequent CBs without relying on access to the memory 206 (e.g., by using the data 232 which is a typical on-chip memory instead). The reference samples represented by the arrow 464 are obtained from the reconstructed sample cache 460 and supplied to the reference sample filter 468 to generate the filtered reference samples indicated by the arrow 472. The filtered reference samples 472 are supplied to the intra-frame prediction module 476. The module 476 generates a block of intra prediction samples indicated by the arrow 480 according to the intra prediction mode parameter 458 signaled in the bitstream 133, and is decoded by the entropy decoder 420.

[0119] When the prediction mode of the CB is instructed to use intra prediction in the bitstream 133, the intra predicted samples 480 form the decoded PB 452 via the multiplexing module 484. Intra prediction generates a block in one color component, i.e., a block in one color component derived using the predicted block (PB) of samples, i.e., the "neighboring samples" in the same color component. The neighboring samples are samples adjacent to the current block and have already been reconstructed by preceding in the block decoding order. When luma blocks and chroma blocks are arranged, the luma blocks and chroma blocks can use different intra prediction modes. However, two chroma CBs share the same intra prediction mode.

[0120] If it is shown that the prediction mode of the CB is an inter prediction in the bitstream 133, the motion compensation module 434 selects a block 498 of samples from the frame buffer 496 and uses the motion vector and the reference frame index (decoded from the bitstream 133 by the entropy decoder 420) to generate a block of inter-predicted samples represented as 438 for filtering. The block 498 of samples is obtained from the previously decoded frames stored in the frame buffer 496. For bidirectional prediction, two blocks of samples are generated and blended together to generate samples for the decoded PB452. Filtered block data 492 from the in-loop filtering module 488 is input to the frame buffer 496. Similar to the in-loop filtering module 368 of the video encoder 114, the in-loop filtering module 488 applies any of the filtering operations of DBF, ALF, and SAO. Generally, the motion vector is applied to both the luma channel and the chroma channel, but the filtering processes for sub-sample interpolation in the luma channel and the chroma channel are different.

[0121] Figure 5 is a schematic block diagram showing a collection 500 of available partitions or partitions that divide one region into one or more sub-regions at each node of the encoding tree structure for multi-purpose video encoding. The partitions shown in the collection 500 are available to the block partitioner 310 of the encoder 114 to divide each CTU into one or more CUs or CBs according to the encoding tree, as determined by Lagrange optimization, as described with reference to FIG. 3.

[0122] Collection 500 only shows that a square region is divided into other, perhaps non-square, sub-regions. However, collection 500 shows the possibility that a parent node of the coding tree is divided into child nodes of the coding tree, and it should be understood that it is not necessary for the parent node to correspond to a square region. If the region included is non-square, the size of the blocks resulting from the division is scaled according to the aspect ratio of the included blocks. If a region is not further divided, i.e., at a leaf node of the coding tree, the CU occupies that region.

[0123] The process of subdividing a region into sub-regions ends when the resulting sub-regions reach the minimum CU size (generally 4×4 luma samples). In addition to being restricted such that a CU prohibits a block region smaller than a predetermined minimum size (e.g., 16 samples), the minimum value of the width or height is restricted to be 4. Additionally, it is also possible to set both the width and height, or the minimum values of both the width and height. The subdivision process may end before the deepest level of decomposition, resulting in a CU larger than the minimum CU size. It is also possible that no division is performed and a single CU occupies the entire CTU. A single CU that occupies the entire CTU is the largest available coding unit size. The use of a sub-sampling chroma format such as 4:2:0 allows the arrangement of the video encoder 114 and the video decoder 134 to end the division of the chroma channel region earlier than the luma channel, including in the case of a shared coding tree that defines the block structure of the luma and chroma channels. When separate coding trees are used for luma and chroma, the restrictions on the available division operations ensure that the minimum chroma CU region is 16 samples even if such a CU is co-located with a larger luma region, e.g., 64 luma samples.

[0124] There is a CU in the leaf node of the symbol tree. For example, leaf node 510 contains one CU. In the non-leaf nodes of the symbol tree, there are splits into two or more further nodes, each of which can be a leaf node forming one CU or a non-leaf node containing a further split into smaller regions. In each leaf node of the symbol tree, there is one CB for each color channel of the symbol tree. A split that ends at the same depth for both the luma and chroma of the shared tree results in one CU having three conjugate CBs.

[0125] Quad-tree split 512 divides the enclosing region into four regions of equal size, as shown in FIG. 5. Compared to HEVC, Versatile Video Coding (VVC) achieves additional flexibility with additional splits including horizontal 2-split 514 and vertical 2-split 516. Each of splits 514 and 516 divides the enclosed region into two regions of the same size. The splits are performed along a horizontal boundary (514) or a vertical boundary (516) within the enclosing block.

[0126] In Versatile Video Coding, further flexibility is obtained by adding a 3-split horizontal split 518 and a 3-split vertical split 520. The 3-splits 518 and 520 divide the block into three regions bounded either horizontally (518) or vertically (520) along 1 / 4 and 3 / 4 of the width or height of the containing region. The combination of quadtree, binary tree, and ternary tree is called "QTBTTT". The root of the tree contains zero or more quadtree splits (the "QT" section of the tree). When the QT section ends, zero or more 2-splits or 3-splits occur (the "multi-tree" or "MT" section of the tree), and finally end with a CB or a CU at the leaf node of the tree. If the tree describes all color channels, the leaf node of the tree is a CU. If the tree describes the luma channel or the chroma channel, the leaf node of the tree is a CB.

[0127] Compared with HEVC which only supports quad-trees and thus only supports square blocks, QTBTTT offers more possible CU sizes, especially considering the possibility of recursively applying binary and / or ternary tree partitions. When only quadtree partitioning is available, as the depth of the coding tree increases, it is equivalent to the CU size being reduced to 1 / 4 of the parent region. In VVC, since binary and ternary tree partitions are possible, the depth of the coding tree no longer directly corresponds to the CU area. By restricting the partitioning options to exclude partitions where the width or height of the block is less than 4 samples or not a multiple of 4 samples, the possibility of non-square block sizes can be reduced. By restricting the partitioning options to exclude partitions where the width or height of the block is less than 4 samples or not a multiple of 4 samples, the possibility of an unusual (non-square) block size can be reduced.

[0128] Figure 6 is a schematic flow diagram showing the data flow 600 of the QTBTTT (or "coding tree") structure used in versatile video coding. The QTBTTT structure is used for each CTU to define the partitioning of the CTU into one or more ECUs. The QTBTTT structure of each CTU is determined by the block partitioner 310 within the video coder 114 and is encoded into the bitstream 115 or decoded from the bitstream 133 by the entropy decoder 420 within the video decoder 134. The data flow 600 further features the acceptable combinations available to the block partitioner 310 for partitioning the CTU into one or more CUs according to the partitioning shown in FIG. 5.

[0129] Starting from the top - most level of the hierarchy, i.e., the CTU, first zero or more quad - tree partitions are performed. Specifically, the quad - tree (QT) partition decision 610 is made by the block partitioner 310. The decision at 610 to return a "1" symbol indicates a decision to split the current node into four sub - nodes according to the quad - tree partition 512. As a result, as at 620, four new nodes are generated, and for each new node, a recursive call is made to the QT partition decision 610. Each new node is considered in raster (or Z - scan) order. Alternatively, if the QT partition decision 610 indicates no further splitting (returns a "0" symbol), the quad - tree partition stops, and subsequently, the multi - tree (MT) partition is considered.

[0130] First, the MT partition decision 612 is made by the block partitioner 310. At 612, a decision on whether to perform an MT partition is indicated. Returning a "0" symbol at decision 612 indicates not to perform a further split of the node into sub - nodes. If no further split of the node is performed, that node is a leaf node of the coding tree and corresponds to a CU. The leaf node is output at 622. Alternatively, if the MT partition 612 indicates a decision to perform an MT partition (returns a "1" symbol), the block partitioner 310 proceeds to the direction decision 614.

[0131] The direction decision 614 indicates the direction of the MT partition as either horizontal ("H" or "0") or vertical ("V" or "1"). If the block partitioner 310 returns a "0" indicating a horizontal direction at decision 614, it proceeds to decision 616. If the block partitioner 310 returns a "1" indicating a vertical direction at decision 614, it proceeds to decision 618.

[0132] In each of the decisions 616 and 618, the number of MT partitions is indicated as either two (two partitions or "BT" nodes) or three (three partitions or "TT") in the BT / TT partition. That is, when the indication direction from 614 is horizontal, the BT / TT partition decision 616 is made by the block partitioner 310, and when the indication direction from 614 is vertical, the BT / TT partition decision 618 is made by the block partitioner 310.

[0133] The BT / TT partition decision 616 indicates whether the horizontal partition is a two-partition 514 indicated by returning "0" or a three-partition 518 indicated by returning "1". When the BT / TT partition decision 616 indicates a two-partition, in the HBT CTU node generation step 625, two nodes are generated by the block partitioner 310 according to the horizontal two-partition 514. When the BT / TT partition 616 indicates a three-partition, in the generate HTT_CTU node step 626, three nodes are generated by the block partitioner 310 according to the horizontal three-partition 518.

[0134] The BT / TT partition decision 618 indicates whether the vertical partition is a two-partition 516 indicated by returning "0" or a three-partition 520 indicated by returning "1". When the BT / TT partition 618 indicates a two-partition, in the VBT_CTU node generation step 627, two nodes are generated by the block partitioner 310 according to the vertical two-partition 516. When the BT / TT partition 618 indicates a three-partition, in the generate VTT_CTU node step 628, three nodes are generated by the block partitioner 310 according to the vertical three-partition 520. For each node obtained from steps 625 to 628, according to the direction 614, the recursion of the data flow 600 returning to the MT partition decision 612 is applied in the order from left to right or from top to bottom. As a result, the binary tree and ternary tree partitions can be applied to generate CUs of various sizes.

[0135] FIG. 7A and FIG. 7B provide an example 700 in which CTU 710 is divided into a number of CUs or CBs. In FIG. 7A, an exemplary CU 712 is shown. FIG. 7A shows the spatial arrangement of the CUs in CTU 710. The exemplary division 700 is also shown as an encoding tree 720 in FIG. 7B.

[0136] At each non-leaf node of CTU 710 in FIG. 7A, such as nodes 714, 716, and 718, the included nodes (which may be further divided or may be CUs) are scanned or traversed in "Z-order" to create a list of nodes represented as columns of the encoding tree 720. In the case of quadtree division, the Z-order scan is performed in the order from top left to right, followed by from bottom left to right. In the case of horizontal and vertical division, the Z-order scan (traversal) is simplified to a scan from top to bottom and a scan from left to right, respectively. The encoding tree 720 in FIG. 7B lists all the nodes and CUs arranged according to the Z-order scan of the encoding tree. Each division generates a list of two, three, or four new nodes at the next level of the tree until a leaf node (CU) is reached.

[0137] After the image is decomposed into CTUs by the block partitioner 310 and further into CUs, and each residual block (324) is generated using the CUs as described with reference to FIG. 3, the residual blocks are subject to forward transformation and quantization by the video encoder 114. The resulting TB 336 is then scanned to form a sequential list of residual coefficients as part of the operation of the entropy encoding module 338. Equivalent processing is performed in the video decoder 134 to obtain the TB from the bitstream 133.

[0138] Figures 8A, 8B, 8C, and 8D show examples of forward transformation and non-separable inverse secondary transformation performed according to transform blocks (TBs) of different sizes. Figure 8A is a diagram showing a series of relationships 800 between the primary transform coefficients 802 and the secondary transform coefficients 804 for a 4×4 TB size. The primary transform coefficients 802 are composed of 4×4 coefficients, and the secondary transform coefficients 804 are composed of eight coefficients. The eight secondary transform coefficients are arranged in a pattern 806. The pattern 806 corresponds to eight positions adjacent in a backward diagonal scan of the TB and including the DC (upper left) position. The remaining eight positions of the backward diagonal scan shown in Figure 8A are not input by performing the forward secondary transformation and thus remain at zero values. Therefore, the forward non-separable secondary transformation 810 for a 4×4 TB receives 16 primary transform coefficients and generates eight secondary transform coefficients as outputs. Therefore, the forward secondary transformation 810 for a 4×4 TB can be represented by an 8×16 matrix of weights. Similarly, the inverse secondary transformation 812 can be represented by a 16×8 matrix of weights.

[0139] Figure 8B shows a set of relationships 818 between the primary transform coefficients and the secondary transform coefficients for 4×N and N×4 TB sizes (N is greater than 4), and in both cases, the upper left 4×4 sub-block 820 of the primary coefficients is associated with the upper left 4×4 sub-block of the secondary transform coefficients 824. In the video encoder 114, the forward non-separable secondary transformation 830 takes 16 primary transform coefficients and generates 16 secondary transform coefficients as outputs. The remaining primary transform coefficients 822 are not input by the forward secondary transformation and thus remain at zero values. After the forward non-separable secondary transformation 830 is performed, the coefficient positions 826 are associated with the coefficients 822, are not input, and thus remain at zero values.

[0140] The forward secondary transformation 830 for a 4×N or N×4 TB can be represented by a 16×16 matrix of weights. The matrix representing the forward secondary transformation 830 is defined as A. Similarly, the corresponding inverse secondary transformation 832 can be represented by a 16×16 matrix of weights. The matrix representing the inverse secondary transformation 832 is defined as B.

[0141] By reusing a part of A for the forward quadratic transformation 810 and the inverse quadratic transformation 812 for 4×4TB, the storage requirement for the non-separable transformation kernel is further reduced. The first 8 rows of A are used for the forward quadratic transformation 810, and the transpose of the first 8 rows of A is used for the inverse quadratic transformation 812.

[0142] FIG. 8C shows the relationship 855 between the primary transformation coefficients 840 and the secondary transformation coefficients 842 for an 8×8 sized TB. The primary transformation coefficients 840 are composed of 8×8 coefficients, and the secondary transformation coefficients 842 are composed of 8 transformation coefficients. The 8 secondary transformation coefficients 842 are arranged in a pattern corresponding to 8 consecutive positions in the backward diagonal scan of the TB, and the 8 consecutive positions include the DC (upper left) coefficient of the TB. Since all the remaining secondary transformation coefficients of the TB are zero, there is no need to scan them. The forward non-separable quadratic transformation 850 of the 8×8TB takes 48 primary transformation coefficients corresponding to 3 4×4 sub-blocks as input and generates 8 secondary transformation coefficients. The forward quadratic transformation 850 for an 8x8 TB can be represented by an 8×48 matrix of weights. Also, the corresponding inverse quadratic transformation 852 of the 8×8TB can be represented by a 48×8 matrix of weights.

[0143] FIG. 8D is a diagram showing the relationship 875 between the primary transformation coefficients 860 and the secondary transformation coefficients 862 for a TB of size 8×8 or larger. The upper left 8×8 block (arranged as 4 4×4 sub-blocks) of the primary coefficient 860 is associated with the upper left 4×4 sub-block of the secondary transformation coefficients 862. In the video encoder 114, the forward non-separable quadratic transformation 870 calculates 48 primary transformation coefficients to generate 16 secondary transformation coefficients. The remaining primary transformation coefficients 864 are set to zero. The secondary transformation coefficient positions 866 outside the upper left 4×4 sub-block of the secondary transformation coefficients 862 are not input and remain zero.

[0144] The forward secondary transform 870 of the TB with a size larger than 8×8 can be represented by a 16×48 matrix of weights. The matrix representing the forward secondary transform 870 is defined as F. Similarly, the corresponding inverse secondary transform 832 can be represented by a 48×16 matrix of weights. The matrix representing the inverse secondary transform 872 is defined as G. As described above with reference to matrices A, B, F, desirably, it has the property of orthogonality. The property of orthogonality means that G = F T only, and F needs to be stored in the video encoder 114 and the video decoder 134. An orthogonal matrix can be expressed as a matrix whose rows have the property of orthogonality.

[0145] By reusing a part of F, the storage requirement of the non-separable transform kernel is further reduced. F is for the forward secondary transform 850 and the inverse secondary transform 852 of the 8×8 TB. The first 8 rows of it are the transpose of the first 8 rows of F and are used for the forward secondary transform 810. F is used for the inverse secondary transform 812.

[0146] The non-separable secondary transform can sparsify the two-dimensional features of the residual signal such as angular features, so coding improvement can be achieved compared to the case where only the separable primary transform is used. Since the angular features in the residual signal may depend on the type of the selected intra prediction mode 387, it is advantageous that the non-separable secondary transform matrix is adaptively selected according to the intra prediction mode. As described above, the intra prediction mode is composed of the "intra DC" mode, the "intra plane" mode, the "intra angular" mode, and the "matrix intra prediction" mode. The intra prediction mode parameter 458 takes a value of 0 when intra DC prediction is used. The intra prediction mode parameter 458 takes a value of 1 when intra plane prediction is used. The intra prediction mode parameter 458 takes a value between 2 and 66 when intra angular prediction on the square TB is used.

[0147] FIG. 9 shows a set 900 of transform blocks available in the Versatile Video Coding (VVC) standard. FIG. 9 also shows the application of a secondary transform to a subset of the residual coefficients from the transform blocks of set 900. FIG. 9 shows a plurality of TUs with widths and heights in the range of 4 to 32. However, TUs with a width and / or height of 64 are possible but not shown for ease of reference.

[0148] For a set of 4×4 coefficients, a 16-point secondary transform 952 (shown in dark shading) is applied. The 16-point secondary transform 952 is applied to TUs with a width or height of 4, e.g., 4×4 TU 910, 8×4 TU 912, 16×4 TU 914, 32×4 TU 916, 4×8 TU 920, 4×16 TU 930, and 4×32 TU 940. The 16-point secondary transform 952 is also applied to TUs of size 4×64 and 64×4 (not shown in FIG. 9). For TUs with a width or height of 4 and primary coefficients of 16 or more, the 16-point secondary transform is applied only to the upper left 4×4 sub-block of the TU, and the other sub-blocks need to have coefficients of 0 value in order to apply the secondary transform. Generally, when applying the 16-point secondary transform, 8 or 16 secondary transform coefficients are generated as described with reference to FIGS. 8 to 8D. The secondary transform coefficients are packed into the TU for encoding the upper left sub-block of the TU.

[0149] When the conversion size has a width and height greater than 4, as shown in FIG. 9, a 48-point quadratic transform 950 (shown in light hatching) for application to three 4×4 sub-blocks of the residual coefficients in the upper left 8×8 region of the conversion block is available. The 48-point quadratic transform 950 is applied to the 8×8 conversion block 922, 16×8 conversion block 924, 32×8 conversion block 926, 8×16 conversion block 932, 16×16 conversion block 934, 32×16 conversion block 936, 8×32 conversion block 942, 16×32 conversion block 944, 32×32 conversion block 946, in each case, in the regions shown by the light shading and the dashed lines. Also, the 48-point quadratic transform 950 is applicable to TBs (not shown) of sizes 8×64, 16×64, 32×64, 64×64, 64×32, 64×16, 64×8. When applying the 48-point quadratic transform kernel, generally, less than 48 quadratic transform coefficients will be generated. For example, as described with reference to FIGS. 8B - 8D, 8 or 16 quadratic transform coefficients can be generated. Primary transform coefficients that are not subject to the quadratic transform (the "primary-only coefficients"), such as the coefficient 966 of TB 934, are required to be zero values for the quadratic transform to be applied. After applying the 48-point quadratic transform 950 in the forward direction, the region that may contain significant coefficients is reduced from 48 coefficients to 16 coefficients, and the number of coefficient positions that may contain significant coefficients is further reduced. In the inverse quadratic transform, the decoded significant coefficients are transformed to generate coefficients that may be significant in the region subject to the primary inverse transform. When one or more sub-blocks are reduced to a set of 16 quadratic transform coefficients by the quadratic transform, only the upper left 4×4 sub-block can contain significant coefficients. The position of the last significant coefficient at any coefficient position where the quadratic transform coefficients may be stored indicates that the quadratic transform has been applied or that only the primary transform has been applied.

[0150] When the last significant coefficient position indicates a secondary transform coefficient position within the TB, the signaled secondary transform index (i.e., 388 or 474) is required to distinguish whether to apply the secondary transform kernel or bypass the secondary transform. The application of the secondary transform to TBs of various sizes in FIG. 9 has been described from the perspective of the video encoder 114, while the corresponding inverse process is performed in the video decoder 134. The video decoder 134 first decodes the position of the last significant coefficient. If the decoded last significant coefficient position indicates the applicability of the secondary transform, the secondary transform index 474 is decoded to determine whether to apply or bypass the inverse secondary transform.

[0151] FIG. 10 shows the syntax structure 0100 of a bitstream 1001 having a plurality of slices. Each of the slices includes a plurality of coding units. The bitstream 1001 may be generated by the video encoder 114, for example, as the bitstream 115, or may be parsed by the video decoder 134, for example, as the bitstream 133. The bitstream 1001 is divided into parts such as, for example, network abstraction layer (NAL) units, and the separation is achieved by preceding each NAL unit with an NAL unit header such as 1008. The sequence parameter set (SPS) 1010 defines sequence-level parameters such as the profile (set of tools), chroma format, sample bit depth, and frame resolution used for encoding and decoding the bitstream. The parameters are also included in the set 1010 that restricts the application of different types of partitions in the coding tree of each CTU.

[0152] The picture parameter set (PPS) 1012 defines a set of parameters applied to zero or more frames. The picture header (PH) 1015 defines the parameters applied to the current frame. The parameters of PH 1015 may include a list of CU chroma QP offsets, one of which is applied at the CU level and can be used to derive the quantization parameter for use by the chroma blocks from the quantization parameter of the co-located luma CB.

[0153] A slice sequence forming a picture with the picture header 1015 is known as an AU (access unit), such as, for example, AU0_0114. AU0_1014 includes three slices such as slices 0 to 2, and slice 1 is denoted as 1016. Similar to other slices, slice 1 (1016) includes a slice header 0118 and slice data 1020.

[0154] FIG. 11 shows a syntax structure 1100 of slice data (such as slice data 1104 corresponding to 1020) of a bitstream 1001 (such as 115 or 133) by a shared encoding tree of a luma encoding unit and a chroma encoding unit of an encoding tree unit such as CTU 1110. CTU 1110 includes one or more CUs. An example is labeled as CU 1114. CU 1114 includes a signaled prediction mode 1116 followed by a transform tree 1118. When the size of CU 1114 does not exceed the maximum transform size (either 32×32 or 64×64 for the luma channel), the transform tree 1118 includes one transform unit shown as TU 1124. When a 4:2:0 chroma format is used, the corresponding maximum chroma transform size is half of the luma maximum transform size in each direction. That is, when the maximum luma transform size is 32×32 or 64×64, the maximum chroma transform sizes are 16×16 or 32×32, respectively. In the case of a 4:4:4 chroma format, the maximum chroma transform size is the same as the luma maximum transform size. In the case of a 4:2:2 chroma format, the maximum chroma transform size is half in the horizontal direction and the same as the luma maximum transform size in the vertical direction. That is, when the luma maximum transform sizes are 32×32 and 64×64, the maximum chroma transform sizes are 16×32 and 32×64, respectively.

[0155] When the prediction mode 1116 indicates the use of intra prediction for CU 1114, a luma intra prediction mode and a chroma intra prediction mode are specified. Also, for the luma CB of CU 1114, according to the MTS index 1122, the primary transform type is signaled to be either (i) DCT-2 in the horizontal and vertical directions, (ii) transform skip in the horizontal and vertical directions, or (iii) a combination of DST-7 and DCT-8 in the horizontal and vertical directions. When the signaled luma transform type is DCT-2 horizontal and vertical (option (i)), an additional luma secondary transform index 1120, also known as the "low-frequency non-separable transform" (LFNST) index, is signaled in the bitstream under the conditions described with reference to FIGS. 8A - 8D and FIGS. 13 - 16.

[0156] By using a common symbolized tree, TU1124 will include TBs for each color channel shown as luma TB_Y_1128, first chroma TB_Cb_1132, and second chroma TB_Cr_1136. Whether each TB exists depends on a corresponding "coded block flag" (CBF), i.e., one of coded block flags 1123. When a TB exists, the corresponding CBF is equal to 1 and at least one residual coefficient within the TB is non-zero. When a TB does not exist, the corresponding CBF is equal to zero and all residual coefficients within the TB are zero. Luma TB1128, first chroma TB1134, and second chroma TB1136 may use transform skip as signaled by transform skip flags 1126, 1130, and 1134, respectively. A coding mode in which a single chroma TB is transmitted to specify chroma residuals for both Cb and Cr channels is known as the "joint CbCr" coding mode and is available. When the joint CbCr coding mode is active, a single chroma TB is coded.

[0157] Regardless of the color channel, each encoded TB includes one or more residual coefficients following the final position. For example, luma TB 1128 includes the final position 1140 and the residual coefficient 1144. The final position 1140 indicates the position of the last significant residual coefficient in the TB when considering the coefficients of the diagonal scan pattern used to serialize the array of coefficients of the TB in the forward direction (i.e., ahead of the DC coefficient). The two TBs 1132 and 1136 for the chroma channel each have corresponding final position syntax elements used in the same way as described for luma TB 1128. If the final positions of the TBs for the CU, i.e., 1128, 1132, and 1136 respectively, indicate that only the coefficients of the secondary transform region are significant for each TB of the CU and all the remaining coefficients that would only undergo a primary transform are zero, the secondary transform index 1120 may be signaled to specify whether to apply the secondary transform. Further conditioning regarding the signaling of the secondary transform index 1120 is described with reference to FIGS. 14 and 16.

[0158] When the secondary transformation is applied, the secondary transformation index 1120 indicates which kernel is selected. Generally, in the "candidate set" of kernels, two kernels are available. Generally, there are four candidate sets, and one candidate set is selected using the intra prediction mode of the block. The luma intra prediction mode is used to select the candidate set of the luma block, and the chroma intra prediction mode is used to select the candidate sets of the two chroma blocks. As described with reference to FIGS. 8A to 8D, the selected kernel also depends on the TB size, and different kernels are used for 4×4, 4×N / N×4, and other sized TBs. When the 4:2:0 chroma format is used, the chroma TB is generally half the width and height of the corresponding luma TB. As a result, when a luma TB with a width or height of 8 is used, different selected kernels will occur for the chroma blocks. For luma blocks of sizes 4×4, 4×8, and 8×4, the one-to-one correspondence between the luma block and the chroma block in the shared coding tree is changed so that there are no small-sized chroma blocks such as 2×2, 2×4, and 4×2.

[0159] The secondary transformation index 1120 indicates, for example, the following. Index value 0 (not applied), 1 (apply the first kernel of the candidate set), 2 (apply the second kernel of the candidate set). For chroma, the selected secondary transformation kernel of the candidate set derived considering the chroma TB size and the chroma intra prediction mode is applied to each chroma channel. Therefore, the residuals of the Cb block 1224 and the Cr block 1226 need to contain only significant coefficients at the positions where the secondary transformation is performed, as described with reference to FIGS. 8A to 8D. When joint CbCr coding is used, the resulting Cb and Cr residuals contain only significant coefficients at the positions corresponding to the significant coefficients within the joint coding TB. Therefore, the requirement of containing significant coefficients at the positions to be subjected to the secondary transformation is applicable only to the single-coded chroma TB.

[0160] FIG. 12 shows a syntax structure 1200 of slice data 1204 (e.g., 1020) of a bitstream (e.g., 115, 133) in which a luma encoding unit and a chroma encoding unit of an encoding tree unit have separate encoding trees. The separate encoding trees are available for "I-slices". The slice data 1204 includes one or more CTUs such as CTU 1210. CTU 1210 generally has a luma sample size of 128×128 and starts with a shared tree including a single quad tree split common to luma and chroma. In each of the resulting 64×64 nodes, separate encoding trees start for luma and chroma. FIG. 12 shows an example of a node 1214. The node 1214 has a luma node 1214a and a chroma node 1214b. The luma tree starts from the luma node 1214a, and the chroma tree starts from the chroma node 1214b. Since the trees following the node 1214a and the node 1214b are independent between luma and chroma, different splitting options are possible to generate the resulting CUs. The luma CU 1220 belongs to the luma encoding tree and includes a luma prediction mode 1221, a luma transform tree 1222, and a secondary transform index 1224. The luma transform tree 1222 includes TUs 1230. Since the luma encoding tree encodes only samples of the luma channel, the TU 1230 includes a luma TB 1234, and the luma transform skip flag 1232 indicates whether the luma residual should be transformed. The luma TB 1234 includes a final position 1236 and residual coefficients 1238.

[0161] Chroma CU1250 belongs to the chroma coding tree and includes a chroma prediction mode 1251, a chroma transform tree 1252, and a secondary transform index 1254. The chroma transform tree 1252 includes TUs 1260. Since the chroma tree includes chroma blocks, TU1260 includes Cb_TB1264 and Cr_TB1268. The application of transform bypass for Cb_TB1264 and Cr_CB1268 is signaled by a Cb transform skip flag 1262 and a Cr transform skip flag 1266, respectively. Each TB includes a final position and residual coefficients. For example, a final position 1270 and residual coefficients 1272 are associated with Cb_TB1264. The signaling of the secondary transform index 1 of 254 for the chroma TBs of the chroma tree is described with reference to FIGS. 14 and 16.

[0162] FIG. 17 is a diagram showing a 32×32 TB1700. It is shown that a conventional scan pattern 1710 is applied to TB1700. The scan pattern 1710 progresses diagonally backward through TB1700, starting from the last significant coefficient position and proceeding toward the DC (top left) coefficient position. This progression divides TB1700 into 4×4 sub-blocks. Each sub-block is scanned diagonally backward inside, as shown in some sub-blocks of TB1700, for example, sub-block 1750. Other sub-blocks are scanned in the same way. However, in FIG. 17, for ease of reference, a limited number of sub-blocks are shown in full scan. The progression from one 4×4 sub-block to the next also follows the diagonally backward scan that spans the entire TB1700.

[0163] When using MTS, only the coefficients of the upper left 16×16 portion 1740 of the TB1700 may be important. The upper left 16×16 portion forms, or is within the range of, a threshold orthogonal position (in this example, (15, 15)) where MTS can be applied. If the last significant coefficient is outside the threshold orthogonal position at any point in either the X or Y coordinate, MTS cannot be applied. That is, if either the X coordinate or the Y coordinate of the position of the last significant coefficient exceeds 15, MTS cannot be applied and DCT-2 is applied (or the transformation is skipped). The position of the last significant coefficient is represented in orthogonal coordinates relative to the position of the DC coefficient within the TB1700. For example, the position 1730 of the last significant coefficient is 15, 15. The scan pattern 1710 starting from the position 1730 and proceeding towards the DC coefficient results in scan sub-blocks 1720 and 1721 (identified by hatching) that are zeroed out in the video encoder 114 when MTS is applied and not used by the video decoder 134. Since the video decoder 134 needs to decode the residual coefficients of the sub-blocks 1720, 1721 as they are included in the scan, the residual coefficients of the decoded sub-blocks 1720, 1721 are not used when MTS is applied. At least, the residual coefficients of the sub-block 1720 may be required to be zero values for MTS to be applied, reducing the associated coding cost and preventing the bitstream from encoding the significant residual coefficients of the sub-block when MTS is applied. That is, the parsing of the "mts_idx" syntax element can also be conditioned on the last significant position being within the portion 1740 and the sub-blocks 1720 and 1721 containing only zero-valued residual coefficients.

[0164] Figure 18 is a diagram showing the scan pattern 1810 of a 32×32 TB1800 using the described arrangement. The scan pattern 1810 groups 4×4 sub-blocks into several "collections" such as the collection 1840.

[0165] In the context of the present disclosure, with respect to a scan pattern, a collection provides a non-overlapping set of sub-blocks that (i) form an area or region of a size applicable to MTS, or (ii) form an area or region surrounding an area applicable to MTS. The scan pattern traverses a transform block by proceeding through several non-overlapping collections of sub-blocks of residual coefficients, and after completing the scan of the current collection, proceeds to the next collection from the current collection.

[0166] In the example of FIG. 18, each collection is a two-dimensional array of 4x4 sub-blocks having a width and height of up to four sub-blocks (option (i) of the collection). Collection 1840 corresponds to the area of potentially significant coefficients when MTS is being used, i.e., the 16×16 area of TB1800. Scan pattern 1810 proceeds from one collection to the next without re-entry, i.e., when all the residual coefficients within a collection have been scanned, scan pattern 1810 proceeds to the next collection. Scan 1810 effectively completes the scan pattern of the current collection completely before proceeding to the scan of the next collection. The collections are non-overlapping, and each residual coefficient position is scanned once, starting from the final position and proceeding towards the DC (top left) coefficient position.

[0167] Similar to the scan pattern 1710, the scan pattern 1810 also divides TU1800 into 4×4 sub-blocks. For a monotonic progression from one collection to the next, when the scan reaches the upper left collection 1840, no further scan of the residual coefficients outside the collection 1840 occurs. In particular, when the final position is within the collection 1840, for example, at the final position 1830 at the position of 15,15, all residual coefficients outside the collection 1840 are not significant. The fact that the residual coefficients outside 1840 are zero is consistent with the zeroing out performed in the video encoder 114 when MTS is used. Therefore, the video decoder 134 only needs to confirm that the final position is within the collection 1840 in order to enable the parsing of the mts_idx syntax element (1122 when the CU belongs to a single coding tree, 1226 when the CU belongs to the luma branch of a separate coding tree). The use of the scan pattern 1810 eliminates the need to ensure that any residual coefficients outside the collection 1840 are zero values. Whether it is a coefficient outside the collection 1840 is already clear thanks to the scan pattern 1810 having a collection size aligned with the MTS transform coefficient region. By dividing TB1800 into a set of collections, each of the same size, the scan pattern 1810 can also enable a reduction in memory consumption compared to the scan pattern 1710. Since the scan over TB1800 can be composed of scans over one collection, memory reduction is possible. For TBs of sizes 16×32 and 32×16, two collections can be used in the same approach as for 16×16-sized collections. For a 32×8-sized TB, it is possible to divide it into a collection restricted to a size of 16×8 due to the relationship of the TB size. When dividing a 32×8 TB into collections, it results in the same scan pattern as regularly progressing diagonally over an 8×2 array of 4×4 sub-blocks that make up the 32×8 TB. Therefore, in the region of 8×16 coefficients subject to MTS transform of the 32×8 TB, by confirming that the final position is within the left half of the 32×8 TB, the characteristics of significant coefficients are satisfied.

[0168] Figure 19 shows the TB1900 of size 8×32. TB1900 can be divided into collections. In the example of Figure 19, there are cases where the collection size is restricted to 8×16 due to the TB size, such as collection 1940. The division of the 8×32 TB1900 into collections results in a different sub-block order compared to the normal diagonal progression on the 2×8 array of 4×4 sub-blocks that make up the 8×32 TB (as shown in Figure 18 for example). By using an 8×16 collection size, when the last significant coefficient position is within collection 1940, the significant coefficients are only possible in the MTS transform coefficient region, and it is guaranteed to be the last significant position 1930 at, for example, 7,15.

[0169] The scan patterns of Figures 18 and 19 scan the residual coefficients of each sub-block diagonally backward. In the examples of Figures 18 and 19, the sub-blocks of each collection are scanned diagonally backward. The scanning between collections is performed in the diagonally backward direction in Figures 18 and 19.

[0170] FIG. 20 is a diagram showing an alternative scan order 2010 of a 32×32 TB2000. The scan order (scan pattern) 2010 is divided into portions 2010a to 2010f. The scan orders 2010 from 2010a to 2010e relate to option (ii) regarding collection, an area surrounding an area or region applicable to MTS, or a set of sub-blocks forming an area. The scan pattern 2010f relates to a collection covering an area 2040 forming an area applicable to MTS. The scan orders 2010a to 2010f are defined such that a backward diagonal progression from one sub-block to the next occurs across the TB2000 excluding the area 2040, and then a scan of the backward diagonal progression is used for scanning. The area 2040 corresponds to an MTS conversion coefficient area. When the TB2000 is divided into a scan on sub-blocks outside the MTS conversion coefficient area and a scan on sub-blocks within the MTS conversion coefficient area that follows, a progression on the sub-blocks is brought about as shown in 2010a, 2010b, 2010c, 2010d, 2010e, and 2010f. The scan pattern 2010 identifies two collections, a collection defined by the collections defined by 2010a to 2010e and a collection defined by the area 2040 scanned by 2010f. The scan is performed in a way that enables all sub-blocks in contact with the collection 2040 to be scanned before the lower right corner (2030) of the collection 2040. The scan pattern 2010 scans an aggregate of sub-blocks formed using the scans 2010a to 2010e. When the collection covered by 2010a to 2010e is completed, the scan pattern 2010 continues to the next collection 2040 scanned according to 2010f. In order to enable signaling of mts_idx, there is a characteristic that confirms that the position of the last significant coefficient such as 2030 is within the area 2040, and it is not necessary to confirm that the residual coefficients outside the area 2040 are zero values.

[0171] The scanning of the residual coefficients is performed as a variation of the backward diagonal scanning of FIG. 20. The scan pattern scans the collection in the backward raster manner in FIG. 20. In variations of the patterns of FIGS. 18 and 19, the collection may be scanned in the backward raster order.

[0172] The scan patterns shown in FIGS. 18 to 20, namely 1810, 1910, 2010a to f, substantially retain the property of progressing from the coefficient of the highest frequency of the TB towards the coefficient of the lowest frequency of the TB as compared with the scan pattern 1710 of FIG. 17. Therefore, the arrangement of the video encoder 114 and the video decoder 134 using the scan patterns 1810, 1910, and 2010a to f can achieve the same compression efficiency as that achieved when using the scan pattern 1710 while being able to depend on the position of the last significant coefficient without the need to further check for zero-value residual coefficients outside the MTS transform coefficient region.

[0173] FIG. 13 shows a method 1300 for encoding frame data 113 into a bitstream 115, and the bitstream 115 includes one or more slices as a sequence of encoding tree units. The method 1300 can be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, the method 1300 may be executed by the video encoder 114 under the execution of the processor 205. Thus, the method 1300 may be implemented as a module of software 233 stored in a computer-readable storage medium and / or memory 206.

[0174] Method 1300 begins with step 1310 of SPS / PPS encoding. In step 1310, video encoder 114 encodes SPS 1010 and PPS 1012 into bitstream 115 as a sequence of fixed-length and variable-length encoding parameters. Parameters of frame data 113, such as resolution and sample bit depth, are encoded. Also encoded are parameters of the bitstream, such as flags indicating the usage of specific encoding tools. The picture parameter set includes parameters that specify the frequency with which the "delta QP" syntax element is present in bitstream 113, the offset of chroma QP relative to luma QP, etc.

[0175] Method 1300 continues from step 1310 to step 1320 of picture header encoding. In the execution of step 1320, processor 205 encodes a picture header (e.g., 1015) into bitstream 113, and picture header 1015 is applicable to all slices within the current frame. Picture header 1015 includes partitioning constraints indicating the maximum allowable depth of binary, ternary, and quad-tree partitioning, and can override similar constraints included as part of SPS 1010.

[0176] Method 1300 continues from step 1320 to step 1330 of slice header encoding. In step 1330, entropy encoder 338 encodes slice header 1118 into bitstream 115.

[0177] Method 1300 continues from step 1330 to step 1340 of dividing a slice into CTUs. In the execution of step 1340, video encoder 114 divides slice 0116 into a sequence of CTUs. Slice boundaries are aligned with CTU boundaries, and the CTUs within a slice are ordered according to the CTU scan order, generally the raster scan order. The division of a slice into CTUs establishes the order in which portions of frame data 113 are to be processed by video encoder 113 when encoding each current slice.

[0178] Method 1300 continues from step 1340 to step 1350 of encoding tree determination. At step 1350, video encoder 114 determines an encoding tree for the currently selected CTU within a slice. Method 1300 starts from the first CTU in slice 0116 in the first call of step 1350 and proceeds to subsequent CTUs in slice 0116 in subsequent calls. When determining the encoding tree of a CTU, various combinations of quadtree, binary, and ternary partitions are generated and tested by block partitioner 310.

[0179] Method 1300 proceeds from step 1350 to step 1360 of encoding unit determination. In step 1360, video encoder 114 executes to determine an encoding for CUs resulting from various encoding trees under evaluation using known methods. Determining the encoding includes determining a prediction mode (e.g., intra prediction by a specific mode 387 or inter prediction by a motion vector) and a primary transform type 389. If the primary transform type 389 is determined to be DCT-2 and all quantized primary transform coefficients that do not undergo forward secondary transform are insignificant, a secondary transform index 388 is determined, which can indicate the application of secondary transform (e.g., encoded as 1120, 1224, or 1254). Otherwise, the secondary transform index 388 indicates a bypass of the secondary transform. Further, a transform skip flag 390 is determined for each TB within the CU, indicating whether to apply a primary transform (and optionally a secondary transform) or to completely bypass the transform (e.g., 1126 / 1130 / 1134 or 1232 / 1262 / 1266, etc.). For the luma channel, the type of primary transform is determined to be one of DCT-2, transform skip, or the MTS option, and for the chroma channel, DCT-2 or transform skip are the available transform types. Determining the encoding can also include determining a quantization parameter that can change the QP, i.e., the "delta QP" syntax element is encoded in bitstream 115. When determining individual encoding units, the optimal encoding tree is also jointly determined. When the encoding units in the shared encoding tree are encoded using intra prediction, in step 1360, the luma intra prediction mode and the chroma intra prediction are determined. When the encoding units in a separate encoding tree are to be encoded using intra prediction, depending on whether the branches of the encoding tree are luma or chroma respectively, either the luma intra prediction mode or the chroma intra prediction mode is determined in step 1360.

[0180] Step 1360 of symbolization unit determination may prohibit the test application of the secondary transformation when there is no "AC" residual coefficient in the primary region residual resulting from the application of the DCT-2 primary transformation by the forward primary transformation module 326. The AC residual coefficient is the residual coefficient at a position other than the upper left position of the transformation block. The prohibition of the test of the secondary transformation when only the DC primary coefficient exists extends to the block to which the secondary transformation index 388 is applied, that is, Y, Cb, Cr of the shared tree (the Y channel only when the Cb and Cr blocks have a width or height of 2 samples). Regardless of whether the symbolization unit is for the shared tree or the separate tree, if there is at least one significant AC main coefficient, the video encoder 114 tests for the selection of a non-zero secondary transformation index value 388 (i.e., for the application of the secondary transformation).

[0181] Method 1300 continues from step 1360 to step 1370 of symbolization unit encoding. At step 1370, the video encoder 114 encodes the determined symbolization unit in step 1360 into the bitstream 115. An example of how the symbolization unit is encoded will be described in more detail with reference to FIG. 14.

[0182] Method 1300 continues from step 1370 to step 1380 of the last symbolization unit test. At step 1380, the processor 205 tests whether the current symbolization unit is the last symbolization unit of the CTU. If not (at step 1380, "NO"), the control within the processor 205 returns to the symbolization unit determination step 1360. Otherwise, if the current symbolization unit is the last symbolization unit (at step 1380, "YES"), the control within the processor 205 proceeds to step 1390 of the last CTU test.

[0183] In step 1390 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU within slice 0116. If the current CTU is not the last CTU within the slice (1016 "NO" in step 1390), the control within the processor 205 returns to the decision encoding tree step 1350. Otherwise, if the current CTU is the last one ( "YES" in step 1390), the control within the processor 205 proceeds to step 13100 of the last slice test.

[0184] In step 13100 of the last slice test, the processor 205 tests whether the currently encoded slice is the last slice within the frame. If the current slice is not the last slice ( "NO" in step 13100), the control within the processor 205 returns to step 1330 of slice header encoding. Otherwise, if the current slice is the last slice and all slices have been encoded ( "YES" in step 13100), method 1300 ends.

[0185] FIG. 14 is a diagram showing a method 1400 for encoding an encoding unit into a bitstream 115, corresponding to step 1370 of FIG. 13. The method 1400 may be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, the method 1400 may be executed by the video encoder 114 under the execution of the processor 205. Thus, the method 1400 may be stored as a module of software 233 on a computer-readable storage medium and / or in the memory 206.

[0186] Method 1400 results in improved compression efficiency by encoding the secondary transform index 1254 only when it is applicable to the chroma TB of TU1260 and encoding the secondary transform index 1120 only when it is applicable to any of the TBs of TU1124. When a shared coding tree is used, method 1400 is invoked for each CU of the coding tree, e.g., CU1114 in FIG. 11, and the Y, Cb, and Cr color channels are coded. When separate coding trees are used, method 1400 is first invoked for each CU of the luma branch 1214a, e.g., 1220, and method 1400 is also invoked for each chroma CU of the chroma branch 1214b, e.g., 1250.

[0187] Method 1400 starts with step 1410 of generating a prediction block. In step 1410, video encoder 114 generates prediction block 320 according to the prediction mode of the CU determined in step 1360, for example, intra prediction mode 387. Entropy encoder 338 encodes intra prediction mode 387 for the coding unit determined in step 1360 into bitstream 115. The "pred_mode" syntax element is encoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit. When intra prediction is used for the coding unit, if the luma PB is applicable to the CU, the luma intra prediction mode is encoded, and if the chroma PB is applicable to the CU, the chroma intra prediction mode is encoded. That is, for an intra-predicted CU belonging to a common tree such as CU1114, prediction mode 1116 includes the luma intra prediction mode and the chroma intra prediction mode. For an intra-predicted CU belonging to the luma branch of a separate coding tree such as CU1220, prediction mode 1221 includes the luma intra prediction mode. For an intra-predicted CU belonging to the chroma branch of another coding tree such as CU1250, prediction mode 1251 includes the chroma intra prediction mode. Primary transform type 389 is encoded to select from the use of DCT-2 in the horizontal and vertical directions, transform skip in the horizontal and vertical directions, or a combination of DCT-8 and DST-7 in the horizontal and vertical directions for the luma TB of the coding unit.

[0188] Method 1400 continues from step 1410 to step 1420 of residual determination. Prediction block 320 is subtracted from the corresponding block of frame data 312 by difference module 322 to generate difference 324.

[0189] Method 1400 continues from step 1420 to step 1430 of residual transformation. In step 1430 of residual transformation, video encoder 114 bypasses the primary and secondary transformations for the residual of step 1420 under the execution of processor 205, or performs the transformation according to primary transformation type 389 and secondary transformation index 388 for each TB of the CU. The transformation of difference 324 may be performed or bypassed according to transformation skip flag 390. If transformed, secondary transformation may also be applied as determined in step 1350 to generate residual samples 350 as described with reference to FIG. 3. After the operation of quantization module 334, residual coefficients 336 are available.

[0190] Method 1400 continues from step 1430 to step 1440 of luma transformation skip flag encoding. In step 1440, entropy encoder 338 encodes context-encoded transformation skip flag 390 into bitstream 115, indicating whether the residual of the luma TB is transformed according to the primary transformation and optionally the secondary transformation, or whether the primary transformation and the secondary transformation are bypassed. Step 1440 is performed in the luma branch of the shared coding tree (coding 1126) or the dual tree (coding 1232) when the CU contains a luma TB.

[0191] Method 1400 continues from step 1440 to step 1450 of luma residual encoding. In step 1450, entropy encoder 338 encodes the residual coefficients 336 for the luma TB into bitstream 115. Step 1450 operates to select an appropriate scan pattern based on the size of the encoding unit. Examples of scan patterns are described in relation to FIG. 17 (conventional scan pattern) and FIGS. 18-20 (additional scan patterns used for MTS flag determination). In the embodiments described herein, the scan patterns related to the examples of FIGS. 18-20 are used. The residual coefficients 336 are typically scanned into a list according to a backward diagonal scan pattern having 4×4 sub-blocks. For a TB having a width or height greater than 16 samples, the scan pattern is as described with reference to FIGS. 18, 19, and 20. The position of the first non-zero residual coefficient in the list (i.e., 1140) is encoded into bitstream 115 as the Cartesian coordinates for the top-left coefficient of the transform block. The remaining residual coefficients are encoded as residual coefficients 1144 in order from the coefficient at the final position to the DC (top-left) residual coefficient. Step 1450 is executed when the CU contains a luma TB, i.e., belongs to the shared encoding tree (encoding 1128), or when the CU belongs to the luma branch of the dual tree (encoding 1234).

[0192] Method 1400 continues from step 1450 to step 1460 of chroma transform skip flag encoding. In step 1460, entropy encoder 338 encodes two additional context encoding transform skip flags 390 into bitstream 115, one for each chroma TB, indicating whether the corresponding TB undergoes DCT-2 transform and optionally a secondary transform, or whether the transform is bypassed. Step 1460 is executed in the shared encoding tree (encoding 1130 and 1134) or the chroma branches of the dual tree (encoding 1262 and 1266) when the CU contains a chroma TB.

[0193] Method 1400 continues from step 1460 to step 1470 of chroma residual encoding. In step 1470, entropy encoder 338 encodes the residual coefficients of the chroma TB into bitstream 115 as described with reference to step 1450. Step 1460 is executed when the CU contains a chroma TB, i.e., in the shared encoding tree (encodings 1132 and 1136) or the chroma branch of the dual tree (encodings 1264 and 1268). For chroma TBs having a width or height greater than 16 samples, the scan pattern is as described with reference to FIGS. 18, 19, and 20. By using the scan patterns of FIGS. 18 - 20 for both luma TBs and chroma TBs, it is possible to avoid the need to define different scan patterns between luma and chroma for TBs of the same size.

[0194] Method 1400 continues from step 1470 to step 1480 of the LFNST signaling test. At step 1480, the processor 205 determines whether a secondary transform can be applied to any of the CUs' TBs. If all of the CUs' TBs use transform skip, there is no need to encode the secondary transform index 388 ( "NO" at step 1480), and method 1400 proceeds to step 14100 of the MTS signaling test. In the case of a shared encoding tree, for example, each of the luma TB and the two chroma TBs is transform skipped to return "NO" at step 1480. In the case of separate encoding trees, the luma TB in the luma branch of the encoding tree is transform skipped for step 1480 to return "NO" for calls regarding luma and chroma respectively, or the two chroma TBs in the chroma branch of the encoding tree are both transform skipped. For a secondary transform to be performed, the corresponding TB only needs to contain significant residual coefficients at the position of the TB targeted for the secondary transform. That is, all other residual coefficients must be zero, and this condition is achieved if the final position of the TB is within 806, 824, 842, or 862 for the TB sizes shown in FIGS. 8A-8D. If the final position of any TB within the CU is outside 806, 824, 842, or 862 for the considered TB size, no secondary transform is performed ( "NO" at step 1480), and method 1400 proceeds to step 14100 of the MTS signaling test.

[0195] In the case of chroma TB, a width or height of 2 may occur. Since there is no kernel defined for a TB with a width or height of 2, it is not subject to secondary conversion (”NO” in step 1480), and method 1400 proceeds to step 14100 of the MTS signaling test. An additional condition for performing secondary conversion is that at least an AC residual coefficient exists among the corresponding TBs. That is, if there is only a significant residual coefficient at the DC (upper left) position of each applied TB, secondary conversion is not performed (”NO” in step 1480), and method 1400 proceeds to step 14100 of the MTS signaling test. On the condition that at least one TB of the CU is subject to primary conversion (the conversion skip flag indicates not skipping for at least one TB of the CU), the acquisition position constraint regarding the TB subject to primary conversion is satisfied, and at least one AC coefficient is included in one or more of the TBs subject to primary conversion (”YES” in step 1480), the control in the processor 205 proceeds to step 1490 of the LFNST index coding. In step 1490 of the LFNST index coding, the entropy encoder 338 encodes a truncated unary codeword indicating three possible selections regarding the application of secondary conversion. The selections are zero (not applied), 1 (the first kernel of the candidate set is applied), and 2 (the second kernel of the candidate set is applied). The codeword uses a maximum of two bins, and each bin is context encoded. Due to the test executed in step 1480, step 1490 is executed only when secondary conversion can be applied, that is, for non-zero indexes to be encoded. Step 1490 encodes, for example, 1120 or 1224 or 1225.

[0196] Substantially, the operations of steps 1480 and 1490 enable the secondary conversion index 1254 for chroma in a separate tree structure to be encoded only when the secondary conversion is applicable to the chroma TB of TU1260. In the shared tree structure, steps 1480 and 1490 operate to encode the secondary conversion index 1120 only when the secondary conversion can be applied to any of the TBs of TU1124. When excluding the relevant secondary conversion indexes (such as 1254 and 1120), method 1400 operates to improve the encoding efficiency. In particular, in the case of a shared or dual tree, unnecessary flags are avoided, thereby reducing the number of bits required and improving the encoding efficiency. In the case of a separate tree, when the corresponding luma conversion block is conversion skipped, the secondary conversion is not necessarily suppressed for chroma.

[0197] Method 1400 proceeds from step 1490 to step 14100 of the MTS signaling test.

[0198] In step 14100 of the MTS signaling, the video encoder 114 determines whether it is necessary to encode the MTS index into the bitstream 115. If the use of DCT-2 conversion is selected in step 1360, the last significant coefficient position can be anywhere within the upper left 32×32 region of the TB. If the last significant coefficient position is outside the upper left 16×16 region of the TB and the scans of FIGS. 18 and 19 (instead of the scan pattern of FIG. 17) are used, there is no need to explicitly signal mts_idx in the bitstream. Since no last significant coefficient outside the upper left 16x16 region is generated when using MTS, in this case, the signal mts_idx is unnecessary in the bitstream. Step 14100 returns "NO", and method 1400 ends with the use of DCT-2 implied by the position of the last significant coefficient.

[0199] The non-DCT-2 selection of the primary transform type is available only when the width and height of the TB are 32 or less. Therefore, for a TB with a width or height exceeding 32, step 14100 returns "NO" and method 1400 ends at step 14100. The non-DCT-2 selection is also available only when no secondary transform is applied. Thus, if it is determined at step 1360 that the secondary transform type 388 is non-zero, step 14100 returns "NO" and method 1400 ends at step 14100.

[0200] When using the scans of FIGS. 18 and 19, the fact that the position of the last significant coefficient lies within the top-left 16×16 region of the TB can result from either the application of the DCT-2 primary transform or the MTS combination of DST-7 and / or DCT-8. Therefore, explicit signaling of mts_idx is required to encode the selection made at step 1360. Thus, when the position of the last significant coefficient is within the top-left 16×16 region of the TB, step 14100 returns "YES" and method 1400 proceeds to step 14110 of MTS index encoding.

[0201] In step 14110 of MTS index encoding, the entropy encoder 338 encodes the truncated unary bit string representing the primary transform type 389. Step 14110 can encode, for example, 1122 or 1226. Method 1400 ends upon execution of step 14110.

[0202] FIG. 15 shows a method 1500 for decoding a bitstream 133 to generate frame data 135, the bitstream 133 including one or more slices as a sequence of encoding tree units. The method 1500 may be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, the method 1500 may be executed by a video decoder 134 under the execution of a processor 205. Thus, the method 1500 may be stored as one or more modules of software 233 on a computer-readable storage medium and / or within a memory 206.

[0203] The method 1500 begins with step 1510 of SPS / PPS decoding. In step 1510, the video decoder 314 decodes SPS 1010 and PPS 1012 from the bitstream 133 as a sequence of fixed-length and variable-length encoding parameters. Parameters of the frame data 113 such as resolution and sample bit depth are decoded. Also, parameters of the bitstream such as flags indicating the use of specific encoding tools are decoded. Default partition constraints may signal the maximum allowable depth of binary, ternary, and quad tree partitions and may be decoded as part of SPS 1010 by the video decoder 134.

[0204] The method 1500 continues from step 1510 to step 1520 of picture header decoding. In the execution of step 1520, the processor 205 decodes a picture header 1015 applicable to all slices within the current frame from the bitstream 113. The picture parameter set includes parameters specifying, for example, the frequency at which the "delta QP" syntax element is present in the bitstream 313, the offset of the chroma QP relative to the luma QP, etc. Optional overridden partition constraints may signal the maximum allowable depth of binary, ternary, and quad tree partitions and may also be decoded as part of the picture header 1015 by the video decoder 134.

[0205] Method 1500 continues from step 1520 to step 1530 of slice header decoding. At step 1530, the entropy decoder decodes 420 slice header 0118 from bitstream 133.

[0206] Method 1500 continues from step 1530 to step 1540 of splitting the slice into CTUs. In performing step 1540, video encoder 114 splits slice 1016 into a sequence of CTUs. The slice boundaries are aligned with the CTU boundaries, and the CTUs within a slice are ordered according to the CTU scan order, generally the raster scan order. The splitting of the slice into CTUs establishes which portions of frame data 133 should be processed by video encoder 313 when decoding the current slice.

[0207] Method 1500 continues from step 1540 to step 1550 of coding tree decoding. At step 1550, video decoder 314 decodes the coding tree of the currently selected CTU within the slice. Method 1500 starts from the first CTU within slice 1016 in the first call of step 1550 and proceeds to subsequent CTUs within slice 1016 in subsequent calls. When decoding the coding tree of the CTU, the flags indicating the combination of quad-tree, binary, and ternary splits determined at step 1350 in video encoder 114 are decoded.

[0208] Method 1500 continues from step 1550 to step 1570 of coded unit decoding. At step 1570, video decoder 314 decodes the coded unit determined at step 1560 from bitstream 133. An example of how the coded unit is decoded will be described in more detail with reference to FIG. 16.

[0209] Method 1500 continues from step 1570 to step 1580 of the last encoding unit test. At step 1580, the processor 205 tests whether the current encoding unit is the last encoding unit of the CTU. If not (”NO” at step 1580), the control within the processor 205 returns to step 1560 of encoding unit decoding. Otherwise, if the current encoding unit is the last encoding unit (”YES” at step 1580), the control within the processor 205 proceeds to step 1590 of the last CTU test.

[0210] At step 1590 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU of slice 1016. If it is not the last CTU of slice 1016 (”NO” at step 1590), the control within the processor 205 returns to step 1550 of encoding tree decoding. Otherwise, if the current CTU is the last one (”YES” at step 190), the control within the processor proceeds to step 15100 of the last slice test.

[0211] At step 15100 of the last slice test, the processor 205 tests whether the current slice being decoded is the last slice within the frame. If the current slice is not the last slice (”NO” at step 15100), the control within the processor 205 returns to step 1530 of slice header decoding. Otherwise, if the current slice is the last slice and all slices have been decoded (”YES” at step 15100), method 1500 ends.

[0212] FIG. 16 is a diagram illustrating a method 1600 for decoding coded units from a bitstream 133 corresponding to step 1570 of FIG. 15. The method 1600 may be implemented by a device such as a configured FPGA, ASIC, or ASSP. Further, the method 1600 may be executed by a video decoder 314 under the execution of a processor 205. Thus, the method 1600 may be stored on a computer-readable storage medium and / or as one or more modules of software 233 within a memory 206.

[0213] When a shared coding tree is used, the method 1600 is invoked for each CU of the coding tree, e.g., CU1114 of FIG. 11, and the Y, Cb, and Cr color channels are coded in a single invocation. When separate coding trees are used, the method 1600 is first invoked for each CU of the luma branch 1214a, e.g., 1220, and the method 1600 is also separately invoked for each chroma CU of the chroma branch 1214b, e.g., 1250.

[0214] The method 1600 starts from step 1610 of luma transform skip flag decoding. In step 1610, the entropy decoder 420 decodes the context-coded transform skip flag 478 (e.g., coded in the bitstream as 1126 of FIG. 11 or 1232 of FIG. 12) from the bitstream 133. The skip flag indicates whether the transform is applied to the luma TB. The transform skip flag 478 indicates that the residual for the luma TB is transformed according to whether (i) a primary transform, (ii) a primary transform and a secondary transform, or (iii) the primary transform and the secondary transform are bypassed. Step 1610 is executed when the CU includes a luma TB in a shared coding tree (e.g., decoding 1126). Step 1610 is executed when the CU belongs to the luma branch of a dual tree of a separate coding tree CTU (decoding 1232).

[0215] Method 1600 continues from step 1610 to step 1620 of inverse quantization of the luma residual. In step 1620, entropy decoder 420 decodes the 424 residual coefficients for the luma TB from bitstream 115. The residual coefficients 424 are combined into the TB by applying a scan to the list of decoded residual coefficients. Step 1620 operates to select an appropriate scan pattern based on the size of the encoding unit. Examples of scan patterns are described in connection with FIG. 17 (conventional scan pattern) and FIGS. 18-20 (additional scan patterns useful for determining the MTS flag). In the examples described herein, a scan pattern based on the patterns described in connection with FIGS. 18-20 is used. This scan is typically a backward diagonal scan pattern using 4×4 sub-blocks as defined with reference to FIGS. 18 and 19. The position of the first non-zero residual coefficient in the list (i.e., 1140) is decoded as Cartesian coordinates for the coefficient in the upper left of the transform block from bitstream 133. The remaining residual coefficients are decoded as residual coefficients 1144 in order from the coefficient at the final position to the DC (upper left) residual coefficient.

[0216] For each sub-block other than the sub-block in the upper left of the TB and the sub-block containing the last significant residual coefficient, a "coded sub-block flag" indicating that there is at least one significant residual coefficient in each sub-block is decoded. When the coded sub-block flag indicates the presence of at least one significant residual coefficient within the sub-block, a "significance map" (a set of flags) is decoded, indicating the significance of each residual coefficient within the sub-block. If a sub-block is shown to contain at least one significant residual coefficient from the decoded coded sub-block flag and the scan reaches the last scan position of the sub-block without encountering a significant residual coefficient, the residual coefficient at the last scan position of the sub-block is presumed to be significant. The coded sub-block flag and the significance map (each flag is named "sig_coeff_flag") are coded using context-coded bins. For each significant residual coefficient within the sub-block, an "abs_level_gtx_flag" indicating whether the magnitude of the corresponding residual coefficient is greater than 1 is decoded. For each residual coefficient within the sub-block having a magnitude greater than 1, a "par_level_flag" and an "abs_level2_gtx_flag" are decoded to further determine the magnitude of the residual coefficient according to Equation (1). AbsLevelPass1 = sig_coeff_flag + par_level_flag + abs_level_gtx_flag + 2×abs_level_gtx_flag2 (1)

[0217] The syntax elements of abs_level_gtx_flag and abs_level_gtx_flag2 are encoded using context - encoded bins. For each residual coefficient having abs_level_gtx_flag2 equal to 1, the bypass - encoded syntax element "abs_remainder" is decoded using Rice - Golomb coding. The decoded magnitude of the residual coefficient is determined as follows. AbsLevel = AbsLevelPass1+2×abs_remainder. To obtain the value of the residual coefficient from the magnitude of the residual coefficient, sign bits are decoded for each significant residual coefficient. The orthogonal coordinates of each sub - block of the scan pattern can be derived from the scan pattern by adjusting (right - shifting) the orthogonal coordinates of the X and Y residual coefficients by log2 of the width and height of the sub - block, respectively. In the case of luma TB, the sub - block size is always 4×4, and X and Y are right - shifted by 2 bits. The scan patterns of FIGS. 18 to 20 can also be applied to chroma TB to avoid storing different scan patterns for blocks of the same size but different color channels. Step 1620 is executed when the CU contains a luma TB, i.e., in the shared encoding tree (decoding 1128), or for a call to the luma branch of the dual tree (e.g., decoding 1234).

[0218] Method 1600 continues from step 1620 to step 1630 of chroma conversion skip flag decoding. In step 1630, the entropy decoder decodes the context-encoded flag from bitstream 1 for each chroma 33TB. For example, the context-encoded flag may be encoded as 1130 and 1134 in FIG. 11, or 1262 and 1266 in FIG. 12). At least one flag is decoded one by one for each of the chroma TBs. The flag decoded in step 1630 indicates whether a conversion is to be applied to the corresponding chroma TB, particularly whether a DCT-2 conversion and optionally a secondary conversion are to be applied to the corresponding chroma TB, or whether all conversions for the corresponding chroma TB are to be bypassed. Step 1630 is executed when the CU contains a chroma TB, i.e., when the CU belongs to a shared coding tree (decoding 1130 and 1134) or a chroma branch of a dual tree (decoding 1262 and 1266).

[0219] Method 1600 continues from step 1630 to step 1640 of chroma residue decoding. In step 1640, the entropy decoder 420 decodes the residual coefficients of the chroma TB from bitstream 133. Step 1640 operates according to the scan pattern defined in FIGS. 18 and 19 in the same way as described with reference to step 1620. Step 1640 is executed when the CU contains a chroma TB, i.e., when the CU belongs to a shared coding tree (decoding 1132 and 1136) or a chroma branch of a dual tree (decoding 1264 and 1268).

[0220] Method 1600 continues from step 1640 to step 1650 of the LFNST signaling test. In step 1650, the processor 205 determines whether the secondary transformation is applicable to any TB of the CU. The luma transformation skip flag can have a value different from that of the chroma transformation skip flag. If all of the TBs of the CU use transformation skip, the secondary transformation is not applicable and there is no need to encode the secondary transformation index ( "NO" in step 1650), and method 1600 proceeds to step 1660 of LFNST index determination. For example, in the case of a shared coding tree, each of the luma TB and the two chroma TBs is transformation skipped to return "NO" in step 1650. For a CU belonging to the luma branch of a separate coding tree (e.g., 1220), step 1650 returns "NO" when the luma TB is transformation skipped. For a CU (e.g., 1250) belonging to the chroma branch of a separate coding tree, step 1650 returns "NO" when both chroma TBs are transformation skipped. For a CU belonging to the chroma branch of another coding tree (e.g., 1250) and having a width or height of less than 4 samples, step 1650 returns "NO". To perform the secondary transformation, the corresponding TB only needs to include significant residual coefficients at the positions of the TBs targeted for the secondary transformation. That is, all other residual coefficients must be zero, and this condition is achieved when the final position of the TB is within 806, 824, 842, or 862 for the TB sizes shown in FIGS. 8A to 8D. If the final position of any TB within the CU is outside 806, 824, 842, or 862 for the considered TB size, the secondary transformation is not performed ( "NO" in step 1650), and method 1600 proceeds to step 1660 of LFNST index determination. For chroma TBs, a width or height of 2 may occur. A TB with a width or height of 2 is not targeted for secondary transformation because there is no kernel defined for such a size of TB. As an additional condition for performing the secondary transformation, there must be at least one AC residual coefficient in the corresponding TB.That is, if the significant residual coefficient is only at the position of DC (upper left) for each TB, the secondary conversion is not executed (\"NO\" in step 1650), and the method 1600 proceeds to step 1660 of LFNST index determination. The constraints regarding the position of the last significant coefficient and the existence of non-DC residual coefficients are applied only to TBs of applicable sizes, that is, TBs having a width and height greater than 2 samples. Control within the processor 205 proceeds to step 1670 of LFNST index decoding on condition that at least one applicable TB is converted, the final position constraint is satisfied, and the non-DC coefficient requirement is satisfied (\"YES\" in step 1650).

[0221] Step 1660 of LFNST index determination is performed when the secondary conversion cannot be applied to any of the TBs associated with the CU. In step 1660, the processor 205 determines that the secondary conversion index has a value of zero indicating the absence of the application of the secondary conversion. Control in the processor 205 proceeds from step 1660 to step 1672 of MTS signaling.

[0222] In step 1670 of LFNST index decoding, the entropy decoder 420 decodes the truncated monomial codewords as a secondary transformation index 474 indicating three possible choices for the application of the secondary transformation. The choices are 0 (not applied), 1 (the first kernel of the candidate set is applied), and 2 (the second kernel of the candidate set is applied). The codewords use at most two bins, and each bin is context encoded. By the test performed in step 1650, step 1670 is only executed if it is possible for the secondary transformation to be applied, i.e., a non-zero index to be decoded. When method 1600 is called as part of a common encoding tree, step 1670 decodes 1120 from bitstream 133. When method 1600 is called as part of a luma branch of a separate encoding tree, step 1670 decodes 1224 from bitstream 133. When step 1670 is called as part of a chroma branch of a separate encoding tree, step 1670 decodes 1254 from bitstream 133. Control within the processor 205 proceeds from step 1670 to step 1672 of MTS signaling.

[0223] Steps 1650, 1660, and 1670 operate to determine the LFNST index, i.e., 474. The LFNST index is decoded from the video bitstream (e.g., decoded 1120, 1224, or 1254) when at least one of the luma transform skip flag and the chroma transform skip flag applied to the CU indicates that the transform of each transform block is not skipped (''YES'' in step 1650, step 1670 is executed). The LFNST index is determined to indicate not to apply the secondary transform when all of the luma transform skip flag and the chroma transform skip flag applicable to the CU indicate to skip the transform of each transform block (''NO'' in step 1650, and step 1660 is executed). In the case of the shared tree, the luma skip value, the chroma skip value, and the LFNST index can be different. For example, the LFNST index decoded for the chroma transform block can be based on the decoded chroma skip flag even if the decoded luma transform skip flag in the collocated block indicates that the transform for the luma block is skipped. Encoding steps 1480 and 1490 operate in a similar manner.

[0224] In step 1672 of MTS signaling, the video decoder 114 determines whether it is necessary to decode the MTS index from the bitstream 133. When encoding the bitstream, if the use of DCT-2 transform is selected in step 1360, the position of the last significant coefficient can be anywhere within the upper left 32×32 region of the TB. If the position of the last significant coefficient decoded in step 1620 is outside the upper left 16×16 region of the TB and the scans of FIGS. 18 and 19 are used, since a primary transform other than DCT-2 does not generate a last significant coefficient outside this region, it is not necessary to explicitly decode the mts_idx. Step 1672 returns "NO", and method 1600 proceeds from step 1672 to step 1674 of MTS index determination. The non-DCT2 primary transform is available only when the secondary transform type 474 indicates bypassing the application of the secondary transform kernel. Accordingly, when the secondary transform type 474 has a non-zero value, method 1600 proceeds from step 1672 to step 1674. When using the scans of FIGS. 18 and 19, the presence of the last significant coefficient position within the upper left 16×16 region of the TB can result from either the application of the DCT-2 primary transform or an MTS combination of DST-7 and / or DCT-8. Therefore, explicit signaling of the mts_idx is required to encode the selection made in step 1360. Thus, when the last significant coefficient position is within the upper left 16×16 region of the TB, step 1672 returns "YES", and method 1600 proceeds to step 1676 of MTS index decoding.

[0225] The non-DCT-2 primary transform is available only when the secondary transform type 474 indicates bypassing the application of the secondary transform kernel. Accordingly, when the secondary transform type 474 has a non-zero value, method 1600 proceeds from step 1672 to step 1674. When using the scans of FIGS. 18 and 19, the presence of the last significant coefficient position within the upper left 16×16 region of the TB can result from either the application of the DCT-2 primary transform or an MTS combination of DST-7 and / or DCT-8. Therefore, explicit signaling of the mts_idx is required to encode the selection made in step 1360. Thus, when the last significant coefficient position is within the upper left 16×16 region of the TB, step 1672 returns "YES", and method 1600 proceeds to step 1676 of MTS index decoding.

[0226] In step 1674 of MTS index determination, the video decoder 134 determines to use DCT-2 as the primary transform. The primary transform type 476 is set to zero. Method 1400 proceeds from step 1674 to the transform residue step 1680.

[0227] In step 1676 of MTS index decoding, the entropy decoder 420 decodes the truncated unary bit string from the bitstream 133 to determine the primary transform type 476. The truncated string is in the bitstream, for example, like 1122 in FIG. 11 or 1226 in FIG. 12. Method 1400 proceeds from step 1676 to the residue transform step 1680.

[0228] Steps 1670, 1672, and 1674 operate to determine the MTS index of the encoding unit. The MTS index is decoded from the video bitstream if the last significant coefficient is at or within the threshold coordinates (15, 15) ( "YES" in steps 1672 and 1676). The MTS index is determined to indicate that MTS is not applied if the last significant coefficient is outside the threshold coordinates ( "NO" in steps 1672 and 1674). Encoding steps 14100 and 14110 operate in a similar manner.

[0229] In an alternative arrangement of the video encoder 114 and the video decoder 134, the appropriately sized chroma TB (to which MTS is not applied) is scanned according to the scan pattern as described with reference to FIG. 17, and the luma TB utilizes scanning according to FIGS. 18 and 19, and only the DST-7 / DCT-8 combination is applied to the luma TB.

[0230] In step 1680 of the residual transformation, video decoder 314 either bypasses the inverse primary and inverse secondary transforms on the residual of step 1420 under the execution of processor 205, or performs an inverse transform according to the primary transform type 476 and the secondary transform index 474. The transform is performed for each TB of the CU according to the decoded transform skip flag 478 for each TB of the CU, as described with reference to FIG. 4. The primary transform type 476 selects whether to use DCT-2 in the horizontal and vertical directions for the luma TB of the coding unit, or to use a combination of DCT-8 and DST-7 in the horizontal and vertical directions. Substantially, step 1680 transforms the luma transform block of the CU and decodes the coding unit according to the decoded luma transform skip flag, the primary transform type 476, and the secondary transform index determined by the operations of steps 1610 and 1650 to 1670. Also, step 1680 can transform the chroma transform block of the CU and decode the coding unit according to the respective decoded chroma transform skip flag and secondary transform index determined by the operations of steps 1630 and 1650 to 1670. For TBs belonging to the chroma channel (e.g., 1132 and 1136 in the case of a shared coding tree, 1264 and 1268 of the chroma branch in the case of a separate coding tree), since there is no available secondary transform kernel for TBs having a width or height of less than 4 samples, the secondary transform is only performed when the width and height of the TB are 4 samples or more. For TBs belonging to the chroma channel, it is difficult to process such small-sized TBs at the block throughput speed required to support video formats such as UHD and 8K. Therefore, the VVC standard has a restriction on the splitting operation that prohibits intra prediction CUs with a TB size of 2×2, 2×4, 4×2. Furthermore, there is a restriction that prohibits intra prediction CUs with a TB of width 2 because it is difficult to access the on-chip memory normally used to generate samples reconstructed as part of the intra prediction operation. Therefore, Table 1 shows the chroma TB sizes (chroma sample units) to which the secondary transform is not applied.

Table 1

[0231] As described in this specification, different scan patterns can be used in encoding and decoding. Step 1680 converts the transform block of the CU according to the MTS index and decodes the coding unit.

[0232] Method 1600 continues from step 1680 to step 1690 of generating a prediction block. In step 1690, video decoder 134 generates a prediction block 452 according to the prediction mode of the CU determined in step 1360 and decoded from bitstream 113 by entropy decoder 420. Entropy decoder 420 decodes the prediction mode for the coding unit as determined in step 1360 from bitstream 133. The "pred_mode" syntax element is decoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit. If intra prediction is used for the coding unit, the luma intra prediction mode is decoded if luma PB is applicable to the CU, and the chroma intra prediction mode is decoded if chroma PB is applicable to the CU.

[0233] Method 1600 continues from step 1690 to step 16100 of coding unit reconstruction. In step 16100, the prediction block is added 452 to the residual samples 424 for each color channel of the CU to generate the reconstructed samples 456. Additional in-loop filtering steps such as deblocking may be applied to the reconstructed samples 456 before being output as frame data 135. Method 1600 ends with the execution of step 16100.

[0234] As described above, in the case of separate coding trees, method 1600 is first called for each CU in the luma branch 1214a, e.g., 1220, and method 1600 is also separately called for each chroma CU in the chroma branch 1214b, e.g., 1250. The call to method 1600 for chroma determines the LFNST index 1254 at steps 1650 to 1670 with respect to whether all of the chroma transform skip flags of CU 1250 are set. Similarly, in the call to method 1600 for luma, the luma LFNST index 1224 is determined at steps 1650 to 1670 with respect to the luma transform skip flag of only CU 1220.

[0235] The scan patterns shown in FIGS. 18 to 20, i.e., 1810, 1910, and 2010a to f, which are performed at steps 1450 and 1620, substantially retain the characteristic of proceeding from the coefficient of the highest frequency of the TB toward the coefficient of the lowest frequency of the TB as compared with the scan pattern 1710 of FIG. 17. Therefore, the arrangement of the video encoder 114 and the video decoder 134 using the scan patterns 1810, 1910, and 2010a to f can achieve the same compression efficiency as achieved when using the scan pattern 1710 while being able to depend on the last significant coefficient position without the further need to check for zero-value residual coefficients outside the MTS transform coefficient region. The final positions used in the scan patterns of FIGS. 18 to 20 enable the use of MTS only when all significant coefficients are present in an appropriate upper left region such as the upper left 16x16 region. The burden on the decoder 134 to check for flags outside the appropriate region, e.g., outside the 16×16 coefficient region of the TB, to confirm that no further insignificant coefficients exist is removed. The operation in the decoder does not require specific changes to implement MTS. Further, as described above, the use of the scan patterns of FIGS. 18 and 19, i.e., for transform blocks of sizes 16×32, 32×16, and 32×32, is replicated from the 16×16 scan, thereby making it possible to reduce memory requirements.

Industrial Applicability

[0236] This method is applicable to computers and the data processing industry, particularly digital signal processing for encoding and decoding video and image signals, and achieves high compression efficiency.

[0237] Some of the conventions described in this specification improve compression efficiency by signaling the secondary transformation index when the available choices include at least one option other than bypassing the secondary transformation. The improvement in compression efficiency is achieved both when the CTU is divided into CUs spanning all color channels ("shared coding tree") and when the CTU is divided into a set of luma CUs and a set of chroma CUs ("separate coding tree"). In the case of a separate tree, redundant signaling of the secondary transformation index is avoided when the secondary transformation index cannot be used. In a shared tree, the LFNST index can be signaled even when the luma uses transform skip in the primary case of chroma DCT-2. Other arrangements maintain compression efficiency while allowing MTS index signaling to depend on the last significant coefficient position without the further need to check for zero-valued residual coefficients outside the MTS transform coefficient region of the TB.

[0238] It should be noted that the above only describes some embodiments of the present invention, and modifications and / or changes can be made without departing from the scope and spirit of the present invention, and the embodiments are illustrative and not restrictive.

Claims

1. A method for decoding an encoded unit from a bitstream, wherein the encoded unit is split from an encoded tree unit of an image using a tree structure, the encoded unit can have at least a luma component or a plurality of chroma components, the plurality of chroma components include a Cb component and a Cr component, and the method includes: a first decoding step of decoding a luma transform skip flag for the luma component from the bitstream when the encoded unit has the luma component; a second decoding step of decoding a first chroma transform skip flag for the Cb component and a second chroma transform skip flag for the Cr component from the bitstream when the encoded unit has the plurality of chroma components; a determination step of determining for the encoded unit whether to decode an index for a specific transform process from the bitstream, wherein the kernel used in the specific transform process can be selected from a candidate set of a plurality of kernels, and the index is an index for specifying the kernel used, the determination step; a third decoding step of decoding the index for the specific transform process from the bitstream for the encoded unit according to the result of the determination in the determination step; and the luma transform skip flag indicates whether the luma transform process for the luma component is skipped; the first chroma transform skip flag indicates whether the first chroma transform process for the Cb component is skipped; the second chroma transform skip flag indicates whether the second chroma transform process for the Cr component is skipped. When the encoding tree unit has a size of 128×128 and the encoding tree structure for the luma component in the encoding tree unit is separate from the encoding tree structure for the plurality of chroma components in the encoding tree unit, (a) the encoding tree unit is divided into four regions each having a size of 64×64 in common for the luma component and the plurality of chroma components, (b) the dual-tree structure for the luma component and the dual-tree structure for the plurality of chroma components start for each of the four regions, and (c) before the determination of whether to decode the index for each of the encoded units divided using the dual-tree structure for the plurality of chroma components from a certain region among the four regions is executed, the determination of whether to decode the index for each of the encoded units divided using the dual-tree structure for the luma component from the certain region is executed. When the encoded unit is divided from the encoding tree unit using a single-tree structure and each transform block in the encoded unit has significant coefficients only at the DC position, the index for the encoded unit is not always decoded from the bitstream. When the luma transform process, the first chroma transform process, and the second chroma transform process are skipped and the encoded unit is divided from the encoding tree unit using a single-tree structure, the specific transform process for the encoded unit is not always executed regardless of other conditions, the index for the encoded unit is not decoded from the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the luma transform process is skipped and the encoded unit is divided from the encoding tree unit using the dual-tree structure for the luma component, the index for the encoded unit is not decoded from the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the first chroma conversion process and the second chroma conversion process are skipped and the encoding unit is split from the encoding tree unit using the dual-tree structure for the plurality of chroma components, the index for the encoding unit is not decoded from the bitstream, and the value of the index for the encoding unit is presumed to be 0. It is possible to use ternary splitting to split the encoding units in the encoding tree unit into a plurality of encoding units. A method characterized by this.

2. The image has a 4:2:0 chroma format. The method according to claim 1, characterized by this.

3. When intra prediction and the single-tree structure are used, the use of chroma blocks having a size of 2×2, 2×4, or 4×2 is not permitted. The method according to claim 1, characterized by this.

4. The index having a value of 0 indicates that the specific conversion process is not used. The method according to claim 1, characterized by this.

5. The DC position in the transform block is the upper left position among a plurality of positions in the transform block. The method according to claim 1, characterized by this.

6. The DC position in the transform block is the position that is scanned last in a predetermined scan order among a plurality of positions in the transform block. The method according to claim 1, characterized by this.

7. Before the dual-tree structure for the plurality of chroma components is decoded for a certain region, the dual-tree structure for the luma component is decoded for the certain region. The method according to claim 1, characterized by this.

8. When the transform block included in the encoding unit has significant coefficients, a flag indicating whether the magnitude of the significant coefficients is greater than 1 is decoded, and the magnitude of the significant coefficients is determined using the decoded flag. The method according to claim 1, characterized by this.

9. Context-encoded bins are used for the flag. The method according to claim 8, characterized by this.

10. The information on the signs of the significant coefficients is decoded. The method according to claim 8, characterized by this.

11. The single-tree structure is a tree structure in which the encoding tree structure is common between the luma component and the plurality of chroma components. The method according to claim 1, characterized in that...

12. A method for encoding an encoding unit into a bitstream, wherein the encoding unit is divided from an encoding tree unit of an image using a tree structure, the encoding unit can have at least a luma component or a plurality of chroma components, the plurality of chroma components include a Cb component and a Cr component, and the method includes: A first encoding step of encoding a luma transform skip flag for the luma component into the bitstream when the encoding unit has the luma component; A second encoding step of encoding a first chroma transform skip flag for the Cb component and a second chroma transform skip flag for the Cr component into the bitstream when the encoding unit has the plurality of chroma components; A determination step of determining for the encoding unit whether to encode an index for a specific transform process into the bitstream, wherein the kernel used in the specific transform process can be selected from a candidate set of a plurality of kernels, and the index is an index that specifies the kernel to be used; the determination step; A third encoding step of encoding the index for the specific transform process into the bitstream for the encoding unit according to the result of the determination in the determination step; comprising The luma transform skip flag indicates whether the luma transform process for the luma component is skipped; The first chroma transform skip flag indicates whether the first chroma transform process for the Cb component is skipped; The second chroma transform skip flag indicates whether the second chroma transform process for the Cr component is skipped. When the encoding tree unit has a size of 128×128 and the encoding tree structure for the luma component in the encoding tree unit is separate from the encoding tree structure for the plurality of chroma components in the encoding tree unit, (a) the encoding tree unit is divided into four regions each having a size of 64×64 in common between the luma component and the plurality of chroma components, (b) the dual-tree structure for the luma component and the dual-tree structure for the plurality of chroma components start for each of the four regions, and (c) before the determination of whether to encode the index for each of the encoded units divided using the dual-tree structure for the plurality of chroma components from a certain region among the four regions is executed, the determination of whether to encode the index for each of the encoded units divided using the dual-tree structure for the luma component from the certain region is executed. When the encoded unit is divided from the encoding tree unit using a single-tree structure and each transform block in the encoded unit has significant coefficients only at the DC position, the index for the encoded unit is not always encoded in the bitstream. When the luma transform process, the first chroma transform process, and the second chroma transform process are skipped and the encoded unit is divided from the encoding tree unit using a single-tree structure, the specific transform process for the encoded unit is not always executed regardless of other conditions, the index for the encoded unit is not encoded in the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the luma transform process is skipped and the encoded unit is divided from the encoding tree unit using the dual-tree structure for the luma component, the index for the encoded unit is not encoded in the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the first chroma conversion process and the second chroma conversion process are skipped, and the encoding unit is divided from the encoding tree unit using the dual-tree structure for the plurality of chroma components, the index for the encoding unit is not encoded in the bitstream, and the value of the index for the encoding unit is presumed to be 0. It is possible to use ternary splitting to divide the encoding units in the encoding tree unit into a plurality of encoding units. A method characterized by this.

13. The image has a 4:2:0 chroma format. The method according to claim 12, characterized by this.

14. When intra prediction and the single-tree structure are used, the use of chroma blocks having a size of 2×2, 2×4, or 4×2 is not permitted. The method according to claim 12, characterized by this.

15. The index having a value of 0 indicates that the specific conversion process is not used. The method according to claim 12, characterized by this.

16. The DC position in the transform block is the upper left position among a plurality of positions in the transform block. The method according to claim 12, characterized by this.

17. The DC position in the transform block is the position that is scanned last in a predetermined scan order among a plurality of positions in the transform block. The method according to claim 12, characterized by this.

18. Before the dual-tree structure for the plurality of chroma components is encoded for a certain region, the dual-tree structure for the luma component is encoded for the certain region. The method according to claim 12, characterized by this.

19. When the transform block included in the encoding unit has significant coefficients, a flag indicating whether the magnitude of the significant coefficients is greater than 1 is encoded using a context bin. The method according to claim 12, characterized by this.

20. The information on the sign of the significant coefficients is encoded. The method according to claim 19, characterized by this.

21. The single-tree structure is a tree structure in which the encoding tree structure is common between the luma component and the plurality of chroma components. The method according to claim 12, characterized by this.

22. An apparatus for decoding an encoded unit from a bitstream, wherein the encoded unit is divided from an encoded tree unit of an image using a tree structure, the encoded unit can have at least a luma component or a plurality of chroma components, the plurality of chroma components include a Cb component and a Cr component, and the apparatus includes: first decoding means for decoding, from the bitstream, a luma conversion skip flag for the luma component when the encoded unit has the luma component; second decoding means for decoding, from the bitstream, a first chroma conversion skip flag for the Cb component and a second chroma conversion skip flag for the Cr component when the encoded unit has the plurality of chroma components; determination means for determining, for the encoded unit, whether to decode an index for a specific conversion process from the bitstream, wherein a kernel used in the specific conversion process can be selected from a candidate set of a plurality of kernels, and the index is an index for specifying the kernel to be used; third decoding means for decoding, from the bitstream, the index for the specific conversion process for the encoded unit according to a result of determination by the determination means; and has the luma conversion skip flag indicates whether luma conversion processing for the luma component is skipped; the first chroma conversion skip flag indicates whether first chroma conversion processing for the Cb component is skipped; the second chroma conversion skip flag indicates whether second chroma conversion processing for the Cr component is skipped. When the encoding tree unit has a size of 128×128 and the encoding tree structure for the luma component in the encoding tree unit is separate from the encoding tree structure for the plurality of chroma components in the encoding tree unit, (a) the encoding tree unit is divided into four regions each having a size of 64×64, in common between the luma component and the plurality of chroma components, (b) the dual-tree structure for the luma component and the dual-tree structure for the plurality of chroma components start for each of the four regions, (c) before the determination of whether to decode the index for each of the encoded units divided using the dual-tree structure for the plurality of chroma components from a certain one of the four regions is executed, the determination of whether to decode the index for each of the encoded units divided using the dual-tree structure for the luma component from the certain one of the four regions is executed, When the encoded unit is divided from the encoding tree unit using a single-tree structure and each transform block in the encoded unit has significant coefficients only at the DC position, the index for the encoded unit is not always decoded from the bitstream. When the luma transform process, the first chroma transform process, and the second chroma transform process are skipped and the encoded unit is divided from the encoding tree unit using a single-tree structure, the specific transform process for the encoded unit is not always executed regardless of other conditions, the index for the encoded unit is not decoded from the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the luma transform process is skipped and the encoded unit is divided from the encoding tree unit using the dual-tree structure for the luma component, the index for the encoded unit is not decoded from the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the first chroma conversion process and the second chroma conversion process are skipped, and the encoding unit is divided from the encoding tree unit using the dual-tree structure for the plurality of chroma components, the index for the encoding unit is not decoded from the bitstream, and the value of the index for the encoding unit is presumed to be 0. It is possible to use ternary splitting to divide the encoding unit in the encoding tree unit into a plurality of encoding units. An apparatus characterized by the above.

23. An apparatus for encoding an encoding unit into a bitstream, wherein the encoding unit is divided from an encoding tree unit of an image using a tree structure, the encoding unit can have at least a luma component or a plurality of chroma components, the plurality of chroma components include a Cb component and a Cr component, and the apparatus includes: First encoding means for encoding a luma conversion skip flag for the luma component into the bitstream when the encoding unit has the luma component; Second encoding means for encoding a first chroma conversion skip flag for the Cb component and a second chroma conversion skip flag for the Cr component into the bitstream when the encoding unit has the plurality of chroma components; Determination means for determining whether to encode an index for a specific conversion process into the bitstream for the encoding unit, wherein the kernel used in the specific conversion process can be selected from a candidate set of a plurality of kernels, the index is an index for specifying the kernel to be used, and the determination means; Third encoding means for encoding the index for the specific conversion process into the bitstream for the encoding unit according to the result of the determination in the determination means; And has The luma conversion skip flag indicates whether the luma conversion process for the luma component is skipped. The first chroma conversion skip flag indicates whether the first chroma conversion process for the Cb component is skipped. The second chroma conversion skip flag indicates whether the second chroma conversion process for the Cr component is skipped. When the encoding tree unit has a size of 128×128 and the encoding tree structure for the luma component in the encoding tree unit is separate from the encoding tree structure for the plurality of chroma components in the encoding tree unit, (a) the encoding tree unit is divided into four regions each having a size of 64×64 in common between the luma component and the plurality of chroma components, (b) the dual-tree structure for the luma component and the dual-tree structure for the plurality of chroma components start for each of the four regions, and (c) before the determination of whether to encode the index for each of the encoded units divided using the dual-tree structure for the plurality of chroma components from a certain region among the four regions is executed, the determination of whether to encode the index for each of the encoded units divided using the dual-tree structure for the luma component from the certain region is executed. When the encoded unit is divided from the encoding tree unit using a single-tree structure and each transform block in the encoded unit has significant coefficients only at the DC position, the index for the encoded unit is not always encoded in the bitstream. When the luma transform process, the first chroma transform process, and the second chroma transform process are skipped and the encoded unit is divided from the encoding tree unit using a single-tree structure, the specific transform process for the encoded unit is not always executed regardless of other conditions, the index for the encoded unit is not encoded in the bitstream, and the value of the index for the encoded unit is presumed to be 0. When the luma transform process is skipped and the encoded unit is divided from the encoding tree unit using the dual-tree structure for the luma component, the index for the encoded unit is not encoded in the bitstream, and the value of the index for the encoded unit is presumed to be 0. If the first chroma conversion process and the second chroma conversion process are skipped and the encoding unit is split from the encoding tree unit using the dual-tree structure for the plurality of chroma components, the index for the encoding unit is not encoded in the bitstream, and the value of the index for the encoding unit is presumed to be 0. It is possible to use ternary splitting to divide the encoding unit in the encoding tree unit into a plurality of encoding units. An apparatus characterized by this. **Claim 24** A program characterized by causing a computer to execute the method according to any one of claims 1 to 11. **Claim 25** A program characterized by causing a computer to execute the method according to any one of claims 12 to 21.

Citation Information

Patent Citations

  • Image segmentation method and apparatus for image encoding and decoding

    WO2019216710A1

  • Transform-based image coding method and device therefor

    WO2021096290A1

Cited By

  • Method, device and program for decoding and encoding encoded unit

    JP2025157333A