Method and apparatus for decoding coding units from a bit stream, method and apparatus for encoding coding units in a bit stream, storage medium and computer program product

By selecting the appropriate core and transformation method in the video encoding tree unit and decoding for the primary color and secondary color channels, the problem of inefficient high-resolution and high frame rate video encoding in the prior art is solved, and more efficient video compression and feasibility of modern silicon processes are achieved, and real-time encoding of immersive videos is suitable for real-time encoding.

CN120378637APending Publication Date: 2025-07-25CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510641814.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-17
Filing Date
2020-08-04
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing video encoding standards are difficult to achieve efficient compression performance and feasibility in modern silicon processes when processing video data at high resolution and high frame rate, especially immersive video. In addition, inadequate flexibility in intra prediction and inter prediction leads to insufficiency of encoding.

Method used

A new video encoding method is adopted, by decoding the encoding units in the encoding tree unit, the cores of the main color channel and the secondary color channel are selected respectively, different cores and transformation methods are used for decoding, and flexible rate control is performed in intra prediction and inter prediction, and quantization parameters are adjusted to adapt to block partition constraints in different regions.

Benefits of technology

It improves the encoding efficiency of video data, adapts to block partition constraints in different regions, achieves higher compression performance and feasibility in modern silicon processes, supports higher resolution and frame rate video formats, and meets the real-time encoding requirements of immersive videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378637A_ABST
    Figure CN120378637A_ABST
Patent Text Reader

Abstract

A method and apparatus for decoding a coding unit from a bitstream, a method and apparatus for encoding a coding unit in a bitstream, a storage medium and a computer program product are provided. Specifically, a system and method of decoding from a video bitstream a coding unit of a coding tree from coding tree units of an image frame, the coding unit having a primary color channel and at least one secondary color channel. The method comprises: determining a coding unit comprising a primary color channel and at least one secondary color channel according to a decoded split flag of a coding tree unit; decoding the first index to select a core for the primary color channel and decoding the second index to select a core for the at least one secondary color channel; selecting a first core according to the first index, and selecting a second core according to the second index; and decoding the coding unit by applying the first core to the residual coefficient of the primary color channel and applying the second core to the residual coefficient of the at least one secondary color channel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] (This application is a divisional application of the application with the filing date of August 4, 2020, application number 2020800626432, and invention name "Method and apparatus for coding units in a coding tree unit for encoding and decoding images, non-transitory computer-readable storage medium, and computer program product".) Technical Field

[0002] The present invention generally relates to digital video signal processing, and in particular, to methods, apparatuses, and systems for encoding and decoding blocks of video samples. The present invention also relates to a computer program product including a computer-readable medium having recorded thereon a computer program for encoding and decoding blocks of video samples. Background Art

[0003] There currently exist many applications for video coding, including applications for transmitting and storing video data. Many video coding standards have also been developed and other video coding standards are currently under development. The latest progress in video coding standardization has led to the formation of a group known as the "Joint Video Exploration Team" (JVET). The Joint Video Exploration Team (JVET) includes members of Study Group 16, Question 6 (SG16 / Q6) of the Telecommunication Standardization Sector (ITU-T) of the International Telecommunication Union (ITU), also known as the "Video Coding Experts Group" (VCEG); and members of Working Group 11 (WG11) of Subcommittee 29, Committee 1 of the International Organization for Standardization / International Electrotechnical Commission Joint Technical Committee 1 (ISO / IEC JTC1 / SC29 / WG11), also known as the "Moving Picture Experts Group" (MPEG).

[0004] The Joint Video Exploration Team (JVET) issued a Call for Proposals (CfP) and analyzed the responses at its 10th meeting held in San Diego, USA. The responses submitted indicated that the video compression capabilities were significantly better than those of the current state-of-the-art video compression standard, namely, "High Efficiency Video Coding" (HEVC). Based on this excellent performance, it was decided to start a project to develop a new video compression standard named "Versatile Video Coding" (VVC). It is expected that VVC will address the continuous demand for even higher compression performance, especially as the capabilities of video formats increase (e.g., with higher resolutions and higher frame rates), and the growing market demand for service provision over WANs where the bandwidth cost is relatively high. Use cases such as immersive video require real-time encoding and decoding of such higher formats. For example, a cube map projection (CMP) can use an 8K format even if the final rendered "viewport" utilizes a lower resolution. VVC must be implementable in contemporary silicon processes and provide an acceptable trade-off between the implemented performance and the implementation cost. For example, the implementation cost can be considered in one or more of silicon area, CPU processor load, memory utilization, and bandwidth. Higher video formats can be processed by dividing the frame area into parts and processing each part in parallel. The bitstream constructed from multiple parts of the compressed frame is still suitable for decoding by a "single-core" decoder, i.e., frame-level constraints (including bitrate) are assigned to each part according to the application requirements.

[0005] Video data includes a sequence of frames of image data, and each frame includes one or more color channels. Typically, one primary color channel and two secondary color channels are required. The primary color channel is usually referred to as the "luma" channel, and the (one or more) secondary color channels are usually referred to as the "chroma" channels. Although video data is typically displayed in the RGB (Red - Green - Blue) color space, this color space has a high correlation between the three corresponding components. The video data representation seen by an encoder or decoder typically uses a color space such as YCbCr. YCbCr concentrates the luminance (mapped to "luma" according to a transformation equation) in the Y (primary) channel and the chroma in the Cb and Cr (secondary) channels. Due to the use of the decorrelated YCbCr signal, the statistics of the luma channel are significantly different from those of the chroma channels. The main differences are: after quantization, the chroma channels contain relatively fewer significant coefficients for a given block compared to the coefficients of the corresponding luma channel block. In addition, the Cb and Cr channels can be spatially sampled at a lower rate compared to the luma channel (e.g., half in the horizontal direction and half in the vertical direction (referred to as "4:2:0 chroma format")). The 4:2:0 chroma format is commonly used in "consumer" applications such as Internet video streaming, broadcast television, and Blu-ray TMStorage on the disc. Subsampling the Cb and Cr channels at half rate in the horizontal direction instead of vertically is called the "4:2:2 chroma format". The 4:2:2 chroma format is commonly used in professional applications, including the capture of footage for movie production and the like. The higher sampling rate of the 4:2:2 chroma format makes the resulting video more resilient to editing operations such as color grading. Before distribution to consumers, 4:2:2 chroma format material is often converted to the 4:2:0 chroma format and then encoded for distribution to consumers. In addition to the chroma format, video is also characterized by resolution and frame rate. Example resolutions are Ultra High Definition (UD) with a resolution of 3840×2160 or "8K" with a resolution of 7680×4320, and example frame rates are 60 Hz or 120 Hz. The range of the luminance sample rate can be from about 500 megasamples per second to several thousand megasamples per second. For the 4:2:0 chroma format, the sampling rate of each chroma channel is one quarter of the luminance sampling rate, and for the 4:2:2 chroma format, the sampling rate of each chroma channel is half of the luminance sampling rate.

[0006] The VVC standard is a "block-based" codec, where a frame is first segmented into an array of square regions called "Coding Tree Units" (CTUs). A CTU typically occupies a relatively large area, such as 128×128 luminance samples and the like. However, the area of the CTUs at the right and bottom edges of each frame may be smaller. Associated with each CTU is a "coding tree" ("common tree") for both the luminance channel and the chroma channel or separate trees for the luminance channel and the chroma channel respectively. The coding tree defines the decomposition of the area of the CTU into a set of regions, also called "Coding Blocks" (CBs). When using a common tree, a single coding tree specifies the blocks for both the luminance channel and the chroma channel, in which case the set of juxtaposed coding blocks is called a "Coding Unit" (CU), that is, each CU has coding blocks for each color channel. The CBs are processed in a specific order for encoding or decoding. As a result of using the 4:2:0 chroma format, a CTU of the luminance coding tree including a 128×128 luminance sample area has a corresponding chroma coding tree of a 64×64 chroma sample area juxtaposed with the 128×128 luminance sample area. When a single coding tree is used for the luminance channel and the chroma channel, the set of juxtaposed blocks of a given area is commonly called a "unit", such as the above-mentioned CU as well as "Prediction Units" (PUs) and "Transform Units" (TUs). A single tree of a CU with color channels spanning 4:2:0 chroma format video data results in chroma blocks that are half the width and height of the corresponding luminance blocks. When separate coding trees are used for a given area, the above-mentioned CBs as well as "Prediction Blocks" (PBs) and "Transform Blocks" (TBs) will be used.

[0007] Despite the above difference between "units" and "blocks", the term "block" can be used as a general term for an area or region of a frame to which an operation is applied to all color channels.

[0008] For each CU, a prediction unit (PU) ("prediction unit") that generates the content (sample values) of the corresponding region of the frame data. In addition, a representation of the difference between the prediction seen at the input of the encoder and the region content (or the "residual" in the spatial domain) is formed. The difference for each color channel can be transformed and encoded as a sequence of residual coefficients, thereby forming one or more transform units (TUs) for a given CU. The transform applied can be a discrete cosine transform (DCT) or other transform applied to individual blocks of the residual values. The transform is applied separately, i.e., a two-pass two-dimensional transform is performed. First, the block is transformed by applying a one-dimensional transform to the rows of samples in the block. Then, the partial result is transformed by applying a one-dimensional transform to the columns of the partial result to produce a final block of transform coefficients that essentially decorrelates the residual samples. The VVC standard supports transforms of various sizes, including transforms of rectangular blocks (sizes of each side being a power of 2). The transform coefficients are quantized for entropy encoding in the bitstream.

[0009] VVC is characterized by intra prediction and inter prediction. Intra prediction involves using previously processed samples in the frame being used to generate a prediction for the current sample block in that frame. Inter prediction involves using a sample block obtained from a previously decoded frame to generate a prediction for the current sample block in the frame. The sample block obtained from the previously decoded frame is offset from the spatial position of the current block according to a motion vector, which typically has filtering applied. The intra prediction block can be (i) a uniform sample value (“DC intra prediction”), (ii) a plane with an offset and horizontal and vertical gradients (“plane intra prediction”), (iii) a population of blocks with adjacent samples applied in a particular direction (“angular intra prediction”), or (iv) the result of a matrix multiplication using adjacent samples and selected matrix coefficients. Further differences between the prediction block and the corresponding input samples can be corrected to some extent by encoding the ‘residual’ in the bitstream. The residual is typically transformed from the spatial domain to the frequency domain to form residual coefficients (in the “primary transform domain”), and the residual coefficients can be further transformed by applying a “secondary transform” (to produce residual coefficients in the “secondary transform domain”). The residual coefficients are quantized according to a quantization parameter, resulting in a loss of precision in the reconstruction of the samples produced at the decoder, while the bitrate within the bitstream is also reduced. The quantization parameter can vary between frames and within individual frames. For a “rate control” encoder, variation of the intra quantization parameter is typical. Regardless of the statistics of the input samples received (such as noise nature, degree of motion, etc.), the rate control encoder attempts to produce a bitstream with a substantially constant bitrate. Since the bitstream is typically transmitted over a network with a limited bandwidth, rate control is a common technique used to ensure reliable performance on the network regardless of the variation of the original frames input to the encoder. The flexibility of the use of rate control is desirable in cases where the frames are encoded in parallel segments, as different segments may have different requirements in terms of the desired fidelity. SUMMARY OF THE INVENTION

[0010] It is an object of the present invention to substantially overcome or at least ameliorate one or more disadvantages of the existing arrangements.

[0011] One aspect of the present disclosure provides a method for decoding an encoding unit of an encoding tree of an encoding tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and at least one secondary color channel according to a decoded split flag of the encoding tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to the residual coefficients of the primary color channel and applying the second kernel to the residual coefficients of the at least one secondary color channel.

[0012] According to another aspect, the first index or the second index is decoded immediately after decoding the position of the last valid residual coefficient of the encoding unit.

[0013] According to another aspect, a single residual coefficient is decoded for a plurality of secondary color channels.

[0014] According to another aspect, a single residual coefficient is decoded for a single secondary color channel.

[0015] According to another aspect, the first index and the second index are independent of each other.

[0016] According to another aspect, the first kernel and the second kernel respectively depend on the intra prediction mode for the primary color channel and the at least one secondary color channel.

[0017] According to another aspect, the first kernel and the second kernel are respectively related to the block size of the primary channel and the block size of the at least one secondary color channel.

[0018] According to another aspect, the second kernel is related to the chrominance subsampling rate of the encoding bitstream.

[0019] According to another aspect, each kernel in the kernel implements an inseparable quadratic transform.

[0020] According to another aspect, the encoding unit includes two secondary color channels, and a separate index is decoded for each of the secondary color channels in the secondary color channels.

[0021] Another aspect of the present disclosure provides a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoded tree unit; selecting a non-separable transform kernel according to a decoded index of the primary color channel; applying the selected non-separable transform kernel to a decoded residual of the primary color channel to generate secondary transform coefficients; and decoding the encoding unit by applying a separable transform kernel to the secondary transform coefficients and applying a separable transform kernel to decoded residuals of the at least one secondary color channel.

[0022] Another aspect of the present disclosure provides a non-transitory computer-readable medium having stored thereon a computer program for implementing a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoded tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.

[0023] Another aspect of the present disclosure provides a video decoder configured to implement a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoded tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.

[0024] Another aspect of the present disclosure provides a system, which includes: a memory; and a processor, wherein the processor is configured to execute code stored on the memory to implement a method for decoding an encoding unit of an encoding tree from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, and the method includes: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoding tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.

[0025] Another aspect of the present disclosure provides a method for decoding a plurality of encoding units from a bitstream to generate an image frame, the encoding units being the result of decomposition of an encoding tree unit, and the plurality of encoding units forming one or more than one continuous part of the bitstream, and the method includes: determining a subdivision level of each of the one or more than one continuous parts of the bitstream, each subdivision level being applicable to the encoding units of the corresponding continuous part of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on decomposing the encoding tree unit into the encoding units of the respective continuous parts of the bitstream and the corresponding determined subdivision levels; determining a quantization parameter for each region according to the decoded incremental quantization parameter of the region and the quantization parameter of an earlier encoding unit of the image frame; and decoding the plurality of encoding units using the determined quantization parameter for each region to generate an image frame.

[0026] According to another aspect, each region is based on a comparison of the subdivision level associated with the encoding unit and the determined subdivision level of the corresponding continuous part.

[0027] According to another aspect, a quantization parameter increment is determined for each region, wherein the corresponding encoding tree has a subdivision level less than or equal to the determined subdivision level of the corresponding continuous part.

[0028] According to another aspect, a new region is set for any node in the encoding tree unit having a subdivision level less than or equal to the corresponding determined subdivision level.

[0029] According to another aspect, the determined subdivision level for each continuous part includes a first subdivision level for the luminance encoding units of the continuous part and a second subdivision level for the chrominance encoding units of the continuous part.

[0030] According to another aspect, the first subdivision level and the second subdivision level are different.

[0031] According to another aspect, the method further includes: decoding a flag indicating a partition constraint for rewriting a sequence parameter set associated with a bitstream.

[0032] According to another aspect, the determined subdivision level for each of one or more consecutive portions includes a maximum luma coding unit depth for the region.

[0033] According to another aspect, the determined subdivision level for each of one or more consecutive portions includes a maximum chroma coding unit depth for the corresponding region.

[0034] According to another aspect, the determined subdivision level for one of the consecutive portions is adjusted to maintain an offset relative to the deepest allowed subdivision level decoded for the partition constraint of the bitstream.

[0035] Another aspect of the present disclosure provides a non - transitory computer - readable medium having stored thereon a computer program for implementing a method of decoding a plurality of coding units from a bitstream to produce an image frame, the coding units being the result of a decomposition of coding tree units, the plurality of coding units forming one or more consecutive portions of the bitstream, the method including: determining a subdivision level for each of one or more consecutive portions of the bitstream, each subdivision level being applicable to the coding units of the corresponding consecutive portion of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of the coding tree unit into the coding units of the respective consecutive portions of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier - coded unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to produce the image frame.

[0036] Another aspect of the present disclosure provides a video decoder configured to implement a method of decoding a plurality of coding units from a bitstream to produce an image frame, the coding units being the result of a decomposition of coding tree units, the plurality of coding units forming one or more consecutive portions of the bitstream, the method including: determining a subdivision level for each of one or more consecutive portions of the bitstream, each subdivision level being applicable to the coding units of the corresponding consecutive portion of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of the coding tree unit into the coding units of the respective consecutive portions of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier - coded unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to produce the image frame.

[0037] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory to implement a method of decoding a plurality of coding units from a bitstream to generate an image frame, the coding units being the decomposition result of coding tree units, the plurality of coding units forming one or more consecutive portions of the bitstream, the method comprising: determining a subdivision level for each of the one or more consecutive portions of the bitstream, each subdivision level being applicable to the coding units of the corresponding consecutive portion of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on decomposing a coding tree unit into the coding units of each consecutive portion of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to generate an image frame.

[0038] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] At least one embodiment of the present invention will now be described with reference to the following drawings and appendices, wherein:

[0040] Figure 1 is a schematic block diagram showing a video encoding and decoding system;

[0041] Figure 2A and 2B constitute a schematic block diagram of a general-purpose computer system for one or both of the video encoding and decoding systems that can practice Figure 1 ;

[0042] Figure 3 is a schematic block diagram showing the functional modules of a video encoder;

[0043] Figure 4 is a schematic block diagram showing the functional modules of a video decoder;

[0044] Figure 5 is a schematic block diagram showing the available partitions of a block into one or more blocks in the tree structure of general video coding;

[0045] Figure 6 is a schematic diagram of a data stream for implementing a permitted partition of a block into one or more blocks in the tree structure of general video coding;

[0046] Figure 7A and 7B show an example partition of a coding tree unit (CTU) into a plurality of coding units (CUs);

[0047] Figure 8A ,8B 8C shows the subdivision levels obtained from the splitting in the coding tree and their impact on splitting the coding tree unit into quantization groups;

[0048] Figure 9A and 9B shows the 4×4 transform block scan pattern and the associated primary transform coefficients and secondary transform coefficients;

[0049] Figure 9C and 9D shows the 8×8 transform block scan pattern and the associated primary transform coefficients and secondary transform coefficients;

[0050] Figure 10 shows the regions where the secondary transform is applied for transform blocks of various sizes;

[0051] Figure 11 shows the syntax structure of a bitstream having multiple strips, each strip including multiple coding units;

[0052] Figure 12 shows the syntax structure of a bitstream having a common tree for luminance and chrominance coding blocks of a coding tree unit;

[0053] Figure 13 shows a method of encoding a frame in a bitstream including one or more than one strip as a sequence of coding units;

[0054] Figure 14 shows a method of encoding a strip header in a bitstream;

[0055] Figure 15 shows a method of encoding a coding unit in a bitstream;

[0056] Figure 16 shows a method of decoding a frame from a bitstream that is a sequence of coding units arranged as strips;

[0057] Figure 17 shows a method of decoding a strip header from a bitstream;

[0058] Figure 18 shows a method of decoding a coding unit from a bitstream; and

[0059] Figure 19A and 19B shows the rules for applying or bypassing the secondary transform for the luminance and chrominance channels. DETAILED DESCRIPTION

[0060] In the case of referring to steps and / or features having the same reference numerals in any one or more than one of the figures, unless a contrary intention appears, these steps and / or features have the same function or functions and / or operation or operations for the purposes of this specification.

[0061] Rate control video encoders require the flexibility to adjust quantization parameters at a granularity suitable for block partitioning constraints. The block partitioning constraints may vary between portions of a frame, for example, where multiple video encoders operate in parallel to compress individual frames. The granularity of the regions requiring quantization parameter adjustment varies accordingly. Further, control of the applied transform selection, including the potential application of a secondary transform, is applied within the scope of generating a prediction signal for the residual being transformed. In particular, for intra prediction, separate modes may be used for luma blocks and chroma blocks since different intra prediction modes may be used.

[0062] Some portions of a video contribute more to the fidelity of a rendered viewport than other portions and may be allocated a greater bitrate as well as greater flexibility in terms of variance in block structure and quantization parameters. Portions that contribute little to the fidelity of the rendered viewport, such as those at the sides or back of the rendered view, may be compressed using a simpler block structure to reduce the encoding effort and with less flexibility in terms of control of quantization parameters. Generally, larger values are selected to more coarsely quantize transform coefficients at lower bitrates. Additionally, the application of transform selection may be independent between the luma channel and the chroma channel to further simplify the encoding process by avoiding the need to jointly consider luma and chroma for transform selection. In particular, after separately considering intra prediction modes for luma and chroma, the need to jointly consider luma and chroma for secondary transform selection is avoided.

[0063] Figure 1 is a schematic block diagram showing the functional modules of a video encoding and decoding system 100. The system 100 can vary the regions in which quantization parameters are adjusted in different portions of a frame to accommodate different block partitioning constraints that may be in effect in various portions of the frame.

[0064] The system 100 includes a source device 110 and a destination device 130. A communication channel 120 is used to communicate encoded video information from the source device 110 to the destination device 130. In some configurations, one or both of the source device 110 and the destination device 130 may respectively include a cellular phone handset or a “smartphone,” in which case the communication channel 120 is a wireless channel. In other configurations, the source device 110 and the destination device 130 may include video conferencing equipment, in which case the communication channel 120 is typically a wired channel such as an Internet connection. Additionally, the source device 110 and the destination device 130 may include a wide range of any devices, where these devices include those supporting over-the-air television broadcasts, cable television applications, Internet video applications (including streaming), and applications that capture encoded video data on some computer-readable storage medium such as a hard drive in a file server, etc.

[0065] As Figure 1 shown, source device 110 includes video source 112, video encoder 114, and transmitter 116. Video source 112 generally includes a source of captured video frame data (represented as 113), such as a camera sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote camera sensor. Video source 112 can also be the output of a computer graphics card (e.g., the video output that displays an operating system and various applications executing on a computing device (e.g., a tablet computer)). Examples of source device 110 that can include a camera sensor as video source 112 include smart phones, video camcorders, professional cameras, and network video cameras.

[0066] Video encoder 114 converts (or "encodes") the captured frame data (indicated by arrow 113) from video source 112 into a bitstream (indicated by arrow 115). Bitstream 115 is transmitted by transmitter 116 via communication channel 120 as encoded video data (or "encoded video information"). Bitstream 115 can also be stored in a non-transitory storage device 122 such as a "flash" memory or a hard disk drive until subsequently transmitted via communication channel 120 or as an alternative to transmission via communication channel 120. For example, the encoded video data can be supplied to a customer via a wide area network (WAN) for video streaming applications when needed.

[0067] Destination device 130 includes receiver 132, video decoder 134, and display device 136. Receiver 132 receives the encoded video data from communication channel 120 and passes the received video data as a bitstream (indicated by arrow 133) to video decoder 134. Then, video decoder 134 outputs the decoded frame data (indicated by arrow 135) to display device 136. The decoded frame data 135 has the same chrominance format as frame data 113. Examples of display device 136 include a cathode ray tube, a liquid crystal display (such as in a smart phone, a tablet computer, a computer monitor, or a stand-alone television set, etc.). The functions of source device 110 and destination device 130 can also be embodied in a single device, examples of which include a mobile phone handset and a tablet computer. The decoded frame data can be further transformed before being presented to the user. For example, a "viewport" with specific latitude and longitude can be rendered from the decoded frame data using a projection format to represent a 360° view of a scene.

[0068] Although the above illustrates example devices, source device 110 and destination device 130 can each generally be configured within a general-purpose computer system via a combination of hardware components and software components. Figure 2AA computer system 200 is shown, which includes: a computer module 201; input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 that can be configured as a video source 112, and a microphone 280; and output devices including a printer 215, a display device 214 that can be configured as a display device 136, and a speaker 217. The computer module 201 can communicate with a communication network 220 via a connection line 221 using an external modulator-demodulator (modem) transceiver device 216. The communication network 220, which can represent the communication channel 120, can be a WAN, such as the Internet, a cellular telecommunications network, or a private WAN. In the case where the connection line 221 is a telephone line, the modem 216 can be a traditional "dial-up" modem. Alternatively, in the case where the connection line 221 is a high-capacity (e.g., cable or optical) connection line, the modem 216 can be a broadband modem. A wireless modem can also be used for a wireless connection to the communication network 220. The transceiver device 216 can provide the functions of a transmitter 116 and a receiver 132, and the communication channel 120 can be embodied in the connection line 221.

[0069] The computer module 201 generally includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 can have a semiconductor random access memory (RAM) and a semiconductor read-only memory (ROM). The computer module 201 also includes a plurality of input / output (I / O) interfaces, where the plurality of input / output (I / O) interfaces include: an audio-video interface 207, which is connected to the video display 214, the speaker 217, and the microphone 280; an I / O interface 213, which is connected to the keyboard 202, the mouse 203, the scanner 226, the camera 227, and an optional joystick or other human-machine interface device (not shown); and an interface 208 used by the external modem 216 and the printer 215. The signal from the audio-video interface 207 to the computer monitor 214 is typically the output of a computer graphics card. In some implementations, the modem 216 can be built into the computer module 201, for example, built into the interface 208. The computer module 201 also has a local network interface 211, where the local network interface 211 allows the computer system 200 to be connected to a local communication network 222 known as a local area network (LAN) via a connection line 223. As Figure 2A shown, the local communication network 222 can also be connected to the wide area network 220 via a connection line 224, where the local communication network 222 generally includes a so-called "firewall" device or a device with similar functions. The local network interface 211 can include an Ethernet ( TM ) circuit card, Bluetooth ( TM)Wireless configuration or IEEE 802.11 wireless configuration; however, for interface 211, a variety of other types of interfaces can be practiced. The local network interface 211 can also provide the functions of the transmitter 116 and the receiver 132, and the communication channel 120 can also be embodied in the local communication network 222.

[0070] The I / O interfaces 208 and 213 can provide either or both of serial and parallel connections, where the former is typically implemented according to the Universal Serial Bus (USB) standard and has a corresponding USB connector (not shown). A storage device 209 is provided, and the storage device 209 typically includes a hard disk drive (HDD) 210. Other storage devices such as floppy disk drives and tape drives (not shown) can also be used. An optical disc drive 212 is typically provided to serve as a non-volatile source of data. Portable memory devices such as optical discs (e.g., CD-ROM, DVD, Blu-ray Disc TM ), USB-RAM, portable external hard disk drives, and floppy disks can be used as suitable sources of data for the computer system 200. Typically, any of the HDD 210, optical disc drive 212, networks 220 and 222 can also be configured to operate as the video source 112 or as the destination for decoded video data to be stored for reproduction via the display 214. The source device 110 and the destination device 130 of the system 100 can be embodied in the computer system 200.

[0071] The components 205 - 213 of the computer module 201 typically communicate via the interconnect bus 204 and in a manner that results in the conventional operating modes of the computer system 200 known to those skilled in the relevant art. For example, the processor 205 is connected to the system bus 204 using wiring 218. Similarly, the memory 206 and the optical disc drive 212 are connected to the system bus 204 via wiring 219. Examples of computers that can practice the described configuration include IBM-PC and compatible machines, Sun SPARCstation, Apple Mac TM or similar computer systems.

[0072] Where appropriate or desired, the computer system 200 can be used to implement the video encoder 114 and the video decoder 134 and the methods described below. In particular, the video encoder 114, the video decoder 134, and the methods to be described can be implemented as one or more software applications 233 executable within the computer system 200. In particular, using the instructions 231 (reference Figure 2B) to implement the steps of video encoder 114, video decoder 134, and the method. The software instructions 231 can be formed into one or more code modules each for performing one or more specific tasks. The software can also be split into two separate parts, where the first part and the corresponding code modules perform the method, and the second part and the corresponding code modules manage the user interface between the first part and the user.

[0073] For example, the software can be stored in a computer-readable medium including the storage devices described below. The software is loaded from the computer-readable medium into the computer system 200 and then executed by the computer system 200. A computer-readable medium having such software or a computer program recorded on the computer-readable medium is a computer program product. Using the computer program product in the computer system 200 preferably implements an advantageous device for implementing the video encoder 114, video decoder 134, and the method.

[0074] Generally, the software 233 is stored in the HDD 210 or the memory 206. The software is loaded from the computer-readable medium into the computer system 200 and executed by the computer system 200. Thus, for example, the software 233 can be stored on an optically readable disc storage medium (e.g., CD-ROM) 225 read by the optical disc drive 212.

[0075] In some instances, the application 233 is supplied to the user in a manner encoded on one or more CD-ROMs 225 and read via the corresponding drive 212, or alternatively, the user can read the application 233 from the network 220 or 222. Further, the software can also be loaded into the computer system 200 from other computer-readable media. A computer-readable storage medium refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray Discs TM ) (Blu-ray Disc), hard disk drives, ROMs, or integrated circuits, USB memories, magneto-optical discs, or computer-readable cards such as PCMCIA cards, etc., regardless of whether these devices are inside or outside the computer module 201. Examples of transitory or non-tangible computer-readable transmission media that can also participate in providing software, applications, instructions, and / or video data or encoded video data to the computer module 401 include: radio or infrared transmission channels and network wiring to other computers or networked devices, and the Internet or intranet including email transmissions and information recorded on websites.

[0076] The second part of the above-described application 233 and the corresponding code modules can be executed to implement one or more graphical user interfaces (GUIs) to be drawn or otherwise presented on the display 214. By typically operating the keyboard 202 and the mouse 203, the user and the application of the computer system 200 can operate the interface in a functionally applicable manner to provide control commands and / or inputs to the application associated with these one or more GUIs. Other functionally applicable forms of user interfaces can also be implemented, such as an audio interface that utilizes voice prompts output via the speaker 217 and user voice commands input via the microphone 280, etc.

[0077] Figure 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents Figure 2A the logical aggregation of all memory modules accessible to the computer module 201 in (including the HDD 209 and the semiconductor memory 206).

[0078] When initially powering on the computer module 201, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in Figure 2A the ROM 249 of the semiconductor memory 206. Sometimes a hardware device such as the ROM 249 storing software is referred to as firmware. The POST program 250 checks the hardware within the computer module 201 to ensure proper operation, and typically checks the processor 205, the memory 234 (209, 206), and the basic input-output system software (BIOS) module 251 that is also typically stored in the ROM 249 for correct operation. Once the POST program 250 runs successfully, the BIOS 251 starts Figure 2A the hard disk drive 210. Starting the hard disk drive 210 causes the boot loader 252 resident on the hard disk drive 210 to be executed via the processor 205. In this way, the operating system 253 is loaded into the RAM memory 206, where the operating system 253 starts to operate. The operating system 253 is a system-level application executable by the processor 205 to implement various high-level functions including processor management, memory management, device management, storage management, software application interfaces, and a general user interface, etc.

[0079] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory to execute without conflicting with the memory allocated to other processes. In addition, Figure 2Athe different types of memory available in the computer system 200 so that each process can run efficiently. Thus, the aggregated memory 234 is not intended to illustrate how to allocate specific segments of memory (unless otherwise stated), but rather provides an overview of the memory accessible by the computer system 200 and how that memory is used.

[0080] As Figure 2B shown, the processor 205 includes a plurality of functional modules, where the plurality of functional modules include a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a plurality of storage registers 244 - 246 in a register section. One or more internal buses 241 functionally interconnect these functional modules. The processor 205 generally also has one or more interfaces 242 for communicating with external devices via the system bus 204 using wiring 218. The memory 234 is connected to the bus 204 using wiring 219.

[0081] The application program 233 includes an instruction sequence 231 that can include conditional branch instructions and loop instructions. The program 233 can also include data 232 used when executing the program 233. The instructions 231 and data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237 respectively. Depending on the relative sizes of the instructions 231 and the memory locations 228 - 230, a particular instruction can be stored in a single memory location as described by the instruction shown in memory location 230. Optionally, as described by the instruction segments shown in memory locations 228 and 229, an instruction can be split into multiple parts each stored in separate memory locations.

[0082] Typically, a set of instructions is given to the processor 205, and that set of instructions is executed within the processor 205. The processor 205 waits for subsequent input, where the processor 205 reacts to the subsequent input by executing another set of instructions. The input can be provided from one or more of a plurality of sources, where the input includes data generated by one or more of the input devices 202, 203, data received from an external source via one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into the corresponding reader 212 (all of which are shown in Figure 2A ). Executing a set of instructions can in some cases result in output data. Execution may also involve storing data or variables to the memory 234.

[0083] The video encoder 114, the video decoder 134, and the method may use the input variables 254 stored in the respective memory locations 255, 256, 257 within the memory 234. The video encoder 114, the video decoder 134, and the method generate the output variables 261 stored in the respective memory locations 262, 263, 264 within the memory 234. Intermediate variables 258 may be stored in the memory locations 259, 260, 266, and 267.

[0084] Reference Figure 2B The processor 205, registers 244, 245, 246, the arithmetic logic unit (ALU) 240, and the control unit 239 work together to perform a sequence of micro-operations, where these micro-operation sequences are required for the "fetch, decode, and execute" cycles for each instruction in the instruction set that constitutes the program 233. Each fetch, decode, and execute cycle includes:

[0085] A fetch operation for fetching or reading an instruction 231 from the memory locations 228, 229, 230;

[0086] A decode operation in which the control unit 239 determines which instruction has been fetched; and

[0087] An execute operation in which the control unit 239 and / or the ALU 240 execute the instruction.

[0088] After that, a further fetch, decode, and execute cycle for the next instruction can be performed. Similarly, a store cycle can be performed, in which the control unit 239 stores or writes a value to the memory location 232.

[0089] To illustrate Figures 13 to 18 Each step or sub-process in the method to be illustrated is associated with one or more segments of the program 233, and is typically performed by the register section 244, 245, 247, the ALU 240, and the control unit 239 in the processor 205 working together to perform the fetch, decode, and execute cycles for each instruction in the segmented instruction set of the program 233.

[0090] Figure 3 is a schematic block diagram showing the functional modules of the video encoder 114. Figure 4 is a schematic block diagram showing the functional modules of the video decoder 134. Generally, data is transferred between the functional modules within the video encoder 114 and the video decoder 134 in groups of samples or coefficients (such as the division of a block into fixed-size sub-blocks, etc.) or as an array. As Figure 2A and 2BAs shown, a general - purpose computer system 200 can be used to implement the video encoder 114 and the video decoder 134. Various functional modules can be implemented by using the dedicated hardware within the computer system 200, by using software executable within the computer system 200 (such as one or more software code modules of a software application 233 residing on the hard - disk drive 205 and controlled by the processor 205 for its execution, etc.). Alternatively, the video encoder 114 and the video decoder 134 can be implemented by using a combination of dedicated hardware and software executable within the computer system 200. The video encoder 114, the video decoder 134, and the method can be alternatively implemented in dedicated hardware such as one or more integrated circuits performing the functions or sub - functions of the method. Such dedicated hardware can include a graphics processing unit (GPU), a digital signal processor (DSP), an application - specific standard product (ASSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or one or more microprocessors and associated memories. In particular, the video encoder 114 includes modules 310 - 390, and the video decoder 134 includes modules 420 - 496, where each of these modules can be implemented as one or more software code modules of the software application 233.

[0091] Although Figure 3 the video encoder 114 is an example of a general - purpose video coding (VVC) video - coding pipeline, other video codecs can also be used for the processing stages described herein. The video encoder 114 receives captured frame data 113 such as a series of frames (each frame including one or more color channels). The frame data 113 can be in any chroma format, such as 4:0:0, 4:2:0, 4:2:2, or 4:4:4 chroma format. The block partitioner 310 first divides the frame data 113 into CTUs. The CTUs are generally square - shaped and are configured to use a specific size of CTU. For example, the size of the CTU can be 64×64, 128×128, or 256×256 luma samples. The block partitioner 310 further divides each CTU into one or more CUs corresponding to a luma - coding tree or a chroma - coding tree. The luma channel can also be referred to as the primary color channel. Each chroma channel can also be referred to as a secondary color channel. The CUs have various sizes and can include both square and non - square aspect ratios. Refer to Figures 13 - 15 for a further description of the operation of the block partitioner 310. However, in the VVC standard, the CUs, PUs, and TUs always have side lengths that are powers of 2. Thus, the current CU (denoted as 312) is output from the block partitioner 310, advancing according to the iteration of one or more blocks of the CTU, according to the luma - coding tree and chroma - coding tree of the CTU. Refer to the following Figure 5 and 6To further illustrate the options for partitioning a CTU into CBs. Although operations are generally described in terms of CTUs, the video encoder 114 and the video decoder 134 can operate on smaller-sized regions to reduce memory consumption. For example, each CTU can be divided into smaller regions, called "virtual pipeline data units" (VPDUs) of size 64×64. VPDUs form a data granularity that is more suitable for pipelining in a hardware architecture, where the reduced memory footprint compared to operating on a complete CTU reduces the silicon area and thus the cost.

[0092] CTUs obtained from the first partitioning of the frame data 113 can be scanned in raster scan order and can be grouped into one or more than one "slice". A slice can be an "intra" (or "I") slice. An intra slice (I slice) indicates that each CU in the slice is intra-predicted. Optionally, a slice can be single-predicted or bi-predicted (a "P" or "B" slice, respectively), indicating the additional availability of single prediction and bi-prediction in the slice, respectively.

[0093] In an I slice, the coding tree of each CTU can diverge into two separate coding trees below the 64×64 level, one for luminance and the other for chrominance. Using separate trees allows different block structures to exist between luminance and chrominance within the 64×64 luminance region of a CTU. For example, large chrominance CBs can be juxtaposed with many smaller luminance CBs, and vice versa. In a P or B slice, a single coding tree of a CTU defines the block structure common to luminance and chrominance. The resulting blocks of the single tree can be intra-predicted or inter-predicted.

[0094] For each CTU, the video encoder 114 operates in two stages. In the first stage (called the "search" stage), the block partitioner 310 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select a CB that provides relatively high compression efficiency and relatively low distortion. This test typically involves Lagrangian optimization, whereby candidate CBs are evaluated based on a weighted combination of rate (coding cost) and distortion (error with respect to the input frame data 113). The "best" candidate CB (the CB with the lowest evaluated rate / distortion) is selected for subsequent encoding in the bitstream 115. Options included in the evaluation of candidate CBs are: using a CB for a given region, or splitting the region according to various splitting options and encoding each smaller resulting region or further splitting the region using other CBs. As a result, both the coding tree and the CB itself are selected in the search stage.

[0095] Video encoder 114 generates a predicted block (PB) indicated by arrow 320 for each CB (e.g., CB 312). PB 320 is a prediction of the content of the associated CB 312. Subtractor module 322 generates a difference (or "residual", which refers to the difference in the spatial domain) represented as 324 between PB 320 and CB 312. The difference 324 is the block - size difference between the corresponding samples in PB 320 and CB 312. The difference 324 is transformed, quantized, and represented as a transformed block (TB) indicated by arrow 336. PB 320 and the associated TB 336 are typically selected from among multiple possible candidate CBs (e.g., based on the evaluated cost or distortion).

[0096] A candidate coded block (CB) is a CB obtained from one of the prediction modes available to video encoder 114 for the associated PB and the resulting residual. When combined with the predicted PB in video decoder 114, TB 336 reduces the difference between the decoded CB and the original CB 312 at the cost of additional signaling in the bitstream.

[0097] Thus, each candidate coded block (CB) (i.e., the combination of a predicted block (PB) and a transformed block (TB)) has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of a CB is typically estimated as a difference in sample values, such as sum of absolute differences (SAD) or sum of squared differences (SSD). Pattern selector 386 can use difference 324 to determine an estimate obtained from each candidate PB to determine prediction mode 387. Prediction mode 387 indicates the decision to use a particular prediction mode (e.g., intra - prediction or inter - prediction) for the current CB. The estimation of the coding cost associated with each candidate prediction mode and the corresponding residual coding can be performed at a cost significantly lower than that of the entropy coding of the residuals. Thus, even in a real - time video encoder, multiple candidate modes can be evaluated to determine the best mode in terms of rate - distortion.

[0098] Determining the best mode in terms of rate - distortion is typically achieved using a variant of Lagrangian optimization.

[0099] A Lagrangian or similar optimization process can be employed for both the selection of the best partitioning of a CTU into CBs (using block partitioner 310) and the selection of the best prediction mode from multiple possibilities. By applying the Lagrangian optimization process for candidate modes in pattern selector module 386, the intra - prediction mode with the lowest cost measurement is selected as the best mode. The lowest - cost mode is the selected quadratic transform index 388 and is also encoded in bitstream 115 by entropy encoder 338.

[0100] In a second phase of operation of video encoder 114, referred to as the "encoding" phase, an iteration of the (one or more than one) determined coding trees for each CTU is performed in video encoder 114. For a CTU using a separate tree, for each 64×64 luminance region of the CTU, the luminance coding tree is first encoded, followed by the chrominance coding tree. Only luminance CBs are encoded within the luminance coding tree, and only chrominance CBs are encoded within the chrominance coding tree. For a CTU using a common tree, a single tree describes the CUs according to the common block structure of the common tree, i.e., luminance CBs and chrominance CBs.

[0101] Entropy encoder 338 supports both variable length coding of syntax elements and arithmetic coding of syntax elements. Parts of the bitstream such as "parameter sets" (e.g., sequence parameter set (SPS) and picture parameter set (PPS)) use a combination of fixed length codewords and variable length codewords. A slice (also referred to as a consecutive part) has a slice header using variable length coding, followed by slice data using arithmetic coding. The slice header defines parameters specific to the current slice, such as slice-level quantization parameter offset, etc. The slice data includes the syntax elements of each CTU in the slice. Using variable length coding and arithmetic coding requires sequential parsing within each part of the bitstream. These parts can be described with start codes to form "network abstraction layer units" or "NAL units". Context adaptive binary arithmetic coding processing is used to support arithmetic coding. The syntax elements for arithmetic coding consist of a sequence of one or more than one "bin (binary file)". Like bits, the value of a bin is either "0" or "1". However, bins are not encoded as discrete bits in bitstream 115. A bin has an associated prediction (or "likely" or "most probable") value and an associated probability (referred to as "context"). When the actual bin to be encoded matches the prediction value, the "most probable symbol" (MPS) is encoded. Encoding the most probable symbol is relatively inexpensive in terms of the bits consumed in bitstream 115 (including a cost of less than one discrete bit in total). When the actual bin to be encoded does not match the likely value, the "least probable symbol" (LPS) is encoded. Encoding the least probable symbol has a relatively high cost in terms of bits consumed. The bin coding technique enables efficient coding of bins where the probability of "0" vs "1" is skewed. For a syntax element with two possible values (i.e., "flag"), a single bin is sufficient. For a syntax element with many possible values, a sequence of bins is required.

[0102] The existence of a later bin in the sequence can be determined based on the value of an earlier bin in the sequence. Additionally, each bin can be associated with more than one context. A particular context can be selected based on an earlier bin in a syntactic element and the bin values of adjacent syntactic elements (i.e., bin values from adjacent blocks), etc. Each time a context-coded bin is coded, the context selected for that bin (if any) is updated in a way that reflects the new bin value. Thus, the binary arithmetic coding scheme is considered adaptive.

[0103] Video encoder 114 also supports bins that lack context ("bypass bins"). The bypass bins are coded assuming an equiprobable distribution between "0" and "1". Thus, each bin has a coding cost of one bit in bitstream 115. The lack of context saves memory and reduces complexity, and thus bypass bins are used where the distribution of the values of a particular bin is not skewed. An example of an entropy encoder that uses context and is adaptive is known in the art as CABAC (Context-Adaptive Binary Arithmetic Coder), and many variants of this encoder have been adopted in video coding.

[0104] Entropy encoder 338 uses a combination of context-coded bins and bypass-coded bins to code quantization parameter 392 and, if for the current CB, codes the LFNST index 388. Quantization parameter 392 is coded using "delta QP". In each region called a "quantization group", the delta QP is signaled at most once. Quantization parameter 392 is applied to the residual coefficients of the luminance CB. The adjusted quantization parameter is applied to the residual coefficients of the co-located chrominance CB. The adjusted quantization parameter can include a mapping from luminance quantization parameter 392 according to a mapping table and a CU-level offset selected from an offset list. The secondary transform index 388 is signaled when the residual associated with a transform block includes valid residual coefficients only in those coefficient positions that are transformed into primary coefficients by applying a secondary transform.

[0105] The multiplexer module 384 outputs the PB 320 from the intra prediction module 364 according to the determined best intra prediction mode selected from the test prediction modes of the respective candidate CBs. The candidate prediction modes need not include every conceivable prediction mode supported by the video encoder 114. Intra prediction is divided into three types. "DC intra prediction" involves filling the PB with a single value representing the average of nearby reconstructed samples. "Planar intra prediction" involves filling the PB with samples according to a plane, where the DC offset and the vertical and horizontal gradients are derived from nearby reconstructed neighboring samples. The nearby reconstructed samples typically include a row of reconstructed samples above the current PB (extending a certain extent to the right of the PB) and a column of reconstructed samples to the left of the current PB (extending a certain extent downward outside the PB). "Angular intra prediction" involves filling the PB with reconstructed neighboring samples that are filtered and propagated across the PB in a particular direction (or "angle"). In VVC, 65 angles are supported, where rectangular blocks can utilize additional angles not available to square blocks to produce a total of 87 angles. A fourth type of intra prediction can be used for chrominance PBs, thus generating the PB from collocated luma reconstructed samples according to the "cross-component linear model" (CCLM) mode. Three different CCLM modes are available, each mode using a different model derived from neighboring luma and chroma samples. The derived model is used to generate a sample block for the chrominance PB from collocated luma samples.

[0106] In cases where previously reconstructed samples are not available (e.g., at the edges of the frame), a default half-tone value of half the sample range is used. For example, for 10-bit video, a value of 512 is used. Since no previous samples are available for the CB located at the upper left position of the frame, the angular and planar intra prediction modes produce the same output as the DC prediction mode, i.e., a flat plane of samples with the half-tone value as the amplitude.

[0107] For inter-frame prediction, the motion compensation module 380 uses samples from one or two frames before the current frame in the order of the coded frames in the bitstream to generate a prediction block 382 and outputs it as PB 320 by the multiplexer module 384. In addition, for inter-frame prediction, a single coding tree is typically used for both the luminance channel and the chrominance channel. The order of the coded frames in the bitstream may be different from the order of the frames when captured or displayed. When one frame is used for prediction, the block is referred to as "single prediction" and has two associated motion vectors. When two frames are used for prediction, the block is referred to as "dual prediction" and has two associated motion vectors. For P slices, each CU can be intra-frame predicted or single predicted. For B slices, each CU can be intra-frame predicted, single predicted, or dual predicted. A "group of pictures" structure is typically used to code the frames, thus implementing a temporal hierarchy of the frames. A frame can be divided into multiple slices, each slice coding a part of the frame. The temporal hierarchy of the frames allows a frame to reference previous and subsequent pictures in the order of the displayed frames. The images are coded in an order necessary to ensure that the dependencies of each frame are satisfied during decoding.

[0108] Samples are selected according to the motion vector 378 and the reference picture index. The motion vector 378 and the reference picture index apply to all color channels, and thus inter-frame prediction is mainly described in terms of the operation on PUs rather than PBs, i.e., using a single coding tree to describe the decomposition of each CTU into one or more inter-frame prediction blocks. The inter-frame prediction method may vary in the number and accuracy of the motion parameters. The motion parameters typically include a reference frame index (which indicates which reference frames from the reference frame lists will be used plus the respective spatial translations of the reference frames), but may include more frames, special frames, or complex affine parameters such as scaling and rotation. Additionally, a predetermined motion refinement process can be applied to generate a dense motion estimate based on the reference sample block.

[0109] When PB 320 is determined and selected and subtracted from the original sample block at subtractor 322, the residual with the lowest encoding cost (denoted as 324) is obtained and lossily compressed. The lossy compression process includes steps of transformation, quantization, and entropy coding. Forward primary transform module 326 applies a forward transform to difference 324, thereby converting difference 324 from the spatial domain to the frequency domain and generating primary transform coefficients represented by arrow 328. The maximum primary transform size in one dimension is a 32-point DCT-2 or a 64-point DCT-2 transform. If the CB being encoded is larger than the maximum supported primary transform size represented as the block size (i.e., 64×64 or 32×32), the primary transform 326 is applied in a block manner to transform all samples of difference 324. Application of transform 326 results in multiple TBs of the CB. In cases where each transform is operating on a TB of difference 324 larger than 32×32 (e.g., 64×64), all resulting primary transform coefficients 328 outside the upper left 32×32 region of the TB are set to zero, i.e., discarded. The remaining primary transform coefficients 328 are passed to quantizer module 334. The primary transform coefficients 328 are quantized according to quantization parameter 392 associated with the CB to produce primary transform coefficients 332. The quantization parameter 392 can be different for the luminance CB relative to each chrominance CB. The primary transform coefficients 332 are passed to forward secondary transform module 330 to produce transform coefficients represented by arrow 336 by performing a non-separable second transform (NSST) operation or bypassing the second transform. The forward primary transform is typically separable, transforming a set of rows of each TB and then a set of columns. For luminance TBs with a width and height not exceeding 16 samples, forward primary transform module 326 uses a type-II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or bypasses the transform in the horizontal and vertical directions, or uses a combination of a type-VII discrete sine transform (DST-7) and a type-VIII discrete cosine transform (DCT-8) in the horizontal or vertical direction. The use of the combination of DST-7 and DCT-8 is referred to as the "multiple transform selection set" (MTS) in the VVC standard.

[0110] The forward secondary transform of module 330 is typically a non-separable transform that is only applied to the residual of an intra-predicted CU and can still be bypassed. The forward secondary transform operates on 16 samples (arranged as the upper left 4×4 sub-block of the primary transform coefficients 328) or 48 samples (arranged as three 4×4 sub-blocks in the upper left 8×8 coefficients of the primary transform coefficients 328) to produce a set of secondary transform coefficients. The number of sets of secondary transform coefficients can be less than the number of sets of primary transform coefficients from which it is derived. Since the secondary transform is only applied to sets of coefficients that are adjacent to each other and include the DC coefficient, the secondary transform is referred to as the "low-frequency non-separable secondary transform" (LFNST). Additionally, when LFNST is applied, all remaining coefficients in the TB must be zero in both the primary transform domain and the secondary transform domain.

[0111] The quantization parameter 392 is constant for a given TB and thus results in a uniform scaling of the residual coefficients generated in the primary transform domain of the TB. The quantization parameter 392 can be varied periodically by a signaled "delta quantization parameter". For a CU contained within a given region (referred to as a "quantization group"), the delta quantization parameter (delta QP) is signaled once. If the CU is larger than the quantization group size, the delta QP is signaled once by one of the TBs of the CU. That is, for the first quantization group of the CU, the entropy encoder 338 signals the delta QP once, and for any subsequent quantization groups of the CU, the delta QP is not signaled. Non-uniform scaling is also possible by applying a "quantization matrix", whereby the scaling factors applied to the individual residual coefficients result from the combination of the quantization parameter 392 and the corresponding entries in the scaling matrix. The scaling matrix can have a size smaller than the size of the TB, and when applied to the TB, a nearest neighbor method is used to provide scaling values for the individual residual coefficients based on the scaling matrix having a size smaller than the TB size. The residual coefficients 336 are supplied to the entropy encoder 338 for encoding in the bitstream 115. Generally, according to a scan pattern, the residual coefficients of each TB of the TU having at least one valid residual coefficient are scanned to produce an ordered list of values. The scan pattern typically scans the TBs as a sequence of 4×4 "sub-blocks", providing a regular scan operation at the granularity of 4×4 groups of residual coefficients, where the arrangement of the sub-blocks depends on the size of the TB. The scan within each sub-block and the progression from one sub-block to the next generally follow a backward diagonal scan pattern. Additionally, the quantization parameter 392 is encoded in the bitstream 115 using the delta QP syntax element, and the secondary transform index 388 is encoded in the bitstream 115 under the conditions described in Figures 13 to 15 the reference

[0112] As described above, the video encoder 114 needs to access a frame representation corresponding to the encoded frame representation seen in the video decoder 134. Thus, the residual coefficients 336 are inverse-transformed by the inverse secondary transform module 344 (operating based on the secondary transform coefficients 388) to produce intermediate inverse transform coefficients represented by arrow 342. The intermediate inverse transform coefficients 346 are inverse-quantized by the dequantization module 340 according to the quantization parameter 392 to produce residual samples represented by arrow 346. The intermediate inverse transform coefficients 346 are passed to the inverse primary transform module 348 to produce residual samples of the TU represented by arrow 350. The type of inverse transform performed by the inverse secondary transform module 344 corresponds to the type of forward transform performed by the forward secondary transform module 330. The type of inverse transform performed by the inverse primary transform module 348 corresponds to the type of primary transform performed by the primary transform module 326. The summation module 352 adds the residual samples 350 and the PU 320 to produce the reconstructed samples of the CU (indicated by arrow 354).

[0113] The reconstructed sample 354 is passed to the reference sample cache 356 and the in-loop filter module 368. The reference sample cache 356, typically implemented using static RAM on an ASIC (thus avoiding expensive off-chip memory access), provides the minimum sample storage required to satisfy the dependencies for generating intra-PBs for subsequent CUs in a frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of a row of CTUs for use by the next row of CTUs and a column buffer whose extent is set by the height of the CTU. The reference sample cache 356 supplies reference samples (indicated by arrow 358) to the reference sample filter 360. The sample filter 360 applies a smoothing operation to produce filtered reference samples (indicated by arrow 362). The filtered reference samples 362 are used by the intra prediction module 364 to produce an intra prediction block of samples indicated by arrow 366. For each candidate intra prediction mode, the intra prediction module 364 produces a sample block, i.e., 366. The sample block 366 is generated by the module 364 using techniques such as DC, planar, or angular intra prediction.

[0114] The in-loop filter module 368 applies several filtering stages to the reconstructed sample 354. The filtering stages include a "deblocking filter" (DBF) that applies smoothing aligned with CU boundaries to reduce artifacts due to discontinuities. Another filtering stage present in the in-loop filter module 368 is an "adaptive loop filter" (ALF) that applies a Wiener-based adaptive filter to further reduce distortion. Another available filtering stage in the in-loop filter module 368 is a "sample adaptive offset" (SAO) filter. The SAO filter works by first classifying the reconstructed samples into one or more classes and applying an offset at the sample level according to the assigned class.

[0115] Filtered samples indicated by arrow 370 are output from the in-loop filter module 368. The filtered samples 370 are stored in the frame buffer 372. The frame buffer 372 typically has the capacity to store several (e.g., up to 16) pictures and is thus stored in the memory 206. Due to the large memory consumption required, the frame buffer 372 typically does not use on-chip memory for storage. As such, access to the frame buffer 372 is expensive in terms of memory bandwidth. The frame buffer 372 supplies a reference frame (indicated by arrow 374) to the motion estimation module 376 and the motion compensation module 380.

[0116] The motion estimation module 376 estimates a plurality of "motion vectors", denoted as 378, each of which is a Cartesian space offset relative to the position of the current CB, for a block in one of the reference frames in the reference frame buffer 372. A filtering block (denoted as 382) generates reference samples for each motion vector. The filtered reference samples 382 form further candidate modes for potential selection by the mode selector 386. Additionally, for a given CU, the PB 320 can be formed using one reference block ("single prediction") or can be formed using two reference blocks ("dual prediction"). For the selected motion vectors, the motion compensation module 380 generates the PU 320 according to a filtering process that supports sub-pixel accuracy in the motion vectors. Thus, the motion estimation module 376 (which operates on many candidate motion vectors) can perform a simplified filtering process compared to the motion compensation module 380 (which operates only on the selected candidates) to achieve reduced computational complexity. When the video encoder 114 selects inter prediction for a CU, the motion vectors 378 are encoded in the bitstream 115.

[0117] Although the video encoder 114 of Figure 3 is described with reference to Versatile Video Coding (VVC), other video coding standards or implementations may also employ the processing stages of modules 310 - 390. The frame data 113 (and the bitstream 115) can also be read from (or written to) the memory 206, hard disk drive 210, CD-ROM, Blu-ray diskTM, or other computer-readable storage media. Additionally, the frame data 113 (and the bitstream 115) can be received from (or sent to) an external source such as a server connected to the communication network 220 or a radio frequency receiver, etc. The communication network 220 may provide limited bandwidth, thus requiring rate control to be used in the video encoder 114 to avoid saturating the network when the frame data 113 is difficult to compress. Further, the bitstream 115 can be constructed from one or more than one strip representing a spatial portion (set of CTUs) of the frame data 113, which are generated by one or more than one instance of the video encoder 114 and operate in a coordinated manner under the control of the processor 205. In the context of the present invention, a strip may also be referred to as a "continuous portion" of the bitstream. The strips are continuous within the bitstream and can be encoded or decoded as separate parts (e.g., if parallel processing is being used).

[0118] In Figure 4 is shown the video decoder 134. Although Figure 4 the video decoder 134 of Figure 4As shown, the bitstream 133 is input to the video decoder 134. The bitstream 133 can be read from the memory 206, hard disk drive 210, CD-ROM, Blu-ray disc, or other non-transitory computer-readable storage medium. Alternatively, the bitstream 133 can be received from an external source (such as a server connected to the communication network 220 or a radio frequency receiver, etc.). The bitstream 133 contains encoded syntax elements representing the captured frame data to be decoded.

[0119] The bitstream 133 is input to the entropy decoder module 420. The entropy decoder module 420 extracts syntax elements from the bitstream 133 by decoding the "bin" sequence and passes the values of the syntax elements to other modules in the video decoder 134. The entropy decoder module 420 uses variable-length and fixed-length decoding to decode the SPS, PPS, or slice header, and uses an arithmetic decoding engine to decode the syntax elements of the slice data into a sequence of one or more bins. Each bin can use one or more "contexts", where the context describes the probability levels used to encode the "one" and "zero" values for the bin. In cases where multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the available contexts to decode the bin. The process of decoding the bins forms a sequential feedback loop, so that each slice can be decoded in its entirety by a given instance of the entropy decoder 420. A single (or a few) high-performance entropy decoder 420 instances can decode all the slices of a frame from the bitstream 115, and multiple low-performance entropy decoder 420 instances can decode the slices of a frame from the bitstream 133 simultaneously.

[0120] The entropy decoder module 420 applies an arithmetic coding algorithm, such as "Context-Adaptive Binary Arithmetic Coding" (CABAC), to decode the syntax elements from the bitstream 133. The decoded syntax elements are used to reconstruct the parameters within the video decoder 134. The parameters include residual coefficients (represented by the arrow 424), quantization parameter 474, quadratic transform index 470, and mode selection information such as the intra prediction mode (represented by the arrow 458). The mode selection information also includes information such as motion vectors, and the partitioning of each CTU into one or more CBs. The parameters are used to generally generate the PB in combination with the sample data from the previously decoded CB.

[0121] The residual coefficients 424 are passed to the inverse quadratic transform module 436, where according to the reference Figures 16 to 18The described method applies a second transformation or bypasses the operation. The inverse second transformation module 436 generates the reconstructed transform coefficients 432, i.e., the primary transform domain coefficients, from the second transform domain coefficients. The reconstructed transform coefficients 432 are input to the dequantizer module 428. The dequantizer module 428 inverse quantizes (or "scales") the residual coefficients 432, i.e., in the primary transform coefficient domain, to create the reconstructed intermediate transform coefficients represented by arrow 440 according to the quantization parameter 474. If it is indicated in the bitstream 133 to use a non-uniform inverse quantization matrix, the video decoder 134 reads the quantization matrix as a sequence of scaling factors from the bitstream 133 and arranges the scaling factors into a matrix according to. The inverse scaling uses the quantization matrix in combination with the quantization parameter to create the reconstructed intermediate transform coefficients 440.

[0122] The reconstructed transform coefficients 440 are passed to the inverse primary transform module 444. Module 444 transforms the coefficients 440 back from the frequency domain to the spatial domain. The result of the operation of module 444 is a block of residual samples represented by arrow 448. The block of residual samples 448 is equal in size to the corresponding CB. The block of residual samples 448 is supplied to the summing module 450. At the summing module 450, the residual samples 448 are added to the decoded PB represented as 452 to produce a block of reconstructed samples represented by arrow 456. The reconstructed samples 456 are supplied to the reconstructed sample cache 460 and the in-loop filter module 488. The in-loop filter module 488 produces a reconstructed block of frame samples represented as 492. The frame samples 492 are written to the frame buffer 496.

[0123] The reconstructed sample cache 460 operates in a manner similar to the reconstructed sample cache 356 of the video encoder 114. The reconstructed sample cache 460 provides storage for the reconstructed samples required for intra prediction of subsequent CBs in the absence of the memory 206 (e.g., by using data 232 which is typically on-chip memory instead). The reference samples represented by arrow 464 are obtained from the reconstructed sample cache 460 and are supplied to the reference sample filter 468 to produce the filtered reference samples represented by arrow 472. The filtered reference samples 472 are supplied to the intra prediction module 476. Module 476 generates a block of intra prediction samples represented by arrow 480 according to the intra prediction mode parameter 458 represented in the bitstream 133 and decoded by the entropy decoder 420. Modes such as DC, planar, or angular intra prediction are used to generate the block of samples 480.

[0124] When the prediction mode of the CB is indicated as intra prediction in the bitstream 133, the intra prediction samples 480 form the decoded PB 452 via the multiplexer module 484. Intra prediction produces a predicted block (PB) of samples, i.e., a block in one color component derived using "neighboring samples" in the same color component. Neighboring samples are samples adjacent to the current block and have been reconstructed because they are earlier in the block decoding order. In the case of luma and chroma block juxtaposition, the luma and chroma blocks may use different intra prediction modes. However, the two chroma channels share the same intra prediction mode.

[0125] When the prediction mode of the CB is indicated as intra prediction in the bitstream 133, the motion compensation module 434 selects and filters a block of samples 498 from the frame buffer 496 using the motion vectors (decoded from the bitstream 133 by the entropy decoder 420) and the reference frame index to produce a block of inter prediction samples represented as 438. The block of samples 498 is obtained from the previously decoded frames stored in the frame buffer 496. For dual prediction, two blocks of samples are generated and mixed together to produce the samples of the decoded PB 452. The frame buffer 496 is filled with the filtered block data 492 from the in-loop filter module 488. Similar to the in-loop filter module 368 of the video encoder 114, the in-loop filter module 488 applies any of the DBF, ALF, and SAO filtering operations. Generally, the motion vectors are applied to both the luma and chroma channels, but the filtering processes for subsample interpolation are different in the luma and chroma channels.

[0126] Figure 5 is a schematic block diagram showing a set 500 of available partitions or splits of a region in the tree structure of general video coding into one or more sub-regions. As referred to Figure 3 as described, the partitions shown in the set 500 are available for the block partitioner 310 of the encoder 114 to partition each CTU into one or more CUs or CBs according to the coding cost determined by Lagrangian optimization.

[0127] Although the set 500 only shows the partitioning of a square region into other potentially non-square sub-regions, it should be understood that the set 500 is showing the potential partitioning of a parent node in the coding tree into child nodes in the coding tree, and it is not required that the parent node corresponds to a square region. If the containing region is non-square, the sizes of the blocks resulting from the partition are scaled according to the aspect ratio of the containing block. Once a region is not further split, i.e., at the leaf node of the coding tree, the CU occupies that region.

[0128] The process of sub - dividing a region into sub - regions must terminate when the resulting sub - regions reach the minimum CU size (usually 4×4 luma samples). In addition to constraining the CU to prohibit the block region from being less than a predetermined minimum size, e.g., 16 samples, the CU is constrained to have a minimum width or height of four. Other minimum values are possible in terms of both width and height or in terms of width or height alone. The sub - division process can also terminate before the deepest level of decomposition, resulting in a CU larger than the minimum CU size. It is possible that no splitting occurs, resulting in a single CU that occupies the entire CTU. A single CU that occupies the entire CTU is the largest available coding unit size. Due to the use of subsampled chroma formats (such as 4:2:0, etc.), the arrangement of video encoder 114 and video decoder 134 can terminate the splitting of regions in the chroma channel earlier than in the luma channel, including the case of a common coding tree that defines the block structure of both luma and chroma channels. When separate coding trees are used for luma and chroma, the constraints on available splitting operations ensure a minimum chroma CB region of 16 samples, even if such a CB is juxtaposed with a larger luma region (e.g., 64 luma samples).

[0129] In the absence of further sub - division, there is a CU at the leaf node of the coding tree. For example, leaf node 510 contains a CU. At non - leaf nodes of the coding tree, there is a split into two or more other nodes, where each node can be a leaf node forming a CU or a non - leaf node containing a further split into smaller regions. At each leaf node of the coding tree, there is an encoded block for each color channel. A split that terminates at the same depth for both luma and chroma results in three juxtaposed CBs. A split that terminates at a deeper depth for luma than for chroma results in multiple luma CBs juxtaposed with the CBs of the chroma channel.

[0130] As Figure 5 shown, the quadtree split 512 divides the containing region into four regions of equal size. Compared to HEVC, Versatile Video Coding (VVC) achieves additional flexibility through additional splits, including horizontal binary split 514 and vertical binary split 516. Each of splits 514 and 516 divides the containing region into two regions of equal size. The division is along the horizontal boundary (514) or vertical boundary (516) within the containing block.

[0131] In general video coding, further flexibility is achieved by adding ternary horizontal splits 518 and ternary vertical splits 520. The ternary splits 518 and 520 divide a block into three regions that form boundaries along 1 / 4 and 3 / 4 of the width or height of the containing region in the horizontal direction (518) or the vertical direction (520). The combination of quadtree, binary tree, and ternary tree is referred to as "QTBTTT". The root of the tree includes zero or more quadtree splits (the "QT" part of the tree). Once the QT part terminates, zero or more binary or ternary splits (the "multi-tree" or the "MT" part of the tree) can occur, ultimately ending in a CB or CU at the leaf nodes of the tree. When the tree describes all color channels, the leaf node of the tree is a CU. When the tree describes the luminance channel or a chrominance channel, the leaf node of the tree is a CB.

[0132] Compared with HEVC which only supports quadtree and thus only supports square blocks, QTBTTT gets more possible CU sizes especially considering the possible recursive application of binary tree and / or ternary tree splits. When only quadtree splits are available, each increase in the coding tree depth corresponds to a reduction of the CU size to one quarter of the size of the parent region. In VVC, the availability of binary and ternary splits means that the coding tree depth no longer directly corresponds to the CU region. The possibility of abnormal (non-square) block sizes can be reduced by constraining the split options to eliminate splits that would result in a block width or height less than four samples or a split that would result in a size that is not a multiple of four samples. Generally, the constraint will apply when considering luminance samples. However, in the described arrangement, the constraint can be applied separately to blocks of the chrominance channel. The application of the split option constraint to the chrominance channel may result in different minimum block sizes for luminance vs chrominance (e.g., when the frame data is in 4:2:0 chrominance format or 4:2:2 chrominance format). Each split produces sub-regions with edge dimensions that are invariant, bisected, or quartered with respect to the containing region. Then, since the CTU size is a power of 2, the edge dimensions of all CUs are also powers of 2.

[0133] Figure 6 is a schematic flow chart of a data stream 600 showing the QTBTTT (or "coding tree") structure used in general video coding. The QTBTTT structure is used for each CTU to define the division of the CTU into one or more CUs. The QTBTTT structure of each CTU is determined by a block partitioner 310 in the video encoder 114 and is encoded into the bitstream 115 or decoded from the bitstream 133 by an entropy decoder 420 in the video decoder 134. According to Figure 5 the shown division, the data stream 600 further features a permitted combination available for the block partitioner 310 to divide the CTU into one or more CUs.

[0134] Starting from the top level of the hierarchical structure, i.e., at the CTU, zero or more quadtree partitions are first performed. Specifically, the quadtree (QT) split decision 610 is made by the block partitioner 310. The decision at 610 returns a "1" symbol, indicating that it is decided to split the current node into four child nodes according to the quadtree split 512. As a result, four new nodes are generated, such as at 620, etc., and for each new node, the process recurs back to the QT split decision 610. Each new node is considered in raster (or Z-scan) order. Alternatively, if the QT split decision 610 indicates no further split (returns a "0" symbol), the quadtree partition stops, and then the multi-tree (MT) split is considered.

[0135] First, the MT split decision 612 is made by the block partitioner 310. At 612, a decision indicating an MT split is made. A "0" symbol is returned at the decision 612, indicating that no further split of the node into child nodes will be performed. If no further split of the node is to be performed, the node is a leaf node of the coding tree and corresponds to a CU. The leaf node is output at 622. Alternatively, if the MT split 612 indicates a decision to perform an MT split (returns a "1" symbol), the block partitioner 310 enters the direction decision 614.

[0136] The direction decision 614 indicates the direction of the MT split as horizontal ("H" or "0") or vertical ("V" or "1"). If the decision 614 returns a "0" indicating the horizontal direction, the block partitioner 310 enters the decision 616. If the decision 614 returns a "1" indicating the vertical direction, the block partitioner 310 enters the decision 618.

[0137] In each of the decisions 616 and 618, the number of partitions of the MT split is indicated as two (binary split or "BT" node) or three (ternary split or "TT") during the BT / TT split. That is, when the direction indicated from 614 is horizontal, the BT / TT split decision 616 is made by the block partitioner 310, and when the direction indicated from 614 is vertical, the BT / TT split decision 618 is made by the block partitioner 310.

[0138] The BT / TT split decision 616 indicates whether the horizontal split is a binary split 514 indicated by returning a "0" or a ternary split 518 indicated by returning a "1". When the BT / TT split decision 616 indicates a binary split, at the step 625 of generating the HBT CTU node, the block partitioner 310 generates two nodes according to the binary horizontal split 514. When the BT / TT split 616 indicates a ternary split, at the step 626 of generating the HTT CTU node, the block partitioner 310 generates three nodes according to the ternary horizontal split 518.

[0139] The BT / TT split decision 618 indicates whether the vertical split is a binary split 516 indicated by returning "0" or a ternary split 520 indicated by returning "1". When the BT / TT split 618 indicates a binary split, at step 627 of generating the VBT CTU node, the block partitioner 310 generates two nodes according to the vertical binary split 516. When the BT / TT split 618 indicates a ternary split, at step 628 of generating the VTT CTU node, the block partitioner 310 generates three nodes according to the vertical ternary split 520. For each node obtained from steps 625 - 628, the data stream 600 is applied recursively back to the MT split decision 612 in the order from left to right or from top to bottom according to the direction 614. As a result, binary tree and ternary tree splits can be applied to generate CUs of various sizes.

[0140] Figure 7A and 7B Provide an example split 700 of the CTU 710 into multiple CUs or CBs. In Figure 7A an example CU 712 is shown. Figure 7A Show the spatial arrangement of the CUs in the CTU 710. The example split 700 is also shown as an encoding tree 720 in Figure 7B .

[0141] At each non - leaf node (e.g., nodes 714, 716, and 718) in the CTU 710 of Figure 7A , the included nodes (which can be further split or can be CUs) are scanned or traversed in "Z - order" to create a list of nodes represented as columns in the encoding tree 720. For quadtree splits, the Z - order scan gives the order from the upper left to the right and then from the lower left to the right. For horizontal and vertical splits, the Z - order scan (traversal) simplifies to a scan from the top to the bottom and a scan from the left to the right, respectively. Figure 7B The encoding tree 720 of

[0142] lists all the nodes and CUs according to the applied scan order. Each split generates a list of two, three, or four new nodes at the next level of the tree until the leaf nodes (CUs) are reached.

[0142] In the case of decomposing an image into CTUs and further into CUs using the block partitioner 310 as described in reference Figure 3 and generating each residual block (324) using the CUs, the video encoder 114 performs forward transformation and quantization on the residual blocks. Subsequently, the resulting TB 336 is scanned to form an ordered list of residual coefficients as part of the operation of the entropy encoding module 338. An equivalent process is performed in the video decoder 134 to obtain the TB from the bitstream 133.

[0143] Figure 8A 、 8BFigures 8C illustrate the levels of subdivision resulting from splits in the coding tree and the corresponding effect on partitioning the coding tree units into quantization groups. The residual of the TB signals an incremental QP (392) for each quantization group at most once. In HEVC, the definition of the quantization group corresponds to the coding tree depth since this defines regions of fixed size. In VVC, the additional splits mean that the coding tree depth is no longer a suitable proxy for the CTU region. In VVC, a "level of subdivision" is defined where each increment corresponds to half of the region encompassed.

[0144] Figure 8A Figure 800 shows a set of splits in the coding tree and the corresponding levels of subdivision. At the root node of the coding tree, the level of subdivision is initialized to zero. When the coding tree includes a quadtree split (e.g., 810), the level of subdivision is incremented by two for any CU contained therein. When the coding tree includes a binary split (e.g., 812), the level of subdivision is incremented by one for any CU contained therein. When the coding tree includes a ternary split (e.g., 814), the level of subdivision is incremented by two for the two outer CUs and incremented by one for the inner CU resulting from the ternary split. When traversing the coding tree of each CTU, as referenced Figure 6 as described, the level of subdivision of each resulting CU is determined according to set 800.

[0145] Figure 8B Figure 840 shows an example set of CU nodes and shows the effect of the splits. The example parent node 820 with a zero level of subdivision in set 840 corresponds to Figure 8B a CTU of size 64×64 in the example of

[0146] In Figure 8BIn the example, the quantization group threshold is set to 1, corresponding to half of the 64×64 region, i.e., a region corresponding to 2048 samples. A flag tracks the start of a new QG. For any node with a subdivision level less than or equal to the quantization group threshold, the flag tracking the new QG is reset. When traversing the parent node 820 with a zero subdivision level, the flag is set. Although the central CU 822 of size 32×64 has a region of 2048 samples, the two sibling CUs 821 and 823 have a subdivision level of 2, i.e., a region of 1024, so the flag is not reset when traversing the central CU, and the quantization group does not start at the central CU. Instead, following the initial flag reset, the flag starts at the parent node as shown at 824. Effectively, the QP can only change at boundaries aligned with multiples of the quantization group region. The delta QP is signaled together with the residual of the TB associated with the CB. If there are no valid coefficients, there is no opportunity to code the delta QP.

[0147] Figure 8C Example 860 shows the division of CTU 862 into multiple CUs and QGs to illustrate the relationship between the subdivision level, QG, and the signaling of delta QP. The vertical binary split divides CTU 862 into two halves, with the left half 870 containing one CU, CU0, and the right half 872 containing several CUs (CU1 - CU4). In Figure 8C the example, the quantization group threshold is set to 2, such that the quantization group typically has a region equal to one - quarter of the CTU region. Since the subdivision level of the parent node (i.e., the root node of the coding tree) is zero, the QG flag is reset, and a new QG will start from the next coding CU (i.e., the CU at arrow 868). CU0 (870) has coding coefficients, so the delta QP 864 is coded together with the residual of CU0. The right half 872 undergoes a horizontal binary split and is further split in the upper and lower parts of the right half 872, resulting in CUs CU1 - CU4. The subdivision levels of the coding tree nodes corresponding to the upper (877, including CU1 and CU2) and lower (878, including CU3 and CU4) parts of the right half 872 are 2. The subdivision level 2 is equal to the quantization group threshold 2, so new QGs start in each part, labeled 874 and 876 respectively. CU1 has no coding coefficients (no residuals), and CU2 is a "skip" CU, which also has no coding coefficients. Thus, for the upper part, no delta QP is coded. CU3 is a skip CU, and CU4 has coding residuals, so for the QG including CU3 and CU4, the delta QP 866 is coded using the residual of CU4.

[0148] Figure 9A and 9BShows a 4×4 transform block scan pattern and associated primary and secondary transform coefficients. The operation of the secondary transform module 330 on the primary residual coefficients is described from the perspective of the video encoder 114. The 4×4 TB 900 is scanned according to the backward diagonal scan pattern 910. The scan pattern 910 travels from the "last significant coefficient" position towards the DC (top left) coefficient position. All coefficient positions that are not scanned, such as the residual coefficients that are located after the last significant coefficient position when considering scanning in the forward direction, are implicitly non-effective. When secondary transform is used, all remaining coefficients are non-effective. That is, all secondary-domain residual coefficients for which no secondary transform is performed are non-effective, and all primary-domain residual coefficients that are not filled by the application of the secondary transform need to be non-effective. Additionally, after the forward secondary transform is applied by module 330, there may be fewer secondary transform coefficients than the number of primary transform coefficients processed by the secondary transform module 330. For example, Figure 9B shows a set of blocks 920. In Figure 9B it, sixteen (16) primary coefficients are arranged as a 4×4 sub-block, i.e., 924 of the 4×4 TB 920. In Figure 9B the example, the primary residual coefficients can be subjected to a secondary transform to produce a secondary transform block 926. The secondary transform block 926 contains eight secondary transform coefficients 928. The eight secondary transform coefficients 928 are stored in the TB according to the scan pattern 910, packed forward from the DC coefficient position. The remaining coefficient positions of the 4×4 sub-block (shown as region 930) contain the quantized residual coefficients from the primary transform and need to be non-effective for the secondary transform to be applied. Thus, the last significant coefficient position indicating one of the first eight scan positions of the 4×4 TB designated as TB 920 indicates (i) the application of the secondary transform, or (ii) after quantization, the output of the primary transform does not have effective coefficients beyond the eighth scan position of TB 920.

[0149] When a TB can be subjected to a secondary transform, a secondary transform index (i.e., 388) is encoded to indicate the possible application of the secondary transform. The secondary transform index can also indicate which kernel will be applied as the secondary transform at module 330 in the case where multiple transform kernels are available. Accordingly, when the last significant coefficient position is at any one of the scan positions reserved for holding secondary transform coefficients (e.g., 928), the video decoder 134 decodes the secondary transform index 470.

[0150] Although a quadratic transform kernel that maps 16 primary coefficients to eight quadratic coefficients has been described, different kernels are possible, including kernels that map to a different number of quadratic transform coefficients. The number of quadratic transform coefficients can be the same as the number of primary transform coefficients, e.g., 16. For a TB with a width of 4 and a height greater than 4, the behavior described for the 4×4 TB case applies to the top sub-block of the TB. When the quadratic transform is applied, the other sub-blocks of the TB have zero-valued residual coefficients. For a TB with a width greater than 4 and a height equal to 4, the behavior described for the 4×4 TB case applies to the leftmost sub-block of the TB, and the other sub-blocks of the TB have zero-valued residual coefficients, allowing the use of the last significant coefficient position to determine whether the quadratic transform index needs to be decoded.

[0151] Figure 9C and 9D illustrates an 8×8 transform block scan pattern and example associated primary and quadratic transform coefficients. Figure 9C illustrates a backward diagonal scan pattern 950 based on 4×4 sub-blocks for an 8×8 TB 940. The 8×8 TB 940 is scanned in the backward diagonal scan pattern 950 based on 4×4 sub-blocks. Figure 9D illustrates a set 960 that shows the operation effect of the quadratic transform. The scan 950 returns from the last significant coefficient position to the DC (top left) coefficient position. When the remaining 16 primary coefficients (shown as 964) are zero-valued, it is possible to apply a forward quadratic transform kernel to 48 primary coefficients (the region 962 shown as 940). Applying the quadratic transform to the region 962 results in 16 quadratic transform coefficients shown as 966. The other coefficient positions of the TB are zero-valued, labeled 968. If the last significant position of the 8×8 TB 940 indicates that the quadratic transform coefficients are within 966, the quadratic transform index 388 is encoded to indicate that the module 330 applies a specific transform kernel (or bypasses the kernel). The video decoder 134 uses the last significant position of the TB to determine whether to decode the quadratic transform index, i.e., index 470. For transform blocks with a width or height exceeding eight samples, Figure 9C and 9D the method of

[0152] as Figures 9A to 9DAs described in , two sizes of secondary transform kernels are available. One size of secondary transform kernel is used for transform blocks with a width or height of 4, and another size of secondary transform is used for transform blocks with a width and height greater than 4. Within kernels of each size, multiple sets (e.g., four) of secondary transform kernels are available. A set is selected based on the intra prediction mode of the block, and the set may be different between luma blocks and chroma blocks. Within the selected set, one or two kernels are available. Independent of the luma blocks and chroma blocks in the coding units belonging to the common tree of the coding tree unit, the use of a kernel within the selected set or bypassing the secondary transform is signaled via the secondary transform index. In other words, the index for the luma channel and the index for the chroma channel are independent of each other.

[0153] Figure 10 A set 1000 of transform blocks available in the Versatile Video Coding (VVC) standard is shown. Figure 10 Also shown is the application of a secondary transform to a subset of the residual coefficients of the transformed blocks from set 1000 . Figure 10 TBs are shown with widths and heights ranging from 4 to 32. However, TBs with a width and / or height of 64 are possible but not shown for ease of reference.

[0154] A 16-point secondary transform 1052 (shown with darker shading) is applied to a 4×4 set of coefficients. The 16-point secondary transform 1052 is applied to TBs with a width or height of 4, such as 4×4 TB 1010, 8×4 TB 1012, 16×4 TB 1014, 32×4 TB 1016, 4×8 TB 1020, 4×16 TB 1030, and 4×32 TB 1040. If a 64-point primary transform is available, the 16-point secondary transform 1052 is applied to TBs of size 4×64 and 64×4 ( Figure 10 (not shown in the figure). For a TB with a width or height of four but with more than 16 primary coefficients, the 16-point secondary transform is applied only to the top left 4×4 sub-block of the TB, and other sub-blocks are required to have zero-valued coefficients to apply the secondary transform. Typically, applying the 16-point secondary transform results in 16 secondary transform coefficients, which are packed into the TB for encoding in the sub-blocks that obtain the original 16 primary transform coefficients. For example, as shown in reference Figure 9B As described, the secondary transform kernel may cause creation of secondary transform coefficients which are smaller in number than the number of primary transform coefficients to which the secondary transform is applied.

[0155] For transform sizes with width and height greater than four, such as Figure 10As shown, a 48-point secondary transform 1050 (shown in lighter shading) can be used for application to three 4×4 sub-blocks of residual coefficients in the upper left 8×8 region of a transform block. In each case, in the regions shown in lighter shading and dashed outlines, the 48-point secondary transform 1050 is applied to the 8×8 transform block 1022, 16×8 transform block 1024, 32×8 transform block 1026, 8×16 transform block 1032, 16×16 transform block 1034, 32×16 transform block 1036, 8×32 transform block 1042, 16×32 transform block 1044, and 32×32 transform block 1046. If a 64-point primary transform is available, the 48-point secondary transform 1050 is also applicable to TBs (not shown) of sizes 8×64, 16×64, 32×64, 64×64, 64×32, 64×16, and 64×8. Application of the 48-point secondary transform kernel typically results in fewer than 48 secondary transform coefficients being produced. For example, 8 or 16 secondary transform coefficients can be produced. The secondary transform coefficients are stored in the transform block in the upper left region. For example, Figure 9D shows eight secondary transform coefficients. The primary transform coefficients that are not subjected to the secondary transform (“only primary coefficients”) (e.g., coefficient 1066 of TB 1034 (similar to Figure 9D 964)) need to be zero-valued to apply the secondary transform. After applying the 48-point secondary transform 1050 in the forward direction, the region that can contain valid coefficients is reduced from 48 coefficients to 16 coefficients, thus further reducing the number of coefficient positions that can contain valid coefficients. For example, 968 will only contain non-valid coefficients. For the inverse secondary transform, for example, the decoded valid coefficients that exist only in 966 of the TB are transformed to produce any coefficients that may be valid in a region (e.g., 962), and then the primary inverse transform is applied to these coefficients. When the secondary transform reduces one or more sub-blocks to a set of 16 secondary transform coefficients, only the upper left 4×4 sub-block can contain valid coefficients. The last valid coefficient position at any coefficient position where the secondary transform coefficients can be stored indicates whether the secondary transform or only the primary transform is applied. However, after quantization, the resulting valid coefficients are in the same region as if the secondary transform kernel had been applied.

[0156] When the last valid coefficient position indicates a secondary transform coefficient position in a TB (e.g., 922 or 962), a signaled secondary transform index is needed to distinguish between applying the secondary transform kernel or bypassing the secondary transform. Although the application of the secondary transform has been described from the perspective of the video encoder 114 to Figure 10TBs of various sizes in [the relevant context], but corresponding inverse processing is performed in the video decoder 134. The video decoder 134 first decodes the last valid coefficient position. If the decoded last valid coefficient position indicates a potential application of the secondary transform, i.e., the position is within 928 or 966 of the secondary transform kernel that generates 8 or 16 secondary transform coefficients respectively, then the secondary transform index is decoded to determine whether to apply or bypass the inverse secondary transform.

[0157] Figure 11 Illustrates the syntax structure 1100 of the bitstream 1101 with multiple slices. Each slice in the slice contains multiple coding units. The bitstream 1101 can be generated by the video encoder 114 as, for example, the bitstream 115, or can be parsed by the video decoder 134 as, for example, the bitstream 133. The bitstream 1101 is segmented into multiple parts, such as Network Abstraction Layer (NAL) units, where the description is achieved by setting a NAL unit header (such as 1108, etc.) before each NAL unit. The Sequence Parameter Set (SPS) 1110 defines sequence-level parameters, such as the profile (toolset) for encoding and decoding the bitstream, chroma format, sample bit depth, and frame resolution, etc. The set 1110 also includes parameters that constrain the application of different types of splits in the coding tree of each CTU. For example, using the log2 base for block size constraints and expressing the parameters relative to other parameters (such as the minimum CTU size, etc.), the encoding of the parameters that constrain the split type can be optimized for a more compact representation. Several parameters encoded in the SPS 1110 are as follows:

[0158] · log2_ctu_size_minus5: Specifies the CTU size, where the encoded values 0, 1, and 2 specify CTU sizes of 32×32, 64×64, and 128×128 respectively.

[0159] · partition_constraints_override_enabled_flag: Enables the slice-level override of several parameters, collectively referred to as the partition constraint parameters 1130.

[0160] · log2_min_luma_coding_block_size_minus2: Specifies the minimum coding block size (in luma samples), where the values 0, 1, 2,... specify minimum luma CB sizes of 4×4, 8×8, 16×16,.... The maximum encoded value is constrained by the specified CTU size, i.e., such that log2_min_luma_coding_block_size_minus2 ≤ log2_ctu_size_minus5 + 3. The available chroma block sizes correspond to the available luma block sizes, scaled according to the chroma channel subsampling of the chroma format in use.

[0161] ·sps_max_mtt_hierarchy_depth_inter_slice: Specifies the maximum hierarchical depth of coding units in the coding tree for multi-tree splitting (i.e., binary and ternary splitting) relative to the quadtree nodes in the coding tree of an inter (P or B) slice (i.e., once the quadtree splitting stops in the coding tree), and is one of the parameters 1130.

[0162] ·sps_max_mtt_hierarchy_depth_intra_slice_luma: Specifies the maximum hierarchical depth of coding units in the coding tree for multi-tree splitting (i.e., binary and ternary) relative to the quadtree nodes in the coding tree of an intra (I) slice (i.e., once the quadtree splitting stops in the coding tree), and is one of the parameters 1130.

[0163] ·partition_constraints_override_flag: When partition_constraints_override_enabled_flag in the SPS is equal to 1, this parameter is signaled in the slice header, and this parameter indicates that the partition constraints signaled in the SPS will be overridden for the corresponding slice.

[0164] Picture Parameter Set (PPS) 1112 defines a set of parameters applicable to zero or more frames. Parameters included in PPS 1112 include parameters for splitting a frame into one or more "regions" and / or "blocks". The parameters of PPS1112 may also include a list of CU chroma QP offsets, one of which may be applied at the CU level to derive the quantization parameter for use by the chroma block from the quantization parameter of the collocated luma CB.

[0165] The sequence of slices forming a picture is called an Access Unit (AU), such as AU 0 1114, etc. AU 0 1114 includes three slices, such as slices 0 to 2, etc. Slice 1 is labeled 1116. Like other slices, slice 1 (1116) includes a slice header 1118 and slice data 1120.

[0166] The slice header includes parameters grouped as 1134. Group 1134 includes:

[0167] ·slice_max_mtt_hierarchy_depth_luma: Signaled in slice header 1118 when partition_constraints_override_flag in the slice header is equal to 1 and overrides the value derived from the SPS. For I slices, instead of using sps_max_mtt_hierarchy_depth_intra_slice_luma to set MaxMttDepth at 1134, slice_max_mtt_hierarchy_depth_luma is used. For P or B slices, instead of using sps_max_mtt_hierarchy_depth_inter_slice, slice_max_mtt_hierarchy_depth_luma is used.

[0168] Variable MinQtLog2SizeIntraY (not shown) is derived from the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma decoded from SPS 1110, which specifies the minimum coded block size resulting from zero or more quadtree splits of an I slice (i.e., no further MTT splits occur in the coding tree). Variable MinQtLog2SizeInterY (not shown) is derived from the syntax element sps_log2_diff_min_qt_min_cb_inter_slice decoded from SPS 1110. Variable MinQtLog2SizeInterY specifies the minimum coded block size resulting from zero or more quadtree splits of P and B slices (i.e., no further MTT splits occur in the coding tree). Since the CUs resulting from quadtree splits are square, variables MinQtLog2SizeIntraY and MinQtLog2SizeInterY each specify both the width and height (as log2 of the CU width / height).

[0169] The parameter cu_qp_delta_subdiv may be optionally signaled in the slice header 1118, and indicates the maximum subdivision level for signaling the delta QP in the coding tree for the luminance branch in a common tree or in a separate tree slice. For an I slice, the range of cu_qp_delta_subdiv is from 0 to 2*(log2_ctu_size_minus5 + 5 - MinQtLog2SizeIntraY + MaxMttDepthY 1134). For a P or B slice, the range of cu_qp_delta_subdiv is from 0 to 2*(log2_ctu_size_minus5 + 5 - MinQtLog2SizeInterY + MaxMttDepthY 1134). Since the range of cu_qp_delta_subdiv depends on the value MaxMttDepthY 1134 derived from the partitioning constraints obtained since the SPS 1110 or the slice header 1118, there is no parsing problem.

[0170] The parameter cu_chroma_qp_offset_subdiv may be optionally signaled in the slice header 1118, and indicates the maximum subdivision level for signaling the chroma CU QP offset in the chroma branch in a common tree or in a separate tree slice. The range constraints for cu_chroma_qp_offset_subdiv for an I or P / B slice are the same as the corresponding range constraints for cu_qp_delta_subdiv.

[0171] The subdivision levels 1136 are derived for the CTUs in the slice 1120, which specify cu_qp_delta_subdiv for the luminance CB and cu_chroma_qp_offset_subdiv for the chroma CB. As described in reference Figure 8A -C, the subdivision levels are used to determine at which points in the CTU the delta QP syntax elements are coded. For the chroma CB, the method of Figure 8A -C is also used to signal the chroma CU level offset enable (and index, if enabled).

[0172] Figure 12Shows the syntax structure 1200 of the strip data 1120 of the bitstream 1101 (e.g., 115 or 133), which has a common tree for encoding the luminance and chrominance coding blocks of tree units (such as CTU 1210, etc.). CTU 1210 includes one or more CUs, exemplified as CU 1214. CU 1214 includes a signaled prediction mode 1216a, followed by a transform tree 1216b. When the size of CU 1214 does not exceed the maximum transform size (32×32 or 64×64), the transform tree 1216b includes one transform unit, exemplified as TU 1218.

[0173] If the prediction mode 1216a indicates the use of intra prediction for CU 1214, the intra prediction mode for luminance and the intra prediction mode for chrominance are specified. For the luminance CB of CU 1214, the primary transform type is also signaled as (i) horizontal and vertical DCT-2, (ii) horizontal and vertical transform skip, or (iii) a combination of horizontal and vertical DST-7 and DCT-8. If the signaled luminance transform type is horizontal and vertical DCT-2 (option (i)), then under the conditions described in reference Figure 9A -D, an additional luminance secondary transform type 1220, also known as the "low-frequency non-separable transform" (LFNST) index, is signaled in the bitstream. The chrominance secondary transform type 1221 is also signaled. The chrominance secondary transform type 1221 is signaled independently of whether the luminance primary transform type is DCT-2.

[0174] The use of the common coding tree results in a TU 1218 that includes TBs for the respective color channels, shown as luminance TB Y 1222, first chrominance TB Cb 1224, and second chrominance TB Cr 1226. It is available to send a single chrominance TB to specify the coding mode for the chrominance residuals of both the Cb and Cr channels, called the "joint CbCr" coding mode. When the joint CbCr coding mode is enabled, a single chrominance TB is encoded.

[0175] Regardless of the color channel, each TB includes a last position 1228. The last position 1228 indicates the last valid residual coefficient position in the TB when considering the coefficients in the diagonal scan mode, which is used to serialize the coefficient array of the TB in the forward direction (i.e., from the DC coefficient forward). If the last position 1228 of the TB indicates that only the coefficients in the secondary transform domain (i.e., all the remaining coefficients that will only undergo the primary transform) are valid, then a secondary transform index is signaled to specify whether the secondary transform is applied.

[0176] If a second transform is to be applied and if more than one second transform kernel is available, the second transform index indicates which kernel to select. Typically, in a "candidate set", one kernel is available, or two kernels are available. The candidate set is determined from the intra prediction mode of the block. Typically, there are four candidate sets, but there can be fewer candidate sets. As described above, the second transform is used for luminance and chrominance, and thus the selected kernel depends on the intra prediction mode for the luminance and chrominance channels respectively. The kernel can also depend on the block size of the corresponding luminance and chrominance TBs. The kernel selected for chrominance also depends on the chrominance subsampling rate of the bitstream. If only one kernel is available, signaling is restricted to applying or not applying the second transform (index range 0 to 1). If two kernels are available, the index value is 0 (not apply), 1 (apply the first kernel), or 2 (apply the second kernel). For chrominance, the same second transform kernel is applied to each chrominance channel, so the residuals of the Cb block 1224 and the Cr block 1226 only need to include valid coefficients in the positions that undergo the second transform, as described in reference Figure 9A -D. If joint CbCr coding is used, the requirement to only include valid coefficients in the positions that undergo the second transform only applies to a single coded chrominance TB, since the resulting Cb and Cr residuals only contain valid coefficients in the positions corresponding to the valid coefficients in the joint coded TB. If the (one or more) applicable color channels for a given second index are described by a single TB (a single last position, e.g., 1228), i.e., when using joint CbCr coding, luminance always only requires one TB and chrominance requires one TB, then the second transform index can be signaled immediately after the coded last position instead of after the TU, i.e., as index 1230 instead of 1220 (or 1221). Signaling the second transform earlier in the bitstream allows the video decoder 134 to start applying the second transform when each residual coefficient in the residual coefficients 1232 is decoded, thus reducing the delay in the system 100.

[0177] In the arrangement of the video encoder 114 and the video decoder 134, when joint CbCr coding is not used, a separate second transform index is signaled for each chrominance TB (i.e., 1224 and 1226), resulting in independent control of the second transform for each color channel. If each TB is controlled independently, the second transform index for each TB can be signaled immediately after the last position of the corresponding TB for luminance and chrominance (regardless of whether the joint CbCr mode is applied).

[0178] Figure 13A method 1300 for encoding frame data 113 into a bitstream 115 is shown. The bitstream 115 includes one or more strips as a sequence of coding tree units. The method 1300 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, the method 1300 may be performed by a video encoder 114 under the execution of a processor 205. Due to the workload of encoding frames, the steps of the method 1300 may be performed in different processors to, for example, share the workload using a contemporary multi-core processor, such that different strips are encoded by different processors. Further, when encoding respective parts (strips) of the bitstream 115, the partition constraints and quantization group definitions may vary from one strip to another, which is considered beneficial for rate control purposes. For additional flexibility in encoding the residuals of respective coding units, not only may the quantization group subdivision levels vary from one strip to another, but also the application of the secondary transform is independently controllable for luminance and chrominance. Thus, the method 1300 may be stored on a computer-readable storage medium and / or in a memory 206.

[0179] The method 1300 begins with an encoding SPS / PPS step 1310. At step 1310, the video encoder 114 encodes the SPS 1110 and PPS 1112 as a sequence of fixed and variable length coding parameters into the bitstream 115. The partition_constraints_override_enabled_flag is encoded as part of the SPS 1110, indicating that the partition constraints can be overridden in the strip header (1118) of a corresponding strip (such as 1116, etc.). The default partition constraints are also encoded by the video encoder 114 as part of the SPS 1110.

[0180] The method 1300 continues from step 1310 to a step 1320 of splitting a frame into strips. In the execution of step 1320, the processor 205 splits the frame data 113 into one or more strips or consecutive parts. In cases where parallelism is desired, separate instances of the video encoder 114 encode the respective strips somewhat independently. A single video encoder 114 may process the respective strips sequentially, or some intermediate degree of parallelism may be implemented. Generally, splitting the frame into strips (consecutive parts) is aligned with the boundaries of splitting the frame into regions such as “sub-pictures” or blocks, etc.

[0181] The method 1300 continues from step 1320 to a step 1330 of encoding a strip header. At step 1330, an entropy encoder 338 encodes the strip header 1118 into the bitstream 115. The following provides an example implementation of step 1330. Figure 14 Provide an example implementation of step 1330.

[0182] Method 1300 continues from step 1330 to step 1340 of splitting the strip into CTUs. In the execution of step 1340, video encoder 114 splits strip 1116 into a sequence of CTUs. The strip boundaries are aligned with the CTU boundaries, and the CTUs in the strip are sorted according to the CTU scan order (usually raster scan order). Splitting the strip into CTUs determines which part of the frame data 113 will be processed by video encoder 113 when encoding the current strip.

[0183] Method 1300 continues from step 1340 to step 1350 of determining the coding tree. At step 1350, video encoder 114 determines the coding tree of the currently selected CTU in the strip. Method 1300 starts from the first CTU in strip 1116 on the first call of step 1350 and advances to the subsequent CTUs in strip 1116 on subsequent calls. When determining the coding tree of a CTU, various combinations of quadtree, binary, and ternary splits are generated and tested by block partitioner 310.

[0184] Method 1300 continues from step 1350 to step 1360 of determining the coding unit. At step 1360, video encoder 114 performs to determine the "best" coding of the CUs obtained from the various coding trees under evaluation using known methods. Determining the best coding involves determining the prediction mode (e.g., intra prediction with a specific mode or inter prediction with a motion vector), transform selection (primary transform type and optional secondary transform type). If the primary transform type of the luminance TB is determined to be DCT-2 or any quantized primary transform coefficients that are not subject to forward secondary transform are valid, the secondary transform index of the luminance TB can indicate the application of the secondary transform. Otherwise, the secondary transform index of the luminance indicates bypassing the secondary transform. For the luminance channel, the primary transform type is determined to be one of the MTS options of the chrominance channel, DCT-2, or transform skip, and DCT-2 is the available transform type. Refer to Figure 19A and 19B the determination of the secondary transform type is further described. Determining the coding may also include determining the quantization parameter that can change the QP, i.e., the quantization parameter at the quantization group boundary. When determining each coding unit, the best coding tree is also determined in a joint manner. When encoding a coding unit using intra prediction, the luminance intra prediction mode and the chrominance intra prediction are determined.

[0185] When there are no "AC" (coefficients in positions other than the upper left position of the transform block) residual coefficients in the primary domain residual obtained by the application of the DCT-2 primary transform, step 1360 of determining the coding unit may prohibit the application of the test secondary transform. If the test secondary transform application is performed on a transform block that includes only DC coefficients (the last position indicates that only the upper left coefficient of the transform block is valid), an increase in coding efficiency is seen. The prohibition of the test secondary transform when only DC primary coefficients are present spans the blocks to which the secondary transform index applies, i.e., Y, Cb, and Cr of the common tree when encoding a single index (only the Y channel when the Cb and Cr blocks are two samples wide or high). Although the residual with only DC coefficients has a lower coding cost compared to a residual with at least one AC coefficient, applying the secondary transform to a residual with only valid DC coefficients also results in a further reduction in the magnitude of the finally encoded DC coefficient. Even after further quantization and / or rounding operations before encoding, the magnitudes of the other (AC) coefficients after the secondary transform are not sufficient to obtain (one or more) valid encoded residual coefficients in the bitstream. In the common or separate tree coding tree, assuming there is at least one valid primary coefficient, even if there is only (one or more) DC coefficient(s) of the corresponding transform block within the application range of the secondary transform index, the video encoder 114 tests the selection of non-zero secondary transform index values (i.e., the application of the secondary transform).

[0186] Method 1300 continues from step 1360 to step 1370 of encoding the coding unit. At step 1370, the video encoder 114 encodes the determined coding unit of step 1360 in the bitstream 115. Refer to Figure 15 An example of how to encode the coding unit is described in more detail.

[0187] Method 1300 continues from step 1370 to step 1380 of the last coding unit test. At step 1380, the processor 205 tests whether the current coding unit is the last coding unit in the CTU. If not (step 1380 is "no"), the control in the processor 205 advances to step 1360 of determining the coding unit. Otherwise, if the current coding unit is the last coding unit (step 1380 is "yes"), the control in the processor 205 advances to step 1390 of the last CTU test.

[0188] At step 1390 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU in the slice 1116. If it is not the last CTU in the slice 1116, the control in the processor 205 returns to step 1350 of determining the coding tree. Otherwise, if the current CTU is the last one (step 1390 is "yes"), the control in the processor advances to step 13100 of the last slice test.

[0189] At step 13100 of the last stripe test, the processor 205 tests whether the current stripe being encoded is the last stripe in the frame. If it is not the last stripe (step 13100 is "No"), then control in the processor 205 advances to step 1330 of encoding the stripe header. Otherwise, if the current stripe is the last one and all stripes (consecutive parts) have been encoded (step 13100 is "Yes"), then method 1300 terminates.

[0190] Figure 14 Method 1400 for encoding stripe header 1118 into bitstream 115 as implemented at step 1330 is shown. Method 1400 can be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1400 can be performed by video encoder 114 under the execution of processor 205. Thus, method 1400 can be stored on a computer-readable storage medium and / or in memory 206.

[0191] Method 1400 begins at step 1410 of the partition constraint override enable test. At step 1410, the processor 205 tests whether the partition constraint override enable flag encoded in SPS 1110 indicates that the partition constraint can be overridden at the stripe level. If the partition constraint can be overridden at the stripe level (step 1410 is "Yes"), then control in the processor 205 advances to step 1420 of determining the partition constraint. Otherwise, if the partition constraint cannot be overridden at the stripe level (step 1410 is "No"), then control in the processor 205 advances to step 1480 of encoding other parameters.

[0192] At step 1420 of determining the partition constraint, the processor 205 determines the partition constraint (e.g., maximum MTT split depth) suitable for the current stripe 1116. In one example, frame data 310 contains a projection of a 360-degree view of a scene mapped into a 2D box and segmented into a number of sub-pictures. Depending on the selected viewport, some stripes may require higher fidelity and other stripes may require lower fidelity. The partition constraint for the stripe can be set based on the fidelity requirements of the portion of the frame data 310 encoded by the given stripe (e.g., according to step 1340). In cases where lower fidelity is considered acceptable, a shallower encoding tree with larger CUs is acceptable, and thus the maximum MTT depth can be set to a lower value. Accordingly, at least the subdivision level 1136 signaled with flag cu_qp_delta_subdiv is determined within the range resulting from the determined maximum MTT depth 1134. The corresponding chroma subdivision level is also determined and signaled.

[0193] Method 1400 continues from step 1420 to step 1430 of encoding a partition constraint override flag. At step 1430, entropy encoder 338 encodes the flag in bitstream 115, which indicates whether the partition constraint signaled in SPS 1110 is to be overridden for slice 1116. If a strip-specific partition constraint is derived at step 1420, the flag value will indicate the use of the partition constraint override feature. If the constraint determined at step 1420 matches the constraint already encoded in SPS 1110, there is no need to override the partition constraint because there is no change to be signaled, and the flag value is encoded accordingly.

[0194] Method 1400 continues from step 1430 to step 1440 of a partition constraint override test. At step 1440, processor 205 tests the flag value encoded at step 1430. If the flag indicates that the partition constraint is to be overridden (step 1440 is "yes"), then control in processor 205 advances to step 1450 of encoding the slice partition constraint. Otherwise, if the partition constraint is not overridden (step 1440 is "no"), then control in processor 205 advances to step 1480 of encoding other parameters.

[0195] Method 1400 continues from step 1440 to step 1450 of encoding the slice partition constraint. In the execution of step 1450, entropy encoder 338 encodes the determined partition constraint of the slice in bitstream 115. The partition constraint of the slice includes "slice_max_mtt_hierarchy_depth_luma", from which MaxMttDepthY 1134 is derived.

[0196] Method 1400 continues from step 1450 to step 1460 of encoding the QP subdivision level. At step 1460, entropy encoder 338 encodes the subdivision level of the luma CB using the "cu_qp_delta_subdiv" syntax element, as referenced Figure 11 as described.

[0197] Method 1400 continues from step 1460 to step 1470 of encoding the chroma QP subdivision level. At step 1470, entropy encoder 338 encodes the signaled subdivision level for the CU chroma QP offset using the "cu_chroma_qp_offset_subdiv" syntax element, as referenced Figure 11 as described.

[0198] Steps 1460 and 1470 operate to encode the overall QP subdivision levels of the stripes (successive portions) of a frame. The overall subdivision levels include both the subdivision level for the luminance coding units of the stripe and the subdivision level for the chrominance coding units of the stripe. For example, since separate coding trees are used for luminance and chrominance in an I stripe, the chrominance and luminance subdivision levels can be different.

[0199] Method 1400 continues from step 1470 to step 1480 of encoding other parameters. At step 1480, entropy encoder 338 encodes the other parameters in the stripe header 1118, such as parameters required to control specific tools such as deblocking, adaptive loop filtering, optionally selecting a scaling list from previously signaled scaling lists (for non-uniformly applying quantization parameters to transform blocks), etc. Method 1400 terminates when step 1480 is executed.

[0200] Figure 15 Method 1500 for encoding a coding unit in bitstream 115 is shown, which corresponds to Figure 13 step 1370. Method 1500 can be embodied by a device such as a configured FPGA, ASIC, or ASSP, etc. Additionally, method 1500 can be performed by video encoder 114 under the execution of processor 205. Thus, method 1500 can be stored on a computer-readable storage medium and / or in memory 206.

[0201] Method 1500 begins at step 1510 of encoding a prediction mode. At step 1510, entropy encoder 338 encodes the prediction mode for the coding unit determined at step 1360 in bitstream 115. The "pred_mode" syntax element is encoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit. If intra prediction is used for the coding unit, the luminance intra prediction mode is encoded, and the chrominance intra prediction mode is encoded. If inter prediction is used for the coding unit, the "merge index" can be encoded to select a motion vector from adjacent coding units for use by the coding unit, and the motion vector delta can be encoded to introduce an offset to the motion vector derived from spatially adjacent blocks. The primary transform type is encoded to select between using DCT-2 horizontally and vertically for the luminance TB of the coding unit, using transform skip horizontally and vertically, or using a combination of DCT-8 and DST-7 horizontally and vertically.

[0202] Method 1500 continues from step 1510 to step 1520 of the coded residual test. At step 1520, the processor 205 determines whether the residuals need to be coded for the coding unit. If there are any valid residual coefficients to be coded for the coding unit (step 1520 is "yes"), then the control in the processor 205 advances to the new QG test step 1530. Otherwise, if there are no valid residual coefficients for coding (step 1520 is "no"), then method 1500 terminates because all the information required to decode the coding unit is present in the bitstream 115.

[0203] At the new QG test step 1530, the processor 205 determines whether the coding unit corresponds to a new quantization group. If the coding unit corresponds to a new quantization group (step 1530 is "yes"), then the control in the processor 205 proceeds to the step 1540 of coding the incremental QP. Otherwise, if the coding unit is not related to a new quantization group (step 1530 is "no"), then the control in the processor 205 advances to the step 1550 of performing the main transform. When coding each coding unit, the nodes of the coding tree of the CTU are traversed at step 1530. When, as determined by "cu_qp_delta_subdiv", any child node of the current node has a subdivision level less than or equal to the subdivision level 1136 of the current slice, a new quantization group starts in the CTU region corresponding to that node, and step 1530 returns "yes". The first CU in the quantization group that includes the coded residuals will also include the coded incremental QP, thereby signaling any change in the quantization parameter applicable to the residual coefficients in that quantization group.

[0204] At the step 1540 of coding the incremental QP, the entropy encoder 338 codes the incremental QP in the bitstream 115. The incremental QP codes the difference between the predicted QP and the expected QP used in the current quantization group. The predicted QP is derived by averaging the QPs of adjacent earlier (above and to the left) quantization groups. When the subdivision level is low, the quantization group is large and the incremental QP is coded less frequently. The less frequent coding of the incremental QP results in a lower overhead for signaling changes in the QP, but also leads to less flexibility in rate control. The selection of the quantization parameter for each quantization group is performed by the QP controller module 390, which generally implements a rate control algorithm for a specific bitrate of the bitstream 115, which is to some extent independent of the variation in the statistics of the underlying frame data 113. Method 1500 continues from step 1540 to the step 1550 of performing the main transform.

[0205] At step 1550 of performing the main transform, the forward main transform module 326 performs the main transform according to the main transform type of the coding unit, thereby obtaining the main transform coefficients 328. The main transform is performed on each color channel. First, the main transform is performed on the luminance channel (Y), and then the main transform is performed on the Cb TB and Cr TB in subsequent calls to step 1550 for the current TU. For the luminance channel, the main transform types (DCT-2, transform skip, MTS option) are performed, and for the chrominance channels, DCT-2 is performed.

[0206] Method 1500 continues from step 1550 to step 1560 of quantizing the main transform coefficients. At step 1560, the quantizer module 334 quantizes the main transform coefficients 328 according to the quantization parameter 392 to generate the quantized main transform coefficients 332. The delta QP (when present) is used to encode the transform coefficients 328.

[0207] Method 1500 continues from step 1560 to step 1570 of performing the secondary transform. At step 1570, the secondary transform module 330 performs the secondary transform on the quantized main transform coefficients 332 according to the secondary transform index 388 of the current transform block to generate the secondary transform coefficients 336. Although the secondary transform is performed after quantization, the main transform coefficients 328 can maintain a higher precision compared to the final expected quantizer step size of the quantization parameter 392. For example, the magnitude can be 16 times the magnitude directly produced by the application of the quantization parameter 392, that is, four additional bits of precision will be retained. Retaining the additional bit precision in the quantized main transform coefficients 332 allows the secondary transform module 330 to operate on the coefficients in the main coefficient domain with higher precision. After applying the secondary transform, the final scaling (e.g., right shift by four bits) at step 1560 quantizes to the expected quantizer step size of the quantization parameter 392. The "scaling list" is applied to the main transform coefficients (which correspond to well-known transform basis functions (DCT-2, DCT-8, DST-7)), rather than operating on the secondary transform coefficients produced by the trained secondary transform kernel. When the secondary transform index 388 of the transform block indicates that no secondary transform is to be applied (index value equal to zero), the secondary transform is bypassed. That is, the main transform coefficients 332 are propagated through the secondary transform module 330 unchanged to become the secondary transform coefficients 336. The luminance secondary transform index is used in combination with the luminance intra prediction mode to select the secondary transform kernel to be applied to the luminance TB. The chrominance secondary transform index is used in combination with the chrominance intra prediction mode to select the secondary transform kernel to be applied to the chrominance TB.

[0208] Method 1500 continues from step 1570 to step 1580 of encoding the last position. At step 1580, entropy encoder 338 encodes the position of the last significant coefficient among the secondary transform coefficients 336 of the current transform block in bitstream 115. When step 1580 is called for the first time, the luminance TB is considered, and subsequent calls consider the Cb TB and then the Cr TB.

[0209] In an arrangement where the secondary transform index 388 is encoded immediately after the last position, method 1500 continues to step 1590 of encoding the LFNST index. If the secondary transform index is not inferred to be zero based on the last position encoded at step 1580, then at step 1590, entropy encoder 338 encodes the secondary transform index 338 as "lfnst_index" in bitstream 115 using a truncated unary codeword. Each CU has one luminance TB, thus allowing step 1590 for the luminance block, and when the "joint" encoding mode is used for chrominance, step 1590 can be performed for a single chrominance TB. Knowing the secondary transform index before decoding each residual coefficient enables the application of the secondary transform coefficient-by-coefficient, for example using multiply-accumulate logic when the coefficients are decoded. Method 1500 continues from step 1590 to step 15100 of encoding sub-blocks.

[0210] If the secondary transform index 388 is not encoded immediately after the last position, method 1500 continues from step 1580 to step 15100 of encoding sub-blocks. At step 15100 of encoding sub-blocks, the residual coefficients (336) of the current transform block are encoded in bitstream 115 as a series of sub-blocks. Moving back from the sub-block containing the last significant coefficient position to the sub-block containing the DC residual coefficient, the residual coefficients are encoded.

[0211] Method 1500 continues from step 15100 to step 15110 of the last TB test. At this step, processor 205 tests whether the current transform block is the last transform block in progress on the color channels (i.e., Y, Cb, and Cr). If the just-encoded transform block is for the Cr TB (step 15110 is "yes"), then the control in processor 205 advances to step 15120 of encoding the luminance LFNST index. Otherwise, if the current TB is not the last one (step 15110 is "no"), then the control in processor 205 returns to step 1550 of performing the main transform and the next TB (selecting Cb or Cr).

[0212] Steps 1550 through 15110 are described with an example of a common coding tree structure where the prediction mode is intra prediction and DCT-2 is used. In addition to the common coding tree structure using known methods, operations of steps such as performing a primary transform (1550), quantizing the primary transform coefficients (1560), and coding the final position (1590) can be implemented for an inter prediction mode or an intra prediction mode. Steps 1510 through 1540 can be implemented regardless of the prediction mode or the coding tree structure.

[0213] Method 1500 continues from step 15110 to step 15120 of coding the luminance LFNST index. At step 15120, if the secondary transform index applied to the luminance TB is not inferred to be zero (no secondary transform is applied), the entropy encoder 338 codes it in the bitstream 115. The luminance secondary transform index is inferred to be zero if the last valid position of the luminance TB indicates valid only primary residual coefficients or if a primary transform other than DCT-2 is performed. Additionally, the secondary transform index applied to the luminance TB is coded in the bitstream only for coding units using intra prediction and the common coding tree structure. The secondary transform index applied to the luminance TB is coded using flag 1220 (or flag 1230 for the combined CbCr mode).

[0214] Method 1500 continues from step 15120 to step 15130 of coding the chrominance LFNST index. At step 1530, if the secondary transform index applied to the chrominance TB is not inferred to be zero (no secondary transform is applied), the chrominance secondary transform index is coded in the bitstream 115 by the entropy encoder 338. The chrominance secondary transform index is inferred to be zero if the last valid position of any chrominance TB indicates valid only primary residual coefficients. Method 1500 terminates after performing step 15130, where control in the processor 205 returns to method 1300. The secondary transform index applied to the chrominance TB is coded in the bitstream only for coding units using intra prediction and the common coding tree structure. The secondary transform index applied to the chrominance TB is coded using flag 1221 (or flag 1230 for the combined CbCr mode).

[0215] Figure 16 Method 1600 for decoding a frame from a bitstream that is a sequence of coding units arranged as strips is shown. Method 1600 can be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1600 can be performed by the video decoder 134 under the execution of the processor 205. Thus, method 1600 can be stored on a computer-readable storage medium and / or in the memory 206.

[0216] Method 1600 decodes a bitstream encoded using Method 1300, in which partition constraints and quantization group definitions may vary from one strip to another, which is considered beneficial for rate control purposes when encoding respective portions (strips) of the encoded bitstream 115. Not only can the quantization group subdivision level vary from one strip to another, but also the application of the secondary transform is independently controllable for luminance and chrominance.

[0217] Method 1600 begins with the SPS / PPS decoding step 1610. In the execution of step 1610, the video decoder 134 decodes the SPS 1110 and PPS 1112 from the bitstream 133 into a sequence of fixed and variable length parameters. The partition_constraints_override_enabled_flag is decoded as part of the SPS 1110, indicating whether the partition constraints can be overridden in the strip header (e.g., 1118) of the corresponding strip (e.g., 1116). The default (i.e., as signaled in the SPS 1110 and used in the strip without a subsequent override) partition constraint parameter 1130 is also decoded by the video decoder 134 as part of the SPS 1110.

[0218] Method 1600 continues from step 1610 to the step 1620 of determining strip boundaries. In the execution of step 1620, the processor 205 determines the position of the strip in the current access unit in the bitstream 133. Generally, strip boundaries are identified by determining NAL unit boundaries (by detecting "start codes") and reading the NAL unit header including the "NAL unit type" for each NAL unit. A particular NAL unit type identifies the strip type, such as "I-strip", "P-strip", "B-strip", etc. After identifying the strip boundaries, the application 233 may distribute the execution of the subsequent steps of Method 1600 across different processors in a multi-processor architecture, for example, for parallel decoding. Each processor in the multi-processor system may decode different strips to obtain a higher decoding throughput.

[0219] Method 1600 continues from step 1610 to the step 1630 of decoding the strip header. At step 1630, the entropy decoder 420 decodes the strip header 1118 from the bitstream 133. An example method of decoding the strip header 1118 from the bitstream 133 as implemented at step 1630 is described below with reference to Figure 17 Describe an example method of decoding the strip header 1118 from the bitstream 133 as implemented at step 1630.

[0220] Method 1600 continues from step 1630 to step 1640 of splitting the strip into CTUs. At step 1640, video decoder 134 splits strip 1116 into a sequence of CTUs. The strip boundaries are aligned with the CTU boundaries, and the CTUs in the strip are sorted according to the CTU scan order. The CTU scan order is typically a raster scan order. Splitting the strip into CTUs determines which part of the frame data 113 will be processed by video decoder 134 when decoding the current strip.

[0221] Method 1600 continues from step 1640 to step 1650 of decoding the coding tree. In the execution of step 1650, video decoder 133 starts from the first CTU in strip 1116 when first invoking step 1650, and decodes the coding tree of the current CTU in the strip from bitstream 133. By Figure 6 decoding the split flag to decode the coding tree of the CTU. For subsequent iterations of step 1650 for the CTU, subsequent CTUs in strip 1116 are decoded. If the coding tree is encoded using an intra prediction mode and a common coding tree structure, the coding unit has a primary color channel (luminance or Y) and at least one secondary color channel (chrominance, Cb and Cr or CbCr). In this case, decoding the coding tree involves decoding a coding unit including the primary color channel and at least one secondary color channel according to the split flag of the coding tree unit.

[0222] Method 1600 continues from step 1660 to step 1670 of decoding the coding unit. At step 1670, video decoder 134 decodes the coding unit from bitstream 133. An example method of decoding the coding unit as implemented at step 1670 is described below with reference to Figure 18 the description.

[0223] Method 1600 continues from step 1610 to step 1680 of the last coding unit test. At step 1680, processor 205 tests whether the current coding unit is the last coding unit in the CTU. If it is not the last coding unit (step 1680 is "no"), the control in processor 205 returns to step 1670 of decoding the coding unit to decode the next coding unit of the coding tree unit. If the current coding unit is the last coding unit (step 1680 is "yes"), the control in processor 205 advances to step 1690 of the last CTU test.

[0224] At step 1690 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU in slice 1116. If it is not the last CTU in the slice (step 1690 is "No"), then control in the processor 205 returns to step 1650 of decoding the coding tree to decode the next coding tree unit of slice 1116. If the current CTU is the last CTU in slice 1116 (step 1690 is "Yes"), then control in the processor 205 advances to step 16100 of the last slice test.

[0225] At step 16100 of the last slice test, the processor 205 tests whether the current slice being decoded is the last slice in the frame. If it is not the last slice in the frame (step 16100 is "No"), then control in the processor 205 returns to step 1630 of decoding the slice header, and step 1630 operates to decode the slice header of the next slice in the frame (e.g., Figure 11 "slice 2") of the frame. If the current slice is the last slice in the frame (step 1600 is "Yes"), then method 1600 terminates.

[0226] As described with respect to Figure 1 apparatus 130 in, method 1600 of operating on multiple coding units operates to produce an image frame.

[0227] Figure 17 Illustrated is method 1700 for decoding a slice header into a bitstream as implemented at step 1630. Method 1700 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1700 may be performed by video decoder 134 under the execution of processor 205. Thus, method 1700 may be stored on a computer-readable storage medium and / or in memory 206.

[0228] Similar to method 1500, method 1700 is performed on the current slice or contiguous portion (1116) in a frame (e.g., frame 1101). Method 1700 begins with a partition constraint override enable test step 1710. At step 1710, the processor 205 tests whether the partition constraint override enable flag decoded from the SPS 1110 indicates that the partition constraint can be overridden at the slice level. If the partition constraint can be overridden at the slice level (step 1710 is "Yes"), then control in the processor 205 advances to step 1720 of decoding the partition constraint override flag. Otherwise, if the partition constraint override enable flag indicates that the constraint cannot be overridden at the slice level (step 1710 is "No"), then control in the processor 205 advances to step 1770 of decoding other parameters.

[0229] At step 1720 of decoding the partition constraint override flag, entropy decoder 420 decodes the partition constraint override flag from bitstream 133. The decoded flag indicates whether the partition constraint signaled in SPS 1110 will be overridden for the current slice 1116.

[0230] Method 1700 continues from step 1720 to step 1730 of the partition constraint override test. In the execution of step 1730, processor 205 tests the decoded flag value at step 1720. If the decoded flag indicates that the partition constraint will be overridden (step 1730 is "yes"), then control in processor 205 advances to step 1740 of decoding the slice partition constraint. Otherwise, if the decoded flag indicates that the partition constraint will not be overridden (step 1730 is "no"), then control in processor 205 advances to step 1770 of decoding other parameters.

[0231] At step 1740 of decoding the slice partition constraint, entropy decoder 420 decodes the determined partition constraint for the slice from bitstream 133. The partition constraint for the slice includes "slice_max_mtt_hierarchy_depth_luma", from which MaxMttDepthY 1134 is derived.

[0232] Method 1700 continues from step 1740 to step 1750 of decoding the QP subdivision level. At step 1720, entropy decoder 420 decodes the subdivision level of the luma CB using the "cu_qp_delta_subdiv" syntax element as described in Figure 11 the reference.

[0233] Method 1700 continues from step 1750 to step 1760 of decoding the chroma QP subdivision level. At step 1760, entropy decoder 420 decodes the subdivision level for signaling the CU chroma QP offset using the "cu_chroma_qp_offset_subdiv" syntax element as described in Figure 11 the reference.

[0234] Steps 1750 and 1760 operate to determine the subdivision levels of a specific consecutive portion (slice) of the bitstream. Repeated iterations between steps 1630 and 16100 operate to determine the subdivision levels of individual consecutive portions (slices) in the bitstream. As described below, each subdivision level applies to the coding units of the corresponding slice (consecutive portion).

[0235] Method 1700 continues from step 1760 to step 1770 of decoding other parameters. At step 1770, entropy decoder 420 decodes other parameters from the slice header 1118, such as parameters required to control specific tools such as deblocking, adaptive loop filter, optionally selecting a scaling list from a previously signaled scaling list (for non-uniformly applying quantization parameters to transform blocks), etc. Method 1700 terminates when step 1770 is executed.

[0236] Figure 18 Method 1800 for decoding a coding unit from a bitstream is shown. Method 1800 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1800 may be performed by video decoder 134 under the execution of processor 205. Thus, method 1800 may be stored on a computer-readable storage medium and / or in memory 206.

[0237] Method 1800 is implemented for the current coding unit of the current CTU (e.g., CTU0 of slice 1116). Method 1800 begins with step 1810 of decoding a prediction mode. At step 1800, entropy decoder 420 decodes the prediction mode of the coding unit determined at step 1360 of the bitstream 133 as Figure 13 such. At step 1810, the "pred_mode" syntax element is decoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit.

[0238] If intra prediction is used for the coding unit, then at step 1810, the luminance intra prediction mode and the chrominance intra prediction mode are also decoded. If inter prediction is used for the coding unit, then at step 1810, the "merge index" may also be decoded to determine the motion vector from adjacent coding units for use by this coding unit, and the motion vector difference may be decoded to introduce an offset to the motion vector derived from spatially adjacent blocks. The primary transform type is also decoded at step 1810 to select between using DCT-2 horizontally and vertically, transform skip horizontally and vertically, or a combination of DCT-8 and DST-7 horizontally and vertically for the luminance TB of the coding unit.

[0239] Method 1800 continues from step 1810 to step 1820 of the coded residual test. During the execution of step 1820, the processor 205 determines whether the residual needs to be decoded for the coding unit by decoding the "root coding block flag" of the coding unit using the entropy decoder 420. If there are any valid residual coefficients to be decoded for the coding unit (step 1820 is "yes"), then the control in the processor 205 proceeds to step 1830 of the new QG test. Otherwise, if there are no residual coefficients to be decoded (step 1820 is "no"), then method 1800 terminates because all the information required to decode the coding unit has been obtained in the bitstream 115. When method 1800 terminates, subsequent steps such as PB generation, applying in-loop filtering, etc. are performed, thereby generating decoded samples, as referenced Figure 4 as described.

[0240] At the new QG test step 1830, the processor 205 determines whether the coding unit corresponds to a new quantization group. If the coding unit corresponds to a new quantization group (step 1830 is "yes"), then the control in the processor 205 proceeds to step 1840 of decoding the delta QP. Otherwise, if the coding unit does not correspond to a new quantization group (step 1830 is "no"), then the control in the processor 205 proceeds to step 1850 of decoding the last position. The new quantization group is related to the current mode or the subdivision level of the coding unit. When decoding each coding unit, the nodes of the coding tree of the CTU are traversed. The new quantization group starts in the region of the CTU corresponding to the node when any child node of the current node has a subdivision level less than or equal to the subdivision level 1136 of the current stripe (i.e., as determined from "cu_qp_delta_subdiv"). The first CU in the quantization group that includes coded residual coefficients will also include the coded delta QP, thereby signaling any change in the quantization parameter applicable to the residual coefficients in the quantization group. Effectively, a single (at most one) quantization parameter increment is decoded for each region (quantization group). As described with respect to Figures 8A to 8C as described, each region (quantization group) is based on the decomposition of the coding tree units of each stripe and the corresponding subdivision levels (e.g., as coded at steps 1460 and 1470). In other words, each region or quantization group is based on the comparison of the subdivision level associated with the coding unit with the determined subdivision levels for the corresponding consecutive parts.

[0241] At the step 1840 of decoding the delta QP, the entropy decoder 420 decodes the delta QP from the bitstream 133. The delta QP encodes the difference between the predicted QP and the expected QP used in the current quantization group. The predicted QP is derived by averaging the QPs of the adjacent (above and left) quantization groups.

[0242] Method 1800 continues from step 1840 to step 1850 of decoding the last position. When performing step 1850, entropy decoder 420 decodes the position of the last significant coefficient among the secondary transform coefficients 424 of the current transform block from bitstream 133. When step 1850 is called for the first time, this step is performed for the luminance TB. In subsequent calls to step 1850 for the current CU, this step is performed for the Cb TB. If the last position indicates that the significant coefficient is outside the set of secondary transform coefficients of the luminance block or chrominance block (i.e., outside 928 or 966), the secondary transform index for the luminance or chrominance channel is respectively inferred to be zero. After the iteration for Cb, this step is implemented for the Cr TB in the iteration.

[0243] As described with respect to Figure 15 step 1590, in some arrangements, the secondary transform index is encoded immediately after the last significant coefficient position of the coding unit. When decoding the same coding unit, if the secondary transform index 470 is not inferred to be zero based on the positioning of the last position of the TB decoded in step 1840, the secondary transform index 470 is decoded immediately after decoding the position of the last significant residual coefficient of the coding unit. In the arrangement where the secondary transform index 470 is decoded immediately after the last significant coefficient position of the coding unit, method 1800 continues from step 1850 to step 1860 of decoding the LFNST index. When performing step 1860, entropy decoder 420 decodes the secondary transform index 470 from bitstream 133 as "lfnst_index" using a truncated unary codeword when all significant coefficients are subject to the inverse secondary transform (e.g., within 928 or 966). When performing joint coding of the chrominance TB using a single transform block, the secondary transform index 470 can be decoded for the luminance TB or chrominance. Method 1800 continues from step 1860 to step 1870 of decoding the sub-blocks.

[0244] If the secondary transform index 470 is not decoded immediately after the last significant position of the coding unit, method 1800 continues from step 1850 to step 1870 of decoding the sub-blocks. At step 1870, the residual coefficients (i.e., 424) of the current transform block are decoded from bitstream 133 as a series of sub-blocks, advancing from the sub-block containing the last significant coefficient position back to the sub-block containing the DC residual coefficient.

[0245] Method 1800 continues from step 1870 to step 1880 of the last TB test. In the execution of step 1880, the processor 205 tests whether the current transform block is the last transform block in progress on a color channel (i.e., Y, Cb, and Cr). If the just-decoded (current) transform block is for the Cr TB, then in the control in the processor 205, all TBs have been decoded (step 1880 is "yes"), and method 1800 proceeds to step 1890 of decoding the luminance LFNST index. Otherwise, if the TB has not been decoded (step 1880 is "no"), then the control in the processor 205 returns to step 1850 of decoding the last position. In the iteration of step 1850, the next TB (following the order of Y, Cb, Cr) is selected for decoding.

[0246] Method 1800 continues from step 1880 to step 1890 of decoding the luminance LFNST index. In the execution of step 1890, if the last position of the luminance TB is within the set of coefficients (e.g., 928 or 966) undergoing the second inverse transform and the luminance TB is using DCT-2 as the main transform both horizontally and vertically, then the entropy decoder 420 decodes from the bitstream 133 the second transform index 470 to be applied to the luminance TB. If the last valid position of the luminance TB indicates that there are valid primary coefficients outside the set of coefficients undergoing the second inverse transform (e.g., outside 928 or 966), then the luminance second transform index is inferred to be zero (no second transform is applied). The second transform index decoded at step 1890 is indicated as 1220 (or 1230 in the combined CbCr mode) in Figure 12 ...

[0247] Method 1800 continues from step 1890 to step 1895 of decoding the chrominance LFNST index. At step 1895, if the last position of each chrominance TB is within the set of coefficients (e.g., 928 or 966) undergoing the second inverse transform, then the entropy decoder 420 decodes from the bitstream 133 the second transform index 470 to be applied to the chrominance TB. If the last valid position of any chrominance TB indicates that there are valid primary coefficients outside the set of coefficients undergoing the second inverse transform (e.g., outside 928 or 966), then the chrominance second transform index is inferred to be zero (no second transform is applied). The second transform index decoded at step 1895 is indicated as 1221 (or 1230 in the combined CbCr mode) in Figure 12 ... When decoding the separate indexes for luminance and chrominance, separate arithmetic contexts for each truncated unary codeword can be used or the contexts can be shared such that the nth bin in each of the luminance and chrominance truncated unary codewords shares the same context.

[0248] Effectively, steps 1890 and 1895 respectively involve: decoding a first index (such as 1220, etc.) to select a kernel for the luminance (primary color) channel, and decoding a second index (such as 1221, etc.) to select a kernel for at least one chrominance (secondary color channel).

[0249] Method 1800 continues from step 1895 to step 18100 where an inverse secondary transform is performed. At this step, the inverse secondary transform module 436 performs an inverse secondary transform on the decoded residual transform coefficients 424 according to the secondary transform index 470 of the current transform block to generate secondary transform coefficients 432. The secondary transform index decoded at step 1890 is applied to the luminance TB, and the secondary transform index decoded at step 1895 is applied to the chrominance TB. The kernel selection for luminance and chrominance also depends respectively on the luminance intra prediction mode and the chrominance intra prediction mode (each decoded at step 1810). Step 18100 selects a kernel according to the LFNST index for luminance and selects a kernel according to the LFNST index for chrominance.

[0250] Method 1800 continues from step 18100 to step 18110 where the primary transform coefficients are inverse quantized. At step 18110, the inverse quantizer module 428 inverse quantizes the secondary transform coefficients 432 according to the quantization parameter 474 to generate inverse quantized primary transform coefficients 440. If the incremental QP is decoded at step 1840, the entropy decoder 420 determines the quantization parameter according to the incremental QP of the quantization group (region) and the quantization parameter of the earlier encoded unit of the image frame. As described above, the earlier encoded unit generally refers to the adjacent upper left encoded unit.

[0251] Method 1800 continues from step 1870 to step 18120 where the primary transform is performed. At step 18120, the inverse primary transform module 444 performs an inverse primary transform according to the primary transform type of the coding unit, such that the transform coefficients 440 are converted into residual samples 448 in the spatial domain. The inverse primary transform is performed on each color channel. First, the inverse primary transform is performed on the luminance channel (Y), and then the inverse primary transform is performed on the Cb and Cr TBs when the subsequent call of step 1650 for the current TU is made. Steps 18100 to 18120 effectively operate to decode the current coding unit by applying the kernel selected according to the LFNST index for luminance at step 1890 to the decoded residual coefficients of the luminance channel and applying the kernel selected according to the LFNST index for chrominance at step 1890 to the decoded residual coefficients of at least one chrominance channel.

[0252] Method 1800 terminates after performing step 18120, where the control in the processor 205 returns to method 1600.

[0253] Steps 1850 through 18120 are described with an example of a common coding tree structure where the prediction mode is intra prediction and the transform is DCT-2. For example, for coding units using intra prediction and the common coding tree structure only, the secondary transform index (1890) applied to the luminance TB is decoded from the bitstream. Similarly, for coding units using intra prediction and the common coding tree structure only, the secondary transform index (1895) applied to the chrominance TB is decoded from the bitstream. Except for the common coding tree structure using known methods, operations of steps such as decoding sub-blocks (1870), inverse quantizing the primary transform coefficients (18110), and performing the primary transform can be implemented for either the inter prediction mode or the intra prediction mode. Regardless of the prediction mode or structure, steps 1810 through 1840 are performed in the described manner.

[0254] Once method 1800 terminates, subsequent steps for decoding the coding unit are performed (including generating intra prediction samples 480 by module 476, summing the decoded residual samples 448 with the prediction block 452 by module 450, and applying the in-loop filter module 488 to produce filtered samples 492), which are output as frame data 135.

[0255] Figure 19A and 19B Shows rules for applying or bypassing the secondary transform to the luminance and chrominance channels. Figure 19A Shows Table 1900 illustrating the conditions for applying the secondary transform to the luminance and chrominance channels in the CUs generated by the common coding tree.

[0256] If the last significant coefficient position of the luminance TB indicates decoded significant coefficients that are not generated by the forward secondary transform and thus not subject to inverse secondary transform, there is condition 1901. If the last significant coefficient position of the luminance TB indicates decoded significant coefficients that are indeed generated by the forward secondary transform and thus subject to inverse secondary transform, there is condition 1902. Additionally, for the luminance channel, the primary transform type needs to be DCT-2 for condition 1902 to exist, otherwise condition 1901 exists.

[0257] If the last significant coefficient position of one or both chrominance TBs indicates decoded significant coefficients that are not generated by the forward secondary transform and thus not subject to inverse secondary transform, there is condition 1910. If the last significant coefficient position of one or both chrominance TBs indicates decoded significant coefficients that are indeed generated by the forward secondary transform and thus subject to inverse secondary transform, there is condition 1911. Additionally, the width and height of the chrominance block need to be at least four samples (e.g., chrominance subsampling when using 4:2:0 or 4:2:2 chrominance formats can result in a width or height of two samples) for condition 1911 to exist.

[0258] If conditions 1901 and 1910 exist, the secondary transform index is not signaled (either independently or jointly), and the secondary transform index is not applied in luminance or chrominance, i.e., 1920. If conditions 1901 and 1911 exist, a secondary transform index is signaled to indicate the application of the selected kernel or bypass for only the luminance channel, i.e., 1921. If conditions 1902 and 1910 exist, a secondary transform index is signaled to indicate the application of the selected kernel or bypass for only the chrominance channel, i.e., 1922. If conditions 1911 and 1902 exist, an arrangement with independent signaling signals two secondary transform indexes, one for luminance TB and one for chrominance TB, i.e., 1923. When conditions 1902 and 1911 exist, an arrangement with a single signaled secondary transform index uses one index to control the selection for both luminance and chrominance, although the selected kernel also depends on the intra prediction mode within the luminance and chrominance frames, which may vary. The ability to apply the secondary transform to luminance or chrominance (i.e., 1921 and 1922) results in increased coding efficiency.

[0259] Figure 19B Table 1950 showing the search options available to the video encoder 114 at step 1360. The secondary transform indexes for luminance (1952) and chrominance (1953) are shown as 1952 and 1953, respectively. An index value of 0 indicates bypassing the secondary transform, and index values of 1 and 2 indicate which of two kernels to use for a candidate set derived from the intra prediction mode in the luminance or chrominance. There are nine combinations ("0,0" to "2,2") of the resulting search space, which may be constrained by the constraints described in the reference Figure 19A A simplified search (1951) of three combinations can test only combinations where the luminance and chrominance secondary transform indexes are the same, subject to zeroing the index indicating the last valid coefficient position to a channel having only primary coefficients, compared to searching all allowable combinations. For example, when condition 1921 exists, the options "1,1" and "2,2" become "0,1" and "0,2" respectively (i.e., 1954). When condition 1922 exists, the options "1,1" and "2,2" become "1,0" and "2,0" respectively (i.e., 1955). When condition 1920 exists, no secondary transform index needs to be signaled, and the option "0,0" is used. In fact, conditions 1921 and 1922 allow the options "0,1", "0,2", "1,0", and "2,0" in the common tree CU, resulting in higher compression efficiency. If these options are prohibited, either of conditions 1901 or 1910 will cause condition 1920 (i.e., the options "1,1" and "2,2") to be prohibited, such that "0,0" is used (see 1956).

[0260] Signaling the quantization group subdivision level in the slice header provides a higher level of granularity of control below the picture level. The higher level of granularity of control is beneficial for applications where the coding fidelity requirements vary from one part of the picture to another, and in particular for applications where multiple encoders may need to operate slightly independently to provide real-time processing capabilities. Signaling the quantization group subdivision level in the slice header is also consistent with the partition override settings and the scaling list application settings in the slice header.

[0261] In one arrangement of the video encoder 114 and the video decoder 134, the secondary transform index for a chrominance intra prediction block is always set to zero, i.e., the secondary transform is not applied to the chrominance intra prediction block. In this case, there is no need to signal the chrominance secondary transform index, so steps 15130 and 1895 can be omitted, and steps 1360, 1570, and 18100 can be simplified accordingly.

[0262] If a node in the coding tree in the common tree has a region of 64 luma samples, further splitting using a binary or quadtree split will result in smaller luma CUs, such as 4×4 blocks, etc., but will not result in smaller chroma CUs. Instead, there is a single chroma CU with a size corresponding to the region of 64 luma samples, such as a 4×4 chroma CU, etc. Similarly, a coding tree node with a region of 128 luma samples and subject to a ternary split results in a set of smaller luma CUs and one chroma CU. Each luma CU has a corresponding luma secondary transform index, and the chroma CU has a chroma secondary transform index.

[0263] When a node in the coding tree has a region of 64 and signaling a further split or has a region of 128 luma samples and signaling a ternary split, the split is applied only in the luma channel, and the resulting CUs (a number of luma CUs and one chroma CU for each chroma channel) are all intra-predicted or all inter-predicted. When a CU has a width or height of four luma samples and includes one CU for each of the color channels (Y, Cb, and Cr), the chroma CU of the CU has a width or height of two samples. A CU with a width or height of two samples does not utilize the 16-point or 48-point LFNST kernel operations, so no secondary transform is needed. For a block with a width or height of two samples, steps 15130, 1895, 1360, 1570, and 18100 are not required.

[0264] In another arrangement of the video encoder 114 and the video decoder 134, when either or both of the luma and chroma contain non-significant residual coefficients only in the regions of the respective TBs that are only subjected to the primary transform, a single secondary transform index is signaled. If the luma TB contains significant residual coefficients in the non-secondary transform region of the decoded residuals (e.g., 1066, 968), or is indicated to not use DCT-2 as the primary transform, the indicated secondary transform kernel (or secondary transform bypass) is only applied to the chroma TBs. If any of the chroma TBs contain significant residual coefficients in the non-secondary transform region of the decoded residuals, the indicated secondary transform kernel (or secondary transform bypass) is only applied to the luma TB. Even when it is not possible for the chroma TBs, it becomes possible for the luma TB to apply the secondary transform and vice versa, thus giving an increase in coding efficiency compared to the requirement that the last positions of all TBs be within the secondary coefficient domain before any TB of a CU can be subjected to the secondary transform. Additionally, only one secondary transform index is required for a CU in a common coding tree. When the luma primary transform is DCT-2, the secondary transform can be inferred to be disabled for both chroma and luma.

[0265] In another arrangement of the video encoder 114 and the video decoder 134, the secondary transform (by modules 330 and 436 respectively) is only applied to the luma TBs of a CU and not to any of the chroma TBs of that CU. The absence of secondary transform logic for the chroma channels results in lower complexity, such as lower execution time or reduced silicon area. The absence of secondary transform logic for the chroma channels enables only one secondary transform index to be signaled, which can be signaled after the last position of the luma TB. That is, instead of steps 15120 and 1890, steps 1590 and 1860 are performed for the luma TB. In this case, steps 15130 and 1895 are omitted.

[0266] In another arrangement of the video encoder 114 and the video decoder 134, syntax elements that define the quantization group size (i.e., cu_chroma_qp_offset_subdiv and cu_qp_delta_subdiv) are signaled in the PPS 1112. Even if the partitioning constraints are overridden in the slice header 1118, the range of the sub - level values is defined according to the partitioning constraints signaled in the SPS 1110. For example, the ranges of cu_qp_delta_subdiv and cu_chroma_qp_offset_subdiv are defined as 0 to 2*(log2_ctu_size_minus5 + 5 - (MinQtLog2SizeInterY or MinQtLog2SizeIntraY)+MaxMttDepthY_SPS). The value MaxMttDepthY is derived from the SPS 1110. That is, when the current slice is an I - slice, MaxMttDepthY is set to be equal to sps_max_mtt_hierarchy_depth_intra_slice_luma, and when the current slice is a P - or B - slice, MaxMttDepthY is set to be equal to sps_max_mtt_hierarchy_depth_inter_slice. For a slice where the partitioning constraint is overridden to be shallower than the depth signaled in the SPS 1110, if the quantization group subdivision level determined from the PPS 1112 is higher (deeper) than the highest achievable subdivision level at the shallower coding tree depth determined from the slice header, the quantization group subdivision level of the slice is clipped to be equal to the highest achievable subdivision level of the slice. For example, for a particular slice, cu_qp_delta_subdiv and cu_chroma_qp_offset_subdiv are clipped to be within 0 to 2*(log2_ctu_size_minus5 + 5 - (MinQtLog2SizeInterY or MinQtLog2SizeIntraY)+MaxMttDepthY_slice_header), and the clipped values are used for that slice. The value MaxMttDepthY_slice_header is derived from the slice header 1118, i.e., MaxMttDepthY_slice_header is set to be equal to slice_max_mtt_hierarchy_depth_luma.

[0267] In yet another arrangement of the video encoder 114 and the video decoder 134, the subdivision levels are determined based on cu_chroma_qp_offset_subdiv and cu_qp_delta_subdiv decoded from the PPS 1112 to derive the luminance and chrominance subdivision levels. When the partition constraints decoded from the slice header 1118 result in different ranges of the slice's subdivision levels, the subdivision levels applied to the slice are adjusted according to the partition constraints decoded from the SPS 1110 to maintain the same offset relative to the deepest allowed subdivision level. For example, if the SPS 1110 indicates a maximum subdivision level of 4, the PPS 1112 indicates a subdivision level of 3, and the slice header 1118 reduces the maximum value to 3, the subdivision level applied within the slice is set to 2 (maintaining an offset of 1 relative to the maximum allowed subdivision level). Adjusting the quantization group regions to correspond to changes in the partition constraints of a particular slice allows the subdivision levels to be signaled less frequently (i.e., at the PPS level), while providing the granularity to adapt to changes in the slice-level partition constraints. Using an arrangement where the subdivision levels are signaled in the PPS 1112 according to the ranges defined by the partition constraints decoded from the SPS 1110 (where adjustments may be made later based on the overriding partition constraints decoded from the slice header 1118) avoids the parsing dependency problem of making the PPS syntax elements dependent on the partition constraints done in the slice header 1118.

[0268] Industrial Applicability

[0269] The described arrangement is applicable to the computer and data processing industries and, in particular, to digital signal processing for encoding or decoding signals such as video and image signals, thereby achieving high compression efficiency.

[0270] The arrangements described herein increase the flexibility provided to the video encoder when generating a highly compressed bitstream from incoming video data. Quantization of different regions or sub-pictures within a frame can be controlled with varying granularity and different granularity from one region to another, thereby reducing the amount of encoded residual data. When needed, for example, for 360-degree images as described above, a higher granularity can be achieved accordingly.

[0271] In some arrangements, as described with respect to steps 15120 and 15130 (and correspondingly steps 1890 and 1895), the application of the secondary transform can be controlled independently for luminance and chrominance, thereby achieving a further reduction in the encoded residual data. A video decoder is described that has the necessary functionality to decode the bitstream produced by such a video encoder.

[0272] The foregoing merely illustrates some embodiments of the present invention, and the present invention can be modified and / or changed without departing from the scope and spirit of the present invention, where these embodiments are merely exemplary and not restrictive.

[0273] Citation of Related Applications

[0274] This application claims the benefit of priority of Australian Patent Application No. 2019232801, filed on September 17, 2019, under 35 U.S.C. § 119, the entire content of which is hereby incorporated by reference for all purposes.

Claims

1. A method for decoding a coding unit in a coding tree unit of an image from a bitstream, the coding unit having a luminance channel and a chrominance channel, the method comprising: Determining the coding unit having the luminance channel and the chrominance channel according to one or more split flags of the coding tree unit, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a quadtree split; Decoding, from the bitstream, a first index for selecting a non-separable transform kernel for the luminance channel; Selecting the non-separable transform kernel according to the first index; Decoding, from the bitstream, coefficients of a luminance transform block of the luminance channel in the coding unit and coefficients of a chrominance transform block of the chrominance channel in the coding unit; Performing non-separable transform on the coefficients of the luminance transform block by applying the selected non-separable transform kernel to derive non-separable transform coefficients of the luminance transform block; And Decoding the coding unit by performing separable transform on the non-separable transform coefficients of the luminance transform block and on the coefficients of the chrominance transform block, wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting a non-separable transform kernel for the luminance channel can be decoded, (b) a second index for selecting a non-separable transform kernel for the chrominance channel is not decoded, (c) only the non-separable transform can be performed on the coefficients of the luminance transform block in the coding unit, and (d) the non-separable transform is not performed on the coefficients of the chrominance transform block in the coding unit, and the width and height of the chrominance transform block are both equal to or greater than 4; wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there are chrominance coding blocks corresponding to the given area, (a) the first index for selecting a non-separable transform kernel for the luminance channel can exist separately for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting a non-separable transform kernel for the chrominance channel can exist for the chrominance coding blocks corresponding to the given area, and wherein the image has a predetermined chrominance format.

2. A method for encoding a coding unit in a coding tree unit of an image in a bitstream, the coding unit having a luminance channel and a chrominance channel, the method comprising: Determining the coding unit having the luminance channel and the chrominance channel, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a quadtree split; (a) Perform a separable transform on the coefficients of the luminance transform block of the luminance channel in the coding unit to derive the coefficients after the separable transform of the luminance transform block, and (b) perform a separable transform on the coefficients of the chrominance transform block of the chrominance channel in the coding unit to derive the coefficients after the separable transform of the chrominance transform block; Select a non-separable transform kernel for the luminance channel; Perform a non-separable transform on the coefficients after the separable transform of the luminance transform block by applying the selected non-separable transform kernel; And Encode a first index for selecting the non-separable transform kernel for the luminance channel in the bitstream, wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting the non-separable transform kernel for the luminance channel can be encoded, (b) the second index for selecting the non-separable transform kernel for the chrominance channel is not encoded, (c) only the coefficients after the separable transform of the luminance transform block in the coding unit can be subjected to the non-separable transform, and (d) the coefficients after the separable transform of the chrominance transform block in the coding unit are not subjected to the non-separable transform, and the width and height of the chrominance transform block are both equal to or greater than 4; wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there is a chrominance coding block corresponding to the given area, (a) the first index for selecting the non-separable transform kernel for the luminance channel can exist individually for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting the non-separable transform kernel for the chrominance channel can exist for the chrominance coding block corresponding to the given area, and wherein the image has a predetermined chrominance format.

3. An apparatus for decoding a coding unit in a coding tree unit of an image from a bitstream, the coding unit having a luminance channel and a chrominance channel, the apparatus comprising: A determination unit configured to determine a coding unit having the luminance channel and the chrominance channel according to one or more split flags of the coding tree unit, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a quadtree split; A first decoding unit configured to decode from the bitstream a first index for selecting a non-separable transform kernel for the luminance channel; A selection unit configured to select the non-separable transform kernel according to the first index; A second decoding unit configured to decode from the bitstream the coefficients of the luminance transform block of the luminance channel in the coding unit and the coefficients of the chrominance transform block of the chrominance channel in the coding unit; An execution unit configured to perform a non-separable transform on coefficients of the luminance transform block by applying a selected non-separable transform kernel to derive non-separable transform coefficients of the luminance transform block; And A third decoding unit configured to decode the coding unit by performing a separable transform on the non-separable transform coefficients of the luminance transform block and on the coefficients of the chrominance transform block, wherein, when a coding tree of a luminance channel in the coding tree unit is the same as a coding tree of a chrominance channel in the coding tree unit, (a) the first index for selecting a non-separable transform kernel for the luminance channel can be decoded, (b) a second index for selecting a non-separable transform kernel for the chrominance channel is not decoded, (c) the non-separable transform can be performed only on coefficients of a luminance transform block in the coding unit, and (d) the non-separable transform is not performed on coefficients of a chrominance transform block in the coding unit, and a width and a height of the chrominance transform block are both equal to or greater than 4; wherein, when a coding tree of a luminance channel in the coding tree unit is separated from a coding tree of a chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there are chrominance coding blocks corresponding to the given area, (a) the first index for selecting a non-separable transform kernel for the luminance channel can exist separately for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting a non-separable transform kernel for the chrominance channel can exist for the chrominance coding blocks corresponding to the given area, and wherein the image has a predetermined chrominance format.

4. A device for encoding a coding unit in a coding tree unit of an image into a bitstream, the coding unit having a luminance channel and a chrominance channel, the device comprising: A determination unit configured to determine a coding unit having the luminance channel and the chrominance channel, wherein the coding unit can be one of a plurality of coding units obtained by splitting one or more than one from the coding tree unit, and the one or more than one split can include a quadtree split; A first execution unit configured to (a) perform a separable transform on coefficients of a luminance transform block of the luminance channel in the coding unit to derive separable transform coefficients of the luminance transform block, and (b) perform a separable transform on coefficients of a chrominance transform block of the chrominance channel in the coding unit to derive separable transform coefficients of the chrominance transform block; A selection unit configured to select a non-separable transform kernel for the luminance channel; A second execution unit configured to perform a non-separable transform on the separable transform coefficients of the luminance transform block by applying the selected non-separable transform kernel; And A coding unit configured to encode a first index for selecting a non-separable transform kernel for the luminance channel in the bitstream, Wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting the non-separable transform kernel for the luminance channel can be coded, (b) the second index for selecting the non-separable transform kernel for the chrominance channel is not coded, (c) the non-separable transform can be performed only on the coefficients after separable transform of the luminance transform block in the coding unit, and (d) the non-separable transform is not performed on the coefficients after separable transform of the chrominance transform block in the coding unit, and the width and height of the chrominance transform block are both equal to or greater than 4. Wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there are chrominance coding blocks corresponding to the given area, (a) the first index for selecting the non-separable transform kernel for the luminance channel can exist individually for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting the non-separable transform kernel for the chrominance channel can exist for the chrominance coding blocks corresponding to the given area. Wherein, the image has a predetermined chrominance format.

5. A non-transitory computer-readable storage medium, which contains computer-executable instructions that cause a computer to perform the method according to claim 1.

6. A non-transitory computer-readable storage medium, which contains computer-executable instructions that cause a computer to perform the method according to claim 2.

7. A computer program product, which includes a program that causes the computer to perform the method according to claim 1 when the program is executed by the computer.

8. A computer program product, which includes a program that causes the computer to perform the method according to claim 2 when the program is executed by the computer.