Method and apparatus for decoding coding units from a bit stream, method and apparatus for encoding coding units in a bit stream, storage medium and computer program product
Through the universal video encoding (VVC) standard, encoding tree units and flexible quantization parameter adjustments are adopted to solve the efficient encoding problem of high-resolution and high frame rate video data, achieving higher encoding efficiency and flexibility.
Patent Information
- Application Number
- CN202510641807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-17
- Filing Date
- 2020-08-04
- Publication Date
- 2025-07-11
AI Technical Summary
Existing video encoding standards are difficult to achieve efficient compression performance and feasibility in contemporary silicon processes when processing high resolution and high frame rate video data, especially immersive video, and the flexibility of intra- and inter-prediction leads to insufficiency of encoding.
Using the universal video encoding (VVC) standard, efficient encoding of video data is achieved by segmenting the frame into a coding tree unit (CTU) and using a shared or separate coding tree to process the brightness and chrominance channels respectively, combining inseparable secondary transformation and flexible quantization parameter adjustment.
It improves the compression performance of video encoding, adapts to block partition constraints in different parts, reduces encoding workload, improves encoding efficiency, and is suitable for real-time encoding and decoding of high-resolution and high-frame-rate video data.
Smart Images

Figure CN120302068A_ABST
Abstract
Description
[0001] (This application is a divisional application of an application with an application date of August 4, 2020, an application number of 2020800626432, and an invention title of "Method and apparatus for coding units in a coding tree unit for encoding and decoding images, non-transitory computer-readable storage medium, and computer program product".) Technical Field
[0002] The present invention generally relates to digital video signal processing, and in particular, to methods, apparatuses, and systems for encoding and decoding blocks of video samples. The present invention also relates to a computer program product including a computer-readable medium having recorded thereon a computer program for encoding and decoding blocks of video samples. Background Art
[0003] There are currently many applications for video coding, including applications for transmitting and storing video data. Many video coding standards have also been developed and other video coding standards are currently under development. The latest progress in video coding standardization has led to the formation of a group known as the "Joint Video Exploration Team" (JVET). The Joint Video Exploration Team (JVET) includes: members of Study Group 16, Question 6 (SG16 / Q6) of the Telecommunication Standardization Sector (ITU-T) of the International Telecommunication Union (ITU), also known as the "Video Coding Experts Group" (VCEG); and members of Working Group 11 (WG11) of Subcommittee 29, Committee 1, of the Joint Technical Committee 1 of the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC), also known as the "Moving Picture Experts Group" (MPEG).
[0004] The Joint Video Exploration Team (JVET) issued a Call for Proposals (CfP) and analyzed the responses at its 10th meeting held in San Diego, USA. The responses submitted demonstrated that the video compression capabilities were significantly better than those of the current state-of-the-art video compression standard, namely, "High Efficiency Video Coding" (HEVC). Based on this excellent performance, a project was decided to start for developing a new video compression standard named "Versatile Video Coding" (VVC). It is expected that VVC will address the continuous need for even higher compression performance, especially with the increase in the capabilities of video formats (e.g., having higher resolutions and higher frame rates), and the growing market demand for service provision over WAN where the bandwidth cost is relatively high. Use cases such as immersive video require real-time encoding and decoding of such higher formats. For example, cube map projection (CMP) can use 8K format even if the final rendered "viewport" utilizes a lower resolution. VVC must be implementable in contemporary silicon processes and provide an acceptable trade-off between the implemented performance and the implementation cost. For example, the implementation cost can be considered in one or more aspects of silicon area, CPU processor load, memory utilization, and bandwidth. Higher video formats can be processed by dividing the frame area into parts and processing each part in parallel. The bitstream constructed from multiple parts of the compressed frame is still suitable for decoding by a "single-core" decoder, i.e., frame-level constraints (including bitrate) are assigned to each part according to the application requirements.
[0005] Video data includes a sequence of frames of image data, and each frame includes one or more color channels. Generally, one main color channel and two secondary color channels are required. The main color channel is usually referred to as the "luminance" channel, and the (one or more) secondary color channels are usually referred to as the "chrominance" channels. Although video data is usually displayed in the RGB (Red - Green - Blue) color space, this color space has a high correlation between the three corresponding components. The video data representation seen by an encoder or decoder usually uses a color space such as YCbCr. YCbCr concentrates the luminance (mapped to "luminance" according to a transformation equation) in the Y (main) channel and the chrominance in the Cb and Cr (secondary) channels. Due to the use of decorrelated YCbCr signals, the statistics of the luminance channel are significantly different from those of the chrominance channels. The main differences are that after quantization, the chrominance channels contain relatively fewer significant coefficients for a given block compared to the coefficients of the corresponding luminance channel block. In addition, the Cb and Cr channels can be spatially sampled at a lower rate compared to the luminance channel (e.g., half in the horizontal direction and half in the vertical direction (referred to as "4:2:0 chrominance format")). The 4:2:0 chrominance format is commonly used in "consumer" applications such as Internet video streaming, broadcast television, and Blu-ray TMStorage on disk. Subsampling the Cb and Cr channels at half rate in the horizontal direction rather than vertically is known as "4:2:2 chroma format". The 4:2:2 chroma format is commonly used in professional applications, including the capture of footage for movie production and the like. The higher sampling rate of the 4:2:2 chroma format makes the resulting video more resilient to editing operations such as color grading. Before distribution to consumers, 4:2:2 chroma format material is often converted to 4:2:0 chroma format and then encoded for distribution to consumers. In addition to the chroma format, video is also characterized by resolution and frame rate. Example resolutions are Ultra High Definition (UD) with a resolution of 3840×2160 or "8K" with a resolution of 7680×4320, and example frame rates are 60 Hz or 120 Hz. The range of the luminance sample rate can be from about 500 megasamples per second to several thousand megasamples per second. For the 4:2:0 chroma format, the sampling rate of each chroma channel is one quarter of the luminance sample rate, and for the 4:2:2 chroma format, the sampling rate of each chroma channel is half of the luminance sample rate.
[0006] The VVC standard is a "block-based" codec, in which a frame is first segmented into an array of square regions called "Coding Tree Units" (CTUs). A CTU typically occupies a relatively large area, such as 128×128 luminance samples or the like. However, the areas of CTUs at the right and bottom edges of each frame may be smaller. Associated with each CTU is a "coding tree" ("common tree") for both the luminance channel and the chroma channel or separate trees for the luminance channel and the chroma channel respectively. The coding tree defines the decomposition of the area of the CTU into a set of regions, also called "Coding Blocks" (CBs). When using a common tree, a single coding tree specifies the blocks for both the luminance channel and the chroma channel, in which case the set of juxtaposed coding blocks is called a "Coding Unit" (CU), i.e., each CU has coding blocks for each color channel. The CBs are processed in a specific order for encoding or decoding. As a result of using the 4:2:0 chroma format, a CTU of a luminance coding tree including a 128×128 luminance sample area has a corresponding chroma coding tree of a 64×64 chroma sample area juxtaposed with the 128×128 luminance sample area. When a single coding tree is used for the luminance channel and the chroma channel, the set of juxtaposed blocks of a given area is commonly called a "unit", such as the above-mentioned CU as well as "Prediction Unit" (PU) and "Transform Unit" (TU). A single tree of a CU with color channels spanning 4:2:0 chroma format video data results in chroma blocks that are half the width and height of the corresponding luminance blocks. When separate coding trees are used for a given area, the above-mentioned CBs as well as "Prediction Blocks" (PBs) and "Transform Blocks" (TBs) will be used.
[0007] Despite the above difference between "units" and "blocks", the term "block" can be used as a general term for an area or region of a frame to which an operation is applied to all color channels.
[0008] For each CU, a prediction unit (PU) ("prediction unit") that generates the content (sample values) of the corresponding region of the frame data. In addition, a representation of the difference between the prediction seen at the input of the encoder and the region content (or the "residual" in the spatial domain) is formed. The difference for each color channel can be transformed and encoded as a sequence of residual coefficients, thereby forming one or more TUs for a given CU. The transformation applied can be a discrete cosine transform (DCT) or other transform applied to individual blocks of the residual values. The transformation is applied separately, i.e., a two-pass two-dimensional transformation is performed. First, the block is transformed by applying a one-dimensional transformation to each row of samples in the block. Then, the partial result is transformed by applying a one-dimensional transformation to each column of the partial result to produce a final block of transform coefficients that essentially decorrelates the residual samples. The VVC standard supports transforms of various sizes, including transforms of rectangular blocks (each side dimension is a power of 2). The transform coefficients are quantized for entropy encoding in the bitstream.
[0009] The features of VVC are intra-frame prediction and inter-frame prediction. Intra-frame prediction involves using previously processed samples in the frame being used to generate a prediction for the current sample block in that frame. Inter-frame prediction involves using a sample block obtained from a previously decoded frame to generate a prediction for the current sample block in the frame. The sample block obtained from the previously decoded frame is offset from the spatial position of the current block according to a motion vector, which has typically been filtered. The intra-frame prediction block can be (i) a uniform sample value (“DC intra-frame prediction”), (ii) a plane with an offset and horizontal and vertical gradients (“plane intra-frame prediction”), (iii) a group of blocks with adjacent samples applied in a specific direction (“angular intra-frame prediction”), or (iv) the result of a matrix multiplication using adjacent samples and selected matrix coefficients. By encoding the ‘residual’ in the bitstream, the further difference between the prediction block and the corresponding input samples can be corrected to some extent. The residual is typically transformed from the spatial domain to the frequency domain to form residual coefficients (in the “primary transform domain”), and the residual coefficients can be further transformed by applying a “secondary transform” (to produce residual coefficients in the “secondary transform domain”). The residual coefficients are quantized according to a quantization parameter, resulting in a loss of precision in the reconstruction of the samples produced at the decoder, while the bitrate within the bitstream is also reduced. The quantization parameter can vary between frames and within individual frames. For a “rate control” encoder, variation of the intra-frame quantization parameter is typical. Regardless of the statistics of the input samples received (such as noise nature, degree of motion, etc.), the rate control encoder attempts to produce a bitstream with a substantially constant bitrate. Since the bitstream is typically transmitted over a network with a limited bandwidth, rate control is a common technique to ensure reliable performance on the network regardless of the variation of the original frames input to the encoder. In the case where frames are encoded in parallel segments, the flexibility of the use of rate control is desirable because different segments may have different requirements in terms of the desired fidelity. SUMMARY OF THE INVENTION
[0010] An object of the present invention is to substantially overcome or at least ameliorate one or more disadvantages of the existing arrangements.
[0011] One aspect of the present disclosure provides a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and at least one secondary color channel according to a decoded split flag of the encoded tree unit; decoding a first index to select a kernel for the primary color channel, and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index, and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.
[0012] According to another aspect, the first index or the second index is decoded immediately after decoding the position of the last valid residual coefficient of the encoding unit.
[0013] According to another aspect, a single residual coefficient is decoded for a plurality of secondary color channels.
[0014] According to another aspect, a single residual coefficient is decoded for a single secondary color channel.
[0015] According to another aspect, the first index and the second index are independent of each other.
[0016] According to another aspect, the first kernel and the second kernel respectively depend on the intra prediction mode for the primary color channel and the at least one secondary color channel.
[0017] According to another aspect, the first kernel and the second kernel are respectively related to the block size of the primary channel and the block size of the at least one secondary color channel.
[0018] According to another aspect, the second kernel is related to the chrominance subsampling rate of the encoded bitstream.
[0019] According to another aspect, each kernel in the kernel implements an inseparable quadratic transform.
[0020] According to another aspect, the encoding unit includes two secondary color channels, and a separate index is decoded for each of the secondary color channels in the secondary color channels.
[0021] Another aspect of the present disclosure provides a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoded tree unit; selecting a non-separable transform kernel according to a decoded index of the primary color channel; applying the selected non-separable transform kernel to a decoded residual of the primary color channel to generate secondary transform coefficients; and decoding the encoding unit by applying a separable transform kernel to the secondary transform coefficients and applying a separable transform kernel to decoded residuals of the at least one secondary color channel.
[0022] Another aspect of the present disclosure provides a non-transitory computer-readable medium having stored thereon a computer program for implementing a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoded tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.
[0023] Another aspect of the present disclosure provides a video decoder configured to implement a method for decoding an encoding unit of an encoding tree of an encoded tree unit from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoded tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.
[0024] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory to implement a method for decoding an encoding unit of an encoding tree from an image frame in a video bitstream, the encoding unit having a primary color channel and at least one secondary color channel, the method comprising: determining an encoding unit comprising the primary color channel and the at least one secondary color channel according to a decoded split flag of the encoding tree unit; decoding a first index to select a kernel for the primary color channel and decoding a second index to select a kernel for the at least one secondary color channel; selecting a first kernel according to the first index and selecting a second kernel according to the second index; and decoding the encoding unit by applying the first kernel to residual coefficients of the primary color channel and applying the second kernel to residual coefficients of the at least one secondary color channel.
[0025] Another aspect of the present disclosure provides a method for decoding a plurality of encoding units from a bitstream to generate an image frame, the encoding units being the result of decomposition of encoding tree units, the plurality of encoding units forming one or more than one continuous part of the bitstream, the method comprising: determining a subdivision level for each of the one or more than one continuous parts of the bitstream, each subdivision level being applicable to the encoding units of the corresponding continuous part of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on decomposing an encoding tree unit into the encoding units of each continuous part of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region according to the decoded incremental quantization parameter of the region and the quantization parameter of an earlier encoding unit of the image frame; and decoding the plurality of encoding units using the determined quantization parameter for each region to generate an image frame.
[0026] According to another aspect, each region is based on a comparison of the subdivision level associated with the encoding unit and the determined subdivision level of the corresponding continuous part.
[0027] According to another aspect, a quantization parameter increment is determined for each region, wherein the corresponding encoding tree has a subdivision level less than or equal to the determined subdivision level of the corresponding continuous part.
[0028] According to another aspect, a new region is set for any node in an encoding tree unit having a subdivision level less than or equal to the determined subdivision level of the corresponding continuous part.
[0029] According to another aspect, the subdivision level determined for each continuous part includes a first subdivision level for the luminance encoding units of the continuous part and a second subdivision level for the chrominance encoding units of the continuous part.
[0030] According to another aspect, the first subdivision level and the second subdivision level are different.
[0031] According to another aspect, the method further includes: decoding a flag indicating a partitioning constraint for a sequence parameter set associated with a bitstream that can be rewritten.
[0032] According to another aspect, the determined subdivision level for each of one or more consecutive portions includes the maximum luminance coding unit depth for the region.
[0033] According to another aspect, the determined subdivision level for each of one or more consecutive portions includes the maximum chrominance coding unit depth for the corresponding region.
[0034] According to another aspect, the determined subdivision level for one of the consecutive portions is adjusted to maintain an offset relative to the deepest allowed subdivision level decoded for the partitioning constraint of the bitstream.
[0035] Another aspect of the present disclosure provides a non-transitory computer-readable medium storing a computer program for implementing a method of decoding a plurality of coding units from a bitstream to produce an image frame, the coding units being the result of a decomposition of coding tree units, the plurality of coding units forming one or more consecutive portions of the bitstream, the method including: determining a subdivision level for each of one or more consecutive portions of the bitstream, each subdivision level being applicable to the coding units of the corresponding consecutive portion of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of coding tree units into the coding units of the respective consecutive portions of the bitstream and the corresponding determined subdivision levels; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to produce an image frame.
[0036] Another aspect of the present disclosure provides a video decoder configured to implement a method of decoding a plurality of coding units from a bitstream to produce an image frame, the coding units being the result of a decomposition of coding tree units, the plurality of coding units forming one or more consecutive portions of the bitstream, the method including: determining a subdivision level for each of one or more consecutive portions of the bitstream, each subdivision level being applicable to the coding units of the corresponding consecutive portion of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of coding tree units into the coding units of the respective consecutive portions of the bitstream and the corresponding determined subdivision levels; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to produce an image frame.
[0037] Another aspect of the present disclosure provides a system, including: a memory; and a processor, wherein the processor is configured to execute code stored on the memory to implement a method of decoding a plurality of coding units from a bitstream to generate an image frame, the coding units being a decomposition result of coding tree units, the plurality of coding units forming one or more consecutive portions of the bitstream, the method including: determining a subdivision level of each of one or more consecutive portions of the bitstream, each subdivision level being applicable to the coding units of the corresponding consecutive portion of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on decomposing a coding tree unit into the coding units of each consecutive portion of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to generate an image frame.
[0038] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] At least one embodiment of the present invention will now be described with reference to the following drawings and appendices, wherein:
[0040] Figure 1 is a schematic block diagram showing a video encoding and decoding system;
[0041] Figure 2A and 2B constitute a schematic block diagram of a general-purpose computer system that can implement one or both of the video encoding and decoding systems that can be practiced Figure 1 ;
[0042] Figure 3 is a schematic block diagram showing the functional modules of a video encoder;
[0043] Figure 4 is a schematic block diagram showing the functional modules of a video decoder;
[0044] Figure 5 is a schematic block diagram showing the available partitions of a block into one or more blocks in the tree structure of general video coding;
[0045] Figure 6 is a schematic diagram of a data stream for implementing a permitted partition of a block into one or more blocks in the tree structure of general video coding;
[0046] Figure 7A and 7B show an example partition of a coding tree unit (CTU) into a plurality of coding units (CUs);
[0047] Figure 8A ,8B And FIGS. 8C illustrate the levels of subdivision resulting from splitting in the coding tree and their impact on partitioning the coding tree units into quantization groups;
[0048] Figure 9A and 9B illustrate the 4×4 transform block scan pattern and the associated primary and secondary transform coefficients;
[0049] Figure 9C and 9D illustrate the 8×8 transform block scan pattern and the associated primary and secondary transform coefficients;
[0050] Figure 10 illustrate the regions where the secondary transform is applied for transform blocks of various sizes;
[0051] Figure 11 illustrate the syntax structure of a bitstream having multiple strips, each strip including multiple coding units;
[0052] Figure 12 illustrate the syntax structure of a bitstream having a common tree for luminance and chrominance coding blocks of coding tree units;
[0053] Figure 13 illustrate a method of encoding a frame in a bitstream including one or more than one strip as a sequence of coding units;
[0054] Figure 14 illustrate a method of encoding a strip header in a bitstream;
[0055] Figure 15 illustrate a method of encoding a coding unit in a bitstream;
[0056] Figure 16 illustrate a method of decoding a frame from a bitstream that is a sequence of coding units arranged as strips;
[0057] Figure 17 illustrate a method of decoding a strip header from a bitstream;
[0058] Figure 18 illustrate a method of decoding a coding unit from a bitstream; and
[0059] Figure 19A and 19B illustrate the rules for applying or bypassing the secondary transform for the luminance and chrominance channels. DETAILED DESCRIPTION
[0060] In the case of referring to steps and / or features having the same reference numerals in any one or more than one of the figures, unless a contrary intention appears, these steps and / or features have the same one or more functions or one or more operations for the purposes of this specification.
[0061] Rate control video encoders require the flexibility to adjust quantization parameters at a granularity suitable for block partitioning constraints. The block partitioning constraints may vary between portions of a frame, e.g., where multiple video encoders operate in parallel to compress respective frames. The granularity of the regions requiring quantization parameter adjustment varies accordingly. Additionally, control over the transform selection applied (including the potential application of a secondary transform) is applied within the scope of generating a prediction signal for the residual being transformed. In particular, for intra prediction, separate modes can be used for luminance blocks and chrominance blocks as they can use different intra prediction modes.
[0062] Some portions of a video contribute more to the fidelity of a rendered viewport than other portions and can be allocated a greater bitrate as well as greater flexibility in terms of block structure and variance of quantization parameters. Portions that contribute little to the fidelity of the rendered viewport (such as those on the sides or behind of the rendered view) can be compressed using a simpler block structure to reduce the encoding effort and with less flexibility in terms of control of quantization parameters. Generally, larger values are selected to more coarsely quantize transform coefficients at lower bitrates. Additionally, the application of transform selection can be independent between the luminance and chrominance channels to further simplify the encoding process by avoiding the need to jointly consider luminance and chrominance for transform selection. In particular, after separately considering intra prediction modes for luminance and chrominance, the need to jointly consider luminance and chrominance for secondary transform selection is avoided.
[0063] Figure 1 is a schematic block diagram showing the functional modules of a video encoding and decoding system 100. System 100 can vary the regions in which quantization parameters are adjusted in different portions of a frame to accommodate different block partitioning constraints that may be in effect in various portions of the frame.
[0064] System 100 includes a source device 110 and a destination device 130. A communication channel 120 is used to communicate encoded video information from the source device 110 to the destination device 130. In some configurations, one or both of the source device 110 and the destination device 130 can include a cellular phone handset or a “smartphone,” in which case the communication channel 120 is a wireless channel. In other configurations, the source device 110 and the destination device 130 can include video conferencing equipment, in which case the communication channel 120 is typically a wired channel such as an Internet connection. Additionally, the source device 110 and the destination device 130 can include a wide range of any devices, where these devices include those supporting over-the-air television broadcasts, cable television applications, Internet video applications (including streaming), and applications for capturing encoded video data on some computer-readable storage media such as a hard drive in a file server, etc.
[0065] As shown Figure 1 in FIG. 1, the source device 110 includes a video source 112, a video encoder 114, and a transmitter 116. The video source 112 generally includes a source of captured video frame data (represented as 113), such as a camera sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote camera sensor. The video source 112 may also be the output of a computer graphics card (e.g., the video output that displays the operating system and various applications executing on a computing device (e.g., a tablet computer)). Examples of source devices 110 that may include a camera sensor as the video source 112 include smart phones, video camcorders, professional cameras, and network video cameras.
[0066] The video encoder 114 converts (or “encodes”) the captured frame data (indicated by arrow 113) from the video source 112 into a bitstream (indicated by arrow 115). The bitstream 115 is transmitted by the transmitter 116 via a communication channel 120 as encoded video data (or “encoded video information”). The bitstream 115 may also be stored in a non-transitory storage device 122, such as a “flash” memory or a hard disk drive, until subsequently transmitted via the communication channel 120 or as an alternative to transmission via the communication channel 120. For example, the encoded video data may be supplied to a customer via a wide area network (WAN) for video streaming applications when needed.
[0067] The destination device 130 includes a receiver 132, a video decoder 134, and a display device 136. The receiver 132 receives the encoded video data from the communication channel 120 and passes the received video data as a bitstream (indicated by arrow 133) to the video decoder 134. The video decoder 134 then outputs the decoded frame data (indicated by arrow 135) to the display device 136. The decoded frame data 135 has the same chrominance format as the frame data 113. Examples of display devices 136 include cathode ray tubes, liquid crystal displays (such as in smart phones, tablet computers, computer monitors, or stand-alone televisions, etc.). The functions of the source device 110 and the destination device 130 may also be embodied in a single device, examples of which include mobile phone handsets and tablet computers. The decoded frame data may be further transformed before being presented to the user. For example, a “viewport” with specific latitude and longitude may be rendered from the decoded frame data using a projection format to represent a 360° view of a scene.
[0068] Although the example devices are described above, the source device 110 and the destination device 130 may each generally be configured within a general-purpose computer system via a combination of hardware components and software components. Figure 2AA computer system 200 is shown, which includes: a computer module 201; input devices such as a keyboard 202, a mouse indicator device 203, a scanner 226, a camera 227 that can be configured as a video source 112, and a microphone 280; and output devices including a printer 215, a display device 214 that can be configured as a display device 136, and a speaker 217. The computer module 201 can communicate with a communication network 220 via a wiring 221 using an external modulator-demodulator (modem) transceiver device 216. The communication network 220, which can represent the communication channel 120, can be a WAN, such as the Internet, a cellular telecommunications network, or a private WAN, etc. In the case where the wiring 221 is a telephone line, the modem 216 can be a traditional "dial-up" modem. Alternatively, in the case where the wiring 221 is a high-capacity (e.g., cable or optical) wiring, the modem 216 can be a broadband modem. A wireless modem can also be used for a wireless connection to the communication network 220. The transceiver device 216 can provide the functions of a transmitter 116 and a receiver 132, and the communication channel 120 can be embodied in the wiring 221.
[0069] The computer module 201 generally includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 can have a semiconductor random access memory (RAM) and a semiconductor read-only memory (ROM). The computer module 201 also includes a plurality of input / output (I / O) interfaces, where the plurality of input / output (I / O) interfaces include: an audio-video interface 207, which is connected to the video display 214, the speaker 217, and the microphone 280; an I / O interface 213, which is connected to the keyboard 202, the mouse 203, the scanner 226, the camera 227, and an optional joystick or other human-machine interface device (not shown); and an interface 208 for the external modem 216 and the printer 215. The signal from the audio-video interface 207 to the computer monitor 214 is generally the output of a computer graphics card. In some implementations, the modem 216 can be built into the computer module 201, for example, built into the interface 208. The computer module 201 also has a local network interface 211, where the local network interface 211 allows the computer system 200 to be connected to a local communication network 222 known as a local area network (LAN) via a wiring 223. As Figure 2A shown, the local communication network 222 can also be connected to the wide area network 220 via a wiring 224, where the local communication network 222 generally includes a so-called "firewall" device or a device with similar functions. The local network interface 211 can include an Ethernet ( TM ) circuit card, Bluetooth ( TM)Wireless configuration or IEEE 802.11 wireless configuration; however, for interface 211, a variety of other types of interfaces can be practiced. The local network interface 211 can also provide the functions of transmitter 116 and receiver 132, and the communication channel 120 can also be embodied in the local communication network 222.
[0070] I / O interfaces 208 and 213 can provide either or both of serial and parallel connections, where the former is typically implemented according to the Universal Serial Bus (USB) standard and has a corresponding USB connector (not shown). A storage device 209 is provided, and the storage device 209 typically includes a hard disk drive (HDD) 210. Other storage devices such as floppy disk drives and tape drives (not shown) can also be used. An optical disc drive 212 is typically provided to serve as a non-volatile source of data. Portable memory devices such as optical discs (e.g., CD-ROM, DVD, Blu-ray Disc TM ), USB-RAM, portable external hard disk drives, and floppy disks can be used as suitable sources of data for the computer system 200. Typically, any of the HDD 210, optical disc drive 212, networks 220 and 222 can also be configured to operate as the video source 112 or as the destination for decoded video data to be stored for playback via the display 214. The source device 110 and destination device 130 of the system 100 can be embodied in the computer system 200.
[0071] The components 205 - 213 of the computer module 201 typically communicate via the interconnect bus 204 and in a manner that results in the traditional operating modes of the computer system 200 known to those skilled in the relevant art. For example, the processor 205 is connected to the system bus 204 using wiring 218. Similarly, the memory 206 and the optical disc drive 212 are connected to the system bus 204 via wiring 219. Examples of computers that can practice the described configuration include IBM-PCs and compatible machines, Sun SPARCstations, Apple Mac TM or similar computer systems.
[0072] In appropriate or desired cases, the computer system 200 can be used to implement the video encoder 114 and video decoder 134 and the methods described below. In particular, the video encoder 114, video decoder 134, and the methods to be described can be implemented as one or more software applications 233 executable within the computer system 200. In particular, using the instructions 231 (reference Figure 2B) to implement the steps of the video encoder 114, the video decoder 134, and the method. The software instructions 231 can be formed into one or more code modules each for performing one or more specific tasks. The software can also be split into two separate parts, where the first part and the corresponding code modules perform the method, and the second part and the corresponding code modules manage the user interface between the first part and the user.
[0073] For example, the software can be stored in a computer-readable medium including the storage devices described below. The software is loaded from the computer-readable medium into the computer system 200 and then executed by the computer system 200. The computer-readable medium having such software or the computer program recorded on the computer-readable medium is a computer program product. Using the computer program product in the computer system 200 preferably implements an advantageous device for implementing the video encoder 114, the video decoder 134, and the method.
[0074] Generally, the software 233 is stored in the HDD 210 or the memory 206. The software is loaded from the computer-readable medium into the computer system 200 and executed by the computer system 200. Thus, for example, the software 233 can be stored on an optically readable disk storage medium (e.g., CD-ROM) 225 read by the optical disk drive 212.
[0075] In some instances, the application program 233 is supplied to the user in a manner encoded on one or more CD-ROMs 225 and read via the corresponding drive 212, or alternatively, the user can read the application program 233 from the network 220 or 222. Further, the software can also be loaded into the computer system 200 from other computer-readable media. A computer-readable storage medium refers to any non-transitory tangible storage medium that provides the recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray Discs TM ) (Blu-ray Disc), hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical disks, or computer-readable cards such as PCMCIA cards, etc., regardless of whether these devices are inside or outside the computer module 201. Examples of transitory or non-tangible computer-readable transmission media that can also participate in providing software, application programs, instructions, and / or video data or encoded video data to the computer module 401 include: radio or infrared transmission channels and network wiring to other computers or networked devices, and the Internet or intranet including email transmissions and information recorded on websites.
[0076] The second part of the above-described application 233 and the corresponding code modules can be executed to implement one or more graphical user interfaces (GUIs) to be drawn or otherwise presented on the display 214. By typically operating the keyboard 202 and the mouse 203, the user and applications of the computer system 200 can operate the interface in a functionally applicable manner to provide control commands and / or inputs to the applications associated with these (one or more) GUIs. Other functionally applicable forms of user interfaces can also be implemented, such as an audio interface that utilizes voice prompts output via the speaker 217 and user voice commands input via the microphone 280, etc.
[0077] Figure 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents Figure 2A the logical aggregation of all memory modules accessible by the computer module 201 in the (including the HDD 209 and the semiconductor memory 206).
[0078] In the case of initially powering on the computer module 201, a power-on self-test (POST) program 250 is executed. The POST program 250 is typically stored in Figure 2A the ROM 249 of the semiconductor memory 206. Sometimes a hardware device such as the ROM 249 storing software is called firmware. The POST program 250 checks the hardware within the computer module 201 to ensure proper operation, and typically checks the processor 205, the memory 234 (209, 206), and the basic input-output system software (BIOS) module 251, which is usually also stored in the ROM 249, for correct operation. Once the POST program 250 runs successfully, the BIOS 251 starts Figure 2A the hard disk drive 210. Starting the hard disk drive 210 causes the boot loader 252 residing on the hard disk drive 210 to be executed via the processor 205. In this way, the operating system 253 is loaded into the RAM memory 206, where the operating system 253 starts to work. The operating system 253 is a system-level application executable by the processor 205 to implement various high-level functions including processor management, memory management, device management, storage management, software application interfaces, and a general user interface.
[0079] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory to execute without conflicting with the memory allocated to other processes. In addition, Figure 2Athe different types of memory available in the computer system 200 so that each process can run efficiently. Thus, the aggregated memory 234 is not intended to illustrate how to allocate a particular segment of memory (unless otherwise stated), but rather provides an overview of the memory accessible to the computer system 200 and how that memory is used.
[0080] As Figure 2B shown, the processor 205 includes a plurality of functional modules, where the plurality of functional modules includes a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a plurality of storage registers 244 - 246 in the register section. One or more internal buses 241 functionally interconnect these functional modules. The processor 205 generally also has one or more interfaces 242 for communicating with external devices via the system bus 204 using wiring 218. The memory 234 is connected to the bus 204 using wiring 219.
[0081] The application program 233 includes an instruction sequence 231 that may include conditional branch instructions and loop instructions. The program 233 may also include data 232 used when executing the program 233. The instructions 231 and data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237, respectively. Depending on the relative sizes of the instructions 231 and the memory locations 228 - 230, a particular instruction may be stored in a single memory location as described by the instruction shown in memory location 230. Optionally, as described by the instruction segments shown in memory locations 228 and 229, an instruction may be split into multiple parts each stored in a separate memory location.
[0082] Typically, a set of instructions is given to the processor 205, where that set of instructions is executed within the processor 205. The processor 205 waits for subsequent input, where the processor 205 reacts to that subsequent input by executing another set of instructions. The input may be provided from one or more of a plurality of sources, where the input includes data generated by one or more of the input devices 202, 203, data received from an external source via one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into the corresponding reader 212 (all of which are shown in Figure 2A ). Executing a set of instructions may in some cases result in output data. Execution may also involve storing data or variables to the memory 234.
[0083] The video encoder 114, the video decoder 134, and the method may use the input variables 254 stored in the respective memory locations 255, 256, 257 within the memory 234. The video encoder 114, the video decoder 134, and the method generate the output variables 261 stored in the respective memory locations 262, 263, 264 within the memory 234. Intermediate variables 258 may be stored in the memory locations 259, 260, 266, and 267.
[0084] Reference Figure 2B The processor 205, registers 244, 245, 246, arithmetic logic unit (ALU) 240, and control unit 239 work together to perform a sequence of micro-operations, where these micro-operation sequences are required for the "fetch, decode, and execute" cycles for each instruction in the instruction set that makes up the program 233. Each fetch, decode, and execute cycle includes:
[0085] A fetch operation for fetching or reading an instruction 231 from the memory locations 228, 229, 230;
[0086] A decode operation, in which the control unit 239 determines which instruction was fetched; and
[0087] An execute operation, in which the control unit 239 and / or the ALU 240 execute the instruction.
[0088] After that, a further fetch, decode, and execute cycle for the next instruction can be performed. Similarly, a store cycle can be performed, in which the control unit 239 stores or writes a value to the memory location 232.
[0089] To illustrate Figures 13 to 18 Each step or sub-process in the method is associated with one or more segments of the program 233, and is typically performed by the register section 244, 245, 247, ALU 240, and control unit 239 in the processor 205 working together to perform the fetch, decode, and execute cycles for each instruction in the segmented instruction set of the program 233.
[0090] Figure 3 is a schematic block diagram showing the functional modules of the video encoder 114. Figure 4 is a schematic block diagram showing the functional modules of the video decoder 134. Generally, data is transferred between the functional modules within the video encoder 114 and the video decoder 134 in groups of samples or coefficients (such as the division of a block into fixed-size sub-blocks, etc.) or as an array. As Figure 2A and 2BAs shown, a general-purpose computer system 200 can be used to implement video encoder 114 and video decoder 134, where various functional modules can be implemented using dedicated hardware within computer system 200, using software executable within computer system 200 (such as one or more software code modules of software application 233 residing on hard disk drive 205 and controlled by processor 205 for its execution), etc. Alternatively, video encoder 114 and video decoder 134 can be implemented using a combination of dedicated hardware and software executable within computer system 200. Video encoder 114, video decoder 134, and the method can alternatively be implemented in dedicated hardware such as one or more integrated circuits performing the functions or sub-functions of the method. Such dedicated hardware can include a graphics processing unit (GPU), a digital signal processor (DSP), an application specific standard product (ASSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or one or more microprocessors and associated memories. In particular, video encoder 114 includes modules 310 - 390, and video decoder 134 includes modules 420 - 496, where each of these modules can be implemented as one or more software code modules of software application 233.
[0091] Although Figure 3 video encoder 114 of is an example of a general video coding (VVC) video coding pipeline, other video codecs can also be used for the processing stages described herein. Video encoder 114 receives captured frame data 113 such as a series of frames (each frame including one or more color channels). Frame data 113 can be in any chroma format, such as 4:0:0, 4:2:0, 4:2:2, or 4:4:4 chroma format. Block partitioner 310 first divides frame data 113 into CTUs, the shape of CTUs is typically square, and is configured such that a specific size of CTU is used. For example, the size of CTU can be 64×64, 128×128, or 256×256 luma samples. Block partitioner 310 further divides each CTU into one or more CBs corresponding to a luma coding tree or a chroma coding tree. The luma channel can also be referred to as the primary color channel. Each chroma channel can also be referred to as a secondary color channel. CBs have various sizes and can include both square and non-square aspect ratios. Refer to Figures 13 - 15 for a further description of the operation of block partitioner 310. However, in the VVC standard, CBs, CUs, PUs, and TUs always have side lengths that are powers of 2. Thus, the current CB (denoted as 312) is output from block partitioner 310, advancing according to the iteration of one or more blocks of the CTU, according to the luma coding tree and chroma coding tree of the CTU. Refer to the following Figure 5 and 6To further illustrate the options for partitioning a CTU into CBs. Although operations are generally described on a CTU-by-CTU basis, the video encoder 114 and the video decoder 134 can operate on smaller-sized regions to reduce memory consumption. For example, each CTU can be divided into smaller regions, called "virtual pipeline data units" (VPDUs) of size 64×64. The VPDUs form a data granularity that is more suitable for pipelining in the hardware architecture, where the reduced memory footprint compared to operating on a full CTU reduces the silicon area and thus the cost.
[0092] The CTUs obtained from the first split of the frame data 113 can be scanned in raster scan order and can be grouped into one or more than one "slice". A slice can be an "intra" (or "I") slice. An intra slice (I slice) indicates that each CU in the slice is intra-predicted. Optionally, a slice can be single-predicted or bi-predicted (a "P" or "B" slice, respectively), indicating the additional availability of single prediction and bi-prediction in the slice, respectively.
[0093] In an I slice, the coding tree of each CTU can diverge into two separate coding trees below the 64×64 level, one for luminance and the other for chrominance. Using separate trees allows different block structures to exist between the luminance and chrominance within the 64×64 luminance region of the CTU. For example, large chrominance CBs can be juxtaposed with many smaller luminance CBs, and vice versa. In a P or B slice, a single coding tree of the CTU defines the block structure common to luminance and chrominance. The resulting blocks of the single tree can be intra-predicted or inter-predicted.
[0094] For each CTU, the video encoder 114 operates in two stages. In the first stage (called the "search" stage), the block partitioner 310 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select the CB that provides a relatively high compression efficiency and a relatively low distortion. This test typically involves Lagrangian optimization, whereby candidate CBs are evaluated based on a weighted combination of rate (coding cost) and distortion (error with respect to the input frame data 113). The "best" candidate CB (the CB with the lowest evaluated rate / distortion) is selected for subsequent encoding in the bitstream 115. Options included in the evaluation of candidate CBs are: using the CB for a given region, or splitting the region according to various split options and encoding each smaller resulting region or further splitting the region using other CBs. As a result, both the coding tree and the CB itself are selected in the search stage.
[0095] Video encoder 114 generates a predicted block (PB) indicated by arrow 320 for each CB (e.g., CB 312). PB 320 is a prediction of the content of the associated CB 312. Subtractor module 322 generates a difference (or "residual," which refers to the difference in the spatial domain) represented as 324 between PB 320 and CB 312. The difference 324 is the block-size difference between the corresponding samples in PB 320 and CB 312. The difference 324 is transformed, quantized, and represented as a transformed block (TB) indicated by arrow 336. PB 320 and the associated TB 336 are typically selected from among multiple possible candidate CBs (e.g., based on an evaluated cost or distortion).
[0096] A candidate coding block (CB) is a CB obtained from one of the prediction modes available to video encoder 114 for the associated PB and the resulting residual. When combined with the predicted PB in video decoder 114, TB 336 reduces the difference between the decoded CB and the original CB 312 at the cost of additional signaling in the bitstream.
[0097] Thus, each candidate coding block (CB) (i.e., the combination of a predicted block (PB) and a transformed block (TB)) has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of a CB is typically estimated as a difference in sample values, such as the sum of absolute differences (SAD) or the sum of squared differences (SSD). Pattern selector 386 can use difference 324 to determine an estimate obtained from each candidate PB to determine prediction mode 387. Prediction mode 387 indicates the decision to use a particular prediction mode (e.g., intra prediction or inter prediction) for the current CB. The estimation of the coding cost associated with each candidate prediction mode and the corresponding residual coding can be performed at a cost significantly lower than that of the entropy coding of the residual. Thus, even in a real-time video encoder, multiple candidate modes can be evaluated to determine the best mode in terms of rate-distortion.
[0098] Determining the best mode in terms of rate-distortion is typically achieved using a variant of Lagrangian optimization.
[0099] A Lagrangian or similar optimization process can be employed for both the selection of the best partition of a CTU into CBs (using block partitioner 310) and the selection of the best prediction mode from among multiple possibilities. By applying the Lagrangian optimization process for candidate modes in mode selector module 386, the intra prediction mode with the lowest cost measurement is selected as the best mode. The lowest-cost mode is the selected secondary transform index 388 and is also encoded in bitstream 115 by entropy encoder 338.
[0100] In a second stage of operation of video encoder 114, referred to as the "encoding" stage, an iteration of the (one or more) determined coding trees for each CTU is performed in video encoder 114. For a CTU using a separate tree, for each 64×64 luminance region of the CTU, the luminance coding tree is first encoded, followed by the chrominance coding tree. Only luminance CUs are encoded within the luminance coding tree, and only chrominance CUs are encoded within the chrominance coding tree. For a CTU using a common tree, a single tree describes the CUs according to the common block structure of the common tree, i.e., the luminance CUs and chrominance CUs.
[0101] Entropy encoder 338 supports both variable length coding of syntax elements and arithmetic coding of syntax elements. Parts of the bitstream such as "parameter sets" (e.g., sequence parameter set (SPS) and picture parameter set (PPS)) use a combination of fixed length codewords and variable length codewords. A slice (also referred to as a consecutive portion) has a slice header using variable length coding, followed by slice data using arithmetic coding. The slice header defines parameters specific to the current slice, such as slice-level quantization parameter offset, etc. The slice data includes the syntax elements of each CTU in the slice. Using variable length coding and arithmetic coding requires sequential parsing within each part of the bitstream. These parts can be described with start codes to form "network abstraction layer units" or "NAL units". Context adaptive binary arithmetic coding processing is used to support arithmetic coding. The syntax elements for arithmetic coding consist of a sequence of one or more "bins (binary files)". Like bits, the value of a bin is "0" or "1". However, the bins are not encoded as discrete bits in bitstream 115. A bin has an associated prediction (or "likely" or "most probable") value and an associated probability (referred to as "context"). When the actual bin to be encoded matches the prediction value, the "most probable symbol" (MPS) is encoded. Encoding the most probable symbol is relatively inexpensive in terms of the bits consumed in bitstream 115 (including a cost of less than one discrete bit in total). When the actual bin to be encoded does not match the likely value, the "least probable symbol" (LPS) is encoded. Encoding the least probable symbol has a relatively high cost in terms of bits consumed. The bin coding technique enables efficient coding of bins where the probability of "0" vs "1" is skewed. For a syntax element with two possible values (i.e., "flag"), a single bin is sufficient. For a syntax element with many possible values, a sequence of bins is required.
[0102] The existence of a later bin in the sequence can be determined based on the value of an earlier bin in the sequence. Additionally, each bin can be associated with more than one context. A particular context can be selected based on the earlier bin in the syntactic element and the bin values of adjacent syntactic elements (i.e., the bin values from adjacent blocks), etc. Each time a context-coded bin is coded, the context selected for that bin (if any) is updated in a way that reflects the new bin value. In this way, the binary arithmetic coding scheme is considered adaptive.
[0103] Video encoder 114 also supports bins that lack context ("bypass bins"). The bypass bins are coded assuming an equiprobable distribution between "0" and "1". Thus, each bin has a coding cost of one bit in bitstream 115. The lack of context saves memory and reduces complexity, thus using bypass bins where the distribution of the values of a particular bin is not skewed. An example of an entropy encoder that uses context and is adaptive is known in the art as CABAC (Context-Adaptive Binary Arithmetic Coder), and many variants of this encoder have been adopted in video coding.
[0104] Entropy encoder 338 uses a combination of context-coded bins and bypass-coded bins to code quantization parameter 392, and if for the current CB, codes the LFNST index 388. The quantization parameter 392 is coded using "delta QP". In each region called a "quantization group", the delta QP is signaled at most once. The quantization parameter 392 is applied to the residual coefficients of the luminance CB. The adjusted quantization parameter is applied to the residual coefficients of the collocated chrominance CB. The adjusted quantization parameter can include a mapping from the luminance quantization parameter 392 according to a mapping table and a CU-level offset selected from an offset list. The secondary transform index 388 is signaled when the residual associated with the transform block includes valid residual coefficients only in those coefficient positions that are transformed into primary coefficients by applying a secondary transform.
[0105] The multiplexer module 384 outputs the PB 320 from the intra prediction module 364 according to the determined best intra prediction mode selected from the test prediction modes of the respective candidate CBs. The candidate prediction modes need not include every conceivable prediction mode supported by the video encoder 114. Intra prediction is divided into three types. "DC intra prediction" involves filling the PB with a single value representing the average of nearby reconstructed samples. "Planar intra prediction" involves filling the PB with samples according to a plane, where the DC offset and the vertical and horizontal gradients are derived from nearby reconstructed neighboring samples. The nearby reconstructed samples typically include a row of reconstructed samples above the current PB (extending a certain extent to the right of the PB) and a column of reconstructed samples to the left of the current PB (extending a certain extent downwards outside the PB). "Angular intra prediction" involves filling the PB with reconstructed neighboring samples that are filtered and propagated across the PB in a particular direction (or "angle"). In VVC, 65 angles are supported, where rectangular blocks can utilize additional angles not available to square blocks to yield a total of 87 angles. A fourth type of intra prediction can be used for chrominance PBs, thus generating the PB from collocated luma reconstructed samples according to the "cross-component linear model" (CCLM) mode. Three different CCLM modes are available, each mode using a different model derived from neighboring luma and chroma samples. The derived model is used to generate a sample block for the chrominance PB from collocated luma samples.
[0106] In cases where previously reconstructed samples are not available (e.g., at the edges of a frame), a default halftone value of half the sample range is used. For example, for 10-bit video, a value of 512 is used. Since no previous samples are available for the CB located at the upper left position of the frame, the angular and planar intra prediction modes produce the same output as the DC prediction mode, i.e., a flat plane of samples with the halftone value as the amplitude.
[0107] For inter-frame prediction, the motion compensation module 380 uses samples from one or two frames before the current frame in the order of the coded frames in the bitstream to generate a prediction block 382 and outputs it as PB 320 by the multiplexer module 384. In addition, for inter-frame prediction, a single coding tree is typically used for both the luminance channel and the chrominance channel. The order of the coded frames in the bitstream may be different from the order of the frames when captured or displayed. When one frame is used for prediction, the block is referred to as "single prediction" and has two associated motion vectors. When two frames are used for prediction, the block is referred to as "dual prediction" and has two associated motion vectors. For P slices, each CU can be intra-frame predicted or single predicted. For B slices, each CU can be intra-frame predicted, single predicted, or dual predicted. A "group of pictures" structure is typically used to code frames, thus implementing a temporal hierarchy of frames. A frame can be divided into multiple slices, each slice coding a part of the frame. The temporal hierarchy of frames allows a frame to reference previous and subsequent pictures in the order of the displayed frames. The images are coded in an order necessary to ensure that the dependencies of each frame are satisfied during decoding.
[0108] Samples are selected according to the motion vector 378 and the reference picture index. The motion vector 378 and the reference picture index apply to all color channels, and thus inter-frame prediction is mainly described in terms of operations on PUs rather than PBs, i.e., a single coding tree is used to describe the decomposition of each CTU into one or more inter-frame prediction blocks. Inter-frame prediction methods may vary in the number and precision of the motion parameters. The motion parameters typically include a reference frame index (which indicates which reference frames from the reference frame list will be used plus the respective spatial translations of the reference frames), but may include more frames, special frames, or complex affine parameters such as scaling and rotation. Additionally, a predetermined motion refinement process can be applied to generate a dense motion estimate based on the reference sample block.
[0109] In the case where PB 320 is determined and selected and subtracted from the original sample block at subtractor 322, the residual with the lowest coding cost (denoted as 324) is obtained and lossy compressed. The lossy compression process includes the steps of transformation, quantization, and entropy coding. The forward primary transformation module 326 applies a forward transformation to the difference 324, thereby converting the difference 324 from the spatial domain to the frequency domain and generating the primary transformation coefficients represented by arrow 328. The maximum primary transformation size in one dimension is a 32-point DCT-2 or a 64-point DCT-2 transformation. If the CB being encoded is greater than the maximum supported primary transformation size represented as the block size (i.e., 64×64 or 32×32), the primary transformation 326 is applied in a block manner to transform all samples of the difference 324. The application of the transformation 326 results in multiple TBs of the CB. In the case where each transformation is operated on a TB of the difference 324 larger than 32×32 (e.g., 64×64), all resulting primary transformation coefficients 328 outside the upper left 32×32 region of the TB are set to zero, i.e., discarded. The remaining primary transformation coefficients 328 are passed to the quantizer module 334. The primary transformation coefficients 328 are quantized according to the quantization parameter 392 associated with the CB to generate the primary transformation coefficients 332. The quantization parameter 392 can be different for the luminance CB relative to each chrominance CB. The primary transformation coefficients 332 are passed to the forward secondary transformation module 330 to generate the transformation coefficients represented by arrow 336 by performing a non-separable second transformation (NSST) operation or bypassing the second transformation. The forward primary transformation is usually separable, transforming a set of rows of each TB and then a set of columns. For a luminance TB with a width and height not exceeding 16 samples, the forward primary transformation module 326 uses a type-II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or bypasses the transformation in the horizontal and vertical directions, or uses a combination of a type-VII discrete sine transform (DST-7) and a type-VIII discrete cosine transform (DCT-8) in the horizontal or vertical direction. The use of the combination of DST-7 and DCT-8 is referred to as the "multiple transform selection set" (MTS) in the VVC standard.
[0110] The forward secondary transformation of module 330 is usually a non-separable transformation, which is only applied to the residual of the intra-predicted CU and can still be bypassed. The forward secondary transformation operates on 16 samples (arranged as the upper left 4×4 sub-block of the primary transformation coefficients 328) or 48 samples (arranged as three 4×4 sub-blocks in the upper left 8×8 coefficients of the primary transformation coefficients 328) to generate a set of secondary transformation coefficients. The number of the set of secondary transformation coefficients can be less than the number of the set of primary transformation coefficients from which it is derived. Since the secondary transformation is only applied to a set of coefficients adjacent to each other and including the DC coefficient, the secondary transformation is referred to as the "low-frequency non-separable secondary transformation" (LFNST). In addition, when the LFNST is applied, all remaining coefficients in the TB must be zero in both the primary transformation domain and the secondary transformation domain.
[0111] The quantization parameter 392 is constant for a given TB, and thus results in a uniform scaling of the residual coefficients generated in the primary transform domain of the TB. The quantization parameter 392 can be varied periodically by a "delta quantization parameter" signaled. For CUs contained within a given region (referred to as a "quantization group"), the delta quantization parameter (delta QP) is signaled once. If a CU is larger than the quantization group size, the delta QP is signaled once by one of the TBs of the CU. That is, for the first quantization group of a CU, the entropy encoder 338 signals the delta QP once, and for any subsequent quantization groups of the CU, the delta QP is not signaled. Non-uniform scaling is also possible by applying a "quantization matrix", whereby the scaling factors applied to the individual residual coefficients result from the combination of the quantization parameter 392 and the corresponding entries in the scaling matrix. The scaling matrix can have a size smaller than the size of the TB, and when applied to the TB, a nearest neighbor method is used to provide scaling values for the individual residual coefficients based on the scaling matrix having a size smaller than the TB size. The residual coefficients 336 are supplied to the entropy encoder 338 for encoding in the bitstream 115. Generally, according to a scan pattern, the residual coefficients of each TB of the scanned TU having at least one valid residual coefficient are scanned to produce an ordered list of values. The scan pattern generally scans the TBs as a sequence of 4×4 "sub-blocks", providing a regular scan operation at the granularity of 4×4 groups of residual coefficients, where the arrangement of the sub-blocks depends on the size of the TB. The scan within each sub-block and the progression from one sub-block to the next generally follow a backward diagonal scan pattern. Additionally, the quantization parameter 392 is encoded in the bitstream 115 using the delta QP syntax element, and the secondary transform index 388 is encoded in the bitstream 115 under the conditions described in the reference Figures 13 to 15 as described.
[0112] As described above, the video encoder 114 needs to access a frame representation corresponding to the encoded frame representation seen in the video decoder 134. Thus, the residual coefficients 336 are inverse-transformed by the inverse secondary transform module 344 (operating according to the secondary transform coefficients 388) to produce intermediate inverse transform coefficients represented by arrow 342. The intermediate inverse transform coefficients 346 are inverse-quantized by the dequantization module 340 according to the quantization parameter 392 to produce residual samples represented by arrow 346. The intermediate inverse transform coefficients 346 are passed to the inverse primary transform module 348 to produce residual samples of the TU represented by arrow 350. The type of inverse transform performed by the inverse secondary transform module 344 corresponds to the type of forward transform performed by the forward secondary transform module 330. The type of inverse transform performed by the inverse primary transform module 348 corresponds to the type of primary transform performed by the primary transform module 326. The summation module 352 adds the residual samples 350 and the PU 320 to produce the reconstructed samples of the CU (indicated by arrow 354).
[0113] The reconstructed sample 354 is passed to the reference sample cache 356 and the in-loop filter module 368. The reference sample cache 356, typically implemented using static RAM on the ASIC (thus avoiding expensive off-chip memory access), provides the minimum sample storage required to satisfy the dependencies for generating intra-PBs for subsequent CUs in a frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of a row of CTUs for use by the next row of CTUs as well as a column buffer with a range set by the height of the CTU. The reference sample cache 356 supplies reference samples (represented by arrow 358) to the reference sample filter 360. The sample filter 360 applies a smoothing operation to produce filtered reference samples (indicated by arrow 362). The filtered reference samples 362 are used by the intra prediction module 364 to produce an intra prediction block of samples represented by arrow 366. For each candidate intra prediction mode, the intra prediction module 364 produces a sample block, i.e., 366. The sample block 366 is generated by the module 364 using techniques such as DC, planar, or angular intra prediction.
[0114] The in-loop filter module 368 applies several filtering stages to the reconstructed sample 354. The filtering stages include a "deblocking filter" (DBF) that applies smoothing aligned with the CU boundaries to reduce artifacts due to discontinuities. Another filtering stage present in the in-loop filter module 368 is the "adaptive loop filter" (ALF) that applies a Wiener-based adaptive filter to further reduce distortion. Another available filtering stage in the in-loop filter module 368 is the "sample adaptive offset" (SAO) filter. The SAO filter works by first classifying the reconstructed samples into one or more classes and applying an offset at the sample level according to the assigned class.
[0115] Filtered samples represented by arrow 370 are output from the in-loop filter module 368. The filtered samples 370 are stored in the frame buffer 372. The frame buffer 372 typically has the capacity to store several (e.g., up to 16) pictures and is thus stored in the memory 206. Due to the large memory consumption required, the frame buffer 372 typically does not use on-chip memory for storage. As such, access to the frame buffer 372 is expensive in terms of memory bandwidth. The frame buffer 372 provides a reference frame (represented by arrow 374) to the motion estimation module 376 and the motion compensation module 380.
[0116] The motion estimation module 376 estimates a plurality of "motion vectors" (denoted as 378), each of which is a Cartesian spatial offset relative to the position of the current CB, thereby referencing a block in one of the reference frames in the frame buffer 372. A filtered block of reference samples (denoted as 382) is generated for each motion vector. The filtered reference samples 382 form a further candidate mode for potential selection by the mode selector 386. In addition, for a given CU, the PB 320 may be formed using one reference block ("uni-prediction"), or may be formed using two reference blocks ("bi-prediction"). For the selected motion vector, the motion compensation module 380 generates the PU 320 according to a filtering process that supports sub-pixel precision in the motion vector. In this way, the motion estimation module 376 (which operates on many candidate motion vectors) can perform a simplified filtering process compared to the motion compensation module 380 (which operates only on the selected candidate) to achieve reduced computational complexity. When the video encoder 114 selects inter-frame prediction for the CU, the motion vector 378 is encoded in the bitstream 115.
[0117] Although the reference to Versatile Video Coding (VVC) describes Figure 3 , but other video coding standards or implementations may also employ the processing stages of modules 310-390. Frame data 113 (and bitstream 115) may also be read from memory 206, hard drive 210, CD-ROM, Blue-ray diskTM, or other computer-readable storage media (or written to memory 206, hard drive 210, CD-ROM, Blue-ray disk, or other computer-readable storage media). In addition, frame data 113 (and bitstream 115) may be received from (or sent to) an external source (such as a server or radio frequency receiver connected to a communication network 220). The communication network 220 may provide limited bandwidth, requiring rate control to be used in the video encoder 114 to avoid saturating the network when the frame data 113 is difficult to compress. In addition, the bitstream 115 may be constructed from one or more stripes representing a spatial portion (CTU set) of the frame data 113, which are generated by one or more instances of the video encoder 114 and operated in a coordinated manner under the control of the processor 205. In the context of the present invention, a slice may also be referred to as a “contiguous portion” of the bitstream. A slice is contiguous within the bitstream and (eg, if parallel processing is being used) may be encoded or decoded as separate portions.
[0118] exist Figure 4 The video decoder 134 is shown in FIG. Figure 4 The video decoder 134 of FIG. 1 is an example of a Versatile Video Coding (VVC) video decoding pipeline, but other video codecs may also be used to perform the processing stages described herein. Figure 4As shown, the bitstream 133 is input to the video decoder 134. The bitstream 133 can be read from a memory 206, a hard disk drive 210, a CD-ROM, a Blu-ray disc, or other non-transitory computer-readable storage media. Alternatively, the bitstream 133 can be received from an external source (such as a server connected to a communication network 220 or a radio frequency receiver, etc.). The bitstream 133 contains encoded syntax elements representing the captured frame data to be decoded.
[0119] The bitstream 133 is input to the entropy decoder module 420. The entropy decoder module 420 extracts syntax elements from the bitstream 133 by decoding "bin" sequences and passes the values of the syntax elements to other modules in the video decoder 134. The entropy decoder module 420 uses variable length and fixed length decoding to decode the SPS, PPS, or slice headers, and uses an arithmetic decoding engine to decode the syntax elements of the slice data into a sequence of one or more bins. Each bin can use one or more "contexts", where the context describes the probability levels used to encode "one" and "zero" values for the bin. In cases where multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the available contexts to decode the bin. The process of decoding the bins forms a sequential feedback loop, so that each slice can be decoded in its entirety by a given instance of the entropy decoder 420. A single (or a few) high-performance entropy decoder 420 instances can decode all slices of a frame from the bitstream 115, and multiple low-performance entropy decoder 420 instances can decode slices of a frame from the bitstream 133 simultaneously.
[0120] The entropy decoder module 420 applies an arithmetic coding algorithm, such as "Context-Adaptive Binary Arithmetic Coding" (CABAC), to decode syntax elements from the bitstream 133. The decoded syntax elements are used to reconstruct parameters within the video decoder 134. The parameters include residual coefficients (represented by arrow 424), quantization parameter 474, quadratic transform index 470, and mode selection information such as intra prediction mode (represented by arrow 458). The mode selection information also includes information such as motion vectors, and the partitioning of each CTU into one or more CBs. The parameters are used to generally generate a PB in combination with sample data from previously decoded CBs.
[0121] The residual coefficients 424 are passed to the inverse quadratic transform module 436, where according to the reference Figures 16 to 18The described method applies a secondary transform or does not perform an operation (bypass). The inverse secondary transform module 436 generates the reconstructed transform coefficients 432 from the secondary transform domain coefficients, i.e., the primary transform domain coefficients. The reconstructed transform coefficients 432 are input to the dequantizer module 428. The dequantizer module 428 inverse quantizes (or "scales") the residual coefficients 432, i.e., in the primary transform coefficient domain, to create the reconstructed intermediate transform coefficients represented by arrow 440 according to the quantization parameter 474. If it is indicated in the bitstream 133 to use a non-uniform inverse quantization matrix, the video decoder 134 reads the quantization matrix as a sequence of scaling factors from the bitstream 133 and arranges the scaling factors into a matrix according to. The inverse scaling uses the quantization matrix in combination with the quantization parameter to create the reconstructed intermediate transform coefficients 440.
[0122] The reconstructed transform coefficients 440 are passed to the inverse primary transform module 444. Module 444 transforms the coefficients 440 back from the frequency domain to the spatial domain. The result of the operation of module 444 is a block of residual samples represented by arrow 448. The block of residual samples 448 is equal in size to the corresponding CB. The block of residual samples 448 is supplied to the summing module 450. At the summing module 450, the residual samples 448 are added to the decoded PB represented as 452 to produce a block of reconstructed samples represented by arrow 456. The reconstructed samples 456 are supplied to the reconstructed sample cache 460 and the in-loop filter module 488. The in-loop filter module 488 produces a reconstructed block of frame samples represented as 492. The frame samples 492 are written to the frame buffer 496.
[0123] The reconstructed sample cache 460 operates in a manner similar to the reconstructed sample cache 356 of the video encoder 114. The reconstructed sample cache 460 provides storage for the reconstructed samples required for intra prediction of subsequent CBs in the absence of the memory 206 (e.g., by using data 232 which is typically on-chip memory as a substitute). The reference samples represented by arrow 464 are obtained from the reconstructed sample cache 460 and are supplied to the reference sample filter 468 to produce the filtered reference samples represented by arrow 472. The filtered reference samples 472 are supplied to the intra prediction module 476. Module 476 generates a block of intra prediction samples represented by arrow 480 according to the intra prediction mode parameter 458 represented in the bitstream 133 and decoded by the entropy decoder 420. Modes such as DC, planar, or angular intra prediction are used to generate the block of samples 480.
[0124] When the prediction mode of the CB is indicated to use intra prediction in the bitstream 133, the intra prediction samples 480 form the decoded PB 452 via the multiplexer module 484. Intra prediction produces a predicted block (PB) of samples, i.e., a block in one color component derived using "neighboring samples" in the same color component. Neighboring samples are samples adjacent to the current block and have been reconstructed because they are earlier in the block decoding order. In the case of luma and chroma block juxtaposition, the luma and chroma blocks may use different intra prediction modes. However, the two chroma channels share the same intra prediction mode.
[0125] When the prediction mode of the CB is indicated to be intra prediction in the bitstream 133, the motion compensation module 434 selects and filters a block of samples 498 from the frame buffer 496 using the motion vector (decoded from the bitstream 133 by the entropy decoder 420) and the reference frame index to produce a block of inter prediction samples represented as 438. The block of samples 498 is obtained from a previously decoded frame stored in the frame buffer 496. For dual prediction, two blocks of samples are generated and mixed together to produce the samples of the decoded PB 452. The frame buffer 496 is filled with the filtered block data 492 from the in-loop filter module 488. Similar to the in-loop filter module 368 of the video encoder 114, the in-loop filter module 488 applies any of the DBF, ALF, and SAO filtering operations. Generally, the motion vector is applied to both the luma and chroma channels, but the filtering processes for subsample interpolation are different in the luma and chroma channels.
[0126] Figure 5 is a schematic block diagram showing a set 500 of available partitions or splits of a region in the tree structure of general video coding into one or more sub-regions. As referred to Figure 3 as described, the partitions shown in the set 500 are available for the block partitioner 310 of the encoder 114 to partition each CTU into one or more CUs or CBs according to the coding cost determined by Lagrangian optimization.
[0127] Although the set 500 only shows the partitioning of a square region into other possibly non-square sub-regions, it should be understood that the set 500 is showing the potential partitioning of a parent node in the coding tree into child nodes in the coding tree, and it is not required that the parent node corresponds to a square region. If the containing region is non-square, the sizes of the blocks resulting from the partition are scaled according to the aspect ratio of the containing block. Once a region is not further split, i.e., at the leaf node of the coding tree, the CU occupies that region.
[0128] The process of sub-dividing a region into sub-regions must terminate when the resulting sub-regions reach the minimum CU size (usually 4×4 luma samples). In addition to constraining the CUs to prohibit block regions smaller than a predetermined minimum size of, for example, 16 samples, the CUs are constrained to have a minimum width or height of four. Other minimum values are possible in terms of both width and height or in terms of either width or height. The sub-division process may also terminate before the deepest level of decomposition, resulting in CUs larger than the minimum CU size. It is possible that no splitting occurs, resulting in a single CU that occupies the entire CTU. A single CU that occupies the entire CTU is the largest available coding unit size. Due to the use of sub-sampled chroma formats (such as 4:2:0, etc.), the arrangement of the video encoder 114 and the video decoder 134 may terminate the splitting of regions in the chroma channel earlier than in the luma channel, including in the case of a common coding tree that defines the block structure of the luma and chroma channels. When separate coding trees are used for luma and chroma, the constraints on the available splitting operations ensure a minimum chroma CB region of 16 samples, even if such a CB is juxtaposed with a larger luma region (e.g., 64 luma samples).
[0129] In the absence of further sub-division, there are CUs at the leaf nodes of the coding tree. For example, leaf node 510 contains a single CU. At non-leaf nodes of the coding tree, there are splits to two or more other nodes, where each node can be a leaf node forming a single CU or a non-leaf node containing a further split to smaller regions. At each leaf node of the coding tree, there is a coding block for each color channel. A split that terminates at the same depth for both luma and chroma results in three juxtaposed CBs. A split that terminates at a deeper depth for luma than for chroma results in multiple luma CBs juxtaposed with the CBs of the chroma channel.
[0130] As Figure 5 shown, the quadtree split 512 divides the containing region into four equally sized regions. Compared to HEVC, Versatile Video Coding (VVC) achieves additional flexibility through additional splits, including horizontal binary split 514 and vertical binary split 516. Each of splits 514 and 516 divides the containing region into two equally sized regions. The division is along a horizontal boundary (514) or a vertical boundary (516) within the containing block.
[0131] In general video coding, further flexibility is achieved by adding ternary horizontal splits 518 and ternary vertical splits 520. The ternary splits 518 and 520 divide a block into three regions that form boundaries in the horizontal direction (518) or vertical direction (520) along 1 / 4 and 3 / 4 of the width or height of the containing region. A combination of quadtree, binary tree, and ternary tree is referred to as "QTBTTT". The root of the tree includes zero or more quadtree splits (the "QT" part of the tree). Once the QT part terminates, zero or more binary or ternary splits ("multi-tree" or the "MT" part of the tree) can occur, finally ending in a CB or CU at the leaf node of the tree. When the tree describes all color channels, the leaf node of the tree is a CU. When the tree describes the luminance channel or a chrominance channel, the leaf node of the tree is a CB.
[0132] Compared to HEVC which only supports quadtree and thus only supports square blocks, QTBTTT gets more possible CU sizes especially considering the possible recursive application of binary tree and / or ternary tree splits. When only quadtree splits are available, each increase in the coding tree depth corresponds to a reduction of the CU size to one quarter of the size of the parent region. In VVC, the availability of binary and ternary splits means that the coding tree depth no longer directly corresponds to the CU region. The possibility of abnormal (non-square) block sizes can be reduced by constraining the split options to eliminate splits that would result in a block width or height less than four samples or a split that would result in a size that is not a multiple of four samples. Generally, the constraint will apply when considering luminance samples. However, in the described arrangement, the constraint can be applied separately to blocks of the chrominance channel. The application of the split option constraint to the chrominance channel may result in different minimum block sizes for luminance vs chrominance (e.g., when the frame data is in 4:2:0 chrominance format or 4:2:2 chrominance format). Each split produces sub-regions with edge dimensions that are invariant, bisected, or quartered with respect to the containing region. Then, since the CTU size is a power of 2, the edge dimensions of all CUs are also powers of 2.
[0133] Figure 6 is a schematic flowchart of a data stream 600 showing the QTBTTT (or "coding tree") structure used in general video coding. The QTBTTT structure is used for each CTU to define the division of the CTU into one or more CUs. The QTBTTT structure of each CTU is determined by a block partitioner 310 in the video encoder 114 and is encoded into the bitstream 115 or decoded from the bitstream 133 by an entropy decoder 420 in the video decoder 134. According to Figure 5 the shown division, the data stream 600 further exhibits the characteristics of a permitted combination for the block partitioner 310 to divide the CTU into one or more CUs.
[0134] Starting from the top level of the hierarchical structure, i.e., at the CTU, zero or more quadtree partitions are first performed. Specifically, the quadtree (QT) split decision 610 is made by the block partitioner 310. The decision at 610 returns a "1" symbol, which indicates that it is decided to split the current node into four child nodes according to the quadtree split 512. As a result, four new nodes are generated, such as at 620, etc., and for each new node, it recurs back to the QT split decision 610. Each new node is considered in raster (or Z-scan) order. Alternatively, if the QT split decision 610 indicates no further split (returns a "0" symbol), the quadtree partition stops, and then the multi-tree (MT) split is considered.
[0135] First, the MT split decision 612 is made by the block partitioner 310. At 612, a decision indicating an MT split is made. A "0" symbol is returned at the decision 612, which indicates that there will be no further split of the node into child nodes. If there will be no further split of the node, the node is a leaf node of the coding tree and corresponds to a CU. The leaf node is output at 622. Alternatively, if the MT split 612 indicates a decision to perform an MT split (returns a "1" symbol), the block partitioner 310 enters the direction decision 614.
[0136] The direction decision 614 indicates the direction of the MT split as horizontal ("H" or "0") or vertical ("V" or "1"). If the decision 614 returns a "0" indicating the horizontal direction, the block partitioner 310 enters the decision 616. If the decision 614 returns a "1" indicating the vertical direction, the block partitioner 310 enters the decision 618.
[0137] In each of the decisions 616 and 618, the number of partitions of the MT split is indicated as two (binary split or "BT" node) or three (ternary split or "TT") during the BT / TT split. That is, when the direction indicated from 614 is horizontal, the block partitioner 310 makes the BT / TT split decision 616, and when the direction indicated from 614 is vertical, the block partitioner 310 makes the BT / TT split decision 618.
[0138] The BT / TT split decision 616 indicates whether the horizontal split is a binary split 514 indicated by returning a "0" or a ternary split 518 indicated by returning a "1". When the BT / TT split decision 616 indicates a binary split, at the step 625 of generating the HBT CTU node, the block partitioner 310 generates two nodes according to the binary horizontal split 514. When the BT / TT split 616 indicates a ternary split, at the step 626 of generating the HTT CTU node, the block partitioner 310 generates three nodes according to the ternary horizontal split 518.
[0139] The BT / TT split decision 618 indicates whether the vertical split is a binary split 516 indicated by returning "0" or a ternary split 520 indicated by returning "1". When the BT / TT split 618 indicates a binary split, at step 627 of generating the VBT CTU node, the block partitioner 310 generates two nodes according to the vertical binary split 516. When the BT / TT split 618 indicates a ternary split, at step 628 of generating the VTT CTU node, the block partitioner 310 generates three nodes according to the vertical ternary split 520. For each node obtained from steps 625 - 628, the data stream 600 is applied recursively back to the MT split decision 612 in the order from left to right or from top to bottom according to the direction 614. As a result, binary tree and ternary tree splits can be applied to generate CUs of various sizes.
[0140] Figure 7A and 7B Provide an example split 700 of the CTU 710 into multiple CUs or CBs. In Figure 7A an example CU 712 is shown. Figure 7A Show the spatial arrangement of the CUs in the CTU 710. The example split 700 is also shown as an encoding tree 720 in Figure 7B .
[0141] At each non - leaf node (e.g., nodes 714, 716, and 718) in the CTU 710 of Figure 7A , scan or traverse the included nodes (which can be further split or can be CUs) in "Z - order" to create a list of nodes represented as columns in the encoding tree 720. For quadtree splits, the Z - order scan results in the order from the upper left to the right and then from the lower left to the right. For horizontal and vertical splits, the Z - order scan (traversal) simplifies to a scan from the top to the bottom and a scan from the left to the right, respectively. Figure 7B The encoding tree 720 of
[0142] lists all the nodes and CUs according to the applied scan order. Each split generates a list of two, three, or four new nodes at the next level of the tree until reaching the leaf nodes (CUs).
[0142] In the case of decomposing an image into CTUs and further into CUs using the block partitioner 310 as described in reference Figure 3 and generating each residual block (324) using the CUs, the video encoder 114 performs forward transformation and quantization on the residual blocks. Subsequently, the resulting TB 336 is scanned to form an ordered list of residual coefficients as part of the operation of the entropy encoding module 338. An equivalent process is performed in the video decoder 134 to obtain the TB from the bitstream 133.
[0143] Figure 8A 、 8BFigs. 8C illustrate the levels of subdivision resulting from splits in the coding tree and the corresponding impact on partitioning the coding tree units into quantization groups. The residual of the TB signals the incremental QP (392) for each quantization group at most once. In HEVC, the definition of the quantization group corresponds to the coding tree depth, since this definition results in regions of fixed size. In VVC, the additional splits mean that the coding tree depth is no longer a suitable proxy for the CTU region. In VVC, the "level of subdivision" is defined, where each increment corresponds to half of the region included.
[0144] Figure 8A Fig. 800 shows a set of splits in the coding tree and the corresponding levels of subdivision. At the root node of the coding tree, the level of subdivision is initialized to zero. When the coding tree includes a quadtree split (e.g., 810), the level of subdivision is incremented by two for any CU included therein. When the coding tree includes a binary split (e.g., 812), the level of subdivision is incremented by one for any CU included therein. When the coding tree includes a ternary split (e.g., 814), the level of subdivision is incremented by two for the two outer CUs and incremented by one for the inner CU resulting from the ternary split. When traversing the coding tree of each CTU, as referred to Figure 6 as described, the level of subdivision of each resulting CU is determined according to set 800.
[0145] Figure 8B Fig. 840 shows an example set of CU nodes and shows the effect of the split. The example parent node 820 with a zero level of subdivision in set 840 corresponds to Figure 8B a CTU of size 64×64 in the example of The parent node 820 is ternary split to produce three child nodes 821, 822, and 823 with sizes 16×64, 32×64, and 16×64, respectively. The child nodes 821, 822, and 823 have levels of subdivision 2, 1, and 2, respectively.
[0146] In Figure 8BIn the example, the quantization group threshold is set to 1, corresponding to half of the 64×64 region, i.e., a region corresponding to 2048 samples. A flag tracks the start of a new QG. For any node with a subdivision level less than or equal to the quantization group threshold, the flag tracking the new QG is reset. When traversing the parent node 820 with a zero subdivision level, the flag is set. Although the central CU 822 of size 32×64 has a region of 2048 samples, the two sibling CUs 821 and 823 have a subdivision level of 2, i.e., a region of 1024, so the flag is not reset when traversing the central CU, and the quantization group does not start at the central CU. Instead, following the initial flag reset, the flag starts at the parent node as shown at 824. Effectively, the QP can only change at boundaries aligned with multiples of the quantization group region. The incremental QP is signaled together with the residual of the TB associated with the CB. If there are no valid coefficients, there is no opportunity to encode the incremental QP.
[0147] Figure 8C Example 860 is shown to illustrate the relationship between the subdivision level, QG, and the signaling of incremental QP by splitting the CTU 862 into multiple CUs and QGs. Vertical binary splitting divides the CTU 862 into two halves, with the left half 870 containing one CU, CU0, and the right half 872 containing several CUs (CU1 - CU4). In Figure 8C the example, the quantization group threshold is set to 2, such that the quantization group generally has a region equal to one - quarter of the CTU region. Since the subdivision level of the parent node (i.e., the root node of the coding tree) is zero, the QG flag is reset, and a new QG will start from the next coding CU (i.e., the CU at arrow 868). CU0 (870) has coding coefficients, so the incremental QP 864 is encoded together with the residual of CU0. The right half 872 undergoes horizontal binary splitting and is further split in the upper and lower parts of the right half 872, resulting in CUs CU1 - CU4. The subdivision levels of the coding tree nodes corresponding to the upper (877 including CU1 and CU2) and lower (878 including CU3 and CU4) parts of the right half 872 are 2. The subdivision level 2 is equal to the quantization group threshold 2, so new QGs start in each part, labeled 874 and 876 respectively. CU1 has no coding coefficients (no residual), and CU2 is a "skip" CU, which also has no coding coefficients. Therefore, for the upper part, no incremental QP is encoded. CU3 is a skip CU, and CU4 has coding residuals, so for the QG including CU3 and CU4, the incremental QP 866 is encoded using the residual of CU4.
[0148] Figure 9A and 9BShows a 4×4 transform block scan pattern and associated primary and secondary transform coefficients. The operation of the secondary transform module 330 on the primary residual coefficients is described from the perspective of the video encoder 114. The 4×4 TB 900 is scanned according to the backward diagonal scan pattern 910. The scan pattern 910 proceeds from the "last significant coefficient" position towards the DC (top left) coefficient position. All coefficient positions that are not scanned, e.g., residual coefficients that are located after the last significant coefficient position when considering scanning in the forward direction, are implicitly non-significant. When secondary transform is used, all remaining coefficients are non-significant. That is, all secondary-domain residual coefficients for which secondary transform is not performed are non-significant, and all primary-domain residual coefficients that are not filled by the application of the secondary transform need to be non-significant. Additionally, after the forward secondary transform is applied by module 330, there may be fewer secondary transform coefficients than the number of primary transform coefficients processed by the secondary transform module 330. For example, Figure 9B shows a set of blocks 920. In Figure 9B it, sixteen (16) primary coefficients are arranged as a 4×4 sub-block, i.e., 924 of the 4×4 TB 920. In Figure 9B the example, the primary residual coefficients can be subjected to secondary transform to produce a secondary transform block 926. The secondary transform block 926 contains eight secondary transform coefficients 928. The eight secondary transform coefficients 928 are stored in the TB according to the scan pattern 910, packed forward from the DC coefficient position. The remaining coefficient positions of the 4×4 sub-block (shown as region 930) contain the quantized residual coefficients from the primary transform and need to be non-significant for the secondary transform to be applied. Thus, the last significant coefficient position indicating a coefficient that is one of the first eight scan positions of the 4×4 TB designated as TB 920 indicates (i) the application of the secondary transform, or (ii) that after quantization, the output of the primary transform does not have significant coefficients beyond the eighth scan position of TB 920.
[0149] When the TB can be subjected to secondary transform, the secondary transform index (i.e., 388) is encoded to indicate the possible application of the secondary transform. The secondary transform index can also indicate which kernel will be applied as the secondary transform at module 330 in the case where multiple transform kernels are available. Accordingly, when the last significant coefficient position is at any one of the scan positions reserved for holding the secondary transform coefficients (e.g., 928), the video decoder 134 decodes the secondary transform index 470.
[0150] Although a quadratic transform kernel that maps 16 primary coefficients to eight quadratic coefficients has been described, different kernels are possible, including kernels that map to a different number of quadratic transform coefficients. The number of quadratic transform coefficients can be the same as the number of primary transform coefficients, e.g., 16. For a TB with a width of 4 and a height greater than 4, the behavior described for the 4×4 TB case applies to the top sub-block of the TB. When the quadratic transform is applied, the other sub-blocks of the TB have zero-valued residual coefficients. For a TB with a width greater than 4 and a height equal to 4, the behavior described for the 4×4 TB case applies to the leftmost sub-block of the TB, and the other sub-blocks of the TB have zero-valued residual coefficients, allowing the last significant coefficient position to be used to determine whether the quadratic transform index needs to be decoded.
[0151] Figure 9C and 9D illustrates an 8×8 transform block scan pattern and example associated primary and quadratic transform coefficients. Figure 9C Illustrates a backward diagonal scan pattern 950 based on 4×4 sub-blocks for an 8×8 TB 940. The 8×8 TB 940 is scanned in the backward diagonal scan pattern 950 based on 4×4 sub-blocks. Figure 9D Illustrates a set 960 that shows the operation effect of the quadratic transform. The scan 950 returns from the last significant coefficient position to the DC (top left) coefficient position. When the remaining 16 primary coefficients (shown as 964) are zero-valued, it is possible to apply the forward quadratic transform kernel to 48 primary coefficients (the region 962 shown as 940). Applying the quadratic transform to the region 962 results in 16 quadratic transform coefficients shown as 966. The other coefficient positions of the TB are zero-valued, labeled 968. If the last significant position of the 8×8 TB 940 indicates that the quadratic transform coefficients are within 966, the quadratic transform index 388 is encoded to indicate to module 330 to apply a specific transform kernel (or bypass the kernel). The video decoder 134 uses the last significant position of the TB to determine whether to decode the quadratic transform index, i.e., index 470. For transform blocks with a width or height exceeding eight samples, Figure 9C and 9D the method of
[0152] such as Figures 9A to 9DAs described, two sizes of quadratic transformation kernels are available. One size of quadratic transformation kernel is for transformation blocks with a width or height of 4, and the other size of quadratic transformation is for transformation blocks with a width and height greater than 4. Within each size of kernel, multiple sets (e.g., four) of quadratic transformation kernels are available. One set is selected based on the block-based intra prediction mode, and this one set can be different between the luminance block and the chrominance block. Within the selected set, one or two kernels are available. Independent of the luminance block and the chrominance block in the coding units belonging to the common tree of the coding tree unit, signaling the use of one kernel within the selected set or bypassing the quadratic transformation via the quadratic transformation index. In other words, the index for the luminance channel and the index for the chrominance channel are independent of each other.
[0153] Figure 10 Fig. 1000 shows a set of transformation blocks available in the Versatile Video Coding (VVC) standard. Figure 10 Also shown is applying a quadratic transformation to a subset of the residual coefficients of the transformation blocks from set 1000. Figure 10 Shows TBs with widths and heights in the range from 4 to 32. However, TBs with a width and / or height of 64 are possible but not shown for ease of reference.
[0154] A 16-point quadratic transformation 1052 (shown in darker shading) is applied to the 4×4 coefficient set. The 16-point quadratic transformation 1052 is applied to TBs with a width or height of 4, e.g., 4×4 TB 1010, 8×4 TB 1012, 16×4 TB 1014, 32×4 TB 1016, 4×8 TB 1020, 4×16 TB 1030, and 4×32 TB 1040. If a 64-point primary transformation is available, the 16-point quadratic transformation 1052 is applied to TBs of size 4×64 and 64×4 ( Figure 10 not shown in Fig. 1). For TBs with a width or height of four but having more than 16 primary coefficients, the 16-point quadratic transformation is only applied to the upper left 4×4 sub-block of the TB, and the other sub-blocks need to have zero-valued coefficients to apply the quadratic transformation. Generally, applying the 16-point quadratic transformation results in 16 quadratic transformation coefficients, which are packed into the TB to be coded in the sub-block that obtained the original 16 primary transformation coefficients. For example, as referenced Figure 9B described, the quadratic transformation kernel can be such that fewer quadratic transformation coefficients are created than the number of primary transformation coefficients to which the quadratic transformation is applied.
[0155] For transformation sizes with a width and height greater than four, as Figure 10As shown, a 48-point secondary transform 1050 (shown in lighter shading) can be used for application to three 4×4 sub-blocks of residual coefficients in the upper-left 8×8 region of a transform block. In each case, in the regions shown with light shading and dashed outlines, the 48-point secondary transform 1050 is applied to 8×8 transform block 1022, 16×8 transform block 1024, 32×8 transform block 1026, 8×16 transform block 1032, 16×16 transform block 1034, 32×16 transform block 1036, 8×32 transform block 1042, 16×32 transform block 1044, and 32×32 transform block 1046. If a 64-point primary transform is available, the 48-point secondary transform 1050 is also applicable to TBs (not shown) of sizes 8×64, 16×64, 32×64, 64×64, 64×32, 64×16, and 64×8. Application of the 48-point secondary transform kernel typically results in fewer than 48 secondary transform coefficients being produced. For example, 8 or 16 secondary transform coefficients can be produced. The secondary transform coefficients are stored in the transform block in the upper-left region. For example, Figure 9D shows eight secondary transform coefficients. The primary transform coefficients that are not subjected to the secondary transform ("only primary coefficients") (e.g., coefficient 1066 of TB 1034 (similar to Figure 9D 's 964)) need to be zero-valued to apply the secondary transform. After applying the 48-point secondary transform 1050 in the forward direction, the region that can contain valid coefficients is reduced from 48 coefficients to 16 coefficients, thus further reducing the number of coefficient positions that can contain valid coefficients. For example, 968 will only contain non-valid coefficients. For the inverse secondary transform, for example, the decoded valid coefficients that exist only in 966 of the TB are transformed to produce any coefficients that may be valid in the region (e.g., 962), and then the primary inverse transform is applied to these coefficients. When the secondary transform reduces one or more sub-blocks to a set of 16 secondary transform coefficients, only the upper-left 4×4 sub-block can contain valid coefficients. The last valid coefficient position at any coefficient position where the secondary transform coefficients can be stored indicates whether the secondary transform or only the primary transform is applied. However, after quantization, the resulting valid coefficients are in the same region as if the secondary transform kernel had been applied.
[0156] When the last valid coefficient position indicates a secondary transform coefficient position in a TB (e.g., 922 or 962), a signalized secondary transform index is needed to distinguish between applying the secondary transform kernel or bypassing the secondary transform. Although the application of the secondary transform has been described from the perspective of video encoder 114 to Figure 10TBs of various sizes in [the relevant context], but corresponding inverse processing is performed in the video decoder 134. The video decoder 134 first decodes the last valid coefficient position. If the decoded last valid coefficient position indicates a potential application of the secondary transform, i.e., the position is within 928 or 966 of the secondary transform kernel that generates 8 or 16 secondary transform coefficients respectively, then the secondary transform index is decoded to determine whether to apply or bypass the inverse secondary transform.
[0157] Figure 11 Illustrates the syntax structure 1100 of the bitstream 1101 having multiple strips. Each strip in the strip contains a plurality of coding units. The bitstream 1101 can be generated by the video encoder 114 as, for example, the bitstream 115, or can be parsed by the video decoder 134 as, for example, the bitstream 133. The bitstream 1101 is segmented into multiple parts, such as Network Abstraction Layer (NAL) units, where the description is achieved by setting a NAL unit header (such as 1108, etc.) before each NAL unit. The Sequence Parameter Set (SPS) 1110 defines sequence-level parameters, such as the profile (toolset) for encoding and decoding the bitstream, chroma format, sample bit depth, and frame resolution, etc. The set 1110 also includes parameters that constrain the application of different types of splits in the coding tree of each CTU. For example, using a log2 base for block size constraints and expressing the parameters relative to other parameters (such as the minimum CTU size, etc.), the encoding of the parameters that constrain the split type can be optimized for a more compact representation. Several parameters encoded in the SPS 1110 are as follows:
[0158] · log2_ctu_size_minus5: Specifies the CTU size, where the encoded values 0, 1, and 2 specify CTU sizes of 32×32, 64×64, and 128×128 respectively.
[0159] · partition_constraints_override_enabled_flag: Enables strip-level override of several parameters, collectively referred to as the partition constraint parameters 1130.
[0160] · log2_min_luma_coding_block_size_minus2: Specifies the minimum coding block size (in luma samples), where the values 0, 1, 2,... specify minimum luma CB sizes of 4×4, 8×8, 16×16,.... The maximum coding value is constrained by the specified CTU size, i.e., such that log2_min_luma_coding_block_size_minus2 ≤ log2_ctu_size_minus5 + 3. The available chroma block sizes correspond to the available luma block sizes, scaled according to the chroma channel subsampling of the chroma format in use.
[0161] ·sps_max_mtt_hierarchy_depth_inter_slice: Specifies the maximum hierarchical depth of coding units in the coding tree of multi-tree splitting (i.e., binary and ternary splitting) relative to the quadtree nodes in the coding tree of an inter (P or B) slice (i.e., once the quadtree splitting stops in the coding tree), and is one of the parameters 1130.
[0162] ·sps_max_mtt_hierarchy_depth_intra_slice_luma: Specifies the maximum hierarchical depth of coding units in the coding tree of multi-tree splitting (i.e., binary and ternary) relative to the quadtree nodes in the coding tree of an intra (I) slice (i.e., once the quadtree splitting stops in the coding tree), and is one of the parameters 1130.
[0163] ·partition_constraints_override_flag: When partition_constraints_override_enabled_flag in the SPS is equal to 1, this parameter is signaled in the slice header, and this parameter indicates that the partition constraints signaled in the SPS will be overridden for the corresponding slice.
[0164] The picture parameter set (PPS) 1112 defines a set of parameters applicable to zero or more frames. The parameters included in the PPS 1112 include parameters for splitting a frame into one or more "regions" and / or "blocks". The parameters of the PPS1112 can also include a list of CU chroma QP offsets, one of which can be applied at the CU level to derive the quantization parameter for use by the chroma blocks from the quantization parameter of the collocated luma CB.
[0165] The sequence of slices forming a picture is called an access unit (AU), such as AU 0 1114, etc. AU 0 1114 includes three slices, such as slices 0 to 2, etc. Slice 1 is labeled 1116. Like other slices, slice 1 (1116) includes a slice header 1118 and slice data 1120.
[0166] The slice header includes parameters grouped as 1134. Group 1134 includes:
[0167] · slice_max_mtt_hierarchy_depth_luma: Signaled in slice header 1118 when partition_constraints_override_flag in the slice header is equal to 1 and overrides the value derived from the SPS. For I slices, instead of using sps_max_mtt_hierarchy_depth_intra_slice_luma to set MaxMttDepth at 1134, slice_max_mtt_hierarchy_depth_luma is used. For P or B slices, instead of using sps_max_mtt_hierarchy_depth_inter_slice, slice_max_mtt_hierarchy_depth_luma is used.
[0168] The variable MinQtLog2SizeIntraY (not shown) is derived from the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma decoded from the SPS 1110, which specifies the minimum coded block size obtained by zero or more quadtree splits of an I slice (i.e., no further MTT splits occur in the coding tree). The variable MinQtLog2SizeInterY (not shown) is derived from the syntax element sps_log2_diff_min_qt_min_cb_inter_slice decoded from the SPS 1110. The variable MinQtLog2SizeInterY specifies the minimum coded block size obtained by zero or more quadtree splits of P and B slices (i.e., no further MTT splits occur in the coding tree). Since the CUs obtained by quadtree splits are square, the variables MinQtLog2SizeIntraY and MinQtLog2SizeInterY each specify both the width and height (as log2 of the CU width / height).
[0169] The parameter cu_qp_delta_subdiv may optionally be signaled in the slice header 1118, and indicates the maximum subdivision level for signaling the incremental QP in the coding tree for the luma branch in a common tree or in a separate tree slice. For an I slice, the range of cu_qp_delta_subdiv is from 0 to 2*(log2_ctu_size_minus5+5-MinQtLog2SizeIntraY+MaxMttDepthY 1134). For a P or B slice, the range of cu_qp_delta_subdiv is from 0 to 2*(log2_ctu_size_minus5+5-MinQtLog2SizeInterY+MaxMttDepthY 1134). Since the range of cu_qp_delta_subdiv depends on the value MaxMttDepthY 1134 derived from the partitioning constraints obtained from the SPS 1110 or the slice header 1118 there is no parsing problem.
[0170] The parameter cu_chroma_qp_offset_subdiv may optionally be signaled in the slice header 1118, and indicates the maximum subdivision level for signaling the chroma CU QP offset in the chroma branch in a common tree or in a separate tree slice. The range constraints for cu_chroma_qp_offset_subdiv for I or P / B slices are the same as the corresponding range constraints for cu_qp_delta_subdiv.
[0171] For a CTU in slice 1120, a subdivision level 1136 is derived, which specifies cu_qp_delta_subdiv for the luma CB and cu_chroma_qp_offset_subdiv for the chroma CB. As described in Ref. Figure 8A -C the subdivision level is used to determine at which points in the CTU the incremental QP syntax element is coded. For the chroma CB the method of Figure 8A -C is also used to signal chroma CU level offset enablement (and index if enabled).
[0172] Figure 12Shows the syntax structure 1200 of the strip data 1120 of the bitstream 1101 (e.g., 115 or 133), which has a common tree for encoding the luminance and chrominance coding blocks of tree units (such as CTU 1210, etc.). CTU 1210 includes one or more CUs, an example being shown as CU 1214. CU 1214 includes a signaled prediction mode 1216a, followed by a transform tree 1216b. When the size of CU 1214 does not exceed the maximum transform size (32×32 or 64×64), the transform tree 1216b includes one transform unit, shown as TU 1218.
[0173] If the prediction mode 1216a indicates the use of intra prediction for CU 1214, the intra prediction modes for luminance and chrominance are specified. For the luminance CB of CU 1214, the primary transform type is also signaled as (i) horizontal and vertical DCT-2, (ii) horizontal and vertical transform skip, or (iii) a combination of horizontal and vertical DST-7 and DCT-8. If the signaled luminance transform type is horizontal and vertical DCT-2 (option (i)), then under the conditions described in reference Figure 9A -D, an additional luminance secondary transform type 1220, also known as the "low-frequency non-separable transform" (LFNST) index, is signaled in the bitstream. The chrominance secondary transform type 1221 is also signaled. The chrominance secondary transform type 1221 is signaled independently of whether the luminance primary transform type is DCT-2.
[0174] The use of the common coding tree results in TU 1218 including TBs for the respective color channels shown as luminance TB Y 1222, first chrominance TB Cb 1224, and second chrominance TB Cr 1226. It is available to send a single chrominance TB to specify the coding mode for the chrominance residuals of both the Cb and Cr channels, called the "joint CbCr" coding mode. When the joint CbCr coding mode is enabled, a single chrominance TB is encoded.
[0175] Regardless of the color channel, each TB includes a last position 1228. The last position 1228 indicates the last valid residual coefficient position in the TB when considering the coefficients in the diagonal scan mode, which is used to serialize the coefficient array of the TB in the forward direction (i.e., from the DC coefficient forward). If the last position 1228 of the TB indicates that only the coefficients in the secondary transform domain (i.e., all remaining coefficients that will only undergo the primary transform) are valid, a secondary transform index is signaled to specify whether the secondary transform is applied.
[0176] If a second transform is to be applied and if more than one second transform kernel is available, the second transform index indicates which kernel to select. Typically, in a "candidate set", one kernel is available, or two kernels are available. The candidate set is determined from the intra prediction mode of the block. Typically, there are four candidate sets, but there can be fewer candidate sets. As described above, the second transform is used for luminance and chrominance, and thus the selected kernel depends on the intra prediction mode for the luminance and chrominance channels respectively. The kernel can also depend on the block size of the corresponding luminance and chrominance TBs. The kernel selected for chrominance also depends on the chrominance subsampling rate of the bitstream. If only one kernel is available, signaling is restricted to applying or not applying the second transform (index range 0 to 1). If two kernels are available, the index value is 0 (not apply), 1 (apply the first kernel), or 2 (apply the second kernel). For chrominance, the same second transform kernel is applied to each chrominance channel, so the residuals of the Cb block 1224 and the Cr block 1226 only need to include valid coefficients at positions that undergo the second transform, as described with reference Figure 9A -D. If joint CbCr coding is used, the requirement to include only valid coefficients at positions that undergo the second transform applies only to a single coded chrominance TB, since the resulting Cb and Cr residuals contain valid coefficients only at positions corresponding to the valid coefficients in the jointly coded TB. If the (one or more) applicable color channels for a given second index are described by a single TB (a single last position, e.g., 1228), i.e., when using joint CbCr coding, luminance always requires only one TB and chrominance requires one TB, then the second transform index can be signaled immediately after the coding of the last position instead of after the TU, i.e., as index 1230 instead of 1220 (or 1221). Signaling the second transform earlier in the bitstream allows the video decoder 134 to start applying the second transform when each residual coefficient in the residual coefficients 1232 is decoded, thus reducing the delay in the system 100.
[0177] In the arrangement of the video encoder 114 and the video decoder 134, when joint CbCr coding is not used, a separate second transform index is signaled for each chrominance TB (i.e., 1224 and 1226), resulting in independent control of the second transform for each color channel. If each TB is controlled independently, the second transform index for each TB can be signaled immediately after the last position of the corresponding TB for luminance and chrominance (regardless of whether the joint CbCr mode is applied).
[0178] Figure 13A method 1300 for encoding frame data 113 into a bitstream 115 is shown. The bitstream 115 includes one or more slices as a sequence of coding tree units. The method 1300 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, the method 1300 may be performed by a video encoder 114 under the execution of a processor 205. Due to the workload of encoding frames, the steps of the method 1300 may be performed in different processors to, for example, share the workload using contemporary multi-core processors such that different slices are encoded by different processors. Further, when encoding respective parts (slices) of the bitstream 115, the partition constraints and quantization group definitions may vary from one slice to another, which is considered beneficial for rate control purposes. For additional flexibility in encoding the residuals of respective coding units, not only may the quantization group subdivision level vary from one slice to another, but also the application of the secondary transform is independently controllable for luminance and chrominance. Thus, the method 1300 may be stored on a computer-readable storage medium and / or in a memory 206.
[0179] The method 1300 begins with an SPS / PPS encoding step 1310. At step 1310, the video encoder 114 encodes the SPS 1110 and PPS 1112 as a sequence of fixed and variable length coding parameters into the bitstream 115. The partition_constraints_override_enabled_flag is encoded as part of the SPS 1110, indicating that the partition constraints can be overridden in the slice header (1118) of a corresponding slice (such as 1116, etc.). The default partition constraints are also encoded by the video encoder 114 as part of the SPS 1110.
[0180] The method 1300 continues from step 1310 to a step 1320 of slicing a frame. In the execution of step 1320, the processor 205 slices the frame data 113 into one or more slices or consecutive parts. In cases where parallelism is desired, separate instances of the video encoder 114 independently encode respective slices to some extent. A single video encoder 114 may process the respective slices sequentially or may implement some intermediate degree of parallelism. Generally, slicing the frame (into consecutive parts) is aligned with the boundaries of slicing the frame into regions referred to as "sub-pictures" or blocks, etc.
[0181] The method 1300 continues from step 1320 to a step 1330 of encoding a slice header. At step 1330, an entropy encoder 338 encodes the slice header 1118 into the bitstream 115. The following provides an example implementation of step 1330. Figure 14 Provide an example implementation of step 1330.
[0182] Method 1300 proceeds from step 1330 to step 1340 of splitting the strip into CTUs. In the execution of step 1340, video encoder 114 splits strip 1116 into a sequence of CTUs. The strip boundaries are aligned with the CTU boundaries, and the CTUs in the strip are sorted according to the CTU scan order (usually raster scan order). Splitting the strip into CTUs determines which part of the frame data 113 will be processed by video encoder 113 when encoding the current strip.
[0183] Method 1300 proceeds from step 1340 to step 1350 of determining the coding tree. At step 1350, video encoder 114 determines the coding tree of the currently selected CTU in the strip. Method 1300 starts from the first CTU in strip 1116 on the first call of step 1350 and advances to the subsequent CTUs in strip 1116 on subsequent calls. When determining the coding tree of the CTU, various combinations of quadtree, binary, and ternary splits generated and tested by block partitioner 310 are involved.
[0184] Method 1300 proceeds from step 1350 to step 1360 of determining the coding unit. At step 1360, video encoder 114 performs to determine the "best" coding of the CUs obtained from the various coding trees under evaluation using known methods. Determining the best coding involves determining the prediction mode (e.g., intra prediction with a specific mode or inter prediction with a motion vector), transform selection (primary transform type and optional secondary transform type). If the primary transform type of the luminance TB is determined to be DCT-2 or any quantized primary transform coefficients that do not undergo forward secondary transform are valid, the secondary transform index of the luminance TB can indicate the application of the secondary transform. Otherwise, the secondary transform index of the luminance indicates bypassing the secondary transform. For the luminance channel, the primary transform type is determined to be one of the MTS options of the chrominance channel, DCT-2, or transform skip, and DCT-2 is the available transform type. Refer to Figure 19A and 19B the determination of the secondary transform type is further described. Determining the coding may also include determining the quantization parameter that can change the QP, i.e., the quantization parameter at the quantization group boundary. When determining each coding unit, the best coding tree is also determined in a joint manner. When encoding the coding unit using intra prediction, the luminance intra prediction mode and chrominance intra prediction are determined.
[0185] When there are no "AC" (coefficients in positions other than the upper left position of the transform block) residual coefficients in the primary domain residual obtained by the application of the DCT-2 primary transform, step 1360 of determining the coding unit may prohibit the application of the test secondary transform. If the test secondary transform application is performed for a transform block that includes only DC coefficients (the last position indicates that only the upper left coefficient of the transform block is valid), an increase in coding efficiency is seen. The prohibition of the test secondary transform when only DC main coefficients are present spans the blocks to which the secondary transform index applies, i.e., the Y, Cb, and Cr of the common tree (only the Y channel when the Cb and Cr blocks are two samples wide or high) when encoding a single index. Although the residual with only DC coefficients has a lower coding cost compared to a residual with at least one AC coefficient, applying the secondary transform to a residual with only valid DC coefficients also results in a further reduction in the magnitude of the finally encoded DC coefficients. Even after further quantization and / or rounding operations before encoding, the magnitude of the other (AC) coefficients after the secondary transform is not sufficient to obtain (one or more) valid coded residual coefficients in the bitstream. In the common or separate tree coding tree, assuming there is at least one valid main coefficient, even if there is only (one or more) DC coefficients of the corresponding transform block within the application range of the secondary transform index, the video encoder 114 tests the selection of non-zero secondary transform index values (i.e., the application of the secondary transform).
[0186] Method 1300 continues from step 1360 to step 1370 of encoding the coding unit. At step 1370, the video encoder 114 encodes the determined coding unit of step 1360 in the bitstream 115. Refer to Figure 15 An example of how to encode a coding unit is described in more detail.
[0187] Method 1300 continues from step 1370 to step 1380 of the last coding unit test. At step 1380, the processor 205 tests whether the current coding unit is the last coding unit in the CTU. If not (step 1380 is "no"), the control in the processor 205 advances to step 1360 of determining the coding unit. Otherwise, if the current coding unit is the last coding unit (step 1380 is "yes"), the control in the processor 205 advances to step 1390 of the last CTU test.
[0188] At step 1390 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU in the slice 1116. If it is not the last CTU in the slice 1116, the control in the processor 205 returns to step 1350 of determining the coding tree. Otherwise, if the current CTU is the last one (step 1390 is "yes"), the control in the processor advances to step 13100 of the last slice test.
[0189] At step 13100 of the last stripe test, the processor 205 tests whether the current stripe being encoded is the last stripe in the frame. If it is not the last stripe (step 13100 is "No"), then control in the processor 205 advances to step 1330 of encoding the stripe header. Otherwise, if the current stripe is the last one and all stripes (continuous portions) have been encoded (step 13100 is "Yes"), then method 1300 terminates.
[0190] Figure 14 Method 1400 for encoding stripe header 1118 into bitstream 115 as implemented at step 1330 is shown. Method 1400 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1400 may be performed by video encoder 114 under the execution of processor 205. Thus, method 1400 may be stored on a computer-readable storage medium and / or in memory 206.
[0191] Method 1400 begins at step 1410 of the partition constraint override enable test. At step 1410, the processor 205 tests whether the partition constraint override enable flag encoded in SPS 1110 indicates that the partition constraint can be overridden at the stripe level. If the partition constraint can be overridden at the stripe level (step 1410 is "Yes"), then control in the processor 205 advances to step 1420 of determining the partition constraint. Otherwise, if the partition constraint cannot be overridden at the stripe level (step 1410 is "No"), then control in the processor 205 advances to step 1480 of encoding other parameters.
[0192] At step 1420 of determining the partition constraint, the processor 205 determines the partition constraint (e.g., maximum MTT split depth) suitable for the current stripe 1116. In one example, the frame data 310 contains a projection of a 360-degree view of a scene mapped into a 2D box and segmented into a number of sub-pictures. Depending on the selected viewport, some stripes may require higher fidelity and other stripes may require lower fidelity. The partition constraint for the stripe may be set based on the fidelity requirements of the portion of the frame data 310 encoded by the given stripe (e.g., in accordance with step 1340). In cases where lower fidelity is considered acceptable, a shallower encoding tree with larger CUs is acceptable, and thus the maximum MTT depth may be set to a lower value. Accordingly, at least in the range resulting from the determined maximum MTT depth 1134, the subdivision level 1136 signaled with flag cu_qp_delta_subdiv is determined. The corresponding chroma subdivision level is also determined and signaled.
[0193] Method 1400 proceeds from step 1420 to step 1430 of encoding a partition constraint override flag. At step 1430, entropy encoder 338 encodes the flag in bitstream 115, which indicates whether the partition constraint signaled in SPS 1110 is to be overridden for slice 1116. If a strip-specific partition constraint is derived at step 1420, the flag value will indicate the use of the partition constraint override feature. If the constraint determined at step 1420 matches the constraint already encoded in SPS 1110, there is no need to override the partition constraint as there is no change to be signaled, and the flag value is encoded accordingly.
[0194] Method 1400 proceeds from step 1430 to step 1440 of a partition constraint override test. At step 1440, processor 205 tests the flag value encoded at step 1430. If the flag indicates that the partition constraint is to be overridden (step 1440 is "yes"), control in processor 205 advances to step 1450 of encoding strip partition constraints. Otherwise, if the partition constraint is not overridden (step 1440 is "no"), control in processor 205 advances to step 1480 of encoding other parameters.
[0195] Method 1400 proceeds from step 1440 to step 1450 of encoding strip partition constraints. In the execution of step 1450, entropy encoder 338 encodes the determined partition constraint of the strip in bitstream 115. The partition constraint of the strip includes "slice_max_mtt_hierarchy_depth_luma", from which MaxMttDepthY 1134 is derived.
[0196] Method 1400 proceeds from step 1450 to step 1460 of encoding the QP subdivision level. At step 1460, entropy encoder 338 encodes the subdivision level of the luma CB using the "cu_qp_delta_subdiv" syntax element, as referenced Figure 11 as described.
[0197] Method 1400 proceeds from step 1460 to step 1470 of encoding the chroma QP subdivision level. At step 1470, entropy encoder 338 encodes the signaled subdivision level for the CU chroma QP offset using the "cu_chroma_qp_offset_subdiv" syntax element, as referenced Figure 11 as described.
[0198] Steps 1460 and 1470 operate to encode the overall QP subdivision levels for the stripes (successive portions) of a frame. The overall subdivision levels include both the subdivision level for the luminance coding units of the stripe and the subdivision level for the chrominance coding units of the stripe. For example, since separate coding trees are used for luminance and chrominance in an I stripe, the chrominance and luminance subdivision levels can be different.
[0199] Method 1400 continues from step 1470 to step 1480 of encoding other parameters. At step 1480, entropy encoder 338 encodes the other parameters in the stripe header 1118, such as parameters required to control specific tools such as deblocking, adaptive loop filtering, optionally selecting a scaling list from a previously signaled scaling list (for non-uniformly applying quantization parameters to transform blocks), etc. Method 1400 terminates when step 1480 is executed.
[0200] Figure 15 Method 1500 for encoding a coding unit in a bitstream 115 is shown, which corresponds to Figure 13 step 1370. Method 1500 can be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1500 can be performed by video encoder 114 under the execution of processor 205. Thus, method 1500 can be stored on a computer-readable storage medium and / or in memory 206.
[0201] Method 1500 begins at step 1510 of encoding a prediction mode. At step 1510, entropy encoder 338 encodes the prediction mode for the coding unit determined at step 1360 in the bitstream 115. The "pred_mode" syntax element is encoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit. If intra prediction is used for the coding unit, the luminance intra prediction mode is encoded and the chrominance intra prediction mode is encoded. If inter prediction is used for the coding unit, a "merge index" can be encoded to select a motion vector from adjacent coding units for use by the coding unit, and a motion vector delta can be encoded to introduce an offset into the motion vector derived from spatially adjacent blocks. The primary transform type is encoded to select between using DCT-2 horizontally and vertically, transform skip horizontally and vertically, or a combination of DCT-8 and DST-7 horizontally and vertically for the luminance TB of the coding unit.
[0202] Method 1500 proceeds from step 1510 to step 1520 of the coded residual test. At step 1520, the processor 205 determines whether the residuals need to be coded for the coding unit. If there are any valid residual coefficients to be coded for the coding unit (step 1520 is "yes"), then the control in the processor 205 advances to the new QG test step 1530. Otherwise, if there are no valid residual coefficients for coding (step 1520 is "no"), then method 1500 terminates because all the information required to decode the coding unit is present in the bitstream 115.
[0203] At the new QG test step 1530, the processor 205 determines whether the coding unit corresponds to a new quantization group. If the coding unit corresponds to a new quantization group (step 1530 is "yes"), then the control in the processor 205 proceeds to the step 1540 of coding the incremental QP. Otherwise, if the coding unit is not related to a new quantization group (step 1530 is "no"), then the control in the processor 205 advances to the step 1550 of performing the main transform. When coding each coding unit, the nodes of the coding tree of the CTU are traversed at step 1530. When, as determined by "cu_qp_delta_subdiv", any child node of the current node has a subdivision level less than or equal to the subdivision level 1136 of the current slice, a new quantization group starts in the CTU region corresponding to that node, and step 1530 returns "yes". The first CU in the quantization group that includes the coded residual will also include the coded incremental QP, thus signaling any change in the quantization parameter applicable to the residual coefficients in that quantization group.
[0204] At the step 1540 of coding the incremental QP, the entropy encoder 338 codes the incremental QP in the bitstream 115. The incremental QP codes the difference between the predicted QP and the expected QP used in the current quantization group. The predicted QP is derived by averaging the QPs of adjacent earlier (above and to the left) quantization groups. When the subdivision level is low, the quantization group is large and the incremental QP is coded less frequently. The less frequent coding of the incremental QP results in a lower overhead for signaling changes in the QP, but also leads to less flexibility in rate control. The selection of the quantization parameter for each quantization group is performed by the QP controller module 390, which generally implements a rate control algorithm for a specific bitrate of the bitstream 115, to some extent independent of the variations in the statistics of the underlying frame data 113. Method 1500 proceeds from step 1540 to the step 1550 of performing the main transform.
[0205] At step 1550 where the primary transform is performed, the forward primary transform module 326 performs a primary transform according to the primary transform type of the coding unit, thereby obtaining primary transform coefficients 328. The primary transform is performed on each color channel. First, the primary transform is performed on the luminance channel (Y), and then the primary transform is performed on the Cb TB and Cr TB in subsequent calls to step 1550 for the current TU. For the luminance channel, the primary transform types (DCT-2, transform skip, MTS option) are performed, and for the chrominance channels, DCT-2 is performed.
[0206] Method 1500 continues from step 1550 to step 1560 where the primary transform coefficients are quantized. At step 1560, the quantizer module 334 quantizes the primary transform coefficients 328 according to the quantization parameter 392 to produce quantized primary transform coefficients 332. The delta QP (when present) is used to encode the transform coefficients 328.
[0207] Method 1500 continues from step 1560 to step 1570 where a secondary transform is performed. At step 1570, the secondary transform module 330 performs a secondary transform on the quantized primary transform coefficients 332 according to the secondary transform index 388 of the current transform block to produce secondary transform coefficients 336. Although the secondary transform is performed after quantization, the primary transform coefficients 328 can maintain a higher precision compared to the final expected quantizer step size of the quantization parameter 392. For example, the magnitude can be 16 times the magnitude directly produced by the application of the quantization parameter 392, that is, four additional bits of precision will be retained. Retaining the additional bit precision in the quantized primary transform coefficients 332 allows the secondary transform module 330 to operate on the coefficients in the primary coefficient domain with higher precision. After applying the secondary transform, the final scaling (e.g., right shift by four bits) at step 1560 quantizes to the expected quantizer step size of the quantization parameter 392. The "scaling list" is applied to the primary transform coefficients (which correspond to well-known transform basis functions (DCT-2, DCT-8, DST-7)), rather than operating on the secondary transform coefficients produced by the trained secondary transform kernel. When the secondary transform index 388 of the transform block indicates that no secondary transform is to be applied (index value equal to zero), the secondary transform is bypassed. That is, the primary transform coefficients 332 are propagated through the secondary transform module 330 unchanged to become the secondary transform coefficients 336. The luminance secondary transform index is used in combination with the luminance intra prediction mode to select the secondary transform kernel to be applied to the luminance TB. The chrominance secondary transform index is used in combination with the chrominance intra prediction mode to select the secondary transform kernel to be applied to the chrominance TB.
[0208] Method 1500 proceeds from step 1570 to step 1580 which encodes the last position. At step 1580, entropy encoder 338 encodes the position of the last significant coefficient among the secondary transform coefficients 336 of the current transform block in bitstream 115. When step 1580 is called for the first time, the luminance TB is considered, and subsequent calls consider the Cb TB and then the Cr TB.
[0209] In an arrangement where the secondary transform index 388 is encoded immediately after the last position, method 1500 proceeds to step 1590 which encodes the LFNST index. If the secondary transform index is not inferred to be zero based on the last position encoded at step 1580, at step 1590, entropy encoder 338 encodes the secondary transform index 338 as "lfnst_index" in bitstream 115 using a truncated unary codeword. Each CU has one luminance TB, thus allowing step 1590 for the luminance block, and when the "joint" encoding mode is used for chrominance, for a single chrominance TB, so step 1590 can be performed for chrominance. Knowing the secondary transform index before decoding each residual coefficient enables the application of the secondary transform coefficient by coefficient, for example using multiply - add logic when the coefficients are decoded. Method 1500 proceeds from step 1590 to step 15100 which encodes sub - blocks.
[0210] If the secondary transform index 388 is not encoded immediately after the last position, method 1500 proceeds from step 1580 to step 15100 which encodes sub - blocks. At step 15100 which encodes sub - blocks, the residual coefficients (336) of the current transform block are encoded in bitstream 115 as a series of sub - blocks. Proceeding back from the sub - block containing the last significant coefficient position to the sub - block containing the DC residual coefficient, the residual coefficients are encoded.
[0211] Method 1500 proceeds from step 15100 to step 15110 which is the last TB test. At this step, processor 205 tests whether the current transform block is the last transform block in progress on the color channels (i.e., Y, Cb, and Cr). If the just - encoded transform block is for the Cr TB (step 15110 is "yes"), the control in processor 205 advances to step 15120 which encodes the luminance LFNST index. Otherwise, if the current TB is not the last (step 15110 is "no"), the control in processor 205 returns to step 1550 which performs the main transform and the next TB (select Cb or Cr).
[0212] Steps 1550 to 15110 are described with an example of a common coding tree structure where the prediction mode is intra prediction and DCT-2 is used. Except for the common coding tree structure using known methods, operations of steps such as performing the primary transform (1550), quantizing the primary transform coefficients (1560), and coding the last position (1590) can be implemented for inter prediction modes or intra prediction modes. Steps 1510 to 1540 can be implemented regardless of the prediction mode or coding tree structure.
[0213] Method 1500 continues from step 15110 to step 15120 of coding the luminance LFNST index. At step 15120, if the secondary transform index applied to the luminance TB is not inferred to be zero (no secondary transform is applied), the entropy encoder 338 encodes it in the bitstream 115. The luminance secondary transform index is inferred to be zero if the last valid position of the luminance TB indicates valid only primary residual coefficients or if a primary transform other than DCT-2 is performed. Additionally, the secondary transform index applied to the luminance TB is encoded in the bitstream only for coding units using intra prediction and the common coding tree structure. The secondary transform index applied to the luminance TB is encoded using flag 1220 (or flag 1230 for the joint CbCr mode).
[0214] Method 1500 continues from step 15120 to step 15130 of coding the chrominance LFNST index. At step 1530, if the secondary transform index applied to the chrominance TB is not inferred to be zero (no secondary transform is applied), the chrominance secondary transform index is encoded in the bitstream 115 by the entropy encoder 338. The chrominance secondary transform index is inferred to be zero if the last valid position of any chrominance TB indicates valid only primary residual coefficients. Method 1500 terminates after performing step 15130, where the control in the processor 205 returns to method 1300. The secondary transform index applied to the chrominance TB is encoded in the bitstream only for coding units using intra prediction and the common coding tree structure. The secondary transform index applied to the chrominance TB is encoded using flag 1221 (or flag 1230 for the joint CbCr mode).
[0215] Figure 16 Method 1600 for decoding a frame from a bitstream that is a sequence of coding units arranged as strips is shown. Method 1600 can be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1600 can be performed by the video decoder 134 under the execution of the processor 205. Thus, method 1600 can be stored on a computer-readable storage medium and / or in the memory 206.
[0216] Method 1600 decodes a bitstream encoded using Method 1300, in which partition constraints and quantization group definitions can vary from one strip to another, which is considered beneficial for rate control purposes when encoding the respective parts (strips) of the encoded bitstream 115. Not only can the quantization group subdivision level vary from one strip to another, but also the application of the secondary transform is independently controllable for luminance and chrominance.
[0217] Method 1600 begins with the SPS / PPS decoding step 1610. In the execution of step 1610, the video decoder 134 decodes the SPS 1110 and PPS 1112 from the bitstream 133 into a sequence of fixed and variable length parameters. The partition_constraints_override_enabled_flag is decoded as part of the SPS 1110, indicating whether the partition constraints can be overridden in the strip header (e.g., 1118) of the corresponding strip (e.g., 1116). The default (i.e., as signaled in the SPS 1110 and used in the strip without subsequent override) partition constraint parameter 1130 is also decoded by the video decoder 134 as part of the SPS 1110.
[0218] Method 1600 continues from step 1610 to the step 1620 of determining the strip boundaries. In the execution of step 1620, the processor 205 determines the position of the strip in the current access unit in the bitstream 133. Generally, the strip is identified by determining the NAL unit boundaries (by detecting the "start code") and reading the NAL unit header including the "NAL unit type" for each NAL unit. A particular NAL unit type identifies the strip type, such as "I-strip", "P-strip", and "B-strip", etc. After identifying the strip boundaries, the application 233 can, for example, distribute the execution of the subsequent steps of Method 1600 across different processors in a multi-processor architecture for parallel decoding. Each processor in the multi-processor system can decode different strips to obtain a higher decoding throughput.
[0219] Method 1600 continues from step 1610 to the step 1630 of decoding the strip header. At step 1630, the entropy decoder 420 decodes the strip header 1118 from the bitstream 133. The following refers to Figure 17 Describes an example method of decoding the strip header 1118 from the bitstream 133 as implemented at step 1630.
[0220] Method 1600 continues from step 1630 to step 1640 of splitting the strip into CTUs. At step 1640, video decoder 134 splits strip 1116 into a sequence of CTUs. The strip boundaries are aligned with the CTU boundaries, and the CTUs in the strip are sorted according to the CTU scan order. The CTU scan order is typically a raster scan order. Splitting the strip into CTUs determines which part of frame data 113 will be processed by video decoder 134 when decoding the current strip.
[0221] Method 1600 continues from step 1640 to step 1650 of decoding the coding tree. In the execution of step 1650, video decoder 133 starts from the first CTU in strip 1116 when first invoking step 1650, and decodes the coding tree of the current CTU in the strip from bitstream 133. By Figure 6 decoding the split flag to decode the coding tree of the CTU. For subsequent iterations of step 1650 for the CTU, the subsequent CTUs in strip 1116 are decoded. If the coding tree is encoded using an intra prediction mode and a common coding tree structure, the coding unit has a primary color channel (luminance or Y) and at least one secondary color channel (chrominance, Cb and Cr or CbCr). In this case, decoding the coding tree involves decoding a coding unit including the primary color channel and at least one secondary color channel according to the split flag of the coding tree unit.
[0222] Method 1600 continues from step 1660 to step 1670 of decoding the coding unit. At step 1670, video decoder 134 decodes the coding unit from bitstream 133. An example method of decoding the coding unit implemented at step 1670 is described below with reference to Figure 18 description.
[0223] Method 1600 continues from step 1610 to step 1680 of the last coding unit test. At step 1680, processor 205 tests whether the current coding unit is the last coding unit in the CTU. If it is not the last coding unit (step 1680 is "no"), the control in processor 205 returns to step 1670 of decoding the coding unit to decode the next coding unit of the coding tree unit. If the current coding unit is the last coding unit (step 1680 is "yes"), the control in processor 205 advances to step 1690 of the last CTU test.
[0224] At step 1690 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU in slice 1116. If it is not the last CTU in the slice (step 1690 is "No"), then control in the processor 205 returns to step 1650 of decoding the coding tree to decode the next coding tree unit of slice 1116. If the current CTU is the last CTU of slice 1116 (step 1690 is "Yes"), then control in the processor 205 advances to step 16100 of the last slice test.
[0225] At the last slice test step 16100, the processor 205 tests whether the current slice being decoded is the last slice in the frame. If it is not the last slice in the frame (step 16100 is "No"), then control in the processor 205 returns to step 1630 of decoding the slice header, and step 1630 operates to decode the slice header of the next slice in the frame (e.g., Figure 11 "Slice 2") of the frame. If the current slice is the last slice in the frame (step 1600 is "Yes"), then method 1600 terminates.
[0226] As described regarding Figure 1 the apparatus 130 in
[0227] Figure 17 Method 1600 of operating on multiple coding units operates to produce an image frame.
[0228] Similar to method 1500, method 1700 is performed on the current slice or consecutive portions (1116) in a frame (e.g., frame 1101). Method 1700 begins with a partition constraint override enable test step 1710. At step 1710, the processor 205 tests whether the partition constraint override enable flag decoded from the SPS 1110 indicates that the partition constraint can be overridden at the slice level. If the partition constraint can be overridden at the slice level (step 1710 is "Yes"), then control in the processor 205 advances to step 1720 of decoding the partition constraint override flag. Otherwise, if the partition constraint override enable flag indicates that the constraint cannot be overridden at the slice level (step 1710 is "No"), then control in the processor 205 advances to step 1770 of decoding other parameters.
[0229] At step 1720 of decoding the slice constraint override flag, entropy decoder 420 decodes the slice constraint override flag from bitstream 133. The decoded flag indicates whether the slice constraints signaled in SPS 1110 are to be overridden for the current slice 1116.
[0230] Method 1700 continues from step 1720 to step 1730 of the slice constraint override test. In the execution of step 1730, processor 205 tests the value of the flag decoded at step 1720. If the decoded flag indicates that the slice constraints are to be overridden (step 1730 is "yes"), then control in processor 205 advances to step 1740 of decoding the slice partition constraints. Otherwise, if the decoded flag indicates that the slice constraints are not to be overridden (step 1730 is "no"), then control in processor 205 advances to step 1770 of decoding other parameters.
[0231] At step 1740 of decoding the slice partition constraints, entropy decoder 420 decodes the determined slice partition constraints from bitstream 133. The slice partition constraints for the slice include "slice_max_mtt_hierarchy_depth_luma" from which MaxMttDepthY 1134 is derived.
[0232] Method 1700 continues from step 1740 to step 1750 of decoding the QP subdivision level. At step 1720, entropy decoder 420 decodes the subdivision level of the luma CB using the "cu_qp_delta_subdiv" syntax element as described in reference Figure 11 as mentioned.
[0233] Method 1700 continues from step 1750 to step 1760 of decoding the chroma QP subdivision level. At step 1760, entropy decoder 420 decodes the subdivision level for signaling the CU chroma QP offset using the "cu_chroma_qp_offset_subdiv" syntax element as described in reference Figure 11 as mentioned.
[0234] Steps 1750 and 1760 operate to determine the subdivision levels of a particular consecutive portion (slice) of the bitstream. Repeated iterations between steps 1630 and 16100 operate to determine the subdivision levels of individual consecutive portions (slices) of the bitstream. As described below, each subdivision level applies to the coding units of the corresponding slice (consecutive portion).
[0235] Method 1700 continues from step 1760 to step 1770 of decoding other parameters. At step 1770, entropy decoder 420 decodes other parameters from slice header 1118, such as parameters required to control specific tools such as deblocking, adaptive loop filter, optionally selecting a scaling list from a previously signaled scaling list (for non-uniformly applying quantization parameters to transform blocks), etc. Method 1700 terminates when step 1770 is executed.
[0236] Figure 18 Method 1800 for decoding a coding unit from a bitstream is shown. Method 1800 can be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, method 1800 can be performed by video decoder 134 under the execution of processor 205. Thus, method 1800 can be stored on a computer-readable storage medium and / or in memory 206.
[0237] Method 1800 is implemented for a current coding unit of a current CTU (e.g., CTU0 of slice 1116). Method 1800 begins with step 1810 of decoding a prediction mode. At step 1800, entropy decoder 420 decodes the prediction mode of the coding unit determined at step 1360 of Figure 13 as from bitstream 133. At step 1810, the "pred_mode" syntax element is decoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit.
[0238] If intra prediction is used for the coding unit, then at step 1810, the luminance intra prediction mode and the chrominance intra prediction mode are also decoded. If inter prediction is used for the coding unit, then at step 1810, the "merge index" can also be decoded to determine a motion vector from neighboring coding units for use by this coding unit, and a motion vector delta can be decoded to introduce an offset to the motion vector derived from spatially neighboring blocks. Also at step 1810, the primary transform type is decoded to select between using DCT-2 horizontally and vertically, transform skip horizontally and vertically, or a combination of DCT-8 and DST-7 horizontally and vertically for the luminance TB of the coding unit.
[0239] Method 1800 continues from step 1810 to step 1820 of the coded residual test. In the execution of step 1820, the processor 205 determines whether the residual needs to be decoded for the coding unit by decoding the "root coding block flag" of the coding unit using the entropy decoder 420. If there are any valid residual coefficients to be decoded for the coding unit (step 1820 is "yes"), then the control in the processor 205 proceeds to step 1830 of the new QG test. Otherwise, if there are no residual coefficients to be decoded (step 1820 is "no"), then method 1800 terminates because all the information required to decode the coding unit has been obtained in the bitstream 115. When method 1800 terminates, subsequent steps such as PB generation, applying in-loop filtering, etc. are performed, thereby generating the decoded samples, as referenced Figure 4 as described.
[0240] At the new QG test step 1830, the processor 205 determines whether the coding unit corresponds to a new quantization group. If the coding unit corresponds to a new quantization group (step 1830 is "yes"), then the control in the processor 205 proceeds to step 1840 of decoding the delta QP. Otherwise, if the coding unit does not correspond to a new quantization group (step 1830 is "no"), then the control in the processor 205 proceeds to step 1850 of decoding the last position. The new quantization group is related to the current mode or the subdivision level of the coding unit. When decoding each coding unit, the nodes of the coding tree of the CTU are traversed. A new quantization group starts in the region of the CTU corresponding to the node when any child node of the current node has a subdivision level less than or equal to the subdivision level 1136 of the current stripe (i.e., as determined from "cu_qp_delta_subdiv"). The first CU in the quantization group that includes the coded residual coefficients will also include the coded delta QP, thereby signaling any change in the quantization parameter applicable to the residual coefficients in the quantization group. Effectively, a single (at most one) quantization parameter increment is decoded for each region (quantization group). As described with respect to Figures 8A to 8C as described, each region (quantization group) is based on the decomposition of the coding tree units of each stripe and the corresponding subdivision levels (e.g., as coded at steps 1460 and 1470). In other words, each region or quantization group is based on the comparison of the subdivision level associated with the coding unit with the determined subdivision levels for the corresponding consecutive parts.
[0241] At the step 1840 of decoding the delta QP, the entropy decoder 420 decodes the delta QP from the bitstream 133. The delta QP encodes the difference between the predicted QP and the expected QP used in the current quantization group. The predicted QP is derived by averaging the QPs of the adjacent (above and to the left) quantization groups.
[0242] Method 1800 continues from step 1840 to step 1850 of decoding the last position. When performing step 1850, entropy decoder 420 decodes the position of the last significant coefficient among the secondary transform coefficients 424 of the current transform block from bitstream 133. When step 1850 is called for the first time, this step is performed for the luminance TB. In subsequent calls to step 1850 for the current CU, this step is performed for the Cb TB. If the last position indicates that the significant coefficient is outside the set of secondary transform coefficients of the luminance block or chrominance block (i.e., outside 928 or 966), the secondary transform index of the luminance or chrominance channel is respectively inferred to be zero. After the iteration for Cb, this step is implemented for the Cr TB in the iteration.
[0243] As described for Figure 15 step 1590, in some arrangements, the secondary transform index is encoded immediately after the position of the last significant coefficient of the coding unit. When decoding the same coding unit, if the secondary transform index 470 is not inferred to be zero based on the positioning of the last position of the TB decoded in step 1840, the secondary transform index 470 is decoded immediately after decoding the position of the last significant residual coefficient of the coding unit. In the arrangement where the secondary transform index 470 is decoded immediately after the position of the last significant coefficient of the coding unit, method 1800 continues from step 1850 to step 1860 of decoding the LFNST index. When performing step 1860, entropy decoder 420 decodes the secondary transform index 470 from bitstream 133 as "lfnst_index" using truncated unary codewords when all significant coefficients are subject to inverse secondary transform (e.g., within 928 or 966). When performing joint coding of the chrominance TB using a single transform block, the secondary transform index 470 can be decoded for the luminance TB or chrominance. Method 1800 continues from step 1860 to step 1870 of decoding the sub-blocks.
[0244] If the secondary transform index 470 is not decoded immediately after the last significant position of the coding unit, method 1800 continues from step 1850 to step 1870 of decoding the sub-blocks. At step 1870, the residual coefficients (i.e., 424) of the current transform block are decoded from bitstream 133 as a series of sub-blocks, advancing from the sub-block containing the last significant coefficient position back to the sub-block containing the DC residual coefficient.
[0245] Method 1800 continues from step 1870 to step 1880 of the last TB test. During the execution of step 1880, the processor 205 tests whether the current transform block is the last transform block in progress on a color channel (i.e., Y, Cb, and Cr). If the just-decoded (current) transform block is for the Cr TB, then in the control in the processor 205, all TBs have been decoded (step 1880 is "yes"), and method 1800 proceeds to step 1890 of decoding the luminance LFNST index. Otherwise, if the TB has not been decoded (step 1880 is "no"), then the control in the processor 205 returns to step 1850 of decoding the last position. In the iteration of step 1850, the next TB (following the order of Y, Cb, Cr) is selected for decoding.
[0246] Method 1800 continues from step 1880 to step 1890 of decoding the luminance LFNST index. During the execution of step 1890, if the last position of the luminance TB is within the set of coefficients (e.g., 928 or 966) undergoing the inverse secondary transform and the luminance TB is using DCT-2 as the primary transform horizontally and vertically, then the entropy decoder 420 decodes from the bitstream 133 the secondary transform index 470 to be applied to the luminance TB. If the last valid position of the luminance TB indicates that there are valid primary coefficients outside the set of coefficients undergoing the inverse secondary transform (e.g., outside 928 or 966), then the luminance secondary transform index is inferred to be zero (no secondary transform is applied). The secondary transform index decoded at step 1890 is indicated as 1220 (or 1230 in the combined CbCr mode) in Figure 12 ...
[0247] Method 1800 continues from step 1890 to step 1895 of decoding the chrominance LFNST index. At step 1895, if the last position of each chrominance TB is within the set of coefficients (e.g., 928 or 966) undergoing the inverse secondary transform, then the entropy decoder 420 decodes from the bitstream 133 the secondary transform index 470 to be applied to the chrominance TB. If the last valid position of any chrominance TB indicates that there are valid primary coefficients outside the set of coefficients undergoing the inverse secondary transform (e.g., outside 928 or 966), then the chrominance secondary transform index is inferred to be zero (no secondary transform is applied). The secondary transform index decoded at step 1895 is indicated as 1221 (or 1230 in the combined CbCr mode) in Figure 12 ... When decoding the individual indices for luminance and chrominance, separate arithmetic contexts for the respective truncated unary codes can be used or the contexts can be shared such that the n-th bin in each of the luminance and chrominance truncated unary codes shares the same context.
[0248] Effectively, steps 1890 and 1895 respectively involve: decoding a first index (such as 1220 etc.) to select a kernel for the luminance (primary color) channel, and decoding a second index (such as 1221 etc.) to select a kernel for at least one chrominance (secondary color channel).
[0249] Method 1800 continues from step 1895 to step 18100 for performing an inverse secondary transform. At this step, the inverse secondary transform module 436 performs an inverse secondary transform on the decoded residual transform coefficients 424 according to the secondary transform index 470 of the current transform block to generate secondary transform coefficients 432. The secondary transform index decoded at step 1890 is applied to the luminance TB, and the secondary transform index decoded at step 1895 is applied to the chrominance TB. The kernel selection for luminance and chrominance also respectively depends on the luminance intra prediction mode and the chrominance intra prediction mode (which are each decoded at step 1810). Step 18100 selects a kernel according to the LFNST index of luminance, and selects a kernel according to the LFNST index of chrominance.
[0250] Method 1800 continues from step 18100 to step 18110 for inverse quantizing the primary transform coefficients. At step 18110, the inverse quantizer module 428 inverse quantizes the secondary transform coefficients 432 according to the quantization parameter 474 to generate inverse quantized primary transform coefficients 440. If the incremental QP is decoded at step 1840, the entropy decoder 420 determines the quantization parameter according to the incremental QP of the quantization group (region) and the quantization parameter of the earlier encoded unit of the image frame. As described above, the earlier encoded unit generally refers to the adjacent upper left encoded unit.
[0251] Method 1800 continues from step 1870 to step 18120 for performing the primary transform. At step 18120, the inverse primary transform module 444 performs an inverse primary transform according to the primary transform type of the coding unit, such that the transform coefficients 440 are converted into residual samples 448 in the spatial domain. The inverse primary transform is performed on each color channel. First, the inverse primary transform is performed on the luminance channel (Y), and then the inverse primary transform is performed on the Cb and Cr TBs when making a subsequent call to step 1650 for the current TU. Steps 18100 to 18120 effectively operate to decode the current coding unit by applying the kernel selected according to the LFNST index of luminance to the decoded residual coefficients of the luminance channel, and applying the kernel selected according to the LFNST index of chrominance to the decoded residual coefficients of at least one chrominance channel.
[0252] Method 1800 terminates after executing step 18120, where the control in the processor 205 returns to method 1600.
[0253] Steps 1850 to 18120 are described with an example of a common coding tree structure where the prediction mode is intra prediction and the transform is DCT-2. For example, for coding units using intra prediction and the common coding tree structure only, the secondary transform index (1890) applied to the luminance TB is decoded from the bitstream. Similarly, for coding units using intra prediction and the common coding tree structure only, the secondary transform index (1895) applied to the chrominance TB is decoded from the bitstream. Except for the common coding tree structure using known methods, operations of steps such as decoding sub-blocks (1870), inverse quantizing the primary transform coefficients (18110), and performing the primary transform can be implemented for inter prediction modes or for intra prediction modes. Regardless of the prediction mode or structure, steps 1810 to 1840 are performed in the described manner.
[0254] Once method 1800 terminates, subsequent steps for decoding the coding unit are performed (including generating intra prediction samples 480 by module 476, summing the decoded residual samples 448 with the prediction block 452 by module 450, and applying the in-loop filter module 488 to produce filtered samples 492), which are output as frame data 135.
[0255] Figure 19A and 19B Shows rules for applying or bypassing the secondary transform to the luminance and chrominance channels. Figure 19A Shows Table 1900 illustrating conditions for applying the secondary transform in the luminance and chrominance channels in a CU generated by the common coding tree.
[0256] If the last significant coefficient position of the luminance TB indicates decoded significant coefficients that are not generated by the forward secondary transform and thus not subject to inverse secondary transform, there is condition 1901. If the last significant coefficient position of the luminance TB indicates decoded significant coefficients that are indeed generated by the forward secondary transform and thus subject to inverse secondary transform, there is condition 1902. Additionally, for the luminance channel, the primary transform type needs to be DCT-2 for condition 1902 to exist, otherwise condition 1901 exists.
[0257] If the last significant coefficient position of one or both chrominance TBs indicates decoded significant coefficients that are not generated by the forward secondary transform and thus not subject to inverse secondary transform, there is condition 1910. If the last significant coefficient position of one or both chrominance TBs indicates decoded significant coefficients that are indeed generated by the forward secondary transform and thus subject to inverse secondary transform, there is condition 1911. Additionally, the width and height of the chrominance block need to be at least four samples (e.g., chrominance subsampling when using 4:2:0 or 4:2:2 chrominance formats can result in a width or height of two samples) for condition 1911 to exist.
[0258] If conditions 1901 and 1910 exist, the secondary transform index is not signaled (either independently or jointly), and the secondary transform index is not applied in the luma or chroma, i.e., 1920. If conditions 1901 and 1911 exist, a secondary transform index is signaled to indicate the application of the selected kernel or bypass for only the luma channel, i.e., 1921. If conditions 1902 and 1910 exist, a secondary transform index is signaled to indicate the application of the selected kernel or bypass for only the chroma channel, i.e., 1922. If conditions 1911 and 1902 exist, an arrangement with independent signaling signals two secondary transform indexes, one for luma TB and one for chroma TB, i.e., 1923. When conditions 1902 and 1911 exist, an arrangement with a single signaled secondary transform index uses one index to control the selection for both luma and chroma, although the selected kernel also depends on the luma and chroma intra prediction modes, which may vary. The ability to apply the secondary transform to luma or chroma (i.e., 1921 and 1922) results in increased coding efficiency.
[0259] Figure 19B Table 1950 showing the search options available to video encoder 114 at step 1360. The secondary transform indexes for luma (1952) and chroma (1953) are shown as 1952 and 1953, respectively. An index value of 0 indicates bypassing the secondary transform, and index values of 1 and 2 indicate which of two kernels to use for a candidate set derived from the luma or chroma intra prediction mode. There are nine combinations ("0,0" to "2,2") of the resulting search space, which may be constrained by the constraints described in reference Figure 19A Compared to searching all admissible combinations, a simplified search (1951) of three combinations can test only combinations where the luma and chroma secondary transform indexes are the same, subject to zeroing the index indicating the last significant coefficient position for the channel with only primary coefficients. For example, when condition 1921 exists, the options "1,1" and "2,2" become "0,1" and "0,2", respectively (i.e., 1954). When condition 1922 exists, the options "1,1" and "2,2" become "1,0" and "2,0", respectively (i.e., 1955). When condition 1920 exists, no secondary transform index needs to be signaled, and option "0,0" is used. In fact, conditions 1921 and 1922 allow the options "0,1", "0,2", "1,0", and "2,0" in a common tree CU, resulting in higher compression efficiency. If these options are prohibited, either condition 1901 or 1910 will cause condition 1920 (i.e., options "1,1" and "2,2") to be prohibited, such that "0,0" is used (see 1956).
[0260] Signaling the quantization group subdivision level in the slice header provides a higher level of granularity of control below the picture level. The higher level of granularity of control is beneficial for applications where the encoding fidelity requirements vary from one part of the picture to another, and in particular for applications where multiple encoders may need to operate slightly independently to provide real-time processing capabilities. Signaling the quantization group subdivision level in the slice header is also consistent with the partition override setting and the Scaling List application setting in the slice header.
[0261] In one arrangement of the video encoder 114 and the video decoder 134, the secondary transform index of the chrominance intra prediction block is always set to zero, i.e., the secondary transform is not applied to the chrominance intra prediction block. In this case, signaling the chrominance secondary transform index is not required, and thus steps 15130 and 1895 can be omitted, and steps 1360, 1570, and 18100 can be simplified accordingly.
[0262] If a node in the coding tree in the common tree has a region of 64 luma samples, further splitting using a binary or quadtree split will result in smaller luma CUs, such as 4×4 blocks, etc., but will not result in smaller chroma CUs. Instead, there is a single chroma CU corresponding to the region of 64 luma samples, such as a 4×4 chroma CU, etc. Similarly, a coding tree node with a region of 128 luma samples and subject to a ternary split results in a set of smaller luma CUs and one chroma CU. Each luma CU has a corresponding luma secondary transform index, and the chroma CU has a chroma secondary transform index.
[0263] When a node in the coding tree has a region of 64 and signaling a further split or has a region of 128 luma samples and signaling a ternary split, the split is applied only in the luma channel, and the resulting CUs (a number of luma CUs and one chroma CU for each chroma channel) are all intra-predicted or all inter-predicted. When a CU has a width or height of four luma samples and includes one CU for each of the color channels (Y, Cb, and Cr), the chroma CU of the CU has a width or height of two samples. A CU with a width or height of two samples does not utilize 16-point or 48-point LFNST kernel operations and thus does not require a secondary transform. For a block with a width or height of two samples, steps 15130, 1895, 1360, 1570, and 18100 are not required.
[0264] In another arrangement of video encoder 114 and video decoder 134, a single secondary transform index is signaled when either or both of the luma and chroma contain non-significant residual coefficients only in regions of the respective TBs that are only subject to the primary transform. If a luma TB contains significant residual coefficients in a non-secondary-transform region of the decoded residuals (e.g., 1066, 968), or is indicated to not use DCT-2 as the primary transform, the indicated secondary transform kernel (or secondary transform bypass) is only applied to the chroma TBs. If any of the chroma TBs contain significant residual coefficients in a non-secondary-transform region of the decoded residuals, the indicated secondary transform kernel (or secondary transform bypass) is only applied to the luma TB. Applying the secondary transform becomes possible for the luma TB even when it is not possible for the chroma TBs, and vice versa, resulting in an increase in coding efficiency compared to requiring that the last positions of all TBs be within the secondary coefficient domain before any TB of a CU can be subject to the secondary transform. Additionally, only one secondary transform index is required for CUs in a common coding tree. When the luma primary transform is DCT-2, the secondary transform can be inferred as disabled for both chroma and luma.
[0265] In another arrangement of video encoder 114 and video decoder 134, the secondary transform (by modules 330 and 436, respectively) is only applied to the luma TBs of a CU and not to any of the chroma TBs of that CU. The absence of secondary transform logic for the chroma channels results in lower complexity, such as lower execution time or reduced silicon area. The absence of secondary transform logic for the chroma channels allows for only one secondary transform index to be signaled, which can be signaled after the last position of the luma TBs. That is, instead of steps 15120 and 1890, steps 1590 and 1860 are performed for the luma TBs. In this case, steps 15130 and 1895 are omitted.
[0266] In another arrangement of the video encoder 114 and the video decoder 134, syntax elements defining the quantization group size (i.e., cu_chroma_qp_offset_subdiv and cu_qp_delta_subdiv) are signaled in the PPS 1112. Even if the partitioning constraint is overridden in the slice header 1118, the range of the finely-granulated levels is defined according to the partitioning constraint signaled in the SPS 1110. For example, the ranges of cu_qp_delta_subdiv and cu_chroma_qp_offset_subdiv are defined as 0 to 2*(log2_ctu_size_minus5 + 5 - (MinQtLog2SizeInterY or MinQtLog2SizeIntraY) + MaxMttDepthY_SPS). The value MaxMttDepthY is derived from the SPS 1110. That is, when the current slice is an I slice, MaxMttDepthY is set to be equal to sps_max_mtt_hierarchy_depth_intra_slice_luma, and when the current slice is a P or B slice, MaxMttDepthY is set to be equal to sps_max_mtt_hierarchy_depth_inter_slice. For a slice whose partitioning constraint is overridden to be shallower than the depth signaled in the SPS 1110, if the quantization group subdivision level determined from the PPS 1112 is higher (deeper) than the highest achievable subdivision level at the shallower coding tree depth determined from the slice header, the quantization group subdivision level of the slice is clipped to be equal to the highest achievable subdivision level of the slice. For example, for a specific slice, cu_qp_delta_subdiv and cu_chroma_qp_offset_subdiv are clipped to be within 0 to 2*(log2_ctu_size_minus5 + 5 - (MinQtLog2SizeInterY or MinQtLog2SizeIntraY) + MaxMttDepthY_slice_header), and the clipped values are used for the slice. The value MaxMttDepthY_slice_header is derived from the slice header 1118, i.e., MaxMttDepthY_slice_header is set to be equal to slice_max_mtt_hierarchy_depth_luma.
[0267] In yet another arrangement of the video encoder 114 and the video decoder 134, the subdivision levels are determined based on cu_chroma_qp_offset_subdiv and cu_qp_delta_subdiv decoded from the PPS 1112 to derive the luminance and chrominance subdivision levels. When the partition constraints decoded from the slice header 1118 result in different ranges of the slice's subdivision levels, the subdivision levels applied to the slice are adjusted according to the partition constraints decoded from the SPS 1110 to maintain the same offset relative to the deepest allowed subdivision level. For example, if the SPS 1110 indicates a maximum subdivision level of 4, the PPS 1112 indicates a subdivision level of 3, and the slice header 1118 reduces the maximum value to 3, the subdivision level applied within the slice is set to 2 (maintaining an offset of 1 relative to the maximum allowed subdivision level). Adjusting the quantization group regions to correspond to changes in the partition constraints of a particular slice allows the subdivision levels to be signaled less frequently (i.e., at the PPS level), while providing the granularity to adapt to changes in the slice-level partition constraints. Using an arrangement where the subdivision levels are signaled in the PPS 1112 according to the ranges defined by the partition constraints decoded from the SPS 1110 (where adjustments may be made later based on the override partition constraints decoded from the slice header 1118) avoids the parsing dependency problem of making the PPS syntax elements dependent on the partition constraints done in the slice header 1118.
[0268] Industrial Applicability
[0269] The described arrangement is applicable to the computer and data processing industries and, in particular, to digital signal processing for encoding or decoding signals such as video and image signals, thereby achieving high compression efficiency.
[0270] The arrangement described here increases the flexibility provided to the video encoder when generating a highly compressed bitstream from incoming video data. The quantization of different regions or sub-pictures in a frame can be controlled with varying granularity and different granularity from one region to another, thereby reducing the amount of encoded residual data. When needed, for example, for a 360-degree image as described above, a higher granularity can be achieved accordingly.
[0271] In some arrangements, as described with respect to steps 15120 and 15130 (and correspondingly steps 1890 and 1895), the application of the secondary transformation can be controlled independently for luminance and chrominance, thereby achieving a further reduction in the encoded residual data. A video decoder is described that has the necessary functionality to decode the bitstream produced by such a video encoder.
[0272] The foregoing merely illustrates some embodiments of the present invention, and the present invention can be modified and / or changed without departing from the scope and spirit of the present invention, where these embodiments are merely exemplary and not restrictive.
[0273] Citation of Related Applications
[0274] This application claims the benefit of priority of Australian Patent Application No. 2019232801, filed on September 17, 2019, under 35 U.S.C. § 119, the entire contents of which are incorporated herein by reference for all purposes.
Claims
1. A method for decoding a coding unit in a coding tree unit of an image from a bitstream, the coding unit having a luminance channel and a chrominance channel, the method comprising: Determining the coding unit having the luminance channel and the chrominance channel according to one or more split flags of the coding tree unit, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a horizontal ternary split; Decoding, from the bitstream, a first index for selecting a non-separable transform kernel for the luminance channel; Selecting the non-separable transform kernel according to the first index; Decoding, from the bitstream, coefficients of a luminance transform block of the luminance channel in the coding unit and coefficients of a chrominance transform block of the chrominance channel in the coding unit; Performing non-separable transform on the coefficients of the luminance transform block by applying the selected non-separable transform kernel to derive non-separable transform coefficients of the luminance transform block; And Decoding the coding unit by performing separable transform on the non-separable transform coefficients of the luminance transform block and on the coefficients of the chrominance transform block, wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting the non-separable transform kernel for the luminance channel can be decoded, (b) a second index for selecting the non-separable transform kernel for the chrominance channel is not decoded, (c) only the non-separable transform can be performed on the coefficients of the luminance transform block in the coding unit, and (d) the non-separable transform is not performed on the coefficients of the chrominance transform block in the coding unit, and the width and height of the chrominance transform block are both equal to or greater than 4, and wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there are chrominance coding blocks corresponding to the given area, (a) the first index for selecting the non-separable transform kernel for the luminance channel can exist separately for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting the non-separable transform kernel for the chrominance channel can exist for the chrominance coding blocks corresponding to the given area.
2. A method for encoding a coding unit in a coding tree unit of an image in a bitstream, the coding unit having a luminance channel and a chrominance channel, the method comprising: Determining the coding unit having the luminance channel and the chrominance channel, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a horizontal ternary split; (a) Perform a separable transform on the coefficients of the luminance transform block of the luminance channel in the coding unit to derive the coefficients after the separable transform of the luminance transform block, and (b) perform a separable transform on the coefficients of the chrominance transform block of the chrominance channel in the coding unit to derive the coefficients after the separable transform of the chrominance transform block; Select a non-separable transform kernel for the luminance channel; Perform a non-separable transform on the coefficients after the separable transform of the luminance transform block by applying the selected non-separable transform kernel; And Encode a first index for selecting the non-separable transform kernel for the luminance channel in the bitstream, wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting the non-separable transform kernel for the luminance channel can be encoded, (b) the second index for selecting the non-separable transform kernel for the chrominance channel is not encoded, (c) the non-separable transform can be performed only on the coefficients after the separable transform of the luminance transform block in the coding unit, and (d) the non-separable transform is not performed on the coefficients after the separable transform of the chrominance transform block in the coding unit, the width and height of the chrominance transform block are both equal to or greater than 4, and wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there is a chrominance coding block corresponding to the given area, (a) the first index for selecting the non-separable transform kernel for the luminance channel can exist separately for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting the non-separable transform kernel for the chrominance channel can exist for the chrominance coding block corresponding to the given area.
3. An apparatus for decoding a coding unit in a coding tree unit of an image from a bitstream, the coding unit having a luminance channel and a chrominance channel, the apparatus comprising: A determination unit configured to determine a coding unit having the luminance channel and the chrominance channel according to one or more split flags of the coding tree unit, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a horizontal ternary split; A first decoding unit configured to decode from the bitstream a first index for selecting a non-separable transform kernel for the luminance channel; A selection unit configured to select the non-separable transform kernel according to the first index; A second decoding unit configured to decode from the bitstream the coefficients of the luminance transform block of the luminance channel in the coding unit and the coefficients of the chrominance transform block of the chrominance channel in the coding unit; An execution unit configured to perform a non-separable transform on coefficients of the luminance transform block by applying a selected non-separable transform kernel to derive non-separable transform coefficients of the luminance transform block; And A third decoding unit configured to decode the coding unit by performing a separable transform on the non-separable transform coefficients of the luminance transform block and on the coefficients of the chrominance transform block, Wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting the non-separable transform kernel for the luminance channel can be decoded, (b) the second index for selecting the non-separable transform kernel for the chrominance channel is not decoded, (c) the non-separable transform can be performed only on the coefficients of the luminance transform block in the coding unit, and (d) the non-separable transform is not performed on the coefficients of the chrominance transform block in the coding unit, the width and height of the chrominance transform block are both equal to or greater than 4, and Wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there are chrominance coding blocks corresponding to the given area, (a) the first index for selecting the non-separable transform kernel for the luminance channel can exist separately for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting the non-separable transform kernel for the chrominance channel can exist for the chrominance coding blocks corresponding to the given area.
4. A device for encoding a coding unit in a coding tree unit of an image in a bitstream, the coding unit having a luminance channel and a chrominance channel, the device comprising: A determination unit configured to determine a coding unit having the luminance channel and the chrominance channel, wherein the coding unit can be one of a plurality of coding units obtained from one or more splits in the coding tree unit, and the one or more splits can include a horizontal ternary split; A first execution unit configured to (a) perform a separable transform on coefficients of a luminance transform block of the luminance channel in the coding unit to derive separable transform coefficients of the luminance transform block, and (b) perform a separable transform on coefficients of a chrominance transform block of the chrominance channel in the coding unit to derive separable transform coefficients of the chrominance transform block; A selection unit configured to select a non-separable transform kernel for the luminance channel; A second execution unit configured to perform a non-separable transform on the separable transform coefficients of the luminance transform block by applying the selected non-separable transform kernel; And A coding unit configured to encode a first index for selecting the non-separable transform kernel for the luminance channel in the bitstream, Wherein, when the coding tree of the luminance channel in the coding tree unit is the same as the coding tree of the chrominance channel in the coding tree unit, (a) the first index for selecting the non-separable transform kernel for the luminance channel can be coded, (b) the second index for selecting the non-separable transform kernel for the chrominance channel is not coded, (c) only the coefficients after separable transformation of the luminance transform block in the coding unit can be subjected to the non-separable transformation, and (d) the coefficients after separable transformation of the chrominance transform block in the coding unit are not subjected to the non-separable transformation, the width and height of the chrominance transform block are both equal to or greater than 4, and Wherein, when the coding tree of the luminance channel in the coding tree unit is separated from the coding tree of the chrominance channel in the coding tree unit, a given area in the coding tree unit is split into luminance coding blocks, and there are chrominance coding blocks corresponding to the given area, (a) the first index for selecting the non-separable transform kernel for the luminance channel can exist individually for each luminance coding block in the luminance coding blocks, and (b) the second index for selecting the non-separable transform kernel for the chrominance channel can exist for the chrominance coding blocks corresponding to the given area.
5. A non-transitory computer-readable storage medium comprising computer-executable instructions that cause a computer to perform the method according to claim 1.
6. A non-transitory computer-readable storage medium comprising computer-executable instructions that cause a computer to perform the method according to claim 2.
7. A computer program product comprising a program that, when executed by a computer, causes the computer to perform the method according to claim 1.
8. A computer program product comprising a program that, when executed by a computer, causes the computer to perform the method according to claim 2.