Method, apparatus, medium, and computer program product for encoding and decoding coding units in a coding tree unit of an image
By introducing a decoding coding tree unit method in the video encoding technology, selecting appropriate core decoding video bitstreams, the problem of difficult to balance the compression performance and implementation costs of high-resolution and high-frame-rate videos in the prior art is solved, and efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202080062643.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-17
- Filing Date
- 2020-08-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-08-04
AI Technical Summary
Existing video encoding technologies are difficult to find a suitable balance between compression performance and implementation costs when processing high resolution and high frame rate videos, especially in network environments with high bandwidth costs.
A method of decoding a coding tree unit from a video bit stream is proposed, the method comprising determining the coding unit according to the decoded split flag of the coding tree unit and decoding the residual coefficients of the main color channel and the secondary color channel by selecting an appropriate core.
Through this method, the compression performance and decoding efficiency of video encoding can be effectively improved, and it is suitable for high resolution and high frame rate video data, especially in network environments with high bandwidth cost.
Smart Images

Figure CN114342391B_ABST
Abstract
Description
[0001] Citation of Related Applications
[0002] This application claims the benefit of priority under 35 U.S.C. §119 to Australian patent application No. 2019232801 filed on September 17, 2019, the entire contents of which are hereby incorporated by reference for all purposes. Technical Field
[0003] The present invention generally relates to digital video signal processing, and more particularly to methods, devices and systems for encoding and decoding blocks of video samples. The present invention also relates to a computer program product comprising a computer readable medium having recorded thereon a computer program for encoding and decoding blocks of video samples. Background Art
[0004] There are currently many applications for video encoding, including applications for transmitting and storing video data. Many video encoding standards have also been developed and other video encoding standards are currently under development. Recent advances in video coding standardization have led to the formation of a group known as the "Joint Video Experts Group" (JVET). The Joint Video Experts Group (JVET) includes members of Study Group 16, Question 6 (SG16 / Q6) of the Telecommunication Standardization Sector (ITU-T) of the International Telecommunication Union (ITU), also known as the "Video Coding Experts Group" (VCEG); and members of the International Organization for Standardization / International Electrotechnical Commission Joint Technical Committee 1 / Subcommittee 29 / Working Group 11 (ISO / IEC JTC1 / SC29 / WG11), also known as the "Moving Picture Experts Group" (MPEG).
[0005] The Joint Video Experts Team (JVET) issued a call for proposals (CfP) and analyzed the responses at its 10th meeting in San Diego, USA. The responses submitted showed that the video compression capabilities are significantly better than those of the current state-of-the-art video compression standard, namely "High Efficiency Video Coding" (HEVC). Based on this excellent performance, it was decided to start a project to develop a new video compression standard named "Versatile Video Coding" (VVC). It is expected that VVC will address the continued demand for even higher compression performance, especially as the capabilities of video formats increase (e.g., with higher resolutions and higher frame rates), and the growing market demand for service provision over WANs (where bandwidth costs are relatively high). Use cases such as immersive video require real-time encoding and decoding of such higher formats, for example, a cube map projection (CMP) can use an 8K format even if the final rendered "viewport" utilizes a lower resolution. VVC must be implementable in contemporary silicon processes and provide an acceptable tradeoff between the achieved performance and the implementation cost. For example, the implementation cost can be considered in one or more of silicon area, CPU processor load, memory utilization, and bandwidth. Higher video formats can be processed by splitting the frame area into parts and processing the parts in parallel. A bitstream constructed from multiple parts of a compressed frame is still suitable for decoding by a "single-core" decoder, i.e., frame-level constraints (including bit rate) are assigned to the parts according to the application needs.
[0006] The video data comprises a sequence of frames of image data, each frame comprising one or more color channels. Typically, one primary color channel and two secondary color channels are required. The primary color channel is often referred to as the "luminance" channel, and the (one or more) secondary color channels are often referred to as the "chrominance" channels. Although video data is often displayed in an RGB (red-green-blue) color space, this color space has a high degree of correlation between the three corresponding components. The video data representation seen by the encoder or decoder often uses a color space such as YCbCr. YCbCr concentrates luminance (mapped to "luminance" according to the transformation equation) in the Y (primary) channel and concentrates chrominance in the Cb and Cr (secondary) channels. Due to the use of decorrelated YCbCr signals, the statistics of the luminance channel are significantly different from those of the chrominance channels. The main difference is that after quantization, the chrominance channel contains relatively fewer significant coefficients for a given block compared to the coefficients of the corresponding luminance channel block. Additionally, the Cb and Cr channels may be spatially sampled at a lower rate than the luma channel, for example, half horizontally and half vertically (referred to as a "4:2:0 chroma format"). The 4:2:0 chroma format is commonly used in "consumer" applications such as Internet video streaming, broadcast television, and Blu-ray. TMStorage on disk. Subsampling the Cb and Cr channels at half rate horizontally instead of vertically is referred to as the "4:2:2 chroma format". The 4:2:2 chroma format is commonly used in professional applications, including the capture of footage for filmmaking, etc. The higher sampling rate of the 4:2:2 chroma format makes the resulting video more resilient to editing operations (such as color grading, etc.). Before distribution to consumers, 4:2:2 chroma format material is often converted to 4:2:0 chroma format and then encoded for distribution to consumers. In addition to the chroma format, video is also characterized by resolution and frame rate. Example resolutions are ultra-high definition (UD) with a resolution of 3840×2160 or "8K" with a resolution of 7680×4320, and example frame rates are 60 Hz or 120 Hz. The luminance sample rate can range from about 500 megasamples / second to several gigasamples / second. For the 4:2:0 chroma format, the sampling rate of each chroma channel is one quarter of the luma sampling rate, and for the 4:2:2 chroma format, the sampling rate of each chroma channel is half of the luma sampling rate.
[0007] The VVC standard is a "block-based" codec in which a frame is first partitioned into an array of square areas called "coding tree units" (CTUs). A CTU typically occupies a relatively large area, such as 128×128 luma samples. However, the areas of the CTUs at the right and bottom edges of each frame may be smaller. Associated with each CTU is a "coding tree" ("common tree") for both the luma channel and the chroma channel, or separate trees for each luma channel and chroma channel. The coding tree defines a decomposition of the area of the CTU into a set of areas, also called "coding blocks" (CBs). When a common tree is used, a single coding tree specifies blocks for both the luma channel and the chroma channel, in which case a collection of juxtaposed coding blocks is called a "coding unit" (CU), i.e., each CU has a coding block for each color channel. CBs are processed in a specific order for encoding or decoding. As a result of using the 4:2:0 chroma format, a CTU that includes a luma coding tree for a 128×128 luma sample area has a corresponding chroma coding tree for a 64×64 chroma sample area collocated with the 128×128 luma sample area. When a single coding tree is used for both the luma channel and the chroma channels, the collection of collocated blocks for a given area is often referred to as a "unit", such as the above-mentioned CU as well as the "prediction unit" (PU) and "transform unit" (TU). A single tree with a CU spanning the color channels of 4:2:0 chroma format video data produces chroma blocks of half the width and height of the corresponding luma blocks. When separate coding trees are used for a given area, the above-mentioned CBs as well as the "prediction blocks" (PB) and "transform blocks" (TB) are used.
[0008] Despite the above distinction between a “unit” and a “block,” the term “block” may be used as a general term for an area or region of a frame to which an operation is applied to all color channels.
[0009] For each CU, a prediction unit (PU) ("prediction unit") is generated for the content (sample values) of the corresponding region of the frame data. In addition, a representation of the difference between the prediction and the region content (or "residual" in the spatial domain) seen at the input of the encoder is formed. The differences of each color channel can be transformed and encoded as a sequence of residual coefficients, thereby forming one or more TUs for a given CU. The applied transform can be a discrete cosine transform (DCT) or other transform applied to each block of residual values. The transform is applied separately, that is, a two-dimensional transform is performed in two passes. The block is first transformed by applying a one-dimensional transform to each row of samples in the block. The partial result is then transformed by applying a one-dimensional transform to each column of the partial result to produce a final block of transform coefficients that essentially decorrelate the residual samples. The VVC standard supports transforms of various sizes, including transforms of rectangular blocks (each side size is a power of 2). The transform coefficients are quantized for entropy encoding in the bitstream.
[0010] VVC is characterized by intra-frame prediction and inter-frame prediction. Intra-frame prediction involves using previously processed samples in the frame being used to generate a prediction of the current sample block in the frame. Inter-frame prediction involves using a block of samples obtained from a previously decoded frame to generate a prediction of the current sample block in the frame. The block of samples obtained from the previously decoded frame is offset from the spatial position of the current block according to a motion vector, which motion vector usually has filtering applied. The intra-frame prediction block can be (i) a uniform sample value ("DC intra-frame prediction"), (ii) a plane with an offset and horizontal and vertical gradients ("plane intra-frame prediction"), (iii) a group of blocks with neighboring samples applied in a specific direction ("angle intra-frame prediction"), or (iv) the result of a matrix multiplication using neighboring samples and selected matrix coefficients. By encoding the 'residual' in the bitstream, further differences between the predicted block and the corresponding input sample can be corrected to some extent. The residual is usually transformed from the spatial domain to the frequency domain to form residual coefficients (in the "primary transform domain"), and the residual coefficients can be further transformed by applying a "secondary transform" (to produce residual coefficients in the "secondary transform domain"). The residual coefficients are quantized according to a quantization parameter, resulting in a loss of precision in the reconstruction of the samples produced at the decoder, while the bit rate within the bitstream is also reduced. The quantization parameter can vary between frames and within individual frames. For a "rate controlled" encoder, variations in the quantization parameter within a frame are typical. Regardless of the statistics of the received input samples (such as noise properties, degree of motion, etc.), a rate controlled encoder attempts to produce a bitstream with a substantially constant bit rate. Since the bitstream is typically transmitted over a network with limited bandwidth, rate control is a common technique used to ensure reliable performance on the network regardless of variations in the original frames input to the encoder. In the case where frames are encoded in parallel segments, flexibility in the use of rate control is desirable because different segments may have different requirements in terms of desired fidelity. Summary of the invention
[0011] It is an object of the present invention to substantially overcome or at least ameliorate one or more disadvantages of existing arrangements.
[0012] One aspect of the present disclosure provides a method for decoding a coding unit of a coding tree from a coding tree unit of an image frame from a video bitstream, the coding unit having a primary color channel and at least one secondary color channel, the method comprising: determining a coding unit including a primary color channel and at least one secondary color channel according to a decoded split flag of the coding tree unit; decoding a first index to select a core for the primary color channel, and decoding a second index to select a core for the at least one secondary color channel; selecting a first core according to the first index, and selecting a second core according to the second index; and decoding the coding unit by applying the first core to a residual coefficient of the primary color channel and applying the second core to a residual coefficient of at least one secondary color channel.
[0013] According to another aspect, the first index or the second index is decoded immediately after the position of the last significant residual coefficient of the coding unit is decoded.
[0014] According to another aspect, a single residual coefficient is decoded for multiple secondary color channels.
[0015] According to another aspect, a single residual coefficient is decoded for a single secondary color channel.
[0016] According to another aspect, the first index and the second index are independent of each other.
[0017] According to another aspect, the first kernel and the second kernel depend on intra prediction modes for the primary color channel and the at least one secondary color channel, respectively.
[0018] According to another aspect, the first kernel and the second kernel are associated with a block size of the primary channel and a block size of the at least one secondary color channel, respectively.
[0019] According to another aspect, the second kernel is related to a chroma subsampling rate of the encoded bitstream.
[0020] According to another aspect, individual ones of the cores implement inseparable secondary transforms.
[0021] According to another aspect, the encoding unit includes two secondary color channels, and a separate index is decoded for each of the secondary color channels.
[0022] Another aspect of the present disclosure provides a method for decoding a coding unit of a coding tree from a coding tree unit of an image frame from a video bitstream, the coding unit having a primary color channel and at least one secondary color channel, the method comprising: determining a coding unit including the primary color channel and the at least one secondary color channel according to a decoded split flag of the coding tree unit; selecting an inseparable transform kernel according to a decoded index of the primary color channel; applying the selected inseparable transform kernel to a decoded residual of the primary color channel to generate a secondary transform coefficient; and decoding the coding unit by applying a separable transform kernel to the secondary transform coefficient and applying the separable transform kernel to the decoded residual of at least one secondary color channel.
[0023] Another aspect of the present disclosure provides a non-transitory computer-readable medium having a computer program stored thereon to implement a method for decoding a coding unit of a coding tree from a coding tree unit of an image frame from a video bitstream, the coding unit having a primary color channel and at least one secondary color channel, the method comprising: determining a coding unit including a primary color channel and at least one secondary color channel based on a decoded split flag of the coding tree unit; decoding a first index to select a core for the primary color channel, and decoding a second index to select a core for the at least one secondary color channel; selecting a first core based on the first index, and selecting a second core based on the second index; and decoding the coding unit by applying the first core to a residual coefficient of the primary color channel and applying the second core to a residual coefficient of at least one secondary color channel.
[0024] Another aspect of the present disclosure provides a video decoder configured to implement a method for decoding a coding unit of a coding tree from a coding tree unit of an image frame from a video bitstream, the coding unit having a primary color channel and at least one secondary color channel, the method comprising: determining a coding unit including a primary color channel and at least one secondary color channel based on a decoded split flag of the coding tree unit; decoding a first index to select a core for the primary color channel, and decoding a second index to select a core for the at least one secondary color channel; selecting a first core based on the first index, and selecting a second core based on the second index; and decoding the coding unit by applying the first core to a residual coefficient of the primary color channel and applying the second core to a residual coefficient of at least one secondary color channel.
[0025] Another aspect of the present disclosure provides a system comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory to implement a method for decoding a coding unit of a coding tree from a coding tree unit of an image frame from a video bitstream, the coding unit having a primary color channel and at least one secondary color channel, the method comprising: determining a coding unit including the primary color channel and the at least one secondary color channel based on a decoded split flag of the coding tree unit; decoding a first index to select a core for the primary color channel, and decoding a second index to select a core for the at least one secondary color channel; selecting a first core based on the first index, and selecting a second core based on the second index; and decoding the coding unit by applying the first core to a residual coefficient of the primary color channel and applying the second core to a residual coefficient of at least one secondary color channel.
[0026] Another aspect of the present disclosure provides a method for decoding multiple coding units from a bitstream to generate an image frame, wherein the coding units are the result of decomposition of a coding tree unit, and the multiple coding units form one or more continuous parts of the bitstream, the method comprising: determining a subdivision level for each of the one or more continuous parts of the bitstream, each subdivision level being applicable to the coding units of the corresponding continuous part of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of the coding tree unit into the coding units of the respective continuous parts of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding multiple coding units using the determined quantization parameter for each region to generate an image frame.
[0027] According to another aspect, the respective regions are based on a comparison of a subdivision level associated with the coding unit and a determined subdivision level of the respective continuous portion.
[0028] According to another aspect, a quantization parameter increment is determined for each region, wherein the corresponding coding tree has a subdivision level that is smaller than or equal to the determined subdivision level of the corresponding contiguous portion.
[0029] According to another aspect, a new region is set for any node in a coding tree unit having a subdivision level less than or equal to the corresponding determined subdivision level.
[0030] According to another aspect, the subdivision levels determined for the respective contiguous portions include a first subdivision level for luma coding units of the contiguous portions and a second subdivision level for chroma coding units of the contiguous portions.
[0031] According to another aspect, the first subdivision level and the second subdivision level are different.
[0032] According to another aspect, the method further comprises decoding a flag indicating that a partition constraint of a sequence parameter set associated with the bitstream may be overwritten.
[0033] According to another aspect, the determined subdivision level for each of the one or more continuous portions comprises a maximum luma coding unit depth for the region.
[0034] According to another aspect, the determined subdivision level for each of the one or more continuous portions comprises a maximum chroma coding unit depth for the corresponding region.
[0035] According to another aspect, the determined subdivision level for one of the consecutive portions is adjusted to maintain an offset relative to a deepest allowed subdivision level decoded for a partition constraint of the bitstream.
[0036] Another aspect of the present disclosure provides a non-transitory computer-readable medium having a computer program stored thereon to implement a method for decoding a plurality of coding units from a bitstream to generate an image frame, the coding units being the result of a decomposition of a coding tree unit, the plurality of coding units forming one or more continuous parts of the bitstream, the method comprising: determining a subdivision level for each of the one or more continuous parts of the bitstream, each subdivision level being applicable to the coding units of the corresponding continuous parts of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of the coding tree unit into the coding units of the respective continuous parts of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to generate an image frame.
[0037] Another aspect of the present disclosure provides a video decoder configured to implement a method for decoding a plurality of coding units from a bitstream to generate an image frame, wherein the coding units are the result of a decomposition of a coding tree unit, and the plurality of coding units form one or more continuous parts of the bitstream, the method comprising: determining a subdivision level for each of the one or more continuous parts of the bitstream, each subdivision level being applicable to the coding units of the corresponding continuous parts of the bitstream; decoding a quantization parameter increment for each of a plurality of regions, each region being based on the decomposition of the coding tree unit into the coding units of the respective continuous parts of the bitstream and the corresponding determined subdivision level; determining a quantization parameter for each region based on the decoded incremental quantization parameter of the region and the quantization parameter of an earlier coding unit of the image frame; and decoding the plurality of coding units using the determined quantization parameter for each region to generate an image frame.
[0038] Another aspect of the present disclosure provides a system, comprising: a memory; and a processor, wherein the processor is configured to execute code stored on the memory to implement a method for decoding multiple coding units from a bitstream to generate an image frame, wherein the coding units are decomposed results of coding tree units, and the multiple coding units form one or more continuous parts of the bitstream, the method comprising: determining a subdivision level for each of the one or more continuous parts of the bitstream, each subdivision level being applicable to the coding units of the corresponding continuous parts of the bitstream; decoding quantization parameter increments for each of a plurality of regions, each region being based on decomposing the coding tree units into coding units of the respective continuous parts of the bitstream and the corresponding determined subdivision levels; determining a quantization parameter for each region based on a decoded incremental quantization parameter of the region and a quantization parameter of an earlier coding unit of the image frame; and decoding the multiple coding units using the determined quantization parameters for each region to generate an image frame.
[0039] Other aspects are also disclosed. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] At least one embodiment of the present invention will now be described with reference to the following drawings and appendices, in which:
[0041] Figure 1 is a schematic block diagram illustrating a video encoding and decoding system;
[0042] Figure 2A and 2B Composition can be practiced Figure 1 A schematic block diagram of a general computer system for one or both of the video encoding and decoding systems;
[0043] Figure 3 is a schematic block diagram showing the functional modules of a video encoder;
[0044] Figure 4 is a schematic block diagram showing the functional modules of a video decoder;
[0045] Figure 5 is a schematic block diagram illustrating possible partitioning of a block into one or more blocks in a tree structure for general video coding;
[0046] Figure 6 A schematic diagram of a data flow for implementing a permissible partitioning of a block into one or more blocks in a tree structure of a general video coding;
[0047] Fig. 7A and 7B shows an example partitioning of a coding tree unit (CTU) into multiple coding units (CUs);
[0048] Fig. 8A ,8B and 8C show the subdivision levels resulting from the split in the coding tree and their impact on the partitioning of the coding tree units into quantization groups;
[0049] Fig.9A and 9B shows a 4×4 transform block scan pattern and associated primary and secondary transform coefficients;
[0050] Fig. 9C and 9D shows an 8×8 transform block scan pattern and associated primary and secondary transform coefficients;
[0051] Fig.10 shows the region where the secondary transform is applied for transform blocks of various sizes;
[0052] Fig.11 shows a syntax structure of a bitstream having multiple slices, each slice comprising multiple coding units;
[0053] Fig.12 shows the syntax structure of a bitstream with a common tree for luma and chroma coding blocks of a coding tree unit;
[0054] Fig.13 A method of encoding a frame in a bitstream comprising one or more than one slice as a sequence of coding units is shown;
[0055] Fig.14 A method of encoding a slice header in a bitstream is shown;
[0056] Fig.15 A method of encoding a coding unit in a bitstream is shown;
[0057] Fig.16 A method of decoding a frame from a bitstream as a sequence of coding units arranged into slices is shown;
[0058] Fig.17 A method of decoding a slice header from a bitstream is shown;
[0059] Fig.18 A method of decoding a coding unit from a bitstream is shown; and
[0060] Fig.19A and 19B The rules for applying or bypassing the secondary transform to the luma and chroma channels are shown. DETAILED DESCRIPTION
[0061] Where steps and / or features having the same reference numerals are referenced in any one or more of the accompanying drawings, these steps and / or features have the same function(s) or operation(s) for the purposes of this specification, unless otherwise intended.
[0062] A rate controlled video encoder requires the flexibility to adjust the quantization parameter at a granularity appropriate to the block partition constraints. The block partition constraints may differ between one part of a frame and another, for example, where multiple video encoders operate in parallel to compress individual frames. The granularity of the areas where quantization parameter adjustment is required varies accordingly. In addition, control over the selection of the transform applied (including the potential application of a secondary transform) is applied within the context of generating a prediction signal for the residual being transformed. In particular, for intra prediction, separate modes may be used for luma blocks and chroma blocks, since they may use different intra prediction modes.
[0063] Some parts of the video contribute more to the fidelity of the rendered viewport than others and can be allocated a larger bitrate and more flexibility in the variance of block structures and quantization parameters. Portions that contribute little to the fidelity of the rendered viewport (such as those to the sides or behind the rendered view) can be compressed with a simpler block structure to reduce the encoding workload and have less flexibility in the control of the quantization parameters. Typically, larger values are selected to more coarsely quantize the transform coefficients of the lower bitrate. In addition, the application of transform selection can be independent between the luma channel and the chroma channels to further simplify the encoding process by avoiding the need to jointly consider luma and chroma for transform selection. In particular, the need to jointly consider luma and chroma for secondary transform selection is avoided after considering the intra-frame prediction modes of luma and chroma separately.
[0064] Figure 1 is a schematic block diagram illustrating functional modules of the video encoding and decoding system 100. The system 100 may vary the region in which the quantization parameter is adjusted in different parts of a frame to accommodate different block partitioning constraints that may be valid in various parts of the frame.
[0065] The system 100 includes a source device 110 and a destination device 130. A communication channel 120 is used to communicate encoded video information from the source device 110 to the destination device 130. In some configurations, one or both of the source device 110 and the destination device 130 may each include a mobile phone handset or "smart phone", in which case the communication channel 120 is a wireless channel. In other configurations, the source device 110 and the destination device 130 may include video conferencing equipment, in which case the communication channel 120 is typically a wired channel such as an Internet connection. In addition, the source device 110 and the destination device 130 may include any of a wide range of devices, including devices that support over-the-air television broadcasts, cable television applications, Internet video applications (including streaming), and applications that capture encoded video data on some computer-readable storage medium (such as a hard drive in a file server).
[0066] like Figure 1 As shown, source device 110 includes video source 112, video encoder 114, and transmitter 116. Video source 112 typically includes a source of captured video frame data (denoted as 113), such as a camera sensor, a previously captured video sequence stored on a non-transitory recording medium, or a video feed from a remote camera sensor. Video source 112 may also be the output of a computer graphics card (e.g., display operating system and video output of various applications executed on a computing device (e.g., tablet computer)). Examples of source device 110 that may include a camera sensor as video source 112 include smart phones, video camcorders, professional video cameras, and web video cameras.
[0067] The video encoder 114 converts (or "encodes") the captured frame data (indicated by arrow 113) from the video source 112 into a bitstream (indicated by arrow 115). The bitstream 115 is transmitted by a transmitter 116 as encoded video data (or "encoded video information") via a communication channel 120. The bitstream 115 may also be stored in a non-transitory storage device 122, such as a "flash" memory or a hard drive, until or in lieu of subsequent transmission via the communication channel 120. For example, the encoded video data may be supplied to a customer via a wide area network (WAN) for video streaming applications when desired.
[0068] The destination device 130 includes a receiver 132, a video decoder 134, and a display device 136. The receiver 132 receives the encoded video data from the communication channel 120 and passes the received video data to the video decoder 134 as a bit stream (indicated by arrow 133). The video decoder 134 then outputs the decoded frame data (indicated by arrow 135) to the display device 136. The decoded frame data 135 has the same chroma format as the frame data 113. Examples of the display device 136 include a cathode ray tube, a liquid crystal display (such as in a smart phone, a tablet computer, a computer monitor, or a stand-alone television). The functions of the source device 110 and the destination device 130 can also be embodied in a single device, examples of which include a mobile phone handset and a tablet computer. The decoded frame data can be further transformed before being presented to the user. For example, a "viewport" with a specific latitude and longitude can be rendered from the decoded frame data using a projection format to represent a 360° view of the scene.
[0069] Although example devices are described above, source device 110 and destination device 130 may each typically be configured within a general purpose computer system via a combination of hardware and software components. Figure 2ASuch a computer system 200 is shown, which includes: a computer module 201; input devices such as a keyboard 202, a mouse pointer device 203, a scanner 226, a camera 227 that can be configured as a video source 112, and a microphone 280; and output devices including a printer 215, a display device 214 that can be configured as a display device 136, and a speaker 217. The computer module 201 can use an external modulator-demodulator (modem) transceiver device 216 to communicate with a communication network 220 via a connection 221. The communication network 220, which can represent a communication channel 120, can be a WAN, such as the Internet, a cellular telecommunications network, or a private WAN. In the case where the connection 221 is a telephone line, the modem 216 can be a traditional "dial-up" modem. Alternatively, in the case where the connection 221 is a high-capacity (e.g., cable or optical) connection, the modem 216 can be a broadband modem. Wireless modems can also be used to make wireless connections to the communication network 220. The transceiver device 216 may provide the functionality of the transmitter 116 and the receiver 132 , and the communication channel 120 may be embodied in the wiring 221 .
[0070] The computer module 201 typically includes at least one processor unit 205 and a memory unit 206. For example, the memory unit 206 may have a semiconductor random access memory (RAM) and a semiconductor read-only memory (ROM). The computer module 201 also includes a plurality of input / output (I / O) interfaces, wherein the plurality of input / output (I / O) interfaces include: an audio-video interface 207 connected to a video display 214, a speaker 217, and a microphone 280; an I / O interface 213 connected to a keyboard 202, a mouse 203, a scanner 226, a camera 227, and an optional joystick or other human-machine interface device (not shown); and an interface 208 for an external modem 216 and a printer 215. The signal from the audio-video interface 207 to the computer monitor 214 is typically the output of a computer graphics card. In some implementations, the modem 216 may be built into the computer module 201, such as built into the interface 208. The computer module 201 also has a local network interface 211, which allows the computer system 200 to be connected to a local area communication network 222, known as a local area network (LAN), via a connection 223. Figure 2A As shown, the local area communication network 222 can also be connected to the wide area network 220 via the connection 224, wherein the local area communication network 222 usually includes a so-called "firewall" device or a device with similar functions. The local network interface 211 may include an Ethernet (Ethernet TM )Circuit card, Bluetooth TM) wireless configuration or IEEE 802.11 wireless configuration; however, for the interface 211, a variety of other types of interfaces can be implemented. The local network interface 211 can also provide the functions of the transmitter 116 and the receiver 132, and the communication channel 120 can also be embodied in the local communication network 222.
[0071] I / O interfaces 208 and 213 may provide either or both serial and parallel connections, the former of which is typically implemented in accordance with the Universal Serial Bus (USB) standard and having a corresponding USB connector (not shown). A storage device 209 is provided, and the storage device 209 typically includes a hard disk drive (HDD) 210. Other storage devices (not shown) such as floppy disk drives and tape drives may also be used. An optical disk drive 212 is typically provided to serve as a non-volatile source of data. For example, an optical disk (e.g., CD-ROM, DVD, Blu-ray Disc) may be used. TM )), portable memory devices such as USB-RAM, portable external hard drives and floppy disks, etc., as suitable sources of data for computer system 200. Generally, any of HDD 210, optical drive 212, networks 220 and 222 may also be configured to operate as video source 112, or as a destination for decoded video data to be stored for reproduction via display 214. Source device 110 and destination device 130 of system 100 may be embodied in computer system 200.
[0072] The components 205-213 of the computer module 201 typically communicate via an interconnect bus 204 and in a manner that results in conventional operating modes of the computer system 200 known to those skilled in the relevant art. For example, the processor 205 is connected to the system bus 204 using a connection 218. Similarly, the memory 206 and the optical drive 212 are connected to the system bus 204 via a connection 219. Examples of computers that can practice the described configuration include IBM-PC and compatible machines, Sun SPARCstation, Apple Mac TM or similar computer system.
[0073] Where appropriate or desired, the video encoder 114 and the video decoder 134 and the methods described below may be implemented using the computer system 200. In particular, the video encoder 114, the video decoder 134 and the methods to be described may be implemented as one or more software applications 233 executable within the computer system 200. In particular, instructions 231 (see instructions 231 and 233) executed within the computer system 200 may be used to implement the video encoder 114 and the video decoder 134 and the methods to be described. Figure 2B) to implement the video encoder 114, the video decoder 134 and the steps of the method. The software instructions 231 may be formed into one or more code modules, each for performing one or more specific tasks. The software may also be split into two separate parts, with a first part and corresponding code modules performing the method, and a second part and corresponding code modules managing a user interface between the first part and a user.
[0074] For example, the software may be stored in a computer readable medium including a storage device as described below. The software is loaded from the computer readable medium into the computer system 200 and then executed by the computer system 200. A computer readable medium with such software or a computer program recorded on the computer readable medium is a computer program product. The use of the computer program product in the computer system 200 preferably implements an advantageous device for implementing the video encoder 114, the video decoder 134 and the method.
[0075] The software 233 is typically stored in the HDD 210 or the memory 206. The software is loaded into the computer system 200 from a computer readable medium and executed by the computer system 200. Thus, for example, the software 233 may be stored on an optically readable disk storage medium (e.g., a CD-ROM) 225 read by the optical drive 212.
[0076] In some examples, the application 233 is provided to the user in a manner encoded on one or more CD-ROMs 225 and read via corresponding drives 212, or alternatively, the application 233 can be read by the user from the network 220 or 222. Furthermore, the software can also be loaded into the computer system 200 from other computer-readable media. Computer-readable storage media refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to the computer system 200 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray Discs, and the like. TM ), a hard disk drive, a ROM or integrated circuit, a USB memory, a magneto-optical disk, or a computer readable card such as a PCMCIA card, etc., regardless of whether these devices are internal or external to the computer module 201. Examples of temporary or non-tangible computer-readable transmission media that can also participate in providing software, applications, instructions and / or video data or encoded video data to the computer module 401 include: radio or infrared transmission channels and network connections to other computers or networked devices, and the Internet or intranet including email transmission and information recorded on websites.
[0077] The second part of the above-mentioned application program 233 and the corresponding code module can be executed to implement one or more graphical user interfaces (GUIs) to be drawn or otherwise presented on the display 214. Users and applications of the computer system 200 can operate the interface in a functionally applicable manner by typically operating the keyboard 202 and the mouse 203 to provide control commands and / or input to the applications associated with these (one or more) GUIs. Other forms of user interfaces that are functionally applicable can also be implemented, such as audio interfaces that utilize voice prompts output via the speaker 217 and user voice commands input via the microphone 280, etc.
[0078] Figure 2B is a detailed schematic block diagram of the processor 205 and the "memory" 234. The memory 234 represents Figure 2A A logical aggregation of all memory modules (including HDD 209 and semiconductor memory 206) that can be accessed by computer module 201 in.
[0079] When the computer module 201 is initially powered on, a power-on self-test (POST) program 250 is executed. The POST program 250 is usually stored in Figure 2A 249 of semiconductor memory 206. Hardware devices such as ROM 249 storing software are sometimes referred to as firmware. POST program 250 checks the hardware within computer module 201 to ensure proper operation, and typically checks processor 205, memory 234 (209, 206), and basic input-output system software (BIOS) module 251, which is also typically stored in ROM 249, for correct operation. Once POST program 250 runs successfully, BIOS 251 starts Figure 2A The hard disk drive 210 is started so that the boot loader 252 residing on the hard disk drive 210 is executed via the processor 205. This loads the operating system 253 into the RAM memory 206, where the operating system 253 starts to work on the RAM memory 206. The operating system 253 is a system-level application executable by the processor 205 to implement various high-level functions including processor management, memory management, device management, storage management, software application interface, and general user interface.
[0080] The operating system 253 manages the memory 234 (209, 206) to ensure that each process or application running on the computer module 201 has sufficient memory to execute without conflicting with memory allocated to other processes. In addition, the appropriate use of Figure 2AThe different types of memory available in the computer system 200 are described so that each process can run efficiently. Therefore, the aggregate memory 234 is not intended to illustrate how to allocate specific segments of memory (unless otherwise specified), but rather to provide an overview of the memory accessible to the computer system 200 and how to use it.
[0081] like Figure 2B As shown, the processor 205 includes a plurality of functional modules, wherein the plurality of functional modules include a control unit 239, an arithmetic logic unit (ALU) 240, and a local or internal memory 248 sometimes referred to as a cache memory. The cache memory 248 typically includes a plurality of storage registers 244-246 in a register section. One or more internal buses 241 functionally interconnect these functional modules. The processor 205 also typically has one or more interfaces 242 for communicating with external devices via the system bus 204 using a connection 218. The memory 234 is connected to the bus 204 using a connection 219.
[0082] The application program 233 includes an instruction sequence 231 that may include conditional branch instructions and loop instructions. The program 233 may also include data 232 used when executing the program 233. The instructions 231 and data 232 are stored in memory locations 228, 229, 230 and 235, 236, 237, respectively. Depending on the relative size of the instructions 231 and the memory locations 228-230, a particular instruction may be stored in a single memory location, as described by the instruction shown in memory location 230. Alternatively, the instruction may be split into multiple parts, each stored in a separate memory location, as described by the instruction segments shown in memory locations 228 and 229.
[0083] Typically, a processor 205 is given a set of instructions, which are executed within the processor 205. The processor 205 awaits a subsequent input, which the processor 205 reacts to by executing another set of instructions. Each input may be provided from one or more of a plurality of sources, including data generated by one or more of the input devices 202, 203, data received from an external source via one of the networks 220, 202, data retrieved from one of the storage devices 206, 209, or data retrieved from a storage medium 225 inserted into a corresponding reader 212 (all of which are described in detail in the accompanying drawings). Figure 2A Execution of a set of instructions may result in outputting data in some cases. Execution may also involve storing data or variables to memory 234.
[0084] The video encoder 114, the video decoder 134, and the method may use input variables 254 stored in corresponding memory locations 255, 256, 257 within the memory 234. The video encoder 114, the video decoder 134, and the method produce output variables 261 stored in corresponding memory locations 262, 263, 264 within the memory 234. Intermediate variables 258 may be stored in memory locations 259, 260, 266, and 267.
[0085] refer to Figure 2B The processor 205, registers 244, 245, 246, arithmetic logic unit (ALU) 240 and control unit 239 work together to perform micro-operation sequences required to perform a "fetch, decode and execute" cycle for each instruction in the instruction set that constitutes the program 233. Each fetch, decode and execute cycle includes:
[0086] A fetch operation for fetching or reading an instruction 231 from a memory location 228 , 229 , 230 ;
[0087] a decode operation in which the control unit 239 determines which instruction was fetched; and
[0088] An execution operation, wherein in the execution operation, the control unit 239 and / or the ALU 240 executes the instruction.
[0089] Thereafter, further fetch, decode and execute cycles for the next instruction may be performed. Likewise, a store cycle may be performed, whereby the control unit 239 stores or writes a value to the memory location 232 .
[0090] To be explained Figures 13 to 18 Each step or sub-process in the method is associated with one or more segments of program 233, and is typically performed by the register units 244, 245, 247, ALU 240 and control unit 239 in the processor 205 working together to perform fetch, decode and execute cycles for each instruction in the instruction set of the segment of program 233.
[0091] Figure 3 is a schematic block diagram illustrating the functional modules of the video encoder 114 . Figure 4 1 is a schematic block diagram showing the functional modules of the video decoder 134. Typically, data is transferred between the functional modules within the video encoder 114 and the video decoder 134 in groups of samples or coefficients (such as partitioning of a block into fixed-size sub-blocks, etc.) or as an array. Figure 2A and 2BAs shown, the video encoder 114 and the video decoder 134 may be implemented using a general purpose computer system 200, wherein various functional modules may be implemented using dedicated hardware within the computer system 200, using software executable within the computer system 200 (such as one or more software code modules of a software application 233 residing on a hard drive 205 and controlled by a processor 205 for execution, etc.). Alternatively, the video encoder 114 and the video decoder 134 may be implemented using a combination of dedicated hardware and software executable within the computer system 200. The video encoder 114, the video decoder 134 and the method may alternatively be implemented in dedicated hardware such as one or more integrated circuits that perform the functions or sub-functions of the method. Such dedicated hardware may include a graphics processing unit (GPU), a digital signal processor (DSP), an application specific standard product (ASSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or one or more microprocessors and associated memory. In particular, the video encoder 114 includes modules 310 - 390 , and the video decoder 134 includes modules 420 - 496 , where each of these modules may be implemented as one or more software code modules of the software application 233 .
[0092] although Figure 3 The video encoder 114 is an example of a Versatile Video Coding (VVC) video encoding pipeline, but other video codecs may also be used to perform the processing stages described herein. The video encoder 114 receives captured frame data 113, such as a series of frames (each frame including one or more color channels). The frame data 113 can be in any chroma format, such as a 4:0:0, 4:2:0, 4:2:2, or 4:4:4 chroma format. The block partitioner 310 first partitions the frame data 113 into CTUs, which are generally square in shape and are configured so that a specific size of the CTU is used. For example, the size of the CTU can be 64×64, 128×128, or 256×256 luma samples. The block partitioner 310 further partitions each CTU into one or more CBs corresponding to a luma coding tree or a chroma coding tree. The luma channel may also be referred to as a primary color channel. The individual chroma channels may also be referred to as secondary color channels. CBs have various sizes and may include both square and non-square aspect ratios. Reference Figure 13-15 The operation of the block partitioner 310 is further described. However, in the VVC standard, CBs, CUs, PUs, and TUs always have side lengths that are powers of 2. Thus, the current CB (denoted as 312) is output from the block partitioner 310, thereby advancing according to the iteration of one or more blocks of the CTU, according to the luma coding tree and the chroma coding tree of the CTU. Figure 5 and 61 to further illustrate the option for partitioning a CTU into CBs. Although operations are generally described in units of CTUs, the video encoder 114 and the video decoder 134 may operate on smaller sized regions to reduce memory consumption. For example, each CTU may be partitioned into smaller regions, referred to as "virtual pipeline data units" (VPDUs) of size 64×64. The VPDUs form a data granularity that is more suitable for pipeline processing in a hardware architecture, where the reduction in memory footprint reduces silicon area and therefore reduces cost compared to operations on a full CTU.
[0093] The CTUs resulting from the first partitioning of the frame data 113 may be scanned in raster scan order and may be grouped into one or more "slices". A slice may be an "intra" (or "I") slice. An intra slice (I slice) indicates that each CU in the slice is intra predicted. Optionally, a slice may be uni-predicted or bi-predicted ("P" or "B" slices, respectively), indicating the additional availability of uni-prediction and bi-prediction in the slice, respectively.
[0094] In an I slice, the coding tree for each CTU may diverge below the 64×64 level into two separate coding trees, one for luma and the other for chroma. Using separate trees allows different block structures to exist between luma and chroma within the luma 64×64 region of the CTU. For example, a large chroma CB may be co-located with many smaller luma CBs, and vice versa. In a P or B slice, a single coding tree for a CTU defines a block structure common to luma and chroma. The resulting blocks of a single tree may be intra-predicted or inter-predicted.
[0095] For each CTU, the video encoder 114 operates in two stages. In the first stage (referred to as the "search" stage), the block partitioner 310 tests various potential configurations of the coding tree. Each potential configuration of the coding tree has an associated "candidate" CB. The first stage involves testing various candidate CBs to select a CB that provides relatively high compression efficiency and relatively low distortion. This test typically involves Lagrangian optimization, whereby the candidate CBs are evaluated based on a weighted combination of rate (coding cost) and distortion (error with respect to the input frame data 113). The "best" candidate CB (the CB with the lowest evaluated rate / distortion) is selected for subsequent encoding in the bitstream 115. The evaluation of the candidate CBs includes the following options: use the CB for a given region, or split the region according to various splitting options and use other CBs to encode each smaller resulting region or further split the region. As a result, both the coding tree and the CB itself are selected in the search stage.
[0096] The video encoder 114 generates a prediction block (PB) indicated by arrow 320 for each CB (e.g., CB 312). PB 320 is a prediction of the content of the associated CB 312. Subtractor module 322 generates a difference (or "residual," which means that the difference is in the spatial domain) between PB 320 and CB 312, represented as 324. Difference 324 is the block size difference between corresponding samples in PB 320 and CB 312. Difference 324 is transformed, quantized, and represented as a transform block (TB) indicated by arrow 336. PB 320 and associated TB 336 are typically selected from one of a plurality of possible candidate CBs (e.g., based on an evaluated cost or distortion).
[0097] A candidate coding block (CB) is a CB derived from one of the prediction modes available to the video encoder 114 for the associated PB and the resulting residual. When combined with the predicted PB in the video decoder 114, the TB 336 reduces the difference between the decoded CB and the original CB 312 at the expense of additional signaling in the bitstream.
[0098] Thus, each candidate coding block (CB), i.e., a combination of a prediction block (PB) and a transform block (TB), has an associated coding cost (or "rate") and an associated difference (or "distortion"). The distortion of a CB is typically estimated as a difference in sample values, such as a sum of absolute differences (SAD) or a sum of squared differences (SSD), etc. The mode selector 386 can use the difference 324 to determine the estimates obtained from each candidate PB to determine a prediction mode 387. The prediction mode 387 indicates a decision to use a particular prediction mode (e.g., intra-frame prediction or inter-frame prediction) for the current CB. The estimation of the coding cost associated with each candidate prediction mode and the corresponding residual encoding can be performed at a significantly lower cost than entropy encoding of the residual. Therefore, even in a real-time video encoder, multiple candidate modes can be evaluated to determine the best mode in terms of rate-distortion.
[0099] Determining the best model in terms of rate-distortion is usually achieved using a variation of Lagrangian optimization.
[0100] A Lagrangian or similar optimization process may be employed for both the selection of the best partitioning of a CTU into a CB (using the block partitioner 310) and the selection of the best prediction mode from multiple possibilities. The intra-prediction mode with the lowest cost measure is selected as the best mode by applying the Lagrangian optimization process of the candidate modes in the mode selector module 386. The lowest cost mode is the selected secondary transform index 388 and is also encoded in the bitstream 115 by the entropy encoder 338.
[0101] In the second stage of the operation of the video encoder 114, referred to as the "encoding" stage, iterations of the determined coding tree(s) for each CTU are performed in the video encoder 114. For CTUs using separate trees, the luma coding tree is first encoded, followed by the chroma coding tree, for each 64x64 luma region of the CTU. Only luma CBs are encoded within the luma coding tree, and only chroma CBs are encoded within the chroma coding tree. For CTUs using a common tree, a single tree describes the CU, i.e., luma CBs and chroma CBs, according to the common block structure of the common tree.
[0102] The entropy encoder 338 supports both variable length encoding of syntax elements and arithmetic encoding of syntax elements. Portions of the bitstream such as "parameter sets" (e.g., sequence parameter sets (SPS) and picture parameter sets (PPS)) use a combination of fixed length codewords and variable length codewords. A slice (also referred to as a continuous portion) has a slice header using variable length encoding, followed by slice data using arithmetic encoding. The slice header defines parameters specific to the current slice, such as slice-level quantization parameter offsets, etc. The slice data includes syntax elements for each CTU in the slice. The use of variable length encoding and arithmetic coding requires sequential parsing within the various portions of the bitstream. These portions can be described with start codes to form "network abstraction layer units" or "NAL units." Arithmetic coding is supported using a context-adaptive binary arithmetic coding process. Arithmetic coded syntax elements consist of a sequence of one or more "bins". Like bits, bin values are "0" or "1". However, bins are not encoded as discrete bits in the bitstream 115. A bin has an associated predicted (or "likely" or "maximum probability") value and an associated probability (called a "context"). When the actual bin to be encoded matches the predicted value, a "maximum probability symbol" (MPS) is encoded. Encoding the maximum probability symbol is relatively cheap in terms of consumed bits in the bitstream 115 (including a total cost of less than one discrete bit). When the actual bin to be encoded does not match the possible value, a "minimum probability symbol" (LPS) is encoded. Encoding the minimum probability symbol has a relatively high cost in terms of consumed bits. Bin encoding techniques enable efficient encoding of bins with skewed probabilities of "0" vs "1". For syntactic elements with two possible values (i.e., "flag"), a single bin is sufficient. For syntactic elements with many possible values, a sequence of bins is required.
[0103] The presence of a later bin in the sequence may be determined based on the value of an earlier bin in the sequence. In addition, each bin may be associated with more than one context. A particular context may be selected based on, for example, an earlier bin in a syntax element and the bin values of adjacent syntax elements (i.e., bin values from adjacent blocks). Each time a context coding bin is encoded, the context selected for that bin (if any) is updated to reflect the new bin value. In this way, the binary arithmetic coding scheme is considered to be adaptive.
[0104] The video encoder 114 also supports bins lacking context ("bypass bins"). Bypass bins are encoded assuming an equal probability distribution between "0" and "1". Thus, each bin has an encoding cost of one bit in the bitstream 115. The absence of context saves memory and reduces complexity, so bypass bins where the distribution of values for a particular bin is not skewed are used. An example of an entropy encoder that employs context and adaptation is known in the art as CABAC (Context Adaptive Binary Arithmetic Coder), and many variations of this encoder are employed in video encoding.
[0105] The entropy encoder 338 encodes a quantization parameter 392 using a combination of context-encoded bins and bypass-encoded bins, and, if used for the current CB, an LFNST index 388. The quantization parameter 392 is encoded using a "delta QP". The delta QP is signaled at most once in each region called a "quantization group". The quantization parameter 392 is applied to the residual coefficients of the luma CB. The adjusted quantization parameter is applied to the residual coefficients of the juxtaposed chroma CB. The adjusted quantization parameter may include mapping from the luma quantization parameter 392 according to a mapping table and a CU-level offset selected from an offset list. The secondary transform index 388 is signaled when the residual associated with the transform block includes valid residual coefficients only in those coefficient positions that are transformed into primary coefficients by applying a secondary transform.
[0106] The multiplexer module 384 outputs the PB 320 from the intra prediction module 364 according to the determined best intra prediction mode selected from the test prediction modes of each candidate CB. The candidate prediction modes do not need to include every conceivable prediction mode supported by the video encoder 114. Intra prediction is divided into three types. "DC intra prediction" involves filling the PB with a single value representing the average of nearby reconstructed samples. "Plane intra prediction" involves filling the PB with samples according to a plane, where the DC offset and vertical and horizontal gradients are derived from nearby reconstructed neighboring samples. Nearby reconstructed samples typically include a row of reconstructed samples above the current PB (extending to a certain extent to the right of the PB) and a column of reconstructed samples on the left side of the current PB (extending to a certain extent downward outside the PB). "Angle intra prediction" involves filling the PB with reconstructed neighboring samples, which are filtered and propagated across the PB in a specific direction (or "angle"). In VVC, 65 angles are supported, where rectangular blocks can use additional angles that are not available to square blocks to produce a total of 87 angles. A fourth type of intra prediction can be used for chroma PBs, generating PBs from collocated luma reconstruction samples according to a "cross component linear model" (CCLM) mode. Three different CCLM modes are available, each using a different model derived from adjacent luma and chroma samples. The derived model is used to generate a block of samples for a chroma PB from collocated luma samples.
[0107] In cases where previously reconstructed samples are not available (e.g., at the edge of a frame), a default halftone value of half the sample range is used. For example, for 10-bit video, a value of 512 is used. Since no previous samples are available for a CB located at the top left position of the frame, the angular and planar intra prediction modes produce the same output as the DC prediction mode, i.e., a flat plane of samples with halftone values as amplitudes.
[0108] For inter-frame prediction, a prediction block 382 is generated by a motion compensation module 380 using samples from one or two frames preceding the current frame in the order of coded frames in the bitstream and output as a PB 320 by a multiplexer module 384. In addition, for inter-frame prediction, a single coding tree is typically used for both the luminance channel and the chrominance channel. The order of coded frames in the bitstream may be different from the order of frames when captured or displayed. When one frame is used for prediction, the block is called "single prediction" and has two associated motion vectors. When two frames are used for prediction, the block is called "double prediction" and has two associated motion vectors. For P slices, each CU can be intra-predicted or single-predicted. For B slices, each CU can be intra-predicted, single-predicted, or double-predicted. Frames are typically encoded using a "picture group" structure, thereby achieving a temporal hierarchical structure of frames. Frames can be divided into multiple slices, each encoding a portion of a frame. The temporal hierarchical structure of frames allows frames to reference previous and subsequent pictures in the order in which the frames are displayed. Images are encoded in the order necessary to ensure that the dependencies for decoding each frame are met.
[0109] Samples are selected based on the motion vector 378 and the reference picture index. The motion vector 378 and the reference picture index apply to all color channels, so inter prediction is described primarily in terms of operations on PUs rather than PBs, i.e., a single coding tree is used to describe the decomposition of each CTU into one or more inter prediction blocks. Inter prediction methods may vary in the number of motion parameters and their precision. The motion parameters typically include a reference frame index (which indicates which reference frames from a reference frame list will be used plus their respective spatial translations), but may include more frames, special frames, or complex affine parameters such as scaling and rotation. In addition, a predetermined motion refinement process may be applied to generate a dense motion estimate based on a block of reference samples.
[0110] When the PB 320 is determined and selected and subtracted from the original sample block at the subtractor 322, a residual with the lowest coding cost (denoted as 324) is obtained and lossily compressed. The lossy compression process includes the steps of transform, quantization and entropy coding. The forward main transform module 326 applies a forward transform to the difference 324, thereby converting the difference 324 from the spatial threshold to the frequency domain and producing the main transform coefficients represented by arrow 328. The maximum main transform size in one dimension is a 32-point DCT-2 or a 64-point DCT-2 transform. If the CB being encoded is larger than the maximum supported main transform size (i.e., 64×64 or 32×32) denoted as a block size, the main transform 326 is applied in a block manner to transform all samples of the difference 324. The application of the transform 326 results in multiple TBs of the CB. In the case where the respective transform applications operate on a difference 324 TB that is larger than 32×32 (e.g., 64×64), all resulting main transform coefficients 328 outside the upper left 32×32 region of the TB are set to zero, i.e., discarded. The remaining main transform coefficients 328 are passed to a quantizer module 334. The main transform coefficients 328 are quantized according to quantization parameters 392 associated with the CBs to produce main transform coefficients 332. The quantization parameters 392 may be different for the luma CB relative to the respective chroma CBs. The main transform coefficients 332 are passed to a forward secondary transform module 330 to produce transform coefficients represented by arrow 336, either by performing a non-separable secondary transform (NSST) operation or by bypassing the secondary transform. The forward main transform is typically separable, transforming a set of rows of each TB and then a set of columns. For luma TBs with a width and height not exceeding 16 samples, the forward main transform module 326 uses a type II discrete cosine transform (DCT-2) in the horizontal and vertical directions, or bypasses the transform in the horizontal and vertical directions, or uses a combination of a type VII discrete sine transform (DST-7) and a type VIII discrete cosine transform (DCT-8) in the horizontal or vertical directions. The use of a combination of DST-7 and DCT-8 is referred to as a "multi-transform selection set" (MTS) in the VVC standard.
[0111] The forward secondary transform of module 330 is typically a non-separable transform that is only applied to the residual of the intra-predicted CU and can still be bypassed. The forward secondary transform operates on 16 samples (arranged in the upper left 4×4 sub-block of the main transform coefficients 328) or 48 samples (arranged in three 4×4 sub-blocks of the upper left 8×8 coefficients of the main transform coefficients 328) to produce a set of secondary transform coefficients. The number of secondary transform coefficient sets can be less than the number of main transform coefficient sets from which they are derived. Since the secondary transform is only applied to coefficient sets that are adjacent to each other and include the DC coefficient, the secondary transform is called a "low-frequency non-separable secondary transform" (LFNST). In addition, when LFNST is applied, all remaining coefficients in the TB must be zero in both the main transform domain and the secondary transform domain.
[0112] The quantization parameter 392 is constant for a given TB and thus results in uniform scaling of the residual coefficients produced in the main transform domain of the TB. The quantization parameter 392 may be varied periodically by a signaled "delta quantization parameter". The delta quantization parameter (delta QP) is signaled once for a CU contained in a given region (referred to as a "quantization group"). If the CU is larger than the quantization group size, the delta QP is signaled once by one of the TBs of the CU. That is, the entropy encoder 338 signals the delta QP once for the first quantization group of the CU, and does not signal the delta QP for any subsequent quantization groups of the CU. Non-uniform scaling is also possible by applying a "quantization matrix", whereby the scaling factors applied to each residual coefficient are derived from a combination of the quantization parameter 392 and the corresponding entries in the scaling matrix. The scaling matrix may have a size less than the size of the TB, and when applied to a TB, a nearest neighbor method is used to provide scaling values for each residual coefficient according to a scaling matrix whose size is less than the size of the TB. The residual coefficients 336 are supplied to the entropy encoder 338 for encoding in the bitstream 115. Typically, the residual coefficients of each TB of a TU having at least one significant residual coefficient are scanned to produce an ordered list of values, according to a scan mode. The scan mode typically scans the TBs as a sequence of 4×4 "sub-blocks", providing a regular scan operation with a granularity of 4×4 groups of residual coefficients, where the arrangement of the sub-blocks depends on the size of the TB. The scanning within each sub-block and the progression from one sub-block to the next typically follows a backward diagonal scan pattern. In addition, the quantization parameter 392 is encoded in the bitstream 115 using the delta QP syntax element, and the secondary transform index 388 is encoded in the reference TB 115. Figures 13 to 15 The described conditions are encoded in the bitstream 115 .
[0113] As described above, the video encoder 114 needs access to a frame representation that corresponds to the encoded frame representation seen in the video decoder 134. Thus, the residual coefficients 336 pass through the inverse secondary transform module 344 (operating according to the secondary transform coefficients 388) to produce intermediate inverse transform coefficients represented by arrows 342. The intermediate inverse transform coefficients 346 are inversely quantized by the dequantization module 340 according to the quantization parameters 392 to produce residual samples represented by arrows 346. The intermediate inverse transform coefficients 346 are passed to the inverse main transform module 348 to produce residual samples for the TU represented by arrows 350. The type of inverse transform performed by the inverse secondary transform module 344 corresponds to the type of forward transform performed by the forward secondary transform module 330. The type of inverse transform performed by the inverse main transform module 348 corresponds to the type of main transform performed by the main transform module 326. The summation module 352 adds the residual samples 350 and the PU 320 to produce reconstructed samples for the CU (indicated by arrows 354).
[0114] The reconstructed samples 354 are passed to the reference sample cache 356 and the in-loop filter module 368. The reference sample cache 356, which is typically implemented using static RAM on an ASIC (thus avoiding expensive off-chip memory accesses), provides the minimum sample storage required to satisfy the dependencies used to generate intra PBs for subsequent CUs in the frame. The minimum dependencies typically include a "line buffer" of samples along the bottom of a row of CTUs for use by the next row of CTUs and a column buffer whose range is set by the height of the CTU. The reference sample cache 356 feeds reference samples (represented by arrow 358) to the reference sample filter 360. The sample filter 360 applies a smoothing operation to produce filtered reference samples (indicated by arrow 362). The filtered reference samples 362 are used by the intra prediction module 364 to produce an intra prediction block of samples represented by arrow 366. For each candidate intra prediction mode, the intra prediction module 364 produces a sample block, namely 366. The sample block 366 is generated by the module 364 using a technique such as DC, planar or angular intra prediction.
[0115] The in-loop filter module 368 applies several filtering stages to the reconstructed samples 354. The filtering stages include a "deblocking filter" (DBF), which applies smoothing aligned with CU boundaries to reduce artifacts caused by discontinuities. Another filtering stage present in the in-loop filter module 368 is an "adaptive loop filter" (ALF), which applies a Wiener-based adaptive filter to further reduce distortion. Another available filtering stage in the in-loop filter module 368 is a "sample adaptive offset" (SAO) filter. The SAO filter works by first classifying the reconstructed samples into one or more categories and applying an offset at the sample level according to the assigned category.
[0116] Filtered samples, represented by arrow 370, are output from the in-loop filter module 368. The filtered samples 370 are stored in a frame buffer 372. The frame buffer 372 typically has the capacity to store several (e.g., up to 16) pictures and is therefore stored in the memory 206. Due to the large memory consumption required, the frame buffer 372 is typically not stored using on-chip memory. As such, access to the frame buffer 372 is expensive in terms of memory bandwidth. The frame buffer 372 provides a reference frame (represented by arrow 374) to the motion estimation module 376 and the motion compensation module 380.
[0117] The motion estimation module 376 estimates a plurality of "motion vectors" (denoted as 378), each of which is a Cartesian spatial offset relative to the position of the current CB, thereby referencing a block in one of the reference frames in the frame buffer 372. A filtered block of reference samples (denoted as 382) is generated for each motion vector. The filtered reference samples 382 form a further candidate mode for potential selection by the mode selector 386. In addition, for a given CU, the PB 320 may be formed using one reference block ("uni-prediction"), or may be formed using two reference blocks ("bi-prediction"). For the selected motion vector, the motion compensation module 380 generates the PU 320 according to a filtering process that supports sub-pixel precision in the motion vector. In this way, the motion estimation module 376 (which operates on many candidate motion vectors) can perform a simplified filtering process compared to the motion compensation module 380 (which operates only on the selected candidate) to achieve reduced computational complexity. When the video encoder 114 selects inter-frame prediction for the CU, the motion vector 378 is encoded in the bitstream 115.
[0118] Although the reference to Versatile Video Coding (VVC) describes Figure 3 , but other video coding standards or implementations may also employ the processing stages of modules 310-390. Frame data 113 (and bitstream 115) may also be read from memory 206, hard drive 210, CD-ROM, Blue-ray diskTM, or other computer-readable storage media (or written to memory 206, hard drive 210, CD-ROM, Blue-ray disk, or other computer-readable storage media). In addition, frame data 113 (and bitstream 115) may be received from (or sent to) an external source (such as a server or radio frequency receiver connected to a communication network 220). The communication network 220 may provide limited bandwidth, requiring rate control to be used in the video encoder 114 to avoid saturating the network when the frame data 113 is difficult to compress. In addition, the bitstream 115 may be constructed from one or more stripes representing a spatial portion (CTU set) of the frame data 113, which are generated by one or more instances of the video encoder 114 and operated in a coordinated manner under the control of the processor 205. In the context of the present invention, a slice may also be referred to as a “contiguous portion” of the bitstream. A slice is contiguous within the bitstream and (eg, if parallel processing is being used) may be encoded or decoded as separate portions.
[0119] exist Figure 4 The video decoder 134 is shown in FIG. Figure 4 The video decoder 134 of FIG. 1 is an example of a Versatile Video Coding (VVC) video decoding pipeline, but other video codecs may also be used to perform the processing stages described herein. Figure 4As shown, the bitstream 133 is input to the video decoder 134. The bitstream 133 can be read from the memory 206, the hard disk drive 210, the CD-ROM, the Blu-ray disc or other non-transitory computer-readable storage medium. Alternatively, the bitstream 133 can be received from an external source (such as a server or a radio frequency receiver connected to the communication network 220). The bitstream 133 contains the encoding syntax elements representing the captured frame data to be decoded.
[0120] The bitstream 133 is input to the entropy decoder module 420. The entropy decoder module 420 extracts syntax elements from the bitstream 133 by decoding a sequence of "bins" and passes the values of the syntax elements to other modules in the video decoder 134. The entropy decoder module 420 uses variable length and fixed length decoding to decode SPS, PPS or slice headers, and uses an arithmetic decoding engine to decode the syntax elements of slice data into a sequence of one or more bins. Each bin can use one or more "contexts", where the context describes the probability level of "one" and "zero" values for encoding the bin. In the case where multiple contexts are available for a given bin, a "context modeling" or "context selection" step is performed to select one of the available contexts to decode the bin. The process of decoding the bin forms a sequential feedback loop, so that each slice can be decoded as a whole by a given entropy decoder 420 instance. A single (or a few) high-performing entropy decoder 420 instances can decode all slices of a frame from the bitstream 115, and multiple low-performing entropy decoder 420 instances can decode slices of a frame from the bitstream 133 at the same time.
[0121] The entropy decoder module 420 applies an arithmetic coding algorithm, such as "context adaptive binary arithmetic coding" (CABAC), to decode syntax elements from the bitstream 133. The decoded syntax elements are used to reconstruct parameters within the video decoder 134. The parameters include residual coefficients (represented by arrow 424), quantization parameters 474, secondary transform indexes 470, and mode selection information such as intra-frame prediction modes (represented by arrow 458). The mode selection information also includes information such as motion vectors, and partitioning of each CTU into one or more CBs. The parameters are used to generate PBs, usually in combination with sample data from previously decoded CBs.
[0122] The residual coefficients 424 are passed to the inverse secondary transform module 436, where they are transformed according to the reference Figures 16 to 18The described method applies a secondary transform or does not operate (bypass). The inverse secondary transform module 436 produces reconstructed transform coefficients 432, i.e., main transform domain coefficients, from the secondary transform domain coefficients. The reconstructed transform coefficients 432 are input to the dequantizer module 428. The dequantizer module 428 inverse quantizes (or "scales") the residual coefficients 432, i.e., in the main transform coefficient domain, to create a reconstructed intermediate transform coefficient represented by arrow 440 according to the quantization parameter 474. If the use of a non-uniform inverse quantization matrix is indicated in the bitstream 133, the video decoder 134 reads the quantization matrix from the bitstream 133 as a sequence of scaling factors and arranges the scaling factors into a matrix according to the quantization parameters. Inverse scaling uses the quantization matrix in combination with the quantization parameters to create the reconstructed intermediate transform coefficients 440.
[0123] The reconstructed transform coefficients 440 are passed to an inverse main transform module 444. Module 444 transforms the coefficients 440 from the frequency domain back to the spatial domain. The result of the operation of module 444 is a block of residual samples represented by arrow 448. The block of residual samples 448 is equal in size to the corresponding CB. The block of residual samples 448 is supplied to a summing module 450. At the summing module 450, the residual samples 448 are added to the decoded PB represented as 452 to produce a block of reconstructed samples represented by arrow 456. The reconstructed samples 456 are supplied to a reconstructed sample cache 460 and an in-loop filtering module 488. The in-loop filtering module 488 produces a reconstructed block of frame samples represented as 492. The frame samples 492 are written to a frame buffer 496.
[0124] The reconstructed sample cache 460 operates in a manner similar to the reconstructed sample cache 356 of the video encoder 114. The reconstructed sample cache 460 provides storage for the reconstructed samples needed for intra prediction of subsequent CBs without the memory 206 (e.g., by using data 232, which is typically on-chip memory, instead). Reference samples, represented by arrows 464, are obtained from the reconstructed sample cache 460 and are supplied to a reference sample filter 468 to produce filtered reference samples, represented by arrows 472. The filtered reference samples 472 are supplied to an intra prediction module 476. The module 476 produces a block of intra prediction samples, represented by arrows 480, based on the intra prediction mode parameters 458 represented in the bitstream 133 and decoded by the entropy decoder 420. The block of samples 480 is generated using a mode such as DC, planar, or angular intra prediction.
[0125] When the prediction mode of the CB is indicated in the bitstream 133 to use intra prediction, the intra prediction samples 480 form the decoded PB 452 via the multiplexer module 484. Intra prediction produces a prediction block (PB) of samples, i.e., a block in one color component derived using "neighboring samples" in the same color component. Neighboring samples are samples that are adjacent to the current block and have been reconstructed because they are at the front in the block decoding order. In the case where luma and chroma blocks are juxtaposed, luma and chroma blocks can use different intra prediction modes. However, the two chroma channels share the same intra prediction mode.
[0126] When the prediction mode of the CB is indicated in the bitstream 133 as intra prediction, the motion compensation module 434 uses the motion vector (decoded from the bitstream 133 by the entropy decoder 420) and the reference frame index to select and filter a block of samples 498 from the frame buffer 496 to produce a block of inter-prediction samples indicated as 438. The block of samples 498 is obtained from a previously decoded frame stored in the frame buffer 496. For bi-prediction, two blocks of samples are generated and blended together to produce samples of the decoded PB 452. The frame buffer 496 is populated with filter block data 492 from the in-loop filtering module 488. Like the in-loop filtering module 368 of the video encoder 114, the in-loop filtering module 488 applies any of the DBF, ALF, and SAO filtering operations. Typically, the motion vector is applied to both the luma and chroma channels, but the filtering process for sub-sample interpolation in the luma and chroma channels is different.
[0127] Figure 5 is a schematic block diagram showing a set 500 of possible partitions or splits of a region into one or more sub-regions in a tree structure for general video coding. Figure 3 As described, the partitioning shown in set 500 may be utilized by block partitioner 310 of encoder 114 to partition each CTU into one or more than one CU or CB according to the encoding number as determined by Lagrangian optimization.
[0128] Although set 500 only shows the partitioning of a square region into other possible non-square sub-regions, it should be understood that set 500 is showing the potential partitioning of a parent node in the coding tree to a child node in the coding tree, and there is no requirement that the parent node corresponds to a square region. If the containing region is non-square, the size of the block resulting from the partition is scaled according to the aspect ratio of the containing block. Once a region is not further split, that is, at a leaf node of the coding tree, a CU occupies the region.
[0129] The process of subdividing a region into subregions must terminate when the resulting subregions reach a minimum CU size (typically 4×4 luma samples). In addition to constraining the CU to prohibit block regions from being smaller than a predetermined minimum size of, for example, 16 samples, the CU is constrained to have a minimum width or height of four. Other minimum values are also possible in terms of width and height or in terms of both width or height. The subdivision process can also terminate before the deepest level of decomposition, resulting in a CU larger than the minimum CU size. It is possible that no splitting occurs, resulting in a single CU occupying the entire CTU. A single CU occupying the entire CTU is the maximum available coding unit size. Due to the use of subsampled chroma formats (such as 4:2:0, etc.), the arrangement of the video encoder 114 and the video decoder 134 can terminate the splitting of regions in the chroma channel earlier than in the luma channel, including the case of a common coding tree defining the block structure of the luma and chroma channels. When separate coding trees are used for luma and chroma, the constraints on the available splitting operations ensure a minimum chroma CB area of 16 samples, even if such a CB is juxtaposed with a larger luma area (e.g., 64 luma samples).
[0130] In the absence of further sub-division, there is a CU at the leaf node of the coding tree. For example, leaf node 510 contains a CU. At the non-leaf node of the coding tree, there is a split to two or more other nodes, where each node can be a leaf node forming a CU, or a non-leaf node containing further splits to smaller areas. At each leaf node of the coding tree, there is a coding block for each color channel. The split that terminates at the same depth for both brightness and chrominance obtains three juxtaposed CBs. The split that terminates at a deeper depth for brightness than for chrominance obtains multiple brightness CBs juxtaposed with the CBs of the chrominance channels.
[0131] like Figure 5 As shown, quadtree split 512 splits the containing area into four equal-sized areas. Compared to HEVC, Versatile Video Coding (VVC) achieves additional flexibility through additional splits, including horizontal binary split 514 and vertical binary split 516. Splits 514 and 516 each split the containing area into two equal-sized areas. The splits are along horizontal boundaries (514) or vertical boundaries (516) within the containing block.
[0132] Further flexibility is achieved in general video coding by adding ternary horizontal split 518 and ternary vertical split 520. Ternary split 518 and 520 divide the block into three regions that form boundaries in the horizontal direction (518) or vertical direction (520) along 1 / 4 and 3 / 4 of the width or height of the containing region. The combination of quadtree, binary tree and ternary tree is called "QTBTTT". The root of the tree includes zero or more quadtree splits (the "QT" part of the tree). Once the QT part is terminated, zero or more binary or ternary splits ("multi-tree" or "MT" part of the tree) may occur, eventually ending in a CB or CU at the leaf node of the tree. In the case where the tree describes all color channels, the leaf node of the tree is a CU. In the case where the tree describes a luminance channel or a chrominance channel, the leaf node of the tree is a CB.
[0133] Compared to HEVC, which only supports quadtrees and therefore only supports square blocks, QTBTTT, in particular, takes into account the possible recursive application of binary and / or ternary tree splits to obtain more possible CU sizes. When only quadtree splitting is available, each increase in the coding tree depth corresponds to a reduction in the CU size to one-quarter of the size of the parent region. In VVC, the availability of binary and ternary splits means that the coding tree depth no longer corresponds directly to the CU region. The possibility of abnormal (non-square) block sizes can be reduced by constraining the splitting options to eliminate the block width or height that will be less than four samples or the split that will not be a multiple of four samples. Typically, the constraints will apply when considering luminance samples. However, in the described arrangement, the constraints can be applied separately to the blocks of the chrominance channel. The application of the constraints of the splitting options to the chrominance channel may obtain different minimum block sizes for luminance vs chrominance (for example, when the frame data adopts a 4:2:0 chrominance format or a 4:2:2 chrominance format). Each split produces a sub-region with a side size that is unchanged, bisected or quartered relative to the containing area. Then, since the CTU size is a power of 2, the side dimensions of all CUs are also a power of 2.
[0134] Figure 6 6 is a schematic flow diagram showing a data flow 600 of a QTBTTT (or "coding tree") structure used in general video coding. The QTBTTT structure is used for each CTU to define the partitioning of the CTU into one or more CUs. The QTBTTT structure for each CTU is determined by the block partitioner 310 in the video encoder 114 and encoded into the bitstream 115 or decoded from the bitstream 133 by the entropy decoder 420 in the video decoder 134. Figure 5 As shown partitioned, data flow 600 further features permitted combinations that may be used by block partitioner 310 to partition a CTU into one or more CUs.
[0135] Starting from the top level of the hierarchy, i.e., at the CTU, zero or more quadtree partitions are first performed. Specifically, a quadtree (QT) split decision 610 is made by the block partitioner 310. The decision at 610 returns a "1" symbol, which indicates a decision to split the current node into four child nodes according to the quadtree split 512. The result is that four new nodes are generated, such as at 620, and for each new node, recursively return to the QT split decision 610. Each new node is considered in raster (or Z-scan) order. Alternatively, if the QT split decision 610 indicates that no further splitting is to be performed (returns a "0" symbol), the quadtree partitioning stops, and multi-tree (MT) splitting is then considered.
[0136] First, an MT split decision 612 is made by the block partitioner 310. At 612, a decision to perform an MT split is indicated. A "0" symbol is returned at the decision 612, indicating that no further splitting of the node into child nodes will be performed. If no further splitting of the node will be performed, the node is a leaf node of the coding tree and corresponds to a CU. The leaf node is output at 622. Alternatively, if the MT split 612 indicates a decision to perform an MT split (a "1" symbol is returned), the block partitioner 310 proceeds to a direction decision 614.
[0137] Direction decision 614 indicates the direction of the MT split as horizontal ("H" or "0") or vertical ("V" or "1"). If decision 614 returns "0" indicating a horizontal direction, block partitioner 310 proceeds to decision 616. If decision 614 returns "1" indicating a vertical direction, block partitioner 310 proceeds to decision 618.
[0138] In each of decisions 616 and 618, the number of partitions into which the MT is split is indicated as two (binary split or "BT" nodes) or three (ternary split or "TT") in the case of a BT / TT split. That is, the BT / TT split decision 616 is made by the block partitioner 310 when the direction indicated from 614 is horizontal, and the BT / TT split decision 618 is made by the block partitioner 310 when the direction indicated from 614 is vertical.
[0139] The BT / TT split decision 616 indicates whether the horizontal split is a binary split 514 indicated by returning “0” or a ternary split 518 indicated by returning “1”. When the BT / TT split decision 616 indicates a binary split, at a generate HBT CTU node step 625, the block partitioner 310 generates two nodes according to the binary horizontal split 514. When the BT / TT split 616 indicates a ternary split, at a generate HTT CTU node step 626, the block partitioner 310 generates three nodes according to the ternary horizontal split 518.
[0140] The BT / TT split decision 618 indicates whether the vertical split is a binary split 516 indicated by returning "0" or a ternary split 520 indicated by returning "1". When the BT / TT split 618 indicates a binary split, at step 627 of generating VBT CTU nodes, the block partitioner 310 generates two nodes according to the vertical binary split 516. When the BT / TT split 618 indicates a ternary split, at step 628 of generating VTT CTU nodes, the block partitioner 310 generates three nodes according to the vertical ternary split 520. For each node obtained from steps 625-628, the recursion of the data flow 600 back to the MT split decision 612 is applied in a left-to-right or top-to-bottom order according to the direction 614. As a result, binary and ternary tree splits can be applied to generate CUs of various sizes.
[0141] Fig. 7A and 7B An example partition 700 of a CTU 710 into multiple CUs or CBs is provided. Fig. 7A An example CU 712 is shown in FIG. Fig. 7A 7 shows the spatial arrangement of CUs in a CTU 710. Example partitioning 700 in Figure 7B Also shown in FIG. 7 is a coding tree 720 .
[0142] exist Fig. 7A At each non-leaf node (e.g., nodes 714, 716, and 718) in the CTU 710 of the coding tree 720, the contained nodes (which may be further partitioned or may be CUs) are scanned or traversed in "Z order" to create a list of nodes represented as columns in the coding tree 720. For quadtree splits, the Z order scan results in an order from top left to right followed by a bottom left to right order. For horizontal and vertical splits, the Z order scan (traversal) is simplified to a scan from top to bottom and a scan from left to right, respectively. Figure 7B The coding tree 720 lists all nodes and CUs according to the scanning order applied. Each split generates a list of two, three or four new nodes at the next level of the tree until a leaf node (CU) is reached.
[0143] In reference Figure 3 Where the image is decomposed into CTUs and further into CUs using block partitioner 310, and the CUs are used to generate respective residual blocks (324), the residual blocks are forward transformed and quantized using video encoder 114. The resulting TBs 336 are then scanned to form a sequential list of residual coefficients as part of the operation of entropy encoding module 338. Equivalent processing is performed in video decoder 134 to obtain the TBs from bitstream 133.
[0144] Fig. 8A , 8B8C show the subdivision levels resulting from the split in the coding tree and the corresponding impact on the partitioning of the coding tree units into quantization groups. The incremental QP (392) is signaled at most once for each quantization group through the residual of the TB. In HEVC, the definition of quantization groups corresponds to the coding tree depth, since this definition results in regions of fixed size. In VVC, the additional splits mean that the coding tree depth is no longer a suitable proxy for CTU regions. In VVC, "subdivision levels" are defined, where each increment corresponds to half of the contained region.
[0145] Fig. 8A A set 800 of splits in a coding tree and corresponding subdivision levels are shown. At the root node of the coding tree, the subdivision level is initialized to zero. When the coding tree includes a quadtree split (e.g., 810), the subdivision level is incremented by two for any CU contained therein. When the coding tree includes a binary split (e.g., 812), the subdivision level is incremented by 1 for any CU contained therein. When the coding tree includes a ternary split (e.g., 814), the subdivision level is incremented by 2 for the outer two CUs and by 1 for the inner CU resulting from the ternary split. When traversing the coding tree for each CTU, as shown in reference Figure 6 As described above, the subdivision level of each resulting CU is determined according to the set 800 .
[0146] Figure 8B An example set 840 of CU nodes is shown, and the effect of the split is shown. An example parent node 820 of the set 840 with a subdivision level of zero corresponds to Figure 8B The example of FIG. 8 is a CTU of size 64×64. The parent node 820 is ternary split to produce three child nodes 821, 822, and 823 of sizes 16×64, 32×64, and 16×64, respectively. The child nodes 821, 822, and 823 have subdivision levels 2, 1, and 2, respectively.
[0147] exist Figure 8BIn the example of , the quantization group threshold is set to 1, corresponding to half of the 64×64 region, i.e., corresponding to a region of 2048 samples. A flag tracks the start of a new QG. For any node with a subdivision level less than or equal to the quantization group threshold, the flag that tracks the new QG is reset. The flag is set when traversing the parent node 820 with a subdivision level of zero. Although the center CU 822 of size 32×64 has a region of 2048 samples, the two sibling CUs 821 and 823 have a subdivision level of 2, i.e., a region of 1024, so the flag is not reset when traversing the center CU, and the quantization group does not start at the center CU. Instead, following the initial flag reset, the flag starts at the parent node as shown in 824. Effectively, the QP can only change on boundaries that are aligned with multiples of the quantization group region. The incremental QP is signaled along with the residual of the TB associated with the CB. If there are no significant coefficients, there is no opportunity to encode the incremental QP.
[0148] Figure 8C An example 860 of splitting a CTU 862 into multiple CUs and QGs is shown to illustrate the relationship between subdivision levels, QGs, and the signaling of delta QPs. The vertical binary split splits the CTU 862 into two halves, with the left half 870 containing one CU CU0 and the right half 872 containing several CUs (CU1-CU4). Figure 8C In the example of , the quantization group threshold is set to 2 so that the quantization group generally has an area equal to one-fourth of the CTU area. Since the subdivision level of the parent node (i.e., the root node of the coding tree) is zero, the QG flag is reset and a new QG starts with the next coded CU (i.e., the CU at arrow 868). CU0 (870) has coded coefficients, so the delta QP 864 is encoded with the residual of CU0. The right half 872 is horizontally binary split and further split in the upper and lower parts of the right half 872, resulting in CU1-CU4. The subdivision level of the coding tree nodes corresponding to the upper (877 including CU1 and CU2) and lower (878 including CU3 and CU4) parts of the right half 872 is 2. Subdivision level 2 is equal to quantization group threshold 2, so new QGs start in the respective parts, labeled 874 and 876, respectively. CU1 has no coded coefficients (no residual), and CU2 is a "skipped" CU, which also has no coded coefficients. Therefore, for the upper part, the delta QP is not encoded. CU3 is a skip CU, and CU4 has a coded residual, so for the QG including CU3 and CU4, the delta QP 866 is encoded using the residual of CU4.
[0149] Fig.9A and 9BA 4×4 transform block scan pattern and associated primary and secondary transform coefficients are shown. The operation of the secondary transform module 330 on the primary residual coefficients is described from the perspective of the video encoder 114. The 4×4 TB 900 is scanned according to a backward diagonal scan pattern 910. The scan pattern 910 proceeds from the "last significant coefficient" position toward the DC (upper left) coefficient position. All coefficient positions that are not scanned, such as when considering scanning in the forward direction, the residual coefficients located after the last significant coefficient position are implicitly non-significant. When a secondary transform is used, all remaining coefficients are non-significant. That is, all secondary domain residual coefficients that are not subjected to a secondary transform are non-significant, and all primary domain residual coefficients that are not filled in by the application of the secondary transform need to be non-significant. In addition, after the forward secondary transform is applied by the module 330, there may be fewer secondary transform coefficients than the number of primary transform coefficients processed by the secondary transform module 330. For example, Fig. 9B A collection of blocks 920 is shown. Fig. 9B In , sixteen (16) primary coefficients are arranged as a 4×4 sub-block, i.e., 924 of 4×4 TB 920. Fig. 9B In the example of , the main residual coefficients can be secondary transformed to produce secondary transform block 926. Secondary transform block 926 contains eight secondary transform coefficients 928. The eight secondary transform coefficients 928 are stored in the TB according to scan pattern 910, packed from the DC coefficient position forward. The remaining coefficient positions of the 4×4 sub-block (shown as area 930) contain quantized residual coefficients from the main transform and need to be non-significant for the secondary transform to be applied. Therefore, the last significant coefficient position of the 4×4 TB, which is a coefficient in one of the first eight scan positions of TB 920, indicates that (i) the secondary transform is applied, or (ii) after quantization, the output of the main transform has no significant coefficients beyond the eighth scan position of TB 920.
[0150] When a secondary transform may be applied to a TB, a secondary transform index (i.e., 388) is encoded to indicate the possible application of the secondary transform. The secondary transform index may also indicate which core will be applied as the secondary transform at module 330 if multiple transform cores are available. Accordingly, when the last significant coefficient position is located at any scan position reserved for holding secondary transform coefficients (e.g., 928), the video decoder 134 decodes the secondary transform index 470.
[0151] Although a secondary transform kernel that maps 16 primary coefficients to eight secondary coefficients has been described, different kernels are possible, including kernels that map to different numbers of secondary transform coefficients. The number of secondary transform coefficients can be the same as the number of primary transform coefficients, for example 16. For TBs with a width of 4 and a height greater than 4, the behavior described for the 4×4 TB case applies to the top sub-block of the TB. When the secondary transform is applied, the other sub-blocks of the TB have zero-valued residual coefficients. For TBs with a width greater than 4 and a height equal to 4, the behavior described for the 4×4 TB case applies to the leftmost sub-block of the TB, and the other sub-blocks of the TB have zero-valued residual coefficients, allowing the last significant coefficient position to be used to determine whether a secondary transform index needs to be decoded.
[0152] Fig. 9C and 9D An 8x8 transform block scan pattern and example associated primary and secondary transform coefficients are shown. Fig. 9C A 4x4 sub-block based backward diagonal scan pattern 950 is shown for the 8x8 TB 940. The 8x8 TB 940 is scanned in the 4x4 sub-block based backward diagonal scan pattern 950. Fig.9D Set 960 is shown, which shows the operating effects of the secondary transform. Scan 950 returns from the last significant coefficient position to the DC (upper left) coefficient position. When the remaining 16 main coefficients (shown as 964) are zero values, it is possible to apply the forward secondary transform kernel to 48 main coefficients (shown as area 962 of 940). Applying the secondary transform to area 962 results in 16 secondary transform coefficients shown as 966. The other coefficient positions of the TB are zero values, labeled 968. If the last significant position of the 8×8 TB 940 indicates that the secondary transform coefficient is within 966, then the secondary transform index 388 is encoded to instruct the module 330 to apply a specific transform kernel (or bypass the kernel). The video decoder 134 uses the last significant position of the TB to determine whether to decode the secondary transform index, i.e., index 470. For transform blocks with a width or height greater than eight samples, Fig. 9C and 9D The method is applied to the upper left 8×8 region, that is, the upper left 2×2 sub-block of the TB.
[0153] like 9A to 9DAs described in , two sizes of secondary transform kernels are available. One size of secondary transform kernel is used for transform blocks with a width or height of 4, and another size of secondary transform is used for transform blocks with a width and height greater than 4. Within kernels of each size, multiple sets (e.g., four) of secondary transform kernels are available. A set is selected based on the intra prediction mode of the block, and the set may be different between luma blocks and chroma blocks. Within the selected set, one or two kernels are available. Independent of the luma blocks and chroma blocks in the coding units belonging to the common tree of the coding tree unit, the use of a kernel within the selected set or bypassing the secondary transform is signaled via the secondary transform index. In other words, the index for the luma channel and the index for the chroma channel are independent of each other.
[0154] Fig.10 A set 1000 of transform blocks available in the Versatile Video Coding (VVC) standard is shown. Fig.10 Also shown is the application of a secondary transform to a subset of the residual coefficients of the transformed blocks from set 1000 . Fig.10 TBs are shown with widths and heights ranging from 4 to 32. However, TBs with a width and / or height of 64 are possible but not shown for ease of reference.
[0155] A 16-point secondary transform 1052 (shown with darker shading) is applied to a 4×4 set of coefficients. The 16-point secondary transform 1052 is applied to TBs with a width or height of 4, such as 4×4 TB 1010, 8×4 TB 1012, 16×4 TB 1014, 32×4 TB 1016, 4×8 TB 1020, 4×16 TB 1030, and 4×32 TB 1040. If a 64-point primary transform is available, the 16-point secondary transform 1052 is applied to TBs of size 4×64 and 64×4 ( Fig.10 (not shown in the figure). For a TB with a width or height of four but with more than 16 primary coefficients, the 16-point secondary transform is applied only to the top left 4×4 sub-block of the TB, and other sub-blocks are required to have zero-valued coefficients to apply the secondary transform. Typically, applying the 16-point secondary transform results in 16 secondary transform coefficients, which are packed into the TB for encoding in the sub-blocks that obtain the original 16 primary transform coefficients. For example, as shown in reference Fig. 9B As described, the secondary transform kernel may cause creation of secondary transform coefficients which are smaller in number than the number of primary transform coefficients to which the secondary transform is applied.
[0156] For transform sizes with width and height greater than four, such as Fig.10As shown, a 48-point secondary transform 1050 (shown with lighter shading) is available for application to three 4×4 sub-blocks of residual coefficients in the upper left 8×8 region of the transform block. In each case in the region shown with light shading and dashed outline, the 48-point secondary transform 1050 is applied to 8×8 transform block 1022, 16×8 transform block 1024, 32×8 transform block 1026, 8×16 transform block 1032, 16×16 transform block 1034, 32×16 transform block 1036, 8×32 transform block 1042, 16×32 transform block 1044, and 32×32 transform block 1046. If a 64-point primary transform is available, the 48-point secondary transform 1050 is also applicable to TBs of sizes 8×64, 16×64, 32×64, 64×64, 64×32, 64×16, and 64×8 (not shown). The application of a 48-point secondary transform kernel typically results in fewer than 48 secondary transform coefficients being generated. For example, 8 or 16 secondary transform coefficients may be generated. The secondary transform coefficients are stored in the transform block in the upper left region, e.g. Fig.9D The eight secondary transform coefficients are shown in FIG. The main transform coefficients that have not been subjected to the secondary transform (“main coefficients only”) (e.g., coefficient 1066 of TB 1034 (similar to Fig.9D 964)) needs to be zero value to apply the secondary transform. After applying the 48-point secondary transform 1050 in the forward direction, the area that can contain significant coefficients is reduced from 48 coefficients to 16 coefficients, thereby further reducing the number of coefficient positions that may contain significant coefficients. For example, 968 will contain only non-significant coefficients. For the inverse secondary transform, for example, the decoded significant coefficients that only exist in 966 of TB are transformed to produce any coefficient that may be significant in the area (e.g., 962), and then these coefficients are subjected to the main inverse transform. When the secondary transform reduces one or more sub-blocks to a set of 16 secondary transform coefficients, only the upper left 4×4 sub-block can contain significant coefficients. The last significant coefficient position located at any coefficient position that can store a secondary transform coefficient indicates that a secondary transform or only a main transform is applied. However, after quantization, the resulting significant coefficients are in the same area as if the secondary transform kernel had been applied.
[0157] When the last significant coefficient position indicates a secondary transform coefficient position in a TB (e.g., 922 or 962), a signaled secondary transform index is needed to distinguish between applying the secondary transform kernel or bypassing the secondary transform. Although the application of the secondary transform to the video encoder 114 has been described from the perspective of the video encoder 114, Fig.10TBs of various sizes in 928 or 966, but the corresponding inverse processing is performed in the video decoder 134. The video decoder 134 first decodes the last significant coefficient position. If the decoded last significant coefficient position indicates a potential application of a secondary transform, that is, the position is within a secondary transform kernel 928 or 966 that produces 8 or 16 secondary transform coefficients, respectively, then the secondary transform index is decoded to determine whether to apply or bypass the inverse secondary transform.
[0158] Fig.11 A syntactic structure 1100 of a bitstream 1101 with multiple slices is shown. Each slice in the slice contains multiple coding units. The bitstream 1101 can be generated by the video encoder 114 as, for example, a bitstream 115, or can be parsed by the video decoder 134 as, for example, a bitstream 133. The bitstream 1101 is divided into multiple parts, such as network abstraction layer (NAL) units, where the depiction is achieved by setting a NAL unit header (such as 1108, etc.) before each NAL unit. The sequence parameter set (SPS) 1110 defines sequence-level parameters, such as profiles (tool sets), chroma formats, sample bit depths, and frame resolutions for encoding and decoding bitstreams. Also included in the set 1110 are parameters for constraining the application of different types of splits in the coding tree of each CTU. For example, using a log2 basis for block size constraints and expressing parameters relative to other parameters (such as the minimum CTU size, etc.), the encoding of parameters of the constrained split type can be optimized for a more compact representation. Several parameters encoded in the SPS 1110 are as follows:
[0159] • log2_ctu_size_minus5: specifies the CTU size, where the encoded values 0, 1, and 2 specify CTU sizes of 32x32, 64x64, and 128x128, respectively.
[0160] partition_constraints_override_enabled_flag: enables stripe-level override of several parameters, collectively referred to as partition constraint parameters 1130 , to be applied.
[0161] log2_min_luma_coding_block_size_minus2: specifies the minimum coding block size (in luma samples), where values 0, 1, 2, ... specify minimum luma CB sizes of 4x4, 8x8, 16x16, ... The maximum coded value is constrained by the specified CTU size, i.e., such that log2_min_luma_coding_block_size_minus2 ≤ log2_ctu_size_minus5+3. The available chroma block sizes correspond to the available luma block sizes, scaled according to the chroma channel subsampling of the chroma format in use.
[0162] sps_max_mtt_hierarchy_depth_inter_slice: specifies the maximum hierarchy depth of coding units in the coding tree of multi-tree splits (ie, binary and ternary splits) relative to quadtree nodes in the coding tree of inter (P or B) slices (ie, once a quadtree split stops in the coding tree), and is one of the parameters 1130 .
[0163] sps_max_mtt_hierarchy_depth_intra_slice_luma: specifies the maximum hierarchy depth of coding units in the coding tree of multi-tree splits (ie, binary and ternary) relative to quadtree nodes in the coding tree of intra (I) slices (ie, once quadtree splitting stops in the coding tree), and is one of the parameters 1130 .
[0164] partition_constraints_override_flag: When partition_constraints_override_enabled_flag in the SPS is equal to 1, this parameter is signaled in the slice header and indicates that the partition constraints signaled in the SPS are to be overwritten for the corresponding slice.
[0165] A picture parameter set (PPS) 1112 defines a set of parameters applicable to zero or more frames. The parameters included in the PPS 1112 include parameters for partitioning a frame into one or more "regions" and / or "blocks". The parameters of the PPS 1112 may also include a list of CU chroma QP offsets, one of which may be applied at the CU level to derive a quantization parameter for use by a chroma block from the quantization parameter of a collocated luma CB.
[0166] A sequence of slices forming one picture is called an access unit (AU), such as AU 0 1114. AU 0 1114 includes three slices, such as slices 0 to 2. Slice 1 is labeled 1116. Slice 1 (1116) includes a slice header 1118 and slice data 1120 like other slices.
[0167] The slice header includes parameters grouped into 1134. Group 1134 includes:
[0168] slice_max_mtt_hierarchy_depth_luma: signaled in the slice header 1118 when partition_constraints_override_flag in the slice header is equal to 1 and overrides the value derived from the SPS. For an I slice, instead of using sps_max_mtt_hierarchy_depth_intra_slice_luma to set the MaxMttDepth at 1134, slice_max_mtt_hierarchy_depth_luma is used. For a P or B slice, instead of using sps_max_mtt_hierarchy_depth_inter_slice, slice_max_mtt_hierarchy_depth_luma is used.
[0169] The variable MinQtLog2SizeIntraY (not shown) is derived from the syntax element sps_log2_diff_min_qt_min_cb_intra_slice_luma decoded from SPS1110, which specifies the minimum coding block size resulting from zero or more quadtree splits of the I slice (i.e., no further MTT splitting occurs in the coding tree). The variable MinQtLog2SizeInterY (not shown) is derived from the syntax element sps_log2_diff_min_qt_min_cb_inter_slice decoded from SPS1110. The variable MinQtLog2SizeInterY specifies the minimum coding block size resulting from zero or more quadtree splits of the P and B slices (i.e., no further MTT splitting occurs in the coding tree). Since the CU resulting from the quadtree split is a square, the variables MinQtLog2SizeIntraY and MinQtLog2SizeInterY each specify both the width and height (as the log2 of the CU width / height).
[0170] The parameter cu_qp_delta_subdiv may optionally be signaled in the slice header 1118 and indicates the maximum subdivision level at which delta QP is signaled in the coding tree for the luma branch in the common tree or separate tree slices. For an I slice, the range of cu_qp_delta_subdiv is 0 to 2*(log2_ctu_size_minus5+5-MinQtLog2SizeIntraY+MaxMttDepthY 1134). For a P or B slice, the range of cu_qp_delta_subdiv is 0 to 2*(log2_ctu_size_minus5+5-MinQtLog2SizeInterY+MaxMttDepthY 1134). Since the range of cu_qp_delta_subdiv depends on the value MaxMttDepthY 1134 derived from the partition constraints obtained from the SPS 1110 or the slice header 1118, there is no parsing issue.
[0171] The parameter cu_chroma_qp_offset_subdiv may optionally be signaled in the slice header 1118 and indicates the maximum subdivision level for signaling chroma CU QP offsets in the common tree or in the chroma branch in a separate tree slice. The range constraints of cu_chroma_qp_offset_subdiv for I or P / B slices are the same as the corresponding range constraints for cu_qp_delta_subdiv.
[0172] A subdivision level 1136 is derived for the CTUs in the slice 1120, which specifies cu_qp_delta_subdiv for luma CB and cu_chroma_qp_offset_subdiv for chroma CB. Fig. 8A -C, the subdivision level is used to determine at which points in the CTU the delta QP syntax elements are encoded. For chroma CBs, the Fig. 8A -C method to signal chroma CU level offset enablement (and index, if enabled).
[0173] Fig.12A syntax structure 1200 of slice data 1120 of a bitstream 1101 (e.g., 115 or 133) is shown with a common tree for luma and chroma coding blocks of a coding tree unit, such as a CTU 1210. CTU 1210 includes one or more CUs, an example of which is shown as CU 1214. CU 1214 includes a signaled prediction mode 1216a, followed by a transform tree 1216b. When the size of CU 1214 does not exceed the maximum transform size (32x32 or 64x64), the transform tree 1216b includes one transform unit, shown as TU 1218.
[0174] If the prediction mode 1216a indicates that intra prediction is used for CU 1214, the luma intra prediction mode and the chroma intra prediction mode are specified. For the luma CB of CU 1214, the primary transform type is also signaled as (i) horizontal and vertical DCT-2, (ii) horizontal and vertical transform skip, or (iii) a combination of horizontal and vertical DCT-7 and DCT-8. If the luma transform type signaled is horizontal and vertical DCT-2 (option (i)), then in the reference Fig.9A -D, an additional luma secondary transform type 1220, also called "low frequency non-separable transform" (LFNST) index, is signaled in the bitstream. A chroma secondary transform type 1221 is also signaled. The chroma secondary transform type 1221 is signaled independently of whether the luma primary transform type is DCT-2.
[0175] The use of a common coding tree results in a TU 1218 including TBs for each color channel, shown as luma TB Y 1222, a first chroma TB Cb 1224, and a second chroma TB Cr 1226. A coding mode is available that sends a single chroma TB to specify the chroma residual for both Cb and Cr channels, called "joint CbCr" coding mode. When the joint CbCr coding mode is enabled, a single chroma TB is encoded.
[0176] Regardless of the color channel, each TB includes a last position 1228. The last position 1228 indicates the last valid residual coefficient position in the TB when considering the diagonal scan mode (which is used to serialize the coefficient array of the TB in the forward direction (i.e., from the DC coefficient forward). If the last position 1228 of the TB indicates that only the coefficients in the secondary transform domain (i.e., all the remaining coefficients that are subject to the main transform only) are valid, a secondary transform index is signaled to specify whether the secondary transform is applied.
[0177] If a secondary transform is to be applied and if more than one secondary transform kernel is available, the secondary transform index indicates which kernel is selected. Typically, in a "candidate set", one kernel is available, or two kernels are available. The candidate set is determined from the intra prediction mode of the block. Typically, there are four candidate sets, but there may be fewer candidate sets. As described above, a secondary transform is used for luma and chroma, and therefore the selected kernel depends on the intra prediction mode used for the luma and chroma channels, respectively. The kernel may also depend on the block size of the corresponding luma and chroma TBs. The kernel selected for chroma also depends on the chroma subsampling rate of the bitstream. If only one kernel is available, a signal is used to restrict the application or non-application of the secondary transform (index range 0 to 1). If two kernels are available, the index value is 0 (not applied), 1 (apply the first kernel), or 2 (apply the second kernel). For chroma, the same secondary transform kernel is applied to each chroma channel, so the residuals of the Cb block 1224 and the Cr block 1226 only need to include the valid coefficients in the positions subjected to the secondary transform, as shown in reference Fig.9A -D. If joint CbCr encoding is used, the requirement to include only significant coefficients in positions that are subject to the secondary transform only applies to a single encoded chroma TB, since the resulting Cb and Cr residuals contain significant coefficients only in positions that correspond to significant coefficients in the jointly encoded TB. If the applicable color channel(s) for a given secondary index are described by a single TB (a single last position, e.g., 1228), i.e., only one TB is always required for luma and one TB for chroma when joint CbCr encoding is used, the secondary transform index can be encoded immediately after the encoded last position instead of after the TU, i.e., as index 1230 instead of 1220 (or 1221). Signaling the secondary transform earlier in the bitstream allows the video decoder 134 to start applying the secondary transform as individual ones of the residual coefficients 1232 are decoded, thereby reducing latency in the system 100.
[0178] In the arrangement of the video encoder 114 and the video decoder 134, when joint CbCr encoding is not used, a separate secondary transform index is signaled for each chroma TB (i.e., 1224 and 1226), thereby obtaining independent control of the secondary transform of each color channel. If each TB is controlled independently, the secondary transform index of each TB can be signaled immediately after the last position of the corresponding TB of luma and chroma (regardless of whether the joint CbCr mode is applied).
[0179] Fig.13A method 1300 is shown for encoding frame data 113 in a bitstream 115, the bitstream 115 including one or more slices as a sequence of coding tree units. The method 1300 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. In addition, the method 1300 may be performed by the video encoder 114 under execution by the processor 205. Due to the workload of encoding a frame, the steps of the method 1300 may be performed in different processors to share the workload, for example using a contemporary multi-core processor, so that different slices are encoded by different processors. In addition, when encoding various portions (slices) of the bitstream 115, the partition constraints and quantization group definitions may vary from one slice to another, which is considered to be beneficial for rate control purposes. For additional flexibility in encoding the residuals of various coding units, not only can the quantization group subdivision level vary from one slice to another, but the application of the secondary transform is independently controllable for luma and chroma. Therefore, the method 1300 may be stored on a computer-readable storage medium and / or in the memory 206.
[0180] The method 1300 begins at an encode SPS / PPS step 1310. At step 1310, the video encoder 114 encodes the SPS 1110 and PPS 1112 as fixed and variable length coding parameter sequences in the bitstream 115. A partition_constraints_override_enabled_flag is encoded as part of the SPS 1110, indicating that the partition constraints can be overridden in the slice header (1118) of the corresponding slice (such as 1116). The default partition constraints are also encoded by the video encoder 114 as part of the SPS 1110.
[0181] From step 1310, the method 1300 continues to step 1320 of segmenting the frame into slices. In performing step 1320, the processor 205 segments the frame data 113 into one or more slices or contiguous portions. Where parallelism is desired, separate instances of the video encoder 114 encode the slices to some extent independently. A single video encoder 114 may process the slices sequentially, or some intermediate degree of parallelism may be achieved. Typically, the segmentation of the frame into slices (contiguous portions) is aligned with the boundaries of the segmentation of the frame into regions called "sub-pictures" or blocks, etc.
[0182] From step 1320, the method 1300 continues to step 1330 of encoding a slice header. At step 1330, the entropy encoder 338 encodes the slice header 1118 in the bitstream 115. Fig.14 An example implementation of step 1330 is provided.
[0183] From step 1330, the method 1300 continues to step 1340 of partitioning the slice into CTUs. In performing step 1340, the video encoder 114 partitions the slice 1116 into a sequence of CTUs. The slice boundaries are aligned with the CTU boundaries, and the CTUs in the slice are ordered according to a CTU scan order (usually a raster scan order). The partitioning of the slice into CTUs determines which portion of the frame data 113 will be processed by the video encoder 113 when encoding the current slice.
[0184] From step 1340, the method 1300 continues to a determine coding tree step 1350. At step 1350, the video encoder 114 determines a coding tree for the currently selected CTU in the slice. The method 1300 starts with the first CTU in the slice 1116 on the first invocation of step 1350 and proceeds to subsequent CTUs in the slice 1116 on subsequent invocations. Various combinations of quadtree, binary, and ternary splits are generated and tested by the block partitioner 310 when determining the coding tree for the CTU.
[0185] From step 1350, method 1300 continues to step 1360 of determining coding units. At step 1360, the video encoder 114 performs to determine the "best" encoding of the CU obtained from the various coding trees under evaluation using known methods. Determining the best encoding involves determining a prediction mode (e.g., intra-frame prediction with a particular mode or inter-frame prediction with motion vectors), transform selection (primary transform type and optional secondary transform type). If the primary transform type of the luma TB is determined to be DCT-2 or any quantized primary transform coefficients that are not subjected to a forward secondary transform are valid, the secondary transform index of the luma TB may indicate the application of a secondary transform. Otherwise, the secondary transform index of the luma indicates bypassing the secondary transform. For the luma channel, the primary transform type is determined to be one of the MTS options for the chroma channels, DCT-2, or transform skip, with DCT-2 being an available transform type. Reference Fig.19A and 19B Determination of the secondary transform type is further described. Determining the encoding may also include determining a quantization parameter that may change the QP, i.e., a quantization parameter at a quantization group boundary. When determining each coding unit, an optimal coding tree is also determined in a joint manner. When encoding the coding unit using intra prediction, a luma intra prediction mode and a chroma intra prediction are determined.
[0186] The step 1360 of determining the coding units may prohibit testing the application of the secondary transform when there are no "AC" (coefficients in positions other than the upper left position of the transform block) residual coefficients in the main domain residual resulting from the application of the DCT-2 main transform. If the application of the secondary transform is tested for a transform block that only includes a DC coefficient (the last position indicates that only the upper left coefficient of the transform block is valid), an increase in coding efficiency is seen. The prohibition of testing the secondary transform when only the DC main coefficient is present spans the blocks for which the secondary transform index applies, i.e., Y, Cb and Cr of the common tree when encoding a single index (only the Y channel when the Cb and Cr blocks are two samples wide or high). Although the residual with only DC coefficients is low in coding cost compared to the residual with at least one AC coefficient, applying the secondary transform to the residual with only valid DC coefficients also results in a further reduction in the amplitude of the final encoded DC coefficient. Even after further quantization and / or rounding operations before encoding, the amplitudes of the other (AC) coefficients after the secondary transform are insufficient to obtain (one or more than one) valid encoded residual coefficients in the bitstream. In a common or separate tree coding tree, assuming that there is at least one valid primary coefficient, the video encoder 114 tests the selection of a non-zero secondary transform index value (i.e., application of the secondary transform) even if there is only (one or more than one) DC coefficient of the corresponding transform block within the application range of the secondary transform index.
[0187] From step 1360, method 1300 proceeds to step 1370 of encoding the coding unit. At step 1370, video encoder 114 encodes the determined coding unit of step 1360 in bitstream 115. Fig.15 An example of how to encode a coding unit is described in more detail.
[0188] From step 1370, method 1300 continues to step 1380 of the last coding unit test. At step 1380, processor 205 tests whether the current coding unit is the last coding unit in the CTU. If not ("No" at step 1380), control in processor 205 proceeds to step 1360 of determining the coding unit. Otherwise, if the current coding unit is the last coding unit ("Yes" at step 1380), control in processor 205 proceeds to step 1390 of the last CTU test.
[0189] At step 1390 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU in the slice 1116. If it is not the last CTU in the slice 1116, control in the processor 205 returns to step 1350 of determining the coding tree. Otherwise, if the current CTU is the last ("yes" at step 1390), control in the processor proceeds to step 13100 of the last slice test.
[0190] At the last slice test step 13100, the processor 205 tests whether the current slice being encoded is the last slice in the frame. If it is not the last slice ("No" at step 13100), control in the processor 205 proceeds to the step 1330 of encoding the slice header. Otherwise, if the current slice is the last and all slices (continuous parts) have been encoded ("Yes" at step 13100), the method 1300 terminates.
[0191] Fig.14 A method 1400 is shown for encoding the slice header 1118 in the bitstream 115 as implemented at step 1330. The method 1400 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, the method 1400 may be performed by the video encoder 114 under execution of the processor 205. Thus, the method 1400 may be stored on a computer-readable storage medium and / or in the memory 206.
[0192] Method 1400 begins at step 1410 of a partition constraint override enable test. At step 1410, processor 205 tests whether a partition constraint override enable flag, as encoded in SPS 1110, indicates that partition constraints can be overridden at the stripe level. If the partition constraints can be overridden at the stripe level ("yes" at step 1410), control in processor 205 proceeds to step 1420 of determining the partition constraints. Otherwise, if the partition constraints cannot be overridden at the stripe level ("no" at step 1410), control in processor 205 proceeds to step 1480 of encoding other parameters.
[0193] At step 1420 of determining partition constraints, the processor 205 determines the partition constraints (e.g., maximum MTT split depth) that are appropriate for the current slice 1116. In one example, the frame data 310 contains a projection of a 360-degree view of a scene mapped into a 2D box and partitioned into several sub-pictures. Depending on the selected viewport, some slices may require higher fidelity and other slices may require lower fidelity. The partition constraints for a given slice may be set based on the fidelity requirements of the portion of the frame data 310 encoded by the slice (e.g., as per step 1340). In the case where lower fidelity is considered acceptable, a shallower coding tree with a larger CU is acceptable, so the maximum MTT depth may be set to a lower value. Accordingly, a subdivision level 1136 signaled by the flag cu_qp_delta_subdiv is determined at least in the range generated by the determined maximum MTT depth 1134. A corresponding chroma subdivision level is also determined and signaled.
[0194] From step 1420, method 1400 continues to step 1430 of encoding a partition constraint override flag. At step 1430, entropy encoder 338 encodes a flag in bitstream 115 indicating whether the partition constraints as signaled in SPS 1110 are to be overwritten for slice 1116. If partition constraints specific to the current slice were derived at step 1420, the flag value will indicate the use of the partition constraint override functionality. If the constraints determined at step 1420 match the constraints already encoded in SPS 1110, then there is no need to overwrite the partition constraints since there are no changes to signal, and the flag value is encoded accordingly.
[0195] From step 1430, method 1400 continues to step 1440 of partition constraint overwrite testing. At step 1440, processor 205 tests the flag value encoded at step 1430. If the flag indicates that the partition constraint is to be overwritten ("yes" at step 1440), control in processor 205 proceeds to step 1450 of encoding the stripe partition constraint. Otherwise, if the partition constraint is not to be overwritten ("no" at step 1440), control in processor 205 proceeds to step 1480 of encoding other parameters.
[0196] From step 1440, the method 1400 continues to step 1450 of encoding the slice partition constraints. In performing step 1450, the entropy encoder 338 encodes the determined partition constraints for the slice in the bitstream 115. The partition constraints for the slice include "slice_max_mtt_hierarchy_depth_luma", from which the MaxMttDepthY 1134 is derived.
[0197] From step 1450, method 1400 proceeds to step 1460 of encoding the QP subdivision level. At step 1460, entropy encoder 338 encodes the subdivision level of luma CB using the "cu_qp_delta_subdiv" syntax element, as shown in FIG. Fig.11 described.
[0198] From step 1460, the method 1400 proceeds to step 1470 of encoding the chroma QP subdivision level. At step 1470, the entropy encoder 338 encodes the signaled subdivision level for the CU chroma QP offset using the "cu_chroma_qp_offset_subdiv" syntax element, as described in reference to Fig.11 described.
[0199] Steps 1460 and 1470 operate to encode the overall QP subdivision level for the slice (contiguous portion) of the frame. The overall subdivision level includes both the subdivision level for the luma coding units of the slice and the subdivision level for the chroma coding units of the slice. For example, since separate coding trees are used for luma and chroma in an I slice, the chroma and luma subdivision levels may be different.
[0200] From step 1470, the method 1400 continues to step 1480 of encoding other parameters. At step 1480, the entropy encoder 338 encodes other parameters in the slice header 1118, such as parameters required to control specific tools such as deblocking, adaptive loop filtering, optionally selecting a scaling list from a previously signaled scaling list (for non-uniform application of quantization parameters to transform blocks), etc. The method 1400 terminates when performing step 1480.
[0201] Fig.15 A method 1500 for encoding a coding unit in a bitstream 115 is shown, which corresponds to Fig.13 Step 1370 of . The method 1500 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. In addition, the method 1500 may be performed by the video encoder 114 under the execution of the processor 205. Therefore, the method 1500 may be stored on a computer-readable storage medium and / or in the memory 206.
[0202] The method 1500 begins at step 1510 of encoding a prediction mode. At step 1510, the entropy encoder 338 encodes the prediction mode for the coding unit determined at step 1360 in the bitstream 115. The "pred_mode" syntax element is encoded to distinguish the use of intra prediction, inter prediction, or other prediction modes for the coding unit. If intra prediction is used for the coding unit, the luma intra prediction mode is encoded and the chroma intra prediction mode is encoded. If inter prediction is used for the coding unit, a "merge index" may be encoded to select a motion vector from a neighboring coding unit for use by the coding unit, and a motion vector delta may be encoded to introduce an offset to a motion vector derived from a spatial neighboring block. The main transform type is encoded to select between using DCT-2 horizontally and vertically, using transform skip horizontally and vertically, or using a combination of DCT-8 and DST-7 horizontally and vertically for the luma TB of the coding unit.
[0203] From step 1510, method 1500 continues to an encoded residual test step 1520. At step 1520, processor 205 determines whether the residual needs to be encoded for the coding unit. If there are any valid residual coefficients to be encoded for the coding unit ("yes" at step 1520), control in processor 205 proceeds to a new QG test step 1530. Otherwise, if there are no valid residual coefficients for encoding ("no" at step 1520), method 1500 terminates because all information required to decode the coding unit is present in bitstream 115.
[0204] At a new QG test step 1530, the processor 205 determines whether the coding unit corresponds to a new quantization group. If the coding unit corresponds to a new quantization group ("yes" at step 1530), control in the processor 205 proceeds to a step 1540 of encoding delta QP. Otherwise, if the coding unit is not associated with a new quantization group ("no" at step 1530), control in the processor 205 proceeds to a step 1550 of performing a main transform. As each coding unit is encoded, the nodes of the coding tree of the CTU are traversed at step 1530. When any child node of the current node has a subdivision level less than or equal to the subdivision level 1136 of the current slice as determined by "cu_qp_delta_subdiv", a new quantization group begins in the CTU region corresponding to the node, and step 1530 returns "yes". The first CU in the quantization group that includes a coded residual will also include a coded delta QP, thereby signaling any changes in the quantization parameters applicable to the residual coefficients in the quantization group.
[0205] At step 1540 of encoding delta QP, the entropy encoder 338 encodes the delta QP in the bitstream 115. The delta QP encodes the difference between the predicted QP and the expected QP used in the current quantization group. The predicted QP is derived by averaging the QPs of the adjacent earlier (above and left) quantization groups. When the subdivision level is low, the quantization group is larger and the delta QP is encoded less frequently. The less frequent encoding of the delta QP results in lower overhead for signaling changes in QP, but also results in less flexibility in rate control. The selection of quantization parameters for the various quantization groups is performed by the QP controller module 390, which typically implements a rate control algorithm to target a specific bit rate for the bitstream 115, which is somewhat independent of changes in the statistics of the underlying frame data 113. From step 1540, the method 1500 continues to step 1550 of performing a main transform.
[0206] At the perform main transform step 1550, the forward main transform module 326 performs the main transform according to the main transform type of the coding unit, thereby obtaining the main transform coefficients 328. The main transform is performed on each color channel, first on the luma channel (Y), and then on the Cb TB and Cr TB in subsequent calls to step 1550 for the current TU. For the luma channel, the main transform type (DCT-2, transform skip, MTS option) is performed, and for the chroma channels, DCT-2 is performed.
[0207] From step 1550, method 1500 proceeds to step 1560 of quantizing the primary transform coefficients. At step 1560, quantizer module 334 quantizes primary transform coefficients 328 according to quantization parameters 392 to produce quantized primary transform coefficients 332. Transform coefficients 328 are encoded using delta QP (when present).
[0208] From step 1560, the method 1500 continues to step 1570 of performing a secondary transform. At step 1570, the secondary transform module 330 performs a secondary transform on the quantized primary transform coefficients 332 according to the secondary transform index 388 of the current transform block to produce secondary transform coefficients 336. Although the secondary transform is performed after quantization, the primary transform coefficients 328 can maintain higher precision than the final expected quantizer step size of the quantization parameter 392, for example, the amplitude can be 16 times the amplitude directly resulting from the application of the quantization parameter 392, that is, four additional bits of precision will be retained. Retaining additional bits of precision in the quantized primary transform coefficients 332 allows the secondary transform module 330 to operate on the coefficients in the primary coefficient domain with higher precision. After applying the secondary transform, the final scaling (e.g., right shifting by four bits) at step 1560 results in quantization to the expected quantizer step size of the quantization parameter 392. The "scaling list" is applied to the primary transform coefficients (which correspond to the well-known transform basis functions (DCT-2, DCT-8, DST-7)) instead of operating on the secondary transform coefficients generated by the trained secondary transform kernel. When the secondary transform index 388 of the transform block indicates that the secondary transform is not applied (index value equal to zero), the secondary transform is bypassed. That is, the primary transform coefficients 332 are propagated unchanged through the secondary transform module 330 to become secondary transform coefficients 336. The luma secondary transform index is used in conjunction with the luma intra prediction mode to select the secondary transform kernel to apply to the luma TB. The chroma secondary transform index is used in conjunction with the chroma intra prediction mode to select the secondary transform kernel to apply to the chroma TB.
[0209] From step 1570, method 1500 continues to step 1580 of encoding the last position. At step 1580, entropy encoder 338 encodes the position of the last significant coefficient in secondary transform coefficients 336 of the current transform block in bitstream 115. At the first call of step 1580, luma TB is considered, and subsequent calls consider Cb TB and then Cr TB.
[0210] In an arrangement where the secondary transform index 388 is encoded immediately after the last position, the method 1500 continues to step 1590 of encoding the LFNST index. If the secondary transform index is not inferred to be zero based on the last position encoded at step 1580, then at step 1590, the entropy encoder 338 encodes the secondary transform index 338 in the bitstream 115 as "lfnst_index" using a truncated unary codeword. Each CU has one luma TB, allowing step 1590 to be performed on luma blocks, and when the "joint" coding mode is used for chroma, a single chroma TB is encoded, so step 1590 can be performed on chroma. Knowing the secondary transform index before decoding each residual coefficient enables the secondary transform to be applied coefficient by coefficient as the coefficients are decoded, for example using multiply-add logic. From step 1590, the method 1500 continues to step 15100 of encoding the sub-blocks.
[0211] If the secondary transform index 388 is not encoded immediately after the last position, the method 1500 proceeds from step 1580 to an encoding sub-block step 15100. At the encoding sub-block step 15100, the residual coefficients (336) for the current transform block are encoded as a series of sub-blocks in the bitstream 115. The residual coefficients are encoded proceeding from the sub-block containing the last significant coefficient position back to the sub-block containing the DC residual coefficient.
[0212] From step 15100, method 1500 continues to step 15110 of last TB test. At this step, processor 205 tests whether the current transform block is the last transform block in progress on the color channels (i.e., Y, Cb, and Cr). If the transform block just encoded is for the Cr TB ("yes" at step 15110), control in processor 205 proceeds to step 15120 of encoding the luma LFNST index. Otherwise, if the current TB is not the last ("no" at 15110), control in processor 205 returns to step 1550 of performing the main transform and the next TB (select Cb or Cr).
[0213] Steps 1550 to 15110 are described with respect to an example in which the prediction mode is intra prediction and a common coding tree structure of DCT-2 is used. In addition to the common coding tree structure using a known method, operations such as steps of performing a main transform (1550), quantizing a main transform coefficient (1560), and encoding a last position (1590) may be implemented for an inter prediction mode or an intra prediction mode. Steps 1510 to 1540 may be implemented regardless of the prediction mode or the coding tree structure.
[0214] From step 15110, the method 1500 proceeds to step 15120 of encoding the luma LFNST index. At step 15120, if the secondary transform index applied to the luma TB is not inferred to be zero (no secondary transform is applied), the entropy encoder 338 encodes it in the bitstream 115. If the last significant position of the luma TB indicates a valid main-only residual coefficient or if a main transform other than DCT-2 is performed, the luma secondary transform index is inferred to be zero. In addition, the secondary transform index applied to the luma TB is encoded in the bitstream only for coding units using intra prediction and a common coding tree structure. The secondary transform index applied to the luma TB is encoded using flag 1220 (or flag 1230 for joint CbCr mode).
[0215] From step 15120, the method 1500 continues to step 15130 of encoding the chroma LFNST index. At step 1530, if the secondary transform index applied to the chroma TB is not inferred to be zero (no secondary transform is applied), the chroma secondary transform index is encoded in the bitstream 115 by the entropy encoder 338. If the last significant position of any chroma TB indicates a valid only main residual coefficient, the chroma secondary transform index is inferred to be zero. The method 1500 terminates after performing step 15130, where control in the processor 205 returns to the method 1300. The secondary transform index applied to the chroma TB is encoded in the bitstream only for coding units using intra prediction and a common coding tree structure. The secondary transform index applied to the chroma TB is encoded using flag 1221 (or flag 1230 for joint CbCr mode).
[0216] Fig.16 A method 1600 of decoding a frame from a bitstream as a sequence of coding units arranged into slices is shown. The method 1600 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. In addition, the method 1600 may be performed by the video decoder 134 under execution of the processor 205. Thus, the method 1600 may be stored on a computer-readable storage medium and / or in the memory 206.
[0217] The method 1600 decodes a bitstream encoded using the method 1300, in which the partition constraints and quantization group definitions can vary from one slice to another, which is believed to be beneficial for rate control purposes when encoding various portions (slices) of the bitstream 115. Not only can the quantization group subdivision level vary from one slice to another, but the application of the secondary transform is independently controllable for luma and chroma.
[0218] The method 1600 begins with a decode SPS / PPS step 1610. In performing step 1610, the video decoder 134 decodes the SPS 1110 and PPS 1112 from the bitstream 133 as a sequence of fixed and variable length parameters. A partition_constraints_override_enabled_flag is decoded as part of the SPS 1110, indicating whether the partition constraints can be overridden in the slice header (e.g., 1118) of the corresponding slice (e.g., 1116). The default (i.e., as signaled in the SPS 1110 and used in the slice without subsequent overwrite) partition constraint parameters 1130 are also decoded by the video decoder 134 as part of the SPS 1110.
[0219] From step 1610, method 1600 continues to step 1620 of determining slice boundaries. In the execution of step 1620, processor 205 determines the location of the slice in the current access unit in bitstream 133. Typically, slices are identified by determining NAL unit boundaries (by detecting "start codes") and reading the NAL unit header including the "NAL unit type" for each NAL unit. A specific NAL unit type identifies the slice type, such as "I slice", "P slice", and "B slice", etc. After identifying the slice boundaries, application 233 can distribute the subsequent steps of method 1600 on different processors, for example in a multi-processor architecture, for parallel decoding. Each processor in a multi-processor system can decode different slices to obtain higher decoding throughput.
[0220] From step 1610, the method 1600 proceeds to step 1630 of decoding the slice header. At step 1630, the entropy decoder 420 decodes the slice header 1118 from the bitstream 133. Fig.17 An example method of decoding the slice header 1118 from the bitstream 133 as implemented at step 1630 is described.
[0221] From step 1630, the method 1600 continues to step 1640 of partitioning the slice into CTUs. At step 1640, the video decoder 134 partitions the slice 1116 into a sequence of CTUs. The slice boundaries are aligned with the CTU boundaries, and the CTUs in the slice are ordered according to a CTU scanning order. The CTU scanning order is typically a raster scanning order. Partitioning the slice into CTUs determines which portion of the frame data 113 will be processed by the video decoder 134 when decoding the current slice.
[0222] From step 1640, method 1600 proceeds to step 1650 of decoding the coding tree. In the execution of step 1650, video decoder 133 decodes the coding tree of the current CTU in the slice from bitstream 133 starting from the first CTU in slice 1116 when step 1650 is first called. Figure 6 The split flag is decoded to decode the coding tree of the CTU. In a subsequent iteration of step 1650 for the CTU, a subsequent CTU in the slice 1116 is decoded. If the coding tree is encoded using intra prediction mode and a common coding tree structure, the coding unit has a primary color channel (luminance or Y) and at least one secondary color channel (chrominance, Cb and Cr or CbCr). In this case, decoding the coding tree involves decoding the coding unit including the primary color channel and the at least one secondary color channel according to the split flag of the coding tree unit.
[0223] From step 1660, method 1600 proceeds to step 1670 of decoding the coding unit. At step 1670, video decoder 134 decodes the coding unit from bitstream 133. Fig.18 An example method of decoding a coding unit as implemented at step 1670 is described.
[0224] From step 1610, method 1600 continues to step 1680 of the last coding unit test. At step 1680, processor 205 tests whether the current coding unit is the last coding unit in the CTU. If it is not the last coding unit ("No" at step 1680), control in processor 205 returns to step 1670 of decoding coding units to decode the next coding unit of the coding tree unit. If the current coding unit is the last coding unit ("Yes" at step 1680), control in processor 205 proceeds to step 1690 of the last CTU test.
[0225] At step 1690 of the last CTU test, the processor 205 tests whether the current CTU is the last CTU in the slice 1116. If it is not the last CTU in the slice (“No” at step 1690), control in the processor 205 returns to step 1650 of decoding the coding tree to decode the next coding tree unit of the slice 1116. If the current CTU is the last CTU of the slice 1116 (“Yes” at step 1690), control in the processor 205 proceeds to step 16100 of the last slice test.
[0226] At the last slice test step 16100, the processor 205 tests whether the current slice being decoded is the last slice in the frame. If it is not the last slice in the frame ("No" at step 16100), control in the processor 205 returns to the decode slice header step 1630, and step 1630 operates to decode the next slice in the frame (e.g., Fig.11 If the current slice is the last slice in the frame ("yes" at step 1600), the method 1600 terminates.
[0227] As about Figure 1 As described in the device 130 in FIG. 1 , the operation method 1600 for multiple coding units operates to generate an image frame.
[0228] Fig.17 A method 1700 for decoding a slice header into a bitstream as implemented at step 1630 is shown. The method 1700 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, the method 1700 may be performed by the video decoder 134 under execution of the processor 205. Thus, the method 1700 may be stored on a computer-readable storage medium and / or in the memory 206.
[0229] Similar to method 1500, method 1700 is performed for the current slice or contiguous portion (1116) in a frame (e.g., frame 1101). Method 1700 begins at a partition constraint override enable test step 1710. At step 1710, processor 205 tests whether a partition constraint override enable flag, as decoded from SPS 1110, indicates that the partition constraint can be overwritten at the slice level. If the partition constraint can be overwritten at the slice level ("yes" at step 1710), control in processor 205 proceeds to step 1720 of decoding the partition constraint override flag. Otherwise, if the partition constraint override enable flag indicates that the constraint cannot be overwritten at the slice level ("no" at step 1710), control in processor 205 proceeds to step 1770 of decoding other parameters.
[0230] At decode partition constraint override flag step 1720, the entropy decoder 420 decodes the partition constraint override flag from the bitstream 133. The decoded flag indicates whether the partition constraints as signaled in the SPS 1110 are to be overwritten for the current slice 1116.
[0231] From step 1720, method 1700 continues to step 1730 of partition constraint overwrite testing. In executing step 1730, processor 205 tests the flag value decoded at step 1720. If the decoded flag indicates that the partition constraint will be overwritten ("yes" at step 1730), control in processor 205 proceeds to step 1740 of decoding the stripe partition constraint. Otherwise, if the decoded flag indicates that the partition constraint will not be overwritten ("no" at step 1730), control in processor 205 proceeds to step 1770 of decoding other parameters.
[0232] At a decode slice partition constraints step 1740, the entropy decoder 420 decodes the determined partition constraints for the slice from the bitstream 133. The partition constraints for the slice include "slice_max_mtt_hierarchy_depth_luma", from which the MaxMttDepthY 1134 is derived.
[0233] From step 1740, method 1700 proceeds to step 1750 of decoding the QP subdivision level. At step 1720, entropy decoder 420 uses the QP subdivision level as described in reference Fig.11 The "cu_qp_delta_subdiv" syntax element decodes the subdivision level of the luma CB.
[0234] From step 1750, method 1700 proceeds to step 1760 of decoding the chroma QP subdivision level. At step 1760, entropy decoder 420 uses the Fig.11 The "cu_chroma_qp_offset_subdiv" syntax element is used to decode the subdivision level used to signal the CU chroma QP offset.
[0235] Steps 1750 and 1760 operate to determine the subdivision level for a particular continuous portion (slice) of the bitstream. Repeated iterations between steps 1630 and 16100 operate to determine the subdivision level for each continuous portion (slice) in the bitstream. As described below, each subdivision level applies to the coding units of the corresponding slice (continuous portion).
[0236] From step 1760, the method 1700 continues to step 1770 of decoding other parameters. At step 1770, the entropy decoder 420 decodes other parameters from the slice header 1118, such as parameters required to control specific tools such as deblocking, adaptive loop filter, optionally selecting a scaling list from a previously signaled scaling list (for non-uniformly applying quantization parameters to transform blocks), etc. The method 1700 terminates when performing step 1770.
[0237] Fig.18 A method 1800 for decoding a coding unit from a bitstream is shown. The method 1800 may be embodied by a device such as a configured FPGA, ASIC, or ASSP. Additionally, the method 1800 may be performed by the video decoder 134 under execution of the processor 205. Thus, the method 1800 may be stored on a computer-readable storage medium and / or in the memory 206.
[0238] Method 1800 is implemented for a current coding unit of a current CTU (e.g., CTU0 of slice 1116). Method 1800 begins at step 1810 of decoding a prediction mode. At step 1800, entropy decoder 420 decodes a prediction mode from bitstream 133 such as Fig.13 The prediction mode of the coding unit determined at step 1360. The "pred_mode" syntax element is decoded at step 1810 to distinguish the use of intra prediction, inter prediction or other prediction modes for the coding unit.
[0239] If intra prediction is used for the coding unit, the luma intra prediction mode and the chroma intra prediction mode are also decoded at step 1810. If inter prediction is used for the coding unit, the "merge index" can also be decoded at step 1810 to determine the motion vector from the neighboring coding unit for use by the coding unit, and the motion vector delta can be decoded to introduce an offset to the motion vector derived from the spatial neighboring block. The main transform type is also decoded at step 1810 to select between using DCT-2 horizontally and vertically, using transform skip horizontally and vertically, or using a combination of DCT-8 and DST-7 horizontally and vertically for the luma TB of the coding unit.
[0240] Method 1800 continues from step 1810 to step 1820 of encoding residual test. In the execution of step 1820, the processor 205 determines whether the residual needs to be decoded for the coding unit by decoding the "root coding block flag" of the coding unit using the entropy decoder 420. If there are any valid residual coefficients to be decoded for the coding unit (step 1820 is "yes"), then control in the processor 205 proceeds to step 1830 of the new QG test. Otherwise, if there are no residual coefficients to be decoded (step 1820 is "no"), then method 1800 terminates because all information required to decode the coding unit has been obtained in the bitstream 115. When method 1800 terminates, subsequent steps such as PB generation, application of in-loop filtering, etc. are performed to produce decoded samples, as shown in reference. Figure 4 described.
[0241] At the new QG test step 1830, the processor 205 determines whether the coding unit corresponds to a new quantization group. If the coding unit corresponds to a new quantization group ("yes" at step 1830), control in the processor 205 proceeds to a step 1840 of decoding the delta QP. Otherwise, if the coding unit does not correspond to a new quantization group ("no" at step 1830), control in the processor 205 proceeds to a step 1850 of decoding the last position. The new quantization group is related to the subdivision level of the current mode or coding unit. When decoding each coding unit, the nodes of the coding tree of the CTU are traversed. When any child node of the current node has a subdivision level less than or equal to the subdivision level 1136 of the current slice (i.e., as determined from "cu_qp_delta_subdiv"), a new quantization group starts in the region of the CTU corresponding to the node. The first CU in the quantization group that includes coded residual coefficients will also include a coded delta QP, thereby signaling any changes in the quantization parameters applicable to the residual coefficients in the quantization group. Effectively, a single (at most one) quantization parameter increment is decoded for each region (quantization group). Figures 8A to 8C As described, the respective regions (quantization groups) are based on the decomposition of the coding tree units of the respective slices and the corresponding subdivision levels (e.g., as encoded at steps 1460 and 1470). In other words, the respective regions or quantization groups are based on a comparison of the subdivision level associated with the coding unit with the subdivision level determined for the corresponding continuous portion.
[0242] At decode delta QP step 1840, the entropy decoder 420 decodes delta QP from the bitstream 133. The delta QP encodes the difference between the predicted QP and the expected QP used in the current quantization group. The predicted QP is derived by averaging the QPs of the adjacent (above and left) quantization groups.
[0243] From step 1840, method 1800 continues to step 1850 of decoding the last position. In performing step 1850, entropy decoder 420 decodes the position of the last significant coefficient in the secondary transform coefficients 424 of the current transform block from bitstream 133. In the first call of step 1850, this step is performed for the luma TB. In subsequent calls of step 1850 for the current CU, this step is performed for the Cb TB. If the last position indicates that the significant coefficient is outside the set of secondary transform coefficients of the luma block or chroma block (i.e., outside 928 or 966), the secondary transform index of the luma or chroma channel, respectively, is inferred to be zero. This step is implemented for the Cr TB in an iteration after the iteration for Cb.
[0244] As about Fig.15 As described in step 1590 of , in some arrangements, the secondary transform index is encoded immediately after the last significant coefficient position of the coding unit. When decoding the same coding unit, if the secondary transform index 470 is not inferred to be zero based on the location of the last position of the TB decoded in step 1840, the secondary transform index 470 is decoded immediately after the position of the last significant residual coefficient of the coding unit is decoded. In an arrangement where the secondary transform index 470 is decoded immediately after the last significant coefficient position of the coding unit, the method 1800 continues from step 1850 to step 1860 of decoding the LFNST index. When performing step 1860, the entropy decoder 420 decodes the secondary transform index 470 from the bitstream 133 as "lfnst_index" using a truncated unary codeword when all significant coefficients are subjected to a secondary inverse transform (e.g., within 928 or 966). When joint encoding of chroma TBs using a single transform block is performed, the secondary transform index 470 can be decoded for luma TBs or chroma. From step 1860, method 1800 continues to step 1870 of decoding the sub-blocks.
[0245] If the secondary transform index 470 is not decoded immediately after the last significant position of the coding unit, the method 1800 continues from step 1850 to a decoding sub-block step 1870. At step 1870, the residual coefficients (i.e., 424) of the current transform block are decoded from the bitstream 133 as a series of sub-blocks, proceeding from the sub-block containing the last significant coefficient position back to the sub-block containing the DC residual coefficient.
[0246] From step 1870, the method 1800 continues to step 1880 of the last TB test. In the execution of step 1880, the processor 205 tests whether the current transform block is the last transform block in progress on the color channels (i.e., Y, Cb, and Cr). If the just decoded (current) transform block is for the Cr TB, then in the control in the processor 205, all TBs have been decoded (step 1880 is "yes"), and the method 1800 proceeds to step 1890 of decoding the luma LFNST index. Otherwise, if the TB has not been decoded (step 1880 is "no"), the control in the processor 205 returns to step 1850 of decoding the last position. In the iteration of step 1850, the next TB (following the order of Y, Cb, Cr) is selected for decoding.
[0247] From step 1880, method 1800 proceeds to step 1890 of decoding luma LFNST indices. In execution of step 1890, if the last position of the luma TB is within the set of coefficients that are subject to a secondary inverse transform (e.g., 928 or 966) and the luma TB is using DCT-2 as the primary transform horizontally and vertically, the secondary transform indices 470 to be applied to the luma TB are decoded from the bitstream 133 by the entropy decoder 420. If the last significant position of the luma TB indicates that there are significant primary coefficients outside the set of coefficients that are subject to a secondary inverse transform (e.g., outside 928 or 966), the luma secondary transform indices are inferred to be zero (no secondary transform is applied). The secondary transform indices decoded at step 1890 are in Fig.12 Indicated as 1220 (or 1230 in joint CbCr mode).
[0248] From step 1890, method 1800 proceeds to step 1895 of decoding chroma LFNST indices. At step 1895, if the last position of the respective chroma TB is within the set of coefficients that are subject to secondary inverse transform (e.g., 928 or 966), the secondary transform indices 470 to be applied to the chroma TB are decoded from the bitstream 133 by entropy decoder 420. If the last significant position of any chroma TB indicates that there are significant primary coefficients outside the set of coefficients that are subject to secondary inverse transform (e.g., outside 928 or 966), the chroma secondary transform indices are inferred to be zero (no secondary transform is applied). The secondary transform indices decoded at step 1895 are in Fig.12 Indicated as 1221 in (or 1230 in joint CbCr mode). When decoding separate indices for luma and chroma, separate arithmetic contexts for each truncated unary codeword may be used or the context may be shared so that the respective nth bins in the luma and chroma truncated unary codewords share the same context.
[0249] Effectively, steps 1890 and 1895 involve respectively decoding a first index (such as 1220, etc.) to select a kernel for a luma (primary color) channel, and decoding a second index (such as 1221, etc.) to select a kernel for at least one chroma (secondary color channel).
[0250] From step 1895, the method 1800 proceeds to step 18100 of performing an inverse secondary transform. At this step, the inverse secondary transform module 436 performs an inverse secondary transform on the decoded residual transform coefficient 424 according to the secondary transform index 470 of the current transform block to generate a secondary transform coefficient 432. The secondary transform index decoded at step 1890 is applied to the luma TB, and the secondary transform index decoded at step 1895 is applied to the chroma TB. The kernel selection for luma and chroma also depends on the luma intra prediction mode and the chroma intra prediction mode, respectively, which are each decoded at step 1810. Step 18100 selects a kernel according to the LFNST index of luma, and selects a kernel according to the LFNST index of chroma.
[0251] From step 18100, the method 1800 proceeds to step 18110 of inverse quantizing the primary transform coefficients. At step 18110, the inverse quantizer module 428 inverse quantizes the secondary transform coefficients 432 according to the quantization parameters 474 to produce inverse quantized primary transform coefficients 440. If the delta QP is decoded at step 1840, the entropy decoder 420 determines the quantization parameter according to the delta QP of the quantization group (region) and the quantization parameter of the earlier coding unit of the image frame. As described above, the earlier coding unit generally refers to the adjacent upper left coding unit.
[0252] From step 1870, method 1800 proceeds to step 18120 of performing a main transform. At step 18120, the inverse main transform module 444 performs an inverse main transform according to the main transform type of the coding unit, so that the transform coefficients 440 are converted into residual samples 448 in the spatial domain. The inverse main transform is performed on each color channel, first on the luma channel (Y), and then on the Cb and Cr TBs in a subsequent call to step 1650 for the current TU. Steps 18100 to 18120 effectively operate to decode the current coding unit by applying the kernel selected at step 1890 according to the LFNST index of luma to the decoded residual coefficients of the luma channel, and applying the kernel selected at step 1890 according to the LFNST index of chroma to the decoded residual coefficients of at least one chroma channel.
[0253] Method 1800 terminates after executing step 18120, where control in processor 205 returns to method 1600.
[0254] Steps 1850 to 18120 are described with respect to an example of a common coding tree structure in which the prediction mode is intra prediction and the transform is DCT-2. For example, a secondary transform index (1890) applied to a luma TB is decoded from a bitstream only for a coding unit using intra prediction and a common coding tree structure. Similarly, a secondary transform index (1895) applied to a chroma TB is decoded from a bitstream only for a coding unit using intra prediction and a common coding tree structure. In addition to a common coding tree structure using a known method, operations such as decoding subblocks (1870), inverse quantizing main transform coefficients (18110), and performing main transforms can be implemented for an inter prediction mode or for an intra prediction mode. Regardless of the prediction mode or structure, steps 1810 to 1840 are performed in the described manner.
[0255] Once method 1800 terminates, subsequent steps for decoding the coding unit (including generating intra-frame prediction samples 480 by module 476, summing the decoded residual samples 448 with the prediction block 452 by module 450, and applying the in-loop filter module 488 to produce filtered samples 492) are performed and output as frame data 135.
[0256] Fig.19A and 19B The rules for applying or bypassing the secondary transform on the luma and chroma channels are shown. Fig.19A Table 1900 illustrating conditions for applying secondary transforms in luma and chroma channels in a CU generated by a common coding tree is shown.
[0257] If the last significant coefficient position of the luma TB indicates a decoded significant coefficient that is not generated by a forward secondary transform and is therefore not subjected to an inverse secondary transform, then condition 1901 exists. If the last significant coefficient position of the luma TB indicates a decoded significant coefficient that is indeed generated by a forward secondary transform and is therefore subjected to an inverse secondary transform, then condition 1902 exists. In addition, for the luma channel, the main transform type needs to be DCT-2 in order for condition 1902 to exist, otherwise condition 1901 exists.
[0258] If the last significant coefficient position of one or both chroma TBs indicates a decoded significant coefficient that is not generated by a forward secondary transform and is therefore not subject to an inverse secondary transform, condition 1910 exists. If the last significant coefficient position of one or both chroma TBs indicates a decoded significant coefficient that is indeed generated by a forward secondary transform and is therefore subject to an inverse secondary transform, condition 1911 exists. In addition, the width and height of the chroma block need to be at least four samples (e.g., chroma subsampling when using a 4:2:0 or 4:2:2 chroma format may result in a width or height of two samples) for condition 1911 to exist.
[0259] If conditions 1901 and 1910 are present, no secondary transform index is signaled (independently or jointly) and the secondary transform index is not applied in luma or chroma, i.e. 1920. If conditions 1901 and 1911 are present, one secondary transform index is signaled to indicate the application of the selected kernel or bypass only for the luma channel, i.e. 1921. If conditions 1902 and 1910 are present, one secondary transform index is signaled to indicate the application of the selected kernel or bypass only for the chroma channels, i.e. 1922. If conditions 1911 and 1902 are present, the arrangement with independent signaling signals two secondary transform indexes, one for luma TB and one for chroma TB, i.e. 1923. When conditions 1902 and 1911 are present, the arrangement with a single signaled secondary transform index uses one index to control the selection of luma and chroma, although the selected kernel also depends on the luma and chroma intra prediction modes, which may be different. The ability to apply a secondary transform to either luma or chroma (ie, 1921 and 1922) results in increased coding efficiency.
[0260] Fig.19B Table 1950 shows the search options available to the video encoder 114 at step 1360. The secondary transform indices for luma (1952) and chroma (1953) are shown as 1952 and 1953, respectively. Index value 0 indicates bypassing the secondary transform, and index values 1 and 2 indicate which of the two kernels is used for the candidate set derived from luma or chroma intra prediction mode. There are nine combinations ("0,0" to "2,2") of the resulting search space, which may be subject to reference Fig.19A The constraints described are constrained. Compared with searching all allowable combinations, the simplified search of three combinations (1951) can test only combinations with the same luminance and chrominance secondary transform indices, subject to zeroing the indices of the channels where the last significant coefficient position indicates that only the main coefficients exist. For example, when condition 1921 exists, options "1,1" and "2,2" become "0,1" and "0,2" respectively (i.e., 1954). When condition 1922 exists, options "1,1" and "2,2" become "1,0" and "2,0" respectively (i.e., 1955). When condition 1920 exists, there is no need to signal the secondary transform index, and option "0,0" is used. In fact, conditions 1921 and 1922 allow options "0,1", "0,2", "1,0" and "2,0" in the common tree CU, thereby obtaining higher compression efficiency. If these options are disabled, then either of conditions 1901 or 1910 will cause condition 1920 (ie, options "1,1" and "2,2") to be disabled, causing "0,0" to be used (see 1956).
[0261] Signaling the quantization group subdivision level in the slice header provides a higher granularity control below the picture level. The higher granularity control is advantageous for applications where the encoding fidelity requirements vary from one part of the picture to another, and is particularly advantageous for applications where multiple encoders may need to operate somewhat independently to provide real-time processing capabilities. Signaling the quantization group subdivision level in the slice header is also consistent with the partition override setting and scaling list application setting in the signaling slice header.
[0262] In one arrangement of the video encoder 114 and the video decoder 134, the secondary transform index of the chroma intra prediction block is always set to zero, that is, the secondary transform is not applied to the chroma intra prediction block. In this case, the chroma secondary transform index does not need to be signaled, so steps 15130 and 1895 can be omitted, and steps 1360, 1570 and 18100 are simplified accordingly.
[0263] If a node in a coding tree in a common tree has an area of 64 luma samples, further splitting with binary or quadtree splitting will produce smaller luma CBs, such as 4×4 blocks, etc., but will not produce smaller chroma CBs. Instead, there is a single chroma CB of a size corresponding to an area of 64 luma samples, such as a 4×4 chroma CB, etc. Similarly, a coding tree node having an area of 128 luma samples and subjected to a ternary split produces a set of smaller luma CBs and one chroma CB. Each luma CB has a corresponding luma secondary transform index, and the chroma CB has a chroma secondary transform index.
[0264] When a node in the coding tree has an area of 64 and further splitting is signaled, or has an area of 128 luma samples and a ternary split is signaled, the split is applied only in the luma channel, and the resulting CBs (several luma CBs and one chroma CB for each chroma channel) are all intra-predicted or all inter-predicted. When a CU has a width or height of four luma samples and includes one CB for each of the color channels (Y, Cb, and Cr), the chroma CB of the CU has a width or height of two samples. CBs with a width or height of two samples do not utilize 16-point or 48-point LFNST kernel operations, so no secondary transform is required. For blocks with a width or height of two samples, steps 15130, 1895, 1360, 1570, and 18100 are not required.
[0265] In another arrangement of the video encoder 114 and the video decoder 134, a single secondary transform index is signaled when either or both of the luma and chroma contain non-significant residual coefficients only in the region of the corresponding TB that is only subjected to the primary transform. If the luma TB contains valid residual coefficients in the non-secondary transform region of the decoded residual (e.g., 1066, 968), or is indicated as not using DCT-2 as the primary transform, the indicated secondary transform kernel (or secondary transform bypass) is applied only to the chroma TB. If any chroma TB contains valid residual coefficients in the non-secondary transform region of the decoded residual, the indicated secondary transform kernel (or secondary transform bypass) is applied only to the luma TB. Even when it is not possible for the chroma TB, applying the secondary transform becomes possible for the luma TB, and vice versa, thereby providing an increase in coding efficiency compared to requiring that the last position of all TBs be in the secondary coefficient domain before any TB of the CU can be subjected to the secondary transform. In addition, only one secondary transform index is required for the CU in the common coding tree. When the luma primary transform is DCT-2, the secondary transform may be inferred to be disabled for chroma as well as for luma.
[0266] In another arrangement of the video encoder 114 and the video decoder 134, the secondary transform (by modules 330 and 436, respectively) is applied only to the luma TB of the CU, and not to any chroma TB of the CU. The absence of secondary transform logic for the chroma channels results in lower complexity, such as lower execution time or reduced silicon area. The absence of secondary transform logic for the chroma channels requires only one secondary transform index to be signaled, which can be signaled after the last position of the luma TB. That is, instead of steps 15120 and 1890, steps 1590 and 1860 are performed for the luma TB. In this case, steps 15130 and 1895 are omitted.
[0267] In another arrangement of the video encoder 114 and the video decoder 134, the syntax elements defining the quantization group sizes (i.e., cu_chroma_qp_offset_subdiv and cu_qp_delta_subdiv) are signaled in the PPS 1112. The range of values for the subdivision levels is defined according to the partition constraints signaled in the SPS 1110, even if the partition constraints are overridden in the slice header 1118. For example, the range of cu_qp_delta_subdiv and cu_chroma_qp_offset_subdiv is defined as 0 to 2*(log2_ctu_size_minus5+5−(MinQtLog2SizeInterY or MinQtLog2SizeIntraY)+MaxMttDepthY_SPS). The value MaxMttDepthY is derived from the SPS 1110. That is, when the current slice is an I slice, MaxMttDepthY is set equal to sps_max_mtt_hierarchy_depth_intra_slice_luma, and when the current slice is a P or B slice, MaxMttDepthY is set equal to sps_max_mtt_hierarchy_depth_inter_slice. For slices whose partition constraints are overwritten to a shallower depth than signaled in the SPS 1110, if the quantization group subdivision level determined from the PPS 1112 is higher (deeper) than the highest achievable subdivision level at the shallower coding tree depth determined from the slice header, the quantization group subdivision level of the slice is clipped to be equal to the highest achievable subdivision level of the slice. For example, for a particular slice, cu_qp_delta_subdiv and cu_chroma_qp_offset_subdiv are clipped to be within 0 to 2*(log2_ctu_size_minus5+5−(MinQtLog2SizeInterY or MinQtLog2SizeIntraY)+MaxMttDepthY_slice_header), and the clipped values are used for that slice. The value MaxMttDepthY_slice_header is derived from the slice header 1118, i.e., MaxMttDepthY_slice_header is set equal to slice_max_mtt_hierarchy_depth_luma.
[0268] In yet another arrangement of the video encoder 114 and the video decoder 134, the subdivision level is determined based on cu_chroma_qp_offset_subdiv and cu_qp_delta_subdiv decoded from the PPS 1112 to derive the luma and chroma subdivision levels. When the partition constraints decoded from the slice header 1118 result in different ranges of subdivision levels for a slice, the subdivision level applied to the slice is adjusted to maintain the same offset relative to the deepest allowed subdivision level based on the partition constraints decoded from the SPS 1110. For example, if the SPS 1110 indicates a maximum subdivision level of 4 and the PPS 1112 indicates a subdivision level of 3 and the slice header 1118 reduces the maximum value to 3, the subdivision level applied within the slice is set to 2 (maintaining an offset of 1 relative to the maximum allowed subdivision level). Adjusting the quantization group area to correspond to changes in the partition constraints for a particular slice allows the subdivision level to be signaled less frequently (i.e., at the PPS level) while providing granularity to adapt to changes in the slice-level partition constraints. Using ranges defined according to partition constraints decoded from SPS 1110 to signal the arrangement of subdivision levels in PPS 1112 (which may be adjusted later based on overridden partition constraints decoded from slice header 1118) avoids parsing dependency issues that make PPS syntax elements dependent on partition constraints done in slice header 1118.
[0269] Industrial Applicability
[0270] The described arrangement is suitable for use in the computer and data processing industries and is particularly suitable for use in digital signal processing for encoding or decoding signals such as video and image signals, thereby achieving high compression efficiency.
[0271] The arrangement described herein increases the flexibility provided to a video encoder in generating a highly compressed bitstream from incoming video data. Quantization of different regions or sub-pictures in a frame can be controlled with varying granularity and with different granularity from one region to another, thereby reducing the amount of encoded residual data. Higher granularity can be achieved accordingly when required, for example, for 360-degree images as described above.
[0272] In some arrangements, as described with respect to steps 15120 and 15130 (and steps 1890 and 1895, respectively), the application of the secondary transform may be independently controlled for luma and chroma, thereby achieving further reduction of the encoded residual data. A video decoder is described having the necessary functionality to decode a bitstream produced by such a video encoder.
[0273] The foregoing describes only some embodiments of the present invention, and modifications and / or changes may be made to the present invention without departing from the scope and spirit of the present invention, wherein the embodiments are intended to be illustrative and not restrictive.
Claims
1. A method for decoding a coding unit in a coding tree unit of an image from a bitstream, the coding unit having a luma channel and a chroma channel, the method comprising: Determining the coding unit having the luma channel and the chroma channels according to one or more split flags of the coding tree unit; decoding from the bitstream an index for selecting a non-separable transform kernel for the luma channel; selecting the inseparable transform kernel according to the index; Decoding coefficients of a luma transform block of the luma channel in the coding unit and coefficients of a chroma transform block of the chroma channels in the coding unit from the bitstream; performing a non-separable transform on the coefficients of the luminance transform block by applying the selected non-separable transform kernel to derive non-separable transformed coefficients of the luminance transform block; as well as decoding the coding unit by performing a separable transformation on coefficients of the luminance transformation block after the non-separable transformation and on coefficients of the chrominance transformation block, Wherein, when the coding tree of the luma channel in the coding tree unit is the same as the coding tree of the chroma channel in the coding tree unit, the inseparable transform can be performed only on the coefficients of the luma transform block in the coding unit, and the inseparable transform can not be performed on the coefficients of the chroma transform block in the coding unit, and the width and height of the chroma transform block are both equal to or greater than 4, and In which, when the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, a given area in the coding tree unit is split into luma coding blocks, and there is a chroma coding block corresponding to the given area, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each luma coding block in the luma coding block, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given area.
2. The method according to claim 1, wherein: The non-separable transform kernel depends on the intra prediction mode for the luma channel.
3. The method according to claim 1, wherein: The non-separable transform kernel is related to the block size of the luminance channel.
4. The method according to claim 1, wherein: When the coding tree of the luma channel in the coding tree unit is the same as the coding tree of the chroma channel in the coding tree unit, the index for selecting the inseparable transform kernel for the luma channel can be decoded for the coding unit, and the index for selecting the inseparable transform kernel for the chroma channel can not be decoded for the coding unit.
5. The method according to claim 1, wherein: The luma channel is a luma component and the chroma channels are chroma components.
6. The method according to claim 1, wherein: In the case where the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, the given area in the coding tree unit has an area of 64 luma samples, the given area is split into four luma coding blocks by quadtree splitting, each of the four luma coding blocks has a size of 4×4, and the chroma coding block corresponding to the given area has a size of 4×4, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each of the four luma coding blocks, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given area.
7. The method according to claim 1, wherein: In a case where the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, the given region is split into three luma coding blocks by ternary splitting, and there is a chroma coding block corresponding to the given region, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each of the three luma coding blocks, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given region.
8. A method of encoding a coding unit in a coding tree unit of an image in a bitstream, the coding unit having a luma channel and a chroma channel, the method comprising: Determining a coding unit having the luma channel and the chroma channel; Performing a separable transformation on coefficients of a luma transform block of the luma channel in the coding unit to derive separable transformed coefficients of the luma transform block, and performing a separable transformation on coefficients of a chroma transform block of the chroma channel in the coding unit to derive separable transformed coefficients of the chroma transform block; selecting a non-separable transform kernel for the luminance channel; performing the non-separable transform on the separably transformed coefficients of the luminance transform block by applying the selected non-separable transform kernel; as well as encoding in the bitstream an index for selecting a non-separable transform kernel for the luma channel, Wherein, when the coding tree of the luma channel in the coding tree unit is the same as the coding tree of the chroma channel in the coding tree unit, the inseparable transformation can be performed only on the separable transformed coefficients of the luma transform block in the coding unit, and the inseparable transformation is not performed on the separable transformed coefficients of the chroma transform block in the coding unit, and the width and height of the chroma transform block are both equal to or greater than 4, and In which, when the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, a given area in the coding tree unit is split into luma coding blocks, and there is a chroma coding block corresponding to the given area, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each luma coding block in the luma coding block, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given area.
9. The method according to claim 8, wherein: When the coding tree of the luma channel in the coding tree unit is the same as the coding tree of the chroma channel in the coding tree unit, an index for selecting an inseparable transform kernel for the luma channel can be encoded for the coding unit, and an index for selecting an inseparable transform kernel for the chroma channel can not be encoded for the coding unit.
10. The method according to claim 8, wherein: The luma channel is a luma component and the chroma channels are chroma components.
11. The method according to claim 8, wherein: In the case where the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, the given area in the coding tree unit has an area of 64 luma samples, the given area is split into four luma coding blocks by quadtree splitting, each of the four luma coding blocks has a size of 4×4, and the chroma coding block corresponding to the given area has a size of 4×4, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each of the four luma coding blocks, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given area.
12. The method according to claim 8, wherein: In a case where the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, the given region is split into three luma coding blocks by ternary splitting, and there is a chroma coding block corresponding to the given region, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each of the three luma coding blocks, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given region.
13. An apparatus for decoding a coding unit in a coding tree unit of an image from a bitstream, the coding unit having a luma channel and a chroma channel, the apparatus comprising: a determining unit configured to determine a coding unit having the luma channel and the chroma channel according to one or more split flags of the coding tree unit; a first decoding unit configured to decode from the bitstream an index for selecting a non-separable transform kernel for the luma channel; A selection unit configured to select the inseparable transform kernel according to the index; A second decoding unit configured to decode coefficients of a luma transform block of the luma channel in the coding unit and coefficients of a chroma transform block of the chroma channel in the coding unit from the bitstream; a performing unit configured to perform an inseparable transformation on the coefficients of the luma transform block by applying the selected inseparable transformation kernel to derive inseparable transformed coefficients of the luma transform block; as well as a third decoding unit configured to decode the coding unit by performing a separable transformation on the coefficients of the luminance transformation block after the non-separable transformation and on the coefficients of the chrominance transformation block, Wherein, when the coding tree of the luma channel in the coding tree unit is the same as the coding tree of the chroma channel in the coding tree unit, the inseparable transform can be performed only on the coefficients of the luma transform block in the coding unit, and the inseparable transform can not be performed on the coefficients of the chroma transform block in the coding unit, and the width and height of the chroma transform block are both equal to or greater than 4, and In which, when the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, a given area in the coding tree unit is split into luma coding blocks, and there is a chroma coding block corresponding to the given area, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each luma coding block in the luma coding block, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given area.
14. An apparatus for encoding a coding unit in a coding tree unit of an image in a bitstream, the coding unit having a luma channel and a chroma channel, the apparatus comprising: a determining unit configured to determine a coding unit having the luma channel and the chroma channel; a first performing unit configured to perform a separable transformation on coefficients of a luma transform block of the luma channel in the coding unit to derive separable transformed coefficients of the luma transform block, and to perform a separable transformation on coefficients of a chroma transform block of the chroma channel in the coding unit to derive separable transformed coefficients of the chroma transform block; a selection unit configured to select a non-separable transform kernel for the luma channel; a second performing unit configured to perform the non-separable transform on the separable transformed coefficients of the luma transform block by applying the selected non-separable transform kernel; as well as an encoding unit configured to encode in the bitstream an index for selecting a non-separable transform kernel for the luma channel, Wherein, when the coding tree of the luma channel in the coding tree unit is the same as the coding tree of the chroma channel in the coding tree unit, the inseparable transformation can be performed only on the separable transformed coefficients of the luma transform block in the coding unit, and the inseparable transformation is not performed on the separable transformed coefficients of the chroma transform block in the coding unit, and the width and height of the chroma transform block are both equal to or greater than 4, and In which, when the coding tree of the luma channel in the coding tree unit is separated from the coding tree of the chroma channel in the coding tree unit, a given area in the coding tree unit is split into luma coding blocks, and there is a chroma coding block corresponding to the given area, an index for selecting an inseparable transform kernel for the luma channel can exist separately for each luma coding block in the luma coding block, and an index for selecting an inseparable transform kernel for the chroma channel can exist for the chroma coding block corresponding to the given area.
15. A non-transitory computer-readable storage medium comprising computer-executable instructions, the computer-executable instructions causing a computer to perform the method according to claim 1.
16. A non-transitory computer-readable storage medium comprising computer-executable instructions, the computer-executable instructions causing a computer to perform the method according to claim 8.
17. A computer program product comprising a program which, when executed by a computer, causes the computer to perform the method according to claim 1.
18. A computer program product comprising a program which, when executed by a computer, causes the computer to perform the method according to claim 8.
Citation Information
Patent Citations
Method and apparatus of video coding
CN109076221A
Method and apparatus for encoding / decoding video signal by using graph-based separable transform
WO2017135661A1