Video Encoder, Video Decoder, and Corresponding Methods

A multi-type tree split method for video blocks, incorporating binary and ternary tree structures, enhances video compression efficiency and maintains image quality, addressing the challenge of further compressing video data in limited bandwidth scenarios.

JP7702549B2Active Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024152708
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-15
Filing Date
2024-09-04
Publication Date
2025-07-03
Estimated Expiration
2039-09-03

AI Technical Summary

Technical Problem

Existing video coding techniques struggle to achieve further compression of video data without sacrificing image quality, particularly in the context of limited network resources and increasing demands for higher video quality, necessitating improved compression and decompression methods.

Method used

The implementation of a multi-type tree split for video blocks, including binary and ternary tree splits, based on minimum and maximum allowable node sizes, to efficiently partition and encode video data, especially at block boundaries, enhancing compression efficiency.

Benefits of technology

This approach promotes efficient splitting and signaling of video blocks, leading to improved compression ratios without significant degradation in image quality, addressing the need for higher video quality in limited bandwidth scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702549000004
    Figure 0007702549000004
  • Figure 0007702549000005
    Figure 0007702549000005
  • Figure 0007702549000006
    Figure 0007702549000006
Patent Text Reader

Abstract

To provide a method that is used for a video coding and decoding for performing a coding unit division and partitioning of an image or a video signal.SOLUTION: A coding method contains steps of: determining whether or not a size of a current block is larger than a minimum permission quadtree leaf node size; applying a binary tree division to the current block on the basis of the maximum permission binary tree root node size in the case where the size of the current block is not larger than the minimum permission quadtree leaf node size; generating a prediction block of the current block after the application of the binary tree division to the current block; acquiring a residual block of the current block on the basis of the current block and the prediction block; acquiring a conversion factor of the current block by applying a conversion to a sample value of the residual block; and coding a bit stream.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application generally relate to the field of video coding, and more specifically to coding unit splitting and partitioning.

Background Art

[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat, video conferencing, DVDs and Blu-ray discs, video content capture and editing systems, and camcorders for security applications.

[0003] The development of the block-based hybrid video coding approach in the 1990 H.261 standard has led to the development of new video coding techniques and tools, laying the foundation for new video coding standards. Further video coding standards include MPEG-1 video, MPEG-2 video, ITU-T H.262 / MPEG-2, ITU-T H.263, ITU-T H.264 / MPEG-4 Part10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC) and extensions, such as scalability and / or three-dimensional (3D) extensions of such standards. As video production and consumption become increasingly widespread, video traffic is the largest burden on communication networks and data storage, and thus one of the aims of many video coding standards has been to achieve bit rate reduction without sacrificing picture quality compared to previous ones. The latest High Efficiency Video Coding (HEVC) can compress video at about twice the rate of AVC without sacrificing picture quality, yet there is still a desire for new techniques to further compress video compared to HEVC.

[0004] The amount of video data required to depict even relatively short videos can be substantial and can pose difficulties when the data is streamed or otherwise communicated over a communication network having a limited bandwidth capacity. Thus, video data is generally compressed before being communicated over today's telecommunications networks. Also, when video is stored on a storage device, the size of the video can be an issue since memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode video data prior to transmission or storage, thereby reducing the amount of data necessary to represent a digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much in the way of image quality are desired due to limited network resources and the ever-increasing demands for higher video quality. SUMMARY OF THE INVENTION

[0005] Embodiments of the present application (or present disclosure) provide an apparatus and method for encoding and decoding according to independent claims. The above and other objects are achieved by the subject matter of the independent claims. Further implementations are apparent from the dependent claims, the description, and the drawings.

[0006] According to a first aspect, the present invention relates to a video decoding method. The method is performed by a decoding device. The method includes determining whether the size of a current block is greater than a minimum allowable quadtree leaf node size, and applying a multi-type tree split to the current block when the size of the current block is not greater than the minimum allowable quadtree leaf node size, where the minimum allowable quadtree leaf node size is not greater than a maximum allowable binary tree root node size or the minimum allowable quadtree leaf node size is not greater than a maximum allowable ternary tree root node size.

[0007] The current block can be obtained by splitting an image or a coding tree unit (CTU).

[0008] The method may include two cases: 1) treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA; 2) treeType is equal to DUAL_TREE_CHROMA. In case 1), the current block is a luma block, and in case 2), the current block is a chroma block.

[0009] The maximum allowable binary tree root node size may be the maximum luma size in the luma samples of a luma coding root block that can be split using binary tree splitting.

[0010] The maximum allowable ternary tree root node size may be the maximum luma size in the luma samples of a luma coding root block that can be split using ternary tree splitting.

[0011] The minimum allowable quadtree leaf node size may be the minimum luma size in the luma samples of a luma leaf block resulting from quadtree splitting.

[0012] This approach promotes efficient splitting or signaling of the splitting parameters for image / video blocks.

[0013] Furthermore, in a possible implementation of the method according to the first aspect, the method further includes determining whether the current block of the picture is a boundary block. When the size of the current block is not greater than the minimum allowable quadtree leaf node size, the step of applying the multi-type tree split to the current block includes the step of applying a binary split to the current block when the current block is a boundary block and the size of the current block is not greater than the minimum allowable quadtree leaf node size. In this case, it should be noted that the minimum allowable quadtree leaf node size is not greater than the maximum allowable binary tree root node size. Therefore, when the size of the current block is not greater than the minimum allowable quadtree leaf node size and the size of the current block is not greater than the maximum allowable binary tree root node size, the above-described step of applying the multi-type tree split to the current block includes the step of applying a binary split to the current block when the current block is a boundary block and the size of the current block is not greater than the minimum allowable quadtree leaf node size.

[0014] The method can further include obtaining a reconstructed block of the block directly or indirectly obtained from applying the binary split to the current block.

[0015] Bringing about a binary split can be particularly beneficial for blocks at the image / video frame boundary, for example, blocks cut by the boundary. Therefore, in some implementations, it may be beneficial to apply this approach to boundary blocks but not to the remaining blocks. However, the present disclosure is not limited to this. As described above, this approach of applying a binary split is also applied to non-boundary blocks and is signaled efficiently.

[0016] In a possible implementation of the method according to the first aspect or the above-described embodiments, the minimum allowable quadtree leaf node size is not greater than the maximum allowable binary tree root node size, and the minimum allowable quadtree leaf node size is not greater than the maximum allowable ternary tree root node size.

[0017] In a possible implementation of the method according to the first aspect or the above-described embodiments, the step of applying a multi-type tree split to the current block may include the step of applying a ternary split to the current block or the step of applying a binary split to the current block. However, the present disclosure is not limited thereto, and generally, a multi-type tree split may include further or other different types of splits.

[0018] In a possible implementation of the method according to the first aspect or the above-described embodiments, the method may further include the step of determining the maximum allowable binary tree root node size based on the minimum allowable quadtree leaf node size. This promotes efficient signaling / memory of the parameters. For example, the maximum allowable binary tree root node size may be considered equal to the minimum allowable quadtree leaf node size. Regarding another example, the lower limit of the maximum allowable binary tree root node size may be considered equal to the minimum allowable quadtree leaf node size, and the minimum allowable quadtree leaf node size can be used to determine the validity of the maximum allowable binary tree root node size. However, the present disclosure is not limited thereto, and another relationship may be assumed to derive the maximum allowable binary tree root node size.

[0019] According to an exemplary embodiment, in a first aspect or additionally or alternatively to the above-described embodiments, the method may further include a step of dividing an image into blocks, where the blocks include the current block. The step of applying a binary split to the current block may include applying the binary split to a boundary block having a maximum boundary multi-type partition depth, where the maximum boundary multi-type partition depth is at least the sum of a maximum multi-type tree depth and a maximum multi-type tree depth offset, and the maximum multi-type tree depth is greater than 0. Further, in some implementations, when applying a binary split to a boundary block, the maximum multi-type tree depth is greater than 0.

[0020] In a possible implementation form of the method according to the first aspect or the above-described embodiments, the method may further include a step of dividing an image into blocks (the blocks include the current block). The step of applying a multi-type tree split to the current block may include applying a multi-type tree split to the current block of a block having a final maximum multi-type tree depth, where the final maximum multi-type tree depth is at least the sum of a maximum multi-type tree depth and a maximum multi-type tree depth offset, and the maximum multi-type tree depth is greater than or equal to the result of subtracting the Lo2 value of a minimum allowable quadtree leaf node size from the Lo2 value of a minimum allowable transformation block size, or the maximum multi-type tree depth is greater than or equal to the result of subtracting the Lo2 value of a minimum allowable coding block size from the Lo2 value of a minimum allowable quadtree leaf node size. This encourages further splitting even for larger partition depths.

[0021] The current block may be a non-boundary block. The maximum multi-type tree depth offset may be 0. Alternatively or additionally, the current block may be a boundary block, and the multi-type tree split is a binary split. The multi-type tree split may be a ternary split or may include it.

[0022] According to a second aspect, the present invention relates to a method for encoding, the method being performed by an encoding device. The method includes determining whether the size of a current block is greater than a minimum allowable quadtree leaf node size, and applying a multi-type tree split to the current block when the size of the current block is not greater than the minimum allowable quadtree leaf node size, where the minimum allowable quadtree leaf node size is not greater than the maximum allowable binary tree root node size or the minimum allowable quadtree leaf node size is not greater than the maximum allowable ternary tree root node size.

[0023] The encoding method can apply any of the above-described rules and constraints explained with respect to the decoding method. Since the encoder side and the decoder side must share a bitstream, in particular, the encoding side generates a bitstream after coding the partitions resulting from the above-described partitioning, while the decoding side analyzes the bitstream and reconstructs the decoded partitions accordingly. The same applies to the embodiments regarding the encoding device (encoder) and the decoding device (decoder) described below.

[0024] According to a third aspect, the present invention relates to a decoding device including a circuit, the circuit being configured to determine whether the size of a current block is greater than a minimum allowable quadtree leaf node size and to apply a multi-type tree split to the current block when the size of the current block is not greater than the minimum allowable quadtree leaf node size, where the minimum allowable quadtree leaf node size is not greater than the maximum allowable binary tree root node size or the minimum allowable quadtree leaf node size is not greater than the maximum allowable ternary tree root node size. It should be noted that determining whether the size of the current block is greater than the minimum allowable quadtree leaf node size may be performed based on signaling in the bitstream on the decoding side.

[0025] According to a fourth aspect, the present invention relates to an encoding device including a circuit, the circuit being configured to determine whether the size of a current block is greater than a minimum allowable quadtree leaf node size, and to apply a multi-type tree split to the current block when the size of the current block is not greater than the minimum allowable quadtree leaf node size, the minimum allowable quadtree leaf node size being not greater than the maximum allowable binary tree root node size or the minimum allowable quadtree leaf node size being not greater than the maximum allowable ternary tree root node size.

[0026] The method according to the first aspect of the present invention can be executed by an apparatus or device according to the third aspect of the present invention. Further features and implementation forms of the method according to the third aspect of the present invention correspond to the features and implementation forms of the apparatus according to the first aspect of the present invention.

[0027] The method according to the second aspect of the present invention can be executed by an apparatus or device according to the fourth aspect of the present invention. Further features and implementation forms of the method according to the fourth aspect of the present invention correspond to the features and implementation forms of the apparatus according to the second aspect of the present invention.

[0028] According to a fifth aspect, the present invention relates to an apparatus for decoding a video stream including a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the first aspect.

[0029] According to a sixth aspect, the present invention relates to an apparatus for encoding a video stream including a processor and a memory. The memory stores instructions for causing the processor to execute the method according to the second aspect.

[0030] According to a seventh aspect, there is proposed a computer-readable storage medium storing instructions that, when executed, cause one or more processors configured to code video data to operate. The instructions cause the one or more processors to execute the method according to the first or second aspect or any possible embodiment of the first or second aspect.

[0031] According to an eighth aspect, the present invention relates to a computer program including program code for executing, when executed on a computer, a method according to the first or second aspect or any possible embodiment of the first or second aspect.

[0032] According to a ninth aspect, there is provided a non-transitory computer-readable storage medium storing programming for execution by a processing circuit, the programming configuring the processing circuit to execute any of the above methods when executed by the processing circuit.

[0033] For purposes of clarity, any one embodiment disclosed herein can be combined with any one or more other embodiments so as to generate new embodiments within the scope of the present disclosure.

[0034] Details of one or more embodiments will be apparent in the following, from the accompanying drawings and the specification. Other features, objects, and advantages will be apparent from the specification, the drawings, and the claims.

Brief Description of the Drawings

[0035] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings and figures.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Best Mode for Carrying Out the Invention

[0036] In the following description, reference is made to the accompanying drawings that form a part of this disclosure and illustrate specific aspects of embodiments of the invention or specific aspects that may be used in embodiments of the invention. It is understood that embodiments of the invention may be used in other aspects and may include structural or logical changes not shown in the accompanying drawings. Accordingly, the following detailed description should not be taken in a limiting sense, and the scope of the invention is defined by the appended claims.

[0037] For example, it is understood that the disclosure related to the described method may also apply to a corresponding device or system configured to perform the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., one unit for performing one or more of the described method steps, or multiple units each performing one or more of the multiple steps) such as functional units for performing the one or more described method steps, even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific device is described based on one or more units such as functional units, the corresponding method may include one step (e.g., one step for performing the functions of one or more units, or multiple steps each performing the functions of one or more of the multiple units) for performing the functions of the one or more units, even if such one or more steps are not explicitly described or illustrated in the drawings. Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described in this application may be combined with each other, unless otherwise specified.

[0038] Video coding typically refers to the processing of a series of pictures that form a video or a video sequence. In the field of video coding, the terms "frame" or "image" may be used synonymously instead of the term "picture". The video coding used in this application (or this disclosure) refers to video encoding or video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves performing the reverse process compared to the encoder to reconstruct the video picture. Embodiments regarding the "coding" of a video picture (or, generally, a picture as described below) are to be understood as being related to the "encoding" or "decoding" of a video sequence. The combination of the encoding part and the decoding part is also referred to as a CODEC (Coding and Decoding).

[0039] In the case of lossless video coding, it is possible to reconstruct the original video picture, i.e., the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss occurs during storage and transmission). In the case of lossy video coding, additional compression is performed, e.g., by quantization, to reduce the amount of data representing the video picture, and the video picture cannot be completely reconstructed on the decoder side, i.e., the quality of the reconstructed video picture is lower or worse than that of the original video picture.

[0040] Several video coding standards since H.261 belong to the group of "non-lossy hybrid video codexes" (i.e., combining spatial and temporal prediction in the sample domain and 2D transform coding for applying quantization in the transform domain). Each picture of a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, on the encoder side, the video is typically processed, i.e., encoded, at the block (video block) level by generating prediction blocks using, for example, spatial (intra-picture) prediction and temporal (inter-picture) prediction, subtracting the prediction block from the current block (the current processed / to-be-processed block) to obtain a residual block, transforming the residual block, and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression). On the decoder side, the reverse process compared to the encoder is partially applied to the encoded or compressed block to reconstruct the current block for presentation. Further, the encoder repeats the decoder's processing loop, and as a result, both will generate the same prediction (e.g., intra and inter prediction) and / or reconstruction for processing, i.e., coding, subsequent blocks.

[0041] As used herein, the term "block" may be part of a picture or frame. For ease of explanation, embodiments of the present invention are described with reference to the reference software of Versatile Video Coding (VVC) or High Efficiency Video Coding (HEVC) developed by the Joint Collaborative Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC. This may refer to CU (Coding Units), PU (Prediction Units), or TU (Transform Units). In HEVC, a Coding Tree Unit (CTU) is divided into a plurality of CUs by using a quadtree structure shown as a coding tree. The decision of whether to code a picture using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can further be divided into one, two, or four PUs according to the PU split type. Within one PU, the same prediction process is applied, and the relevant information is sent to the decoder based on the PU. After obtaining the residual block by applying the prediction process based on the PU split type, the CU can be divided into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU. In the development of the latest video compression technology, a quadtree and binary tree (QTBT) partitioning frame is used to divide the coding block. In the QTBT block structure, the CU can be square or rectangular. For example, a Coding Tree Unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes are further partitioned by a binary tree structure. The binary tree leaf nodes are referred to as Coding Units (CUs), and their segmentation is used for prediction and conversion processing without any further partitioning.This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. At the same time, it has been proposed that multiple partitions, such as a ternary tree (TT) partition, be used together with the QTBT block structure. The term "device" may also be an "apparatus", "decoder", or "encoder".

[0042] Hereinafter, embodiments of the encoder 20, decoder 30, and coding system 10 will be described based on FIGS. 1-3.

[0043] FIG. 1A is a conceptual or schematic block diagram showing an exemplary coding system 10, such as a video coding system 10 capable of using the technology of the present application (this disclosure). The encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) of the video coding system 10 represent example devices that may be configured to execute the techniques according to various examples described in the present application. As shown in FIG. 1A, the coding system 10 includes a source device 12 configured to provide encoded data 13, such as an encoded picture 13, to a destination device 14 that decodes the encoded data 13, for example.

[0044] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include a picture source 16, a preprocessing unit 18 such as a picture preprocessing unit 18, and a communication interface or communication unit 22.

[0045] Video source 16 may include any kind of picture capture device, such as capturing a real-world picture, and / or, for example, a picture or comment generation device (in the case of screen content coding, any text on the screen is also considered part of the picture or image to be coded), for example, a computer graphics processor that generates computer animation pictures, or any kind of device that acquires and / or provides real-world pictures or computer animation pictures (such as screen content or virtual reality (VR) pictures), and / or any combination thereof (such as augmented reality (AR) pictures), or it may be those. The picture source may be any kind of memory or storage that stores any of the above pictures.

[0046] (Digital) A picture is or can be considered to be a two-dimensional array or matrix of samples having intensity values. The samples in the array may be called pixels (a shortening of picture elements) or pels. The amount of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used, i.e., the picture may be represented by or may include three sample arrays. In the RGB format or color space, the picture includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance / chrominance format or color space, such as YCbCr, which includes a luminance component indicated by Y (sometimes L is used instead) and two chrominance components indicated by Cb and Cr. The luminance (or abbreviation, luma) component Y represents the luminance or gray-level intensity (as in, for example, a gray-scale picture), and the two chrominance (or abbreviation, chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format can be converted or transformed to the YCbCr format, and vice versa, and the process is also known as color conversion or transformation. If the picture is monochrome, the picture may include only a luminance sample array.

[0047] The picture source 16 (e.g., video source 16) may be, for example, a camera that captures pictures, a memory such as a picture memory that includes or stores pre-captured or generated pictures, and / or any kind of (internal or external) interface for acquiring or receiving pictures. The camera may be, for example, a local or integrated camera integrated into the source device, and the memory may be, for example, a local or integrated memory integrated into the source device. The interface may be, for example, an external interface for receiving pictures from an external video source, an external picture capture device such as a camera, an external memory, or an external picture generation device, for example, an external computer graphics processor, a computer or a server. The interface can be any kind of interface that follows any proprietary or standardized interface protocol, for example, a wired or wireless interface, an optical interface. The interface for acquiring the picture data 17 may be the same interface as, or a part of, the communication interface 22.

[0048] Unlike the preprocessing unit 18 and the processing executed by the preprocessing unit 18, the pictures and the picture data 17 (e.g., video data 16) may also be referred to as unprocessed pictures or unprocessed picture data 17.

[0049] The preprocessing unit 18 is configured to receive the (unprocessed) picture data 17 and perform preprocessing on the picture data 17 to obtain the preprocessed picture 19 or the preprocessed picture data 19. The preprocessing executed by the preprocessing unit 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise reduction. It is possible to understand that the preprocessing unit 18 may be an optional component.

[0050] The encoder 20 (e.g., video encoder 20) is configured to receive the pre - processed picture data 19 and provide encoded picture data 21 (further details will be described, for example, based on FIG. 2 or FIG. 4 below).

[0051] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further - processed version thereof) via the communication channel 13 to other devices, for example, the destination device 14 or any other device, for storage or direct reconstruction.

[0052] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21, process the encoded picture data 21 before storing it or transmitting it to other devices, for example, the destination device 14 or any other device, for storage or direct reconstruction, or before storing the encoded data 13 and / or transmitting the encoded data 13 to other devices, for example, the destination device 14 or any other device, for decoding or storage.

[0053] The destination device 14 includes a decoder 30 (e.g., video decoder 30) and additionally, i.e., optionally, may include a communication interface or communication unit 28, a post - processing unit 32, and a display device 34.

[0054] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further - processed version thereof) directly from, for example, the source device 12 or from any other source, for example, from a source device, for example, from a storage device of the encoded picture data, and provide the encoded picture data 21 to the decoder 30.

[0055] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 or the encoded data 13 directly from, for example, the source device 12, or from any other source, for example, a storage device, for example, an encoded picture data storage device.

[0056] The communication interface 22 and the communication interface 28 can be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, for example, a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or any combination thereof, or via any type of private and public network, or any combination thereof.

[0057] The communication interface 22 can be configured to, for example, package the encoded picture data 21 in a suitable format, such as a packet, and / or use any type of transmission encoding or processing for transmission via the communication link or communication network to process the encoded picture data.

[0058] The communication interface 28 that forms the corresponding part of the communication interface 22 can be configured to, for example, unpack the encoded data 13 to obtain the encoded picture data 21.

[0059] The communication interface 28 that forms the corresponding part of the communication interface 22 processes the transmitted data, for example, by receiving the transmitted data and using any type of corresponding transmission decoding or processing and / or unpacking to obtain the encoded picture data 21.

[0060] Both communication interface 22 and communication interface 28 can be configured as a unidirectional communication interface or a bidirectional communication interface indicated by an arrow regarding the encoded picture data 13 from the source device 12 to the destination device 14 in FIG. 1A, and can be configured to, for example, send and receive messages, for example, set up a connection, and confirm and exchange other arbitrary information related to a communication link and / or data transmission such as, for example, encoded picture data transmission.

[0061] Decoder 30 is configured to receive the encoded picture data 21 and provide the decoded picture data 31 or the decoded picture 31 (further details will be described below, for example, based on FIG. 3 or FIG. 5).

[0062] The post - processing processor 32 of the destination device 14 is configured to post - process the decoded picture data 31 (also called reconstructed picture data) such as, for example, the decoded picture 31 to obtain post - processed picture data 33 such as the post - processed picture 33. The post - processing executed by the post - processing unit 32 may include, for example, color - format conversion (e.g., from YCbCr to RGB), color correction, trimming, resampling, or any other arbitrary processing for preparing, for example, the decoded picture data 31 for display by the display device 34.

[0063] The display device 34 of the destination device 14 is configured to receive the post - processed picture data 33 and display the picture to, for example, a user or viewer. The display device 34 may be any type of display that represents the reconstructed picture, such as an integrated or external display or monitor, or may include them. The display may include, for example, a liquid crystal display (LCD), an organic light - emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0064] FIG. 1A depicts the source device 12 and the destination device 14 as separate devices, but embodiments of the device may also include both or both functions, the source device 12 or corresponding functions and the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions and the destination device 14 or corresponding functions may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0065] As will be apparent to those skilled in the art based on the specification, the presence and (exact) partitioning of the functions in the source device 12 and / or destination device 14 shown in FIG. 1A may vary depending on the actual device and application.

[0066] The encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) can each be implemented as various suitable circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented partially in software, the device can store software instructions in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Any of the above (including hardware, software, combinations of hardware and software, etc.) may be considered to be one or more processors. The video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders and may both be integrated as part of a combined encoder / decoder (CODEC) within an individual device.

[0067] Encoder 20 may be implemented by processing circuit 46 to embody various modules as described with respect to encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented via processing circuit 46 to embody various modules as described with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations as described hereinafter. As shown in FIG. 5, when the technology is implemented partially in software, the device stores software instructions in a suitable non-transitory computer-readable storage medium and uses one or more processors to execute the instructions in hardware to implement the technology of the present disclosure. Either video encoder 20 or video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) within a single device, for example, as shown in FIG. 1B.

[0068] Source device 12 may be referred to as a video encoding device or a video encoder. Destination device 14 may be referred to as a video decoding device or a video decoder. Source device 12 and destination device 14 may be an example of a video coding device or a video coder.

[0069] Source device 12 and destination device 14 may include any device within a wide range including any type of handheld or stationary device, for example, notebook or laptop computers, mobile phones, smartphones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, display devices, digital media players, video game consoles, video streaming devices (such as content service servers or content delivery servers, etc.), broadcast receiving devices, broadcast transmitting devices, etc., and may use or not use any type of operating system.

[0070] In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Accordingly, the source device 12 and the destination device 14 may be wireless communication devices.

[0071] In some cases, the video coding system 10 shown in FIG. 1A is merely an example, and the technology of the present application may be applicable to video coding settings (e.g., video encoding or video decoding) that do not necessarily involve any data communication between the encoding device and the decoding device. In other examples, data is retrieved from local memory and streamed over a network, etc. The video encoding device can encode data and store it in memory, and / or the video decoding device can retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to memory and / or retrieve and decode data from memory.

[0072] For the sake of convenience of explanation, embodiments of the present invention are described herein by referring to the reference software of the next-generation video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) for video coding such as High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Moving Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC. For each of the above examples described in relation to the video encoder 20, it should be understood that the video decoder 30 can be configured to perform the reverse process. Regarding signaling syntax elements, the video decoder 30 can be configured to receive and analyze such syntax elements and decode the relevant video data accordingly. In some examples, the video encoder 20 can entropy encode one or more slices into the encoded video bitstream. In such examples, the video decoder 30 can analyze such syntax elements and decode the relevant video data accordingly.

[0073] FIG. 1B is an exemplary diagram of another exemplary video coding system 40 including the encoder 20 of FIG. 2 and / or the decoder 30 of FIG. 3 according to an exemplary embodiment. The system 40 can implement the technology according to various examples described in the present application. In the illustrated implementation, the video coding system 40 can include an imaging device 41, a video encoder 100, a video decoder 30 (and / or a video coder implemented by the logic circuit 47 of the processing unit 46), an antenna 42, one or more processors 43, one or more memory stores 44, and / or a display device 45.

[0074] As shown, the imaging device 41, antenna 42, processing unit 46, logic circuit 47, video encoder 20, video decoder 30, processor 43, memory store 44, and / or display device 45 can communicate with each other. Although both the video encoder 20 and the video decoder 30 are shown as described above, the video coding system 40 may include only the video encoder 20 or only the video decoder 30 in various examples.

[0075] As shown, in some examples, video coding system 40 may include antenna 42. Antenna 42 can be configured, for example, to transmit or receive an encoded bitstream of video data. Further, in some examples, video coding system 40 may include display device 45. Display device 45 can be configured to present video data. As shown, in some examples, logic circuit 47 may be implemented by processing unit 46. Processing unit 46 can include, for example, application specific integrated circuit (ASIC) logic, a graphics processor, a general purpose processor, etc. Also, video coding system 40 may also include optional processor 43, which may similarly include application specific integrated circuit (ASIC) logic, a graphics processor, a general purpose processor, etc. In some examples, logic circuit 47 can be implemented by hardware, video coding dedicated hardware, etc., and processor 43 can implement general purpose software, an operating system, etc. Further, memory store 44 can be any type of memory such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM)), or non-volatile memory (e.g., flash memory), etc. By way of non-limiting example, memory store 44 can be implemented by cache memory. In some examples, logic circuit 47 can access memory store 44 (e.g., for the implementation of an image buffer). In other examples, logic circuit 47 and / or processing unit 46 may include a memory store (e.g., a cache, etc.) for the implementation of an image buffer, etc.

[0076] In some examples, the video encoder 100 implemented by a logic circuit may include an image buffer (e.g., by the processing unit 46 or the memory store 44) and a graphics processing unit (e.g., by the processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit may include a video encoder 100 implemented by the logic circuit 47 to implement various modules as described with respect to FIG. 2 and / or any other encoder system or subsystem described in the present application. The logic circuit can be configured to perform various operations as described in the present application.

[0077] The video decoder 30 can be implemented in a similar manner as implemented by the logic circuit 47 to implement various modules as described with respect to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described in the present application. In some examples, the video decoder 30 may be implemented by a logic circuit and may include an image buffer (e.g., by the processing unit 420 or the memory storage store 44) and a graphics processing unit (e.g., by the processing unit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit may include a video decoder 30 implemented by the logic circuit 47 to implement various modules as described with respect to FIG. 3 and / or any other decoder system or subsystem described in the present application.

[0078] In some examples, the antenna 42 of the video coding system 40 can be configured to receive an encoded bitstream of video data. As described above, the encoded bitstream may include data related to encoding a video frame as described in this application, such as data related to the coding partition, indicators, index values, mode selection data, etc. (e.g., transform coefficients or quantization transform coefficients, optional indicators as discussed, and / or data defining the coding partition). The video coding system 40 may also include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present the video frame.

[0079] FIG. 2 is a schematic / conceptual block diagram of an exemplary video encoder 20 configured to implement the technology of this application (disclosure). In the example of FIG. 2, the video encoder 20 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a buffer for decoded pictures (DPB) 230, a prediction processing unit 260, and an entropy coding unit 270. The prediction processing unit 260 can include an inter prediction unit 244, an intra prediction unit 254, and a mode selection unit 262. The inter prediction unit 244 can include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder by a hybrid video codec.

[0080] For example, the residual calculation unit 204, the conversion processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy encoding unit 270 form the forward signal path of the encoder 20. For example, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the buffer for decoded pictures (DPB) 230, and the prediction processing unit 260 form the reverse signal path of the encoder. The reverse signal path of the encoder corresponds to the signal path of the decoder (see the decoder 30 in FIG. 3).

[0081] Also, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the loop filter 220, the buffer for decoded pictures (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also referred to as forming the "built-in decoder" of the video encoder 20. The encoder 20 is configured to receive, for example, a picture from the input 202, the picture 201, or the block 203 of the picture 201, for example, a picture in a sequence of pictures forming a video or a video sequence. The picture block 203 is also referred to as the current picture block or the picture block to be coded, and the picture 201 is referred to as the current picture or the picture to be coded (in particular, in video coding, the current picture is distinguished from other pictures, for example, pictures previously encoded and / or decoded in the same video sequence, i.e., the video sequence including the current picture).

[0082] (Digital) pictures are, or can be considered to be, two-dimensional arrays or matrices of samples having intensity values. Samples in the array may also be called pixels (abbreviation for picture elements) or pels. The amount of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used, i.e., the picture may be represented by, or may include, three sample arrays. In the RGB format or color space, the picture includes corresponding red, green, and blue sample arrays. However, in video coding, each sample is typically represented in a luminance / chrominance format or color space, e.g., YCbCr, which includes a luminance component represented by Y (sometimes L is used instead) and two chrominance components represented by Cb and Cr. The luminance (or abbreviation, luma) component Y represents the luminance or gray-level intensity (e.g., as in a gray-scale picture), and the two chrominance (or abbreviation, chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format can be converted or transformed to the YCbCr format, and vice versa, and the process is also known as color conversion or transformation. If the picture is monochrome, the picture may only include a luminance sample array. Thus, the picture may be, for example, an array of luma samples in a monochrome format, or in 4:2:0, 4:2:2, and 4:4:4 color formats of an array of luma samples and two corresponding arrays of chroma samples.

[0083] Partitioning An embodiment of the encoder 20 can include a partitioning unit (not shown in FIG. 2) configured to partition a picture 201 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The partitioning unit can be configured to use the same block size and a corresponding grid defining the block size for all pictures of a video sequence, or to vary the block size between pictures, subsets, or groups of pictures and partition each picture into corresponding blocks.

[0084] In a further embodiment, the video encoder may be configured to directly receive the blocks 203 of the picture 201, such as one, several, or all of the blocks forming the picture 201. Also, the picture block 203 may also be referred to as the current picture block or the picture block to be encoded. In one example, the prediction processing unit 260 of the video encoder 20 can be configured to perform any combination of the partitioning techniques described above.

[0085] Similar to picture 201, block 203 can again be considered as a two-dimensional array or matrix of samples having intensity values (sample values), or as such, although it has dimensions smaller than those of picture 201. In other words, block 203 may include, for example, one sample array (e.g., the luma array in the case of a monochrome picture 201) or three sample arrays (e.g., the luma and two chroma arrays in the case of a color picture 201), or any other quantity and / or type of array depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an MxN (M columns and N rows) array of samples, or an MxN array of transform coefficients.

[0086] As shown in FIG. 2, encoder 20 is configured to encode blocks of picture 201 on a block-by-block basis, and for example, encoding and prediction are performed for each block 203.

[0087] The embodiment of video encoder 20 shown in FIG. 2 can further be configured to partition and / or encode a picture by using slices (also called video slices), and the picture can be partitioned into one or more slices (typically non-overlapping) or encoded using them, and each slice can include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., tiles (H.265 / HEVC and VVC) or bricks (VVC)).

[0088] The embodiment of the video encoder 20 shown in FIG. 2 can be further configured to partition and / or encode pictures by using slices / tile groups (also called video tile groups) and / or tiles (also called video tiles), where a picture can be partitioned into and / or encoded using one or more slices / tile groups (typically non-overlapping), each slice / tile group can include, for example, one or more blocks (e.g., CTUs) or one or more tiles, each tile can be, for example, rectangular in shape, and can include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.

[0089] Residual calculation The residual calculation unit 204 is configured to calculate a residual block 205 based on a picture block 203 and a prediction block 265 (more details about the prediction block 265 will be described later) by subtracting the sample values of the prediction block 265 from the sample values of the picture block 203 on a sample-by-sample (pixel-by-pixel) basis, and to obtain the residual block 205 in the sample domain.

[0090] Transformation The transformation processing unit 206 is configured to apply a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 in order to obtain transformation coefficients 207 in the transform domain. The transformation coefficients 207 are also referred to as transform residual coefficients and represent the residual block 205 in the transform domain.

[0091] The conversion processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the conversion specified for HEVC / H.265. Compared to the orthogonal DCT transform, such an integer approximation is typically scaled by a factor. To ensure the norm of the residual block processed by the forward and inverse transforms, an additional scaling factor is applied as part of the conversion process. The scaling factor is typically selected based on specific constraints such as a scaling factor that is a power of two for shift operations, the bit depth of the transform coefficients, the trade-off between accuracy and implementation cost, etc. A specific scaling factor is specified, for example, for the inverse transform by the inverse transform processing unit 212 in the decoder 30 (and the corresponding inverse transform by the inverse transform processing unit 212 in the encoder 20), and the corresponding scaling factor for the forward transform by the conversion processing unit 206 in the encoder 20 may be specified accordingly.

[0092] Embodiments of the video encoder 20 (individual conversion processing units 206) may be configured to directly or encode or compress conversion parameters, such as a certain type of conversion or multiple conversions, by the entropy encoding unit 270 and output the result, such that the video decoder 30 can receive and use the conversion parameters for decoding.

[0093] Quantization The quantization unit 208 is configured to quantize the transform coefficients 207 and obtain quantized transform coefficients 209 by applying, for example, scalar quantization or vector quantization. The quantized transform coefficients 209 may also be referred to as quantized residual coefficients 209. The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization can be adjusted by adjusting the quantization parameter (QP). For example, for scalar quantization, different scalings may be applied to achieve finer or coarser quantization. Smaller quantization steps correspond to finer quantization, and larger quantization steps correspond to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). For example, the quantization parameter may be an index for a given set of applicable quantization step sizes. For example, a smaller quantization parameter can correspond to finer quantization (smaller quantization step size), and a larger quantization parameter can correspond to coarser quantization (larger quantization step size), and vice versa. Quantization may include division by the quantization step size and corresponding or inverse quantization by an inverse quantization unit 210, etc., and may include multiplication by the quantization step size. Embodiments according to some standards, such as HEVC, may be configured to use the quantization parameter to determine the quantization step size. Generally, the quantization step size can be calculated based on the quantization parameter using a fixed-point approximation of an expression that includes division. An additional scaling factor is introduced for quantization and inverse quantization, and it is possible to restore the norm of the residual block, which may be modified due to the scaling used in the fixed-point approximation of the expression related to the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and inverse quantization may be combined.Alternatively, a customized quantization table may be used and signaled, e.g., in a bitstream, from the encoder to the decoder. Quantization is a lossy operation, and the loss increases as the quantization step size increases.

[0094] An embodiment of the video encoder 20 (individual conversion processing units 206) can be configured to encode and output a quantization parameter (QP), e.g., directly or by the entropy encoding unit 270, such that the video decoder 30 can receive and apply the quantization parameter for decoding.

[0095] The inverse quantization unit 210 is configured to apply the inverse quantization of the quantization unit 208 to the quantization coefficients, e.g., by applying the inverse of the quantization method applied by the quantization unit 208, based on or using the same quantization step as the quantization unit 208, to obtain the inverse quantization coefficients 211. The inverse quantization coefficients 211 may also be referred to as inverse quantization residual coefficients 211 and correspond to the transform coefficients 207 but typically do not match the transform coefficients due to the loss due to quantization.

[0096] The inverse transform processing unit 212 is configured to apply the inverse transform of the transform applied by the transform processing unit 206, e.g., inverse discrete cosine transform (DCT) or inverse discrete sine transform (DST), to obtain the inverse transform block 213 in the sample domain. The inverse transform block 213 may also be referred to as the inverse transform inverse quantization block 213 or the inverse transform residual block 213.

[0097] The reconstruction unit 214 (e.g., adder 214) is configured to add the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, e.g., by adding the sample values of the reconstructed residual block 213 and the sample values of the prediction block 265, to obtain the reconstructed block 215 in the sample domain.

[0098] As an option, a buffer unit 216 (or simply referred to as "buffer" 216), such as a line buffer 216 for example, is configured to buffer or store the reconstructed block 215 and individual sample values for intra prediction for example. In a further embodiment, the encoder may be configured to use the unfiltered reconstructed block and / or corresponding sample values stored in the buffer unit 216 for any kind of estimation and / or prediction, such as intra prediction.

[0099] An embodiment of the encoder 20 is such that, for example, the buffer unit 216 is used not only for storing the reconstructed block 215 for intra prediction 254, but also for a loop filter unit 220 (not shown in FIG. 2), and / or such that, for example, the buffer unit 216 and the buffer unit 230 of the decoded picture form one buffer. Other embodiments may be configured to use blocks or samples from the filtered block 221 and / or the buffer 230 of the decoded picture (both not shown in FIG. 2) as input or basis for intra prediction 254.

[0100] The loop filter unit 220 (or, abbreviated as "loop filter" 220) is configured to filter the reconstructed block 215 in order to obtain the filtered block 221, for example, to smooth pixel transitions or improve video quality in another way. The loop filter unit 220 represents one or more loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, or other filters, for example, a bilateral filter, an adaptive loop filter (ALF), an edge enhancement or smoothing filter, or a collaborative filter. The loop filter unit 220 is shown in FIG. 2 as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstructed block 221. The buffer 230 of the decoded picture can store the reconstructed coding block after the loop filter unit 220 performs a filtering process on the reconstructed coding block.

[0101] The loop filter unit 220 (or, abbreviated as "loop filter" 220) filters the reconstruction block 215 to obtain a filtered block 221, and is generally configured to filter reconstruction samples to obtain filtered sample values. The loop filter unit is configured to, for example, smooth pixel transitions or improve video quality in another way. The loop filter unit 220 can include one or more loop filters such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, for example, an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In one example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, SAO, and ALF. In another example, a process called luma mapping by chroma scaling (LMCS) (i.e., an adaptive in-loop resampler) is added. This process is executed before deblocking. In another example, the deblocking filter process may also be applied to internal sub-block edges, for example, affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra-sub-partition (ISP) edges. The loop filter unit 220 is shown as an in-loop filter in FIG. 2, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstruction block 221.

[0102] An embodiment of the video encoder 20 (individual loop filter units 220) can be configured to encode and output loop filter parameters (such as SAO filter parameters or ALF filter parameters or LMCS parameters), for example, directly or by the entropy decoding unit 270, so that, for example, the decoder 30 can receive and apply the same loop filter parameters or individual loop filters for decoding.

[0103] An embodiment of the encoder 20 (individual loop filter units 220) can be configured to entropy encode and output output filter parameters (such as sample adaptive offset information), for example, directly or by the entropy encoding unit 270 or any other entropy coding unit, so that the decoder 30 can receive and apply the same loop filter parameters for decoding.

[0104] The buffer (DPB) 230 of the decoded picture may be a reference picture memory that stores reference picture data used when the video encoder 20 encodes video data. The DPB 230 may be formed by any of various memory devices such as synchronous DRAM (SDRAM), magnetic resistance RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The DPB 230 and the buffer 216 may be provided by the same memory device or separate memory devices. In some examples, the buffer (DPB) 230 of the decoded picture is configured to store the filtered block 221. The buffer 230 of the decoded picture is further configured to store other previously filtered blocks, such as the same current picture or different pictures such as previously reconstructed pictures, such as the previously reconstructed and filtered block 221, and it is possible to provide a completely previously reconstructed, i.e., decoded picture (and corresponding reference blocks and corresponding samples) and / or a currently partially reconstructed picture (and corresponding reference blocks and corresponding samples) for, for example, inter prediction. In some examples, when the reconstruction block 215 is reconstructed without in-loop filtering, the buffer (DPB) 230 of the decoded picture stores one or more unfiltered reconstruction blocks 215, or generally, for example, when the reconstructed block 215 is not filtered by the loop filter unit 220, it is configured to store unfiltered reconstructed samples, or any other further processed version of the reconstructed block or sample.

[0105] The prediction processing unit 260, also referred to as the block prediction processing unit 260, receives or obtains the block 203 (the current block 203 of the current picture 201), the reconstructed picture data, for example, the reference samples of the same (current) picture from the buffer 216, and / or the reference picture data 231 from one or more previously decoded pictures from the buffer 230 of the decoded pictures, processes that data for prediction, i.e., is configured to provide a prediction block 265 that may be an inter prediction block 245 or an intra prediction block 255.

[0106] The mode selection unit 262 is configurable to select a prediction mode (e.g., an intra or inter prediction mode) and / or a corresponding prediction block 245 or 255 to be used as the prediction block 265 for the calculation of the residual block 205 and for the reconstruction of the reconstruction block 215.

[0107] An embodiment of the mode selection unit 262 may be configured to select a prediction mode (e.g., from the prediction modes supported by the prediction processing unit 260), where the prediction mode provides the best match or, in other words, the minimum residual (the minimum residual means better compression for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or considers or balances both. The mode selection unit 262 may be configured to determine the prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides minimum rate distortion optimization or a prediction mode whose associated rate distortion at least meets the prediction mode selection criteria.

[0108] The prediction processing (e.g., the prediction processing unit 260 and the mode selection (e.g., by the mode selection unit 262)) performed by the exemplary encoder 20 will be described in detail below.

[0109] Additionally or alternatively to the above-described embodiments, in another embodiment according to FIG. 15, the mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, such as an original block 203 (current block 203 of current picture 17), and reconstructed picture data, such as from the same (current) picture and / or from one or more previously decoded pictures, for example, filtered and / or unfiltered samples or blocks from a buffer 230 of a decoded picture or other buffer (e.g., a line buffer not shown). The reconstructed picture data is used as reference picture data for prediction, such as inter prediction or intra prediction, to obtain a prediction block 265 or predictor 265.

[0110] The mode selection unit 260 can be configured to determine or select a partitioning for the prediction mode (including non-partitioned) of the current block and the prediction mode (e.g., intra or inter prediction mode), and generate a corresponding prediction block 265 used for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.

[0111] An embodiment of the mode selection unit 260 is configured to select a partitioning and prediction mode that provides the best match, or in other words, the minimum residual (the minimum residual means better compression for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or both, taking into account or balancing them (e.g., from those supported or available by the mode selection unit 260). The mode selection unit 260 may be configured to determine the partitioning and prediction mode based on rate distortion optimization (RDO), i.e., to select a prediction mode that provides the minimum rate distortion. Terms such as "best", "lowest", "optimal", etc. in this context do not necessarily refer to the overall "best", "lowest", "optimal", etc., but may refer to an end or selection criterion such as a value exceeding or falling below a threshold, or may result in a "second-best choice" but also refer to the achievement of other constraints that reduce complexity and processing time.

[0112] In other words, the partitioning unit 262 can be configured to partition pictures from a video sequence into a sequence of coding tree units (CTUs), and the CTU 203 can be further partitioned into smaller block partitions or sub-blocks (which form blocks again) by repeatedly using, for example, quadtree partitioning (QT), binary tree partitioning (BT), or ternary tree partitioning (TT) or any combination thereof, and it is possible to perform prediction for each of the block partitions or sub-blocks, and the mode selection includes the selection of the tree structure of the partitioned blocks 203, and the prediction mode is applied to each of the block partitions or sub-blocks.

[0113] The partitioning (e.g., by partitioning unit 260) and prediction processing (by inter prediction unit 244 and intra prediction unit 254) executed by the exemplary video encoder 20 will be described in more detail below.

[0114] Partitioning The partitioning unit 262 can be configured to partition pictures from a video sequence into a sequence of coding tree units (CTUs), and the partitioning unit 262 can partition (or split) the coding tree unit (CTU) 203 into smaller partitions, such as smaller blocks of square or rectangular size. For a picture having three sample arrays, a CTU is composed of an N×N block of luma samples and two corresponding blocks of chroma samples. The maximum allowable size of the luma block of a CTU is specified to be 128×128 in the developing Versatile Video Coding (VVC), but may be specified to be a value other than 128×128 in the future, such as 256×256. The CTUs of a picture can be clustered / grouped as a slice / tile group, a tile, or a brick. A tile covers a rectangular region of a picture and a tile can be divided into one or more bricks. A brick is composed of a plurality of CTU rows within a tile. A tile that is not divided into a plurality of bricks can be referred to as a brick. However, a brick is a true subset of a tile and is not referred to as a tile. There are two modes for the tile groups supported in VVC, namely the raster scan slice / tile group mode and the rectangular slice mode. In the raster scan tile group mode, a slice / tile group includes the tile sequence in the tile raster scan of a picture. In the rectangular slice mode, a slice includes a number of bricks of a picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are in the order of the brick raster scan of the slice. These smaller blocks (which may also be called sub-blocks) can be further partitioned into even smaller partitions.This is also called tree partitioning or hierarchical tree partitioning. For example, a root block at the root tree level 0 (hierarchical level 0, depth 0) can be recursively partitioned into, for example, two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchical level 1, depth 1). These blocks can be further partitioned into two or more blocks at the next lower level, such as at tree level 2 (hierarchical level 2, depth 2), etc., until the partitioning ends, for example, due to an end criterion being met, such as reaching the maximum tree depth or the minimum block size. Blocks that are no longer partitioned are also called leaf blocks or leaf nodes of the tree. A tree that uses partitioning into two partitions is called a binary tree (BT), a tree that uses partitioning into three partitions is called a ternary tree (TT), and a tree that uses partitioning into four partitions is called a quaternary tree (QT).

[0115] For example, a coding tree unit (CTU) may be, or may include, a CTB of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a CTB of samples of a picture or a monochrome picture coded using three separate color planes and syntax structures used to code the samples. Correspondingly, a coding tree block (CTB) may be an NxN block of samples for a certain value N such that the splitting into component CTBs is a partitioning. A coding unit (CU) may be, or may include, a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture having three sample arrays, or a coding block of samples of a picture or a monochrome picture coded using three separate color planes and syntax structures used to code the samples. Correspondingly, a coding block (CB) may be an MxN block of samples for certain values of M and N such that the splitting into coding blocks of the CTB is a partitioning.

[0116] In an embodiment, for example according to HEVC, a coding tree unit (CTU) can be split into CUs by using a quadtree structure shown as a coding tree. The decision of whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the leaf CU level. Each leaf CU can be further split into one, two, or four PUs depending on the PU split type. Within one PU, the same prediction process is applied and the relevant information is sent to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the PU split type, the leaf CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU.

[0117] In an embodiment, for example, according to the currently evolving state-of-the-art video coding standard called Versatile Video Coding (VVC), a combined quadtree nested multi-type tree using, for example, binary and ternary segmentation structures is used to partition coding tree units. In the coding tree structure in a coding tree unit, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree. Then, a quadtree leaf node can be further partitioned by a multi-type tree structure. There are four split types in the multi-type tree structure: vertical binary split (SPLIT_BT_VER), horizontal binary split (SPLIT_BT_HOR), vertical ternary split (SPLIT_TT_VER), and horizontal ternary split (SPLIT_TT_HOR). A multi-type tree leaf node is called a coding unit (CU), and this segmentation is used for prediction and transformation processing without any further partitioning as long as the CU is not too large for the maximum transform length. This means that, in most cases, the CU, PU, and TU have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color components of the CU. VVC is developing a unique signaling mechanism for partition split information in a quadtree with a nested multi-type tree coding tree structure. In the signaling mechanism, a coding tree unit (CTU) is treated as the root of the quadtree and is first partitioned by the quadtree structure. Each quadtree leaf node (if large enough to allow it) is then further partitioned by the multi-type tree structure.In a multi-type tree structure, the first flag (mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned. If the node is further partitioned, the second flag (mtt_split_cu_vertical_flag) is signaled to indicate the split direction, and then the third flag (mtt_split_cu_binary_flag) is signaled to indicate whether the split is a binary split or a ternary split. Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree split mode (MttSplitMode) of the CU can be derived by the decoder based on predefined rules or tables. In a specific design, such as the 64×64 luma block and 32×32 chroma pipeline design in a VVC hardware decoder, as shown in Figure 6, it should be noted that the TT split is prohibited when either the width or height of the luma coding block is greater than 64. The TT split is also prohibited when the width or height of the chroma coding block exceeds 32. The pipeline design divides the picture into virtual pipeline data units (VPDUs), which are defined as non-overlapping units within the picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. Since the VPDU size is roughly proportional to the buffer size of most pipeline stages, it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) and binary tree (BT) partitions may cause an increase in the VPDU size. Furthermore, it should be noted that if a part of the tree node block exceeds the lower or right picture boundary, the tree node block is forced to be split until all samples of all coded CUs are placed within the picture boundary.

[0118] As an example, an Intra-Sub-Partition (ISP) tool can divide a luma intra prediction block into two or four sub-partitions either vertically or horizontally, depending on the block size.

[0119] In one example, the mode selection unit 260 of the video encoder 20 can be configured to perform any combination of the partitioning techniques described in this application. As described above, the encoder 20 is configured to determine or select the best or optimal prediction mode from a set of (predetermined) prediction modes. The set of prediction modes can include, for example, an intra prediction mode and / or an inter prediction mode.

[0120] The set of intra prediction modes can include 35 different intra prediction modes, such as non-directional modes like the DC (or average) mode and the planar mode, or directional modes as defined, for example, in H.265, or can include 67 different intra prediction modes, such as non-directional modes like the DC (or average) mode and the planar mode, or directional modes as defined, for example, in VVC. As an example, some conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks, as defined, for example, in VVC. As another example, only the long side is used to calculate the average for non-square blocks in order to avoid the division operation for DC prediction. Also, the result of intra prediction in the planar mode can be further corrected by the position-dependent intra prediction combination (PDPC) method. The intra prediction unit 254 is configured to generate an intra prediction block 265 using the reconstructed samples of adjacent blocks of the same current picture according to the intra prediction mode of the set of intra prediction modes.

[0121] The intra prediction unit 254 (or, generally, the mode selection unit 260) is further configured to output an intra prediction parameter (or, generally, information indicating the selected intra prediction mode for a block) to the entropy encoding unit 270 in the form of a syntax element 266 for inclusion in the encoded picture data 21, such that, for example, the video decoder 30 can receive and use the prediction parameter for decoding.

[0122] A set of (or possible) intra prediction modes depends on available reference pictures (i.e., previously decoded pictures that are at least partially decoded, e.g., those stored in the DBP 230) and other inter prediction parameters, such as whether only the whole or a part of the reference picture, e.g., the search window area around the area of the current block of the reference picture, is used to search for the best matching reference block, and / or whether pixel interpolation such as half / semi - pel, quarter - pel, and / or 1 / 16 - pel interpolation is applied, etc.

[0123] In addition to the above prediction modes, a skip mode, a direct mode, and / or other inter prediction modes may be applied.

[0124] For example, in extended merge prediction, the merge candidate list in such a mode is constructed by including, in order, the following five types of candidates: spatial MVP from spatially adjacent CUs, temporal MVP from CUs at the same position, history-based MVP from the FIFO table, pairwise average MVP, and zero MV. Also, to improve the accuracy of the MV in the merge mode, bilateral matching-based decoder-side motion vector refinement (DMVR) may be applied. The merge mode with MVD (MMVD) is derived from the merge mode with motion vector difference. The MMVD flag is signaled to specify whether the MMVD mode is used for the CU immediately after sending the skip flag and the merge flag. And a CU-level adaptive motion vector resolution (AMVR) scheme may be applied. AMVR allows the MVD of the CU to be decoded with different accuracies. Depending on the prediction mode of the current CU, the MVD of the current CU can be adaptively selected. When the CU is encoded in the merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. To obtain the CIIP prediction, a weighted average of the inter and intra prediction signals is performed. For affine motion compensation prediction, the affine motion field of the block is described by the motion information of 2 control points (4 parameters) or 3 control point motion vectors (6 parameters). Sub-block-based temporal motion vector prediction (SbTMVP) is similar to the temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vectors of the sub-CUs in the current CU. The bi-directional optical flow (BDOF), previously called BIO, is a simpler version that requires much less computational effort, especially with respect to the number of multiplications and the size of the multiplier. In a mode such as the triangular partition mode, the CU is evenly divided into two triangular partitions using either a diagonal split or an anti-diagonal split. Additionally, the bi-directional prediction mode is extended beyond simple averaging to allow a weighted average of two prediction signals.

[0125] In addition to the above prediction mode, a skip mode and / or a direct mode may be applied.

[0126] The prediction processing unit 260 may further partition block 203 into smaller block partitions or sub-blocks, for example, by repeatedly using a quadtree partition (QT), a binary partition (BT), a ternary partition (TT), or a combination thereof, and is configured to perform prediction for each of the block partitions or sub-blocks, for example. The mode selection includes the selection of the tree structure of the partitioned block 203 and the prediction mode applied to each of the block partitions or sub-blocks.

[0127] The inter prediction unit 244 may include a motion estimation (ME) unit (not shown in FIG. 2) and a motion compensation (MC) unit (not shown in FIG. 2). The motion estimation unit is configured to receive or obtain, for motion estimation, the picture block 203 (the current picture block 203 of the current picture 201) and the decoded picture 231, or at least one or a plurality of previously reconstructed blocks, for example, the reconstructed blocks of one or a plurality of other / different previously decoded pictures 231. For example, the video sequence may include the current picture and the previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of the pictures forming the video sequence, or may form a sequence of pictures.

[0128] The encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different pictures of a plurality of other pictures, for example, and provide an offset (spatial offset) between the reference picture (or reference picture index,...) and / or the position (x, y coordinates) of the reference block and the position of the current block to a motion estimation unit (not shown in FIG. 2) as an inter prediction parameter. This offset is also called a motion vector (MV).

[0129] The motion compensation unit is configured to obtain, for example receive, an inter prediction parameter, perform an inter prediction based on or using the inter prediction parameter, and obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit can include fetching or generating a prediction block based on the motion / block vector determined by motion estimation, and potentially performing interpolation up to sub-pixel accuracy. Interpolation filtering can generate additional pixel samples from known pixel samples, and thus potentially increase the number of candidate prediction blocks that can be used to code a picture block. When receiving a motion vector for the PU of the current picture block, the motion compensation unit can locate the prediction block indicated by the motion vector in one of the reference picture lists.

[0130] The intra prediction unit 254 is configured to obtain, for example receive, the picture block 203 (current picture block) and one or more previously reconstructed blocks of the same picture, such as reconstructed adjacent blocks, for intra estimation. The encoder 20 can be configured to select an intra prediction mode, for example from a plurality of (predetermined) intra prediction modes.

[0131] Embodiments of the encoder 20 may be configured to select an intra prediction mode based on an optimization criterion, such as minimum residual (e.g., the intra prediction mode provides the prediction block 255 that is most similar to the current picture block 203) or minimum rate distortion.

[0132] The intra prediction unit 254 is further configured to determine based on intra prediction parameters, for example, a selected intra prediction mode, an intra prediction block 255. In any case, after selecting an intra prediction mode for a block, the intra prediction unit 254 is also configured to provide the intra prediction parameters, that is, information indicating the selected intra prediction mode for the block, to the entropy encoding unit 270. In one example, the intra prediction unit 254 may be configured to execute any combination of the intra prediction techniques described below.

[0133] The entropy encoding unit 270 applies an entropy encoding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC scheme (CALVC), arithmetic coding scheme, context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy encoding methods or techniques) to the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, and / or loop filter parameters, and individually or together (or not all), obtains the encoded picture data 21 that can be output by the output 272, for example, in the form of an encoded bitstream 21. The encoded bitstream 21 may be transmitted to the video decoder 30 or may be stored for later transmission or retrieval by the video decoder 30. The entropy encoding unit 270 can be further configured to entropy encode other syntax elements for the current video slice to be encoded.

[0134] Other structural variations of the video encoder 20 are possible for use in encoding the video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal for a particular block or frame without the transform processing unit 206. In another implementation, the encoder 20 can combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.

[0135] FIG. 3 shows an exemplary video decoder 30 configured to implement the realization of the present application. The video decoder 30 is configured to receive encoded picture data (e.g., an encoded bitstream) 21, such as that encoded by the encoder 100, to obtain a decoded picture 131. During the decoding process, the video decoder 30 receives video data, such as an encoded video slice and an encoded video bitstream representing the picture blocks of the associated syntax elements, from the video encoder 100.

[0136] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a buffer 330 for the decoded picture, and a prediction processing unit 360. The prediction processing unit 360 may include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. The video decoder 30 can execute a decoding path that is generally inverse to the encoding path described with respect to the video encoder 100 in FIG. 2 in some examples.

[0137] As described with respect to the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 344, and the intra prediction unit 354 are also referred to as forming the "built-in decoder" of the video encoder 20. Thus, the inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 212, the reconstruction unit 314 may be functionally identical to the reconstruction unit 214, the loop filter 320 may be functionally identical to the loop filter 220, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 230. Thus, the description made for each unit and function of the video encoder 20 is correspondingly applicable to each unit and function of the video decoder 30.

[0138] The entropy decoding unit 304 is configured to perform entropy decoding on the encoded picture data 21 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), such as (decoded) inter prediction parameters, intra prediction parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 is further configured to transfer the inter prediction parameters, intra prediction parameters, and / or other syntax elements to the prediction processing unit 360. The video decoder 30 can receive syntax elements at the video slice level and / or the video block level.

[0139] Entropy decoding unit 304 analyzes the bitstream 21 (or generally the encoded picture data 21), and performs, for example, entropy decoding on the encoded picture data 21 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), such as inter prediction parameters (e.g., reference picture index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. Entropy decoding unit 304 can be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme as described with respect to the entropy encoding unit 270 of encoder 20. Entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to mode application unit 360 and other parameters to other units of decoder 30. Video decoder 30 can receive syntax elements at the video slice level and / or video block level. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.

[0140] Inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 112, reconstruction unit 314 may be functionally identical, buffer 316 may be functionally identical to buffer 116, loop filter 320 may be functionally identical to loop filter 120, and buffer 330 for the decoded picture may be functionally identical to buffer 130 for the decoded picture.

[0141] An embodiment of the decoder 30 may include a partitioning unit (not shown in FIG. 3). In one example, the prediction processing unit 360 of the video decoder 30 may be configured to perform any combination of the partitioning techniques described above.

[0142] The prediction processing unit 360 includes an inter prediction unit 344 and an intra prediction unit 354. The inter prediction unit 344 may be functionally similar to the inter prediction unit 144, and the intra prediction unit 354 may be functionally similar to the intra prediction unit 154. The prediction processing unit 360 is typically configured to perform block prediction and / or obtain a prediction block 365 from the encoded data 21, and receive or obtain information regarding prediction-related parameters and / or selected prediction modes (explicitly or implicitly) from, for example, the entropy decoding unit 304.

[0143] When a video slice is encoded as an intra-coded (I) slice, the intra prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on data from previously decoded blocks of the current frame or picture and the signaled intra prediction mode. When a video frame is encoded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., the motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, the prediction block may be generated from one reference picture within one reference picture list. The video decoder 30 can configure the reference frame lists, List0 and List1, by using a default configuration technique based on the reference pictures stored in the DPB 330.

[0144] The prediction processing unit 360 is configured to determine prediction information for a video block of a current video slice by analyzing motion vectors and other syntax elements, and to generate a prediction block for the current video block to be decoded using the prediction information. For example, the prediction processing unit 360 uses a part of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used for coding a video block of a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists for the slice, a motion vector for each inter-coded video block of the slice, an inter prediction status for each inter-coded video block of the slice, and other information for decoding a video block within the current video slice.

[0145] The inverse quantization unit 310 is configured to inverse-quantize, i.e., de-quantize, the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The inverse quantization process may include using quantization parameters calculated by the video encoder 100 for each video block in the video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.

[0146] The inverse quantization unit 310 can be configured to receive a quantization parameter (QP) (or generally information regarding inverse quantization) and quantized coefficients from the encoded picture data 21 (e.g., by parsing and / or decoding, e.g., by the entropy decoding unit 304), and to apply inverse quantization to the decoded quantized coefficients based on the quantization parameter to obtain non-quantized coefficients 311, also referred to as transform coefficients 311.

[0147] The inverse transformation processing unit 312 is configured to apply an inverse transformation, such as an inverse DCT, an inverse integer transformation, or a conceptually similar inverse transformation process, to the transformation coefficients to generate a residual block in the pixel domain.

[0148] The inverse transformation processing unit 312 can be configured to receive the dequantized coefficients 311, also referred to as transformation coefficients 311, and apply a transformation to the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may also be referred to as the transformation block 313. The transformation may be an inverse transformation, such as an inverse DCT, an inverse DST, an inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may further be configured to receive transformation parameters or corresponding information from the encoded picture data 21 (e.g., by parsing and / or decoding, e.g., by the entropy decoding unit 304) to determine the transformation applied to the dequantized coefficients 311.

[0149] The reconstruction unit 314 (e.g., the adder 314) is configured to add the inverse transformation block 313 (i.e., the reconstructed residual block 313) to the prediction block 365 to obtain a reconstructed block 315 in the sample domain, e.g., by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365.

[0150] The loop filter unit 320 (either within or after the encoding loop) filters the reconstructed block 315 to obtain a filtered block 321, and is configured to, for example, smooth pixel transitions or otherwise improve video quality. The loop filter unit 320 can include a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as one or more loop filters like an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In one example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, SAO, and ALF. In another example, a process called luma mapping by chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is performed before deblocking. In another example, the deblocking filter process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra-sub-partition (ISP) edges. The loop filter unit 320 is shown as an in-loop filter in FIG. 3, but in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0151] Subsequently, the decoded video block 321 in a given frame or picture is stored in a decoded picture buffer 330 that stores reference pictures for subsequent motion compensation.

[0152] Subsequently, the decoded video block 321 of the picture is stored in the decoded picture buffer 330, which stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or outputs respectively.

[0153] The decoder 30 is configured to output, for example, via output 332, the decoded picture 331 for presentation or display to a user.

[0154] Other variations of the video encoder 30 can be used to decode the compressed bitstream. For example, the decoder 30 can generate an output video stream without the loop filtering unit 320. For example, a non-transform-based decoder 30 can directly inverse quantize the residual signal for a particular block or frame without the inverse transform processing unit 312. In another implementation, the video decoder 30 can combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.

[0155] Additionally or alternatively to the above-described embodiments, in another embodiment according to FIG. 16, the inter prediction unit 344 may be the same as the inter prediction unit 244 (in particular, the motion compensation unit), the intra prediction unit 354 may be functionally the same as the inter prediction unit 254, and based on each piece of information received from the partitioning and / or prediction parameters or the encoded picture data 21 (e.g., by parsing and / or decoding, by the entropy decoding unit 304), perform partitioning or partitioning determination and prediction. The mode application unit 360 may be configured to perform prediction (intra prediction or inter prediction) for each block based on the reconstructed picture, block, or each sample (filtered or unfiltered) to obtain the prediction block 365.

[0156] When a video slice is encoded as an intra-coded (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on data from previously decoded blocks of the current picture and the signaled intra prediction mode. When a video picture is encoded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vectors received from the entropy decoding unit 304 and other syntax elements. For inter prediction, the prediction block may be generated from one reference picture within one reference picture list. The video decoder 30 can construct the reference frame lists, List0 and List1, using a default construction technique based on the reference pictures stored in the DPB 330. The same or similar can be applied additionally or alternatively to a slice (e.g., video slice), to an embodiment using tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles), e.g., the video can be encoded using I, P or B tile groups and / or tiles.

[0157] The mode application unit 360 is configured to determine prediction information for a video block of a current video slice by analyzing motion vectors or related information and other syntax elements, and to generate a prediction block for the current video block to be decoded using the prediction information. For example, the mode application unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to code a video block of a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists of the slice, a motion vector for each inter-coded video block of the slice, an inter prediction status for each inter-coded video block of the slice, and other information for decoding a video block within the current video slice. The same or similar can be applied additionally or alternatively to a tile group (e.g., video tile group) and / or a tile (e.g., video tile) for a slice (e.g., video slice), for example, a video can be coded using I, P, or B tile groups and / or tiles.

[0158] The embodiment of the video decoder 30 shown in FIG. 3 can be configured to partition and / or decode a picture by using slices (also called video slices), and a picture can be partitioned or decoded into one or more slices (typically non-overlapping), and each slice can include one or more blocks (e.g., CTUs) or a group of one or more blocks (e.g., tiles (H.265 / HEVC and VVC) or bricks (VC)).

[0159] The embodiment of the video decoder 30 shown in FIG. 3 can be configured to partition and / or decode a picture by using slice / tile groups (also called video tile groups) and / or tiles (also called video tiles), where the picture can be partitioned or decoded into one or more slice / tile groups (typically non-overlapping), each slice / tile group can include, for example, one or more blocks (e.g., CTUs) or one or more tiles, each tile can be, for example, rectangular in shape and can include one or more blocks (e.g., CTUs), such as complete or partial blocks.

[0160] Other variations of the video decoder 30 can be used to decode the encoded picture data 21. For example, the decoder 30 can generate an output video stream without the loop filtering unit 320. For example, the non-transform-based decoder 30 can directly inverse quantize the residual signal for a particular block or frame without the inverse transform processing unit 312. In another implementation, the video decoder 30 can combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.

[0161] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further processing such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.

[0162] FIG. 4 is a schematic diagram of a video encoding device 400 according to an embodiment of the present disclosure. The video encoding device 400 is suitable for implementing the disclosed embodiments as described in the present application. In an embodiment, the video encoding device 400 may be a decoder such as the video decoder 30 of FIG. 1A, or an encoder such as the video encoder 20 of FIG. 1A. In an embodiment, the video encoding device 400 may be one or more components of the video decoder 30 or the video encoder 20 of FIG. 1A as described above.

[0163] The video encoding device 400 includes an input port 410 and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit 430 for processing data; a transmitter unit (Tx) 440 and an output port 450 for transmitting data; and a memory 460 for storing data. The video encoding device 400 may also include optical - electrical (OE) components and electro - optical (EO) components coupled to the input port 410, the receiver unit 420, the transmitter unit 440, and the output port 450 for the entry and exit of optical or electrical signals.

[0164] Processor 430 is implemented by hardware and software. Processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. Processor 430 communicates with an input port 410, a receiver unit 420, a transmitter unit 440, an output port 450, and a memory 460. Processor 430 includes a coding module 470. Coding module 470 implements the disclosed embodiments described above. For example, coding module 470 realizes, processes, prepares, or provides various coding operations. Therefore, including coding module 470 provides a significant improvement to the functions of video coding device 400 and affects the conversion of video coding device 400 to different states. Alternatively, coding module 470 is implemented as instructions stored in memory 460 and executed by processor 430.

[0165] Memory 460 includes one or more disks, tape drives, and solid state drives and is used as an overflow data storage device, storing programs when such programs are selected for execution and capable of storing instructions and data read during program execution. Memory 460 may be volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).

[0166] FIG. 5 is a simplified block diagram of an apparatus 500 that can be used as either or both of the source device 310 and the destination device 320 of FIG. 1 according to an exemplary embodiment. The apparatus 500 can implement the technology of the present application described above. The apparatus 500 can be in the form of a computing system including a plurality of computing devices, or in the form of a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.

[0167] The processor 502 within the device 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices capable of manipulating or processing information that currently exists or will be developed in the future. The disclosed implementation can be carried out using a single processor, such as processor 502 as illustrated, but multiple processors can be used to achieve advantages in speed and efficiency. The memory 504 within the device 500 can be a read-only memory (ROM) device or a random access memory (RAM) device in the implementation. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that enables the processor 502 to execute the methods described in this application. For example, the application program 510 can include applications 1 through N, which further include a video coding application that executes the methods described in this application. The device 500 can also include additional memory in the form of secondary storage 514, which can be, for example, a memory card used with a mobile computing device. Since video communication sessions can contain a significant amount of information, they can be stored, in whole or in part, in the secondary storage 514 and loaded into the memory 504 as needed for processing. Also, the device 500 can include one or more output devices such as a display 518. The display 518 can be, in one example, a touch-sensitive display that combines the display with a touch-sensitive element that operates to sense touch input. The display 518 can be coupled to the processor 502 via the bus 512.

[0168] Device 500 may also include one or more output devices such as display 518. In one example, display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element capable of sensing touch input. Display 518 can be coupled to processor 502 via bus 512. Other output devices that allow a user to program or otherwise use device 500 can be provided in addition to or in place of display 518. If the output device is or includes a display, the display can be implemented in a variety of ways including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light emitting diode (LED) display such as an organic LED (OLED) display.

[0169] Device 500 may also include or be communicable with an image sensing device 520, such as a camera, or any other existing or future developed image sensing device 520 capable of sensing an image such as an image of a user operating device 500. Image sensing device 520 can be positioned to be directed toward a user operating device 500. In one example, the position and optical axis of image sensing device 520 can be set such that the field of view includes an area directly adjacent to and from which display 518 is visible.

[0170] Device 500 may also include or be communicable with an acoustic sensing device 522, such as a microphone, or any other existing or future developed acoustic sensing device capable of sensing sound in the vicinity of device 500. Acoustic sensing device 522 can be positioned to be directed toward a user operating device 500 and configured to receive sound generated by the user, such as speech or other utterances, while the user operates device 500.

[0171] FIG. 5 depicts the processor 502 and the memory 504 of the apparatus 500 integrated into a single unit, although other configurations are possible. The operation of the processor 502 can be directly or distributed among a plurality of machines (each having one or more processors) that can be coupled via a local area or other network. The memory 504 can be distributed among a plurality of machines such as network-based memory or memory within a plurality of machines that execute the operation of the apparatus 500. Although depicted here as a single bus, the bus 512 of the apparatus 500 can be composed of a plurality of buses. Further, the secondary storage 514 can be directly coupled to other components of the apparatus 500 or accessed via a network and can include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Accordingly, the apparatus 500 can be implemented in a wide variety of configurations.

[0172] Next-generation video coding (NGVC) removes the distinction of the concepts of CU, PU, and TU and supports more flexibility with respect to the CU partition shape. The size of the CU corresponds to the size of the coding node and may be square or non-square (e.g., rectangular) in shape.

[0173] J. An et al., “Block partitioning structure for next generation video coding”, International Telecommunication Union, COM16-C966, September 2015 (hereinafter, “VCEG proposal COM16-C966”), In [reference], the quadtree binary tree (QTBT) partitioning technique has been proposed for future video coding standards beyond HEVC. Simulations have shown that the proposed QTBT structure is more efficient than the quadtree structure in the HEVC being used. In HEVC, for reducing the memory access of motion compensation, the inter prediction for small blocks is restricted, and the inter prediction for 4×4 blocks is not supported. In the QTBT of JEM, these restrictions have been removed.

[0174] In QTBT, a CU can have either a square or rectangular shape. As shown in Figure 6, a coding tree unit (CTU) is first partitioned by a quadtree structure. A quadtree leaf node can be further partitioned by a binary tree structure. In binary tree splitting, there are two splitting types: symmetric horizontal splitting and symmetric vertical splitting. In each case, the node is split by dividing it horizontally or vertically from the center of the node. A binary tree leaf node is called a coding unit (CU), and its segmentation is used for prediction and transform processing without any further partitioning. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. A CU is often composed of coding blocks (CBs) of different color components. For example, in the case of P and B slices with a 4:2:0 chroma format, one CU includes one luma CB and two chroma CBs, and it can also often be composed of a single-component CB. For example, in the case of an I slice, one CU includes only one luma CB or only two chroma CBs.

[0175] The following parameters are defined for the QTBT partitioning scheme. - CTU size: The size of the root node of the quadtree, which is the same concept as in HEVC. - MinQTSize: The minimum allowable quadtree leaf node size - MaxBTSize: Maximum allowable binary tree root node size - MaxBTDepth: Maximum allowable binary tree depth - MinBTSize: Minimum allowable binary tree leaf node size

[0176] In an example of the QTBT partitioning structure, when a quadtree node has a size less than or equal to MinQTSize, no further quadtree is considered. Since the size (MinQTSize) does not exceed MaxBTSize, it will not be further divided by a binary tree. Otherwise, the leaf quadtree node can be further partitioned by a binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree, which has a binary tree depth of 0 (zero). When the depth of the binary tree reaches MaxBTDepth (i.e., 4), no further division is considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal division is considered. Similarly, when a binary tree node has a height equal to MinBTSize, no further vertical division is considered. The leaf node of the binary tree is further processed by prediction and conversion processing without any further partitioning. In JEM, the maximum CTU size is 256×256 luma samples. The leaf node of the binary tree (CU) may be further processed without any further partitioning (e.g., by performing a prediction process and a conversion process).

[0177] FIG. 6 shows an example of a block 30 (e.g., CTB) partitioned using QTBT partitioning technology. As shown in FIG. 6, using QTBT partitioning technology, each block is symmetrically divided through the center of each block. FIG. 7 shows a tree structure corresponding to the block partitioning of FIG. 6. The solid lines in FIG. 7 indicate quadtree partitioning, and the dotted lines indicate binary tree partitioning. In one example, at each split (i.e., non-leaf) node of the binary tree, a syntax element (e.g., a flag) is signaled to indicate the type of split (e.g., horizontal or vertical) to be performed, where 0 indicates a horizontal split and 1 indicates a vertical split. In the case of quadtree partitioning, since the quadtree partitioning always divides the block into four sub-blocks of equal size horizontally and vertically, there is no need to specify the split type.

[0178] As shown in FIG. 7, at node 50, block 30 (corresponding to root 50) is divided into the four blocks 31, 32, 33, and 34 shown in FIG. 6 using QT partitioning. Block 34 is not further divided and is thus a leaf node. At node 52, block 31 is further divided into two blocks using BT partitioning. As shown in FIG. 7, node 52 is marked with a 1 indicating a vertical split. Thus, the split at node 52 results in block 37 and a block containing both blocks 35 and 36. Blocks 35 and 36 are generated by a further vertical split at node 54. At node 56, block 32 is further divided into two blocks 38 and 39 using BT partitioning.

[0179] At node 58, block 33 is divided into four equal-sized blocks using QT partitioning. Blocks 43 and 44 are generated from this QT partitioning and are not further divided. At node 60, the upper-left block is first divided using a vertical binary tree split, resulting in block 40 and a right vertical block. The right vertical block is then divided into blocks 41 and 42 using a horizontal binary tree split. The lower-right block created from the quadtree split at node 58 is divided into blocks 45 and 46 using a horizontal binary tree split at node 62. As shown in FIG. 7, node 62 is marked with a 0 indicating a horizontal split.

[0180] In addition to QTBT, a block partitioning structure named multi-type tree (MTT) is proposed to replace BT in the QTBT-based CU structure, which means that first, the CTU may be divided by QT partitioning to obtain the blocks of the CTU, and then the blocks may be secondarily divided by MTT partitioning.

[0181] The MTT partitioning structure is still a recursive tree structure. In MTT, multiple different partitioning structures (e.g., two or more) are used. For example, according to the MTT technique, two or more different partitioning structures may be used for each non-leaf node of the tree structure at each depth of the tree structure. The depth of a node in the tree structure can indicate the length of the path (e.g., the number of splits) from the node to the root of the tree structure.

[0182] In MTT, there are two partition types: BT partitioning and ternary tree (TT) partitioning. The partition type can be selected from BT partitioning and TT partitioning. The TT partition structure is different from those of the QT and BT structures in that it does not divide the blocks from the center. The central region of the block remains together within the same sub-block. Different from the QT that generates four blocks or the binary tree that generates two blocks, the division by the TT partition structure generates three blocks. Exemplary partition types by the TT partition structure include asymmetric partition types (both horizontal and vertical) in addition to symmetric partition types (both horizontal and vertical). Further, the symmetric partition types by the TT partition structure may be non-uniform / non-uniform or uniform / uniform. The asymmetric partition types by the TT partition structure are non-uniform / non-uniform. In one example, the TT partition structure can include at least one of the following partition types: horizontal uniform / uniform symmetric ternary tree, vertical uniform / uniform symmetric ternary tree, horizontal non-uniform / non-uniform symmetric ternary tree, vertical non-uniform / non-uniform symmetric ternary tree, horizontal non-uniform / non-uniform asymmetric ternary tree, or vertical non-uniform / non-uniform asymmetric ternary tree.

[0183] Generally, a non-uniform / non-uniform symmetric ternary partition type is a partition type that is symmetric with respect to the center line of the block, but in this case, at least one of the resulting three blocks is not the same size as the other two. One preferred example is when the side blocks are 1 / 4 the size of the block and the central block is 1 / 2 the size of the block. A uniform / uniform symmetric ternary partition type is a partition type that is symmetric with respect to the center line of the block, and all of the resulting blocks are the same size. Such a partition is possible when the height or width of the block is a multiple of 3, depending on whether it is a vertical or horizontal split. A non-uniform / non-uniform asymmetric ternary partition type is a partition type that is not symmetric with respect to the center line of the block, and in this case, at least one of the resulting blocks is not the same size as the other two.

[0184] FIG. 8 is a conceptual diagram showing an exemplary horizontal ternary partition type of an option. FIG. 9 is a conceptual diagram showing an exemplary vertical ternary partition type of an option. In both FIGS. 8 and 9, h represents the height of the block in the luma or chroma samples, and w represents the width of the block in the luma or chroma samples. Note that the center line of each block does not represent the boundary of the block (i.e., the ternary partition does not divide the block through the center line). Rather, the center line \ is used to indicate whether a particular partition type is symmetric or asymmetric with respect to the center line of the original block. Also, the center line runs along the direction of the split.

[0185] As shown in FIG. 8, block 71 is partitioned into a horizontal uniform / uniform symmetric partition type. The horizontal uniform / uniform symmetric partition type generates upper and lower halves that are symmetric with respect to the center line of block 71. The horizontal uniform / uniform symmetric partition type generates three sub-blocks of equal size, each having a height of h / 3 and a width of w. The horizontal uniform / uniform symmetric partition type is possible when the height of block 71 is evenly divisible by 3.

[0186] Block 73 is partitioned into a horizontally non-uniform / non-uniform symmetric partition type. The horizontally non-uniform / non-uniform symmetric partition type generates upper and lower halves that are symmetric with respect to the center line of block 73. The horizontally non-uniform / non-uniform symmetric partition type generates two blocks of equal size (e.g., upper and lower blocks having a height of h / 4) and a central block of a different size (e.g., a central block having a height of h / 2). In one example, according to the horizontally non-uniform / non-uniform symmetric partition type, the area of the central block is equal to the total area of the upper and lower blocks. In some examples, the horizontally non-uniform / non-uniform symmetric partition type may be preferred for blocks having a height that is a power of 2 (e.g., 2, 4, 8, 16, 32, etc.).

[0187] Block 75 is partitioned into a horizontally non-uniform / non-uniform asymmetric partition type. The horizontally non-uniform / non-uniform asymmetric partition type does not generate upper and lower halves that are symmetric with respect to the center line of block 75 (i.e., the upper and lower halves are asymmetric). In the example of FIG. 8, the horizontally non-uniform / non-uniform asymmetric partition type generates an upper block having a height of h / 4, a central block having a height of 3h / 8, and a lower block having a height of 3h / 8. Of course, other asymmetric arrangements may be used.

[0188] As shown in FIG. 9, block 81 is partitioned into a vertically uniform / uniform symmetric partition type. The vertically uniform / uniform symmetric partition type generates left and right halves that are symmetric with respect to the center line of block 81. The vertically uniform / uniform symmetric partition type generates three sub-blocks of equal size, each having a width of w / 3 and a height of h. The vertically uniform / uniform symmetric partition type is possible when the width of block 81 is evenly divisible by 3.

[0189] Block 83 is partitioned into a vertically non-uniform / non-uniform symmetric partition type. The vertically non-uniform / non-uniform symmetric partition type generates left and right halves that are symmetric with respect to the center line of block 83. The vertically non-uniform / non-uniform symmetric partition type generates left and right halves that are symmetric with respect to the center line of 83. The vertically non-uniform / non-uniform symmetric partition type generates two blocks of equal size (e.g., a left and a right block with a width of w / 4) and a central block of a different size (e.g., a central block with a width of w / 2). In one example, according to the vertically non-uniform / non-uniform symmetric partition type, the area of the central block is equal to the total area of the left and right blocks. In some examples, the vertically non-uniform / non-uniform symmetric partition type may be preferred for blocks having a width that is a power of 2 (e.g., 2, 4, 8, 16, 32, etc.).

[0190] Block 85 is partitioned into a vertically non-uniform / non-uniform asymmetric partition type. The vertically non-uniform / non-uniform asymmetric partition type does not generate left and right halves that are symmetric with respect to the center line of block 85 (i.e., the left and right halves are asymmetric). In the example of FIG. 9, the vertically non-uniform / non-uniform asymmetric partition type generates a left block with a width of w / 4, a central block with a width of 3w / 8, and a right block with a width of 3w / 8. Of course, other asymmetric arrangements may be used.

[0191] In addition to (or alternatively to) the parameters of the QTBT defined above, the following parameters are defined for the MTT partitioning scheme: - MaxBTSize: Maximum allowable binary tree root node size - MinBtSize: Minimum allowable binary tree root node size - MaxMttDepth: Maximum multi-type tree depth - MaxMttDepth offset: Maximum multi-type tree depth offset - MaxTtSize: Maximum allowable ternary tree root node size - MinTtSize: Minimum allowable ternary tree root node size - MinCbSize: Minimum allowable coding block size

[0192] Embodiments of the present disclosure can be implemented by a video encoder or a video decoder such as the video encoder 20 of FIG. 2 or the video decoder 30 of FIG. 3 according to the embodiments of the present application. One or more structural elements of the video encoder 20 or the video decoder 30 including a partition unit can be configured to execute the technology of the embodiments of the disclosure.

[0193] In the embodiments of the disclosure: In JVET-K1001-v4, log2_ctu_size_minus2, log2_min_qt_size_intra_slices_minus2, and log2_min_qt_size_inter_slices_minus2 are signaled in the SPS (as syntax elements).

[0194] The parameter log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU. Specifically: CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-5) CtbSizeY = 1 << CtbLog2SizeY (7-6)

[0195] In other words, CtbLog2SizeY specifies the log2 value of the CTU size CtbSizeY corresponding to the coding tree block (CTB) size for luma (Y).

[0196] Further settings are provided as follows: MinCbLog2SizeY = 2 (7-7) MinCbSizeY = 1 << MinCbLog2SizeY (7-8) MinTbSizeY = 4 (7-9) MaxTbSizeY = 64 (7-10)

[0197] The parameter log2_min_qt_size_intra_slices_minus2 plus 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of a CTU in a slice having slice_type equal to 2(I), i.e., an intra slice. The value of log2_min_qt_size_intra_slices_minus2 shall be within the range including both ends of 0 to CtbLog2SizeY - 2. MinQtLog2SizeIntraY = log2_min_qt_size_intra_slices_minus2 + 2 (7-22)

[0198] The parameter log2_min_qt_size_inter_slices_minus2 plus 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of a CTU in a slice having slice_type equal to 0(B) or 1(P), i.e., an inter slice. The value of log2_min_qt_size_inter_slices_minus2 shall be within the range including both ends of 0 to CtbLog2SizeY - 2. MinQtLog2SizeInterY = log2_min_qt_size_inter_slices_minus2 + 2 (7-23)

[0199] MinQtSizeY is defined in (7-30), which means the minimum allowable quadtree partition size in luma samples. When the coding block size is less than or equal to MinQtSizeY, quadtree partitioning is not allowed. Further settings are provided as follows: MinQtLog2SizeY = ( slice_type == I )? MinQtLog2SizeIntraY : MinQtLog2SizeInterY (7-25) MaxBtLog2SizeY = CtbLog2SizeY - log2_diff_ctu_max_bt_size (7-26) MinBtLog2SizeY = MinCbLog2SizeY (7-27) MaxTtLog2SizeY = ( slice_type == I )? 5 : 6 (7-28) MinTtLog2SizeY = MinCbLog2SizeY (7-29) MinQtSizeY = 1 << MinQtLog2SizeY (7-30) MaxBtSizeY = 1 << MaxBtLog2SizeY (7-31) MinBtSizeY = 1 << MinBtLog2SizeY (7-32) MaxTtSizeY = 1 << MaxTtLog2SizeY (7-33) MinTtSizeY = 1 << MinTtLog2SizeY (7-34) MaxMttDepth = ( slice_type == I )? max_mtt_hierarchy_depth_intra_slices : max_mtt_hierarchy_depth_inter_slices (7-35)

[0200] The parameters max_mtt_hierarchy_depth_intra_slices and max_mtt_hierarchy_depth_inter_slices indicate the maximum hierarchical depth of MTT type splitting for intra and inter slices, respectively.

[0201] Based on the semantics of log2_min_qt_size_intra_slices_minus2 and log2_min_qt_size_inter_slices_minus2, the ranges of log2_min_qt_size_intra_slices_minus2 and log2_min_qt_size_inter_slices_minus2 are from 0 to CtbLog2SizeY - 2. Here, CtbLog2SizeY is defined by the semantics of log2_ctu_size_minus2, which means the log2 value of the luma coding tree block size of each CTU, and CtbLog2SizeY in VTM2.0 is equal to 7.

[0202] Based on (7-22) and (7-23), the ranges of MinQtLog2SizeIntraY and MinQtLog2SizeInterY are from 2 to CtbLog2SizeY.

[0203] Based on (7-25), the range of MinQtLog2SizeY is from 2 to CtbLog2SizeY.

[0204] Based on (7-30), the range of MinQtSizeY is from (1<<2) to (1<<CtbLog2SizeY) in JVET-K1001-v4, and the range in VTM2.0 is from (1<<2) to (1<<7), which is equal to 4 to 128.

[0205] In JVET-K1001-v4, log2_diff_ctu_max_bt_size is conditionally signaled in the slice header.

[0206] The parameter log2_diff_ctu_max_bt_size specifies the difference between the maximum luma size (width or height) of a coding block that can be split using binary splitting and the luma CTB size. The value of log2_diff_ctu_max_bt_size shall be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY.

[0207] If log2_diff_ctu_max_bt_size does not exist, the value of log2_diff_ctu_max_bt_size is assumed to be equal to 2.

[0208] MinCbLog2SizeY is defined in (7-7), which means the minimum allowable coding block size.

[0209] Based on the semantics of log2_diff_ctu_max_bt_size, the range of log2_diff_ctu_max_bt_size is from 0 to CtbLog2SizeY - MinCbLog2SizeY.

[0210] Based on (7-26), the range of MaxBtLog2SizeY is from CtbLog2SizeY to MinCbLog2SizeY.

[0211] Based on (7-31), the range of MaxBtSizeY is from (1<< CtbLog2SizeY ) to (1<< MinCbLog2SizeY).

[0212] Based on (7-7), in JVET-K1001-v4, the range of MaxBtSizeY is from (1<< CtbLog2SizeY ) to (1<< 2), and since CtbLog2SizeY is equal to 7 in VTM2.0, the range of MaxBtSizeY in VTM2.0 is equal to 128 to 4.

[0213] Therefore, MinQtSizeY ranges from 4 to (1 << CtbLog2SizeY), which is from 4 to 128 in VTM2.0, and MaxBtSizeY ranges from (1 << CtbLog2SizeY) to 4, which is from 128 to 4 in VTM2.0.

[0214] Therefore, MinQtSizeY can be larger than MaxBtSizeY.

[0215] Furthermore, based on the current boundary processing in VVC2.0, only QT and BT partitioning are allowed for blocks located at the boundary (TT is not allowed, and non - splitting is not allowed).

[0216] If the current coding block is on the boundary and the current coding block size cbSizeY satisfies the condition: MinQtSizeY > cbSizeY > MaxBtSizeY then neither QT splitting nor BT splitting is possible for the current coding block. Therefore, there is no available partitioning mode for the current block.

[0217] Embodiment 1 The solutions (embodiments of the present invention) to the above problems including boundary - case issues are described in more detail below.

[0218] According to the embodiment, in order to solve the above problems, the lower limit of MaxBtSizeY should be restricted to MinQtSizeY to ensure that MaxBtSizeY is not smaller than MinQtSizeY. In particular, the lower limit of MaxBtSizeY may be equal to MinQtSizeY. Therefore, the range of MaxBtSizeY should be from (1 << CtbLog2SizeY) to (1 << MinQtLog2SizeY). Accordingly, the range of MaxBtLog2SizeY should be from CtbLog2SizeY to MinQtLog2SizeY. Accordingly, the range of log2_diff_ctu_max_bt_size should be from 0 to CtbLog2SizeY - MinQtLog2SizeY. Therefore, the information of MinQtSizeY may be used to determine the validity of MaxBtSizeY. In other words, MaxBtSizeY can be determined based on the information of MinQtSizeY.

[0219] The corresponding change in the draft text (of the video standard) is as follows in terms of the semantics of log2_diff_ctu_max_bt_size: log2_diff_ctu_max_bt_size specifies the difference between the maximum luma size (width or height) of a coding block that can be split using binary splitting and the luma CTB size. The value of log2_diff_ctu_max_bt_size shall be within the range including both ends of 0 to CtbLog2SizeY - MinQtLog2SizeY.

[0220] The corresponding method of coding implemented by a coding device (decoder or encoder) can be as follows: Determine whether the current block of the picture is a boundary block; Determine whether the size of the current block is larger than the minimum allowable quadtree leaf node size; When the current block is a boundary block and the size of the current block is not greater than the minimum allowable quadtree leaf node size, apply binary splitting to the current block; the minimum allowable quadtree leaf node size (MinQtSizeY) is not greater than the maximum allowable binary tree root node size (MaxBtSizeY). In this case, applying binary splitting to the current block may include applying forced binary splitting to the current block. Here, the coding corresponds to image, video, or video coding.

[0221] Being a boundary block means that the image / frame boundary cuts the block, i.e., in other words, the block is at the image / frame boundary. In the above embodiment, when the current block is a boundary block (condition 1) and its size is not greater than the minimum allowable quadtree leaf node size (condition 2), binary splitting is applied to the current block. Note that in some embodiments, ternary splitting or other splitting may be used instead of binary splitting. Further, in some embodiments, regardless of condition 1, binary splitting may be applied under condition 2. In other words, condition 1 does not need to be evaluated. If the size of the current block is actually greater than the minimum allowable quadtree leaf node size (i.e., condition 2 is not met), quadtree splitting may be applied.

[0222] Note that there are embodiments where binary splitting is used only for boundary blocks (condition 1). For non-boundary blocks, quadtree splitting may be the only splitting used. Applying binary (or ternary) splitting to the image / frame boundary provides the advantage of potentially more efficient splitting, such as horizontal binary / ternary splitting at the horizontal boundary and vertical binary / ternary splitting at the vertical boundary.

[0223] Another corresponding method of coding implemented by a coding device (decoder or encoder) can be as follows: Determine whether the size of the boundary block is larger than the minimum allowable quadtree leaf node size, and when the size of the boundary block is not larger than the minimum allowable quadtree leaf node size, and the minimum allowable quadtree leaf node size is not larger than the maximum allowable binary tree root node size (e.g., as defined by the specification), a binary split is applied to the boundary block.

[0224] Optionally, the boundary block may not include corner blocks. In other words, corner blocks that are cut at both the vertical and horizontal image / frame boundaries are not considered boundary blocks for the purposes of condition 1 above.

[0225] Embodiment 2 Other embodiments of the disclosure (which can be combined with the embodiments described above) are described below.

[0226] In JVET-K1001-v4, max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices are signaled in the SPS. In other words, max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices are syntax elements, meaning that their values are included in the bitstream including the encoded image or video.

[0227] In particular, max_mtt_hierarchy_depth_inter_slices specifies the maximum hierarchical depth for coding units resulting from the multi-type tree partitioning of quadtree leaves in slices having a slice_type equal to 0 (B) or 1 (P). The value of max_mtt_hierarchy_depth_inter_slices shall be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY.

[0228] max_mtt_hierarchy_depth_intra_slices specifies the maximum hierarchical depth for coding units resulting from the multi-type tree partitioning of quadtree leaves in slices having a slice_type equal to 2 (I). The value of max_mtt_hierarchy_depth_intra_slices shall be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY.

[0229] MinTbSizeY is defined by (7 - 9), which is fixed to 4, and thus MinTbLog2SizeY = log2 MinTbSizeY, which is fixed to 2.

[0230] MaxMttDepth is defined, which means the maximum allowable depth of the multi-type tree partition. If the current multi-type tree partition depth is greater than or equal to MaxMttDepth, the multi-type tree partition is not allowed (applied).

[0231] Based on the semantics of max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices, the range of max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices is from 0 to CtbLog2SizeY - MinTbLog2SizeY.

[0232] (7 - 35) Based on this, the range of MaxMttDepth is from 0 to CtbLog2SizeY - MinTbLog2SizeY. In VTM2.0, since CtbLog2SizeY is equal to 7, the range of MaxMttDepth is from 0 to 5.

[0233] Therefore, MaxMttDepth has a range from 0 to CtbLog2SizeY - MinTbLog2SizeY, and in VTM2.0, it has a range from 0 to 5.

[0234] Based on the current boundary processing in VVC2.0, only QT and BT partitioning are allowed for blocks located at the boundary (TT is not allowed, and non - splitting is not allowed).

[0235] If the above - mentioned first problem is solved (MaxBtSizeY >= MinQtSizeY), then the following conditions are further satisfied: cbSizeY <= MinQtSizeY MaxMttDepth = 0

[0236] There is no sufficient level of BT (generally, any MTT including TT) partitioning for boundary processing.

[0237] For example, MinQtSizeY is equal to 16, MinTbSizeY is equal to 4, and MaxMttDepth is 0.

[0238] If the boundary block has cbSizeY = 16, the parent partition is QT, and this block is still located at the boundary, then the Mttdepth of the current block reaches MaxMttDepth, so no further partitioning can be performed.

[0239] Solution to this problem of the boundary case (Embodiment of the present invention): To solve the above problem, the lower limit of MaxMttDepth is restricted to 1 (in other words, it cannot take a value of zero), and after the QT partition, there should be a sufficient level of multi-type tree partition for the boundary case. Alternatively, further, the lower limit of MaxMttDepth is restricted to (MinQtLog2SizeY - MinTbLog2SizeY), and after the QT partition, there should be a sufficient level of multi-type tree partition for both boundary and non-boundary cases.

[0240] The corresponding changes in the (standard) draft text are as follows in terms of the semantics of max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices: max_mtt_hierarchy_depth_inter_slices specifies the maximum hierarchical depth for the coding units resulting from the multi-type tree splitting of the quadtree leaves in slices having a slice_type equal to 0 (B) or 1 (P). The value of max_mtt_hierarchy_depth_inter_slices shall be within the range including both ends of 1 to CtbLog2SizeY - MinTbLog2SizeY. max_mtt_hierarchy_depth_intra_slices specifies the maximum hierarchical depth for the coding units resulting from the multi-type tree splitting of the quadtree leaves in slices having a slice_type equal to 2 (I). The value of max_mtt_hierarchy_depth_intra_slices shall be within the range including both ends of 1 to CtbLog2SizeY - MinTbLog2SizeY. Alternatively, max_mtt_hierarchy_depth_inter_slices specifies the maximum hierarchical depth for coding units resulting from the multi-type tree splitting of quadtree leaves in slices having a slice_type equal to 0(B) or 1(P). The value of max_mtt_hierarchy_depth_inter_slices shall be within the range including both ends of MinQtLog2SizeY - MinTbLog2SizeY to CtbLog2SizeY - MinTbLog2SizeY. max_mtt_hierarchy_depth_intra_slices specifies the maximum hierarchical depth for coding units resulting from the multi-type tree splitting of quadtree leaves in slices having a slice_type equal to 2(I). The value of max_mtt_hierarchy_depth_intra_slices shall be within the range including both ends of MinQtLog2SizeY - MinTbLog2SizeY to CtbLog2SizeY - MinTbLog2SizeY.

[0241] The corresponding method of coding implemented by a coding device (decoder or encoder) can be as follows: Divide the image into blocks, where the blocks include boundary blocks. Apply a binary split to the boundary blocks having the maximum boundary multi-type partition depth. The maximum boundary multi-type partition depth is at least the sum of the maximum multi-type tree depth and the maximum multi-type tree depth offset, and the maximum multi-type tree depth is greater than 0. This embodiment may be combined with Embodiment 1 or applied independently of Embodiment 1.

[0242] Optionally, when applying a binary split to the boundary blocks, the maximum multi-type tree depth is greater than 0.

[0243] Optionally, the boundary blocks may not include corner blocks.

[0244] Embodiment 3 Another embodiment of the disclosure:

[0245] In JVET-K1001-v4, when MinQtSizeY > MaxBtSizeY and MinQtSizeY > MaxTtSizeY

[0246] When cbSize = MinQtsizeY, there is no possible available partitioning mode, so the partition cannot reach MinCbSizeY (MinTbSizeY and MinCbsizeY are fixed and equal to 4).

[0247] Solution to this problem in non-boundary or boundary cases: To solve the above problem, the lower limit of MaxBtSizeY should be restricted to MinQtSizeY to ensure that MaxBtSizeY is not smaller than MinQtSizeY. Alternatively, the lower limit of MaxTtSizeY should be restricted to MinQtSizeY to ensure that MaxTtSizeY is not smaller than MinQtSizeY.

[0248] The corresponding changes in the draft text are semantic log2_diff_ctu_max_bt_size specifies the difference between the maximum luma size (width or height) of the coding block that can be split using binary splitting and the luma CTB size. The value of log2_diff_ctu_max_bt_size shall be within the range including both ends of 0 to CtbLog2SizeY - MinQtLog2SizeY. and / or log2_min_qt_size_intra_slices_minus2 plus 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of a CTU in a slice having a slice_type equal to 2(I). The value of log2_min_qt_size_intra_slices_minus2 shall be within the range including both ends from 0 to MaxTtLog2SizeY - 2. log2_min_qt_size_inter_slices_minus2 plus 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of a CTU in a slice having a slice_type equal to 0(B) or 1(P). The value of log2_min_qt_size_inter_slices_minus2 shall be within the range including both ends from 0 to MaxTtLog2SizeY - 2.

[0249] The corresponding method of coding implemented by a coding device (decoder or encoder) can be as follows: Determine whether the size of the current block is larger than the minimum allowable quadtree leaf node size; Apply multi-type tree partitioning to the current block when the size of the current block is not larger than the minimum allowable quadtree leaf node size; In this case, the leaf node size of the minimum allowable quadtree is not larger than the maximum allowable binary tree root node size, or the leaf node size of the minimum allowable quadtree is not larger than the maximum allowable ternary tree root node size.

[0250] Optionally, the leaf node size of the minimum allowable quadtree is not larger than the maximum allowable binary tree root node size, and the leaf node size of the minimum allowable quadtree is not larger than the maximum allowable ternary tree root node size.

[0251] As an option, applying multi-type tree splitting to the current block includes applying a three-way split to the current block or applying a two-way split to the current block.

[0252] As an option, the boundary block may not include a corner block.

[0253] Embodiment 4 In another embodiment of the disclosure:

[0254] When MaxBtSizeY >= MinQtSizeY, MinQtSizeY > MinTbLog2SizeY and MaxMttDepth < (MinQtLog2SizeY - MinTbLog2SizeY), If cbSize = MinQtsizeY, there is no sufficient level of acceptable multi-type tree partitioning, so the partition cannot reach MinCbSizeY.

[0255] Solution to this problem in non-boundary or boundary cases: To solve the above problem, the lower limit of MaxMttDepth is limited to (MinQtLog2SizeY - MinTbLog2SizeY), and after QT partitioning, there should be a sufficient level of multi-type tree partitioning for both boundary and non-boundary cases.

[0256] The corresponding changes in the draft text are as follows in terms of the semantics of max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices: max_mtt_hierarchy_depth_inter_slices specifies the maximum hierarchical depth for coding units resulting from the multi-type tree partitioning of quadtree leaves in slices having a slice_type equal to 0 (B) or 1 (P). The value of max_mtt_hierarchy_depth_inter_slices shall be within the range including both ends of MinQtLog2SizeY - MinTbLog2SizeY to CtbLog2SizeY - MinTbLog2SizeY. max_mtt_hierarchy_depth_intra_slices specifies the maximum hierarchical depth for coding units resulting from the multi-type tree partitioning of quadtree leaves in slices having a slice_type equal to 2 (I). The value of max_mtt_hierarchy_depth_intra_slices shall be within the range including both ends of MinQtLog2SizeY - MinTbLog2SizeY to CtbLog2SizeY - MinTbLog2SizeY.

[0257] The corresponding method of coding implemented by a coding device (decoder or encoder) can be as follows: Divide the image into blocks; Apply multi-type tree partitioning to the blocks among the blocks having the final maximum multi-type tree depth. The final maximum multi-type tree depth is at least the sum of the maximum multi-type tree depth and the maximum multi-type tree depth offset. The maximum multi-type tree depth is greater than or equal to the Log2 value of the minimum allowable quadtree leaf node size minus the Log2 value of the minimum allowable transform block size, or the maximum multi-type tree depth is greater than or equal to the Log2 value of the minimum allowable quadtree leaf node size minus the Log2 value of the minimum allowable coding block size.

[0258] Optionally, the block is a non-boundary block.

[0259] As an option, the maximum multi-type tree depth offset is 0.

[0260] As an option, the block is a boundary block and the multi-type tree split is a binary split.

[0261] As an option, the multi-type tree split is a ternary split (or includes it).

[0262] As an option, the boundary block may not include a corner block.

[0263] Embodiments 1 to 4 can be applied on the encoder side to divide an image / frame into coding units and to code the coding units. Embodiments 1 to 4 can be applied on the decoder side to provide a partition of an image / frame, i.e., a coding unit, and accordingly to decode the coding unit (e.g., correctly parse the coding units from a stream and decode them).

[0264] According to some embodiments, a decoder is provided that includes one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the decoder to perform any of the methods described above in connection with Embodiments 1 to 4 when executed by the processors.

[0265] Furthermore, an encoder is provided that includes one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the encoder to perform any of the methods described above in connection with Embodiments 1 to 4 when executed by the processors.

[0266] Further Embodiments Related to Boundary Partitioning In VVC, the multi-type (binary / trinary / quaternary) tree (BT / TT / QT or binary tree / trinary tree / quaternary tree) segmentation structure may or may be able to replace the concept of multiple partition unit types, that is, except when required for a CU having a size too large for the maximum transform length, it removes the distinction between the concepts of CU, PU, and TU and supports more flexibility for the CU partition shape. [J]

[0267] FIGS. 10A-F show, as an example, the partition modes currently used in VTM. FIG. 10A shows a non-divided block (no division), FIG. 10B shows a quaternary or quaternary tree (QT) partitioning, FIG. 10C shows a horizontal binary or binary tree (BT) partitioning, FIG. 10D shows a vertical binary or binary tree (BT) partitioning, FIG. 10E shows a horizontal trinary or trinary tree (TT) partitioning, and FIG. 10F shows a vertical trinary or trinary tree (TT) partitioning of a block such as a CU or CTU. Embodiments may be configured to implement the partition mode as shown in FIGS. 10A to 10F.

[0268] In an embodiment, the following parameters can be defined and specified by the sequence parameter set (SPS) syntax elements for the BT / TT / QT coding tree scheme: - CTU size: The root node size of the quaternary tree - MinQTSize: The minimum allowable quaternary tree leaf node size - MaxBTTSize: The maximum allowable binary and trinary tree root node size - MaxBTTDepth: The maximum allowable binary and trinary tree depth, and - MinBTTSize: The minimum allowable binary and trinary tree leaf node size

[0269] In other embodiments, the minimum allowable quad-tree leaf node size MinQTSize parameter may be included in other headers or sets, such as a slice header (SH) or a picture parameter set (PPS).

[0270] In the HEVC standard, coding tree units (CTUs) or coding units (CUs) located on slice / picture boundaries will be forced to be split using a quadtree (QT) until the bottom-right sample of the leaf node is located within the slice / picture boundary. The forced QT partitioning or partitioning does not need to be signaled in the bitstream because both the encoder and the decoder, such as both the video encoder 20 and the video decoder 30, know when to apply the forced QT. The purpose of the forced partitioning is to enable the boundary CTUs / CUs by the video encoder 20 / video decoder 30.

[0271] International Patent Publication No. WO2016 / 090568 discloses a QTBT (quadtree plus binary tree) structure, and in VTM1.0, the forced partitioning process of boundary CTUs / CUs is inherited from HEVC. This means that CTUs / CUs located on frame boundaries are forced to be partitioned by a quadtree (QT) structure without considering rate distortion (RD) optimization until the entire current CU is within the picture boundary. These forced partitions are not signaled in the bitstream.

[0272] FIG. 11A shows an example of the forced partitioning of a bottom boundary CTU (128×128) of a high definition (HD) (1920×1080 pixels) forced by a forced QT. In FIG. 11, the HD picture has or is 1920×1080 pixels, and the CTU has or is 128×128 pixels.

[0273] San Diego Conference (April 2018) [In SubCE2 (picture boundary processing) of CE1 (partitioning) in JVET-J1021, 15 tests were proposed for picture boundary processing using BT, TT, or ABT (asymmetric BT). For example, in JVET-K0280 and JVET-K0376, the boundaries are defined as shown in Figure 12. Figure 12 shows the picture boundary by dot-hash lines and the areas of boundary cases in a straight line, i.e., the bottom boundary case, the corner boundary case, and the right boundary case. The bottom boundary can be divided by horizontal forced BT or forced QT, the right boundary can be divided by vertical forced BT or forced QT, and the corner case can only be divided by forced QT. Here, the decision of which forced BT or forced QT partitioning to use is based on the rate-distortion optimization criterion and is signaled in the bitstream. Forced partitioning means that the block must be divided. For example, forced partitioning is applied to boundary blocks that may not be coded using "non-partitioning" as shown in Figure 10A.

[0274] When forced QT splitting is used in forced boundary partitioning, the partitioning constraint of MinQTSize is ignored. For example, in Figure 13A, if MinQTSize is signaled as 32 in the SPS, QT splitting down to block size 8x8 will be required to match the boundary with the forced QT method. This ignores the constraint of MinQTSize being 32.

[0275] According to embodiments of the present disclosure, when forced QT is used for picture boundary partitioning, the forced QT splitting, for example, follows the splitting constraints signaled in the SPS and, for example, is not ignored. Further, if more forced splitting is required, only forced BT is used, which may also be referred to as forced QTBT in combination. In embodiments of the present disclosure, for example, the partition constraint MinQTSize is considered for forced QT partitioning at picture boundaries, and no additional signaling for forced BT partitioning is required. Also, embodiments make it possible to reconcile partitioning for normal (non-boundary) blocks and boundary blocks. For example, in conventional solutions, two "MinQTSize" parameters are required, one for normal block partitioning and another for boundary block partitioning. Embodiments only require one common "MinQTSize" parameter for both normal block and boundary block partitioning, which can be flexibly set between the encoder and the decoder, for example, by signaling one "MinQTSize" parameter. Further, embodiments require fewer partitions, for example, than forced QT.

[0276] Solutions for the bottom boundary case and the right boundary case In the bottom and right boundary cases, when the block size is larger than MinQTSize, the partition mode for picture boundary partitioning can be selected between forced BT partitioning and forced QT partitioning, for example, based on RDO (rate distortion optimization). Otherwise (i.e., when the block size is less than or equal to MinQTSize), only forced BT partitioning is used for picture boundary partitioning. More specifically, horizontal forced BT is used for the bottom boundary for each boundary block located at the bottom boundary of the picture, and vertical forced BT is used for the right boundary for each boundary block located at the right boundary of the picture.

[0277] Compulsory BT partitioning may involve recursively partitioning the current block by horizontal compulsory boundary partitioning until the sub - partition of the current block is located at the lower boundary of the picture, and then recursively partitioning the sub - partition by vertical compulsory boundary partitioning until the leaf node is completely located at the right boundary of the picture. Alternatively, compulsory BT partitioning may involve recursively partitioning the current block by vertical compulsory boundary partitioning until the sub - partition of the current block is located at the lower boundary, and then recursively partitioning the sub - partition by horizontal compulsory boundary partitioning until the leaf node is completely located at the right boundary. MinQTSize may also be applied to control the partitioning of non - boundary blocks.

[0278] For example, in the case shown in FIG. 11A, if MinQTSize is 32 or limited to 32 and a rectangular (non - square) block size of 8 samples in height or width is required to match the picture boundary, compulsory BT partitioning will be used to partition the 32×32 block where the boundary is located. The BT partitions can be further partitioned using the same type of compulsory BT partitioning. For example, in a case where compulsory vertical BT partitioning is applied, only compulsory vertical BT partitioning will be further applied, and in a case where compulsory horizontal BT partitioning is applied, only compulsory horizontal BT partitioning will be further applied. Compulsory BT partitioning continues until the leaf node is completely within the picture.

[0279] FIG. 11B shows an exemplary partitioning of a bottom boundary CTU having a size of 128×128 samples according to an embodiment of the present invention. The bottom boundary CTU forming the root block or root node of the partitioning tree is divided into smaller partitions, for example, smaller blocks of square or rectangular size. These smaller partitions or blocks can be further divided into even smaller partitions or blocks. In FIG. 11B, the CTU is first quad-tree divided into four square blocks 710, 720, 730, and 740, each having a size of 64×64 samples. Of these blocks, blocks 710 and 720 are again bottom boundary blocks, while blocks 730 and 740 are outside the picture (located outside the picture respectively) and are not processed.

[0280] Block 710 is further divided using a quad-tree that divides it into four square blocks 750, 760, 770, and 780, each having a size of 32×32 samples. Blocks 750 and 760 are located inside the picture, while blocks 770 and 780 again form bottom boundary blocks. Since the size of block 770 is not larger than, for example, MinQTSize which is 32, recursive horizontal forced binary partitioning is applied to block 770 until the leaf node is completely within the picture or located completely within the picture, for example, until the leaf node block 772, a rectangular non-square block having 32×16 samples, is within the picture (after one horizontal bisection), or until the leaf node block 774, a rectangular non-square block located at the bottom boundary of the picture and having 32×8 samples, is within the picture (after two horizontal bisections). The same applies to block 780.

[0281] Embodiments of the present disclosure enable reconciling partitioning for normal blocks that are fully within a picture and partitioning of boundary blocks. A boundary block is a block that is not fully within the picture and not fully outside the picture. In other words, a boundary block is a block that includes a portion located within the picture and a portion located outside the picture. Further, embodiments of the present disclosure allow for reducing signaling since mandatory BT partitioning below MinQTSize need not be signaled.

[0282] Solutions for corner cases In corner cases, some approaches only allow for mandatory QT splitting, which also ignores the MinQTSize constraint. Embodiments of the present disclosure provide two solutions for corner cases. A corner case occurs when the block currently being processed is at the corner of the picture. This is the case when the current block is intersected or adjacent to two picture boundaries (vertical and horizontal).

[0283] Solution 1: The corner case can be considered as a bottom boundary case or a right boundary case. FIG. 14 shows an embodiment of boundary definition. FIG. 14 shows the boundaries of the picture with dot - hash lines and the area of the boundary case with straight lines. As shown, the corner case is defined as a bottom boundary case. Thus, the solution is the same as that described for the bottom boundary case and the right boundary case above. In other words, first, horizontal partitioning is applied (as described for the bottom boundary case) until the block or partition is fully within the picture (in the vertical direction), and then vertical partitioning is applied (as described for the right boundary case) until the leaf node is fully within the picture (in the horizontal direction). The boundary case may be a boundary block.

[0284] Solution 2: The definition of the boundary case remains as it is. When the forced QT is constrained by MinQTSize (the current block size is less than or equal to MinQTSize), a horizontal forced BT is used to match the lower boundary. If the lower boundary matches, a vertical forced BT is used to match the right boundary.

[0285] For example, in FIG. 13A showing an embodiment of forced QTBT for a block located at the corner of a picture, if MinQTSize is 32 or restricted for the forced QT partition of the corner case, additional BT partitions will be used after the 32x32 block partition until the forced partition ends.

[0286] FIG. 13B shows further details of an exemplary partitioning of a boundary CTU at or within the corner of a picture according to an embodiment of the present invention. The CTU has a size of 128×128 samples. The CTU is first quad-tree partitioned into four square blocks, each having a size of 64×64 samples. Of these blocks, only the upper left block 910 is a boundary block, and the other three are located outside (completely outside) the picture and are not further processed. Block 910 is further partitioned using a quad-tree that divides it into four square blocks 920, 930, 940, and 950, each having a size of 32×32 samples. Block 920 is located inside the picture, and blocks 930, 940, and 950 form boundary blocks again. Since the sizes of these blocks 930, 940, and 950 are not greater than MinQTSize, which is 32, forced binary splitting is applied to blocks 930, 940, and 950.

[0287] Block 930, which is located at the right boundary, is partitioned using recursive vertical forced binary splitting until a leaf node enters the picture, for example, until block 932 is located at the right boundary of the picture (here, after two vertical binary splits).

[0288] Block 940 is located at the lower boundary and is partitioned using recursive horizontal forced binary splitting until a leaf node enters the picture, e.g., until block 942 (here, after two horizontal binary splits) is located at the right boundary of the picture.

[0289] Block 950 is located at the corner boundary and is partitioned using a first recursive horizontal forced binary split until a sub - partition or block, here block 952, is located at the lower boundary of the picture (here, after two horizontal binary splits), and then the sub - partitions are recursively partitioned by vertical forced boundary partitioning until a leaf node or block, e.g., block 954, is located at the right boundary of the picture (here, after two vertical binary splits), or until each leaf node is located within the picture.

[0290] The above approach may be applied to both decoding and encoding. In the case of decoding, MinQTSize may be received by the SPS. In the case of encoding, MinQTSize may be sent by the SPS. Embodiments may use boundary definitions as shown in FIG. 12 or FIG. 14, or other boundary definitions.

[0291] Further embodiments of the present disclosure are provided below. It should be noted that the numbers used in the following sections do not necessarily follow the numbers used in the previous sections. Embodiment 1: A partitioning method comprising: determining whether the current block of the picture is a boundary block; when the current block is a boundary block, determining whether the size of the current block is greater than the minimum allowable quadtree leaf node size; when the size of the current block is not greater than the minimum allowable quadtree leaf node size, applying a forced binary split to the current block.

[0292] Embodiment 2: In the partitioning method of Embodiment 1, the forced binary tree partitioning is a recursive horizontal forced binary split when the current block is located at the lower boundary of the picture, or a recursive vertical forced boundary partitioning when the current block is located at the right boundary of the picture.

[0293] Embodiment 3: In the partitioning method of Embodiment 1 or 2, the forced binary split recursively divides the current block by horizontal forced boundary partitioning until the sub - partition of the current block is directly located at the lower boundary of the picture, and recursively divides the sub - partition by vertical forced boundary partitioning until the leaf node is completely directly located at the right boundary of the picture, or vice versa.

[0294] Embodiment 4: In the partitioning method of any one of Embodiments 1 - 3, the minimum allowable quadtree leaf node size is the minimum allowable quadtree leaf node size also applied to control the partitioning of non - boundary blocks.

[0295] Embodiment 5: A decoding method for decoding a block by partitioning the block according to the partitioning method of any one of Embodiments 1 - 4.

[0296] Embodiment 6: In the decoding method of Embodiment 5, the minimum allowable quadtree leaf node size is received by the SPS.

[0297] Embodiment 7: An encoding method for encoding a block by partitioning the block according to the partitioning method of any one of Embodiments 1 - 4.

[0298] Embodiment 8: In the encoding method of Embodiment 7, the minimum allowable quadtree leaf node size is transmitted by the SPS.

[0299] Embodiment 9: A decoding device including a logic circuit configured to execute any one of the methods of Embodiment 5 or 6.

[0300] Embodiment 10: An encoding device including a logic circuit configured to execute any one of the methods of Embodiment 7 or 8.

[0301] Embodiment 11: A non-transitory storage medium storing instructions that, when executed by a processor, cause the processor to execute any one of the methods according to Embodiments 1-8.

[0302] The apparatus includes a memory element and a processor element coupled to the memory element and configured to determine whether a current block of a picture is a boundary block, and if the current block is a boundary block, determine whether the size of the current block is greater than a minimum allowable quadtree (QT) leaf node size (MinQTSize), and if the size of the current block is not greater than MinQTSize, apply a forced binary tree (BT) partitioning to the current block.

[0303] In summary, embodiments of the present application (or the present disclosure) provide apparatuses and methods for encoding and decoding.

[0304] A first aspect relates to a partitioning method including steps of determining whether a current block of a picture is a boundary block and whether the size of the current block is greater than a minimum allowable quadtree leaf node size, and applying a forced binary tree (BT) partitioning to the current block if the current block is a boundary block and the size of the current block is not greater than a minimum allowable quadtree leaf node size (MinQTSize).

[0305] In a first implementation form of the method according to such a first aspect, the forced binary tree partitioning is a recursive horizontal forced bisection when the current block is located at the lower boundary of the picture, or a recursive vertical forced boundary partitioning when the current block is located at the right boundary of the picture.

[0306] In a second implementation form of the method according to such a first aspect or any preceding implementation form of the first aspect, the forced binary tree partitioning continues until the leaf node block is within the picture.

[0307] In a third implementation form of the method according to such a first aspect or any preceding implementation form of the first aspect, the forced bisection includes the steps of recursively dividing the current block by horizontal forced boundary partitioning until the sub - partition of the current block is located at the lower boundary of the picture, and recursively dividing the sub - partition by vertical forced boundary partitioning until the leaf node is completely located at the right boundary of the picture.

[0308] In a fourth implementation form of the method according to such a first aspect or any preceding implementation form of the first aspect, the forced BT partitioning includes the steps of recursively dividing the current block by vertical forced boundary partitioning until the sub - partition of the current block is located at the lower boundary, and recursively dividing the sub - partition by horizontal forced boundary partitioning until the leaf node is completely located at the right boundary.

[0309] In a fifth implementation form of the method according to such a first aspect or any preceding implementation form of the first aspect, the method further includes the step of applying a minimum allowable quadtree leaf node size to control the partitioning of non - boundary blocks.

[0310] In the sixth implementation form of the method according to such a first aspect or any preceding implementation form of the first aspect, the boundary block is a block that is not completely within the picture and not completely outside the picture.

[0311] The second aspect relates to a decoding method for decoding blocks by partitioning the blocks according to such a first aspect or any preceding implementation form of the first aspect.

[0312] In the first implementation form of the method according to such a second aspect, the method further includes the step of receiving a minimum allowable quadtree leaf node size via a sequence parameter set (SPS).

[0313] The third aspect relates to an encoding method for encoding blocks by partitioning the blocks according to such a first aspect or any preceding implementation form of the first aspect.

[0314] In the first implementation form of the method according to such a third aspect, the method further includes the step of transmitting a minimum allowable quadtree leaf node size via a sequence parameter set (SPS).

[0315] The fourth aspect relates to a decoding device including a logic circuit configured to decode blocks by partitioning the blocks according to the partitioning method of such a first aspect or any preceding implementation form of the first aspect.

[0316] In the first implementation form of the decoding device according to such a fourth aspect, the logic circuit is further configured to receive a minimum allowable quadtree leaf node size via a sequence parameter set (SPS).

[0317] The fifth aspect relates to an encoding device including a logic circuit configured to encode blocks by partitioning the blocks according to the partitioning method of such a first aspect or any preceding implementation form of the first aspect.

[0318] In a first implementation form of the decoding device according to such a fifth aspect, the logic circuit is further configured to transmit the minimum allowable quad-tree leaf node size via a sequence parameter set (SPS).

[0319] A sixth aspect is related to a non-transitory storage medium for storing instructions that cause a processor to execute any of such first, second, third aspects or any preceding implementation forms of the first, second, third aspects when executed by the processor.

[0320] A seventh aspect relates to a method including a step of determining that a current block of a picture is a boundary block and that a size of the current block is less than or equal to a minimum allowable quad-tree (QT) leaf node size (MinQTSize), and a step of applying a forced binary tree (BT) partitioning to the current block in response to the determination.

[0321] In a first implementation form of the method according to such a seventh aspect, the current block is located at a bottom boundary of the picture, and the forced BT partitioning is a recursive horizontal forced BT partitioning.

[0322] In a second implementation form of the method according to such a seventh aspect or any preceding implementation form of the seventh aspect, the current block is located at a right boundary of the picture, and the forced BT partitioning is a recursive vertical forced BT partitioning.

[0323] In a third implementation form of the method according to such a seventh aspect or any preceding implementation form of the seventh aspect, the forced BT partitioning includes a step of recursively dividing the current block by horizontal forced boundary partitioning until a sub-partition of the current block is located at the bottom boundary, and a step of recursively dividing the sub-partition by vertical forced boundary partitioning until a leaf node is completely located at the right boundary.

[0324] In a fourth implementation form of the method according to such a seventh aspect or any preceding implementation form of the seventh aspect, the forced BT partitioning recursively divides the current block by vertical forced boundary partitioning until the sub-partition of the current block is located at the lower boundary, and recursively divides the sub-partition by horizontal forced boundary partitioning until the leaf node is completely located at the right boundary.

[0325] In a fifth implementation form of the method according to such a seventh aspect or any preceding implementation form of the seventh aspect, the method further includes applying MinQTSize to control the partitioning of non-boundary blocks.

[0326] In a sixth implementation form of the method according to such a seventh aspect or any preceding implementation form of the seventh aspect, the method further includes receiving MinQTSize via a sequence parameter set (SPS).

[0327] In a seventh implementation form of the method according to such a seventh aspect or any preceding implementation form of the seventh aspect, the method further includes transmitting MinQTSize via a sequence parameter set (SPS).

[0328] The eighth aspect relates to an apparatus including a memory and a processor coupled to the memory and configured to determine whether a current block of a picture is a boundary block, and if the current block is a boundary block, determine whether the size of the current block is larger than a minimum allowable quad-tree (QT) leaf node size (MinQTSize), and apply forced binary tree (BT) partitioning to the current block if the size of the current block is not larger than MinQTSize.

[0329] In a first implementation form of the apparatus according to such an eighth aspect, the forced BT partitioning is a recursive horizontal forced BT partitioning if the current block is located at the lower boundary of the picture, or a recursive vertical forced BT partitioning if the current block is located at the right boundary of the picture.

[0330] In a second implementation form of the apparatus according to such an eighth aspect or any preceding implementation form of the eighth aspect, the forced BT partitioning includes the steps of recursively dividing the current block by horizontal forced boundary partitioning until the sub-partition of the current block is located at the lower boundary, and recursively dividing the sub-partition by vertical forced boundary partitioning until the leaf node is completely located at the right boundary.

[0331] In a third implementation form of the apparatus according to such an eighth aspect or any preceding implementation form of the eighth aspect, the forced BT partitioning includes the steps of recursively dividing the current block by vertical forced boundary partitioning until the sub-partition of the current block is located at the lower boundary, and recursively dividing the sub-partition by horizontal forced boundary partitioning until the leaf node is completely located at the right boundary.

[0332] In a fourth implementation form of the apparatus according to such an eighth aspect or any preceding implementation form of the eighth aspect, the processor is further configured to apply MinQTSize to control the partitioning of non-boundary blocks.

[0333] In a fifth implementation form of the apparatus according to such an eighth aspect or any preceding implementation form of the eighth aspect, the apparatus further includes a receiver coupled to the processor and configured to receive MinQTSize via a sequence parameter set (SPS).

[0334] In a sixth implementation form of the apparatus according to such an eighth aspect or any preceding implementation form of the eighth aspect, the apparatus further includes a transmitter coupled to a processor and configured to transmit MinQTSize via a sequence parameter set (SPS).

[0335] A ninth aspect relates to a computer program product including computer-executable instructions stored on a non-transitory medium, which, when executed by a processor, cause the apparatus to determine whether a current block of a picture is a boundary block, and if the current block is a boundary block, determine whether the size of the current block is greater than a minimum allowable quadtree (QT) leaf node size (MinQTSize), and if the size of the current block is not greater than MinQTSize, cause the apparatus to apply a forced binary tree (BT) partitioning to the current block.

[0336] In a first implementation form of the apparatus according to such an eighth aspect, the forced BT partitioning is a recursive horizontal forced BT partitioning if the current block is located at the lower boundary of the picture, or a recursive vertical forced BT partitioning if the current block is located at the right boundary of the picture.

[0337] In a second implementation form of the apparatus according to such a ninth aspect or any preceding implementation form of the ninth aspect, the forced BT partitioning includes recursively dividing the current block by horizontal forced boundary partitioning until a sub-partition of the current block is located at the lower boundary, and recursively dividing the sub-partition by vertical forced boundary partitioning until a leaf node is completely located at the right boundary.

[0338] In a third implementation form of the apparatus according to such a ninth aspect or any preceding implementation form of the ninth aspect, the forced BT partitioning includes recursively partitioning the current block by vertical forced boundary partitioning until the sub-partition of the current block is located at the lower boundary, and recursively partitioning the sub-partition by horizontal forced boundary partitioning until the leaf node is completely located at the right boundary.

[0339] In a fourth implementation form of the apparatus according to such a ninth aspect or any preceding implementation form of the ninth aspect, the instruction further causes the apparatus to apply MinQTSize to control the partitioning of non-boundary blocks.

[0340] In a fifth implementation form of the apparatus according to such a ninth aspect or any preceding implementation form of the ninth aspect, the instruction further causes the apparatus to receive MinQTSize via a sequence parameter set (SPS).

[0341] In a sixth implementation form of the apparatus according to such a ninth aspect or any preceding implementation form of the ninth aspect, the instruction further causes the apparatus to transmit MinQTSize via a sequence parameter set (SPS).

[0342] According to some embodiments, a decoder is provided that includes one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the decoder to perform any of the methods described above in connection with embodiments 1 through 4 when executed by the processors.

[0343] Furthermore, an encoder is provided that includes one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming configuring the encoder to perform any of the methods described above in connection with Embodiments 1 through 4 when executed by the processors.

[0344] The following is an explanation of an encoding method similar to the decoding method as shown in the above-described embodiments, and the application of a system using the same.

[0345] FIG. 17 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types.

[0346] The capture device 3102 generates data and may encode the data by an encoding method as shown in the above embodiments. Alternatively, the capture device 3102 can distribute the data to a streaming server (not shown), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a pad, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 can actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 can actually perform audio encoding processing. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, such as in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.

[0347] In the content supply system 3100, the terminal device 310 receives and plays back the encoded data. The terminal device 3106 can be a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, or a device having data reception and restoration capabilities such as those capable of decrypting the above-described encoded data. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.

[0348] For a terminal device having a display, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its display. For example, for a terminal device without a display, such as the STB 3116, the video conferencing system 3118, or the video surveillance system 3120, an external display 3126 is attached thereto to receive and display the decoded data.

[0349] When each device in this system performs encoding or decoding, a picture encoding device or a picture decoding device can be used as shown in the above embodiment.

[0350] FIG. 18 is a diagram showing a configuration of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination of these types.

[0351] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.

[0352] By the demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. The video decoder 3206 including the video decoder 30 as described in the above embodiment decodes the video ES by the decoding method as shown in the above embodiment to generate a video frame, and sends this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate an audio frame, and sends this data to the synchronization unit 3212. Alternatively, the video frame may be stored in a buffer (not shown in FIG. 18) before supplying it to the synchronization unit 3212. Similarly, the audio frame may be stored in a buffer (not shown in FIG. 18) before supplying it to the synchronization unit 3212.

[0353] Synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to video / audio display 3214. For example, synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded in syntax using a time stamp related to the presentation of coded audio and visual data and a time stamp related to the delivery of the data stream itself.

[0354] When subtitles are included in the stream, subtitle decoder 3210 decodes the subtitles, synchronizes them with video frames and audio frames, and supplies video / audio / subtitles to video / audio / subtitle display 3216.

[0355] The present invention is not limited to the above-described system, and either the picture encoding device or the picture decoding device in the above-described embodiment can be incorporated into other systems, such as vehicle systems.

[0356] Embodiments of the present invention have mainly been described based on video coding. However, embodiments of the coding system 10, the encoder 20, and the decoder 30 (and corresponding system 10), as well as other embodiments described in the present application, may also be configured for the processing or coding of still images, i.e., individual pictures independent of any preceding or subsequent pictures, as in video coding. It should be noted that generally, when picture processing coding is limited to a single picture 17, only the inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can be used similarly for still image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, inverse quantization 210 / 310, (inverse) transformation 212 / 312, partitioning 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, and entropy coding 270 and entropy decoding 304.

[0357] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described in the present application in relation to, for example, the encoder 20 and the decoder 30, can be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions can be stored in a computer-readable medium or transmitted via a communication medium as one or more instructions or codes and executed by a hardware-based processing unit. The computer-readable medium may correspond to a computer-readable storage medium such as a data storage medium, or may include a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, in accordance with a communication protocol. Thus, the computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.

[0358] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for use in implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0359] For example, but not limited to, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. Also, any connection can be appropriately referred to as a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, wireless, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, wireless, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transient tangible storage media. Disks and discs, as used in this application, include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy discs, and Blu-ray discs, where discs typically read data magnetically and discs read data optically with a laser. The above combinations should also be included within the scope of computer-readable media.

[0360] The commands can be executed by one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used in this application may refer to any of the foregoing structures, or any other structure suitable for implementing the technology described in this application. Further, in some aspects, the functions described in this application may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Also, the technology can be fully implemented with one or more circuits or logic elements.

[0361] The technology of this disclosure can be implemented in a wide variety of devices or apparatuses including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., a chipset). To emphasize the functional aspects of devices configured to execute the disclosed technology, various components, modules, or units are described in this disclosure, but do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined within a codec hardware unit, or provided by a set of interoperable hardware units including one or more processors as described above, along with appropriate software and / or firmware.

[0362] The following logical or mathematical operators are defined as follows. The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations such as exponentiation and real-valued division are specified. The numbering and counting rules generally start from 0. For example, "the first" is equivalent to the 0th, "the second" is equivalent to the 1st, and so on.

[0363] Arithmetic operators The following arithmetic operators are defined as follows:

Table 1

[0364] Logical operators The following logical operators are defined as follows: x&&y The Boolean logical “and” of x and y x||y The Boolean logical “or” of x and y ! Boolean logical “not” x?y:z If x is TRUE or not equal to 0, evaluate the value of y; otherwise, evaluate the value of z.

[0365] Relational operators The following relational operators are defined as follows: > Greater than >= Greater than or equal to < Less than <= Less than or equal to == Equal to != Not equal to When a relational operator is applied to a syntax element or variable specified with the value “na” (not applicable), the value “na” is treated as a distinct value of that syntax element or variable. The value “na” is considered not equal to any other value.

[0366] Bitwise operators The following bitwise operators are defined as follows: & Bitwise “and”. When acting on integer arguments, it acts on the two's complement representation of integer values. When acting on binary arguments containing fewer bits than the other argument, the shorter argument is extended by adding additional high-order bits equal to 0. | Bitwise “or”. When acting on integer arguments, it acts on the two's complement representation of the integer values. When acting on binary arguments that contain fewer bits than the other arguments, the shorter argument is extended by adding additional high-order bits equal to 0. ^ Bitwise “exclusive or”. When acting on integer arguments, it acts on the two's complement representation of the integer values. When acting on binary arguments that contain fewer bits than the other arguments, the shorter argument is extended by adding additional high-order bits equal to 0. x>>y The two's complement integer representation of x arithmetically right-shifted by the number of bits in the binary number y. This function is defined only for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift is equal to the MSB of x before the shift operation. x<<y The two's complement integer representation of x arithmetically left-shifted by the number of bits in the binary number y. This function is defined only for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.

[0367] Assignment operators The following assignment operators are defined as follows: = Assignment operator ++ Increment. That is, x++ is equal to x=x+1. When used as an array index, the value of the variable before the increment operation is evaluated. -- Decrement. That is, x-- is equal to x=x-1. When used as an array index, the value of the variable before the decrement operation is evaluated. += Increment by the specified amount. That is, x+=3 is equal to x=x+3. x+=(-3) is equal to x=x+(-3). -= Decrement by the specified amount. That is, x-=3 is equal to x=x-3. x-=(-3) is equal to x=x-(-3).

[0368] Range notation The following notations are used to specify a range of values: x = y..z means x takes integer values starting from y and including both ends up to z, where x, y, and z are integers and z is greater than y.

[0369] Mathematical functions The following mathematical functions are defined:

Table 2

[0370] Operator precedence If the precedence is not explicitly specified in an expression by using parentheses, the following rules apply: - Operations with higher precedence are evaluated before any operations with lower precedence. - Operations with the same precedence are evaluated in order from left to right.

[0371] The following table shows the operator precedence from highest to lowest, where a higher position in the table indicates a higher precedence.

[0372] For the operators also used in the C programming language, the precedence used in this specification is the same as that used in the C programming language Table: Operator precedence from highest (top of the table) to lowest (bottom of the table)

Table 3

[0373] Text description of logical operators In text, in the following form: if( condition 0 ) statement 0 else if( condition 1 ) statement 1 ... else / * informative remark on remaining condition * / statement n A statement of logical operators, as described mathematically, can be described in the following ways: …as follows / ...the following applies: - If condition 0, statement 0 - Otherwise, if condition 1, statement 1 -... - Otherwise (informative remark on remaining condition), statement n

[0374] Each “If... Otherwise, if... Otherwise,...” statement in the text is introduced with “... as follows” or “... the following applies” immediately following “If...”. The last condition of “If... Otherwise, if... Otherwise,...” is always “Otherwise,...”. Alternating “If... Otherwise, if... Otherwise,...” statements can be verified by matching “... as follows” or “... the following applies” with the trailing “Otherwise,...”.

[0375] In the text, the following form: if( condition 0a && condition 0b ) statement 0 else if( condition 1a || condition 1b ) statement 1 ... else statement n A statement of a logical operator as mathematically described can be described in the following way: ... as follows / ... the following applies: - If all of the following conditions are true, statement 0: -condition 0a -condition 0b - Otherwise, if one or more of the following conditions are true, statement 1: -condition 1a -condition 1b - ... - Otherwise, statement n

[0376] In text, the following form: if( condition 0 ) statement 0 if( condition 1 ) statement 1 A statement of a logical operator as mathematically described can be described in the following way: If condition 0, then statement 0 If condition 1, then statement 1

[0377] In summary, the present disclosure relates to methods and devices used for encoding and decoding image or video signals. These include determining whether the size of the current block is larger than the minimum allowable quadtree leaf node size. If the size of the current block is not larger than the minimum allowable quadtree leaf node size, a multi-type tree split is applied to the current block. The minimum allowable quadtree leaf node size is not larger than the maximum allowable binary tree root node size, or the minimum allowable quadtree leaf node size is not larger than the maximum allowable ternary tree root node size.

Claims

Claim 1 A method of symbolization, comprising: determining whether a size of a current block is greater than a minimum allowable quadtree leaf node size; applying a binary tree split to the current block based on a maximum allowable binary tree root node size, under a condition that the size of the current block is not greater than the minimum allowable quadtree leaf node size, wherein the maximum allowable binary tree root node size is determined based on the minimum allowable quadtree leaf node size, and the minimum allowable quadtree leaf node size is not greater than the maximum allowable binary tree root node size; generating a predicted block of the current block after applying the binary tree split to the current block; obtaining a residual block of the current block based on the current block and the predicted block; obtaining transform coefficients of the current block by applying a transform to sample values of the residual block; encoding the transform coefficients of the current block into a bit stream; A method comprising the above steps. Claim 2 The method according to claim 1, wherein the current block is a boundary block. Claim 3 The method according to claim 1, further comprising: splitting an image into blocks, wherein the blocks include the current block, wherein the step of applying the binary tree split to the current block is applying the binary tree split to the current block of the blocks having a final maximum multi-type tree depth, wherein the final maximum multi-type tree depth is at least a sum of a maximum multi-type tree depth and a maximum multi-type tree depth offset; A method comprising the above steps. Claim 4 The method according to claim 3, wherein the maximum multi-type tree depth offset is 0. Claim 5 The method according to claim 3, wherein the maximum multi-type tree depth is greater than 0. Claim 6 The method according to claim 1, further comprising: Encoding the first syntax element and the second syntax element into the bitstream, wherein the first syntax element is used to derive the minimum allowable quadtree leaf node size, and the second syntax element is used to derive the maximum allowable binary tree root node size; A method comprising. **Claim 7** The method according to claim 6, wherein the first syntax element and the second syntax element are signaled in a sequence parameter set (SPS). **Claim 8** The method according to claim 1, wherein the method further comprises: Transmitting the bitstream; A method comprising. **Claim 9** A method of decoding, comprising: Receiving a bitstream including encoded data of an image; Determining whether the size of the current block of the image is larger than the minimum allowable quadtree leaf node size; Applying binary tree splitting to the current block based on the maximum allowable binary tree root node size, under the condition that the size of the current block is not larger than the minimum allowable quadtree leaf node size, wherein the maximum allowable binary tree root node size is determined based on the minimum allowable quadtree leaf node size, and the minimum allowable quadtree leaf node size is not larger than the maximum allowable binary tree root node size; Obtaining the transform coefficients of the current block from the bitstream; Obtaining a residual block of the current block based on the transform coefficients of the current block; Reconstructing the current block based on the residual block of the current block and the predicted block of the current block; A method comprising. **Claim 10** The method according to claim 9, wherein the current block is a boundary block. **Claim 11** The method according to claim 9, wherein the bitstream further includes a first syntax element, and the method further comprises: Analyzing the first syntax element from the bitstream; Deriving the minimum allowable quadtree leaf node size based on the first syntax element; A method comprising. **Claim 12** In the method according to claim 9, the bitstream includes a second syntax element, and the method further comprises: analyzing the second syntax element from the bitstream; deriving the maximum allowable binary tree root node size based on the second syntax element; A method comprising the steps of:

13. A coding device comprising a processing circuit for performing the method according to any one of claims 1 to 12.

14. In a coding device: one or more processors; and a computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, the programming configuring the coding device to perform the method according to any one of claims 1 - 12 when executed by the one or more processors; A coding device comprising:

15. A computer program comprising program code for performing the method according to any one of claims 1 - 12 when executed on a computer or processor.

16. A computer-readable storage medium storing program code for causing a computer device to perform the method according to any one of claims 1 - 12 when executed by the computer device.

17. A device for storing a video or image bitstream, comprising a communication interface and a storage medium, the communication interface being configured to receive and / or transmit a bitstream, the storage medium being configured to store the bitstream, the bitstream including encoded data of one or more images, a first syntax element for deriving a minimum allowable quadtree leaf node size, and a second syntax element for deriving a maximum allowable binary tree root node size based on the minimum allowable quadtree leaf node size; the minimum allowable quadtree leaf node size is not greater than the maximum allowable binary tree root node size, and the bitstream further includes parameters of a prediction process and a conversion process.