Multi-type tree framework for transforms in video coding

By adopting multi-type tree (MTT) segmentation technology in video encoding, the problem of inflexible video block segmentation and organization in the prior art is solved, and more efficient video encoding is achieved.

CN120017838APending Publication Date: 2025-05-16QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510016633.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-01
Filing Date
2019-04-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing video encoding technology lacks flexibility when segmenting and organizing video data blocks, resulting in low encoding efficiency.

Method used

Multi-type tree (MTT) segmentation technology is adopted to flexibly segment video data blocks by using three or more different segmentation structures at each depth of the encoded tree structure.

Benefits of technology

It improves the flexibility and encoding efficiency of video encoding, and enables more efficient segmentation and organization of video data blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017838A_ABST
    Figure CN120017838A_ABST
Patent Text Reader

Abstract

Embodiments include methods and apparatus for decoding video data, including receiving an encoded video bitstream forming a representation of an encoded picture of the video data, and determining to partition the encoded picture of the video data into a plurality of coding units. The partitioning may be in accordance with a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure. The method further includes determining to recursively divide the residual block of the leaf nodes into a plurality of transform units according to a second tree structure.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application with application date of April 2, 2019 and application number 201980022133.X.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This patent application claims priority to Provisional Application No. 62 / 651,689 filed on April 2, 2018 and Non-Provisional Application No. 16 / 372,249 filed on April 1, 2019, which are assigned to the assignee of the present application and are expressly incorporated herein by reference. Technical Field

[0004] The present disclosure relates to video encoding and video decoding. Background Art

[0005] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless telephones, so-called "smart phones", video teleconferencing devices, video streaming devices, etc. Digital video devices implement video coding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) standards, and extensions of these standards. By implementing such video coding techniques, video devices may more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0006] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture / frame or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding units (CUs), and / or coding nodes. Pictures may be referred to as frames. Reference pictures may be referred to as reference frames.

[0007] Spatial or temporal prediction results in a predictive block for the block to be coded. The residual data represents the pixel difference between the original block to be coded and the predictive block. For further compression, the residual data can be transformed from the pixel domain to the transform domain to obtain residual transform coefficients, which can then be quantized. Entropy coding can be applied to achieve even more compression. Summary of the invention

[0008] This disclosure describes techniques for partitioning blocks of video data using a multi-tree based framework, including recursively partitioning transform blocks or units into a tree structure that may be different from a tree structure of a corresponding coding unit.

[0009] One embodiment includes a method for decoding video data. The method includes receiving an encoded video bitstream that forms a representation of an encoded picture of the video data. The method also includes determining to segment the encoded picture of the video data into a plurality of coding units. The segmentation is performed according to a first tree structure. The plurality of coding units include leaf nodes of the first tree structure. The method also includes determining to recursively divide the residual blocks of the leaf nodes into a plurality of transform units according to a second tree structure. The method also includes reconstructing the encoded picture based on the determined first tree structure and the second tree structure.

[0010] Another embodiment includes a method for encoding video data. The method includes obtaining a picture for encoding into an encoded video bitstream, and determining to segment the picture of the video data into a plurality of coding units. The segmentation is according to a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure. The method also includes determining to recursively divide the residual blocks of the leaf nodes into a plurality of transform units according to a second tree structure. The method also includes encoding the picture into an encoded video bitstream based on the determined first tree structure and the second tree structure.

[0011] Another embodiment includes an apparatus for decoding video data. The apparatus includes a memory configured to store a coded picture of the video data. The apparatus also includes a video processor configured to receive a coded video bitstream forming a representation of a coded picture of the video data and determine to segment the coded picture of the video data into a plurality of coding units. The segmentation is performed according to a first tree structure. The plurality of coding units include leaf nodes of the first tree structure. The processor is further configured to determine to recursively divide the residual blocks of the leaf nodes into a plurality of transform units according to a second tree structure, and to reconstruct the coded picture based on the determined first tree structure and the second tree structure.

[0012] Another embodiment includes an apparatus for encoding video data. The apparatus includes a memory configured to store an encoded picture of the video data. The apparatus also includes a processor configured to obtain a picture for encoding into an encoded video bitstream and determine to segment the picture of the video data into a plurality of coding units. The segmentation is performed according to a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure. The processor is further configured to further determine to recursively divide the residual blocks of the leaf nodes into a plurality of transform units according to a second tree structure. The processor is further configured to encode the picture into an encoded video bitstream based on the determined first tree structure and the second tree structure.

[0013] Another embodiment includes an apparatus for decoding video data. The apparatus includes a component for storing a coded picture of the video data. The apparatus also includes a component for obtaining a coded video bitstream forming a representation of the coded picture of the video data, and a component for determining to segment the coded picture of the video data into a plurality of coding units. The segmentation is performed according to a first tree structure. The plurality of coding units include leaf nodes of the first tree structure. The apparatus also includes a component for determining to recursively divide the residual blocks of the leaf nodes into a plurality of transform units according to a second tree structure, and a component for reconstructing the coded picture based on the determined first tree structure and the second tree structure.

[0014] Another embodiment includes an apparatus for encoding video data. The apparatus includes a component for storing an encoded picture of the video data. The apparatus also includes a component for obtaining a picture for encoding into an encoded video bitstream, and a component for determining to segment the picture of the video data into a plurality of coding units. The segmentation is performed according to a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure. The apparatus also includes a component for determining to recursively divide the residual blocks of the leaf nodes into a plurality of transform units according to a second tree structure. The apparatus also includes a component for encoding the picture into an encoded video bitstream based on the determined first tree structure and the second tree structure.

[0015] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a block diagram illustrating an example video encoding and decoding system configured to implement the techniques of this disclosure.

[0017] Figure 2 is a conceptual diagram showing a coding unit (CU) structure in high efficiency video coding (HEVC).

[0018] Figure 3 is a conceptual diagram illustrating example partition types for inter prediction mode.

[0019] Figure 4A is a conceptual diagram illustrating an example tree structure of block partitioning using a quadtree binary tree (QTBT) structure.

[0020] Figure 4B is shown and used Figure 4A Conceptual diagram of an example tree structure corresponding to the block partitioning of the QTBT structure.

[0021] Figure 5A is a conceptual diagram illustrating an example horizontal ternary tree partition type.

[0022] Figure 5Bis a conceptual diagram illustrating an example horizontal ternary tree partition type.

[0023] Figure 5C is a conceptual diagram illustrating example asymmetric binary tree partition types.

[0024] Fig. 6A is a conceptual diagram showing quadtree partitioning.

[0025] Figure 6B is a conceptual diagram illustrating vertical binary tree partitioning.

[0026] Figure 6C is a conceptual diagram showing horizontal binary tree partitioning.

[0027] Fig.6D is a conceptual diagram illustrating vertical center-side tree partitioning.

[0028] Fig. 6E is a conceptual diagram illustrating horizontal center-side tree partitioning.

[0029] Figure 7 is a conceptual diagram illustrating an example of coding tree unit (CTU) partitioning according to the technology of this disclosure.

[0030] Figure 8 is a block diagram illustrating an example of a video encoder.

[0031] Fig. 9 is a block diagram illustrating an example of a video decoder.

[0032] Fig. 10A is a flow diagram illustrating example operation of a video encoder in accordance with techniques of this disclosure.

[0033] Fig. 10B is a flow diagram illustrating example operation of a video decoder in accordance with techniques of this disclosure.

[0034] Fig.11 is a flow chart illustrating example operation of a video encoder according to another example technique of this disclosure.

[0035] Fig.12 is a flow chart illustrating example operation of a video decoder according to another example technique of this disclosure. DETAILED DESCRIPTION

[0036] The present disclosure relates to the segmentation and / or organization of video data blocks (e.g., coding units) in block-based video coding. The techniques of the present disclosure can be applied in video coding standards. In various examples described below, the techniques of the present disclosure include segmenting video data blocks using three or more different segmentation structures. In some examples, three or more different segmentation structures can be used at each depth of the coding tree structure. Such segmentation techniques can be referred to as multi-type-tree (MTT) segmentation. By using MTT segmentation, video data can be segmented more flexibly, thereby allowing higher coding efficiency.

[0037] Figure 1 is a block diagram illustrating an example video encoding and decoding system 10 that can utilize the techniques of this disclosure to segment blocks of video data, signal and parse segmentation types, and apply transforms and further transform segmentations. Figure 1 As shown, system 10 includes a source device 12 that provides encoded video data to be decoded by a destination device 14 at a later time. In particular, source device 12 provides video data to destination device 14 via computer-readable medium 16. Source device 12 and destination device 14 may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets (such as so-called "smart" phones), tablet computers, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, or similar devices. In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Therefore, source device 12 and destination device 14 may be wireless communication devices. Source device 12 is an example video encoding device (i.e., a device for encoding video data). Destination device 14 is an example video decoding device (e.g., a device or apparatus for decoding video data).

[0038] exist Figure 1 In the example of the embodiment of the present invention, the source device 12 includes a video source 18, a storage medium 20 configured to store video data, a video encoder 22, and an output interface 24. The destination device 14 includes an input interface 26, a storage medium 28 configured to store encoded video data, a video decoder 30, and a display device 32. In other examples, the source device 12 and the destination device 14 include other components or arrangements. For example, the source device 12 can receive video data from an external video source (such as an external camera). Similarly, the destination device 14 can be connected to an external display device interface instead of including an integrated display device.

[0039] Figure 1The system 10 shown is only one example. The techniques for processing video data can be performed by any digital video encoding and / or decoding device or apparatus. Although the techniques of the present disclosure are generally performed by a video encoding device and a video decoding device, these techniques can also be performed by a combined video codec / decoder, commonly referred to as a "codec (CODEC)". The source device 12 and the destination device 14 are merely examples of such encoding devices, wherein the source device 12 generates encoded video data for transmission to the destination device 14. In some examples, the source device 12 and the destination device 14 operate in a substantially symmetrical manner, so that each of the source device 12 and the destination device 14 includes a video encoding and decoding component. Therefore, the system 10 can support one-way or two-way video transmission between the source device 12 and the destination device 14, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0040] The video source 18 of the source device 12 may include a video capture device, such as a video camera, a video archive containing previously captured videos, and / or a video feed interface for receiving video data from a video content provider. As a further option, the video source 18 may generate computer graphics-based data as the source video, or a combination of real-time video, archived video, and computer-generated video. The source device 12 may include one or more data storage media (e.g., storage media 20) configured to store video data. In general, the techniques described in the present disclosure may be generally applicable to video encoding, and may be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by a video encoder 22. The output interface 24 may output the encoded video information to a computer-readable medium 16.

[0041] The destination device 14 may receive the encoded video data to be decoded via the computer-readable medium 16. The computer-readable medium 16 may include any type of medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In some examples, the computer-readable medium 16 includes a communication medium to enable the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard (such as a wireless communication protocol) and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other equipment that may help facilitate communication from the source device 12 to the destination device 14. The destination device 14 may include one or more data storage media configured to store the encoded video data and the decoded video data.

[0042] In some examples, the encoded data (e.g., encoded video data) may be output from the output interface 24 to a storage device. Similarly, the encoded data may be accessed from the storage device via the input interface 26. The storage device may include any of a variety of distributed or locally accessed data storage media (such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data). In a further example, the storage device may correspond to a file server or another intermediate storage device that may store the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and transmitting the encoded video data to the destination device 14. Example file servers include a network server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data via any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, which is suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.

[0043] The techniques of the present disclosure may be applied to video encoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as dynamic adaptive streaming (DASH) over HTTP), digital video encoded to a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0044] Computer-readable medium 16 may include transient media, such as wireless broadcast or wired network transmission, or storage media (that is, non-transitory storage media), such as hard disks, flash drives, optical disks, digital video disks, Blu-ray disks, or other computer-readable media. In some examples, a network server (not shown) can receive encoded video data from source device 12 and provide the encoded video data to destination device 14, for example, via network transmission. Similarly, a computing device of a media production facility (such as a disc stamping facility) can receive encoded video data from source device 12 and produce a disc containing the encoded video data. Therefore, in various examples, computer-readable medium 16 can be understood to include one or more computer-readable media in various forms.

[0045] The input interface 26 of the destination device 14 receives information from the computer-readable medium 16. The information of the computer-readable medium 16 may include syntax information defined by the video encoder 22, which is also used by the video decoder 30, and includes syntax elements that describe the characteristics and / or processing of blocks and other coding units (e.g., groups of pictures (GOPs)). The storage medium 28 can store the encoded video data received by the input interface 26. The display device 32 displays the decoded video data to the user. The display device 32 may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0046] The video encoder 22 and the video decoder 30 can each be implemented as any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store the instructions of the software in a suitable, non-transitory computer-readable medium, and can use one or more processors to run the instructions in hardware to perform the technology of the present disclosure. Each of the video encoder 22 and the video decoder 30 can be included in one or more encoders or decoders, any of which can be integrated in the corresponding device as part of a combined encoder / decoder (CODEC). The video encoder 22 and the video decoder 30 may include an ALU, an EFU, a digital circuit, an analog circuit, and / or a programmable core formed by a programmable circuit. In an example where the operation of the video encoder 22 and / or the video decoder 30 is performed by software running on a programmable circuit, an on-chip or off-chip memory can store instructions (e.g., object code) of the software received and run by the video encoder 22 or the video decoder 30.

[0047] In some examples, the video encoder 22 and the video decoder 30 may operate in accordance with a video coding standard such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC), or an extension thereof such as a multi-view and / or scalable video coding extension. Alternatively, the video encoder 22 and the video decoder 30 may operate in accordance with other proprietary standards or industry standards such as the Joint Exploration Test Model (JEM) or ITU-T H.266, also known as Versatile Video Coding (VVC). A recent draft of the VVC standard is described in Bross et al., “Versatile Video Coding (Draft 3),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC1 / SC 29 / WG 11, 12th Meeting: Macau, China, October 3-12, 2018, JVET-L1001-v9 (hereinafter referred to as “VVC Draft 3”). However, the techniques of this disclosure are not limited to any particular encoding standard.Rather, for purposes of illustration, examples are described with respect to the above examples.

[0048] In HEVC and other video coding specifications, a video sequence usually consists of a series of pictures. Pictures can also be called "frames". Pictures can include three sample arrays, denoted by S L , S Cb and S Cr . S L is a two-dimensional array (i.e., block) of luma samples. Cb It is a two-dimensional array of Cb chroma samples. Cr is a two-dimensional array of Cr chrominance samples. Chrominance samples may also be referred to herein as "chroma" samples. In other cases, a picture may be monochrome and may include only an array of luma samples.

[0049] In addition, in HEVC and other video coding specifications, in order to generate a coded representation of a picture, the video encoder 22 may generate a set of coding tree units (CTUs). Each of the CTUs may include a coding tree block of luminance samples, two corresponding coding tree blocks of chrominance samples, and a syntax structure for encoding the samples of the coding tree block. In a monochrome picture or a picture with three separate color planes, the CTU may include a single coding tree block and a syntax structure for encoding the samples of the coding tree block. The coding tree block may be an NxN sample block. The CTU may also be referred to as a "tree block" or "largest coding unit" (LCU). The CTU of HEVC may be roughly similar to macroblocks of other standards, such as H.264 / AVC. However, the CTU is not necessarily limited to a specific size and may include one or more coding units (CUs). A slice may include an integer number of CTUs arranged consecutively in a raster scan order.

[0050] If operating according to HEVC, in order to generate a coded CTU, the video encoder 22 may recursively quadtree partition the coding tree block of the CTU to divide the coding tree block into coding blocks, hence the name "coding tree unit". A coding block is an NxN sample block. A CU may include a luma sample coding block and two corresponding chroma sample coding blocks of a picture having a luma sample array, a Cb sample array, and a Cr sample array, and a syntax structure for encoding the samples of the coding block. In a monochrome picture or a picture with three independent color layers, a CU may include a single coding block and a syntax structure for encoding the samples of the coding block.

[0051] The syntax data within the bitstream may also define the size of the CTU. A slice includes several consecutive CTUs in coding order. A video frame or picture may be partitioned into one or more slices. As described above, each tree block may be divided into coding units (CUs) according to a quadtree. In general, a quadtree data structure includes one node per CU, where the root node corresponds to a tree block. If the CU is divided into four sub-CUs, the node corresponding to the CU includes four leaf nodes, each of which corresponds to one of the sub-CUs.

[0052] Each node of the quadtree data structure may provide syntax data for the corresponding CU. For example, a node in the quadtree may include a split flag indicating whether the CU corresponding to the node is split into sub-CUs. The syntax elements of a CU may be defined recursively and may depend on whether the CU is split into sub-CUs. If the CU is not further split, it is referred to as a leaf CU. If the CU block is further split, it may generally be referred to as a non-leaf CU. In some examples of the present disclosure, the four sub-CUs of a leaf CU may be referred to as leaf CUs even if the original leaf CU is not explicitly split. For example, if a CU of size 16x16 is not further split, the four 8x8 sub-CUs may be referred to as leaf CUs even though the 16x16 CU has never been split.

[0053] CU has a similar purpose to the macroblock of the H.264 standard, except that CU has no size difference. For example, a treeblock can be divided into four child nodes (also called sub-CUs), each of which can be a parent node and divided into another four child nodes. The last undivided child nodes, called leaf nodes of the quadtree, include coding nodes, also called leaf CUs. The syntax data associated with the coded bitstream can define the maximum number of times a treeblock can be divided, called the maximum CU depth, and can also define the minimum size of a coding node. Therefore, the bitstream can also define a smallest coding unit (SCU). The present disclosure uses the term "block" to refer to any of a CU, PU, ​​or TU in the context of HEVC, or similar data structures in other standard contexts (e.g., macroblocks and their subblocks in H.264 / AVC).

[0054] A CU includes a coding node and a prediction unit (PU) and a transform unit (TU) associated with the coding node. The size of the CU corresponds to the size of the coding node and may be square in some examples. In the example of HEVC, the size of the CU may be from 8x8 pixels to a maximum tree block size of 64x64 pixels or larger. Each CU may contain one or more PUs and one or more TUs. The syntax data associated with the CU may describe, for example, the partitioning of the CU into one or more PUs. The partitioning mode may be different whether the CU is skip mode or direct mode encoded, intra-frame prediction mode encoded, or inter-frame prediction mode encoded. The PU may be partitioned into a non-square shape. The syntax data associated with the CU may also describe, for example, the partitioning of the CU into one or more TUs according to a quadtree. The shape of the TU may be square or non-square (e.g., rectangular).

[0055] The HEVC standard allows transformation based on TUs. TUs can be different for different CUs. The size of a TU is typically based on the size of a PU within a given CU defined for a partitioned LCU, although this is not always the case. A TU is typically the same size as or smaller than a PU. In some examples, the residual samples corresponding to a CU may be subdivided into smaller units using a quadtree structure, sometimes referred to as a "residual quad tree" (RQT). The leaf nodes of an RQT may be referred to as TUs. The pixel difference values ​​associated with a TU may be transformed to produce transform coefficients, which may be quantized.

[0056] A leaf-CU may include one or more PUs. In general, a PU represents a spatial region corresponding to all or part of a corresponding CU, and may include data for retrieving reference samples of the PU. In addition, the PU includes data related to prediction. For example, when the PU is intra-mode encoded, the data of the PU may be included in the RQT, which may include data describing the intra-prediction mode of the TU corresponding to the PU. As another example, when the PU is inter-mode encoded, the PU may include data defining one or more motion vectors of the PU. The data defining the motion vector of the PU may describe, for example, the horizontal component of the motion vector, the vertical component of the motion vector, the resolution of the motion vector (e.g., quarter-pixel precision or eighth-pixel precision), the reference picture to which the motion vector points, and / or a reference picture list of the motion vector (e.g., Table 0, Table 1, or Table C).

[0057] A leaf-CU having one or more PUs may also include one or more TUs. As described above, a TU may be specified using an RQT (also referred to as a TU quadtree structure). For example, a partition flag may indicate whether a leaf-CU is partitioned into four transform units. In some examples, each transform unit may be further partitioned into further sub-TUs. When a TU is not further partitioned, it may be referred to as a leaf-TU. In general, for intra-coding, all leaf-TUs belonging to a leaf-CU contain residual data generated from the same intra-prediction mode. That is, the same intra-prediction mode is generally applied to calculate the predicted values ​​to be transformed in all TUs of the leaf-CU. For intra-coding, the video encoder 22 may calculate the residual value of each leaf-TU using the intra-prediction mode as the difference between the portion of the CU corresponding to the TU and the original block. A TU is not necessarily limited to the size of a PU. Therefore, a TU may be larger or smaller than a PU. For intra-coding, a PU may be juxtaposed with a corresponding leaf-TU of the same CU. In some examples, the maximum size of a leaf-TU may correspond to the size of a corresponding leaf-CU.

[0058] In addition, the TU of the leaf CU may also be associated with a corresponding RQT structure. That is, the leaf CU may include a quadtree indicating how the leaf CU is partitioned into TUs. The root node of the TU quadtree generally corresponds to the leaf CU, while the root node of the CU quadtree generally corresponds to a tree block (or LCU).

[0059] As described above, the video encoder 22 may partition the coding block of a CU into one or more prediction blocks. A prediction block is a rectangular (i.e., square or non-square) block of samples to which the same prediction is applied. The PU of a CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and a syntax structure for predicting the prediction blocks. In a monochrome picture or a picture with three independent color layers, a PU may include a single prediction block and a syntax structure for predicting the prediction block. The video encoder 22 may generate predictive blocks (e.g., luma, Cb, and Cr predictive blocks) for the prediction blocks (e.g., luma, Cb, and Cr prediction blocks) of each PU of the CU.

[0060] Video encoder 22 may use intra prediction or inter prediction to generate the predictive blocks of the PU. If video encoder 22 uses intra prediction to generate the predictive blocks of the PU, video encoder 22 may generate the predictive blocks of the PU based on decoded samples of the picture including the PU.

[0061] After the video encoder 22 generates predictive blocks (e.g., luma, Cb, and Cr predictive blocks) for one or more PUs of a CU, the video encoder 22 may generate one or more residual blocks for the CU. For example, the video encoder 22 may generate a luma residual block for the CU. Each sample in the CU luma residual block indicates the difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. In addition, the video encoder 22 may generate a Cb residual block for the CU. Each sample in the Cb residual block of the CU may indicate the difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU. The video encoder 22 may also generate a Cr residual block for the CU. Each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0062] In addition, as described above, the video encoder 22 may decompose the residual blocks (e.g., luma, Cb, and Cr residual blocks) of a CU into one or more transform blocks (e.g., luma, Cb, and Cr transform blocks) using quadtree partitioning. A transform block is a rectangular (e.g., square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and a syntax structure for transforming the transform block samples. Thus, each TU of a CU may have a luma transform block, a Cb transform block, and a Cr transform block. The luma transform block of a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three independent color layers, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0063] The video encoder 22 may apply one or more transforms to a transform block of a TU to generate a coefficient block of the TU. For example, the video encoder 22 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block of the TU. A coefficient block may be a two-dimensional array of transform coefficients. The transform coefficient may be a scalar. The video encoder 22 may apply one or more transforms to a Cb transform block of a TU to generate a Cb coefficient block of the TU. The video encoder 22 may apply one or more transforms to a Cr transform block of a TU to generate a Cr coefficient block of the TU.

[0064] In some examples, video encoder 22 skips application of a transform to a transform block. In such examples, video encoder 22 may treat residual sample values ​​in the same manner as transform coefficients. Thus, in examples where video encoder 22 skips application of a transform, the following discussion regarding transform coefficients and coefficient blocks may apply to transform blocks of residual samples.

[0065] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 22 may quantize the coefficient block to possibly reduce the amount of data used to represent the coefficient block, possibly providing further compression. Quantization generally refers to the process of compressing a series of values ​​into a single value. For example, quantization may be accomplished by dividing a value by a constant and then rounding to the nearest integer. To quantize the coefficient block, the video encoder 22 may quantize the transform coefficients of the coefficient block. After the video encoder 22 quantizes the coefficient block, the video encoder 22 may entropy encode a syntax element indicating the quantized transform coefficients. For example, the video encoder 22 may perform context-adaptive binary arithmetic coding (CABAC) or other entropy coding techniques on the syntax element indicating the quantized transform coefficients.

[0066] The video encoder 22 may output a bitstream including a sequence of bits forming a representation of a coded picture and associated data. Thus, the bitstream includes a coded representation of video data. The bitstream may include a sequence of network abstraction layer (NAL) units. A NAL unit is a grammatical structure containing an indication of the type of data in a NAL unit and bytes of the data in the form of a raw byte sequence payload (RBSP) interspersed with emulation prevention bits when necessary. Each of the NAL units may include a NAL unit header and may encapsulate an RBSP. The NAL unit header may include a grammatical element indicating a NAL unit type code. The NAL unit type code specified by the NAL unit header of the NAL unit indicates the type of the NAL unit. The RBSP may be a grammatical structure containing an integer number of bytes encapsulated within a NAL unit. In some cases, the RBSP includes zero bits.

[0067] The video decoder 30 may receive a bitstream generated by the video encoder 22. The video decoder 30 may decode the bitstream to reconstruct a picture of the video data. As part of decoding the bitstream, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a picture of the video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data may generally be inverse to the process performed by the video encoder 22. For example, the video decoder 30 may use the motion vector of the PU to determine the predictive block of the PU of the current CU. In addition, the video decoder 30 may inversely quantize the coefficient block of the TU of the current CU. The video decoder 30 may inversely transform the coefficient block to reconstruct the transform block of the TU of the current CU. The video decoder 30 may reconstruct the coding block of the current CU by adding the samples of the predictive block of the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. By reconstructing the coding block of each CU of the picture, the video decoder 30 may reconstruct the picture.

[0068] The following introduces common concepts and certain design aspects of HEVC, focusing on techniques for block segmentation. In HEVC, the largest coding unit in a slice is called CTB. The CTB is partitioned according to a quadtree structure, and its nodes are coding units. The multiple nodes in the quadtree structure include leaf nodes and non-leaf nodes. The leaf node has no child nodes in the tree structure (i.e., the leaf node is not further divided). The non-leaf node includes the root node of the tree structure. The root node corresponds to an initial video block of video data (e.g., CTB). For each corresponding non-root node in the multiple nodes, the corresponding non-root node corresponds to a video block, which is a child block of the video block of the parent node in the tree structure corresponding to the corresponding non-root node. Each corresponding non-leaf node of the multiple non-leaf nodes has one or more child nodes in the tree structure.

[0069] In the HEVC Main Profile, CTBs range in size from 16x16 to 64x64 (although 8x8 CTB size is technically supported). As described in WJ Han et al., "Improved Video Compression Efficiency Through Flexible Unit Representation and Corresponding Extension of Coding Tools," IEEE Transactions on Circuits and Systems for Video Technology, Vol. 20, No. 12, pp. 1709-1720, December 2010, CTBs can be recursively partitioned into CUs in a quadtree fashion and as Figure 2 As shown. Figure 2 As shown, each level of the partitioning is a quadtree divided into four sub-blocks. The black blocks are examples of leaf nodes (i.e., blocks that are not further partitioned).

[0070] In some examples, the CU can be the same size as the CTB, although the CU can be as small as 8x8. Each CU is encoded with a coding mode, which can be, for example, an intra coding mode or an inter coding mode. Other coding modes are also possible, including coding modes for screen content (e.g., intra block copy mode, palette-based coding mode, etc.). When the CU is inter-coded (i.e., inter mode is applied), the CU can be further partitioned into prediction units (PUs). For example, a CU can be partitioned into 2 or 4 PUs. In another example, when no further partitioning is applied, the entire CU is considered a single PU. In the HEVC example, when there are two PUs in a CU, they can be half-sized rectangles or two rectangles with a size of ¼ or ¾ of the CU.

[0071] like Figure 3As shown, in HEVC, for a CU encoded using an inter-frame prediction mode, there are eight partitioning modes, namely PART_2Nx2N, PART_2NxN, PART_Nx2N, PART_NxN, PART_2NxnU, PART_2NxnD, PART_nLx2N, and PART_nRx2N. Figure 3 As shown, the CU encoded with the partition mode PART_2Nx2N is not further divided. That is, the entire CU is regarded as a single PU (PU0). The CU encoded with the partition mode PART_2NxN is symmetrically divided horizontally into two PUs (PU0 and PU1). The CU encoded with the partition mode PART_Nx2N is symmetrically divided vertically into two PUs. The CU encoded with the partition mode PART_NxN is symmetrically divided into 4 PUs of equal size (PU0, PU1, PU2, PU3).

[0072] The CU coded with the partition mode PART_2NxnU is asymmetrically split horizontally into a PU0 (upper PU) with 1 / 4 of the CU size and a PU1 (lower PU) with 3 / 4 of the CU size. The CU coded with the partition mode PART_2NxnD is asymmetrically split horizontally into a PU0 (upper PU) with 3 / 4 of the CU size and a PU1 (lower PU) with 1 / 4 of the CU size. The CU coded with the partition mode PART_nLx2N is asymmetrically split vertically into a PU0 (left PU) with 1 / 4 of the CU size and a PU1 (right PU) with 3 / 4 of the CU size. The CU coded with the partition mode PART_nRx2N is asymmetrically split vertically into a PU0 (left PU) with 3 / 4 of the CU size and a PU1 (right PU) with 1 / 4 of the CU size.

[0073] When a CU is inter-coded, each PU has a set of motion information (such as motion vector, prediction direction, and reference picture). In addition, each PU is encoded with a unique inter-frame prediction mode to obtain the set of motion information. However, it should be understood that even if two PUs are uniquely coded, in some cases they may still have the same motion information.

[0074] In J. An et al., "Block partitioning structure for next generation video coding", International Telecommunication Union, COM16-C966, September 2015 (hereinafter referred to as "VCEG proposal COM16-C966"), a quadtree binary tree (QTBT) partitioning technique is proposed for future video coding standards after HEVC. Simulation results have shown that the proposed QTBT structure is more efficient than the quadtree structure already used in HEVC.

[0075] In the QTBT structure of the proposed VCEG proposal COM16-C966, the CTB is first partitioned using a quadtree partitioning technique, where the quadtree partitioning of a node can be iterated until the node reaches the minimum allowed quadtree leaf node size. The minimum allowed quadtree leaf node size can be indicated to the video decoder by the value of the syntax element MinQTSize. If the quadtree leaf node size is not greater than the maximum allowed binary tree root node size (e.g., represented by the syntax element MaxBTSize), the quadtree leaf node can be further partitioned using binary tree partitioning. The binary tree partitioning of a node can be iterated until the node reaches the minimum allowed binary tree leaf node size (e.g., represented by the syntax element MinBTSize) or the maximum allowed binary tree depth (e.g., represented by the syntax element MaxBTDepth). VCEG proposal COM16-C966 uses the term "CU" to represent a binary tree leaf node. In VCEG proposal COM16-C966, a CU is used for prediction (e.g., intra-frame prediction, inter-frame prediction, etc.) and transformation without any further partitioning. Generally speaking, according to the QTBT technique, there are two types of partitioning for binary tree partitioning: symmetric horizontal partitioning and symmetric vertical partitioning. In each case, the block is partitioned by cutting the block horizontally or vertically from the middle.

[0076] In one example of a QTBT partitioning structure, the CTU size is set to 128x128 (e.g., 128x128 luma block and two corresponding 64x64 chroma blocks), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate a quadtree leaf node. The quadtree leaf node can have a size from 16x16 (i.e., MinQTSize is 16x16) to 128x128 (i.e., CTU size). According to one example of QTBT partitioning, if the leaf quadtree node is 128x128, the leaf quadtree node cannot be further divided by a binary tree because the size of the leaf quadtree node exceeds MaxBTSize (i.e., 64x64). Otherwise, the leaf quadtree node will be further divided by a binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. The binary tree depth reaches MaxBTDepth (e.g., 4), which means that there are no further partitions. The binary tree node has a width equal to MinBTSize (e.g., 4), which means that there are no further horizontal partitions. Similarly, the binary tree node has a height equal to MinBTSize, which means that there are no further vertical partitions. The leaf nodes (CUs) of the binary tree are further processed (e.g., by performing prediction processing and transform processing) without any further partitioning.

[0077] Figure 4A An example of a block 50 (eg, CTB) partitioned using the QTBT partitioning technique is shown. Figure 4A As shown, using the QTBT partitioning technique, each of the resulting blocks is partitioned symmetrically through the center of each block. Figure 4B Shows the corresponding Figure 4B The tree structure of the block partition. Figure 4B The solid lines in the figure represent quadtree partitions, and the dashed lines represent binary tree partitions. In one example, in each partition (i.e., non-leaf) node of the binary tree, a syntax element (e.g., a flag) is signaled to indicate the type of partitioning performed (e.g., horizontal or vertical), where 0 indicates horizontal partitioning and 1 indicates vertical partitioning. For quadtree partitioning, there is no need to indicate the partitioning type, because quadtree partitioning always divides the block horizontally and vertically into 4 sub-blocks of equal size.

[0078] like Figure 4B As shown, at node 70, block 50 is divided into Figure 4A The four blocks 51, 52, 53 and 54 are shown. Block 54 is not further divided and is therefore a leaf node. At node 72, block 51 is further divided into two blocks using BT partitioning. Figure 4BAs shown, node 72 is marked with 1, indicating a vertical partition. Thus, the partition at node 72 results in block 57 and a block including blocks 55 and 56. Blocks 55 and 56 are generated by further vertical partitioning at node 74. At node 76, block 52 is further partitioned into two blocks 58 and 59 using BT partitioning. Figure 4B As shown in FIG. 1 , node 76 is labeled with a 1, indicating a horizontal partition.

[0079] At node 78, block 53 is divided into 4 equal-sized blocks using QT partitioning. Blocks 63 and 66 are created by this QT partitioning and are not partitioned further. At node 80, the top left block is first partitioned using vertical binary tree partitioning, resulting in block 60 and the vertical block on the right. The vertical block on the right is then partitioned into blocks 61 and 62 using horizontal binary tree partitioning. The bottom right block created by quadtree partitioning at node 78 is partitioned into blocks 64 and 65 using horizontal binary tree partitioning at node 84.

[0080] In other embodiments, the MTT-based CU structure may replace the QT-based, BT-based and / or QTBT-based CU structure. The MTT partition structure is still a recursive tree structure. However, multiple different partition structures (e.g., three or more) are used. For example, according to the MTT technology of the present disclosure, three or more different partition structures may be used at each depth of the tree structure. In this context, the depth of a node in the tree structure may refer to the path length (e.g., the number of divisions) from the node to the root of the tree structure. As used in the present disclosure, the partition structure may generally refer to how many different blocks a block can be divided into. For example, a quadtree partition structure may divide a block into four blocks, a binary tree partition structure may divide a block into two blocks, and a ternary tree partition structure may divide a block into three blocks. The partition structure may have a variety of different partition types, as will be explained in more detail below. The partition type may further define how to divide the block, including symmetrical or asymmetrical partitioning, uniform or non-uniform partitioning, and / or horizontal or vertical partitioning.

[0081] In one example of the technology according to the present disclosure, the video encoder 22 may be configured to receive a picture of video data, and use three or more different segmentation structures to segment the picture of the video data into a plurality of blocks, and encode the plurality of blocks of the picture of the video data. Similarly, the video decoder 30 may be configured to receive a bitstream including a bit sequence representing a coded picture of the video data, determine to segment the coded picture of the video data into a plurality of blocks using three or more different segmentation structures, and reconstruct the plurality of blocks of the coded picture of the video data. In one example, segmenting a video data frame includes segmenting the video data frame into a plurality of blocks using three or more different segmentation structures, wherein at least three of the three or more different segmentation structures may be used to represent each depth of a tree structure of how to segment the video data frame. In one example, the three or more different segmentation structures include a ternary tree segmentation structure, and the video encoder 22 and / or the video decoder 30 may be configured to segment one of the plurality of video data blocks using a ternary tree segmentation type of the ternary tree segmentation structure, wherein the ternary tree segmentation structure divides one of the plurality of blocks into three sub-blocks without dividing one of the plurality of blocks through the center. In another embodiment of the present disclosure, the three or more different partitioning structures further include a quadtree partitioning structure and a binary tree partitioning structure.

[0082] Thus, in one example, video encoder 22 may generate an encoded representation of an initial video block (e.g., a coded tree block or CTU) of video data. As part of generating the encoded representation of the initial video block, video encoder 22 determines a tree structure including a plurality of nodes. For example, video encoder 22 may partition the tree block using the MTT partitioning structure of the present disclosure.

[0083] The multiple nodes in the MTT segmentation structure include multiple leaf nodes and multiple non-leaf nodes. The leaf node has no child nodes in the tree structure. The non-leaf node includes the root node of the tree structure. The root node corresponds to the initial video block. For each corresponding non-root node in the multiple nodes, the corresponding non-root node corresponds to a video block (e.g., a coding block), which is a child block of the video block of the parent node in the tree structure corresponding to the corresponding non-root node. Each corresponding non-leaf node of the multiple non-leaf nodes has one or more child nodes in the tree structure. In some examples, due to forced partitioning and one of the child nodes corresponding to a block outside the picture boundary, the non-leaf node at the picture boundary may have only one child node.

[0084] According to the technology of the present disclosure, for each corresponding non-leaf node of the tree structure at each depth level of the tree structure, there are multiple allowed partitioning modes (e.g., segmentation structures) for the corresponding non-leaf node. For example, for each depth of the tree structure, three or more segmentation structures may be allowed. The video encoder 22 may be configured to segment the video block corresponding to the corresponding non-leaf node into video blocks corresponding to child nodes of the corresponding non-leaf node according to one of the multiple allowed segmentation structures. Each corresponding allowed segmentation structure in the multiple allowed segmentation structures may correspond to a different way of segmenting the video block corresponding to the corresponding non-leaf node into video blocks corresponding to child nodes of the corresponding non-leaf node. In addition, in the present example, the video encoder 22 may include an encoded representation of the initial video block in a bitstream, which includes an encoded representation of the video data.

[0085] In a similar example, the video decoder 30 can determine a tree structure including a plurality of nodes. As in the previous example, the plurality of nodes include a plurality of leaf nodes and a plurality of non-leaf nodes. The leaf node has no child nodes in the tree structure. The non-leaf node includes the root node of the tree structure. The root node corresponds to an initial video block of the video data. For each corresponding non-root node in the plurality of nodes, the corresponding non-root node corresponds to a video block, which is a child block of a video block of a parent node in the tree structure corresponding to the corresponding non-root node. Each corresponding non-leaf node of the plurality of non-leaf nodes has one or more child nodes in the tree structure. For each corresponding non-leaf node of the tree structure at each depth level of the tree structure, there are a plurality of allowed partitioning modes for the corresponding non-leaf node, and according to one of the plurality of allowed partitioning modes, the video block corresponding to the corresponding non-leaf node is divided into a video block corresponding to a child node of the corresponding non-leaf node. Each corresponding allowed partitioning mode in the plurality of allowed partitioning modes corresponds to a different way of dividing the video block corresponding to the corresponding non-leaf node into a video block corresponding to a child node of the corresponding non-leaf node. Furthermore, in this example, for each (or at least one) corresponding leaf node of the tree structure, the video decoder 30 reconstructs a video block corresponding to the corresponding leaf node.

[0086] In some such examples, for each corresponding non-leaf node other than the root node in the tree structure, multiple allowed partitioning modes (e.g., a segmentation structure) for the corresponding non-leaf node are independent of the segmentation structure, according to which the video block corresponding to the parent node of the corresponding non-leaf node is partitioned into video blocks corresponding to the child nodes of the parent node of the corresponding non-leaf node.

[0087] In other examples of the present disclosure, at each depth of the tree structure, the video encoder 22 may be configured to further divide the subtree using a specific partition type from one of three or more partition structures. For example, the video encoder 22 may be configured to determine a specific partition type from QT, BT, triple-tree (TT), and other partition structures. In one example, the QT partition structure may include square quadtree and rectangular quadtree partition types. The video encoder 22 may use square quadtree partitioning to partition a square block by dividing the block horizontally and vertically along the center into four square blocks of equal size. Similarly, the video encoder 22 may use rectangular quadtree partitioning to partition a rectangular (e.g., non-square) block by dividing the rectangular block horizontally and vertically along the center into four rectangular blocks of equal size.

[0088] The BT partitioning structure may include horizontal symmetric binary tree, vertical symmetric binary tree, horizontal asymmetric binary tree and vertical asymmetric binary tree partitioning types. For the horizontal symmetric binary tree partitioning type, the video encoder 22 may be configured to divide the block horizontally along the center of the block into two symmetric blocks of the same size. For the vertical symmetric binary tree partitioning type, the video encoder 22 may be configured to divide the block vertically along the center of the block into two symmetric blocks of the same size. For the horizontal asymmetric binary tree partitioning type, the video encoder 22 may be configured to divide the block horizontally into two blocks of different sizes. For example, one block may be 1 / 4 the size of the parent block, and the other block may be 3 / 4 the size of the parent block, such as Figure 3 For the vertical asymmetric binary tree partition type, the video encoder 22 may be configured to vertically divide the block into two blocks of different sizes. For example, one block may be 1 / 4 the size of the parent block, and the other block may be 3 / 4 the size of the parent block, as shown in FIG. Figure 3 The same is true for the PART_nLx2N or PART_nRx2N segmentation types.

[0089] In other examples, an asymmetric binary tree partitioning type can divide the parent block into parts of different sizes. For example, one child block can be 3 / 8 of the parent block, while another child block can be 5 / 8 of the parent block. Of course, this partitioning type can be vertical or horizontal.

[0090] The TT partition structure is different from the QT or BT structure in that the TT partition structure does not divide the block along the center. The central area of ​​the block is kept together in the same sub-block. Unlike the QT that produces four blocks or the binary tree that produces two blocks, the division according to the TT partition structure produces three blocks. Example partition types according to the TT partition structure include symmetric partition types (both horizontal and vertical), and asymmetric partition types (both horizontal and vertical). In addition, the symmetric partition type according to the TT partition structure can be uneven / inconsistent or uniform / consistent. The asymmetric partition type of the TT partition structure according to the present disclosure is uneven / inconsistent. In one example of the present disclosure, the TT partition structure may include the following partition types: a horizontally uniform / consistent symmetric ternary tree, a vertically uniform / consistent symmetric ternary tree, a horizontally uneven / inconsistent symmetric ternary tree, a vertically uneven / inconsistent symmetric ternary tree, a horizontally uneven / inconsistent asymmetric ternary tree, and a vertically uneven / inconsistent asymmetric ternary tree partition type.

[0091] Generally speaking, an uneven / inconsistent symmetric ternary tree partition type is a partition type that is symmetric about the center line of the block, but in which at least one of the three blocks of the resulting block is not the same size as the other two blocks. A preferred example is that the side blocks are 1 / 4 the size of the block, and the center block is 1 / 2 the size of the block. An even / consistent symmetric ternary tree partition type is a partition type that is symmetric about the center line of the block, and the resulting blocks are all the same size. Such a partition is possible if the block height or width is a multiple of 3, depending on the vertical or horizontal partition. An uneven / inconsistent asymmetric ternary tree partition type is a partition type that is asymmetric about the center line of the block, and in which at least one of the resulting blocks is not the same size as the other two.

[0092] Figure 5A is a conceptual diagram illustrating an example horizontal ternary tree partition type. Figure 5B is a conceptual diagram showing an example vertical ternary tree partition type. Figure 5A and Figure 5B In both, h represents the height of the block in luma or chroma samples, and w represents the width of the block in luma or chroma samples. Figure 5A and Figure 5B The corresponding "center line" in each of the ternary tree partitions does not represent the boundary of the block (i.e., the ternary tree partition does not divide the block by the center line). Instead, the center line is shown to depict whether a particular partition type is symmetrical or asymmetrical relative to the center line of the original block. The depicted center line is also along the direction of the partition.

[0093] like Figure 5AAs shown in , block 71 is partitioned using a horizontally uniform / consistent symmetric partitioning type. The horizontally uniform / consistent symmetric partitioning type produces a top half and a bottom half that are symmetric with respect to the center line of block 71. The horizontally uniform / consistent symmetric partitioning type produces three equally sized sub-blocks, each having a height of h / 3 and a width of w. The horizontally uniform / consistent symmetric partitioning type is possible when the height of block 71 is divisible by 3.

[0094] The block 73 is partitioned using a horizontally uneven / inconsistent symmetric partitioning type. The horizontally uneven / inconsistent symmetric partitioning type produces a top half and a bottom half that are symmetrical with respect to a center line of the block 73. The horizontally uneven / inconsistent symmetric partitioning type produces two blocks of equal size (e.g., a top and bottom block with a height of h / 4), and a center block of a different size (e.g., a center block with a height of h / 2). In one example of the present disclosure, according to the horizontally uneven / inconsistent symmetric partitioning type, the area of ​​the center block is equal to the combined area of ​​the top and bottom blocks. In some examples, the horizontally uneven / inconsistent symmetric partitioning type may be preferably used for blocks having a height that is a power of 2 (e.g., 2, 4, 8, 16, 32, etc.).

[0095] The block 75 is segmented using a horizontally uneven / non-uniform symmetrical segmentation type. The horizontally uneven / non-uniform asymmetrical segmentation type does not produce symmetrical top and bottom halves relative to the centerline of the block 75 (i.e., the top and bottom halves are asymmetrical). Figure 5A In the example of , the horizontally uneven / inconsistent asymmetric partitioning type produces a top block with a height of h / 4, a center block with a height of 3h / 8, and a bottom block with a height of 3h / 8. Of course, other asymmetric arrangements can be used.

[0096] like Figure 5B As shown in , block 77 is partitioned using a vertically uniform / consistent symmetric partitioning type. The vertically uniform / consistent symmetric partitioning type produces a left half and a right half that are symmetric with respect to the center line of block 77. The vertically uniform / consistent symmetric partitioning type produces three sub-blocks of equal size, each having a width of w / 3 and a height of h. The vertically uniform / consistent symmetric partitioning type is possible when the width of block 77 is divisible by 3.

[0097] Block 79 is partitioned using a vertically uneven / inconsistent symmetric partitioning type. The vertically uneven / inconsistent symmetric partitioning type produces a left half and a right half that are symmetrical with respect to the center line of block 79. The vertically uneven / inconsistent symmetric partitioning type produces a left half and a right half that are symmetrical with respect to the center line of 79. The vertically uneven / inconsistent symmetric partitioning type produces two blocks of equal size (e.g., left and right blocks with a width of w / 4), and a center block of different size (e.g., a center block with a width of w / 2). In one example of the present disclosure, according to the vertically uneven / inconsistent symmetric partitioning type, the area of ​​the center block is equal to the combined area of ​​the left and right blocks. In some examples, the vertically uneven / inconsistent symmetric partitioning type may be preferably used for blocks having a width that is a power of 2 (e.g., 2, 4, 8, 16, 32, etc.).

[0098] The block 81 is segmented using a vertically uneven / inconsistent asymmetric segmentation type. The vertically uneven / inconsistent asymmetric segmentation type does not produce symmetrical left and right halves relative to the center line of the block 81 (i.e., the left and right halves are asymmetric). Figure 5B In the example of , the vertically uneven / non-uniform asymmetric partitioning type produces a left block having a width of w / 4, a center block having a width of 3w / 8, and a right block having a width of 3w / 8. Of course, other asymmetric arrangements may be used.

[0099] In an example where a block (e.g., at a subtree node) is partitioned into an asymmetric ternary tree partition type, the video encoder 22 and / or the video decoder 30 may apply a restriction so that two of the three partitions have the same size. Such a restriction may correspond to a restriction that the video encoder 22 must comply with when encoding video data. In addition, in some examples, the video encoder 22 and the video decoder 30 may apply a restriction whereby, when partitioning is performed according to the asymmetric ternary tree partition type, the sum of the areas of the two partitions is equal to the area of ​​the remaining partitions. For example, the video encoder 22 may generate or the video decoder 30 may receive an encoded representation of an initial video block that complies with the restriction, the restriction indicating that when a video block corresponding to a node of a tree structure is partitioned according to an asymmetric ternary tree partition mode, the node has a first child node, a second child node, and a third child node, the second child node corresponds to a video block between the video blocks corresponding to the first child node and the third child node, the video blocks corresponding to the first child node and the third child node have the same size, and the sum of the sizes of the video blocks corresponding to the first child node and the third child node is equal to the size of the video block corresponding to the second child node.

[0100] Figure 5Cis a conceptual diagram showing example asymmetric binary tree partition types. An embodiment may include an asymmetric binary tree (ABT) CU encoding tool on top of a QTBT structure. In the ABT structure, the binary tree partitioning mode set of QTBT is extended to support asymmetric binary partitioning of coding units. Specifically, a CU may be split into 2 sub-CUs having 1 / 4 size and 3 / 4 size of the parent CU, respectively. Four new binary tree partitioning modes have been introduced into the QTBT framework, allowing new partitioning configurations. As Figure 5C As shown in , in addition to the partitioning modes already available in QTBT, so-called asymmetric partitioning modes can also be used. Further details of the ABT structure can be found in F. Le Leannec, T. Poirier, F. Urban, "Asymmetric Coding Units in QTBT", International Telecommunication Union, JVET-D0064, October 2016.

[0101] In some examples of the present disclosure, the video encoder 22 may be configured to select from all of the above-described segmentation types for each of the QT, BT, ABT, and TT segmentation structures. In other examples, the video encoder 22 may be configured to determine the segmentation type only from a subset of the above-described segmentation types. For example, a subset of the segmentation types discussed above (or other segmentation types) may be used for certain block sizes of a quadtree structure or for certain depths. The subset of supported segmentation types may be signaled in the bitstream for use by the video decoder 30, or may be predefined so that the video encoder 22 and the video decoder 30 may determine the subset without any signaling.

[0102] In other examples, the number of supported partition types may be fixed for all depths in all CTUs. That is, the video encoder 22 and the video decoder 30 may be preconfigured to use the same number of partition types for any depth of the CTU. In other examples, the number of supported partition types may be different and may depend on the depth, slice type, or other previously encoded information. In one example, at depth 0 or depth 1 of the tree structure, only the QT partition structure is used. At depths greater than 1, each of the QT, BT, and TT partition structures may be used.

[0103] In some examples, video encoder 22 and / or video decoder 30 may apply preconfigured constraints to supported partition types to avoid repeated partitioning of a region of a video image or a region of a CTU. In one example, when a block is partitioned using an asymmetric partition type, video encoder 22 and / or video decoder 30 may be configured to not further partition the largest sub-block partitioned from the current block. For example, when a block is partitioned according to an asymmetric partition type (e.g., Figure 3 When dividing a square block into PART_2NxnU partition types, the largest sub-block among all sub-blocks (for example, Figure 3 PU1 of PART_2NxnU partition type) is a noted leaf node and cannot be further partitioned. However, smaller sub-blocks (e.g., Figure 3 PU0) in PART_2NxnU partition type can be further divided.

[0104] As another example where constraints may be applied to supported partition types to avoid repeating the partitioning of a region, when a block is partitioned using an asymmetric partition type, the largest sub-block partitioned from the current block cannot be further partitioned in the same direction. Figure 3 When dividing a square block into two sub-blocks (PART_2NxnU partition type in the example), the video encoder 22 and / or the video decoder 30 may be configured not to divide a large sub-block (e.g., Figure 3 However, in this example, video encoder 22 and / or video decoder 30 may divide PU1 again in the vertical direction.

[0105] As another example of where constraints may be applied to supported partitioning types to avoid difficulties in further partitioning, video encoder 22 and / or video decoder 30 may be configured to not partition a block horizontally or vertically when the width / height of the block is not a power of 2 (e.g., when the width / height is not 2, 4, 8, 16, etc.).

[0106] The above examples describe how the video encoder 22 can be configured to perform MTT segmentation according to the techniques of the present disclosure. The video decoder 30 can then also apply the same MTT segmentation as performed by the video encoder 22. In some examples, how the video encoder 22 segments the pictures of the video data can be determined by applying the same set of predefined rules at the video decoder 30. However, in many cases, the video encoder 22 can determine the specific segmentation structure and segmentation type to be used based on the rate-distortion metric for a particular picture of the video data being encoded. Thus, in order for the video decoder 30 to determine the segmentation for a particular picture, the video encoder 22 can signal syntax elements in the encoded bitstream that indicate how the picture and the CTU of the picture will be segmented. The video decoder 30 can parse such syntax elements and segment the picture and CTU accordingly.

[0107] In one example of the present disclosure, the video encoder 22 may be configured to signal a specific subset of supported partition types as a high-level syntax element in a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, an adaptive parameter set (APS), or any other high-level syntax parameter set. For example, the maximum number of partition types and which types are supported may be predefined in a sequence parameter set (SPS), a picture parameter set (PPS), or any other high-level syntax parameter set or signaled in the bitstream as a high-level syntax element. The video decoder 30 may be configured to receive and parse such syntax elements to determine the specific subset of partition types used and / or the maximum number of supported partition structures (e.g., QT, BT, TT, etc.) and types.

[0108] In some examples, at each depth, the video encoder 22 may be configured to signal an index indicating the selected split type used at that depth of the tree structure. Additionally, in some examples, the video encoder 22 may adaptively signal such a split type index at each CU, i.e., the index may be different for different CUs. For example, the video encoder 22 may set the index of the split type based on one or more rate-distortion calculations. In one example, signaling of the split type (e.g., the index of the split type) may be skipped if certain conditions are met. For example, when there is only one supported split type associated with a particular depth, the video encoder 22 may skip signaling of the split type. In this example, when approaching the picture boundary, the region to be encoded may be smaller than the CTU. Thus, in this example, the CTU may be forced to be partitioned to fit the picture boundary. In one example, only a symmetric binary tree is used for the forced partitioning and the split type is not signaled. In some examples, at a certain depth, the split type may be derived based on previously encoded information such as slice type, CTU depth, CU position.

[0109] In another example of the present disclosure, for each CU (leaf node), the video encoder 22 may be further configured to signal a syntax element (e.g., a 1-bit transform_partition_flag) to indicate whether a transform will be performed on the same size of the CU (i.e., this flag indicates whether the TU is the same size as the CU or is further partitioned). In the case where the transform_partition_flag is signaled as true, the video encoder 22 may be configured to further divide the residual of the CU into multiple sub-blocks and perform a transform on each sub-block. The video decoder 30 may perform the inverse process.

[0110] In one example, when the transform_partition_flag is signaled as true, the following operations are performed. If the CU corresponds to a square block (i.e., the CU is square), then the video encoder 22 uses a quadtree partition to divide the residual into four square sub-blocks and performs a transform on each square sub-block. If the CU corresponds to a non-square block, e.g., MxN, then the video encoder 22 divides the residual into two sub-blocks, with sub-block sizes of 0.5MxN when M > N and Mx0.5N when M < N. As another example, when the transform_partition_flag is signaled as true and the CU corresponds to a non-square block, e.g., MxN, (i.e., the CU is non-square), the video encoder 22 may be configured to divide the residual into sub-blocks with a size of KxK, and a KxK square transform is used for each sub-block, where K is equal to the greatest common factor (factor) of M and N. As another example, when the CU is a square block, the transform_partition_flag is not signaled.

[0111] In some examples, when there is a residual in the predicted CU, the split flag is not signaled and only a transform with one derived size is used. For example, for a CU with size equal to MxN, a KxK square transform is used, where K is equal to the maximum factor of M and N. Therefore, in this example, for a CU with size 16x8, the same 8x8 transform can be applied to both 8x8 sub-blocks of the residual data of the CU. A "split flag" is a syntax element that indicates that a node in a tree structure has child nodes in the tree structure.

[0112] In some examples, for each CU, if the CU is not partitioned into a square quadtree or a symmetric binary tree, video encoder 22 is configured to always set the transform size equal to the partition size (eg, the size of the CU).

[0113] Simulation results have shown that encoding performance using the disclosed MTT technique has shown improvement in random access cases compared to the JEM-3.1 reference software. On average, simulations have shown that the disclosed MTT technique has provided a 3.18% reduction in bitrate distortion (BD) rate with only a modest increase in encoding time. Simulations have shown that the disclosed MTT technique provides good performance for higher resolutions, for example, 4.20% and 4.89% reduction in luma BD rate for Class A1 and Class A2 tests. Class A1 and Class A2 are example 4K resolution test sequences.

[0114] It should be appreciated that for each of the above examples described with reference to video encoder 22, video decoder 30 may be configured to perform the inverse process. With respect to signaling syntax elements, video decoder 30 may be configured to receive and parse such syntax elements and segment and decode the associated video data accordingly.

[0115] In a specific example of the present disclosure, the video decoder may be configured to segment the video block according to three different segmentation structures (QT, BT, and TT), allowing five different segmentation types at each depth. FIG. 5A to FIG. 5C As shown, the partitioning types include quadtree partitioning (QT partitioning structure), horizontal binary tree partitioning (BT partitioning structure), vertical binary tree partitioning (BT partitioning structure), horizontal center-side ternary tree partitioning (TT partitioning structure) and vertical center-side ternary tree partitioning (TT partitioning structure).

[0116] The five segmentation types are defined as follows. Note that a square is considered a special case of a rectangle.

[0117] Quadtree partitioning: The block is further divided into four rectangular blocks of equal size. Fig. 6A An example of quadtree partitioning is shown.

[0118] Vertical binary tree partitioning: Divide the block vertically into two rectangular blocks of equal size. Figure 6B is an example of vertical binary tree partitioning.

[0119] Horizontal binary tree partitioning: Divide the block horizontally into two rectangular blocks of equal size. Figure 6C is an example of horizontal binary tree partitioning.

[0120] • Vertical center-side ternary tree partitioning: Divide the block vertically into three rectangular blocks, so that the two side blocks share the same size, and the size of the center block is the sum of the two side blocks. Fig.6D is an example of a vertical center-side ternary tree partition.

[0121] Horizontal center-side ternary tree partitioning: Divide the block horizontally into three rectangular blocks, so that the two side blocks share the same size, and the size of the center block is the sum of the two side blocks. Fig. 6E is an example of a horizontal center-side ternary tree partition.

[0122] For a block associated with a particular depth, the video encoder 22 determines which partition type to use (including no further partitioning) and signals the determined partition type to the video decoder 30, either explicitly or implicitly (e.g., the partition type may be derived from a predetermined rule). The video encoder 22 may determine the partition type to use based on examining the rate-distortion cost of blocks using different partition types. To obtain the rate-distortion cost, the video encoder 22 may need to recursively examine possible partition types for the block.

[0123] Figure 7 is a conceptual diagram showing an example of coding tree unit (CTU) partitioning. In other words, Figure 7 The partitioning of CTB 91 corresponding to a CTU is shown. Figure 7 In the example,

[0124] • At depth 0, CTB 91 (ie, the entire CTB) is divided into two blocks using horizontal binary tree partitioning (as indicated by line 93 with connectors separated by a single dot).

[0125] At depth 1:

[0126] • Divide the upper block into three blocks using a vertical center-side tri-tree partition (as indicated by lines 95 and 86 with small connectors).

[0127] • Divide the bottom block into four blocks using quadtree partitioning (as indicated by lines 88 and 90 with connectors separated by two dots).

[0128] At depth 2:

[0129] • Divide the left block of the upper block at depth 1 into three blocks using horizontal center-side ternary tree partitioning (as indicated by lines 92 and 94 with long dashes separated by short dashes).

[0130] The center block and the right block of the upper block at depth 1 are no longer divided.

[0131] The four blocks at the bottom block at depth 1 are not split any more.

[0132] exist Figure 7 As can be seen in the example, three different partitioning structures (BT, QT and TT) are used, along with four different partitioning types (horizontal binary tree partitioning, vertical center-side tritree partitioning, quadtree partitioning and horizontal center-side tritree partitioning).

[0133] In another example, additional constraints may be imposed on blocks at a certain depth or with a certain size. For example, when the height / width of a block is less than 16 pixels, the block cannot be partitioned using a vertical / horizontal center-side tree to avoid blocks with a height / width less than 4 pixels.

[0134] Although the QTBT, ABT, and TT structures discussed herein may have better coding performance than the quadtree structure used in HEVC, further flexibility is desired. For example, in some examples, the CU leaf nodes in the BT, ABT, and TT structures may be further divided into recursive transform units.

[0135] RQT in HEVC only allows the residual block of a CU to be divided into TUs in a quadtree structure, which is not flexible enough. In some examples, the residual block of a CU can be flexibly recursively divided into TUs in QT, BT, ABT, TT or other structures.

[0136] Transformation unit partition structure

[0137] In order to achieve more flexible partitioning of CTU, in some embodiments, the residual block of the CU can be further recursively divided into transform units in the manner of QT, BT, ABT, TT or any other partitioning or any combination thereof. For example, the video encoder 22 or the decoder 30 may determine to partition the coded picture into multiple coding units according to a first tree structure including leaf nodes. The video encoder 22 or the decoder 30 may determine to recursively partition the residual block of the leaf node into multiple transform units according to a second tree structure.

[0138] In one example, a flag is encoded in the video bitstream (encoded into the bitstream via the video encoder and decoded from the bitstream via the video decoder) at the root of the transform unit tree so that only a single transform unit partitioning structure of the second tree is allowed in the corresponding tree of the entire residual block. For example, as an example, if only vertical center-side trees are allowed in the tree, the TUs in the tree can only be divided into vertical center-side trees.

[0139] In another example, a flag is encoded for each partitioned TU or group (eg, subset) of TUs within the same block to determine in which structure the TU will be partitioned.

[0140] Change information signaling notification

[0141] Optionally, for TUs, transform tools such as enhanced multiple transform (EMT), non-separable secondary transform (NSST), and transform skip (TS) can be applied. In one example, the flag of the transform tool is encoded only at the root of the transform unit tree (encoded by the encoder in the bitstream and decoded by the decoder from such a bitstream) so that all TUs in the tree share the same flag. For example, in such an example, the EMT flag and index are signaled only at the root level. The EMT flag and index will determine whether EMT is applied and which EMT set is applied to all TUs in the same tree. In another example, the flag of the transform tool is encoded for each TU or TU group. For example, the EMT flag and index are encoded for each TU. Each TU will perform EMT based on its own flag.

[0142] Coefficient encoding

[0143] In HEVC, the coded block flag (CBF) and the coefficients are encoded for each TU. By signaling the CBF and the last position of the non-zero coefficients, such encoding may suffer from a large overhead. According to some embodiments, in order to reduce this overhead, the coefficients of the entire transform unit tree (in other words, for all TUs) may be signaled together. In such an example, only the CBF at the root level and the last position of the non-zero coefficients are signaled.

[0144] Various examples have been described. Specific examples of the present disclosure may be used alone or in combination with each other.

[0145] Figure 8 is a block diagram illustrating an example video encoder 22 in which the techniques of this disclosure may be implemented. Figure 8It is provided for the purpose of explanation and should not be considered as limiting the techniques broadly exemplified and described in this disclosure. The techniques of this disclosure can be applied to various encoding standards or methods.

[0146] exist Figure 8 In the example of , the video encoder 22 includes a prediction processing unit 100, a video data memory 101, a residual generation unit 102, a transform processing unit 104, a quantization unit 106, an inverse quantization unit 108, an inverse transform processing unit 110, a reconstruction unit 112, a filter unit 114, a decoded picture buffer 116, and an entropy encoding unit 118. The prediction processing unit 100 includes an inter-prediction processing unit 120 and an intra-prediction processing unit 126. The inter-prediction processing unit 120 may include a motion estimation unit and a motion compensation unit (not shown).

[0147] The video data memory 101 may be configured to store video data to be encoded by components of the video encoder 22. The video data stored in the video data memory 101 may, for example, be obtained from a video source 18. The decoded picture buffer 116 may be a reference picture memory that stores reference video data for encoding the video data by the video encoder 22, for example, for use in intra-frame coding or inter-frame coding modes. The video data memory 101 and the decoded picture buffer 116 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 101 and the decoded picture buffer 116 may be provided with the same memory device or with separate memory devices. In various examples, the video data memory 101 may be on-chip with other components of the video encoder 22, or off-chip relative to those components. The video data memory 101 may be provided with a plurality of memory devices. Figure 1 The storage medium 20 in is the same as or a part of it.

[0148] The video encoder 22 receives video data. The video encoder 22 may encode each CTU in a slice of a picture of the video data. Each of the CTUs may be associated with a luma coding tree block (CTB) of equal size and a corresponding CTB of the picture. As part of encoding the CTU, the prediction processing unit 100 may perform partitioning to divide the CTB of the CTU into progressively smaller blocks. The smaller blocks may be coding blocks of the CU. For example, the prediction processing unit 100 may partition the CTB associated with the CTU according to a tree structure. According to one or more techniques of the present disclosure, for each corresponding non-leaf node of the tree structure at each depth level of the tree structure, there are multiple allowed partitioning modes for the corresponding non-leaf node, and according to one of the multiple allowed partitioning modes, the video block corresponding to the corresponding non-leaf node is partitioned into video blocks corresponding to child nodes of the corresponding non-leaf node. In one example, the prediction processing unit 100 or another processing unit of the video encoder 22 may be configured to perform any combination of the above-mentioned MTT partitioning techniques.

[0149] The video encoder 22 may encode the CU of the CTU to generate an encoded representation of the CU (i.e., an encoded CU). As part of encoding the CU, the prediction processing unit 100 may partition the coding blocks associated with the CU among one or more PUs of the CU. According to the technology of the present disclosure, a CU may include only a single PU. That is, in some examples of the present disclosure, the CU is not divided into separate prediction blocks, but prediction processing is performed on the entire CU. Therefore, each CU may be associated with a luminance prediction block and a corresponding chrominance prediction block. The video encoder 22 and the video decoder 30 may support CUs of various sizes. As shown above, the size of the CU may refer to the size of the luminance coding block of the CU or the size of the luminance prediction block. As described above, the video encoder 22 and the video decoder 30 may support CU sizes defined by any combination of the above-described example MTT partitioning types.

[0150] The inter-frame prediction processing unit 120 can generate predictive data for a PU by performing inter-frame prediction on each PU of a CU. As explained above, in some MTT examples of the present disclosure, a CU may contain only a single PU, that is, a CU and a PU may be synonymous. The predictive data of a PU may include a predictive block of the PU and motion information of the PU. The inter-frame prediction processing unit 120 may perform different operations on a PU or a CU depending on whether the PU is in an I slice, a P slice, or a B slice. In an I slice, all PUs are intra-predicted. Therefore, if a PU is in an I slice, the inter-frame prediction processing unit 120 does not perform inter-frame prediction on the PU. Therefore, for a block encoded in I mode, the predicted block is formed using spatial prediction from a neighboring block previously encoded in the same picture. If the PU is in a P slice, the inter-frame prediction processing unit 120 may use unidirectional inter-frame prediction to generate a predictive block for the PU. If the PU is in a B slice, the inter-frame prediction processing unit 120 may use unidirectional or bidirectional inter-frame prediction to generate a predictive block for the PU.

[0151] Intra-prediction processing unit 126 may generate predictive data for a PU by performing intra-prediction on the PU. The predictive data for the PU may include predictive blocks and various syntax elements for the PU. Intra-prediction processing unit 126 may perform intra-prediction on PUs in I slices, P slices, and B slices.

[0152] To perform intra prediction on a PU, the intra prediction processing unit 126 may use multiple intra prediction modes to generate multiple predictive data sets for the PU. The intra prediction processing unit 126 may use samples from sample blocks of neighboring PUs to generate a predictive block for the PU. Assuming a left-to-right, top-to-bottom coding order for PUs, CUs, and CTUs, a neighboring PU may be above, to the upper right, to the upper left, or to the left of the PU. The intra prediction processing unit 126 may use various numbers of intra prediction modes, for example, 33 directional intra prediction modes. In some examples, the number of intra prediction modes may depend on the size of the region associated with the PU.

[0153] Prediction processing unit 100 may select predictive data for PUs of a CU from the predictive data generated for the PU by inter-prediction processing unit 120 or the predictive data generated for the PU by intra-prediction processing unit 126. In some examples, prediction processing unit 100 selects the predictive data for the PUs of the CU based on a rate / distortion metric of the predictive data set. The predictive block of the selected predictive data may be referred to herein as a selected predictive block.

[0154] The residual generation unit 102 may generate a residual block of the CU (e.g., luma, Cb, and Cr residual blocks) based on the coding blocks of the CU (e.g., luma, Cb, and Cr coding blocks) and the selected predictive blocks of the PUs of the CU (e.g., predictive luma, Cb, and Cr blocks). For example, the residual generation unit 102 may generate the residual block of the CU so that each sample in the residual block has a value equal to the difference between the sample in the coding block of the CU and the corresponding sample in the corresponding selected predictive block of the PUs of the CU.

[0155] The transform processing unit 104 may perform quadtree partitioning to partition the residual block associated with the CU into transform blocks associated with the TUs of the CU. Thus, a TU may be associated with a luma transform block and two chroma transform blocks. The size and position of the luma and chroma transform blocks of the TUs of the CU may or may not be based on the size and position of the prediction blocks of the PUs of the CU. A quadtree structure known as a "residual quadtree" (RQT) may include nodes associated with each of the regions. The TUs of the CU may correspond to leaf nodes of the RQT. In other examples, the transform processing unit 104 may be configured to partition the TUs according to the MTT technique described above. For example, the video encoder 22 may not further partition the CU into TUs using the RQT structure. Thus, in one example, the CU includes a single TU.

[0156] The transform processing unit 104 may generate a transform coefficient block for each TU of the CU by applying one or more transforms to the transform block of the TU. The transform processing unit 104 may apply various transforms to the transform block associated with the TU. For example, the transform processing unit 104 may apply a discrete cosine transform (DCT), a directional transform, or a conceptually similar transform to the transform block. In some examples, the transform processing unit 104 does not apply a transform to the transform block. In such examples, the transform block may be treated as a transform coefficient block.

[0157] The quantization unit 106 may quantize the transform coefficients in the coefficient block. The quantization process may reduce the bit depth associated with some or all transform coefficients. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during the quantization process, where n is greater than m. The quantization unit 106 may quantize the coefficient block associated with the TU of the CU based on a quantization parameter (QP) value associated with the CU. The video encoder 22 may adjust the quantization applied to the coefficient block associated with the CU by adjusting the QP value associated with the CU. Quantization may introduce information loss. Therefore, the quantized transform coefficient may have lower precision than the original coefficient.

[0158] Inverse quantization unit 108 and inverse transform processing unit 110 may apply inverse quantization and inverse transform, respectively, to the coefficient block to reconstruct a residual block from the coefficient block. Reconstruction unit 112 may add the reconstructed residual block to corresponding samples of one or more predictive blocks generated by prediction processing unit 100 to produce a reconstructed transform block associated with the TU. By reconstructing the transform block for each TU of a CU in this manner, video encoder 22 may reconstruct the coding block of the CU.

[0159] The filter unit 114 may perform one or more deblocking operations to reduce blocking artifacts in the coding blocks associated with the CU. The decoded picture buffer 116 may store the reconstructed coding blocks after the filter unit 114 performs one or more deblocking operations on the reconstructed coding blocks. The inter-prediction processing unit 120 may use a reference picture containing the reconstructed coding blocks to perform inter-prediction on PUs of other pictures. In addition, the intra-prediction processing unit 126 may use the reconstructed coding blocks in the decoded picture buffer 116 to perform intra-prediction on other PUs in the same picture as the CU.

[0160] The entropy coding unit 118 may receive data from other functional components of the video encoder 22. For example, the entropy coding unit 118 may receive coefficient blocks from the quantization unit 106 and may receive syntax elements from the prediction processing unit 100. The entropy coding unit 118 may perform one or more entropy coding operations on the data to generate entropy-coded data. For example, the entropy coding unit 118 may perform a CABAC operation, a context-adaptive variable length coding (CAVLC) operation, a variable-to-variable (V2V) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioning entropy (PIPE) coding operation, an exponential Golomb coding operation, or another type of entropy coding operation on the data. The video encoder 22 may output a bitstream including the entropy-coded data generated by the entropy coding unit 118. For example, according to the technology of the present disclosure, the bitstream may include data representing the partition structure of the CU.

[0161] Fig. 9 is a block diagram illustrating an example video decoder 30 configured to implement the techniques of this disclosure. Fig. 9It is provided for the purpose of explanation and is not a limitation of the techniques widely exemplified and described in this disclosure. For the purpose of explanation, this disclosure describes the video decoder 30 in the context of HEVC encoding. However, the techniques of this disclosure may be applicable to other encoding standards or methods.

[0162] exist Fig. 9 In the example of , video decoder 30 includes entropy decoding unit 150, video data memory 151, prediction processing unit 152, inverse quantization unit 154, inverse transform processing unit 156, reconstruction unit 158, filter unit 160, and decoded picture buffer 162. Prediction processing unit 152 includes motion compensation unit 164 and intra prediction processing unit 166. In other examples, video decoder 30 may include more, fewer, or different functional components.

[0163] The video data memory 151 may store encoded video data, such as an encoded video bitstream, for decoding by components of the video decoder 30. The video data stored in the video data memory 151 may be obtained, for example, from a computer-readable medium 16, for example, from a local video source (such as a camera), via a wired or wireless network communication of video data, or by accessing a physical data storage medium. The video data memory 151 may form a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. The decoded picture buffer 162 may be a reference picture memory that stores reference video data for use by the video decoder 30 in decoding video data, for example, in intra-frame coding or inter-frame coding mode, or for output. The video data memory 151 and the decoded picture buffer 162 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 151 and the decoded picture buffer 162 may be provided with the same memory device or separate memory devices. In various examples, video data memory 151 may be on-chip with other components of video decoder 30, or off-chip relative to those components. Figure 1 The same as or part of the storage medium 28 in.

[0164] The video data memory 151 receives and stores the encoded video data (e.g., NAL units) of the bitstream. The entropy decoding unit 150 may receive the encoded video data (e.g., NAL units) from the video data memory 151, and may parse the NAL units to obtain syntax elements. The entropy decoding unit 150 may entropy decode the entropy-encoded syntax elements in the NAL units. The prediction processing unit 152, the inverse quantization unit 154, the inverse transform processing unit 156, the reconstruction unit 158, and the filter unit 160 may generate decoded video data based on the syntax elements extracted from the bitstream. The entropy decoding unit 150 may perform a process that is generally inverse to the process of the entropy encoding unit 118.

[0165] According to some examples of the present disclosure, the entropy decoding unit 150 or another processing unit of the video decoder 30 may determine a tree structure as part of obtaining syntax elements from a bitstream. The tree structure may specify how to partition an initial video block (such as a CTB) into smaller video blocks (such as coding units). According to one or more techniques of the present disclosure, for each corresponding non-leaf node of the tree structure at each depth level of the tree structure, there are multiple allowed partitioning types for the corresponding non-leaf node, and according to one of the multiple allowed partitioning modes, the video block corresponding to the corresponding non-leaf node is partitioned into video blocks corresponding to child nodes of the corresponding non-leaf node. In other embodiments, the residual block corresponding to the CU leaf node may be partitioned according to the determined tree structure.

[0166] In addition to obtaining syntax elements from the bitstream, the video decoder 30 can also perform a reconstruction operation on a non-partitioned CU. In order to perform a reconstruction operation on a CU, the video decoder 30 can perform a reconstruction operation on each TU of the CU. By performing a reconstruction operation on each TU of the CU, the video decoder 30 can reconstruct the residual block of the CU. As described above, in one example of the present disclosure, the CU includes a single TU.

[0167] As part of performing a reconstruction operation on a TU of a CU, the inverse quantization unit 154 may inverse quantize, i.e., dequantize, a coefficient block associated with the TU. After the inverse quantization unit 154 inverse quantizes the coefficient block, the inverse transform processing unit 156 may apply one or more inverse transforms to the coefficient block to generate a residual block associated with the TU. For example, the inverse transform processing unit 156 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse directional transform, or another inverse transform to the coefficient block.

[0168] If the CU or PU is encoded using intra prediction, intra prediction processing unit 166 may perform intra prediction to generate predictive blocks for the PU. Intra prediction processing unit 166 may use an intra prediction mode based on sample spatial neighboring blocks to generate predictive blocks for the PU. Intra prediction processing unit 166 may determine the intra prediction mode for the PU based on one or more syntax elements obtained from the bitstream.

[0169] If the PU is encoded using inter prediction, the entropy decoding unit 150 may determine the motion information of the PU. The motion compensation unit 164 may determine one or more reference blocks based on the motion information of the PU. The motion compensation unit 164 may generate a predictive block (e.g., a predictive luma, Cb, and Cr block) of the PU based on the one or more reference blocks. As described above, in one example of the present disclosure using MTT partitioning, the CU may include only a single PU. That is, the CU may not be split into multiple PUs.

[0170] The reconstruction unit 158 ​​may reconstruct the coding blocks (e.g., luma, Cb, and Cr coding blocks) of the CU using the transform blocks (e.g., luma, Cb, and Cr transform blocks) of the TUs of the CU and the predictive blocks (e.g., luma, Cb, and Cr blocks) of the PUs of the CU, i.e., the intra prediction data or the inter prediction data (if applicable). For example, the reconstruction unit 158 ​​may add samples of the transform blocks (e.g., luma, Cb, and Cr transform blocks) to corresponding samples of the predictive blocks (e.g., luma, Cb, and Cr predictive blocks) to reconstruct the coding blocks (e.g., luma, Cb, and Cr coding blocks) of the CU.

[0171] Filter unit 160 may perform a deblocking operation to reduce blocking artifacts associated with the coding blocks of the CU. Video decoder 30 may store the coding blocks of the CU in a decoded picture buffer 162. Decoded picture buffer 162 may be used for subsequent motion compensation, intra-frame prediction, and display on a display device such as Figure 1 For example, the video decoder 30 may perform intra prediction or inter prediction operations for PUs of other CUs based on blocks in the decoded picture buffer 162.

[0172] Fig. 10A is a flowchart illustrating example operation of video encoder 22 according to techniques of this disclosure. Fig. 10AIn an example of , video encoder 22 may generate an encoded representation of an initial video block (e.g., a coding tree block) of video data (200). As part of generating the encoded representation of the initial video block, video encoder 22 determines a tree structure including a plurality of nodes. The plurality of nodes include a plurality of leaf nodes and a plurality of non-leaf nodes. The leaf nodes have no child nodes in the tree structure. The non-leaf nodes include a root node of the tree structure. The root node corresponds to the initial video block. For each corresponding non-root node of the plurality of nodes, the corresponding non-root node corresponds to a video block (e.g., a coding block) that is a child block of a video block of a parent node in the tree structure corresponding to the corresponding non-root node. Each corresponding non-leaf node of the plurality of non-leaf nodes has one or more child nodes in the tree structure. For each corresponding non-leaf node of the tree structure at each depth level of the tree structure, there are a plurality of allowed partition types of three or more partition structures (e.g., BT, QT, and TT partition structures) for the corresponding non-leaf node, and according to one of the plurality of allowed partition types, the video block corresponding to the corresponding non-leaf node is partitioned into video blocks corresponding to the child nodes of the corresponding non-leaf node. Each respective allowed partition type of the plurality of allowed partition types may correspond to a different way of partitioning a video block corresponding to a respective non-leaf node into video blocks corresponding to child nodes of the respective non-leaf node. In addition, in the present example, video encoder 22 may include an encoded representation of the initial video block in a bitstream that includes an encoded representation of the video data (202).

[0173] Fig. 10B is a flowchart illustrating an example operation of video decoder 30 according to techniques of this disclosure. Fig. 10BIn an example of , the video decoder 30 may determine a tree structure (250) including a plurality of nodes. The plurality of nodes include a plurality of leaf nodes and a plurality of non-leaf nodes. The leaf node has no child nodes in the tree structure. The non-leaf nodes include a root node of the tree structure. The root node corresponds to an initial video block of the video data. For each corresponding non-root node among the plurality of nodes, the corresponding non-root node corresponds to a video block, which is a child block of a video block of a parent node in the tree structure corresponding to the corresponding non-root node. Each corresponding non-leaf node of the plurality of non-leaf nodes has one or more child nodes in the tree structure. For each corresponding non-leaf node of the tree structure at each depth level of the tree structure, there are a plurality of allowed segmentation types of three or more segmentation structures (e.g., BT, QT, and TT segmentation structures) for the corresponding non-leaf node, and according to one of the plurality of allowed segmentation types, the video block corresponding to the corresponding non-leaf node is segmented into video blocks corresponding to child nodes of the corresponding non-leaf node. Each corresponding allowed segmentation type among the plurality of allowed segmentation types corresponds to a different way of segmenting the video block corresponding to the corresponding non-leaf node into video blocks corresponding to child nodes of the corresponding non-leaf node. Furthermore, in the present example, for each (or at least one) corresponding leaf node of the tree structure, video decoder 30 reconstructs a video block corresponding to the corresponding leaf node ( 252 ).

[0174] exist Fig. 10A and Fig. 10B In an example, for each corresponding non-leaf node other than the root node in the tree structure, a plurality of allowed partition types for the corresponding non-leaf node may be independent of a partitioning mode according to which a video block corresponding to a parent node of the corresponding non-leaf node is partitioned into video blocks of child nodes corresponding to the parent node of the corresponding non-leaf node. For example, unlike VCEG proposal COM16-C966, if a video block of a specific node is partitioned according to a binary tree partitioning mode, video blocks of child nodes of the specific node may be partitioned according to a quadtree partitioning mode.

[0175] In addition, Fig. 10A and Fig. 10B In the example, for each corresponding non-leaf node of the tree structure, the multiple allowed partitioning modes for the corresponding non-leaf node may include two or more of the following: a square quadtree partitioning mode, a rectangular quadtree partitioning mode, a symmetric binary tree partitioning mode, an asymmetric binary tree partitioning mode, a symmetric ternary tree partitioning mode, or an asymmetric ternary tree partitioning mode.

[0176] In addition, as shown above, only a subset of the above-mentioned partition types are used. The subset of supported partition types can be signaled or predefined in the bitstream. Therefore, in some examples, the video decoder 30 can obtain syntax elements indicating multiple supported partition modes from the bitstream. Similarly, the video encoder 22 can signal multiple supported partition modes in the bitstream. In these examples, for each corresponding non-leaf node of the tree structure, the multiple supported partition modes may include multiple allowed partition modes for the corresponding non-leaf node. In these examples, the syntax elements indicating multiple supported partition modes can be obtained from the bitstream (and signaled therein), such as in a sequence parameter set (SPS) or a picture parameter set (PPS) or a slice header.

[0177] As shown above, in some examples, when the subtree is divided into an asymmetric ternary tree, a restriction that two of the three partitions have the same size is applied. Therefore, in some examples, video decoder 30 can receive an encoded representation of an initial video block that complies with the restriction, which specifies that when the video block corresponding to the node of the tree structure is partitioned according to the asymmetric ternary tree pattern, the video blocks corresponding to the two child nodes of the node have the same size. Similarly, video encoder 22 can generate an encoded representation of the initial video block to comply with the restriction, which specifies that when the video block corresponding to the node of the tree structure is partitioned according to the asymmetric ternary tree pattern, the video blocks corresponding to the two child nodes of the node have the same size.

[0178] As shown above, in some examples, the number of supported partitioning types can be fixed for all depths in all CTUs. For example, among multiple allowed partitioning modes, the same number of allowed partitioning modes can exist for each non-leaf node of the tree structure. In addition, as shown above, in other examples, the number of supported partitioning types can depend on the depth, slice type, CTU type, or other previously encoded information. For example, for at least one non-leaf node in the tree structure, the number of allowed partitioning modes among multiple allowed partitioning modes for the non-leaf node depends on at least one of the following: the depth of the non-leaf node in the tree structure, the size of the video block corresponding to the non-leaf node in the tree structure, the slice type, or previously encoded information.

[0179] In some examples, when using an asymmetric split type (e.g., Figure 3When a block is partitioned using the asymmetric binary tree partition types shown in , including PART_2NxnU, PART 2NxnD, PART_nLx2N, PART_nRx2N), the largest subblock partitioned from the current block cannot be further partitioned. For example, a constraint on how the initial video block is encoded may require that when a video block corresponding to any non-leaf node of the tree structure is partitioned into multiple subblocks according to an asymmetric partitioning pattern, the largest subblock of the multiple subblocks corresponds to a leaf node of the tree structure.

[0180] In some examples, when a block is partitioned using an asymmetric partitioning type, the largest subblock partitioned from the current block cannot be further partitioned in the same direction. For example, a constraint on how the initial video block is encoded may require that when a video block corresponding to any non-leaf node of a tree structure is partitioned into a plurality of subblocks in a first direction according to an asymmetric partitioning pattern, the tree structure cannot contain a node corresponding to a largest subblock of the plurality of subblocks that is partitioned from the largest subblock of the plurality of subblocks in the first direction.

[0181] In some examples, no further horizontal / vertical splitting is allowed when the width / height of the block is not a power of 2. For example, a constraint on how the initial video block is encoded may require that nodes of the tree structure corresponding to video blocks whose height or width is not a power of 2 be leaf nodes.

[0182] In some examples, at each depth, an index of the selected partition type is signaled in the bitstream. Thus, in some examples, video encoder 22 may include in the bitstream an index of a partitioning pattern according to which video blocks corresponding to non-leaf nodes of the tree structure are partitioned into video blocks corresponding to child nodes of the non-leaf node. Similarly, in some examples, video decoder 30 may obtain from the bitstream an index of a partitioning pattern according to which video blocks corresponding to non-leaf nodes of the tree structure are partitioned into video blocks corresponding to child nodes of the non-leaf node.

[0183] In some examples, for each CU (leaf node), a 1-bit transform_partition flag is further signaled to indicate whether the transform is performed with the same size of the CU. If the transform_partition flag is signaled as true, the residual of the CU is further divided into multiple sub-blocks, and each sub-block is transformed. Therefore, in one example, for at least one leaf node of the tree structure, the video encoder 22 can include a syntax element in the bitstream. In this example, the syntax element with a first value indicates that a transform with the same size as the video block corresponding to the leaf node is applied to the residual data of the video block corresponding to the leaf node; the syntax element with a second value indicates that multiple transforms with a smaller size than the video block corresponding to the leaf node are applied to the sub-blocks of the residual data of the video block corresponding to the leaf node. In a similar example, for at least one leaf node of the tree structure, the video decoder 30 can obtain the syntax element from the bitstream.

[0184] In some examples, when a residual exists in a CU, a partition flag is not signaled, and only a transform with one derived size is used. For example, for at least one leaf node of the tree structure, the video encoder 22 may apply the same transform (e.g., discrete cosine transform, discrete sine transform, etc.) to different parts of the residual data corresponding to the video block corresponding to the leaf node to convert the residual data from the sample domain to the transform domain. In the sample domain, the residual data is represented by the value of the sample (e.g., the component of the pixel). In the transform domain, the residual data can be represented by frequency coefficients. Similarly, for at least one leaf node of the tree structure, the video decoder 30 may apply the same transform (i.e., inverse discrete cosine transform, inverse sine transform, etc.) to different parts of the residual data corresponding to the video block corresponding to the leaf node to convert the residual data from the transform domain to the sample domain.

[0185] In some examples, for each CU, if the CU is not partitioned into a square quadtree or a symmetric binary tree, the transform size is always set to be equal to the size of the partition size. For example, for each corresponding non-leaf node of the tree structure corresponding to the video block partitioned according to the square quadtree partitioning pattern or the symmetric binary tree partitioning pattern, the transform size of the transform applied to the residual data of the video block corresponding to the child node of the corresponding non-leaf node is always set to be equal to the size of the video block corresponding to the child node of the corresponding non-leaf node.

[0186] Fig.11 is a flowchart illustrating an example operation of a video encoder according to another example technique of the present disclosure. One or more structural elements of the video encoder 22, including the prediction processing unit 100, may be configured to perform Fig.11 technology.

[0187] In one example of the present disclosure, the video encoder 22 may be configured to obtain a picture of video data. At block 300, the video encoder 22 determines that the coded picture of the video data is divided into a plurality of coding units according to a first tree structure including leaf nodes. The first tree structure may be determined using a search based on a rate-distortion metric. Next, at block 302, the encoder 22 determines that the residual block of the leaf node is recursively divided into a plurality of transform units according to a second tree structure. The determination of dividing the residual block may also be based on a search using a rate-distortion metric. At block 304, the video encoder 22 encodes the picture based on the first tree structure and the second tree structure. Not shown, the video encoder 22 also encodes each of the other coding units using a similar technique.

[0188] Fig.12 is a flowchart illustrating an example operation of a video decoder according to another example technique of the present disclosure. One or more structural elements of the video decoder 30, including the entropy decoding unit 150 and / or the prediction processing unit 152, may be configured to perform Fig.12 technology.

[0189] In one example of the present disclosure, the video decoder 30 is configured to receive an encoded video bitstream that forms a representation of an encoded picture of video data. At block 400, the video decoder 30 determines to partition the encoded picture of the video data into a plurality of coding units according to a first tree structure including leaf nodes. At block 402, the video decoder determines to recursively partition the residual block of the leaf node into a plurality of transform units according to a second tree structure. Moving to block 404, the video decoder 30 reconstructs the encoded picture based on the determined first tree structure and the second tree structure. The video decoder 30 can apply similar techniques to reconstruct each of the coding units of the picture.

[0190] For purposes of illustration, some aspects of the present disclosure have been described with respect to extensions to the HEVC standard. However, the techniques described in this disclosure may be useful for other video encoding processes, including other standards or yet-to-be-developed proprietary video encoding processes, including VVC.

[0191] As described in the present disclosure, a video encoder may refer to a video encoder or a video decoder. Similarly, a video encoding unit may refer to a video encoder or a video decoder. Likewise, if applicable, video encoding may refer to video encoding or video decoding. In the present disclosure, the phrase "based on" may indicate based only on, at least partially on, or based on in some way. The present disclosure may use the term "video unit" or "video block" or "block" to refer to one or more sample blocks and a syntax structure for encoding samples of one or more sample blocks. Example types of video units may include CTU, CU, PU, ​​transform unit (TU), macroblock, macroblock partition, etc. In some contexts, the discussion of PU may be interchangeable with the discussion of macroblock or macroblock partition. Example types of video blocks may include coding tree blocks, coding blocks, and other types of video data blocks.

[0192] It should be appreciated that, depending on the examples, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the techniques). Additionally, in some examples, actions or events may be performed concurrently, e.g., through multithreading, interrupt handling, or multiple processors, rather than sequentially.

[0193] In one or more examples, the functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media, such as data storage media, or communication media including any media that facilitates (e.g., according to a communication protocol) the transfer of a computer program from one place to another. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) communication media, such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.

[0194] As an example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the required program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection is appropriately referred to as a computer-readable medium. For example, if the instruction is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals or other transient media, but are directed to non-transient tangible storage media. The disks and discs used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and blue-ray discs, where disks usually copy data magnetically, while discs use lasers to copy data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0195] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein may refer to any of the above structures or any other structure suitable for the implementation of the techniques described herein. In addition, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Moreover, these techniques may be implemented entirely in one or more circuits or logic elements.

[0196] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperating hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.

[0197] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for decoding video data, the method comprising: receiving an encoded video bitstream, the encoded video bitstream forming a representation of an encoded picture of the video data; Determine to partition a coding picture of the video data into a plurality of coding units, wherein the partitioning is performed according to a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure; decoding a syntax element from the coded video bitstream, the syntax element indicating that only a single transform unit partition structure for the residual block of the leaf node is allowed in the second tree structure; Determining that the residual block of the leaf node is recursively divided into a plurality of transform units according to the second tree structure, wherein the residual block is recursively divided into the second tree structure according to at least one of a quadtree, a binary tree, an asymmetric binary tree, or a ternary tree; as well as The coded image is reconstructed based on the determined first tree structure and second tree structure.

2. The method of claim 1, further comprising decoding, from the coded video bitstream, a syntax element associated with a root of a transform unit tree, the syntax element indicating that a plurality of transform units in the tree are transformed according to a specified transform. 3 . The method of claim 2 , wherein the syntax element indicates that the plurality of TUs share a parameter, the parameter being associated with one of enhanced multiple transform (EMT), non-separable secondary transform (NSST), or transform skip (TS).

4. The method of claim 1, further comprising decoding a syntax element associated with a root of a transform unit tree from the encoded video bitstream, the syntax element indicating at least one of a coded block flag or a last position of a non-zero coefficient of the transform unit tree.

5. A method for encoding video data, the method comprising: obtaining a picture for encoding into a coded video bitstream; Determine to partition a picture of the video data into a plurality of coding units, wherein the partitioning is performed according to a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure; Determining that the residual block of the leaf node is recursively divided into a plurality of transform units according to a second tree structure, wherein the residual block is recursively divided into the second tree structure according to at least one of a quadtree, a binary tree, an asymmetric binary tree, or a ternary tree; as well as The picture is encoded into the encoded video bitstream based on the determined first tree structure and second tree structure, which includes encoding a syntax element into the encoded video bitstream, the syntax element indicating a single transform unit partition structure that only allows residual blocks for the leaf nodes in the second tree structure. 6 . The method of claim 5 , further comprising encoding a syntax element into the coded video bitstream, the syntax element indicating that only a single transform unit partition structure for the residual block is allowed in the tree.

7. The method of claim 6, further comprising encoding a syntax element associated with a root of a transform unit tree into the coded video bitstream, the syntax element indicating that a plurality of transform units in the tree are transformed according to a specified transform.

8. The method of claim 7, wherein the syntax element indicates that the plurality of TUs share a parameter associated with one of enhanced multiple transform (EMT), non-separable secondary transform (NSST), or transform skip (TS).

9. The method of claim 5, further comprising encoding a syntax element associated with a root of a transform unit tree into the encoded video bitstream, the syntax element indicating at least one of a coding block flag or a last position of a non-zero coefficient of the transform unit tree.

10. An apparatus for decoding video data, the apparatus comprising: A memory configured to store encoded pictures of the video data; as well as The video processor is configured as: receiving an encoded video bitstream, the encoded video bitstream forming a representation of an encoded picture of the video data; Determine to partition a coding picture of the video data into a plurality of coding units, wherein the partitioning is performed according to a first tree structure, and the plurality of coding units include leaf nodes of the first tree structure; decoding a syntax element from the coded video bitstream, the syntax element indicating that only a single transform unit partition structure for the residual block of the leaf node is allowed in the second tree structure; Determining that the residual block of the leaf node is recursively divided into a plurality of transform units according to the second tree structure, wherein the residual block is recursively divided into the second tree structure according to at least one of a quadtree, a binary tree, an asymmetric binary tree, or a ternary tree; as well as The coded image is reconstructed based on the determined first tree structure and second tree structure.

11. The apparatus of claim 10, wherein the video processor is further configured to decode a syntax element from the encoded video bitstream, the syntax element indicating that only a single transform unit partition structure for the residual block is allowed in the tree.