Video encoder, video decoder, and corresponding method
The method addresses the challenge of further reducing bitrate in video coding by using an override flag to extract and apply image region-specific partition constraints during decoding, enhancing parsing efficiency and encoding flexibility.
Patent Information
- Application Number
- JP2023107167
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-05
- Filing Date
- 2023-06-29
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2039-09-18
AI Technical Summary
Current video coding standards, such as HEVC, while efficient, still face challenges in further reducing bitrate without compromising picture quality, particularly in handling complex image regions.
The proposed method involves a decoding apparatus that extracts an override flag from the video bitstream, using it to obtain first partition constraint information from the image region header, and partitions blocks of the image region accordingly. This approach allows for flexible block partitioning based on image region-specific constraints.
This method enables efficient bitstream parsing and partition constraint signaling, allowing for more flexible and efficient video encoding and decoding, particularly in handling complex image regions.
Smart Images

Figure 0007689993000026 
Figure 0007689993000027 
Figure 0007689993000028
Abstract
Description
Technical Field
[0001] Embodiments of the present application generally relate to the field of video encoding, and more particularly to block splitting and partitioning.
Background Art
[0002] Video encoding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital television, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat, video conferencing, DVDs and Blu-ray discs, video content acquisition and editing systems, and camcorders for security applications.
[0003] Since the development of the block-based hybrid video coding approach in the H.261 standard in 1990, new video coding techniques and tools have been developed, forming the basis of new video coding standards. Further video coding standards include MPEG-1 video, MPEG-2 video, ITU-T H.262 / MPEG-2, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile video coding (VVC) and extensions, for example including scalability and / or three-dimensional (3D) extensions of these standards. As video generation and usage become increasingly ubiquitous, video traffic becomes the largest load in communication networks and data storage, and correspondingly, one of the goals of most video coding standards has been to achieve a reduction in bitrate compared to its predecessors without sacrificing picture quality. Even though the latest High Efficiency Video Coding (HEVC) can compress about twice as much video as AVC without sacrificing quality, there is a desire for new technologies to compress video even further compared to HEVC.
Summary of the Invention
[0004] Embodiments of the present application (or the present disclosure) provide apparatuses and methods for encoding and decoding according to the independent claims.
[0005] The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementations are apparent from the dependent claims, the description, and the drawings.
[0006] From the standard, the definitions of some features are as follows. Picture Parameter Set (PPS): A syntax structure that contains syntax elements applicable to zero or more coded pictures, as determined by the syntax elements found in each slice header. Sequence Parameter Set (SPS): A syntax structure that contains syntax elements applicable to zero or more CVSs, as determined by the content of the syntax elements found in the PPS, which are referred to by the syntax elements found in each slice header. Slice Header: A part of an encoded slice that contains data elements related to the first or all of the blocks represented within the slice. Subpicture: A rectangular region of one or more slices within a picture. A slice is composed of either the number of complete tiles or only a consecutive sequence of complete blocks of one tile. Tile: A rectangular region of CTUs within a specific tile column and a specific tile row within a picture.
[0007] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular region of the picture. A tile is divided into one or more blocks, each of which is composed of a number of CTU columns within the tile. A tile that is not partitioned into multiple blocks is also called a block. However, a block that is a proper subset of a tile is not called a tile. A slice contains either a number of tiles of a picture or a number of blocks of a tile. A subpicture contains one or more slices that collectively cover a rectangular region of the picture.
[0008] Two modes of slicing are supported, namely the raster scan slice mode and the rectangular slice mode. In the raster scan slice mode, a slice contains a sequence of tiles within the tile raster scan of a picture. In the rectangular slice mode, a slice contains a number of bricks of the picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are in the brick raster scan order of the slice.
[0009] According to a first aspect of the present invention, there is provided a method for decoding a video bit stream implemented by a decoding apparatus, wherein the video bit stream includes data representing an image region and an image region header of the image region, and the decoding method comprises: obtaining an override flag from the video bit stream; when the value of the override flag is an override value, obtaining first partition constraint information of the image region from the image region header; partitioning blocks of the image region according to the first partition constraint information.
[0010] This approach allows each image region to have its own partition constraint information in addition to the partition constraint information of a plurality of image regions in a parameter set. Thus, this approach enables efficient bitstream parsing and, in particular, efficient partition constraint information signaling.
[0011] The step of obtaining the first partition constraint information of the image region from the image region header may include obtaining the first partition constraint information of the image region from the data representing the image region header.
[0012] The override value may be preset.
[0013] The override value includes true, false, 0, or 1.
[0014] The image area header may be a set or structure including the data elements regarding all or part of the image area.
[0015] In a possible implementation form of the method according to the first aspect, therefore, the decoding method further includes a step of obtaining an override enabling flag from the video bitstream, where the value of the override enabling flag is an enabling value.
[0016] The enabling value may be preset.
[0017] The enabling value includes true, false, 0, or 1.
[0018] In a possible implementation form of the method according to the first aspect, therefore, the decoding method further includes a step of obtaining an override enabling flag from the video bitstream, and the step of obtaining the override flag from the video bitstream includes a step of obtaining the override flag from the video bitstream when the value of the override enabling flag is an enabling value.
[0019] The enabling value may be preset.
[0020] The enabling value includes true, false, 0, or 1.
[0021] By providing an override enabling flag, the override may be controlled in an efficient manner, thus increasing flexibility in handling syntax elements related to block partitioning. Note that when the override enabling flag is set to the enabling value, the override flag may be further extracted from the bitstream. Alternatively, the override flag may not be extracted from the bitstream, in which case no override is applied. Instead, second or third partitioning constraints may be used to partition the block.
[0022] In any of the foregoing implementations of the first aspect or in a possible implementation form of the method according to the first aspect, thus, the video bitstream further includes data representing a parameter set of the video bitstream, and the decoding method When the value of the override enabling flag is a disabling value, further including the step of partitioning the block of the image region according to second partitioning constraint information of the video bitstream. The second partitioning constraint information may be from or within the parameter set.
[0023] The parameter set may be a sequence parameter set (SPS) or a picture parameter set (PPS) or any other parameter set.
[0024] The disabling value is different from the enabling value.
[0025] The disabling value may be preset.
[0026] The disabling value includes true, false, 0, or 1.
[0027] When the value of the override enable flag is the disable value, the first partition constraint information may not be present in the video bitstream, and the value of the first partition constraint information may be presumed to be equal to the value of the second partition constraint information.
[0028] The parameter set may be a set or structure including syntax elements applied to all zero or more coded pictures or a coded video sequence including the image region.
[0029] The parameter set is different from the image region header.
[0030] For example, the second partition constraint information may include information on the minimum allowable quadtree leaf node size, information on the maximum multi-type tree depth, information on the maximum allowable ternary tree root node size, or information on the maximum allowable binary tree root node size. Any combination / subset of these and further parameters may be signaled to configure the partition constraint.
[0031] The information on the minimum allowable quadtree leaf node size may be a delta value to obtain the value of the minimum allowable quadtree leaf node size. For example, the information on the minimum allowable quadtree leaf node size may be sps_log2_min_qt_size_intra_slices_minus2, sps_log2_min_qt_size_inter_slices_minus2, or log2_min_qt_size_minus2.
[0032] The information on the maximum allowable ternary tree root node size may be a delta value to obtain the value of the maximum allowable ternary tree root node size. For example, the information on the maximum allowable ternary tree root node size may be sps_log2_diff_ctu_max_tt_size_intra_slices, sps_log2_diff_ctu_max_tt_size_inter_slices, or log2_diff_ctu_max_tt_size.
[0033] The information on the maximum allowable binary tree root node size may be a delta value in order to obtain the value of the maximum allowable binary tree root node size. For example, the information on the maximum allowable binary tree root node size may be sps_log2_diff_ctu_max_bt_size_intra_slices, sps_log2_diff_ctu_max_bt_size_inter_slices, or log2_diff_ctu_max_bt_size.
[0034] For example, the information on the maximum multi-type tree depth may be sps_max_mtt_hierarchy_depth_inter_slices, sps_max_mtt_hierarchy_depth_intra_slices, or max_mtt_hierarchy_depth.
[0035] Additionally or alternatively, the second partitioning constraint information includes the partitioning constraint information of a block in the intra mode or the partitioning constraint information of a block in the inter mode.
[0036] The second partitioning constraint information may include both the partitioning constraint information of a block in the intra mode and the partitioning constraint information of a block in the inter mode, which are signaled separately. However, the present invention is not limited thereto, and there may be one partitioning constraint information common to both the partitioning constraint information of a block in the intra mode and the partitioning constraint information of a block in the inter mode.
[0037] The block in the intra mode or the block in the inter mode refers to the parameter set.
[0038] The parameter set may include a sequence parameter set (SPS) or a picture parameter set (PPS).
[0039] The block in the intra mode may be in a CTU within a slice having a slice_type equal to 2(I) that refers to the parameter set, or the block in the inter mode may be in a CTU within a slice having a slice_type equal to 0(B) or 1(P) that refers to the parameter set.
[0040] Additionally or alternatively, the second partitioning constraint information includes partitioning constraint information of a luma block and / or partitioning constraint information of a chroma block.
[0041] The luma block or the chroma block refers to the parameter set.
[0042] The parameter set may include a sequence parameter set (SPS) or a picture parameter set (PPS).
[0043] The luma block or the chroma block may be in a CTU within a slice that refers to the parameter set.
[0044] In any of the foregoing implementations of the first aspect or a possible implementation form of the method according to the first aspect, therefore, the video bitstream further includes data representing a parameter set of the video bitstream, and the step of obtaining an override enable flag from the video bitstream includes the step of obtaining the override enable flag from the parameter set or the step of obtaining the override enable flag in the parameter set.
[0045] The step of obtaining the override enable flag from the parameter set may include the step of obtaining the override enable flag from the data representing the parameter set. The parameter set may be a sequence parameter set (SPS), a picture parameter set (PPS), or any other parameter set.
[0046] In any of the foregoing implementations of the first aspect or in a possible implementation form of the method according to the first aspect, therefore, the step of obtaining the override flag from the video bitstream includes the step of obtaining the override flag from the image region header.
[0047] The step of obtaining the override flag from the image region header may include the step of obtaining the override flag from the data representing the image region header.
[0048] In any of the foregoing implementations of the first aspect or in a possible implementation form of the method according to the first aspect, therefore, the first partition constraint information includes information on the minimum allowable quadtree leaf node size, information on the maximum multi-type tree depth, information on the maximum allowable ternary tree root node size, or information on the maximum allowable binary tree root node size.
[0049] The information on the minimum allowable quadtree leaf node size may be a delta value in order to obtain the value of the minimum allowable quadtree leaf node size. For example, the information on the minimum allowable quadtree leaf node size may be sps_log2_min_qt_size_intra_slices_minus2, sps_log2_min_qt_size_inter_slices_minus2, or log2_min_qt_size_minus2.
[0050] The information on the maximum allowable ternary tree root node size may be a delta value for obtaining the value of the maximum allowable ternary tree root node size. For example, the information on the maximum allowable ternary tree root node size may be sps_log2_diff_ctu_max_tt_size_intra_slices, sps_log2_diff_ctu_max_tt_size_inter_slices, or log2_diff_ctu_max_tt_size.
[0051] The information on the maximum allowable binary tree root node size may be a delta value for obtaining the value of the maximum allowable binary tree root node size. For example, the information on the maximum allowable binary tree root node size may be sps_log2_diff_ctu_max_bt_size_intra_slices, sps_log2_diff_ctu_max_bt_size_inter_slices, or log2_diff_ctu_max_bt_size.
[0052] For example, the information on the maximum multi-type tree depth may be sps_max_mtt_hierarchy_depth_inter_slices, sps_max_mtt_hierarchy_depth_intra_slices, or max_mtt_hierarchy_depth.
[0053] For example, the image region may include a slice, a tile, or a sub-picture, and the image region header may include the slice header of the slice, the tile header of the tile, or the header of the sub-picture.
[0054] In any of the foregoing implementations of the first aspect or in a possible implementation form of the method according to the first aspect, therefore, the video bitstream may further include data representing a parameter set of the video bitstream, and the decoding method is When the value of the override flag is not the override value, step S230 of partitioning the block of the image region according to the second partition constraint information of the video bit stream from the parameter set, or step S230 of partitioning the block of the image region according to the second partition constraint information of the video bit stream in the parameter set, is further included.
[0055] The parameter set may be a sequence parameter set (SPS) or a picture parameter set (PPS) or any other parameter set.
[0056] When the override value is true, the fact that the value of the override flag is not the override value means that the value of the override flag is false.
[0057] When the override value is 1, the fact that the value of the override flag is not the override value means that the value of the override flag is 0.
[0058] According to a second aspect of the present invention, a method for encoding a video bit stream implemented by an encoding device, wherein the video bit stream includes data representing an image region and an image region header of the image region, and the encoding method includes: Determining whether the partition of the block of the image region follows the first partition constraint information in the image region header; When it is determined that the partition of the block follows the first partition constraint information, partitioning the block of the image region according to the first partition constraint information; Setting the value of the override flag to the override value; Including the data of the override flag in the video bit stream.
[0059] In a possible implementation of the method according to the second aspect, therefore, the encoding method determining whether the partitioning of the block according to the first partitioning constraint information is enabled; when it is determined that the partitioning of the block according to the first partitioning constraint information is enabled, setting the value of the override enable flag to an enabled value; including the data of the override enable flag in the video bitstream.
[0060] The step of determining whether the partitioning of the block in the image area conforms to the first partitioning constraint information in the image area header includes, when it is determined that the partitioning of the block according to the first partitioning constraint information is enabled, determining whether the partitioning of the block in the image area conforms to the first partitioning constraint information in the image area header.
[0061] For example, the video bitstream further includes data representing a parameter set of the video bitstream, and the encoding method when it is determined that the partitioning of the block according to the first partitioning constraint information is not enabled, partitioning the block in the image area according to the second partitioning constraint information of the video bitstream in the parameter set; setting the value of the override enable flag to a disabled value.
[0062] Additionally or alternatively, the second partitioning constraint information includes information on a minimum allowable quadtree leaf node size, information on a maximum multi-type tree depth, information on a maximum allowable ternary tree root node size, or information on a maximum allowable binary tree root node size.
[0063] Additionally or alternatively, the second partition constraint information includes partition constraint information of blocks in the intra mode or partition constraint information of blocks in the inter mode.
[0064] For example, the second partition constraint information includes partition constraint information of luma blocks or partition constraint information of chroma blocks.
[0065] In any of the foregoing implementations of the second aspect or a possible implementation form of the method according to the second aspect, therefore, the video bitstream further includes data representing a parameter set of the video bitstream, and the override enable flag is in the parameter set.
[0066] For example, the override flag is in the image region header.
[0067] Additionally or alternatively to any of the embodiments, the first partition constraint information includes information on the minimum allowable quadtree leaf node size, information on the maximum multi-type tree depth, information on the maximum allowable ternary tree root node size, or information on the maximum allowable binary tree root node size.
[0068] Additionally or alternatively to any of the embodiments, the image region includes a slice, a tile, or a subpicture, and the image region header includes a slice header of the slice, a tile header of the tile, or a header of the subpicture.
[0069] For example, the video bitstream further includes data representing a parameter set of the video bitstream, and the decoding method is when it is determined that the partition of the block does not conform to the first partition constraint information, partitioning the block of the image region according to the second partition constraint information of the video bitstream in the parameter set (step S360); further comprising the step of setting the value of the override flag to a non-override value.
[0070] The method according to the second aspect can be extended to an implementation format corresponding to the implementation format of the first device according to the first aspect. Therefore, the implementation format of the method includes the features of the corresponding implementation format of the first device.
[0071] The advantages of the method according to the second aspect are the same as those of the corresponding implementation format of the first device according to the first aspect.
[0072] According to a third aspect of the present invention, there is provided a decoder comprising: one or more processors; a non-transitory computer-readable storage medium connected to the processor and storing programming for execution by the processor, wherein the programming, when executed by the processor, configures the decoder to execute any of the above-described decoding methods according to the first aspect or any possible implementation of the first aspect.
[0073] According to a fourth aspect of the present invention, there is provided an encoder comprising: one or more processors; a non-transitory computer-readable storage medium connected to the processor and storing programming for execution by the processor, wherein the programming, when executed by the processor, configures the encoder to execute a method according to any of the above-described decoding methods according to the second aspect or any possible implementation of the second aspect.
[0074] According to a fifth aspect, there is proposed a computer-readable storage medium storing instructions that, when executed, cause one or more processors to be configured to encode video data. The instructions cause the one or more processors to execute a method according to the first or second aspect, or any possible implementation of the first or second aspect.
[0075] According to a sixth aspect, the present invention relates to a computer program including program code for executing, when executed on a computer, a method according to the first or second aspect or any possible embodiment of the first or second aspect.
[0076] According to a seventh aspect of the present invention, there is provided a decoder for decoding a video bit stream, the video bit stream including data representing an image region and an image region header of the image region, the decoder comprising: an override determination unit that obtains an override flag from the video bit stream; a partition constraint determination unit that obtains first partition constraint information of the image region from the image region header when the value of the override flag is an override value; a block partition unit that partitions blocks of the image region according to the first partition constraint information; and a decoder including the same is provided.
[0077] The method according to the first aspect of the present invention can be executed by the decoder according to the seventh aspect of the present invention. Further features and implementation forms of the decoder according to the third aspect of the present invention correspond to the features and implementation forms of the method according to the first aspect of the present invention or any possible implementation of the first aspect. According to an eighth aspect of the present invention, there is provided an encoder for encoding a video bit stream, the video bit stream including data representing an image region and an image region header of the image region, the encoder comprising: a block partition unit that partitions blocks of the image region according to first partition constraint information; a bit stream generator that inserts the first partition constraint information of the image region into the image region header, sets the value of the override flag to an override value, and inserts the override flag into the video bit stream; and an encoder including the same is provided.
[0078] The method according to the second aspect of the present invention can be executed by an encoder according to the eighth aspect of the present invention. Further features and implementation forms of the encoder according to the eighth aspect of the present invention correspond to the features and implementation forms of the method according to the second aspect of the present invention or any possible implementation of the second aspect.
[0079] For the purpose of clarity, any one of the embodiments disclosed herein may be combined with any one or more of the other embodiments to generate new embodiments within the scope of the present disclosure.
[0080] According to a ninth aspect of the present invention, there is provided a video bitstream, the video bitstream including data representing an image region and an image region header of the image region, the video bitstream further including an override flag specifying whether first partition constraint information of the image region is present in the image region header.
[0081] In a possible implementation form of the method according to the ninth aspect, therefore, the video bitstream further includes an override enable flag specifying whether the override flag is present in the image region header.
[0082] In any of the foregoing implementations of the first aspect or in a possible implementation form of the method according to the first aspect, therefore, the override enable flag is in the parameter set or data representing the parameter set.
[0083] In any of the foregoing implementations of the first aspect or in a possible implementation form of the method according to the first aspect, therefore, the override flag is in the image region header or data representing the image region header.
[0084] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0085] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings and figures.
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 15
Figure 16
Figure 17
Figure 18
[0086] Hereinafter, unless explicitly stated otherwise, the same reference numerals represent the same or at least functionally equivalent features.
Mode for Carrying Out the Invention
[0087] In the following description, reference is made to the accompanying drawings, which form a part of the present disclosure and illustrate specific aspects of embodiments of the present invention or specific aspects in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other aspects and may include structural or logical changes not shown in the figures. The following detailed description should, therefore, not be considered in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0088] For example, it will be understood that the disclosures related to the described method may well apply to a corresponding apparatus or system configured to perform the method, and vice versa. For example, when one or more specific method steps are described, even if such one or more units are not explicitly described or shown in the figures, the corresponding apparatus may include one or more units, such as functional units, for performing the one or more method steps described (e.g., one unit performs one or more steps, or each of the plurality of units performs one or more of the plurality of steps). On the other hand, for example, when a specific device is described based on one or more units, such as functional units, even if such one or more steps are not explicitly described or shown in the figures, the corresponding method may include one step for performing the functions of the one or more units (e.g., one step performs the functions of one or more units, or each of the plurality of steps performs one or more of the functions of the plurality of units). Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other unless otherwise specified.
[0089] Video coding typically represents the processing of a sequence of pictures that form a video or video sequence. Instead of the term "picture", the terms "frame" or "image" may be used synonymously in the field of video coding. The video coding used in this application (or this disclosure) refers to video coding or video decoding. Video coding is performed on the source side and typically involves processing the original video pictures (e.g., by compression) to reduce the amount of data required to represent the video pictures (for more efficient storage and / or transmission). Video decoding is performed on the destination side and typically involves performing the opposite process to the encoder to reconstruct the video pictures. Embodiments referring to the "coding" of a video picture (or generally a picture as described later) should be understood to be related to either the "coding" or "decoding" of a video sequence. The combination of the coding part and the decoding part is also called a codec (Coding and Decoding (CODEC)).
[0090] In the case of lossless video coding, the original video picture is reconstructible. That is, the reconstructed video picture has the same quality as the original video picture (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, for example, further compression by quantization is performed to reduce the amount of data representing the video picture. This cannot be fully reconstructed on the decoder side. That is, the quality of the reconstructed video picture is lower or worse than that of the original video picture.
[0091] Some video coding standards since H.261 belong to the group of "lossy hybrid video coders" (that is, combining spatial and temporal prediction in the sampled domain and 2D transform coding applying quantization in the transform domain). Each picture of the video sequence is typically partitioned into a set of non-overlapping blocks, and the coding is typically performed at the block level. In other words, in the encoder, for example, prediction blocks are generated using spatial (intrapicture) prediction and temporal (interpicture) prediction, the prediction blocks are subtracted from the current block (the block being currently processed / to be processed) to obtain a residual block, the residual block is transformed, and the residual block is quantized in the transform domain to reduce the amount of data to be transmitted (compressed), so that the video is typically processed, that is, coded, at the block (video block) level. On the other hand, in the decoder, the reverse process compared to the encoder is partially applied to the coded or compressed block to reconstruct the current block for presentation. Further, the encoder duplicates the decoder processing loop so that both generate the same prediction (e.g., intra and inter prediction) and / or reconstruction for processing, that is, coding subsequent blocks.
[0092] As used herein, the term "block" may be part of a picture or frame. For ease of explanation, embodiments of the present invention are described herein with reference to the reference software of High-Efficiency Video Coding (HEVC), or Versatile video coding (VVC) developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC. CU, PU, and TU may be referred to. In HEVC, a CTU is divided into CUs using a quadtree structure shown as a coding tree. The decision of whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into 1, 2, or 4 PUs according to the PU partition type. Within one PU, the same prediction process is applied, and relevant information is sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU. In the latest advancements in video compression technology, quadtree and binary tree (QTBT) partition frames are used to partition coding blocks. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree structure. The leaf nodes of the quadtree are further partitioned by a binary tree structure.The leaf nodes of the binary tree are called coding units (CUs), and are segmented for prediction and transformation processing without any further partitioning. This means that the CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In parallel, multi-partitioning, e.g., ternary tree partitioning, has also been proposed for use with the QTBT block structure. The term "apparatus" may be a "device", "decoder", or "encoder".
[0093] In the following embodiments of the encoder 20, the decoder 30 and the encoding system 10 are described based on FIGS. 1-3.
[0094] FIG. 1A is a conceptual or schematic block diagram showing an exemplary encoding system 10 that can utilize the technology of the present application (the present disclosure), e.g., a video encoding system 10. The encoder 20 (e.g., a video encoder 20) and the decoder 30 (e.g., a video decoder 30) of the video encoding system 10 represent examples of apparatuses that can be configured to execute techniques according to various examples described in the present application. As shown in FIG. 1A, the encoding system 10 includes a source device 12 configured to provide encoded data 13, e.g., an encoded picture 13, to a destination device 14 that decodes the encoded data 13.
[0095] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include a picture source 16, a preprocessing unit 18, e.g., a picture preprocessing unit 18, and a communication interface or communication unit 22.
[0096] The picture source 16 may include or be any kind of picture capture device that captures real pictures, for example, and / or any kind of picture or comment (in screen content encoding, considered as part of a picture or image for which any text on the screen should also be encoded) generation device, for example, a computer graphic processor that generates computer animation pictures, or any kind of device that acquires and / or provides real-world pictures, computer animation pictures (for example, screen content, virtual reality (VR) pictures) and / or any combination thereof (for example, augmented reality (AR) pictures). The picture source may be or be any kind of memory or storage device that stores any of the aforementioned pictures.
[0097] (Digital) pictures are or can be thought of as two-dimensional arrays or matrices of samples having intensity values. Samples in the array may also be called pixels (abbreviation of picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are typically used. That is, a picture may be represented by or may include three sample arrays. In the RGB format or color space, a picture includes corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance / chrominance format, or color space, such as YCbCr, which includes a luminance component represented by Y (sometimes L may be used instead) and two chrominance components represented by Cb and Cr. The luminance (or simply luma) component Y represents brightness or gray-level intensity (such as in a grayscale picture). On the other hand, the two chrominance (or simply chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in the YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted to or from the YCbCr format, and the process is also known as a color conversion or color transformation. If the picture is monochromatic, the picture may include only a luminance sample array.
[0098] The picture source 16 (e.g., video source 16) may be, for example, a camera that captures pictures, a memory that includes or stores previously captured or generated pictures, e.g., a picture memory, and / or any kind of (internal or external) interface for acquiring or receiving pictures. The camera may be, for example, a local or built-in camera integrated into the source device. The memory may be, for example, a local or built-in memory integrated into the source device. The interface may be, for example, an external interface for receiving pictures from an external video source, e.g., an external picture capture device such as a camera, an external memory, or an external picture generation device, e.g., an external computer graphics processor, a computer or a server. The interface may be any kind of interface, e.g., a wired or wireless interface, an optical interface, according to any characteristic or standardized interface protocol. The interface for acquiring the picture data 17 may be the same interface as or a part of the communication interface 22.
[0099] In contrast to the preprocessing unit 18 and the processing performed by the preprocessing unit 18, the picture or picture data 17 (e.g., video data 16) may also be referred to as raw picture or raw picture data 17.
[0100] The preprocessing unit 18 is configured to receive the (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain the preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessing unit 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optical component.
[0101] The encoder 20 (e.g., video encoder 20) is configured to receive the preprocessed picture data 19 and provide the encoded picture data 21 (further details will be described later, for example, based on FIG. 2 or FIG. 4).
[0102] The communication interface 22 of the source device 12 receives the encoded picture data 21 and transmits the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction, or stores the encoded data 13 for decoding or storage and / or processes the encoded picture data 21 respectively before transmitting the encoded data 13 to another device, such as the destination device 14 or any other device, and may be configured to do so.
[0103] The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and additionally, i.e., optionally, may include a communication interface or communication unit 28, a post-processing unit 32, and a display device 34.
[0104] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof) or the encoded data 13, e.g., directly from the source device 12 or from any other source, such as a storage device, e.g., an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.
[0105] The communication interfaces 22 and 28 are configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network, or any combination thereof, or any type of private or public network, or any combination of any type thereof.
[0106] The communication interface 22 may be configured to process the encoded picture data using any type of transmission encoding or processing, e.g., to package the encoded picture data 21 into a suitable format, such as a packet, and / or to transmit it via the communication link or communication network.
[0107] Communication interface 28 forms the counterpart of communication interface 22 and may be configured to, for example, receive the transmitted data, process the transmitted data using any kind of corresponding transmission decoding or processing, and / or unpack the encoded data 13 to obtain the encoded picture data 21.
[0108] Both communication interface 22 and communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface, as indicated by the arrow of the encoded picture data 13 pointing from the source device 12 to the destination device 14 in FIG. 1A. For example, to establish a connection, it may be configured to send and receive messages, for example, to positively respond to and exchange any other information related to the communication link and / or data transmission, for example, the encoded picture data transmission.
[0109] Decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (further details will be described later based on, for example, FIG. 3 or FIG. 5).
[0110] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, post-processed picture 33. The post-processing executed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31 for display by, for example, a display device 34.
[0111] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33, for example, to display a picture to a user or a viewer. The display device 34 may be or include any kind of display that presents a reconstructed picture, such as a built-in or external display or monitor. The display may include, for example, liquid crystal displays (LCDs), organic light emitting diodes (OLED) displays, plasma displays, projectors, micro LED displays, liquid crystal on silicon (LCoS), digital light processors (DLP), or any other kind of display.
[0112] FIG. 1A shows the source device 12 and the destination device 14 as separate devices, but embodiments of the devices may include both the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality, or both. In such embodiments, the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or separate hardware and / or software, or any combination thereof.
[0113] As will be apparent to those skilled in the art based on the description, the presence and (exact) division of the different units or functions within the source device 12 and / or destination device 14 as shown in FIG. 1A may vary depending on the actual device and application.
[0114] The encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) may each be implemented as any of a variety of suitable circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. Where the technology is implemented partially in software, the apparatus may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered to be one or more processors. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders and both may be integrated within their respective apparatuses as part of a combined encoder / decoder (CODEC).
[0115] Encoder 20 may be implemented by processing circuitry 46 to implement various modules as discussed with respect to encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented by processing circuitry 46 to implement various modules as discussed with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuitry may be configured to perform various operations as discussed later. As shown in FIG. 5, if the technology is implemented partially in software, the apparatus may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to execute the technology of the present disclosure. Either video encoder 20 or video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) within a single apparatus, for example, as shown in FIG. 1B.
[0116] Source device 12 may be referred to as a video encoding device or video encoding equipment. Destination device 14 may be referred to as a video decoding device or video decoding equipment. Source device 12 and destination device 14 may be examples of a video encoding device or video encoding equipment.
[0117] Source device 12 and destination device 14 may include any of a wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, mobile phone, smartphone, tablet or tablet computer, camera, desktop computer, set-top box, television, display device, digital media player, video game console, video streaming device (such as a content service server or content delivery server), broadcast receiving device, broadcast transmitting device, etc., and may use or not use any type of operating system.
[0118] In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Accordingly, source device 12 and destination device 14 may be wireless communication devices.
[0119] In some cases, the video encoding system 10 shown in FIG. 1A is merely an example, and the techniques of the present application may be applicable to video encoding settings (e.g., video encoding or video decoding) that do not necessarily involve any data communication between the encoding device and the decoding device. In other examples, the data may be read from local memory, streamed over a network, etc. The video encoding device may encode the data and store it in memory, and / or the video decoding device may read the data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but merely encode data to and / or read and decode data from memory.
[0120] For the sake of convenience of explanation, embodiments of the present invention are described herein with reference to, for example, the reference software of High-Efficiency Video Coding (HEVC), or Versatile Video coding (VVC), and the next-generation video coding standard developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0121] It should be understood that for each of the above examples described with reference to the video encoder 20, the video decoder 30 may be configured to perform reciprocal processing. With respect to signaling syntax elements, the video decoder 30 may be configured to receive and parse such syntax elements and accordingly decode associated video data. In some examples, the video encoder 20 may entropy encode one or more syntax elements into the encoded video bitstream. In such examples, the video decoder 30 may parse such syntax elements and accordingly decode associated video data.
[0122] FIG. 1B is an explanatory diagram of another exemplary video encoding system 40 including the encoder 20 of FIG. 2 and / or the decoder 30 of FIG. 3 according to an exemplary embodiment. The system 40 can implement techniques according to various examples described in the present application. In the illustrated implementation, the video encoding system 40 may include an image device 41, a video encoder 100, a video decoder 30 (and / or a video encoder implemented by the logic circuit 47 of the processing unit 46), an antenna 42, one or more processors 43, one or more memory stores 44, and / or a display device 45.
[0123] As illustrated, the image device 41, the antenna 42, the processing unit 46, the logic circuit 47, the video encoder 20, the video decoder 30, the processor 43, the memory store 44, and / or the display device 45 may be communicable with each other. Although both the video encoder 20 and the video decoder 30 are shown, the video encoding system 40 may include only the video encoder 20 or only the video decoder 30 in various examples.
[0124] As shown, in some examples, the video encoding system 40 may include an antenna 42. The antenna 42 may be configured to transmit or receive, for example, an encoded bitstream of video data. Further, in some examples, the video encoding system 40 may include a display device 45. The display device 45 may be configured to present video data. As shown, in some examples, the logic circuit 47 may be implemented by a processing unit 46. The processing unit 46 may include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. The video encoding system 40 may also include an optional processor 43 that may similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. In some examples, the logic circuit 47 may be implemented by hardware, video-encoding-specific hardware, and the like, and the processor 43 may be implemented by general-purpose software, an operating system, and the like. Further, the memory store 44 may be any type of memory such as volatile memory (e.g., Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). By way of non-limiting example, the memory store 44 may be implemented by cache memory. In some examples, the logic circuit 47 may access the memory store 44 (e.g., for implementation of an image buffer). In other examples, the logic circuit 47 and / or the processing unit 46 may include a memory store (e.g., a cache, etc.) for implementation of an image buffer or the like.
[0125] In some examples, the video encoder 100 implemented by a logic circuit may include an image buffer (e.g., by either the processing unit 46 or the memory store 44), and a graphics processing unit (e.g., by the processing unit 46). The graphics processing unit may be communicatively connected to the image buffer. The graphics processing unit may include a video encoder 100 implemented by the logic circuit 47 to implement various modules as discussed with respect to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuit may be configured to execute various operations as discussed herein.
[0126] The video decoder 30 may be implemented in a similar manner as implemented by the logic circuit 47 to implement various modules as discussed with respect to the decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, the video decoder 30 may be implemented via a logic circuit and may include an image buffer (e.g., by either the processing unit 420 or the memory store 44), and a graphics processing unit (e.g., by the processing unit 46). The graphics processing unit may be communicatively connected to the image buffer. The graphics processing unit may include a video decoder 30 implemented by the logic circuit 47 to implement various modules as discussed with respect to FIG. 3 and / or any other decoder system or subsystem described herein.
[0127] In some examples, the antenna 42 of the video encoding system 40 may be configured to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data related to an encoded partition (e.g., transform coefficients or quantized transform coefficients, optional indicators (as will be discussed), and / or data defining the encoded partition), data, indicators, index values, mode selection data, etc. related to the encoding of the video frames discussed herein. The video encoding system 40 may also include a video decoder 30 connected to the antenna 42 and configured to decode the encoded bitstream. The display device 45 is configured to present the video frames.
[0128] FIG. 2 shows a schematic / conceptual block diagram of an exemplary video encoder 20 configured to implement the technology of the present application. In the example of FIG. 2, the video encoder 20 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy encoding unit 270. The prediction processing unit 260 may include an inter prediction unit 244, an intra prediction processing unit 254, and a mode selection unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 as shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder compliant with a hybrid video codec.
[0129] For example, the residual calculation unit 204, the transformation processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy encoding unit 270 form the forward signal path of the encoder 20. On the other hand, for example, the inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, and the prediction processing unit 260 form the reverse signal path of the encoder, and the reverse signal path of the encoder corresponds to the signal path of the decoder (see the decoder 30 in FIG. 3).
[0130] The inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 are also represented as forming the "built-in decoder" of the video encoder 20.
[0131] The encoder 20 is configured to receive, for example, by the input 202, the picture 201 or the block 203 of the picture 201, for example, a sequence of pictures forming a video or a video sequence. The picture block 203 may also be referred to as the current picture block or the picture block to be encoded (especially in video encoding, to distinguish the current picture from other pictures, for example, pictures encoded and / or decoded before in the same video sequence, i.e., the video sequence including the current picture), and the picture 201 may also be referred to as the current picture or the picture to be encoded.
[0132] (Digital) pictures can be considered or thought of as two-dimensional arrays or matrices of samples having intensity values. Samples in the array may also be called pixels (short for picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are typically used. That is, a picture may be represented by or may include three sample arrays. In the RGB format or color space, a picture includes corresponding red, green, and blue sample arrays. However, in video encoding, each pixel is typically represented in YCbCr, which includes a luminance component indicated by Y (sometimes L is used instead) and two chrominance components indicated by Cb and Cr. The luminance (or simply luma) component Y represents brightness or gray-level intensity (such as in a grayscale picture). On the other hand, the two chrominance (or simply chroma) components Cb and Cr represent chrominance or color information components. Thus, a picture in YCbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in RGB format may be converted or transformed to YCbCr format, and vice versa, and the process is also known as color conversion or color transformation. If a picture is monochromatic, the picture may include only a luminance sample array. Thus, a picture may be, for example, an array of luma samples in monochromatic format or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0133] Partitioning
[0134] An embodiment of the encoder 20 may include a partitioning unit (not shown in FIG. 2) configured to partition a picture 201 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The partitioning unit may use the same block size for all pictures of a video sequence and the corresponding grid that defines the block size, or may vary the block size between pictures or subsets or groups of pictures and be configured to partition each picture into corresponding blocks.
[0135] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of the picture 201, such as one, some, or all of the blocks that form the picture 201. The picture blocks 203 may also be referred to as current picture blocks or picture blocks to be encoded.
[0136] In one example, the prediction processing unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described above.
[0137] Similar to picture 201, block 203 is also here a two-dimensional array or matrix of samples having intensity values (sample values) or can be considered as such, but of a smaller dimension than picture 201. In other words, block 203 may include, for example, one sample array (e.g., the luma array in the case of a monochromatic picture 201), or three sample arrays (e.g., the luma and two chroma arrays in the case of a color picture 201), or any other number and / or kind of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of transform coefficients.
[0138] An encoder 20 as shown in FIG. 2 is configured to encode picture 201 block by block. For example, encoding and prediction are performed for each block 203.
[0139] An embodiment of video encoder 20 as shown in FIG. 2 may be further configured to partition and / or encode a picture using slices (also called video slices). Here, a picture may be partitioned into or encoded using one or more slices (which do not standardly overlap), and each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., tiles (H.265 / HEVC and VVC) or bricks (VVC)).
[0140] An embodiment of the video encoder 20 as shown in FIG. 2 may be further configured to partition and / or encode pictures into slices / tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). Here, a picture may be partitioned into or encoded using one or more slices / tile groups (which do not overlap standardly), each slice / tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, each tile may be, for example, rectangular in shape, and may include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.
[0141] Residual calculation The residual calculation unit 204 is configured to calculate the residual block 205 by, for example, subtracting the sample values of the prediction block 265 from the sample values of the picture block 203 sample-by-sample (pixel-by-pixel) based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 will be provided later) to obtain the residual block 205 in the sample domain.
[0142] Transformation The transformation processing unit 206 is configured to apply a transformation, e.g., a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block 205 to obtain the transform coefficients 207 in the transform domain. The transform coefficients 207, also referred to as transform residual coefficients, may represent the residual block 205 in the transform domain.
[0143] The conversion processing unit 206 may be configured to apply an integer approximation of DCT / DST, such as the conversion specified for HEVC / H.265. Such an integer approximation is typically scaled by a specific factor compared to the orthogonal DCT transform. To maintain the norm of the residual block processed by the forward and inverse transforms, an additional scaling factor is applied as part of the conversion process. The scaling factor is typically selected based on specific constraints such as the scaling factor being a power of two for shift operations, the bit depth of the conversion coefficients, the trade-off between accuracy and implementation cost, etc. A specific scaling factor may be specified, for example, for the inverse transform by the inverse conversion processing unit 212 in the decoder 30 (and the corresponding inverse transform by the inverse conversion processing unit 212 in the encoder 20, for example), and the corresponding scaling factor for the forward transform by the conversion processing unit 206 in the encoder 20 may be correspondingly specified.
[0144] Embodiments of the video encoder 20 (each, the conversion processing unit 206) may be configured to output conversion parameters, such as the type of conversion or conversions, for example, encoded or compressed directly or by the entropy encoding unit 270. As a result, for example, the video decoder 30 may receive and use the conversion parameters for decoding.
[0145] Quantization The quantization unit 208 is configured to obtain the quantized transform coefficient 209 by quantizing the transform coefficient 207, for example, by applying scalar quantization or vector quantization. The quantized transform coefficient 209 may also be referred to as the quantized residual coefficient 209. The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization. Here, n is greater than m. The degree of quantization may be changed by adjusting a quantization parameter (QP). For example, in scalar quantization, different scalings may be applied to achieve finer or coarser quantization. The smaller the quantization step size, the more it corresponds to fine quantization. On the other hand, the larger the quantization step size, the more it corresponds to coarse quantization. The applicable quantization step may be indicated by a quantization parameter (QP). The quantization parameter may be, for example, an index for a predetermined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), a large quantization parameter may correspond to coarse quantization (large quantization step size), and vice versa. Quantization may include division by the quantization step size. For example, the corresponding or inverse inverse quantization by the inverse quantization 210 may include multiplication by the quantization step size. Some standards, for example, embodiments according to HEVC, may be configured to use the quantization parameter to determine the quantization step size. Usually, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation including division. Additional scaling factors for quantization and inverse quantization may be introduced to restore the norm of the residual block that can be changed for the scaling used in the fixed-point approximation of the equations of the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and the inverse quantization may be combined. Alternatively, a customized quantization table may be used and signaled from the encoder to the decoder, for example, in the bitstream.Quantization is a lossy operation, and the loss increases with an increase in the quantization step size.
[0146] Embodiments of the video encoder 20 (each, quantization unit 208) may be configured to output quantization parameters (QP) that are, for example, encoded directly or by the entropy encoding unit 270. As a result, for example, the video decoder 30 may receive and apply the quantization parameters for decoding.
[0147] The inverse quantization unit 210 is configured to apply inverse quantization of the quantization unit 208 to the quantized coefficients, for example, by applying the inverse of the quantization method applied by the quantization unit 208 based on or using the same quantization step size as the quantization unit 208, to obtain inverse quantized coefficients 211. The inverse quantized coefficients 211 are also referred to as inverse quantized residual coefficients 211 and are typically not the same as the transform coefficients due to loss by quantization, but may correspond to the transform coefficients 207.
[0148] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain an inverse transform block 213 in the sample domain. The inverse transform block 213 may also be referred to as an inverse transform inverse quantized block 213 or an inverse transform residual block 213.
[0149] The reconstruction unit 214 (e.g., adder 214) is configured to add the inverse transform block 213 (i.e., the reconstruction residual block 213) to the prediction block 265, for example, by adding the sample values of the reconstruction residual block 213 and the sample values of the prediction block 265, to obtain a reconstruction block 215 in the sample domain.
[0150] Optionally, buffer unit 216 (abbreviated as “buffer” 216), for example line buffer 216, is configured to buffer or store the reconstruction block 215 and respective sample values, for example for intra prediction. In a further embodiment, the encoder may be configured to use the unfiltered reconstruction block and / or respective sample values stored in buffer unit 216 for any kind of estimation and / or prediction, for example for intra prediction.
[0151] An embodiment of encoder 20 may be configured such that, for example, buffer unit 216 is used to store reconstruction block 215 not only for intra prediction 254 but also for loop filter unit 220 (not shown in FIG. 2), and / or such that buffer unit 216 and decoded picture buffer unit 230 form one buffer. A further embodiment may be configured to use blocks or samples (both not shown in FIG. 2) from filtered block 221 and / or decoded picture buffer 230 as input or basis for intra prediction 254.
[0152] The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstruction block 215 to obtain a filtered block 221, for example, to smooth pixel transitions or to improve video quality. The loop filter unit 220 is intended to represent a deblocking filter, a sample-adaptive offset (SAO) filter or other filters, such as one or more filters like a bilateral filter or an adaptive loop filter (ALF) or a sharpening or smoothing filter or a joint filter. The loop filter unit 220 is shown in FIG. 2 as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may be referred to as a filtered reconstruction block 221. The decoded picture buffer 230 may store the reconstructed coded block after the loop filter unit 220 has performed a filtering operation on the reconstructed coded block.
[0153] The loop filter unit 220 (or simply "loop filter" 220) is configured to filter the reconstruction block 215 to obtain the filtered block 221, or typically, to filter the reconstruction samples to obtain the filtered sample values. The loop filter unit is configured to, for example, smooth pixel transitions or, in other cases, improve video quality. The loop filter unit 220 may include a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In an example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering processes may be a deblocking filter, SAO, and ALF. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is executed before deblocking. In another example, the deblocking filter process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 220 is shown in FIG. 2 as an in-loop filter, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as the filtered reconstruction block 221.
[0154] Embodiments of the video encoder 20 (each loop filter unit 220) may be configured to output loop filter parameters (such as SAO filter parameters or ALF filter parameters or LMCS parameters), which are encoded, for example, directly or by the entropy encoding unit 270. As a result, for example, the decoder 30 may receive and apply the same loop filter parameters or respective loop filters for decoding.
[0155] Embodiments of the encoder 20 (each loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information), which are entropy encoded, for example, directly or by any other entropy encoding unit 270 or any other entropy encoding unit. As a result, for example, the decoder 30 may receive and apply the same loop filter parameters for decoding.
[0156] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use in encoding video data by the video encoder 20. The DPB 230 may be formed by any of various memory devices, such as a dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The DPB 230 and the buffer 216 may be provided by the same memory device or separate memory devices. In some examples, the decoded picture buffer (DPB) 230 is configured to store the filtered blocks 221. The decoded picture buffer 230 may be further configured to store other previously filtered blocks of the same current picture or different pictures, such as previously reconstructed pictures, for example, the previously reconstructed and filtered blocks 221, and may provide a complete previous reconstruction, that is, a decoded picture (and corresponding reference blocks and samples), and / or a partially reconstructed current picture (and corresponding reference blocks and samples) for, for example, inter prediction. In some examples, when the reconstruction block 215 is reconstructed without in-loop filtering, the decoded picture buffer (DPB) 230 stores one or more unfiltered reconstruction blocks 215, or generally, unfiltered reconstruction samples, or any other further processed version of the reconstruction block or sample, for example, when the reconstruction block 215 is not filtered by the loop filter unit 220.
[0157] The prediction processing unit 260, also referred to as the block prediction processing unit 260, is configured to receive or obtain the block 203 (the current block 203 of the current picture 201) and the reconstructed picture data, for example, reference samples of the same (current) picture, from the buffer 216, and / or reference picture data 231 from one or more previous decoded pictures from the decoded picture buffer 230, and process such data for prediction, that is, to provide a prediction block 265 which may be an inter prediction block 245 or an intra prediction block 255.
[0158] The mode selection unit 262 may be configured to select a corresponding prediction block 245 or 255 to be used as the prediction block 265 for the calculation of the prediction mode (e.g., intra or inter prediction mode) and / or for the reconstruction of the residual block 205 and the reconstruction of the reconstruction block 215.
[0159] An embodiment of the mode selection unit 262 may be configured to select a prediction mode (e.g., from those supported by the prediction processing unit 260) that provides the most suitable or, in other words, the minimum residual (the minimum residual means better compression for transmission or storage) or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage) or takes both into account or balances them. The mode selection unit 262 may be configured to determine the prediction mode based on rate distortion optimization (RDO), that is, to select a prediction mode that provides the minimum rate distortion optimization or whose associated rate distortion at least meets the prediction mode selection criteria.
[0160] Hereinafter, the prediction processing (e.g., the prediction processing unit 260) and the mode selection (e.g., by the mode selection unit 262) performed by the exemplary encoder 20 are described in more detail.
[0161] As an addition to or an alternative to the above-described embodiments, in another embodiment according to FIG. 17, the mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain original picture data, for example, the original block 203 (the current block 203 of the current picture 17), and reconstructed picture data, for example, filtered and / or unfiltered reconstructed samples or blocks from the same (current) picture and / or one or more previous decoded pictures, for example, from the decoded picture buffer 230 or another buffer (for example, a line buffer not shown). The reconstructed picture data is used as reference picture data for prediction, for example, inter prediction or intra prediction, to obtain the prediction block 265 or predictor 265.
[0162] The mode selection unit 260 may be configured to determine or select a partition of the current block prediction mode (including not partitioning) and a prediction mode (for example, an intra or inter prediction mode), and generate a corresponding prediction block 205 to be used for the calculation of the residual block 205 and the reconstruction of the reconstruction block 215.
[0163] Embodiments of the mode selection unit 260 may be configured to select a partition and prediction mode that provides the best match or, in other words, the minimum residual (the minimum residual means better compression for transmission or storage) or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage) or takes both into account or balances, for example, from those supported or available by the mode selection unit 260. The mode selection unit 260 may be configured to determine a partition and prediction mode based on rate distortion optimization (RDO), that is, to select a prediction mode that provides the minimum rate distortion. Terms such as "best", "minimum", "optimal", etc. in this context do not necessarily represent the overall "best", "minimum", "optimal", etc., but may represent an end or selection criterion such as a value above or below a threshold, or other constraints that are "quasi-optimal selections" but may result in a reduction in complexity and processing time.
[0164] In other words, the partitioning unit 262 partitions a picture from the video sequence into a sequence of coding tree units (CTUs), and the CTU 203 may be further partitioned into smaller block partitions or sub-blocks (which also form blocks) by repeatedly using, for example, quad-tree-partitioning (QT), binary partitioning (BT), or triple-tree-partitioning (TT), or any combination thereof, and may be further configured to perform prediction for each block partition or sub-block. Here, mode selection includes the selection of the tree structure of the partitioned block 203, and the prediction mode is applied to each of the block partitions or sub-blocks.
[0165] The partitioning (e.g., by partition unit 260) and prediction processing (by inter prediction unit 244 and intra prediction unit 254) performed by exemplary video encoder 20 is described in further detail below.
[0166] Partitioning The partition unit 262 may be configured to partition pictures from a video sequence into a sequence of coding tree units (CTUs), and the partition unit 262 may further partition (or split) a coding tree unit (CTU) 203 into smaller partitions, such as smaller square or rectangular block sizes. For a picture having three sample arrays, a CTU is composed of an N×N block of luma samples together with two corresponding blocks of chroma samples. The maximum allowable size of the luma block within a CTU is specified to be 128×128 for the ongoing Versatile Video Coding (VVC), but may be specified to be a value other than 128×128 in the future, such as 256×256. The CTUs of a picture may be clustered / grouped as slices / tile groups, tiles, or bricks. A tile covers a rectangular region of a picture, and a tile can be divided into one or more bricks. A brick is composed of a number of CTU rows within a tile. A tile that is not partitioned into multiple bricks can be called a brick. However, a brick is a proper subset of a tile and is not called a tile. There are two modes of tile groups supported in VVC, namely the raster scan slice / tile group mode and the rectangular slice mode. In the raster scan tile group mode, a slice / tile group contains a sequence of tiles in the tile raster scan of a picture. In the rectangular slice mode, a slice contains a number of bricks of a picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are in the brick raster scan order of the slice. These smaller blocks (which may also be called sub-blocks) may be further partitioned into even smaller partitions. This is also called a tree partition or hierarchical tree partition.Here, for example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively partitioned into, for example, two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchical level 1, depth 1). Here, for example, until the partitioning ends because an end criterion is satisfied, such as reaching the maximum tree depth or the minimum block size, these blocks may again be partitioned into two or more blocks at the next lower tree level, such as tree level 2 (hierarchical level 2, depth 2), and so on. Blocks that are not further partitioned are also called leaf blocks or leaf nodes of the tree. A tree using a partition into two partitions is called a binary-tree (BT), a tree using a partition into three partitions is called a ternary-tree (TT), and a tree using a partition into four partitions is called a quad-tree (QT).
[0167] For example, a coding tree unit (CTU) may be or include the CTB of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or the CTB of samples of a picture encoded using a syntax structure used to encode a monochrome picture or three separate color planes and samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N. As a result, the splitting of components into CTBs is a partition. A coding unit (CU) may be or include the coding block of luma samples, two corresponding coding blocks of chroma samples of a picture having three sample arrays, or the coding block of samples of a picture encoded using a syntax structure used to encode a monochrome picture or three separate color planes and samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values of M and N. As a result, the splitting of CTBs into coding blocks is a partition.
[0168] For example, in an embodiment according to HEVC, a coding tree unit (CTU) may be split into CUs using a quadtree structure shown as a coding tree. The decision of whether to encode a picture area using inter-picture (temporal) or intra-picture (spatial) prediction is made at the leaf CU level. Each leaf CU can be further split into 1, 2, or 4 PUs according to the PU split type. Within one PU, the same prediction process is applied and related information is sent to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU split type, the leaf CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU.
[0169] For example, in an embodiment that complies with the latest video coding standard currently under development, called Versatile Video Coding (VVC), a combined quadtree nested multi-type tree using 2 and 3 partitions divides, for example, a segmentation structure used to partition a coding tree unit. In the coding tree structure within the coding tree unit, the CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quadtree. Next, the quadtree leaf nodes can be further partitioned by a multi-type tree structure. The multi-type tree structure has four split types, vertical 2-way split (SPLIT_BT_VER), horizontal 2-way split (SPLIT_BT_HOR), vertical 3-way split (SPLIT_TT_VER), and horizontal 3-way split (SPLIT_TT_HOR). The multi-type tree leaf nodes are called coding units (CUs), and this segmentation is used for prediction and transform processing without any further partitioning as long as the CU is not too large for the maximum transform length. This means that, in most cases, the CU, PU, and TU have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color component of the CU. VVC develops a unique signaling mechanism for partition information in a quadtree with a nested multi-type tree coding tree structure. In the signaling mechanism, a coding tree unit (CTU) is treated as the root of the quadtree and is first partitioned by the quadtree structure. Each quadtree leaf node (when large enough to allow it) is then further partitioned by the multi-type tree structure. In the multi-type tree structure, a first flag (mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned. When the node is further partitioned, a second flag (mtt_split_cu_vertical_flag) is signaled to indicate the split direction.Next, the third flag (mtt_split_cu_binary_flag) is signaled to indicate whether the split is into two or three parts. Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree split mode (MttSplitMode) of the CU can be derived by the decoder based on certain rules or tables. It should be noted that in a specific design, such as the 64×64 luma block and 32×32 chroma pipeline design in a VVC hardware decoder, as shown in FIG. 6, when either the width or height of the luma encoded block is greater than 64, the TT split is prohibited. The TT split is also prohibited when either the width or height of the chroma encoded block is greater than 32. The pipeline design divides the picture into virtual pipeline data units (VPDUs) defined as non-overlapping units within the picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. The VPDU size is approximately proportional to the buffer size in most pipeline stages. Therefore, it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the ternary tree (TT) and binary tree (BT) partitions can lead to an increase in the VPDU size.
[0170] Furthermore, it should be noted that if a part of the tree node block crosses the lower or right picture boundary, the tree node block is forcibly split until all samples of all encoded CUs are located inside the picture boundary.
[0171] As an example, the Intra Sub-Partition (ISP) tool may split the luma intra prediction block horizontally or vertically into two or four sub-partitions depending on the block size.
[0172] In one example, the mode selection unit 260 of the video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.
[0173] As described above, the encoder 20 is configured to determine the best or optimal prediction mode or select from a set of (predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0174] The set of intra prediction modes may include 35 different intra prediction modes, such as the DC (or average) mode and the non - directional modes such as the planar mode as defined in H.265, or 67 different intra prediction modes, such as the DC (or average) mode and the non - directional modes or directional modes such as the planar mode as defined for VCC. As an example, some conventional angular intra prediction modes are adapted and replaced by, for example, a wide - angle intra prediction mode for non - square blocks as defined in VVC. As another example, to avoid a splitting operation for DC prediction, only the longer side is used to calculate the average for non - square blocks. And the result of the intra prediction in the planar mode may be further modified by a position dependent intra prediction combination (PDPC) method.
[0175] The intra prediction unit 254 is configured to use the reconstructed samples of neighboring blocks of the same current picture to generate an intra prediction block 265 according to the intra prediction mode among the set of intra prediction modes.
[0176] The intra prediction unit 254 (or generally the mode selection unit 260) is further configured to output the intra prediction parameters (or generally the information indicating the intra prediction mode selected for the block) in the form of a syntax element 266 to the entropy encoding unit 270 for inclusion in the encoded picture data 21. As a result, for example, the video decoder 30 may receive and use the prediction parameters for decoding.
[0177] The set of inter prediction modes (or possible inter prediction modes) depends on the available reference pictures (i.e., for example, the previously at least partially decoded pictures stored in the DBP 230) and other inter prediction parameters, for example, whether only the whole or a part of the reference picture is used, for example, to search for the best matching reference block in the search window area around the area of the current block of the reference picture, and / or, for example, whether pixel interpolation, for example, half / semi-pel, quarter-pel, and / or sixteenth-pel interpolation is applied.
[0178] In addition to the prediction modes described above, a skip mode, a direct mode, and / or other inter prediction modes may be applied.
[0179] For example, in extended merge prediction, the merge candidate list for such a mode consists of the following five types of candidates, in order: spatial MVP from spatial neighboring CUs, temporal MVP from CUs at the same position, history-based MVP from the FIFO table, per-pair average MVP, and zero MVP. And decoder side motion vector refinement (DMVR) based on bidirectional consistency may be applied to improve the accuracy of the MV in the merge mode. Merge mode with MVD (MMVD), which is derived from the merge mode with motion vector difference. The MMVD flag is signaled immediately after sending the skip flag and the merge flag to specify whether the MMVD mode is used for the CU. And the adaptive motion vector resolution (AMVR) scheme at the CU level may be applied. AMVR enables the MVD of the CU to be encoded with different accuracies. Depending on the prediction mode of the current CU, the MVD of the current CU can be adaptively selected. When the CU is encoded in the merge mode, the combined inter / intra prediction (CIIP) mode may be applied to the current CU. A weighted average of the inter and intra prediction signals is performed to obtain the CIIP prediction. Affine motion compensation prediction, the affine motion field of the block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). Subblock-based temporal motion vector prediction (SbTMVP), which is similar to the temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vectors of sub CUs within the current CU.The bi-directional optical flow (BDOF), formerly called BIO, is a simpler version that requires far fewer calculations, especially in terms of the number of multiplications and the size of the multipliers. In the triangle partitioning mode, in such a mode, the CU is evenly divided into two triangular-shaped partitions using either diagonal or non-diagonal partitioning. Further, the dual prediction mode is extended beyond simple averaging to enable a weighted average of two prediction signals.
[0180] In addition to the prediction modes described above, a skip mode and / or a direct mode may be applied.
[0181] The prediction processing unit 260 may further partition the block 203 into smaller block partitions or sub-blocks by repeatedly using, for example, quad-tree-partitioning (QT), binary partitioning (BT), ternary-tree-partitioning (TT), or any combination thereof, and may be further configured to perform predictions for each block partition or sub-block, for example. Here, the mode selection includes the selection of the prediction mode applied to each of the tree structure of the partitioned block 203 and the block partitions or sub-blocks.
[0182] The inter prediction unit 244 may include a motion estimation (ME) unit (not shown in FIG. 2) and a motion compensation (MC) unit (not shown in FIG. 2). The motion estimation unit is configured to receive or acquire, for motion estimation, the picture block 203 (the current block 203 of the current picture 201), and at least one or more of the decoded picture 231 or the previous reconstructed blocks, for example, the reconstructed blocks of one or more other / different previous decoded pictures 231. For example, the video sequence may include the current picture and the previous decoded picture 231. Or, in other words, the current picture and the previous decoded picture 231 may be or form part of a sequence of pictures forming the video sequence.
[0183] The encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different pictures of a plurality of other pictures, and provide an offset (spatial offset) between the reference picture (or reference picture index) and / or the position (x, y coordinates) of the reference block and the position of the current block to a motion estimation unit (not shown in FIG. 2) as an inter prediction parameter. This offset is also called a motion vector (MV).
[0184] The motion compensation unit is configured to obtain, for example receive, an inter prediction parameter and perform an inter prediction based on or using the inter prediction parameter to obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit (not shown in FIG. 2) may include fetching or generating a prediction block based on the motion / block vector determined by motion estimation and, optionally, performing interpolation to sub-pixel accuracy. Interpolation filtering may generate additional pixel samples and thus increase the number of candidate prediction blocks that can be used to encode a picture block. When receiving the motion vector of the PU of the current picture block, the motion compensation unit may identify the position of the prediction block pointed to by the motion vector within one of the reference picture lists. The motion compensation unit may also generate syntax elements related to the block and the video slice for use by the video decoder 30 when decoding the picture block of the video slice.
[0185] The intra prediction unit 254 is configured to obtain, for example receive, the picture block 203 (current picture block) and one or more previous reconstructed blocks of the same picture, for example reconstructed neighboring blocks, for intra estimation. The encoder 20 may be configured to select an intra prediction mode, for example, from a plurality of (predetermined) intra prediction modes.
[0186] Embodiments of the encoder 20 may be configured to select an intra prediction mode based on an optimization criterion, for example minimum residual (e.g., the intra prediction mode that provides the prediction block 255 that is most similar to the current picture block 203) or minimum rate distortion.
[0187] The intra prediction unit 254 is further configured to determine an intra prediction block 255 based on intra prediction parameters, such as a selected intra prediction mode. In any case, after selecting the intra prediction mode of the block, the intra prediction unit 254 is also configured to provide the entropy coding unit 270 with the intra prediction parameters, that is, information indicating the selected intra prediction mode for the block. In one example, the intra prediction unit 254 may be configured to perform any combination of the intra prediction techniques described below.
[0188] The entropy coding unit 270 is configured to apply an entropy coding algorithm or method (e.g., variable length coding (VLC) method, context adaptive VLC (CALVC) method, arithmetic coding method, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique) to the quantized residual coefficients 209, inter prediction parameters, intra prediction parameters, and / or loop filter parameters, either individually or together (or not at all), to obtain coded picture data 21 that can be output, for example, in the form of a coded bitstream 21 by an output 272. The coded bitstream 21 may be transmitted to the video decoder 30 or archived for later transmission or reading by the video decoder 30. The entropy coding unit 270 may be further configured to entropy code other syntax elements of the current video slice during coding.
[0189] Other structural variations of the video encoder 20 can be used to encode the video stream. For example, the non-transform-based encoder 20 can directly quantize the residual signal for a particular block or frame without having a transform processing unit 206. In another implementation, the encoder 20 may have a quantization unit 208 and an inverse quantization unit 210 coupled to a single unit.
[0190] FIG. 3 shows an exemplary video decoder 30 configured to implement the technology of the present application. The video decoder 30 is configured to receive encoded picture data (e.g., an encoded bitstream) 21 from, for example, the encoder 100 to obtain a decoded picture 131. During the decoding process, the video decoder 30 receives from the video encoder 100 an encoded video stream representing video data, e.g., picture blocks of an encoded video slice and associated syntax elements.
[0191] In the example of FIG. 3, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., an adder 314), a buffer 316, a loop filter 320, a decoded picture buffer 330, and a prediction processing unit 360. The prediction processing unit 360 may include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. In some examples, the video decoder 30 may perform a decoding path that is typically reciprocal to the encoding path described with respect to the video encoder 100 in FIG. 2.
[0192] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter prediction unit 344, and intra prediction unit 354 are also represented as forming the "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 212, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Accordingly, the descriptions provided for each unit and the functions of video 20 encoder apply correspondingly to each unit and function of video decoder 30.
[0193] Entropy decoding unit 304 is configured to perform entropy decoding on encoded picture data 21 to obtain, for example, quantized coefficients 309 and / or decoded encoded parameters (not shown in FIG. 3), such as inter prediction parameters, intra prediction parameters, loop filter parameters, and / or any or all of other syntax elements. Entropy decoding unit 304 is further configured to transfer inter prediction parameters, intra prediction parameters, and / or other syntax elements to prediction processing unit 360. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.
[0194] Entropy decoding unit 304 parses the bitstream 21 (or generally the encoded picture data 21), and for example, performs entropy decoding on the encoded picture data 21 to obtain, for example, the quantized coefficients 309 and / or decoded encoding parameters (not shown in FIG. 3), such as inter-prediction parameters (e.g., reference picture index and motion vector), intra-prediction parameters (e.g., intra-prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or method corresponding to the encoding method as described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode application unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level. As an addition to or an alternative to slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.
[0195] The inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 112, the reconstruction unit 314 may be functionally identical to the reconstruction unit 114, the buffer 316 may be functionally identical to the buffer 116, the loop filter 320 may be functionally identical to the loop filter 120, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 130.
[0196] An embodiment of the decoder 30 may include a partitioning unit (not shown in FIG. 3). In one example, the prediction processing unit 360 of the video decoder 30 may be configured to perform any combination of the partitioning techniques described above.
[0197] The prediction processing unit 360 may include an inter prediction unit 344 and an intra prediction unit 354. Here, the inter prediction unit 344 may be functionally similar to the inter prediction unit 144, and the intra prediction unit 354 may be functionally similar to the intra prediction unit 154. The prediction processing unit 360 is typically configured to perform block prediction and / or obtain a prediction block 365 from the coded data 21, and receive or obtain prediction-related parameters and / or information regarding the selected prediction mode, for example, from the entropy decoding unit 304 (explicitly or implicitly).
[0198] When a video slice is coded as an intra coded (I) slice, the intra prediction unit 354 of the prediction processing unit 360 is configured to generate a prediction block 365 for the picture blocks of the current video slice based on the signaled intra prediction mode and data from the decoded blocks of the previous frame or picture of the current frame. When a video frame is coded as an inter coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., a motion compensation unit) of the prediction processing unit 360 is configured to generate a prediction block 365 for the video blocks of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. In inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame lists: list 0 and list 1 using a specified configuration technique based on the reference pictures stored in the DPB 330.
[0199] The prediction processing unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing motion vectors and other syntax elements, and use the prediction information to generate a prediction block for the currently decoded video block. For example, the prediction processing unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to encode video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), one or more configuration information of the reference picture lists of the slice, the motion vector of each inter-coded video block of the slice, the inter prediction state of each inter-coded video block of the slice, and other information for decoding video blocks within the current video slice.
[0200] The inverse quantization unit 310 is configured to inverse-quantize, i.e., dequantize, the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 304. The inverse quantization process may include using quantization parameters calculated by the video encoder 100 for each video block within the video slice to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied.
[0201] The inverse quantization unit 310 may receive a quantization parameter (QP) (or generally information regarding inverse quantization) and the quantized coefficients from the coded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304, for example), and apply inverse quantization based on the quantization parameter to the decoded quantized coefficients 309 to obtain inverse-quantized coefficients 311, which may also be referred to as transform coefficients 311.
[0202] The inverse transform processing unit 312 is configured to apply an inverse transform, e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to generate a residual block in the pixel domain.
[0203] The inverse transformation processing unit 312 may be configured to receive the inverse quantized coefficients 311, also referred to as the transformation coefficients 311, and apply a transformation to the inverse quantized coefficients 311 to obtain the reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 may also be referred to as the transformation block 313. The transformation may be an inverse transformation, for example, an inverse DCT, an inverse DST, an inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit 312 may be further configured to receive transformation parameters or corresponding information from the encoded picture data 21 (for example, by parsing and / or decoding by the entropy decoding unit 304, for example) and determine the transformation to be applied to the inverse quantized coefficients 311.
[0204] The reconstruction unit 314 (for example, the adder 314) is configured to add the inverse transformation block 313 (that is, the reconstructed residual block 313) to the prediction block 365, for example, by adding the sample values of the reconstructed residual block 313 and the sample values of the prediction block 365, to obtain the reconstruction block 315 in the sample domain.
[0205] The loop filter unit 320 (either within or after the encoding loop) is configured to filter the reconstructed block 315 to obtain a filtered block 321, for example, to smooth pixel transitions or in other cases to improve video quality. In one example, the loop filter unit 320 may be configured to perform any combination of the filtering techniques described below. The loop filter unit 320 is intended to represent one or more loop filters such as a deblocking filter, a sample-adaptive offset (SAO) filter or other filters, for example, a bilateral filter or an adaptive loop filter (ALF) or a sharpening or smoothing filter or a joint filter. Although the loop filter unit 320 is shown in FIG. 3 as an in-loop filter, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.
[0206] The loop filter unit 320 may also include one or more loop filters such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. In an example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering processes may be a deblocking filter, SAO, and ALF. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is performed before deblocking. In another example, the deblocking filter process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges.
[0207] The decoded video block 321 within a given frame or picture is then stored in a decoded picture buffer 330 that stores reference pictures for use in subsequent motion compensation.
[0208] The decoded video block 321 of the picture is then stored in a decoded picture buffer 330 that stores decoded pictures 331 for use as reference pictures for subsequent motion compensation for other pictures and / or for display output, respectively.
[0209] The decoder 30 is configured to output the decoded picture 331, for example, via an output 332, for presentation or viewing by a user.
[0210] Other variations of the video decoder 30 can be used to decode the compressed video stream. For example, the decoder 30 can generate an output video stream without having a loop filter unit 320. For example, a non-conversion-based decoder 30 can directly inverse quantize the residual signal for a particular block or frame without having an inverse transform processing unit 312. In another implementation, the video decoder 30 can have an inverse quantization unit 310 and an inverse transform processing unit 312 combined in a single unit.
[0211] In addition to or as an alternative to the above-described embodiments, in another embodiment according to FIG. 18, the inter prediction unit 344 may be identical to the inter prediction unit 244 (particularly the motion compensation unit), the intra prediction unit 354 may be functionally identical to the inter prediction unit 254, and performs partitioning or partition determination and prediction based on each piece of information received from the partition and / or prediction parameters or the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304). The mode application unit 360 may be configured to perform prediction (intra or inter prediction) for each block based on the reconstructed picture, block, or each (filtered or unfiltered) sample to obtain a predicted block 365.
[0212] When a video slice is coded as an intra-coded (I) slice, the intra prediction unit 354 of the mode application unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on the coded intra prediction mode and data from the decoded block before the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., motion compensation unit) of the mode application unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. In inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 may configure the reference frame lists: list 0 and list 1 using a prescribed construction technique based on the reference pictures stored in the DPB 330. The same or similar may apply to or be applied in embodiments using tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or as an alternative to slices (e.g., video slices). For example, the video may be coded using I, P, or B tile groups and / or tiles.
[0213] The mode application unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing motion vectors or related information and other syntax elements, and use the prediction information to generate a prediction block for the currently decoded video block. For example, the mode application unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to encode video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of the reference picture lists of the slice, the motion vector of each inter-coded video block of the slice, the inter prediction state of each inter-coded video block of the slice, and other information for decoding video blocks within the current video slice. The same or similar may apply to embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices). For example, the video may be encoded using I, P, or B tile groups and / or tiles.
[0214] An embodiment of the video decoder 30 as shown in FIG. 3 may be configured to partition and / or decode a picture using slices (also referred to as video slices). Here, the picture may be partitioned into or decoded using one or more slices (which do not standardly overlap), and each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., tiles (H.265 / HEVC and VVC) or bricks (VVC)).
[0215] An embodiment of the video decoder 30 as shown in FIG. 3 may be configured to partition and / or decode a picture using slices / tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles). Here, a picture may be partitioned into or decoded using one or more slices / tile groups (which do not overlap standardly), and each slice / tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), such as complete or partial blocks.
[0216] Other variations of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 may be able to generate an output video stream without having a loop filter unit 320. For example, a non-conversion-based decoder 30 may be able to directly inverse quantize the residual signal for a particular block or frame without having an inverse transform processing unit 312. In another implementation, the decoder 30 may have an inverse quantization unit 310 and an inverse transform processing unit 312 coupled to a single unit.
[0217] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0218] FIG. 4 is a schematic diagram of a video encoding device 400 according to an embodiment of the present disclosure. The video encoding device 400 is suitable for implementing the embodiments of the disclosure as described herein. In the embodiment, the video encoding device 400 may be a decoder such as the video decoder 30 of FIG. 1A, or an encoder such as the video encoder 20 of FIG. 1A. In the embodiment, the video encoding device 400 may be one or more components of the video decoder 30 of FIG. 1A or the video encoder 20 of FIG. 1A as described above.
[0219] The video encoding device 400 includes an ingress port 410 and a receiver unit (Rx) 420 for receiving data, a processor, a logic unit, or a central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an egress port 450 for transmitting data, and a memory 460 for storing data. The video encoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components for ingress or egress of optical or electrical signals connected to the ingress port 410, the receiver unit 420, the transmitter unit 440, and the egress port 450.
[0220] Processor 430 is implemented by hardware and software. Processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processor), FPGA, ASIC, and DSP. Processor 430 communicates with ingress port 410, receiver unit 420, transmitter unit 440, ingress port 450, and memory 460. Processor 430 includes an encoding module 470. Encoding module 470 implements the embodiments of the above disclosure. For example, encoding module 470 implements, processes, prepares, or provides various encoding operations. What is included in encoding module 470 thus provides a substantial improvement to the functionality of video encoding device 400 and results in a conversion of video encoding device 400 to different states. Alternatively, encoding module 470 is implemented as instructions stored in memory 460 and executed by processor 430.
[0221] Memory 460 includes one or more disks, tape drives, and solid state drives and may be used as an overflow data storage device for storing a program when the program is selected for execution and for storing instructions and data read during program execution. Memory 460 may be volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0222] FIG. 5 is a simplified block diagram of a device 500 that may be used as one or both of the source device 310 and the destination device 320 from FIG. 1 according to an exemplary embodiment. The device 500 can implement the technology of the present application described above. The device 500 can be in the form of a computing system including a plurality of computing devices, or in the form of a single computing device, such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.
[0223] The processor 502 within the device 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices that can manipulate or process information that exists currently or will be developed in the future. Although the disclosed implementation can be carried out by a single processor, such as the processor 502 as shown, the benefits in terms of speed and efficiency can be achieved using more than one processor.
[0224] In one implementation, the memory 504 within the machine 500 can be a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and application programs 510. The application programs 510 include at least one program that enables the processor 502 to execute the methods described herein. For example, the application programs 510 can include applications 1 - N that further include a video encoding application that executes the methods described herein. The device 500 can also include additional memory in the form of secondary storage 514, which can be, for example, a memory card used with a mobile computing device. Since video communication sessions can include a significant amount of information, they can be stored in whole or in part in the secondary storage 514 and loaded into the memory 504 as needed for processing. The device 500 can also include one or more output devices such as a display 518. In one example, the display 518 can be a touch-sensitive display that combines a display with touch-sensitive elements that operate to sense touch inputs. The display 518 can be coupled to the processor 502 via the bus 512.
[0225] Device 500 may also include one or more output devices such as display 518. In one example, display 518 may be a touch-sensitive display that combines a touch-sensitive element that operates to sense touch input with a display. Display 518 may be coupled to processor 502 via bus 512. Other output devices that enable a user to program or otherwise use device 500 may be provided in addition to or in place of display 518. When the output device is or includes a display, the display may be implemented in various ways, including a liquid crystal display (LCD), a cathode-ray tube (CRT) display, a plasma display, or a light emitting diode (LED) display such as an organic LED (OLED) display.
[0226] Device 500 may also include or communicate with an image sensing device 520, such as a camera, or any other existing or future-developed image sensing device 520 that can sense an image, such as an image of a user operating device 500. Image sensing device 520 may be positioned to face a user operating device 500. In one example, the position and optical axis of image sensing device 520 may be configured such that the field of view includes and can see display 518 in an area immediately adjacent to display 518.
[0227] Device 500 may also include or communicate with an audio sensing device 522, such as a microphone, or any other existing or future-developed audio sensing device that can sense audio near device 500. Audio sensing device 522 may be positioned to face a user operating device 500 and may be configured to receive audio generated by the user, such as conversation or other utterances, while the user is operating device 500.
[0228] FIG. 5 shows the processor 502 and memory 504 of device 500 integrated into a single unit, although other configurations are available. The operation of processor 502 can be distributed across a plurality of machines (each machine having one or more processors) that can be passed to or directly coupled to a local area or other network. Memory 504 can be distributed across a plurality of machines, such as network-based memory or memory among a plurality of machines that execute the operation of device 500. Although shown here as a single bus, bus 512 of device 500 can be composed of a plurality of buses. Further, secondary storage 514 can be directly coupled to other components of device 500 or accessed via a network and can include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Device 500 can, therefore, be implemented in a variety of configurations.
[0229] Next Generation Video Coding (NGVC) removes the separation of the CU, PU, and TU concepts and supports further flexibility in CU partition shapes. The size of the CU corresponds to the size of the coding node and can be square or non-square (e.g., rectangular) in shape.
[0230] Additionally or alternatively, the TU or PU can also be obtained by splitting the CU.
[0231] J. An et al., “Block partitioning structure for next generation video coding”, International Telecommunication Union, COM16-C966, September 2015 (hereinafter referred to as "VCEG proposal COM16-C966") proposed a quad-tree-binary-tree (QTBT) partitioning technique for future video coding standards after HEVC. Simulations have shown that the proposed QTBT structure is more efficient than the quad-tree structure in HEVC where it is used. In HEVC, the inter-prediction of small blocks is restricted to reduce motion compensation memory access, and inter-prediction is not supported for 4×4 blocks. In the QTBT of JEM, these restrictions are removed.
[0232] In QTBT, a CU can have either a square or rectangular shape. As shown in Figure 6, a coding tree unit (CTU) is first partitioned by a quad-tree structure. The leaf nodes of the quad-tree can be further partitioned by a binary-tree structure. There are two types of binary-tree partitions: symmetric horizontal partition and symmetric vertical partition. In each case, the node is partitioned either horizontally or vertically by dividing the node in the center. The leaf nodes of the binary-tree are called coding units (CUs), and without any further partitioning, segmentation is used for prediction and transformation processing. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. A CU is sometimes composed of coding blocks (CBs) of different color components. For example, in the case of P and B slices in 4:2:0 chroma format, one CU contains one luma CB and two chroma CBs, and sometimes it is composed of a single-component CB. For example, in the case of an I slice, one CU contains only one luma CB or just two chroma CBs.
[0233] The following parameters are defined for the QTBT partitioning scheme. - CTU size: The size of the root node of the quadtree, the same concept as in HEVC. - MinQTSize: The minimum allowable quadtree leaf node size. - MaxBTSize: The maximum allowable binary tree root node size. - MaxBTDepth: The maximum allowable binary tree depth. - MinBTSize: The minimum allowable binary tree leaf node size.
[0234] In an example of the QTBT partition structure, when a quadtree node has a size equal to or smaller than MinQTSize, no further quadtree is considered. Since its size exceeds MaxBTSize, it is not further divided by a binary tree. In other cases, a leaf quadtree node may be further partitioned by a binary tree. Thus, a quadtree leaf node is also the root node of a binary tree, which has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further division is considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal division is considered. Similarly, when a binary tree node has a height equal to MinBTSize, no further vertical division is considered. The leaf node of the binary tree is further processed by prediction and transformation processing without any further partitioning. In JEM, the maximum CTU size is 256×256 luma samples. The leaf node (CU) of the binary tree may be further processed (e.g., by performing prediction processing and transformation processing) without any further partitioning.
[0235] Figure 6 shows an example of a block 30 (e.g., CTB) partitioned using the QTBT partitioning technique. As shown in Figure 6, using the QTBT partitioning technique, each of the blocks is symmetrically divided through the center of each block. Figure 7 shows the tree structure corresponding to the block partition of Figure 6. The solid lines in Figure 7 indicate a quadtree division, and the dashed lines indicate a binary tree division. In one example, at each division (i.e., non-leaf) node of the binary tree, a syntax element (e.g., a flag) is signaled to indicate the type of division (e.g., horizontal or vertical) being performed. Here, 0 indicates a horizontal division, and 1 indicates a vertical division. In a quadtree division, since the quadtree division always divides the block into four sub-blocks having equal sizes horizontally and vertically, there is no need to indicate the division type.
[0236] As shown in Figure 7, at node 50, block 30 is divided using the QT partition into the four blocks 31, 32, 33, and 34 shown in Figure 6. Block 34 is not further divided and is thus a leaf node. At node 52, block 31 is further divided into two blocks using the BT partition. As shown in Figure 7, node 52 is marked with a 1 indicating a vertical division. Thus, the division at node 52 results in block 37 and a block containing both blocks 35 and 36. Blocks 35 and 36 are generated by a further vertical division at node 54. At node 56, block 32 is further divided into two blocks 38 and 39 using the BT partition.
[0237] At node 58, block 33 is divided into four equally sized blocks using a QT partition. Blocks 43 and 44 are generated from this QT partition and are not further divided. At node 60, the upper left block is first divided using a vertical binary tree partition to yield block 40 and the right vertical block. The right vertical block is then divided into blocks 41 and 42 using a horizontal binary tree partition. The lower right block generated from the quadtree partition at node 58 is divided into blocks 45 and 46 using a horizontal binary tree partition at node 62. As shown in FIG. 7, node 62 is marked with a 0 indicating a horizontal split.
[0238] In addition to QTBT, a block partition structure called a multi-type-tree (MTT) is proposed to replace the BT in the CU structure based on QTBT. This means that the CTU block can first be divided by a QT partition to obtain the CTU block, and then the block can be secondarily divided by an MTT partition.
[0239] The MTT partition structure is still a recursive tree structure. In MTT, multiple different partition structures (e.g., two or more) are used. For example, according to the MTT technique, two or more different partition structures may be used for each non-leaf node of the tree structure at each depth of the tree structure. The depth of a node in the tree structure may represent the length of the path from the node to the root of the tree structure (e.g., the number of splits).
[0240] In MTT, there are two partition types: the BT partition and the ternary-tree (TT) partition. The partition type can be selected from the BT partition and the TT partition. The TT partition structure is different from the QT or BT structure in that the TT partition structure does not split the block in the center. The central region of the block remains together within the same sub-block. Different from the QT that results in four blocks or the binary tree that results in two blocks, the split by the TT partition structure results in three blocks. Exemplary partition types by the TT partition structure include symmetric partition types (both horizontal and vertical), as well as asymmetric partition types (both horizontal and vertical). Further, the symmetric partition types by the TT partition structure may be non-uniform / non-isomorphic or uniform / isomorphic. The asymmetric partition types by the TT partition structure are non-uniform / non-isomorphic. In one example, the TT partition structure may include at least one of the following partition types: horizontal uniform / isomorphic symmetric ternary tree, vertical uniform / isomorphic symmetric ternary tree, horizontal non-uniform / non-isomorphic symmetric ternary tree, vertical non-uniform / non-isomorphic symmetric ternary tree, horizontal non-uniform / non-isomorphic asymmetric ternary tree, or vertical non-uniform / non-isomorphic asymmetric ternary tree partition type.
[0241] Generally, a non-uniform / non-isomorphic symmetric ternary partition type is a partition type that is symmetric with respect to the center line of the block, but at least one of the resulting three blocks is not the same size as the other two. One suitable example is when the end blocks are one-quarter the size of the block and the central block is one-half the size of the block. A uniform / isomorphic symmetric ternary partition type is a partition type that is symmetric with respect to the center line of the block, and all of the resulting blocks are the same size. Such a partition is possible when the block height or width is a multiple of 3 depending on whether it is a vertical or horizontal split. A non-uniform / non-isomorphic asymmetric ternary partition type is a partition type that is not symmetric with respect to the center line of the block, and at least one of the resulting blocks is not the same size as the other two.
[0242] FIG. 8 is a conceptual diagram showing an exemplary horizontal ternary partition type. FIG. 9 is a conceptual diagram showing an exemplary vertical ternary partition type. In both FIGS. 8 and 9, h represents the height of the block in the luma or chroma sample, and w represents the width of the block in the luma or chroma sample. Note that the center line of each block does not represent the boundary of the block (i.e., the ternary partition does not split through the block at the center line). Rather, the center line is used to indicate whether a particular partition type is symmetric or asymmetric with respect to the center line of the original block. The center line also runs along the direction of the split.
[0243] As shown in FIG. 8, block 71 is partitioned by a horizontal uniform / isomorphic symmetric partition type. The horizontal uniform / isomorphic symmetric partition type generates upper and lower halves that are symmetric with respect to the center line of block 71. The horizontal uniform / isomorphic symmetric partition type generates three equal-sized sub-blocks each having a height of h / 3 and a width of w. The horizontal uniform / isomorphic symmetric partition type is possible when the height of block 71 is evenly divisible by 3.
[0244] Block 73 is partitioned by a horizontal non-uniform / non-conformal symmetric partitioning type. The horizontal non-uniform / non-conformal symmetric partitioning type generates upper and lower halves that are symmetric with respect to the center line of block 73. The horizontal non-uniform / non-conformal symmetric partitioning type generates two blocks of equal size (e.g., upper and lower blocks having a height of h / 4), and a central block of a different size (e.g., a central block having a height of h / 2). In one example, according to the horizontal non-uniform / non-conformal symmetric partitioning type, the area of the central block is equal to the combined area of the upper and lower blocks. In some examples, the horizontal non-uniform / non-conformal symmetric partitioning type may be preferred for blocks having a height that is a power of two (e.g., 2, 4, 8, 16, 32, etc.).
[0245] Block 75 is partitioned by a horizontal non-uniform / non-conformal asymmetric partitioning type. The horizontal non-uniform / non-conformal asymmetric partitioning type does not generate upper and lower halves that are symmetric with respect to the center line of block 75 (i.e., the upper and lower halves are asymmetric). In the example of FIG. 8, the horizontal non-uniform / non-conformal asymmetric partitioning type generates an upper block having a height of h / 4, a central block having a height of 3h / 8, and a lower block having a height of 3h / 8. Of course, other asymmetric configurations may be used.
[0246] As shown in FIG. 9, block 81 is partitioned by a vertical uniform / conformal symmetric partitioning type. The vertical uniform / conformal symmetric partitioning type generates left and right halves that are symmetric with respect to the center line of block 81. The vertical uniform / conformal symmetric partitioning type generates three sub-blocks of equal size, each having a width of w / 3 and a width of h. The vertical uniform / conformal symmetric partitioning type is possible when the width of block 81 is evenly divisible by 3.
[0247] Block 83 is partitioned by a vertical non-uniform / non-isomorphic symmetric partitioning type. The vertical non-uniform / non-isomorphic symmetric partitioning type generates left and right halves that are symmetric with respect to the center line of block 83. The vertical non-uniform / non-isomorphic symmetric partitioning type generates left and right halves that are symmetric with respect to the center line of 83. The vertical non-uniform / non-isomorphic symmetric partitioning type generates two blocks of equal size (e.g., a left and a right block with a width of w / 4), and a center block of a different size (e.g., a center block with a width of w / 2). In one example, according to the vertical non-uniform / non-isomorphic symmetric partitioning type, the area of the center block is equal to the combined area of the left and right blocks. In some examples, the vertical non-uniform / non-isomorphic symmetric partitioning type may be preferred for blocks having a width that is a power of 2 (e.g., 2, 4, 8, 16, 32, etc.).
[0248] Block 85 is partitioned by a vertical non-uniform / non-isomorphic asymmetric partitioning type. The vertical non-uniform / non-isomorphic asymmetric partitioning type does not generate left and right halves that are symmetric with respect to the center line of block 85 (i.e., the left and right halves are asymmetric). In the example of FIG. 9, the vertical non-uniform / non-isomorphic asymmetric partitioning type generates a left block with a width of w / 4, a center block with a width of 3w / 8, and a right block with a width of 3w / 8. Of course, other asymmetric configurations may be used.
[0249] In addition to the parameters of the QTBT, the following parameters are defined for the MIT partitioning scheme. -MaxBTSize: The maximum allowable binary tree root node size. -MinBtSize: The minimum allowable binary tree root node size. -MaxMttDepth: The maximum multi-type tree depth. -MaxMttDepth offset: The maximum multi-type tree depth offset. -MaxTtSize: The maximum allowable ternary tree root node size. -MinTtSize: The minimum allowable ternary tree root node size. -MinCbSize: Minimum allowable quantization block size.
[0250] Embodiments of the present disclosure may be implemented by a video encoder or a video decoder such as the video encoder 20 of FIG. 2 or the video decoder 30 of FIG. 3 according to the embodiments of the present application. One or more structural elements of the video encoder 20 or the video decoder 30 including partition units may be configured to execute the techniques of the embodiments of the present disclosure.
[0251] In [JVET-K1001-v4], JVET AHG report, J.-R. Ohm, G.J. Sullivan, http: / / phenix.int-evry.fr / jvet / , the syntax elements of MinQtSizeY (log2_min_qt_size_intra_slices_minus2 and log2_min_qt_size_inter_slices_minus2), and the syntax elements of MaxMttDepth (max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices) are signaled in the SPS.
[0252] The syntax element (log2_diff_ctu_max_bt_size) of the difference between the luma CTB size and MaxBtSizeY is signaled in the slice header.
[0253] CtbSizeY and the corresponding syntax element log2_ctu_size_minus2 indicate the size of the maximum coding block size in terms of the number of luma samples.
[0254] MinCbSizeY is defined as the minimum luma size of the leaf blocks resulting from the quadtree partitioning of a CTU (coding tree unit). The size can be indicated by the number of samples for either the width or the height of the block. It can also indicate the width and height together in the case of a square block. As an example, when MinQtSizeY is equal to 16, coding blocks having a size smaller than or equal to 16 cannot be partitioned into child blocks using the quadtree partitioning method. In the conventional MinQtSizeY, log2_min_qt_size_intra_slices_minus2 and log2_min_qt_size_inter_slices_minus2 are used to indicate the minimum quadtree block size. Note that the indication of the size can also be an indirect indication, meaning that log2_min_qt_size_intra_slices_minus2 can be the binary logarithm (base 2) of the number of luma samples of the minimum quadtree block. MaxMttDepth is defined as the maximum hierarchical depth of the coding units resulting from the multi-type tree partitioning of a quadtree leaf or a CTU. A coding tree unit (or CTB, Coding Tree Block) describes the maximum block size used to partition a picture frame. MaxMttDepth describes the upper limit of the number of consecutive 2- or 3-way partitions that can be applied to obtain child blocks. As an example, assume that the CTU size is 128×128 (width equal to 128 and height equal to 128) and MaxMttDepth is equal to 1. In this case, the parent block (size 128×128) can first be partitioned into two 128×64 child blocks using a 2-way partition. However, since the maximum number of allowed 2-way partitions is reached, the child blocks cannot apply any consecutive 2-way partitions (resulting in either 128×32 or 64×64 child blocks). Note that MaxMttDepth can control either the maximum 2-way depth or the maximum 3-way depth, or both, simultaneously. When controlling both 2- and 3-way partitions simultaneously, one 2-way partition followed by one 3-way partition can be counted as two hierarchical partitions.In the conventional MaxMttDepth, max_mtt_hierarchy_depth_inter_slices and max_mtt_hierarchy_depth_intra_slices are used to indicate the maximum hierarchical structure depth of coding units resulting from a multi-type tree.
[0255] Note that the names of syntax elements are used as they appear in the prior art. However, the names can be changed, and thus it should be clearly stated that what should be considered important is the logical meaning of the syntax elements.
[0256] MaxBtSizeY is defined, in terms of the number of samples, as the maximum luma size (width or height) of a coding block that can be divided using binary splitting. As an example, when MaxBtSizeY is equal to 64, a coding block larger in size in either width or height cannot be divided using binary splitting. This means that a block with size 128×128 cannot be divided using binary splitting, while a block with size 64×64 can be divided using binary splitting.
[0257] MinBtSizeY is defined, in terms of the number of samples, as the minimum luma size (width or height) of a coding block that can be divided using binary splitting. As an example, when MinBtSizeY is equal to 16, a coding block smaller in size or equal in size in either width or height cannot be divided using binary splitting. This means that a block with size 8×8 cannot be divided using binary splitting, while a block with size 32×32 can be divided using binary splitting.
[0258] MinCbSizeY is defined as the minimum coding block size. As an example, MinCbSizeY can be equal to 8. This means that a parent block with size 8×8 cannot be split using any of the split modes, as it is guaranteed that the resulting child blocks will be smaller than MinCbSizeY in either width or height. According to a second example, when MinCbSizeY is equal to 8, a parent block with size 8×16 cannot be partitioned, for example, using a quadtree split. This is because the resulting four child blocks could have a size of 4×8 (width equal to 4 and height equal to 8), and the width of the resulting child blocks could be smaller than MinCbSizeY. In the second example, two different syntax elements can be used to independently limit the width and height, but it was assumed that MinCbSizeY applies to both the width and height of the block.
[0259] MinTbSizeY is defined as the minimum transform block size of a coding block that can be split using 3-way split in terms of the number of samples. As an example, when MinTbSizeY is equal to 16, a coding block with a size smaller than or equal to it in either width or height cannot be split using 3-way split. This means that a block with size 8×8 cannot be split using 3-way split, while a block with size 32×32 can be split using 3-way split.
[0260] Sequence Parameter Set RBSP (Raw Byte Sequence Payload) Syntax ([JVET-K1001-v4], Section 7.3.2.1) [Ed.(BB): Preliminary basic SPS, subject to further study, and further specification development in progress.]
Table 1
[0261] In these syntax tables, boldface indicates syntax elements included in the bitstream. Elements not shown in boldface are conditions or placeholders for further syntax units.
[0262] Slice header syntax (Section 7.3.3 of [JVET-K1001-v4]) [Ed.(BB): Preliminary basic slice header, subject to further study and under further specification development.]
Table 2
[0263] The semantics of syntax elements, i.e., how the syntax elements included in the bitstream should be interpreted, are also provided by the standard. The semantics of the above-mentioned elements are provided below.
[0264] Sequence parameter set RBSP semantics (Section 7.4.3.1 of [JVET-K1001-v4]) log2_ctu_size_minus2 + 2 specifies the luma coding tree block size of each CTU.
[0265] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows.
Equation
[0266] log2_min_qt_size_intra_slices_minus2 + 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of CTUs in slices having slice_type equal to 2(I). The value of log2_min_qt_size_intra_slices_minus2 should be included in the range from 0 to CtbLog2SizeY - 2, inclusive at both ends. [Number] [Ed.(BB): A leaf of a quadtree can be either an encoding unit or the root of a nested multi-type tree.]
[0267] log2_min_qt_size_inter_slices_minus2 + 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of CTUs in slices having slice_type equal to 0(B) or 1(P). The value of log2_min_qt_size_inter_slices_minus2 should be included in the range from 0 to CtbLog2SizeY - 2, inclusive at both ends. [Number]
[0268] max_mtt_hierarchy_depth_inter_slices specifies the maximum hierarchical structure depth of the encoding units resulting from the multi-type tree partitioning of the quadtree leaves in slices having slice_type equal to 0(B) or 1(P). The value of max_mtt_hierarchy_depth_inter_slices should be included in the range from 0 to CtbLog2SizeY - MinTbLog2SizeY, inclusive at both ends.
[0269] max_mtt_hierarchy_depth_intra_slices specifies the maximum hierarchical depth of coding units resulting from the multi-type tree partitioning of quadtree leaves in slices having slice_type equal to 2(I). The value of max_mtt_hierarchy_depth_intra_slices should be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY.
[0270] Slice header semantics (Section 7.4.4 of [JVET-K1001-v4])
[0271] log2_diff_ctu_max_bt_size specifies the difference between the luma CTB size and the maximum luma size (width or height) of a coded block that can be partitioned using 2-way partitioning. The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY.
[0272] When log2_diff_ctu_max_bt_size does not exist, the value of log2_diff_ctu_max_bt_size is assumed to be equal to 2.
[0273] The variables MinQtLog2SizeY, MaxBtLog2SizeY, MinBtLog2SizeY, MaxTtLog2SizeY, MinTtLog2SizeY, MaxBtSizeY, MinBtSizeY, MaxTtSizeY, MinTtSizeY, and MaxMttDepth are derived as follows.
Number
[0274] In Embodiment 1 of the present disclosure: Embodiment 1 relates to the signaling of individual partition constraint related high-level syntax elements (e.g., MinQtSizeY, MaxMttDepht, MaxBtSizeY) for each slice type in the SPS (sequence parameter sets), and / or the signaling of a partition constraint override enable (or disable) flag.
[0275] Signaling a partition constraint override flag in the slice header means the following: If the flag is true, Override the partition constraint related high-level syntax elements in the slice header, where overriding means re-signaling the elements in the slice header. In other cases, Estimate the partition constraint related high-level syntax elements based on the values signaled from the SPS according to the slice type.
[0276] In other words, the partition constraint override flag is signaled in the slice header to indicate whether one or more partition constraint parameters are signaled in the slice header or in a parameter set such as the SPS. Note that the parameter set does not necessarily have to be the SPS. It can also be the PPS, or any other type of parameter set related to, for example, more than one slice, e.g., one or more pictures of a video.
[0277] Instead, In SPS, partition constraint related high-level syntax elements (e.g., MinQtSizeY, MaxMttDepht, MaxBtSizeY) are signaled individually in groups based on features or indexes, and a partition constraint override enable (or disable) flag is signaled.
[0278] In the slice header, a partition constraint override flag is signaled, and: If the flag is true, Override the partition constraint related high-level syntax elements in the slice header, where overriding means re-signaling the elements in the slice header. In other cases, Estimate the partition constraint related high-level syntax elements based on the values signaled from the SPS according to the features or indexes used to individualize the signaling.
[0279] Regarding the positions of signaling and overriding, instead, for example: The signaling of partition constraint related high-level syntax elements can be performed in the parameter set, and the override operation can be performed in the slice header.
[0280] The signaling of partition constraint related high-level syntax elements can be performed in the parameter set, and the override operation can be performed in the tile header.
[0281] The signaling of partition constraint related high-level syntax elements can be performed in the first parameter set, and the override operation can be performed in the second parameter set.
[0282] The signaling of high-level syntax elements related to partition constraints can be performed in the slice header, and the override operation can be performed in the tile header.
[0283] Generally, when the signaling of high-level syntax elements related to partition constraints is performed in the first parameter set and the override operation is performed in the second parameter set, efficient coding can be achieved in that the first set is related to a larger image / video region than the second parameter set.
[0284] Technical advantages (e.g., signaling in SPS, override in slice header): High-level partition constraints control the trade-off between partition complexity and coding efficiency from the partition. The present invention guarantees the flexibility to control the trade-off for individual slices.
[0285] Both the encoder and decoder perform the same (corresponding) operations.
[0286] The corresponding syntax and semantic changes based on the prior art are shown below: Changed sequence parameter set RBSP syntax (Section 7.3.2.1 of [JVET-K1001-v4]) [Ed.(BB): Preliminary basic SPS, subject to further study, and further specification development in progress.]
Table 3
[0287] Changed slice header syntax (Section 7.3.3 of [JVET-K1001-v4]) [Ed.(BB): Preliminary basic slice header, subject to further study, and further specification development in progress.]
Table 4
[0288] Changed sequence parameter set RBSP semantics (Section 7.4.3.1 of [JVET-K1001-v4])
[0289] The partition_constraint_override_enabled_flag equal to 1 specifies the presence of the partition_constraint_override_flag in the slice header for slices that refer to the SPS. The partition_constraint_override_enabled_flag equal to 0 specifies the absence of the partition_constraint_override_flag in the slice header for slices that refer to the SPS.
[0290] sps_log2_min_qt_size_intra_slices_minus2 + 2 specifies the nominal minimum luma size of leaf blocks resulting from the quadtree partitioning of a CTU in slices with slice_type equal to 2(I) that refer to the SPS, unless overridden by the minimum luma size of leaf blocks resulting from the quadtree partitioning of a CTU present in the slice header of the slice that refers to the SPS. The value of log2_min_qt_size_intra_slices_minus2 should be included in the range that includes both ends of 0 to CtbLog2SizeY - 2.
Number
[0291] sps_log2_min_qt_size_inter_slices_minus2 specifies the nominal minimum luma size of leaf blocks resulting from the quadtree partitioning of CTUs in slices with slice_type equal to 0 (B) or 1 (P) that refer to the SPS, provided that the nominal minimum luma size of leaf blocks resulting from the quadtree partitioning of CTUs in the slice header of the slice that refers to the SPS does not override it. The value of log2_min_qt_size_inter_slices_minus2 should be within the range including both ends of 0 to CtbLog2SizeY - 2.
Number
[0292] sps_max_mtt_hierarchy_depth_inter_slices specifies the nominal maximum hierarchical structure depth of coding units resulting from the multi-type tree partitioning of quadtree leaves in slices with slice_type equal to 0 (B) or 1 (P) that refer to the SPS, provided that the maximum hierarchical structure depth of coding units resulting from the multi-type tree partitioning of quadtree leaves in the slice header of the slice that refers to the SPS does not override it. The value of max_mtt_hierarchy_depth_inter_slices should be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY.
[0293] sps_max_mtt_hierarchy_depth_intra_slices specifies the specified maximum hierarchical structure depth of coding units resulting from the multi-type tree partition of quadtree leaves, provided that the specified maximum hierarchical structure depth of coding units resulting from the multi-type tree partition of quadtree leaves in the slice header of a slice referring to the SPS is not overridden by the maximum hierarchical structure depth of coding units resulting from the multi-type tree partition of quadtree leaves in a slice having a slice_type equal to 2(I) that refers to the SPS. The value of max_mtt_hierarchy_depth_intra_slices should be included in the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY.
[0294] sps_log2_diff_ctu_max_bt_size_intra_slices specifies the specified difference between the luma CTB size and the maximum luma size (width or height) of a coded block that can be partitioned using 2-way partitioning, provided that the specified difference between the luma CTB size and the maximum luma size (width or height) of a coded block that can be partitioned using 2-way partitioning in the slice header of a slice referring to the SPS is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) of a coded block that can be partitioned using 2-way partitioning in a slice having a slice_type equal to 2(I) that refers to the SPS. The value of log2_diff_ctu_max_bt_size should be included in the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY.
[0295] sps_log2_diff_ctu_max_bt_size_inter_slices specifies the defined difference between the luma CTB size and the maximum luma size (width or height) of a codable block that can be partitioned using 2-way partitioning, provided that the defined difference between the luma CTB size and the maximum luma size (width or height) of a codable block that can be partitioned using 2-way partitioning as present in the slice header of a slice that refers to the SPS is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) of a codable block that can be partitioned using 2-way partitioning in a slice that has a slice_type equal to 0(B) or 1(P) and refers to the SPS. The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY.
[0296] Changed slice header semantics ([JVET-K1001-v4], Section 7.4.4) A partition_constraint_override_flag equal to 1 specifies that the partition constraint parameter is present in the slice header. A partition_constraint_override_flag equal to 0 specifies that the partition constraint parameter is not present in the slice header. When not present, the value of partition_cosntraints_override_flag is assumed to be equal to 0.
[0297] log2_min_qt_size_minus2 + 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of the CTU for the current slice. The value of log2_min_qt_size_inter_slices_minus2 should be included in the range that includes both ends of 0 to CtbLog2SizeY - 2. When it does not exist, the value of log2_min_qt_size_minus2 is presumed to be equal to sps_log2_min_qt_size_intra_slices_minus2 when slice_type is equal to 2 (I), and is presumed to be equal to sps_log2_min_qt_size_inter_slices_minus2 when slice_type is equal to 0 (B) or 1 (P).
[0298] max_mtt_hierarchy_depth specifies the maximum hierarchical depth of the coding units resulting from the multi-type tree partitioning of the quadtree leaf for the current slice. The value of max_mtt_hierarchy_depth_intra_slices should be included in the range that includes both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, the value of max_mtt_hierarchy_depth is presumed to be equal to sps_max_mtt_hierarchy_depth_intra_slices for slice_type equal to 2 (I), and is presumed to be equal to sps_max_mtt_hierarchy_depth_inter_slices for slice_type equal to 0 (B) or 1 (P).
[0299] log2_diff_ctu_max_bt_size specifies the difference between the luma CTB size and the maximum luma size (width or height) of the coding block that can be divided using 2-way partitioning for the current slice. The value of log2_diff_ctu_max_bt_size should be included in the range from 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When it does not exist, the value of log2_diff_ctu_max_bt_size is presumed to be equal to sps_log2_diff_ctu_max_bt_size_intra_slices for slice_type equal to 2 (I), and presumed to be equal to sps_log2_diff_ctu_max_bt_size_inter_slices for slice_type equal to 0 (B) or 1 (P).
[0300] The variables MinQtLog2SizeY, MaxBtLog2SizeY, MinBtLog2SizeY, MaxTtLog2SizeY, MinTtLog2SizeY, MaxBtSizeY, MinBtSizeY, MaxTtSizeY, MinTtSizeY, and MaxMttDepth are derived as follows. [Number]
[0301] In an alternative implementation of Embodiment 1 of the present disclosure, it is described as follows: A sequence parameter set (SPS) applies to the entire encoded video sequence and includes parameters that do not change for each picture in the encoded video sequence (abbreviated as CVS). All pictures within the same CVS use the same SPS.
[0302] A PPS includes parameters that may vary for different pictures within the same encoded video sequence. However, multiple pictures may refer to the same PPS even if they have different slice coding types (I, P, and B).
[0303] As mentioned in Embodiment 1 of the present disclosure, high-level partition constraints control the trade-off between partition complexity and coding efficiency from the partition. To address the advantage of flexible control between complexity and coding efficiency in individual pictures / slices, instead of the method in Embodiment 1 (signaling partition constraint syntax elements within the SPS and overriding the partition constraint syntax elements within the slice header based on the partition constraint override flag signaled in the slice header), the partition constraint syntax elements (MinQtSizeY, MaxMttDepth, MaxBtSizeY, MaxTtSizeY, etc.) are signaled within the PPS to adjust the trade-off between partition complexity and coding efficiency from picture-level partitions. If each picture uses an individual PPS, the adjustment is applied to the individual picture. If multiple pictures refer to the same PPS, the same adjustment is applied to the pictures.
[0304] The PPS-level signaling of partition constraint syntax elements can be signaled within one group. For example, within the PPS, one indicator for MinQtSizeY, one indicator for MaxMttDepth, one indicator for MaxBtSizeY, and one indicator for MaxTtSizeY are signaled. In this case, the adjustability of the trade-off between partition complexity and coding efficiency from the partition is individual for each different picture.
[0305] The PPS-level signaling of partition constraint syntax elements can also be signaled within two groups based on the slice type. For example, within the PPS, one intra-slice indicator for MinQtSizeY, one inter-slice indicator for MinQtSizeY, one intra-slice indicator for MaxMttDepth, one inter-indicator for MaxMttDepth, one intra-slice indicator for MaxBtSizeY, one inter-slice indicator for MaxBtSizeY, one intra-slice indicator for MaxTtSizeY, and one inter-slice indicator for MaxTtSizeY are signaled. In this case, the adjustability of the trade-off between the partition complexity and the coding efficiency from the partition is individual for each slice type (intra or inter).
[0306] The PPS-level signaling of partition constraint syntax elements can be signaled within multiple groups based on the slice identification (e.g., index). For example, when one picture is divided into three slices, within the PPS, three different indicators based on the slice identification of MinQtSizeY, three different indicators based on the slice identification of MaxMttDepth, three different indicators based on the slice identification of MaxBtSizeY, and three different indicators based on the slice identification of MaxTtSizeY are signaled. In this case, the adjustability of the trade-off between the partition complexity and the coding efficiency from the partition is individual for each slice.
[0307] Compared with the method in Embodiment 1, the advantage of this alternative implementation is that the indication structure is simplified. In this method, in order to flexibly adjust the trade-off between the partition complexity and the coding gain from the partition, it is not necessary to override the partition constraint syntax elements in the slice header.
[0308] On the other hand, compared with the method in Embodiment 1, this alternative implementation is limited in some scenarios. This method only signals partition constraints in the PPS. This means that when multiple pictures refer to the same PPS, the trade-off between partition complexity and coding gain from the partition cannot be adjusted individually for each picture. Further, if the adjustment is only required for the primary picture, this method will signal redundant information in the PPS.
[0309] Multiple partition constraint syntax elements are signaled at the parameter set level (such as PPS, VPS, SPS) or in headers (such as picture header, slice header, or tile header).
[0310] In Embodiment 2 of the present disclosure: The embodiment means the following: · The partition high-level syntax constraint element can be signaled in the SPS. · The partition high-level syntax constraint element can be overridden in the slice header. · The partition high-level syntax constraint element can use a default value. · BT and TT can be disabled in the SPS. · BT and TT can be disabled in the slice header. · The BT and TT activation (deactivation) flag is signaled in the SPS and can be overridden in the slice header.
[0311] Technical advantages (e.g., signaling in the SPS, overriding in the slice header): High-level partition constraints control the trade-off between partition complexity and coding efficiency from the partition. The present invention guarantees the flexibility to control the trade-off for individual slices. There is more flexibility to control the elements for the default value and the BtTt activation (deactivation function).
[0312] Both the encoder and the decoder perform the same (corresponding) operations.
[0313] The corresponding changes based on the prior art are shown below:
[0314] Changed sequence parameter set RBSP syntax ([JVET-K1001-v4], Section 7.3.2.1) [Ed.(BB): Preliminary basic SPS, subject to further study and under further specification development.] [Table 5]
[0315] Changed slice header syntax ([JVET-K1001-v4], Section 7.3.3) [Ed.(BB): Preliminary basic slice header, subject to further study and under further specification development.] [Table 6]
[0316] Changed sequence parameter set RBSP semantics ([JVET-K1001-v4], Section 7.4.3.1)
[0317] A partition_constraint_control_present_flag equal to 1 specifies the presence of a partition constraint control syntax element in the SPS. A partition_constraint_control_present_flag equal to 0 specifies the absence of a partition constraint control syntax element in the SPS.
[0318] The sps_btt_enabled_flag equal to 1 specifies that the operation of the multi-type tree partition is applied to slices that refer to an SPS without the slice_btt_enable_flag. The sps_btt_enabled_flag equal to 0 specifies that the operation of the multi-type tree partition is not applied to slices that refer to an SPS without the slice_btt_enable_flag. When it does not exist, the value of the sps_btt_enabled_flag is presumed to be equal to 1.
[0319] The partition_constraint_override_enabled_flag equal to 1 specifies the existence of the partition_constraint_override_flag in the slice header for slices that refer to an SPS. The partition_constraint_override_enabled_flag equal to 0 specifies the non-existence of the partition_constraint_override_flag in the slice header for slices that refer to an SPS. When it does not exist, the value of the partition_constraint_override_enabled_flag is presumed to be equal to 0.
[0320] sps_log2_min_qt_size_intra_slices_minus2 + 2 specifies the initial value of the minimum luma size of leaf blocks resulting from the quadtree partitioning of CTUs in slices with a slice_type equal to 2(I) that refer to an SPS, unless overridden by the minimum luma size of leaf blocks resulting from the quadtree partitioning of CTUs in the slice header of the slice that refers to the SPS. The value of log2_min_qt_size_intra_slices_minus2 should be included in the range from 0 to CtbLog2SizeY - 2, inclusive. When it does not exist, the value of sps_log2_min_qt_size_intra_slices_minus2 is presumed to be equal to 0.
Number
[0321] sps_log2_min_qt_size_inter_slices_minus2 + 2 specifies the initial value of the minimum luma size of the leaf blocks resulting from the quadtree partitioning of the CTU in the SPS for slices with slice_type equal to 0(B) or 1(P) that reference the SPS, provided that it is not overridden by the minimum luma size of the leaf blocks resulting from the quadtree partitioning of the CTU in the slice header of the slice that references the SPS. The value of log2_min_qt_size_inter_slices_minus2 should be included in the range that includes both ends of 0 to CtbLog2SizeY - 2. When it does not exist, the value of sps_log2_min_qt_size_inter_slices_minus2 is assumed to be equal to 0.
Number
[0322] sps_max_mtt_hierarchy_depth_inter_slices specifies the initial value of the maximum hierarchical structure depth of the coding units resulting from the multi-type tree partitioning of the quadtree leaf in the SPS for slices with slice_type equal to 0(B) or 1(P) that reference the SPS, provided that it is not overridden by the maximum hierarchical structure depth of the coding units resulting from the multi-type tree partitioning of the quadtree leaf in the slice header of the slice that references the SPS. The value of max_mtt_hierarchy_depth_inter_slices should be included in the range that includes both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, when sps_btt_enabled_flag is equal to 1, The value of sps_max_mtt_hierarchy_depth_inter_slices is presumed to be equal to 3. In other cases, The value of sps_max_mtt_hierarchy_depth_inter_slices is presumed to be equal to 0.
[0323] sps_max_mtt_hierarchy_depth_intra_slices specifies the initial value of the maximum hierarchical structure depth of coding units resulting from the multi-type tree partition of a quadtree leaf, provided that it is not overridden by the maximum hierarchical structure depth of coding units resulting from the multi-type tree partition of a quadtree leaf in the slice header of a slice that refers to the SPS and has a slice_type equal to 2(I) that refers to the SPS. The value of max_mtt_hierarchy_depth_intra_slices should be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, When sps_btt_enabled_flag is equal to 1, The value of sps_max_mtt_hierarchy_depth_intra_slices is presumed to be equal to 3. In other cases, The value of sps_max_mtt_hierarchy_depth_intra_slices is presumed to be equal to 0.
[0324] The initial value of the difference between the luma CTB size and the maximum luma size (width or height) in the SPS of the coded blocks that can be partitioned using 2-way partitioning specifies the difference between the luma CTB size and the maximum luma size (width or height) in the SPS of the coded blocks that can be partitioned using 2-way partitioning in a slice having a slice_type equal to 2(I) that references the SPS, unless it is overridden by the difference between the luma CTB size and the maximum luma size (width or height) in the slice header of the slice that references the SPS. The value of log2_diff_ctu_max_bt_size should be included in the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When not present, if sps_btt_enabled_flag is equal to 1, the value of sps_log2_diff_ctu_max_bt_size_intra_slices is assumed to be equal to 2. otherwise, the value of sps_log2_diff_ctu_max_bt_size_intra_slices is assumed to be equal to CtbLog2SizeY - MinCbLog2SizeY.
[0325] sps_log2_diff_ctu_max_bt_size_inter_slices specifies the initial value of the difference between the luma CTB size and the maximum luma size (width or height) in the SPS of the coded blocks that can be partitioned using 2-way partitioning, provided that the initial value is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) in the SPS of the coded blocks that can be partitioned using 2-way partitioning in the slice header of the slice that refers to the SPS and has a slice_type equal to 0(B) or 1(P). The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When it does not exist, when sps_btt_enabled_flag is equal to 1, the value of sps_log2_diff_ctu_max_bt_size_inter_slices is presumed to be equal to 0. In other cases, the value of sps_log2_diff_ctu_max_bt_size_inter_slices is presumed to be equal to CtbLog2SizeY - MinCbLog2SizeY.
[0326] Changed slice header semantics ([JVET-K1001-v4], Section 7.4.4) partition_constraint_override_flag equal to 1 specifies that the partition constraint parameter exists in the slice header. partition_constraint_override_flag equal to 0 specifies that the partition constraint parameter does not exist in the slice header. When it does not exist, the value of partition_constraints_override_flag is presumed to be equal to 0.
[0327] A slice_btt_enabled_flag equal to 1 specifies that the operation of the multi-type tree partition is applied to the current slice. A slice_btt_enabled_flag equal to 0 specifies that the operation of the multi-type tree partition is not applied to the current slice. When the slice_btt_enabled_flag does not exist, it is assumed to be equal to the sps_btt_enabled_flag.
[0328] log2_min_qt_size_minus2 + 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree splitting of the CTU for the current slice. The value of log2_min_qt_size_inter_slices_minus2 should be within the range including both ends of 0 to CtbLog2SizeY - 2. When it does not exist, the value of log2_min_qt_size_minus2 is assumed to be equal to sps_log2_min_qt_size_intra_slices_minus2 for a slice_type equal to 2(I), and equal to sps_log2_min_qt_size_inter_slices_minus2 for a slice_type equal to 0(B) or 1(P).
[0329] max_mtt_hierarchy_depth specifies the maximum hierarchical structure depth of the coding units resulting from the multi-type tree splitting of the quadtree leaf for the current slice. The value of max_mtt_hierarchy_depth_intra_slices should be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, the value of max_mtt_hierarchy_depth is assumed to be equal to sps_max_mtt_hierarchy_depth_intra_slices for a slice_type equal to 2(I), and equal to sps_max_mtt_hierarchy_depth_inter_slices for a slice_type equal to 0(B) or 1(P).
[0330] log2_diff_ctu_max_bt_size specifies the difference between the luma CTB size and the maximum luma size (width or height) of the coding block that can be divided using 2-way partitioning for the current slice. The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When it does not exist, the value of log2_diff_ctu_max_bt_size is presumed to be equal to sps_log2_diff_ctu_max_bt_size_intra_slices for slice_type equal to 2(I), and presumed to be equal to sps_log2_diff_ctu_max_bt_size_inter_slices for slice_type equal to 0(B) or 1(P).
[0331] The variables MinQtLog2SizeY, MaxBtLog2SizeY, MinBtLog2SizeY, MaxTtLog2SizeY, MinTtLog2SizeY, MaxBtSizeY, MinBtSizeY, MaxTtSizeY, MinTtSizeY, and MaxMttDepth are derived as follows. [Number]
[0332] In Embodiment 3 of the present disclosure: When MaxTTSizeY (the maximum luma size (width or height) of the coding block that can be divided using 3-way partitioning) is signaled in the SPS (or other parameter set or slice header), Embodiment 1 or Embodiment 2 can be applied in the same way as for the above partition parameters.
[0333] Technical Advantage: The present invention that signals an indicator of the MaxTtSizeY syntax element ensures that there is more flexibility in controlling the element.
[0334] Both the encoder and the decoder perform the same (corresponding) operations.
[0335] The syntax change is based on Embodiment 1 or Embodiment 2.
[0336] Changed sequence parameter set RBSP syntax (Section 7.3.2.1 of [JVET-K1001-v4]) [Table 7]
[0337] Changed slice header syntax (Section 7.3.3 of [JVET-K1001-v4]) [Ed.(BB): Preliminary basic slice header, subject to further study and further specification development in progress.] [Table 8]
[0338] In Embodiment 4 of the present disclosure: The btt_enabled_flag in Embodiment 2 is split into bt_enalbed_flag and tt_eabled_flag, and the bt and tt splits are enabled or disabled separately.
[0339] Technical advantage: Signaling the BT enable flag and the TT enable flag separately provides additional flexibility in controlling the partition constraint syntax element.
[0340] Both the encoder and the decoder perform the same (corresponding) operations.
[0341] The syntax and semantics change based on Embodiment 2.
[0342] Changed sequence parameter set RBSP syntax (Section 7.3.2.1 of [JVET-K1001-v4]) [Ed.(BB): Preliminary basic SPS, subject to further study and further specification development in progress.]
Table 9
[0343] Changed slice header syntax ([JVET-K1001-v4] Section 7.3.3) [Ed.(BB): Preliminary basic slice header, subject to further study and further specification development.]
Table 10
[0344] Sequence parameter set RBSP semantics ([JVET-K1001-v4] Section 7.4.3.1) A partition_constraint_control_present_flag equal to 1 specifies the presence of a partition constraint control syntax element in the SPS. A partition_constraint_control_present_flag equal to 0 specifies the absence of a partition constraint control syntax element in the SPS.
[0345] A sps_bt_enabled_flag equal to 1 specifies that the binary tree partitioning operation applies to slices that reference an SPS without a slice_bt_enable_flag. A sps_bt_enabled_flag equal to 0 specifies that the binary tree partitioning operation does not apply to slices that reference an SPS without a slice_bt_enable_flag. When absent, the value of sps_bt_enabled_flag is assumed to be equal to 1.
[0346] The sps_tt_enabled_flag equal to 1 specifies that the operation of the ternary tree partition is applied to slices that refer to an SPS without the slice_tt_enable_flag. The sps_tt_enabled_flag equal to 0 specifies that the operation of the ternary tree partition is not applied to slices that refer to an SPS without the slice_tt_enable_flag. When it does not exist, the value of the sps_tt_enabled_flag is presumed to be equal to 1.
[0347] The partition_constraint_override_enabled_flag equal to 1 specifies the existence of the partition_constraint_override_flag in the slice header for slices that refer to an SPS. The partition_constraint_override_enabled_flag equal to 0 specifies the non - existence of the partition_constraint_override_flag in the slice header for slices that refer to an SPS. When it does not exist, the value of the partition_constraint_override_enabled_flag is presumed to be equal to 0.
[0348] sps_log2_min_qt_size_intra_slices_minus2 + 2 specifies the nominal minimum luma size of leaf blocks resulting from the quadtree partitioning of a CTU in slices with slice_type equal to 2(I) that refer to an SPS, unless overridden by the minimum luma size of leaf blocks resulting from the quadtree partitioning of a CTU present in the slice header of the slice. The value of log2_min_qt_size_intra_slices_minus2 should be included in the range from 0 to CtbLog2SizeY - 2, inclusive. When it does not exist, the value of sps_log2_min_qt_size_intra_slices_minus2 is presumed to be equal to 0.
Number
[0349] sps_log2_min_qt_size_inter_slices_minus2 specifies the specified minimum luma size of a leaf block resulting from the quadtree partitioning of a CTU, provided that it is not overridden by the minimum luma size of a leaf block resulting from the quadtree partitioning of a CTU present in the slice header of a slice that has a slice_type equal to 0 (B) or 1 (P) and that references the SPS. The value of log2_min_qt_size_inter_slices_minus2 should be within the range including both ends of 0 to CtbLog2SizeY - 2. When it does not exist, the value of sps_log2_min_qt_size_inter_slices_minus2 is assumed to be 0.
Number
[0350] sps_max_mtt_hierarchy_depth_inter_slices specifies the specified maximum hierarchical structure depth of a coding unit resulting from the multi-type tree partitioning of a quadtree leaf, provided that it is not overridden by the maximum hierarchical structure depth of a coding unit resulting from the multi-type tree partitioning of a quadtree leaf present in the slice header of a slice that has a slice_type equal to 0 (B) or 1 (P) and that references the SPS. The value of max_mtt_hierarchy_depth_inter_slices should be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, when sps_bt_enabled_flag is equal to 1, or when sps_tt_enabled_flag is equal to 1, The value of sps_max_mtt_hierarchy_depth_inter_slices is presumed to be equal to 3. In other cases, The value of sps_max_mtt_hierarchy_depth_inter_slices is presumed to be equal to 0.
[0351] sps_max_mtt_hierarchy_depth_intra_slices specifies the specified maximum hierarchical structure depth of the coding unit resulting from the multi-type tree partition of the quadtree leaf, provided that it is not overridden by the maximum hierarchical structure depth of the coding unit resulting from the multi-type tree partition of the quadtree leaf existing in the slice header of the slice referring to the SPS, for a slice having a slice_type equal to 2(I) that refers to the SPS. The value of max_mtt_hierarchy_depth_intra_slices should be included in the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, When sps_btt_enabled_flag is equal to 1 and sps_tt_enabled_flag is equal to 1, The value of sps_max_mtt_hierarchy_depth_intra_slices is presumed to be equal to 3. In other cases, The value of sps_max_mtt_hierarchy_depth_intra_slices is presumed to be equal to 0.
[0352] sps_log2_diff_ctu_max_bt_size_intra_slices specifies the defined difference between the luma CTB size and the maximum luma size (width or height) of the coding blocks that can be partitioned using 2-way partitioning, provided that the defined difference between the luma CTB size and the maximum luma size (width or height) of the coding blocks that can be partitioned using 2-way partitioning in the slice header of the slice that refers to the luma CTB size and the SPS is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) of the coding blocks that can be partitioned using 2-way partitioning in the slice having a slice_type equal to 2(I) that refers to the SPS. The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When not present, if sps_bt_enabled_flag is equal to 1, the value of sps_log2_diff_ctu_max_bt_size_intra_slices is presumed to be equal to 2. In other cases, the value of sps_log2_diff_ctu_max_bt_size_intra_slices is presumed to be equal to CtbLog2SizeY - MinCbLog2SizeY.
[0353] sps_log2_diff_ctu_max_bt_size_inter_slices specifies the defined difference between the luma CTB size and the maximum luma size (width or height) of the coded blocks that can be partitioned using 2-way partitioning, provided that the defined difference is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) of the coded blocks that can be partitioned using 2-way partitioning in the slice header of the slice that refers to the SPS and has a slice_type equal to 0(B) or 1(P). The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When not present, if sps_bt_enabled_flag is equal to 1, the value of sps_log2_diff_ctu_max_bt_size_inter_slices is presumed to be 0. In other cases, the value of sps_log2_diff_ctu_max_bt_size_inter_slices is presumed to be equal to CtbLog2SizeY - MinCbLog2SizeY.
[0354] sps_log2_diff_ctu_max_tt_size_intra_slices specifies the defined difference between the luma CTB size and the maximum luma size (width or height) of a codable block that can be partitioned using three partitions, provided that the defined difference between the luma CTB size and the maximum luma size (width or height) of a codable block that can be partitioned using three partitions in the slice header of a slice that refers to the SPS and has a slice_type equal to 2(I) is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) of a codable block that can be partitioned using three partitions in the slice. The value of sps_log2_diff_ctu_max_tt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When not present, if sps_tt_enabled_flag is equal to 1, the value of sps_log2_diff_ctu_max_tt_size_intra_slices is presumed to be equal to 2. In other cases, the value of sps_log2_diff_ctu_max_tt_size_intra_slices is presumed to be equal to CtbLog2SizeY - MinCbLog2SizeY.
[0355] sps_log2_diff_ctu_max_tt_size_inter_slices specifies the defined difference between the luma CTB size and the maximum luma size (width or height) of the coded blocks that can be partitioned using 3 partitions, provided that the defined difference between the luma CTB size and the maximum luma size (width or height) of the coded blocks that can be partitioned using 3 partitions and that is present in the slice header of the slice referring to the SPS is not overridden by the difference between the luma CTB size and the maximum luma size (width or height) of the coded blocks that can be partitioned using 3 partitions in a slice having a slice_type equal to 0(B) or 1(P) that refers to the SPS. The value of log2_diff_ctu_max_tt_size shall be included in the range that includes both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When not present, if sps_tt_enabled_flag is equal to 1, the value of sps_log2_diff_ctu_max_tt_size_inter_slices is presumed to be equal to 1. In other cases, the value of sps_log2_diff_ctu_max_tt_size_inter_slices is presumed to be equal to CtbLog2SizeY - MinCbLog2SizeY.
[0356] Changed slice header semantics ([JVET-K1001-v4], Section 7.4.4) partition_constraint_override_flag equal to 1 specifies that the partition constraint parameter is present in the slice header. partition_constraint_override_flag equal to 0 specifies that the partition constraint parameter is not present in the slice header. When not present, the value of partition_constraints_override_flag is presumed to be equal to 0.
[0357] A slice_btt_enabled_flag equal to 1 specifies that the operation of the multi-type tree partition is not applied to the current slice. A slice_btt_enabled_flag equal to 0 specifies that the operation of the multi-type tree partition is applied to the current slice. When the slice_btt_enabled_flag does not exist, it is assumed to be equal to the sps_btt_enabled_flag.
[0358] log2_min_qt_size_minus2 + 2 specifies the minimum luma size of the leaf blocks resulting from the quadtree partitioning of the CTU for the current slice. The value of log2_min_qt_size_inter_slices_minus2 should be within the range including both ends of 0 to CtbLog2SizeY - 2. When it does not exist, the value of log2_min_qt_size_minus2 is assumed to be equal to sps_log2_min_qt_size_intra_slices_minus2 for a slice_type equal to 2(I), and equal to sps_log2_min_qt_size_inter_slices_minus2 for a slice_type equal to 0(B) or 1(P).
[0359] max_mtt_hierarchy_depth specifies the maximum hierarchical structure depth of the coding units resulting from the multi-type tree partitioning of the quadtree leaf for the current slice. The value of max_mtt_hierarchy_depth_intra_slices should be within the range including both ends of 0 to CtbLog2SizeY - MinTbLog2SizeY. When it does not exist, the value of max_mtt_hierarchy_depth is assumed to be equal to sps_max_mtt_hierarchy_depth_intra_slices for a slice_type equal to 2(I), and equal to sps_max_mtt_hierarchy_depth_inter_slices for a slice_type equal to 0(B) or 1(P).
[0360] log2_diff_ctu_max_bt_size specifies the difference between the luma CTB size and the maximum luma size (width or height) of the coded block that can be divided using 2-way partitioning for the current slice. The value of log2_diff_ctu_max_bt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When it does not exist, the value of log2_diff_ctu_max_bt_size is presumed to be equal to sps_log2_diff_ctu_max_bt_size_intra_slices for slice_type equal to 2(I), and presumed to be equal to sps_log2_diff_ctu_max_bt_size_inter_slices for slice_type equal to 0(B) or 1(P).
[0361] log2_diff_ctu_max_tt_size specifies the difference between the luma CTB size and the maximum luma size (width or height) of the coded block that can be divided using 2-way partitioning for the current slice. The value of log2_diff_ctu_max_tt_size should be within the range including both ends of 0 to CtbLog2SizeY - MinCbLog2SizeY. When it does not exist, the value of log2_diff_ctu_max_tt_size is presumed to be equal to sps_log2_diff_ctu_max_tt_size_intra_slices for slice_type equal to 2(I), and presumed to be equal to sps_log2_diff_ctu_max_tt_size_inter_slices for slice_type equal to 0(B) or 1(P).
[0362] The variables MinQtLog2SizeY, MaxBtLog2SizeY, MinBtLog2SizeY, MaxTtLog2SizeY, MinTtLog2SizeY, MaxBtSizeY, MinBtSizeY, MaxTtSizeY, MinTtSizeY, and MaxMttDepth are derived as follows.
Number
[0363] FIG. 10 shows a corresponding method for decoding a video bit stream implemented by a decoding apparatus, the video bit stream including data representing an image area and an image area header of the image area. The decoding method includes a step S110 of obtaining a partition_constraint_override_flag from the video bit stream, and a step S120 of obtaining first partition constraint information of the image area from the image area header when the value of the override flag is an override value (e.g., 1), and a step S130 of partitioning blocks of the image area according to the first partition constraint information. When the flag is not set, the partition constraint information may be obtained from a source different from the image area header. The image area may be a slice or a tile.
[0364] FIG. 11 shows a flowchart incorporating the flowchart of FIG. 10. Further, the method shown in the flowchart includes a step S210 of obtaining a partition_constraint_override_enabled_flag from the video bit stream, and a step S110 of obtaining an override flag from the video bit stream when the value of the override enable flag is an enable value (e.g., 1). Further, when the value of the override flag is not an override value (e.g., the value of the override flag is 0), the step S230 of partitioning blocks of the image area may be performed according to second partition constraint information for the video bit stream from a parameter set. Further, when the value of the override enable flag is a disable value (e.g., the value of the override enable flag is 0), the step S230 of partitioning blocks of the image area may be performed according to second partition constraint information for the video bit stream from a parameter set.
[0365] For specific features in the embodiments of the present invention, refer to the embodiments of the related decoding method described above. Details are not described again here.
[0366] FIG. 12 shows a decoder 1200 for decoding a video bitstream. The video bitstream includes data representing an image region and an image region header of the image region. The decoder includes an override determination unit 1210 that obtains an override flag from the video bitstream, a partition constraint determination unit 1220 that obtains first partition constraint information of the image region from the image region header when the value of the override flag is an override value, and a block partition unit 1230 that partitions the blocks of the image region according to the first partition constraint information.
[0367] For specific functions of the units in the decoder 1200 in the embodiments of the present invention, refer to the related descriptions of the embodiments of the decoding method of the present invention. Details are not described again here.
[0368] The units in the decoder 1200 may be implemented by software or circuits.
[0369] The decoder 1200 may be a decoder 30, a video encoding device 400, or a device 500, or a part of the decoder 30, the video encoding device 400, or the device 500.
[0370] The encoder 1300 may be an encoder 20, a video encoding device 400, or a device 500, or a part of the encoder 20, the video encoding device 400, or the device 500.
[0371] FIG. 13 shows an encoder 1300 that encodes a video bitstream, the video bitstream including data representing an image area and an image area header for the image area. The encoder includes a partition determination unit 1310 that determines whether the partition of a block of the image area conforms to first partition constraint information in the image area header, a block partition unit 1320 that partitions the block of the image area according to the first partition constraint information when it is determined that the partition of the block conforms to the first partition constraint information, an override flag setting unit 1330 that sets the value of the override flag to an override value, and a bitstream generator 1340 that inserts the override flag into the video bitstream.
[0372] For the specific functions of the units in the encoder 1300 in the embodiments of the present invention, refer to the related descriptions of the embodiments of the encoding method of the present invention. Details are not described here again.
[0373] The units in the encoder 1300 may be implemented by software or circuitry.
[0374] The encoder 1300 may be the encoder 20, the video encoding device 400, or the apparatus 500, or a part of the encoder 20, the video encoding device 400, or the apparatus 500.
[0375] FIG. 14A shows a flowchart of a method for encoding a video bitstream implemented by an encoding device, the video bitstream including data representing an image region and an image region header of the image region. The encoding method includes step S310 of determining whether the partitioning of a block of the image region conforms to first partitioning constraint information in the image region header, and when it is determined that the partitioning of the block conforms to the first partitioning constraint information (YES in step S310), step S320 of partitioning the block of the image region according to the first partitioning constraint information, step S325 of setting the value of an override flag to an override value, and step S330 of including the data of the override flag in the video bitstream.
[0376] In some exemplary embodiments, when it is determined that partitioning the block does not conform to the first partitioning constraint information (NO in step S310), the block of the image region is partitioned according to second partitioning constraint information in S360, and the value of the override flag is set to the override value in S365.
[0377] FIG. 14B shows an encoding method including step S370 of determining whether partitioning the block according to the first partitioning constraint information is enabled. When it is determined that partitioning the block according to the first partitioning constraint information is enabled (should be determined), the method includes step S340 of setting the value of an override enable flag to an enable value, and step S350 of including the data of the override enable flag in the video bitstream. Further, when it is determined that partitioning the block according to the first partitioning constraint information is enabled (should be determined), S310 determines whether partitioning the block of the image region conforms to the first partitioning constraint information in the image region header.
[0378] In some exemplary embodiments, when it is determined that partitioning the blocks according to the first partition constraint information should not be (i.e., should be disabled) enabled, the method includes step S380 of setting the value of the override enable flag to a disabled value.
[0379] For certain features in embodiments of the present invention, refer to the related decoding method embodiments described above. Details are not described again here.
[0380] The following is an explanation of the application of the encoding method, decoding method, and system using them as shown in the above embodiments.
[0381] FIG. 14 is a block diagram showing a content supply system 3100 that realizes a content delivery service. This content supply system 3100 includes a capture device 3102, a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types, etc.
[0382] The capture device 3102 may generate data and encode the data by an encoding method as shown in the above-described embodiments. Alternatively, the capture device 3102 may distribute the data to a streaming server (not shown in the figure), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a Pad, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof, etc. For example, the capture device 3102 may include the source device 102 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. In some practical scenarios, the capture device 3102 distributes the encoded video and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0383] In the content supply system 3100, the terminal device 310 receives and plays the encoded data. The terminal device 3106 can be a device with data reception and restoration capabilities, such as a smartphone or Pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0384] In a terminal device having its own display, such as a smartphone, or Pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video decoder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its own display. In a terminal device without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is connected thereto to receive and display the decoded data.
[0385] When each device in this system performs encoding or decoding, as shown in the above-described embodiments, a picture encoding device or a picture decoding device can be used.
[0386] FIG. 15 is a diagram showing the structure of an example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, the Real Time Streaming Protocol (RTSP), the Hyper Text Transfer Protocol (HTTP), the HTTP Live streaming protocol (HLS), MPEG-DASH, the Real-time Transport protocol (RTP), the Real Time Messaging Protocol (RTMP), or any combination of those of any kind, etc.
[0387] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0388] By inverse multiplexing processing, a video elementary stream (ES), an audio ES, and any alternatives are generated. A video decoder 3206 including a video decoder 30 as described in the above embodiments decodes the video ES by a decoding method as shown in the above embodiments to generate a video frame, and supplies this data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate an audio frame, and supplies this data to the synchronization unit 3212. Alternatively, the video frame may be stored in a buffer (not shown in FIG. 15) before being supplied to the synchronization unit 3212. Similarly, the audio frame may be stored in a buffer (not shown in FIG. 15) before being supplied to the synchronization unit 3212.
[0389] The synchronization unit 3212 synchronizes the video frame and the audio frame, and supplies the video / audio to a video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be encoded in the syntax using time stamps related to the presentation of the encoded audio and visual data, and time stamps related to the delivery of the data stream itself.
[0390] When subtitles are included in the stream, a subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frame and the audio frame, and supplies the video / audio / subtitle to a video / audio / subtitle display 3216.
[0391] The present invention is not limited to the above-described system, and any of the picture encoding device or the picture decoding device in the above embodiments can be incorporated into other systems, such as vehicle systems.
[0392] Mathematical operator The mathematical operators used in this application are the same as those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operators such as exponentiation and division of real values are defined. The numbering and counting rules generally start from 0. For example, "the first" is equivalent to the 0th, "the second" is equivalent to the 1st, and so on.
[0393] Logical Operators The following logical operators are defined as follows.
Number
[0394] Relational Operators The following relational operators are defined as follows.
Number
[0395] When a relational operator is applied to a variable assigned the syntax element or value "na" (not applicable), the value "na" is treated as an individual value of the syntax element or variable. The value "na" is considered not equal to any other value.
[0396] Bitwise Operators The following bitwise operators are defined as follows.
[0397] & Bitwise "logical product". When acting on integer arguments, it acts on the two's complement integer representation. When acting on a two-valued argument containing fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0.
[0398] | Bitwise "logical sum". When acting on integer arguments, it acts on the two's complement integer representation. When acting on a two-valued argument containing fewer bits than the other argument, the shorter argument is extended by adding higher-order bits equal to 0.
[0399] ^ Exclusive OR per bit. When acting on an integer argument, it acts on the complete binary representation of the integer value. When acting on a binary argument that contains fewer bits than another argument, the shorter argument is extended by adding higher-order bits equal to 0.
[0400] x >> y Arithmetic right shift of the two's complement integer representation of x × y in binary. This function is defined only for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has a value equal to the MSB of x before the shift operation.
[0401] x << y Arithmetic left shift of the two's complement integer representation of x × y in binary. This function is defined only for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has a value equal to 0.
[0402] In one or more examples, the functions described may be implemented by hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a processing unit based on hardware. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, the computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to read instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0403] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be accessed by a computer and is capable of storing the required program code in the form of instructions or data structures. Also, any connection can be properly called a computer-readable medium. When instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transient tangible storage media. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data and disc optically reproduces data by laser. The foregoing combinations should also be included within the scope of computer-readable media.
[0404] The commands may be executed by one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the term "processor" as used herein may represent any of the foregoing structures or any other structure suitable for implementation of the techniques described herein. Further, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured to be incorporated into an encoder and decoder or combined codec. Further, the techniques may be implemented entirely with one or more circuits or logic elements.
[0405] The techniques of the present disclosure may be implemented in a variety of apparatus or devices including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). Although various components, modules, or units have been described in the present disclosure to emphasize functional aspects of an apparatus configured to execute the disclosed techniques, implementation by different hardware units is not necessarily required. Rather, as described above, the various units may be combined within a codec hardware unit in combination with appropriate software and / or firmware, or provided by a collection of interoperable hardware units including one or more processors as described above.
[0406] In one example, there is a method of encoding performed by an encoding device, the method including: partitioning blocks of an image region according to partition constraint information; and generating a bitstream including one or more partition constraint syntax elements, where the one or more partition constraint syntax elements indicate the partition constraint information and the one or more partition constraint syntax elements are signaled at a picture parameter set (PPS) level.
[0407] For example, the partition constraint information includes one or more selected from the following: information on a minimum allowable quadtree leaf node size (MinQtSizeY), information on a maximum multi-type tree depth (MaxMttDepth), information on a maximum allowable binary tree root node size (MaxBtSizeY), and information on a maximum allowable ternary tree root node size (MaxTtSizeY).
[0408] For example, in some embodiments, the partition constraint information includes information on a minimum allowable quadtree leaf node size (MinQtSizeY), information on a maximum multi-type tree depth (MaxMttDepth), and information on a maximum allowable binary tree root node size (MaxBtSizeY).
[0409] In some embodiments, the partition constraint information includes information on a minimum allowable quadtree leaf node size (MinQtSizeY), information on a maximum multi-type tree depth (MaxMttDepth), information on a maximum allowable binary tree root node size (MaxBtSizeY), and information on a maximum allowable ternary tree root node size (MaxTtSizeY).
[0410] In any of the above methods, the partition constraint information includes: N sets or groups of partition constraint information corresponding to N slice types, or N sets or groups of partition constraint information corresponding to N slice indexes, where each set or group of partition constraint information includes one or more selected from the following: information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY), and N is a positive integer.
[0411] The method may include the steps of partitioning blocks of an image region according to partition constraint information, and generating a bitstream including a plurality of partition constraint syntax elements, where the plurality of partition constraint syntax elements indicate the partition constraint information and the plurality of partition constraint syntax elements are signaled at a parameter set level or in a header.
[0412] For example, the plurality of partition constraint syntax elements are signaled at any one of a video parameter set (VPS) level, a sequence parameter set (SPS) level, a picture parameter set (PPS) level, a picture header, a slice header, or a tile header.
[0413] In some exemplary implementations, the partition constraint information includes information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), and information on the maximum allowable binary tree root node size (MaxBtSizeY).
[0414] For example, the partition constraint information includes information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY).
[0415] In some embodiments, the partition constraint information includes two or more selected from the following: information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY).
[0416] For example, the partition constraint information includes N sets or groups of partition constraint information corresponding to N slice types, or N sets or groups of partition constraint information corresponding to N slice indexes. Here, each set or group of partition constraint information includes two or more selected from the following: information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY), and N is a positive integer.
[0417] According to an embodiment, there is provided a decoding method implemented by a decoding apparatus, the method including: parsing, from a bitstream, one or more partition constraint syntax elements, where the one or more partition constraint syntax elements indicate partition constraint information and are obtained from a picture parameter set (PPS) level of the bitstream; and partitioning blocks of an image region according to the partition constraint information.
[0418] In some implementations, the partition constraint information includes one or more selected from the following: information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY).
[0419] For example, the partition constraint information includes information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), and information on the maximum allowable binary tree root node size (MaxBtSizeY). The partition constraint information may include information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY).
[0420] In some embodiments, the partition constraint information includes N sets or groups of partition constraint information corresponding to N slice types, or N sets or groups of partition constraint information corresponding to N slice indexes. Here, each set or group of partition constraint information includes one or more selected from the following: information on the minimum allowable quadtree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY), and N is a positive integer.
[0421] According to an embodiment, a decoding method implemented by a decoding apparatus includes: parsing a plurality of partition constraint syntax elements from a bitstream, where the plurality of partition constraint syntax elements indicate partition constraint information and are obtained from a parameter set level of the bitstream or a header of the bitstream; and partitioning blocks of an image region according to the partition constraint information.
[0422] For example, the plurality of partition constraint syntax elements are obtained from any one of a video parameter set (VPS) level, a sequence parameter set (SPS) level, a picture parameter set (PPS) level, a picture header, a slice header, or a tile header.
[0423] For example, the partition constraint information includes information on a minimum allowable quadtree leaf node size (MinQtSizeY), information on a maximum multi-type tree depth (MaxMttDepth), and information on a maximum allowable binary tree root node size (MaxBtSizeY).
[0424] In some embodiments, the partition constraint information includes information on a minimum allowable quadtree leaf node size (MinQtSizeY), information on a maximum multi-type tree depth (MaxMttDepth), information on a maximum allowable binary tree root node size (MaxBtSizeY), and information on a maximum allowable ternary tree root node size (MaxTtSizeY).
[0425] In some implementations, the partition constraint information includes two or more selected from the following: information on the minimum allowable quad tree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY).
[0426] For example, the partition constraint information includes N sets or groups of partition constraint information corresponding to N slice types, or N sets or groups of partition constraint information corresponding to N slice indexes. Here, each set or group of partition constraint information includes two or more selected from the following: information on the minimum allowable quad tree leaf node size (MinQtSizeY), information on the maximum multi-type tree depth (MaxMttDepth), information on the maximum allowable binary tree root node size (MaxBtSizeY), and information on the maximum allowable ternary tree root node size (MaxTtSizeY), and N is a positive integer.
[0427] In some examples, the partition constraint information includes partition constraint information corresponding to different slice types or different slice indexes.
[0428] For example, the partition constraint information includes partition constraint information in the intra mode and / or partition constraint information in the inter mode.
[0429] In any of the embodiments, the image region includes a picture or a part of a picture.
[0430] In some embodiments, when the value of the multi-type tree partition activation flag from the picture parameter set (PPS) activates the multi-type tree partition of a block, the partition constraint information from the picture parameter set is parsed, and the multi-type tree partition is applied to the blocks of the image region according to the partition constraint information.
[0431] According to an embodiment, an encoder is provided that includes a processing circuit for performing any of the above-described methods.
[0432] According to an embodiment, a decoder is provided that includes a processing circuit for performing any of the above-described methods.
[0433] According to an embodiment, a computer program product is provided that includes program code for performing a method according to any of the above-described methods.
[0434] According to an embodiment, a decoder is provided that includes one or more processors and a non-transitory computer-readable storage medium connected to the processor and storing programming for execution by the processor, the programming configuring the decoder to perform a method according to any of the above-described decoding methods when executed by the processor.
[0435] According to an embodiment, an encoder is provided that includes one or more processors and a non-transitory computer-readable storage medium connected to the processor and storing programming for execution by the processor, the programming configuring the encoder to perform a method according to any of the above-described decoding methods when executed by the processor.
[0436] According to a first aspect, the present invention is a method for decoding a video bitstream implemented by a decoding device, wherein the video bitstream includes data representing an image region and an image region header of the image region, and the decoding method includes: obtaining an override flag from the video bitstream; when the value of the override flag (e.g., partition_constraint_override_flag) is an override value, obtaining first partition constraint information of the image region from the image region header; partitioning blocks of the image region according to the first partition constraint information. The method relates to a method including the above steps.
[0437] In a possible implementation, the step of partitioning blocks of the image region according to the first partition constraint information includes partitioning blocks of the image region into sub-blocks according to the first partition constraint information. The decoding method further includes the step of reconstructing the sub-blocks.
[0438] In a possible implementation, the decoding method includes: further obtaining an override enable flag from the video bitstream, and obtaining the override flag from the video bitstream when the value of the override enable flag (e.g., partition_constraint_override_enabled_flag) is an enable value.
[0439] In a possible implementation, the decoding method includes: further obtaining a partition constraint control presence flag from the video bitstream. When the value of the partition constraint control presence flag (e.g., partition_constraint_control_present_flag) is true, obtain the override enable flag from the video bitstream.
[0440] In a possible implementation, the video bitstream further includes data representing a parameter set of the video bitstream, and the fact that the value of the partition constraint control presence flag is false specifies the absence of a partition constraint control syntax element in the parameter set.
[0441] In a possible implementation, the parameter set is a picture parameter set or a sequence parameter set.
[0442] In a possible implementation, the video bitstream further includes data representing a parameter set of the video bitstream, and the decoding method when the value of the override enable flag is a disabling value, further includes the step of partitioning the blocks of the image region according to second partition constraint information of the video bitstream from the parameter set.
[0443] In a possible implementation, the second partition constraint information includes information on a minimum allowable quadtree leaf node size, information on a maximum multi-type tree depth, information on a maximum allowable ternary tree root node size, or information on a maximum allowable binary tree root node size.
[0444] In a possible implementation, the second partition constraint information includes partition constraint information corresponding to different parameters related to the image region or corresponding to different indexes.
[0445] In a possible implementation, the second partition constraint information includes partition constraint information in an intra mode or partition constraint information in an inter mode.
[0446] In a possible implementation, the second partition constraint information includes partition constraint information of luma blocks or partition constraint information of chroma blocks.
[0447] In a possible implementation, the video bitstream further includes data representing a parameter set of the video bitstream, and the step of obtaining an override enable flag from the video bitstream includes the step of obtaining the override enable flag from the parameter set.
[0448] In a possible implementation, the step of obtaining an override flag from the video bitstream includes the step of obtaining the override flag from the image region header.
[0449] In a possible implementation, the first partition constraint information includes information on a minimum allowable quadtree leaf node size, information on a maximum multi-type tree depth, information on a maximum allowable ternary tree root node size, or information on a maximum allowable binary tree root node size.
[0450] In a possible implementation, the image region includes a slice or a tile, and the image region header includes a slice header of the slice or a tile header of the tile.
[0451] In a possible implementation, the video bitstream further includes data representing a parameter set of the video bitstream, and the decoding method further includes the step of partitioning the blocks of the image region according to the second partition constraint information of the video bitstream from the parameter set when the value of the override flag is not the override value.
[0452] In a possible implementation, when the value of the multi-type tree partitioning enabling flag (e.g., slice_btt_enabled_flag) from the image region header enables the multi-type tree partitioning of the block, the first partition constraint information is obtained, and the multi-type tree partitioning is applied to the block in the image region according to the first partition constraint information.
[0453] In a possible implementation, the video bitstream further includes data representing a parameter set of the video bitstream. When the multi-type tree partitioning enabling flag from the image region header does not exist and the value of the multi-type tree partitioning enabling flag (e.g., sps_btt_enabled_flag) from the parameter set enables the multi-type tree partitioning of the block, the second partition constraint information of the video bitstream is obtained from the parameter set, and the multi-type tree partitioning is applied to the block in the image region according to the second partition constraint information.
[0454] According to a second aspect, the present invention is a method for decoding a video bitstream implemented by a decoding device. The video bitstream includes data representing a block and a first parameter set of the video bitstream. The decoding method includes: obtaining an override flag from the video bitstream; when the value of the override flag is an override value, obtaining the first partition constraint information of the block from the first parameter set; partitioning the block according to the first partition constraint information.
[0455] In a possible implementation, the step of partitioning the block according to the first partition constraint information includes the step of partitioning the block into sub-blocks according to the first partition constraint information. The decoding method further includes the step of reconstructing the sub-blocks.
[0456] In a possible implementation, the decoding method further includes the step of obtaining an override enable flag from the video bitstream, and obtaining the override flag from the video bitstream when the value of the override enable flag is an enable value.
[0457] In a possible implementation, the decoding method further includes the step of obtaining a partition constraint control presence flag from the video bitstream, and obtaining the override enable flag from the video bitstream when the value of the partition constraint control presence flag is true.
[0458] In a possible implementation, the video bitstream further includes data representing a second parameter set of the video bitstream, and the fact that the value of the partition constraint control presence flag is false specifies the absence of a partition constraint control syntax element in the parameter set.
[0459] In a possible implementation, the video bitstream further includes data representing a second parameter set of the video bitstream, and the decoding method further includes the step of partitioning the block according to second partition constraint information of the video bitstream from the second parameter set when the value of the override enable flag is a disable value.
[0460] In a possible implementation, the second partition constraint information includes information on the minimum allowable quadtree leaf node size, information on the maximum multi-type tree depth, information on the maximum allowable ternary tree root node size, or information on the maximum allowable binary tree root node size.
[0461] In a possible implementation, the second partition constraint information includes partition constraint information corresponding to different parameter sets related to the image region represented by the video bitstream or corresponding to different indexes.
[0462] In a possible implementation, the second partition constraint information includes partition constraint information in the intra mode or partition constraint information in the inter mode.
[0463] In a possible implementation, the second partition constraint information includes partition constraint information for luma blocks or partition constraint information for chroma blocks.
[0464] In a possible implementation, the video bitstream further includes data representing a second parameter set of the video bitstream, and the step of obtaining an override enable flag from the video bitstream includes the step of obtaining the override enable flag from the second parameter set.
[0465] In a possible implementation, the step of obtaining an override flag from the video bitstream includes the step of obtaining the override flag from the first parameter set.
[0466] In a possible implementation, the first partition constraint information includes information on the minimum allowable quadtree leaf node size, information on the maximum multi-type tree depth, information on the maximum allowable ternary tree root node size, or information on the maximum allowable binary tree root node size.
[0467] In a possible implementation, the video bitstream further includes data representing a second parameter set of the video bitstream, the first parameter set is a picture parameter set, and the second parameter set is a sequence parameter set.
[0468] In a possible implementation, the video bitstream further includes data representing a second parameter set of the video bitstream, and the decoding method includes when the value of the override flag is not an override value, partitioning the block according to second partition constraint information of the video bitstream from the second parameter set.
[0469] In a possible implementation, when the value of the multi-type tree partition activation flag from the first parameter set activates the multi-type tree partition of the block, obtaining first partition constraint information and applying the multi-type tree partition to the block according to the first partition constraint information.
[0470] In a possible implementation, the video bitstream further includes data representing a second parameter set of the video bitstream, the multi-type tree partition activation flag from the first parameter set does not exist, and when the value of the multi-type tree partition activation flag from the second parameter set activates the multi-type tree partition of the block, obtaining second partition constraint information of the video bitstream from the second parameter set and applying the multi-type tree partition to the block according to the second partition constraint information.
[0471] According to a third aspect, the present invention is a method for decoding a video bitstream implemented by a decoding device, where the video bitstream includes data representing a first image area and a first image area header of the first image area, and the decoding method includes The step of obtaining an override flag from the video bitstream; When the value of the override flag is an override value, the step of obtaining first partition constraint information of the first image region from the first image region header; Partitioning blocks of the first image region according to the first partition constraint information. A method including the above steps is related.
[0472] In a possible implementation, the step of partitioning the blocks of the first image region according to the first partition constraint information includes the step of partitioning the blocks of the first image region into sub-blocks according to the first partition constraint information. The decoding method further includes the step of reconstructing the sub-blocks.
[0473] In a possible implementation, the decoding method further includes the step of obtaining an override enable flag from the video bitstream, and obtaining the override flag from the video bitstream when the value of the override enable flag is an enable value.
[0474] In a possible implementation, the decoding method further includes the step of obtaining a partition constraint control presence flag from the video bitstream, and obtaining the override enable flag from the video bitstream when the value of the partition constraint control presence flag is true.
[0475] In a possible implementation, the video bitstream further includes data representing a second image region and a second image region header of the second image region. The fact that the value of the partition constraint control presence flag is false specifies the absence of a partition constraint control syntax element in the second image region header.
[0476] In a possible implementation, the video bitstream further includes data representing a second image region and a second image region header of the second image region, and the decoding method includes: When the value of the override enable flag is an invalidation value, partitioning the block of the first image region according to second partition constraint information of the video bitstream from the second image region header, where the second image region includes the block of the first image region.
[0477] In a possible implementation, the second partition constraint information includes information on a minimum allowable quadtree leaf node size, information on a maximum multi-type tree depth, information on a maximum allowable ternary tree root node size, or information on a maximum allowable binary tree root node size.
[0478] In a possible implementation, the second partition constraint information includes partition constraint information corresponding to different parameter sets or different indexes related to the image region represented by the video bitstream.
[0479] In a possible implementation, the second partition constraint information includes partition constraint information in the intra mode or partition constraint information in the inter mode.
[0480] In a possible implementation, the second partition constraint information includes partition constraint information for luma blocks or partition constraint information for chroma blocks.
[0481] In a possible implementation, the video bitstream further includes data representing a second image region and a second image region header of the second image region, and the step of obtaining the override enable flag from the video bitstream includes obtaining the override enable flag from the second image region header.
[0482] In a possible implementation, the step of obtaining the override flag from the video bitstream includes the step of obtaining the override flag from the first image region header.
[0483] In a possible implementation, the first partitioning constraint information includes information on a minimum allowable quad-tree leaf node size, information on a maximum multi-type tree depth, information on a maximum allowable ternary tree root node size, or information on a maximum allowable binary tree root node size.
[0484] In a possible implementation, the video bitstream further includes data representing a second image region and a second image region header of the second image region. The first image region header is a slice header, the second image region header is a tile header, the first image region is a slice, the second image region is a tile, and the tile includes the slice, or The first image region header is a tile header, the second image region header is a slice header, the first image region is a tile, the second image region is a slice, and the slice includes the tile.
[0485] In a possible implementation, the video bitstream further includes data representing a second image region and a second image region header of the second image region, and the decoding method is When the value of the override flag is not an override value, partitioning the blocks of the first image region according to the second partitioning constraint information of the video bitstream from the second image region header, wherein the second image region includes the blocks of the first image region.
[0486] In a possible implementation, when the value of the multi-type tree partition activation flag from the first image region header activates the multi-type tree partition of the block, the first partition constraint information is obtained, and the multi-type tree partition is applied to the block in the first image region according to the first partition constraint information.
[0487] In a possible implementation, the video bitstream further includes data representing a second image region header of the video bitstream. When the multi-type tree partition activation flag from the first image region header does not exist and the value of the multi-type tree partition activation flag from the second image region header is Front When activating the multi-type tree partition of the block, the second partition constraint information of the video bitstream is obtained from the second image region header, and the multi-type tree partition is applied to the block in the image region according to the second partition constraint information.
[0488] According to a fourth aspect, the present invention relates to an apparatus for decoding a video stream, including a processor and a memory. The memory stores instructions for causing the processor to execute a method according to the first aspect, the second aspect, or the third aspect, or any possible implementation of the first aspect, the second aspect, or the third aspect.
[0489] According to a fifth aspect, there is proposed a computer-readable storage medium storing instructions which, when executed, cause one or more processors to be configured to encode video data. The instructions cause the one or more processors to execute a method according to the first aspect, the second aspect, or the third aspect, or any possible implementation of the first aspect, the second aspect, or the third aspect.
[0490] According to a sixth aspect, the present invention relates to a computer program including program code for executing, when executed on a computer, a method according to the first aspect, the second aspect, or the third aspect or any possible embodiment of the first aspect, the second aspect, or the third aspect.
[0491] In summary, the present disclosure provides an encoding and decoding apparatus and an encoding and decoding method. In particular, the present disclosure relates to signaling partition parameters in a block partition and a bit stream. An override flag in an image region header indicates whether a block is partitioned according to first partition constraint information. The override flag is included in the bit stream and the block is partitioned accordingly.
Claims
Claim 1 A method for decoding a video bitstream implemented by a decoding device, wherein the video bitstream includes data representing an image region and an image region header of the image region, and the video bitstream further includes a sequence parameter set (SPS) of the video bitstream, and the decoding method includes: obtaining an override enable flag from the video bitstream; when the value of the override enable flag is an enable value, obtaining an override flag from the video bitstream, wherein the override flag indicates whether first partition constraint information from the image region header or second partition constraint information from the SPS should be used to partition blocks of the image region; when the value of the override flag is an override value, obtaining first partition constraint information of the image region from the image region header; partitioning the block according to the first partition constraint information; The decoding method includes the above steps. Claim 2 The decoding method further includes: when the value of the override enable flag is a disable value, partitioning the block of the image region according to the second partition constraint information of the video bitstream from the SPS. The decoding method according to claim 1, further including the above step. Claim 3 The decoding method according to claim 1 or 2, wherein the second partition constraint information includes partition constraint information of a block in an intra mode or partition constraint information of a block in an inter mode. Claim 4 The decoding method according to claim 1 or 2, wherein the second partition constraint information includes partition constraint information of a luma block or partition constraint information of a chroma block. Claim 5 The decoding method according to any one of claims 1 to 4, wherein the step of obtaining the override enable flag from the video bitstream includes obtaining the override enable flag from the SPS. Claim 6 The step of obtaining the override flag from the video bit stream includes the step of obtaining the override flag from the image region header, the decoding method according to any one of claims 1 to 5.
7. The image region includes a picture, and the image region header includes the picture header of the picture, the decoding method according to any one of claims 1 to 6.
8. The decoding method is When the value of the override flag is not the override value, partitioning the block of the image region according to the second partition constraint information of the video bit stream from the SPS The decoding method according to any one of claims 1 to 7, further comprising.
9. A method for encoding an image region into a video bit stream, comprising: Determining whether the partitioning of the blocks of the image region according to the first partition constraint information is enabled; When it is determined that the partitioning of the block according to the first partition constraint information is enabled, Setting the value of the override enable flag to the enable value; Encoding the override enable flag having the enable value into the video bit stream; Determining whether to partition the block according to the first partition constraint information; When partitioning the block according to the first partition constraint information, Encoding the first partition constraint information into the image region header of the video bit stream; Partitioning the block according to the first partition constraint information; Setting the value of the override flag to the override value; Encoding the override flag having the override value into the video bit stream; An encoding method including.
10. The encoding method is When it is determined that the partitioning of the block according to the first partition constraint information is not enabled, partitioning the block according to the second partition constraint information and setting the value of the override enable flag to the disable value; Encoding the second partition constraint information into the sequence parameter set (SPS) of the video bit stream; The encoding method according to claim 9, further comprising
11. The encoding method according to claim 10, wherein the second partition constraint information includes partition constraint information in an intra mode or partition constraint information in an inter mode.
12. The encoding method according to claim 10, wherein the second partition constraint information includes partition constraint information for a luma block or partition constraint information for a chroma block.
13. The encoding method according to any one of claims 10 to 12, wherein the override enable flag is in the SPS.
14. The encoding method according to any one of claims 9 to 13, wherein the override flag is in the image region header.
15. The encoding method according to any one of claims 9 to 14, wherein the image region includes a picture, and the image region header includes a picture header of the picture.
16. The encoding method includes When it is determined not to partition the block according to the first partition constraint information, partitioning the block according to the second partition constraint information and setting the value of the override flag to a non-override value, wherein the second partition constraint information is encoded in a sequence parameter set (SPS) of the video bitstream; The encoding method according to any one of claims 9 to 15, further comprising
17. A computer program comprising program code for performing the method according to any one of claims 1 to 16.
18. An encoder, comprising One or more processors; A non-transitory computer-readable storage medium connected to the processor and storing programming for execution by the processor, wherein the programming, when executed by the processor, configures the encoder to perform the method according to any one of claims 1 to 16; An encoder comprising
19. A decoder comprising a processing circuit for performing the method according to any one of claims 1 to 8.
20. An encoder comprising a processing circuit for performing the method according to any one of claims 9 to 16.
21. A decoder for decoding a video bit stream, wherein the video bit stream includes data representing an image region and an image region header of the image region, and the decoder comprises: an override determination unit that obtains an override enable flag from the video bit stream and obtains an override flag from the video bit stream when a value of the override enable flag obtained from the video bit stream is an enable value; a partition constraint determination unit that obtains first partition constraint information of the image region from the image region header when the value of the override flag is an override value; a block partitioning unit that partitions blocks of the image region according to the first partition constraint information; a decoder comprising the above. **Claim 22** A non-transitory computer-readable medium storing program code, wherein when the program code is executed by a computer device, the computer device is caused to execute the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Method and apparatus for cooperative partition coding of region-based filters
JP2012533215A
Signaling of deblocking filter parameters in video coding
JP2015508251A
Signaling of deblocking filter parameters in video coding
US20140369404A1
Cited By
Video encoder, video decoder and corresponding methods
JP2025106300A