Syntax elements for video encoding or video decoding
By incorporating new syntax elements for advanced partitioning, intra prediction, sample adaptive offset, and illumination compensation, the video coding standards achieve enhanced coding efficiency and video quality, addressing limitations in existing technologies.
Patent Information
- Application Number
- JP2025006472
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-02
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-05-28
AI Technical Summary
Existing video coding standards face challenges in achieving high compression efficiency due to limitations in prediction modes, partitioning strategies, and illumination compensation, which affect the flexibility and accuracy of video encoding and decoding processes.
The introduction of new syntax elements that include parameters for advanced partitioning modes, intra prediction modes, sample adaptive offset loop filter adaptations, and illumination compensation for bi-predictive blocks, enhancing the flexibility and efficiency of video coding and decoding processes.
These new syntax elements significantly improve coding efficiency by allowing for more flexible and adaptive encoding and decoding strategies, leading to better compression performance and video quality.
Smart Images

Figure 0007682414000023 
Figure 0007682414000024 
Figure 0007682414000025
Abstract
Description
[Technical Field]
[0001] Technical Field At least one of the present embodiments generally relates to syntax elements for video coding or video decoding. [Background technology]
[0002] background To achieve high compression efficiency, image and video coding schemes typically use prediction and transformation to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original block and the predicted block (often called the prediction error or prediction residual) is transformed, quantized, and entropy coded. To restore the video, the compressed data is decoded by the inverse process corresponding to entropy coding, quantization, transformation, and prediction. Summary of the Invention
[0003] overview According to a first aspect of at least one embodiment, a video signal is presented, the video signal being formatted to include information in accordance with a video coding standard and including a bitstream having video content and high-level syntax information, the high-level syntax information including at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for bi-predictive blocks.
[0004] According to a second aspect of at least one embodiment, a storage medium is presented, the storage medium having coded video signal data, the video signal being formatted to include information according to a video coding standard and including a bitstream having video content and high-level syntax information, the high-level syntax information including at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for bi-predictive blocks.
[0005] According to a third aspect of at least one embodiment, an apparatus is presented, the apparatus including: a video encoder for coding picture data of at least one block in a picture, the coding being performed using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block, the parameters being inserted into a high-level syntax element of the coded picture data.
[0006] According to a fourth aspect of at least one embodiment, a method is presented, the method including: coding picture data of at least one block in a picture using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block; and inserting the parameter into a high-level syntax element of the coded picture data.
[0007] According to a fifth aspect of at least one embodiment, an apparatus is presented, the apparatus including: a video decoder for decoding picture data of at least one block in a picture, the decoding being performed using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block, the parameters being obtained from high-level syntax elements of the coded picture data.
[0008] According to a sixth aspect of at least one embodiment, a method is presented, the method including: obtaining parameters from high-level syntax elements of coded picture data; and decoding picture data of at least one block in the picture, the decoding being performed using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block.
[0009] According to a seventh aspect of at least one embodiment, a non-transitory computer-readable medium is presented, the non-transitory computer-readable medium including data content generated according to the third or fourth aspects.
[0010] According to a seventh aspect of at least one embodiment, there is provided a computer program comprising program code instructions executable by a processor, the computer program performing at least the steps of a method according to the fourth or sixth aspects.
[0011] According to an eighth aspect of at least one embodiment, there is provided a computer program product comprising program code instructions stored on a non-transitory computer readable medium and executable by a processor, the computer program product performing at least the steps of a method according to the fourth or sixth aspects. [Brief explanation of the drawings]
[0012] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] FIG. 1 shows an example of a video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder. [Figure 2] 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. [Figure 3] 1 shows an example of a coding tree unit and a coding tree in the compressed domain. [Figure 4] 1 illustrates an example of division of a CTU into coding units, prediction units, and transform units. [Figure 5] An example of a CTU representation of a Quad-Tree plus Binary-Tree (QTBT) is shown below. [Figure 6] 1 illustrates an exemplary set of extensions for coding unit partitioning. [Figure 7] 10 illustrates an example of an intra prediction mode with two reference layers. [Figure 8] 10 illustrates an example embodiment of a compensation mode for bidirectional illumination compensation. [Figure 9] 1 illustrates the interpretation of bt-split-flag according to an example embodiment. [Figure 10] 1 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. [Figure 11] 1 shows a flowchart of an example of an encoding method according to an embodiment using the new encoding tool. [Figure 12]1 shows a flowchart of an example of a portion of a decoding method according to an embodiment using the new coding tool. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description In at least one embodiment, improved coding efficiency is provided through the use of the following new coding tools: In an example embodiment, efficient signaling of these coding tools conveys information indicating the coding tools utilized for coding, e.g., from an encoder device to a receiver device (e.g., a decoder or display), so that the appropriate tool is used for the decoding stage. These tools include new partitioning modes, new intra-prediction modes, increased flexibility for sample adaptive offset, and new illumination compensation modes for bi-prediction blocks. Thus, the proposed new syntax, which brings together multiple coding tools, provides more efficient video coding.
[0014] Figure 1 shows an example of a video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder. Figure 1 may also show an encoder that incorporates improvements to the HEVC standard or an encoder that uses technology similar to HEVC, such as the Joint Exploration Model (JEM) encoder being developed by the Joint Video Exploration Team (JVET).
[0015] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "coded" or "encoded" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side and "decoded" is used on the decoder side.
[0016] Before being coded, a video sequence may undergo coding preprocessing (101), for example by applying a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or by remapping the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the preprocessing and may be combined into the bitstream.
[0017] In HEVC, to code a video sequence having one or more pictures, a picture is partitioned (102) into one or more slices, where each slice may contain one or more slice segments. Slice segments are organized into coding units, prediction units, and transform units. The HEVC specification distinguishes between "blocks" and "units," where a "block" addresses a particular area of a sample array (e.g., luma, Y), and a "unit" includes all coded color components (Y, Cb, Cr, or monochrome), syntax elements, and array blocks of prediction data associated with the block (e.g., motion vectors).
[0018] For HEVC encoding, a picture is partitioned into square-shaped coding tree blocks (CTBs) with configurable sizes, and a set of consecutive coding tree blocks is grouped into a slice. A coding tree unit (CTU) includes the CTB of a coded color component. The CTB is the root of the quadtree that partitions into coding blocks (CBs), which may be partitioned into one or more prediction blocks (PBs), forming the root of the quadtree that partitions into transform blocks (TBs). Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) includes a set of prediction units (PUs) and tree-structured transform units (TUs), where a PU includes prediction information for all color components, and a TU includes a residual coding syntax structure for each color component. The sizes of the CB, PB, and TB for the luma component are applied to the corresponding CU, PU, and TU. In this application, the term "block" may be used to refer to, for example, any of CTU, CU, PU, TU, CB, PB, and TB. In addition, "block" may also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally to refer to arrays of data of various sizes.
[0019] In the example encoder 100, a picture is coded by the encoder elements as follows: The picture to be coded is processed in units of CUs. Each CU is coded using intra mode or inter mode. When a CU is coded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to code the CU and indicates the intra / inter decision with a prediction mode flag. A prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0020] Intra-mode CUs are predicted from reconstructed neighboring samples within the same slice. A set of 35 intra-prediction modes is available in HEVC, including DC prediction mode, planar prediction mode, and 33 angular prediction modes. Intra-prediction references are reconstructed from rows and columns adjacent to the current block. The references extend across twice the block size in the horizontal and vertical directions, using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra-prediction, the reference samples may be copied along the direction indicated by the angular prediction mode.
[0021] The luma intra-prediction modes applicable to the current block can be coded using two different options: If the applicable mode is included in a constructed list of three most probable modes (MPMs), the mode is signaled by an index in the MPM list; otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra-prediction modes of the top and left neighboring blocks.
[0022] For an inter CU, the corresponding coding block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how the inter prediction is performed. Motion information (e.g., motion vectors and reference picture indexes) can be signaled in two ways: "merge mode" and "advanced motion vector prediction (AMVP)."
[0023] In merge mode, a video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates in the candidate list. At the decoder side, motion vectors (MVs) and reference picture indices are reconstructed based on the signaled candidates.
[0024] In AMVP, a video encoder or decoder assembles a candidate list based on motion vectors determined from previously coded blocks. The video encoder then signals an index within the candidate list to identify a motion vector predictor (MVP) and a motion vector difference (MVD). At the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. Also, the applicable reference picture index is explicitly coded in the AMVP PU syntax.
[0025] The prediction residual is then transformed (125) and quantized (130) (including at least one embodiment for applying a chroma quantization parameter described below). The transform is typically based on a separable transform. For example, a DCT transform is applied first horizontally and then vertically. In modern codecs such as JEM, the transforms used in both directions can be different (e.g., DCT in one direction and DST in the other), which results in a wide variety of 2D transforms, whereas in earlier codecs the variety of 2D transforms for a given block size is typically limited.
[0026] The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also avoid both the transform and quantization (i.e., the residual is coded directly without applying a transform or quantization process). In direct PCM coding, no prediction is applied and coding unit samples are coded directly into the bitstream.
[0027] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. An image block is reconstructed by combining (155) the decoded prediction residual and the predicted block. An in-loop filter (165) is applied to the reconstructed picture to reduce coding artifacts, for example, by performing deblocking / sample adaptive offset (SAO) filtering. The filtered image is stored in a reference picture buffer (180).
[0028] Figure 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. In the example decoder 200, the bitstream is decoded by the following decoder elements: The video decoder 200 generally performs a decoding pass that is an inverse of the coding pass as described in Figure 1, which performs video decoding as part of the coding of the video data. Figure 2 may also show a decoder that has improvements made to the HEVC standard or that uses HEVC-like technology, such as a JEM decoder.
[0029] Specifically, the decoder's input includes a video bitstream, such as may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, picture partitioning information, and other coding information. The picture partitioning information indicates the size of CTUs and how the CTUs are split into CUs and, if applicable, PUs. Thus, the decoder divides (235) the picture into CTUs and each CTU into CUs according to the decoded picture partitioning information. The transform coefficients are inverse quantized (240) (including at least one embodiment for adapting chroma quantization parameters, described below) and inverse transformed (250) to decode the prediction residual.
[0030] An image block is reconstructed by combining (255) the decoded prediction residual and the predicted block. The predicted block may be obtained (270) from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275). As described above, AMVP and merge mode techniques can be used to derive motion vectors for motion compensation, which may use interpolation filters to calculate interpolated values of sub-integer samples of the reference block. An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0031] The decoded picture may further undergo post-decoding processing (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that reverses the remapping process performed in pre-coding processing (101). The post-decoding processing may use metadata derived in pre-coding processing and signaled in the bitstream.
[0032] Figure 3 shows an example of a coding tree unit and a coding tree in the compressed domain. In the HEVC video compression standard, a picture is divided into so-called coding tree units (CTUs), whose sizes are typically 64x64, 128x128, or 256x256 pixels. Each CTU is represented by a coding tree in the compressed domain, which is a quadtree decomposition of the CTU, where each leaf is called a coding unit (CU).
[0033] 4 shows an example of division of a CTU into coding units, prediction units, and transform units. Each CU is given some intra prediction parameters or inter prediction parameters (prediction information). To this end, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. An intra coding mode or an inter coding mode is assigned at the CU level.
[0034] Emerging video compression tools, including a coding tree unit representation in the compressed domain, are proposed to represent picture data in a more flexible way. The advantage of this more flexible representation of the coding tree is that it offers improved compression efficiency compared to the CU / PU / TU arrangement of the HEVC standard.
[0035] FIG. 5 shows an example of a quadtree + binary tree (QTBT) CTU representation. The quadtree + binary tree (QTBT) coding tool provides this increased flexibility. The QTBT consists of a coding tree in which coding units can be split in both quadtree and binary tree fashions. The splitting of coding units is determined on the encoder side by a rate distortion (RD) optimization procedure that determines the QTBT representation of the CTU with the minimum RD cost. In the QTBT technique, CUs have square or rectangular shapes. The size of a coding unit is always a power of two, typically ranging from 4 to 128. In addition to this rectangular variety of coding units, such a CTU representation has the following different characteristics compared to HEVC: The QTBT decomposition of a CTU consists of two stages: first, the CTU is split in a quadtree fashion, and then each quadtree leaf can be further divided in a binary fashion. This is shown on the right side of the diagram, where the solid lines represent the quadtree decomposition phase and the dashed lines represent the binary decomposition spatially embedded in the quadtree leaves. In intra-slice, the luma and chroma block partitioning structures are separated and determined independently. No further CU partitioning into prediction units or transform units is used. That is, each coding unit systematically consists of a single prediction unit (2Nx2N prediction unit partition type) and a single transform unit (no division into a transform tree).
[0036] 6 shows an exemplary set of extensions for coding unit partitioning. In one asymmetric bisection and tree splitting mode (ABT), a rectangular coding unit with size (w, h) (width and height) that is split by one of the asymmetric bisection splitting modes (e.g., HOR_UP (horizontal-up)) is split into rectangular coding units of size (w, h), respectively.
number
[0037] Using one or more embodiments of the new topology described above, significant coding efficiency improvements are achieved.
[0038] FIG. 7 shows an example of an intra prediction mode with two reference layers. In fact, in the improved syntax, new intra prediction modes are considered. The first new intra prediction mode is named multi-reference intra prediction (MRIP). This tool facilitates using multiple reference layers for intra prediction of a block. Generally, two reference layers are used for intra prediction, and each reference layer consists of a left reference array and an upper reference array. Each reference layer is constructed by reference sample permutation and then pre-filtered.
[0039] Each reference layer is used to build a prediction for the block, as is commonly done in HEVC or JEM. The final prediction is formed as a weighted average of the predictions made from the two reference layers. The prediction from the closest reference layer is given more weight than the prediction from the furthest layer. Commonly used weights are 3 and 1.
[0040] In the final part of the prediction process, multiple reference layers may be used to smooth boundary samples for a particular prediction mode using a mode-dependent process.
[0041] In an example embodiment, the video compression tool also includes an adaptive block size for sample adaptive offset (SAO), which is a loop filter specified in HEVC. In HEVC, the SAO process classifies the reconstructed samples of a block into several classes, and samples belonging to some classes are corrected using an offset. The SAO parameters can be coded block by block, or can be inherited only from the left or top neighboring blocks.
[0042] A further improvement is proposed by defining an SAO palette mode. The SAO palette maintains the same SAO parameters per block, but a block can inherit these parameters from all other blocks. This brings more flexibility to SAO by widening the range of possible SAO parameters for a block. The SAO palette consists of different sets of SAO parameters. For each block, an index is coded to indicate which SAO parameters to use.
[0043] 8 illustrates an example embodiment of a compensation mode of bidirectional illumination compensation. Illumination Compensation (IC) allows correcting block prediction samples (SMC) obtained by motion compensation, possibly by taking into account spatial or temporal local illumination variations. S IC =a i .S MC +b i
[0044] In the case of bi-prediction, IC parameters are estimated using samples of two reference CU samples as depicted in Figure 8. First, IC parameters (a, b) are estimated between two reference blocks, and then IC parameters (ai, bi) i = 0, 1 between the reference blocks and the current block are derived as follows:
number
[0045] If the CU size is less than or equal to 8 in width or height, the IC parameters are estimated using a unidirectional process. If the CU size is greater than 8, the choice of using IC parameters derived from the bidirectional or unidirectional process is made by selecting the IC parameters that minimize the difference in the means of the IC-compensated reference blocks. Bidirectional optical flow (BIO) is enabled using a bidirectional illumination compensation tool.
[0046] The tools introduced above need to be selected and used by video encoder 100, and retrieved and used by video decoder 200. To that end, information representing the use of these tools is carried in the coded bitstream generated by the video encoder and retrieved by video decoder 200 in the form of high-level syntax information. To that end, a corresponding syntax is defined and described below. This syntax, which brings together the tools, enables improved coding efficiency.
[0047] This syntax structure uses the HEVC syntax structure as a basis and includes some additional syntax elements. In the table below, syntax elements signaled in bold italic font and highlighted in gray correspond to additional syntax elements according to an example embodiment. Note that syntax elements may have other forms or names than those shown in the syntax table while handling the same functionality. Syntax elements may exist at different levels (e.g., some syntax elements may be located in a sequence parameter set (SPS) and some may be located in a picture parameter set (PPS)).
[0048] Table 1 below shows the Sequence Parameter Set (SPS) and introduces new syntax elements inserted into this parameter set according to at least one example embodiment, more precisely: multi_type_tree_enabled_primary, log2_min_cu_size_minus2, log2_max_cu_size_minus4, log2_max_tu_size_minus2, sep_tree_mode_intra, multi_type_tree_enabled_secondary, sps_bdip_enabled_flag, sps_mrip_enabled_flag, use_erp_aqp_flag, use_high_perf_chroma_qp_table, abt_one_third_flag, sps_num_intra_mode_ratio.
[0049] [Table 1]
[0050] [Table 2]
[0051] [Table 3]
[0052] [Table 4]
[0053] The new syntax elements are defined as follows: - use_high_perf_chroma_qp_table: This syntax element specifies the chroma QP table to be used to derive the QP used for decoding the chroma components associated with a slice as a function of the base chroma QP associated with the chroma components of the considered slice. - multi_type_tree_enabled_primary: This syntax element indicates the type of partitioning allowed in a coded slice. In at least one embodiment, this syntax element is signaled at the sequence level (in an SPS) and therefore applies to all coded slices using this SPS. For example, the first value of this syntax element allows quad-tree and binary tree (QTBT) partitioning (NO-SPLIT, QT-SPLIT, HOR, and VER in Figure 5), the second value allows QTBT + triple-tree (TT) partitioning (NO-SPLIT, QT-SPLIT, HOR, VER, HOR_TRIPLE, and VER_TRIPLE in Figure 5), the third value allows QTBT + asymmetric binary tree (ABT) partitioning (NO-SPLIT, QT-SPLIT, HOR, VER, HOR-UP, HOR_DOWN, VER_LEFT, and VER_RIGHT in Figure 5), and the fourth value allows QTBT + TT + ABT partitioning (all split cases in Figure 5). - sep_tree_mode_intra: This syntax element indicates whether separate coding trees are used for luma and chroma blocks. If sep_tree_mode_intra is equal to 1, the luma and chroma blocks use independent coding trees, and thus the partitioning of the luma and chroma blocks is independent. In at least one embodiment, this syntax element is signaled at the sequence level (in the SPS) and therefore applies to all coded slices using this SPS. - multi_type_tree_enabled_secondary: This syntax element indicates the type of partitioning allowed for chroma blocks in a coded slice. The possible values of this syntax element are generally the same as those of the syntax element multi_type_tree_enabled_primary. In at least one embodiment, this syntax element is signaled at the sequence level (in an SPS) and therefore applies to all coded slices using this SPS. sps_bdip_enabled_flag: This syntax element indicates that bidirectional intra prediction, as described in JVET-J0022, is allowed in coded slices included in the considered coded video sequence. sps_mrip_enabled_flag: This syntax element indicates that multi-reference intra prediction tools are used in decoding code slices contained in the considered coded video bitstream. use_erp_aqp_flag: This syntax element indicates whether spatial adaptive quantization of ERP content in VR360 as described in JVET-J0022 is activated in the coded slice. abt_one_third_flag: This syntax element indicates whether asymmetric partitioning into two partitions with horizontal or vertical dimensions that are 1 / 3 and 2 / 3, or 2 / 3 and 1 / 3, respectively, of the horizontal or vertical dimension of the first CU is activated in the coding slice. sps_num_intra_mode_ratio: This syntax element specifies how to derive the number of intra predictions to be used for block sizes equal to a multiple of 3 in width or height.
[0054] Additionally, the following three syntax elements are used to control the size of the coding unit (CU) and the size of the transform unit (TU): In at least one embodiment, these syntax elements are signaled at the sequence level (in the SPS) and therefore apply to all coded slices using this SPS: - log2_min_cu_size_minus2 specifies the minimum CU size. - log2_max_cu_size_minus4 specifies the maximum CU size. - log2_max_tu_size_minus2 specifies the maximum TU size.
[0055] In another embodiment, the new syntax elements introduced above are introduced in the Picture Parameter Set (PPS) instead of the Sequence Parameter Set.
[0056] Table 2 below shows the Sequence Parameter Set (SPS) and introduces new syntax elements inserted into this parameter set according to at least one example embodiment, more precisely: - slice_sao_size_id: This syntax element specifies the size of the block to which SAO is applied in terms of the coded slice. - slice_bidir_ic_enable_flag: indicates whether bidirectional illumination compensation is activated for the coded slice.
[0057] [Table 5]
[0058] [Table 6]
[0059] [Table 7]
[0060] [Table 8]
[0061] [Table 9]
[0062] Tables 3, 4, 5, and 6 below show the coding tree syntax modified to support additional partitioning modes according to at least one example embodiment. Specifically, the syntax that allows asymmetric partitioning is specified in coding_binary_tree(). New syntax elements, more precisely the following, are inserted into the coding tree syntax: - asymmetricSplitFlag, which indicates whether asymmetric bisection splits are allowed for the current CU, and - asymmetric_type, which indicates the type of asymmetric bisection split, can take two values to allow the horizontal split to be top or bottom, or the vertical split to be left or right
[0063] In addition, the parameter btSplitMode is interpreted differently.
[0064] [Table 10]
[0065] [Table 11]
[0066] [Table 12]
[0067] [Table 13]
[0068] [Table 14]
[0069] FIG. 9 illustrates the interpretation of bt-split-flag according to at least one example embodiment. Two syntax elements, asymmetricSplitFlag and asymmetric_type, are used to specify the value of the parameter btSplitMode. btSplitMode indicates the bisection split mode applied to the current CU. btSplitMode is derived differently than in the prior art. In JEM, it can take values of HOR, VER. In the enhanced syntax, it can additionally take values of HOR_UP, HOR_DOWN, VER_LEFT, VER_RIGHT, which correspond to the partitioning shown in FIG. 6.
[0070] Tables 7, 8, 9, and 10 below show the coding unit syntax elements and introduce new syntax elements for bidirectional intra-prediction modes.
[0071] [Table 15]
[0072] [Table 16]
[0073] [Table 17]
[0074] [Table 18]
[0075] [Table 19]
[0076] [Table 20]
[0077] Some of the syntax elements described in this document are defined as arrays. For example, btSplitFlag is defined by an array of dimension 4 indexed by the horizontal and vertical position within the picture and by the coding block horizontal and vertical size. To simplify the notation, these indices are not maintained in the semantic description of the syntax elements (e.g., btSplitFlag[x0][y0][cbSizeX][cbSizeY] is simply denoted as btSplitFlag).
[0078] FIG. 10 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 1000 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, encoders, transcoders, and servers. Elements of system 1000, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, system 1000 is configured to perform one or more of the aspects described herein.
[0079] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 1010 may include internal memory, input / output interfaces, and various other circuitry known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 1040 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0080] System 1000 includes an encoder / decoder module 1030 configured to process data to provide, for example, coded or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that may be included within a device to perform coding and / or decoding functions. As is known, a device may include one or both of a coding module and a decoding module. Furthermore, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within the processor 1010 as a combination of hardware and software, as known to those skilled in the art.
[0081] Program code loaded into the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and later loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items while performing the processes described herein. Such stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, expressions, operations, and arithmetic logic.
[0082] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040 (e.g., dynamic volatile memory and / or non-volatile flash memory). In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as MPEG-2, HEVC, or Versatile Video Coding (VVC).
[0083] Input to the elements of system 1000 may be provided using various input devices as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted over the air, for example by a broadcast station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0084] In various embodiments, the input devices of block 1130 have associated respective input processing elements known in the art. For example, the RF section may be associated with elements necessary to (i) select a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconvert the selected signal, (iii) bandlimit again to a narrower frequency band to select (for example) a signal frequency band, which in certain embodiments may be referred to as a channel, (iv) demodulate the downconverted and bandlimited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs some of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements (e.g., inserting amplifiers and analog-to-digital converters, etc.). In various embodiments, the RF section includes an antenna.
[0085] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via USB and / or HDMI connections. It is understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 1010, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, as desired, within a separate interface IC or within processor 1010. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and an encoder / decoder 1030, which operates in combination with memory and storage elements to process the data stream as desired for display on an output device.
[0086] The various elements of system 1000 may be provided within an integrated housing in which the various elements may be interconnected and transmit data between them using any suitable connection scheme (e.g., internal buses known in the art), including an I2C bus, wiring, and printed circuit boards.
[0087] System 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. Communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 1060. Communication interface 1050 may include, but is not limited to, a modem or a network card, and communication channel 1060 may be implemented within a wired and / or wireless medium, for example.
[0088] In various embodiments, data is streamed to system 1000 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received by communication interface 1050 and communication channel 1060 suitable for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, enabling streaming usage and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that sends data via HDMI communication in input block 1130. Still other embodiments provide streamed data to system 1000 using an RF connection in input block 1130.
[0089] System 1000 may provide output signals to various output devices, including display 1100, speakers 1110, and other peripheral devices 1120. Other peripheral devices 1120, in various example embodiments, include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices may be communicatively coupled to system 1000 via dedicated connections via respective interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speakers 1110 may be integrated in a single unit with other components of the system 1000 in an electronic device such as a television. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip.
[0090] The display 1100 and speakers 1110 may alternatively be separate from one or more of the other components, for example, if the RF portion of the input 1130 is part of a separate set-top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output. The implementations described herein may be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even when described only in the context of a single form of implementation (e.g., described only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in an apparatus such as a processor, which refers generally to processing devices (e.g., including a computer, microprocessor, integrated circuit, or programmable logic device). Processors also include communication devices such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0091] FIG. 11 shows a flowchart of an example of an encoding method using a new encoding tool, according to an embodiment. Such an encoding method can be performed by the system 1000 described in FIG. 10, and more precisely, can be implemented by the processor 1010. In at least one embodiment, in step 1190, the processor 1010 selects the encoding tool to be used. The selection can be made using different techniques. The selection can be made by a user using encoding configuration parameters (typically flags) that indicate the tool (and corresponding parameters, if necessary) to be used during the encoding and decoding processes. In an example embodiment, the selection is made by setting a value in a file, and the flag is read by the encoding device, which interprets the value of the flag to select which tool to use. In an example embodiment, the selection is made by a manual action of a user selecting the encoding configuration parameters using a graphical user interface operating the encoder device. Once this selection is made, coding is performed in step 1193 using the selected coding tool from among several available coding tools, with the tool selection being signaled in a higher level syntax element (e.g., in the following SPS, PPS, or slice header).
[0092] Figure 12 shows a flowchart of an example of a portion of a decoding method according to one embodiment using the new coding tools. Such a decoding method can be performed by the system 1000 described in Figure 10, and more precisely, can be implemented by the processor 1010. In at least one embodiment, in step 1200, the processor 1010 accesses a signal (e.g., received at an input interface or read from a media support) and extracts and analyzes high-level syntax elements to determine the tools selected by the coding device. In step 1210, decoding is performed using these tools to generate a decoded image that can, for example, be provided to or displayed on a device.
[0093] References to "one embodiment," or "an embodiment," or "one implementation," or "an embodiment," and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment," or "in an embodiment," or "in one embodiment," or "in an embodiment," and other variations thereof, in various places throughout this specification do not necessarily all refer to the same embodiment.
[0094] Additionally, this application or its claims may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0095] Additionally, this application or its claims may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, predicting information, or estimating information.
[0096] Additionally, the application or its claims may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory or optical media storage). Furthermore, "receiving" generally involves in some way an operation such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0097] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selection of only the first listed option (A), or selection of only the second listed option (B), or selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass selection of only the first listed option (A), or selection of only the second listed option (B), or selection of only the third listed option (C), or selection of only the first and second listed options (A and B), or selection of only the first and third listed options (A and C), or selection of only the second and third listed options (B and C), or selection of all three options (A, B, and C). This can be expanded to include as many items as are listed, as would be readily apparent to one skilled in this and related arts.
[0098] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. Information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0099] In variants of at least the first, second, third, fourth, fifth, sixth, seventh, and eighth aspects of an embodiment, the parameter representing the type of partitioning includes a first value of quadtree and binary tree partitioning, a second value of quadtree and binary tree partitioning + ternary tree partitioning, a third value of quadtree and binary tree partitioning + asymmetric binary tree partitioning, and a fourth value of quadtree and binary tree partitioning + ternary tree partitioning + asymmetric binary tree partitioning.
[0100] In a variant of at least the first, second, third, fourth, fifth, sixth, seventh, and eighth aspects of an embodiment, the parameters further include a size of a block to which the sample adaptive offset is applied with respect to the coded slice.
[0101] In a variant of at least the first, second, third, fourth, fifth, sixth, seventh, and eighth aspects of an embodiment, the type of intra prediction mode includes at least one of an intra bi-prediction mode and a multi-reference intra prediction mode.
Claims
1. Apparatus for encoding video picture data, comprising: Obtaining the picture data; determining that the coding of the picture data should use independent coding trees for a plurality of luma and chroma blocks and determining a type of partitioning allowed for the chroma blocks; inserting into a sequence parameter set of the coded picture data a single flag and at least one parameter indicative of a type of partitioning allowed for the chroma blocks, the single flag indicating that the coding uses independent coding trees for the luma and chroma blocks; coding the picture data for at least one block in a picture and the sequence parameter set for a plurality of blocks to generate a coded bitstream, the coding using the determined type of partitioning allowed for the chroma blocks.
23. An apparatus comprising: at least one processor configured to:
2. The apparatus of claim 1, wherein if the single flag is equal to 1, the multiple luma blocks and chroma blocks use independent coding trees, and if the single flag is equal to 0, a common coding tree is used for the multiple luma blocks and chroma blocks.
3. The apparatus of claim 1, wherein the single flag is signaled at a sequence level within a sequence parameter set so that the single flag is applied to all coded slices using the sequence parameter set.
4. The apparatus of claim 1, further comprising a file input interface, wherein the determination is made by reading a value of a flag in a file.
5. The apparatus of claim 1 , further comprising a graphical user interface, and wherein said determining is made by obtaining coded configuration parameters from said graphical user interface.
6. A method for encoding video picture data, comprising the steps of: obtaining said picture data; determining that the coding of the picture data should use independent coding trees for luma and chroma blocks and determining a type of partitioning allowed for the chroma blocks; inserting into a sequence parameter set of the coded picture data a single flag and at least one parameter indicative of a type of partitioning allowed for the chroma blocks, the single flag indicating that the coding uses independent coding trees for the luma and chroma blocks; encoding the picture data for at least one block in a picture and the sequence parameter set for a plurality of blocks to generate a coded bitstream, the coding using the determined type of partitioning allowed for the chroma blocks; A method comprising:
7. The method described in claim 6, wherein if the single flag is equal to 1, the multiple luma blocks and chroma blocks use independent coding trees, and if the single flag is equal to 0, a common coding tree is used for the multiple luma blocks and chroma blocks.
8. The method of claim 6, wherein the single flag is signaled at a sequence level within a sequence parameter set so that the single flag applies to all coded slices using the sequence parameter set.
9. A non-transitory computer-readable medium having stored thereon program code instructions executable by a processor for performing the steps of the method of claim 6.
10. The method of claim 6, wherein the determination is made by reading the value of a flag in a file.
11. The method of claim 6, wherein the determination is made by obtaining coded configuration parameters from a graphical user interface.
12. Apparatus for decoding video picture data, comprising: receiving coded picture data; Decoding from a sequence parameter set of the coded picture data at least one parameter indicative of a single flag and a type of partitioning allowed for chroma blocks; determining that the single flag indicates that the coded picture data uses independent coding trees for multiple luma and chroma blocks; determining the type of partitioning allowed for the chroma block; Decoding picture data for at least one block in a picture using separate coding trees for the plurality of luma blocks and chroma blocks and using the determined type of partitioning allowed for the chroma blocks.
23. An apparatus comprising: at least one processor configured to:
13. The apparatus of claim 12, wherein if the single flag is equal to 1, the multiple luma blocks and chroma blocks use independent coding trees, and if the single flag is equal to 0, a common coding tree is used for the multiple luma blocks and chroma blocks.
14. The apparatus of claim 12, wherein the single flag is signaled at a sequence level within a sequence parameter set such that the single flag is applied to all coded slices using the sequence parameter set.
15. The device of claim 12, wherein a first value of the parameter representing the type of partitioning of the saturation blocks allows quadtree and binary tree partitioning, a second value of the parameter representing the type of partitioning of the saturation blocks allows quadtree, binary tree and ternary tree partitioning, and a third value of the parameter representing the type of partitioning of the saturation blocks allows quadtree, binary tree and asymmetric binary tree partitioning.
16. A method for decoding video picture data, comprising the steps of: receiving coded picture data; decoding at least one parameter indicative of a type of partitioning allowed for a single flag and a chroma block from a sequence parameter set of the coded picture data; determining that the single flag indicates that the coded picture data uses independent coding trees for multiple luma and chroma blocks; determining the type of partitioning allowed for the chroma block; decoding picture data for at least one block in a picture using separate coding trees for the plurality of luma blocks and chroma blocks and using the determined type of partitioning allowed for the chroma blocks; A method comprising:
17. The method described in claim 16, wherein if the single flag is equal to 1, the multiple luma blocks and chroma blocks use independent coding trees, and if the single flag is equal to 0, a common coding tree is used for the multiple luma blocks and chroma blocks.
18. The method of claim 16, wherein the single flag is signaled at a sequence level within a sequence parameter set such that the single flag applies to all coded slices using the sequence parameter set.
19. A non-transitory computer-readable medium having stored thereon program code instructions executable by a processor for performing the steps of the method of claim 16.
20. The method of claim 16, wherein a first value of the parameter representing the type of partitioning of the saturation blocks allows quadtree and binary tree partitioning, a second value of the parameter representing the type of partitioning of the saturation blocks allows quadtree, binary tree and ternary tree partitioning, and a third value of the parameter representing the type of partitioning of the saturation blocks allows quadtree, binary tree and asymmetric binary tree partitioning.
Citation Information
Patent Citations
Using luma information for chroma prediction with separate luma-chroma framework in video coding
US20170272748A1
Method and device for encoding / decoding an image unit comprising image data represented by a luminance channel and at least one chrominance channel
WO2017137311A1