Syntax elements for video encoding or video decoding
By integrating advanced partitioning, multi-reference intra-prediction, and bidirectional illumination compensation, the solution addresses inefficiencies in HEVC, enhancing video coding efficiency through improved partitioning and prediction techniques.
Patent Information
- Application Number
- JP2025080605
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-07-02
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-13
AI Technical Summary
Existing video coding technologies, such as HEVC, face challenges in achieving efficient compression due to limitations in partitioning modes, intra-prediction, and illumination compensation, leading to suboptimal coding efficiency.
Incorporation of new coding tools including enhanced partitioning modes, multi-reference intra-prediction, adaptive block size for sample adaptive offset loop filtering, and bidirectional illumination compensation, with efficient signaling of these tools in the bitstream to facilitate improved decoding.
The proposed solution achieves significant improvements in coding efficiency by optimizing partitioning, prediction, and illumination compensation, resulting in enhanced video compression performance.
Smart Images

Figure 2025118845000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical Field At least one of the present embodiments generally relates to syntax elements for video coding or video decoding. [Background technology]
[0002] background To achieve high compression efficiency, image and video coding schemes typically use prediction and transformation to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original block and the predicted block (often called the prediction error or prediction residual) is transformed, quantized, and entropy coded. To restore the video, the compressed data is decoded by the inverse process corresponding to entropy coding, quantization, transformation, and prediction. Summary of the Invention
[0003] overview According to a first aspect of at least one embodiment, a video signal is presented, the video signal being formatted to include information in accordance with a video coding standard and including a bitstream having video content and high-level syntax information, the high-level syntax information including at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for bi-predictive blocks.
[0004] According to a second aspect of at least one embodiment, a storage medium is presented, the storage medium having coded video signal data, the video signal being formatted to include information according to a video coding standard and including a bitstream having video content and high-level syntax information, the high-level syntax information including at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for bi-predictive blocks.
[0005] According to a third aspect of at least one embodiment, an apparatus is presented, the apparatus including: a video encoder for coding picture data of at least one block in a picture, the coding being performed using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block, the parameters being inserted into a high-level syntax element of the coded picture data.
[0006] According to a fourth aspect of at least one embodiment, a method is presented, the method comprising: coding picture data of at least one block in a picture, the method comprising: selecting at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing a block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block. This involves coding performed using parameters and inserting the parameters into high-level syntax elements of the coded picture data.
[0007] According to a fifth aspect of at least one embodiment, an apparatus is presented, the apparatus including: a video decoder for decoding picture data of at least one block in a picture, the decoding being performed using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block, the parameters being obtained from high-level syntax elements of the coded picture data.
[0008] According to a sixth aspect of at least one embodiment, a method is presented, the method including: obtaining parameters from high-level syntax elements of coded picture data; and decoding picture data of at least one block in the picture, the decoding being performed using at least one parameter from among a parameter representing a type of partitioning, a parameter representing a type of intra-prediction mode, a parameter representing block size adaptation of a sample adaptive offset loop filter, and a parameter representing an illumination compensation mode for a bi-predictive block.
[0009] According to a seventh aspect of at least one embodiment, a non-transitory computer-readable medium is presented, the non-transitory computer-readable medium including data content generated according to the third or fourth aspects.
[0010] According to a seventh aspect of at least one embodiment, there is provided a computer program comprising program code instructions executable by a processor, the computer program performing at least the steps of a method according to the fourth or sixth aspects.
[0011] According to an eighth aspect of at least one embodiment, there is provided a computer program product comprising program code instructions stored on a non-transitory computer readable medium and executable by a processor, the computer program product performing at least the steps of a method according to the fourth or sixth aspects. [Brief explanation of the drawings]
[0012] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] FIG. 1 shows an example of a video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder. [Figure 2] 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. [Figure 3] 1 shows an example of a coding tree unit and a coding tree in the compressed domain. [Figure 4] 1 illustrates an example of division of a CTU into coding units, prediction units, and transform units. [Figure 5] An example of a CTU representation of a Quad-Tree plus Binary-Tree (QTBT) is shown below. [Figure 6] 1 illustrates an exemplary set of extensions for coding unit partitioning. [Figure 7] 10 illustrates an example of an intra prediction mode with two reference layers. [Figure 8] 10 illustrates an example embodiment of a compensation mode for bidirectional illumination compensation. [Figure 9] 1 illustrates the interpretation of bt-split-flag according to an example embodiment. [Figure 10] 1 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. [Figure 11] 1 shows a flowchart of an example of an encoding method according to an embodiment using the new encoding tool. [Figure 12]1 shows a flowchart of an example of a portion of a decoding method according to an embodiment using the new coding tool. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description In at least one embodiment, improved coding efficiency is provided through the use of the following new coding tools: In an example embodiment, efficient signaling of these coding tools conveys information representing the coding tools utilized for coding, e.g., from an encoder device to a receiver device (e.g., a decoder or display), so that the appropriate tool is used for the decoding stage. These tools include new partitioning modes, new intra-prediction modes, increased flexibility for sample adaptive offset, and new illumination compensation modes for bi-prediction blocks. Thus, The proposed new syntax brings together multiple coding tools to provide more efficient video coding.
[0014] FIG. 1 illustrates an example of a video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder. FIG. 1 illustrates an encoder that incorporates improvements to the HEVC standard, or a Joint Exploration Model (JEM) encoder that is being developed by the Joint Video Exploration Team (JVET). It may also refer to an encoder that uses technology similar to HEVC, such as an encoder.
[0015] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "coded" or "encoded" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side and "decoded" is used on the decoder side.
[0016] Before being coded, a video sequence may undergo coding preprocessing (101), for example by applying a color transformation to the input color picture (e.g., from RGB 4:4:4 to YCbCr 4:2:0) or by remapping the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the preprocessing and may be combined into the bitstream.
[0017] In HEVC, to code a video sequence having one or more pictures, a picture is partitioned (102) into one or more slices, where each slice may contain one or more slice segments. Slice segments are organized into coding units, prediction units, and transform units. The HEVC specification distinguishes between "blocks" and "units," where a "block" addresses a particular area of a sample array (e.g., luma, Y), and a "unit" includes all coded color components (Y, Cb, Cr, or monochrome), syntax elements, and array blocks of prediction data associated with the block (e.g., motion vectors).
[0018] For HEVC encoding, a picture is partitioned into square coding tree blocks (CTBs) of configurable size, and consecutive A set of coding tree blocks corresponding to a given color component is grouped into a slice. A coding tree unit (CTU) contains the CTB of a coded color component. The CTB is the root of the quadtree that partitions into coding blocks (CBs), which may be partitioned into one or more prediction blocks (PBs) and transform blocks (TBs). The CU forms the root of a quadtree that encodes the coding blocks, prediction blocks, and transform blocks. Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) includes a set of prediction units (PU) and tree-structured transform units (TU), where a PU includes prediction information for all color components, and a TU includes a residual coding syntax structure for each color component. The sizes of the CB, PB, and TB of the luma component apply to the corresponding CU, PU, and TU. In this application, the term "block" may be used to refer to, for example, any of the CTU, CU, PU, TU, CB, PB, and TB. In addition, "block" may also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally to refer to arrays of data of various sizes.
[0019] In the example encoder 100, a picture is coded by the encoder elements as follows: The picture to be coded is processed in units of CUs. Each CU is coded using intra mode or inter mode. When a CU is coded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to code the CU and indicates the intra / inter decision with a prediction mode flag. A prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0020] Intra-mode CUs are predicted from reconstructed neighboring samples within the same slice. A set of 35 intra-prediction modes is available in HEVC, including DC prediction mode, planar prediction mode, and 33 angular prediction modes. Intra-prediction references are reconstructed from rows and columns adjacent to the current block. The references extend across twice the block size in the horizontal and vertical directions, using samples available from previously reconstructed blocks. When an angular prediction mode is used for intra-prediction, the reference samples may be copied along the direction indicated by the angular prediction mode.
[0021] The luma intra-prediction modes applicable to the current block can be coded using two different options: If the applicable mode is included in a constructed list of three most probable modes (MPMs), the mode is signaled by an index in the MPM list; otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra-prediction modes of the top and left neighboring blocks.
[0022] For an inter CU, the corresponding coding block is further partitioned into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how inter prediction is performed. Motion information (e.g., motion vectors and reference picture indexes) is signaled in two ways: "merge mode" and "advanced motion vector prediction (AMVP)." It can be pinged.
[0023] In merge mode, a video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals an index for one of the candidates in the candidate list. At the decoder side, motion vectors (MVs) and reference picture indices are reconstructed based on the signaled candidates. .
[0024] In AMVP, a video encoder or decoder assembles a candidate list based on motion vectors determined from already coded blocks. is used to identify the candidate list of motion vector predictors (MVPs). The AMVP signaling scheme signals the index within the frame and the motion vector difference (MVD). At the decoder side, the motion vector (MV) is reconstructed as MVP+MVD. Also, the applicable reference picture index is explicitly coded in the PU syntax of AMVP.
[0025] The prediction residual is then transformed (125) and quantized (130) (including at least one embodiment for applying a chroma quantization parameter described below). The transform is typically based on a separable transform. For example, a DCT transform is applied first horizontally and then vertically. In modern codecs such as JEM, the transforms used in both directions can be different (e.g., DCT in one direction and DST in the other), which results in a wide variety of 2D transforms, whereas in earlier codecs the variety of 2D transforms for a given block size is typically limited.
[0026] The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also avoid both the transform and quantization (i.e., the residual is coded directly without applying a transform or quantization process). In direct PCM coding, no prediction is applied and coding unit samples are coded directly into the bitstream.
[0027] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. An image block is reconstructed by combining (155) the decoded prediction residual and the predicted block. An in-loop filter (165) is applied to the reconstructed picture to reduce coding artifacts, for example, by performing deblocking / sample adaptive offset (SAO) filtering. The filtered image is stored in a reference picture buffer (180).
[0028] Figure 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. In the example decoder 200, the bitstream is decoded by the following decoder elements: The video decoder 200 generally performs a decoding pass that is an inverse of the coding pass as described in Figure 1, which performs video decoding as part of the coding of the video data. Figure 2 may also show a decoder that has improvements made to the HEVC standard or that uses HEVC-like technology, such as a JEM decoder.
[0029] Specifically, the decoder's input includes a video bitstream, such as may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, picture partitioning information, and other coding information. The picture partitioning information indicates the size of CTUs and how the CTUs are split into CUs and, if applicable, PUs. Thus, the decoder divides (235) the picture into CTUs and each CTU into CUs according to the decoded picture partitioning information. The transform coefficients are inverse quantized (240) (including at least one embodiment for adapting chroma quantization parameters, described below) and inverse transformed (250) to decode the prediction residual.
[0030] The image block is reconstructed by combining the decoded prediction residual and the predicted block (255). The predicted block can be obtained (270) from intra prediction (260) or motion compensated prediction (i.e., inter prediction) (275). As described above, the AMVP and merge mode techniques are used to interpolate sub-integer samples of the reference block. A motion vector for motion compensation can be derived, which may use an interpolation filter to calculate the value. An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0031] The decoded picture may further undergo post-decoding processing (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4 conversion) or an inverse remapping that reverses the remapping process performed in pre-coding processing (101). The post-decoding processing may use metadata derived in pre-coding processing and signaled in the bitstream.
[0032] Figure 3 shows an example of a coding tree unit and a coding tree in the compressed domain. In the HEVC video compression standard, a picture is divided into so-called coding tree units (CTUs), whose sizes are typically 64x64, 128x128, or 256x256 pixels. Each CTU is represented by a coding tree in the compressed domain, which is a quadtree decomposition of the CTU, where each leaf is called a coding unit (CU).
[0033] 4 shows an example of division of a CTU into coding units, prediction units, and transform units. Each CU is given some intra prediction parameters or inter prediction parameters (prediction information). To this end, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. An intra coding mode or an inter coding mode is assigned at the CU level.
[0034] Emerging video compression tools, including a coding tree unit representation in the compressed domain, are proposed to represent picture data in a more flexible way. The advantage of this more flexible representation of the coding tree is that it offers improved compression efficiency compared to the CU / PU / TU arrangement of the HEVC standard.
[0035] Figure 5 shows an example of a quadtree + binary tree (QTBT) CTU representation. The quadtree + binary tree (QTBT) coding tool provides this increased flexibility. A QTBT consists of a coding tree in which coding units can be split in both quadtree and binary tree fashion. Splitting a coding unit results in a CTU with the minimum rate distortion cost. The QTBT representation of a CU is determined at the encoder side by an RD optimization procedure. In the QTBT technique, CUs have a square or rectangular shape. The size of a coding unit is always a power of two, typically ranging from 4 to 128. In addition to this rectangular variety of coding units, such CTU representations have the following different characteristics compared to HEVC. The QTBT decomposition of a CTU consists of two stages: first, the CTU is split in a quadtree manner, and then each quadtree leaf can be further divided in a binary manner. This is shown on the right side of the diagram, where the solid line represents the quadtree decomposition phase and the dashed line represents the binary decomposition spatially embedded in the quadtree leaf. In intra-slice, the luma and chroma block partitioning structures are separated and determined independently. No further CU partitioning into prediction units or transform units is used. That is, each coding unit systematically consists of a single prediction unit (2N x 2N prediction unit partition type) and a single transform unit (no division into a transform tree).
[0036] 6 shows an exemplary set of extensions for coding unit partitioning. In one asymmetric bisection and tree splitting mode (ABT), a rectangular coding unit with size (w, h) (width and height) that is split by one of the asymmetric bisection splitting modes (e.g., HOR_UP (horizontal-up)) is split into rectangular coding units of size (w, h), respectively.
number
[0037] Using one or more embodiments of the new topology described above, significant coding efficiency improvements are achieved.
[0038] FIG. 7 shows an example of an intra prediction mode with two reference layers. In fact, in the improved syntax, new intra prediction modes are considered. The first new intra prediction mode is named multi-reference intra prediction (MRIP). This tool facilitates using multiple reference layers for intra prediction of a block. Generally, two reference layers are used for intra prediction, and each reference layer consists of a left reference array and an upper reference array. Each reference layer is constructed by reference sample permutation and then pre-filtered.
[0039] Each reference layer is used to build a prediction for the block, as is commonly done in HEVC or JEM. The final prediction is formed as a weighted average of the predictions made from the two reference layers. The prediction from the closest reference layer is given more weight than the prediction from the furthest layer. Commonly used weights are 3 and 1.
[0040] In the final part of the prediction process, multiple reference layers may be used to smooth boundary samples for a particular prediction mode using a mode-dependent process.
[0041] In an example embodiment, the video compression tool also includes an adaptive block size for sample adaptive offset (SAO), which is a loop filter specified in HEVC. In HEVC, the SAO process classifies the reconstructed samples of a block into several classes, and samples belonging to some classes are corrected using an offset. The SAO parameters can be coded block by block, or can be inherited only from the left or top neighboring blocks.
[0042] A further improvement is proposed by defining an SAO palette mode. The SAO palette maintains the same SAO parameters per block, but a block can inherit these parameters from all other blocks. This brings more flexibility to SAO by widening the range of possible SAO parameters for a block. The SAO palette consists of different sets of SAO parameters. For each block, an index is coded to indicate which SAO parameters to use.
[0043] 8 illustrates an example embodiment of a compensation mode for bidirectional illumination compensation. Illumination Compensation (IC) optionally takes into account spatial or temporal local illumination variations. allows correcting the block prediction samples (SMC) obtained by motion compensation. S IC =a i .S MC +b i
[0044] In the case of bi-prediction, the IC parameters are estimated using samples of two reference CU samples as depicted in Figure 8. First, the IC parameters (a, b) are estimated using the samples of two reference blocks Then, the IC parameters (ai,bi)i=0,1 between the reference block and the current block are derived as follows:
number
[0045] If the CU size is less than or equal to 8 in width or height, the IC parameters are estimated using a unidirectional process. If the CU size is greater than 8, the choice of using IC parameters derived from the bidirectional or unidirectional process is made by selecting the IC parameters that minimize the difference in the means of the IC-compensated reference blocks. Bidirectional optical flow (BIO) is enabled using a bidirectional illumination compensation tool.
[0046] The tools introduced above need to be selected and used by video encoder 100, and retrieved and used by video decoder 200. To that end, information representing the use of these tools is carried in the coded bitstream generated by the video encoder and retrieved by video decoder 200 in the form of high-level syntax information. To that end, a corresponding syntax is defined and described below. This syntax, which brings together the tools, enables improved coding efficiency.
[0047] This syntax structure uses the HEVC syntax structure as a basis and includes some additional syntax elements. In the table below, syntax elements signaled in italic bold font and highlighted in gray correspond to additional syntax elements according to an example embodiment. Note that syntax elements may have other forms or names than those shown in the syntax table while handling the same functionality. Syntax elements may exist at different levels (e.g., some syntax elements may be located within a sequence parameter set (SPS), some may be located within a picture parameter set (PPS)). ) may be placed within
[0048] Table 1 below shows the Sequence Parameter Set (SPS) and introduces new syntax elements inserted into this parameter set according to at least one example embodiment, more precisely: multi_type_tree_enabled_primary, log2_min_cu_size_minus2, log2_max_cu_size_minus4, log2_max_tu_size_minus2, sep_tree_mode_intra, multi_type_tree_enabled_secondary, sps_bdip_enabled_flag, sps_mrip_enabled_flag, use_erp_aqp_flag, use_high_perf_chroma_qp_table, abt_one_third_flag, sps_num_intra_mode_ratio.
[0049] [Table 1]
[0050] [Table 2]
[0051] [Table 3]
[0052] [Table 4]
[0053] The new syntax elements are defined as follows: - use_high_perf_chroma_qp_table: This syntax element specifies the Specifies a saturation QP table used to derive the QP used for decoding a given saturation component as a function of the base saturation QP associated with the saturation component of the considered slice. - multi_type_tree_enabled_primary: This syntax element is used in the coding slice. Indicates the type of partitioning allowed. In at least one embodiment, this syntax element is signaled at the sequence level (in the SPS) and therefore applies to all coded slices using this SPS. For example, a first value of this syntax element allows quad-tree and binary tree (QTBT) partitioning (NO-SPLIT, QT-SPLIT, HOR, and VER in FIG. 5), a second value allows QTBT + triple-tree (TT) partitioning, and so on. ) partitioning (NO-SPLIT, QT-SPLIT, HOR, VER, HOR_TRIPLE, and VER_TRIPLE in Figure 5), and the third value allows QTBT + Asymmetric Binary Tree (ABT) partitioning (NO-SPLIT, QT-SPLIT, HOR, VER, HOR-UP, HOR_DOWN, VER_LEFT, and VER_RIGHT in Figure 5). and the fourth value allows QTBT+TT+ABT partitioning (all split cases in Figure 5). - sep_tree_mode_intra: This syntax element is used to separate the luma and chroma blocks. If sep_tree_mode_intra is set to 1, it indicates whether a separate coding tree is used for If they are equal, then the luma and chroma blocks use independent coding trees, and therefore the partitioning of the luma and chroma blocks is independent. In one embodiment, this syntax element is signaled at the sequence level (in an SPS) and therefore applies to all coded slices using this SPS. - multi_type_tree_enabled_secondary: This syntax element specifies whether to enable or disable the partitioning. Indicates the types of tree-encoding allowed for chroma blocks in a coded slice. The possible values of this syntax element are generally the same as those of the syntax element multi_type_tree_enabled_primary. In at least one implementation, this syntax element: It is signaled at the sequence level (SPS) and therefore applies to all coded slices using this SPS. - sps_bdip_enabled_flag: This syntax element specifies which coded video sequences are considered. This indicates that bidirectional intra prediction as described in JVET-J0022 is allowed in the coded slices included in this slice. - sps_mrip_enabled_flag: This syntax element specifies the number of coded video bits that are taken into account. Indicates that multi-reference intra prediction tools are used in decoding the code slices contained in the stream. use_erp_aqp_flag: This syntax element indicates whether spatial adaptive quantization of ERP content in VR360 as described in JVET-J0022 is activated in the coded slice. abt_one_third_flag: This syntax element indicates whether asymmetric partitioning into two partitions with horizontal or vertical dimensions that are 1 / 3 and 2 / 3, or 2 / 3 and 1 / 3, respectively, of the horizontal or vertical dimension of the first CU is activated in the coding slice. sps_num_intra_mode_ratio: This syntax element specifies how to derive the number of intra predictions to be used for block sizes equal to a multiple of 3 in width or height.
[0054] Additionally, the following three syntax elements are used to control the size of the coding unit (CU) and the size of the transform unit (TU): In at least one embodiment, these syntax elements are signaled at the sequence level (in the SPS) and therefore apply to all coded slices using this SPS: - log2_min_cu_size_minus2 specifies the minimum CU size. - log2_max_cu_size_minus4 specifies the maximum CU size. - log2_max_tu_size_minus2 specifies the maximum TU size.
[0055] In another embodiment, the new syntax elements introduced above are introduced in the Picture Parameter Set (PPS) instead of the Sequence Parameter Set.
[0056] Table 2 below shows the Sequence Parameter Set (SPS) and introduces new syntax elements inserted into this parameter set according to at least one example embodiment, more precisely: - slice_sao_size_id: This syntax element specifies the size of the SAO for the coded slice. Specifies the size of the block to be applied. - slice_bidir_ic_enable_flag: indicates whether bidirectional illumination compensation is activated for the coded slice.
[0057] [Table 5]
[0058] [Table 6]
[0059] [Table 7]
[0060] [Table 8]
[0061] [Table 9]
[0062] Tables 3, 4, 5, and 6 below show the coding tree syntax modified to support additional partitioning modes according to at least one example embodiment. Specifically, the syntax that allows asymmetric partitioning is specified in coding_binary_tree(). New syntax elements, more precisely the following, are inserted into the coding tree syntax: - asymmetricSplitFlag, which indicates whether asymmetric bisection splits are allowed for the current CU, and - asymmetric_type indicates the type of asymmetric bisection split, whether the horizontal split is up or down. It can take two values to allow for a vertical split to be either left or right, or for a vertical split to be either right or left.
[0063] In addition, the parameter btSplitMode is interpreted differently.
[0064] [Table 10]
[0065] [Table 11]
[0066] [Table 12]
[0067] [Table 13]
[0068] [Table 14]
[0069] FIG. 9 illustrates the interpretation of bt-split-flag according to at least one example embodiment. The value of the parameter btSplitMode is specified using the tax elements asymmetricSplitFlag and asymmetric_type. btSplitMode indicates the bisection split mode applied to the current CU. btSplitMode is derived differently from the prior art. In JEM, it can take the values HOR, VER. In the improved syntax, it is derived using the partitioning scheme shown in Figure 6. It can also take the values HOR_UP, HOR_DOWN, VER_LEFT, and VER_RIGHT, which correspond to the
[0070] Tables 7, 8, 9, and 10 below show the coding unit syntax elements and introduce new syntax elements for bidirectional intra-prediction modes.
[0071] [Table 15]
[0072] [Table 16]
[0073] [Table 17]
[0074] [Table 18]
[0075] [Table 19]
[0076] [Table 20]
[0077] Some of the syntax elements described in this document are defined as arrays. For example, btSplitFlag is sorted by horizontal and vertical position within the picture and by coding block horizontal size. To simplify the notation, which is defined by an array of dimension 4 indexed by the vertical and horizontal sizes, these indices are not maintained in the semantic description of the syntax elements (e.g., btSplitFlag[x0][y0][cbSizeX][cbSizeY] is simply denoted as btSplitFlag).
[0078] FIG. 10 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 1000 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, encoders, transcoders, and servers. Elements of system 1000, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various embodiments, system 1000 is configured to perform one or more of the aspects described herein.
[0079] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 1010 may include internal memory, input / output interfaces, and other hardware components known in the art. The system 1000 may include various other circuitry. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 1040 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0080] System 1000 includes an encoder / decoder module 1030 configured to process data to provide, for example, coded or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that may be included within a device to perform coding and / or decoding functions. As is known, a device may include one or both of a coding module and a decoding module. Furthermore, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within the processor 1010 as a combination of hardware and software, as known to those skilled in the art.
[0081] Program code loaded into the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and later loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items while performing the processes described herein. Such stored items may include, but are not limited to, input video, decoded video, or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, expressions, operations, and arithmetic logic.
[0082] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040 (e.g., dynamic volatile memory and / or non-volatile flash memory). In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as MPEG-2, HEVC, or Versatile Video Coding (VVC).
[0083] Input to the elements of system 1000 may be provided using various input devices as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted over the air, for example by a broadcast station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0084] In various embodiments, the input devices of block 1130 have associated respective input processing elements known in the art. For example, the RF section may be configured to (i) select a desired frequency (also called selecting a signal or bandlimiting a signal to a frequency band), (ii) downconvert the selected signal, (iii) convert the selected signal to a desired frequency band, and (iv) convert the selected signal to a desired frequency band. The RF section may be associated with elements necessary to (i) bandlimit the downconverted and bandlimited signal again to a narrower frequency band to select (for example) a signal frequency band, which may be referred to as a channel in the set-top box; (ii) demodulate the downconverted and bandlimited signal; (iii) perform error correction; and (iv) demultiplex to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs some of these functions, including downconverting the received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements (e.g., inserting amplifiers and analog-to-digital converters, etc.). In various embodiments, the RF section includes an antenna.
[0085] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via USB and / or HDMI connections. It is understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 1010, as desired. Similarly, aspects of the USB or HDMI interface processing may be implemented, as desired, within a separate interface IC or within processor 1010. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and an encoder / decoder 1030, which operates in combination with memory and storage elements to process the data stream as desired for display on an output device.
[0086] The various elements of system 1000 may be provided within an integrated housing in which the various elements may be interconnected and transmit data between them using any suitable connection scheme (e.g., internal buses known in the art), including an I2C bus, wiring, and printed circuit boards.
[0087] System 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. Communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 1060. Communication interface 1050 may include, but is not limited to, a modem or a network card, and communication channel 1060 may be implemented within a wired and / or wireless medium, for example.
[0088] In various embodiments, data is streamed to system 1000 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received by communication interface 1050 and communication channel 1060 suitable for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, enabling streaming usage and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that sends data via HDMI communication in input block 1130. Still other embodiments provide streamed data to system 1000 using an RF connection in input block 1130.
[0089] System 1000 may provide output signals to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. In various example embodiments, other peripheral devices 1120 may include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, the control signals may be transmitted via AV.Link , CEC, or other communication protocols that allow device-to-device control with or without user intervention. Output devices may be communicatively coupled to system 1000 via dedicated connections via respective interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speaker 1110 may be integrated in a single unit with other components of system 1000 in an electronic device such as a television. In various embodiments, display interface 1070 includes a display driver, such as a timing controller (T Con) chip. Included.
[0090] The display 1100 and speakers 1110 may alternatively be separate from one or more of the other components, for example, if the RF portion of the input 1130 is part of a separate set-top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output. The implementations described herein may be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even when described only in the context of a single form of implementation (e.g., described only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in an apparatus such as a processor, which refers generally to processing devices (e.g., including a computer, microprocessor, integrated circuit, or programmable logic device). Processors also include communication devices such as, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0091] FIG. 11 shows a flowchart of an example of an encoding method using a new encoding tool, according to an embodiment. Such an encoding method can be performed by the system 1000 described in FIG. 10, and more precisely, can be implemented by the processor 1010. In at least one embodiment, in step 1190, the processor 1010 selects the encoding tool to be used. The selection can be made using different techniques. The selection can be made by a user using encoding configuration parameters (typically flags) that indicate the tool (and corresponding parameters, if necessary) to be used during the encoding and decoding processes. In an example embodiment, the selection is made by setting a value in a file, and the flag is read by the encoding device, which interprets the value of the flag to select which tool to use. In an example embodiment, the selection is made by a manual action of a user selecting the encoding configuration parameters using a graphical user interface operating the encoder device. Once this selection is made, coding is performed in step 1193 using the selected coding tool from among several available coding tools, with the tool selection being signaled in a higher level syntax element (e.g., in the following SPS, PPS, or slice header).
[0092] FIG. 12 shows an example of a portion of a decoding method according to one embodiment using the new encoding tool. 12 shows a flowchart of a decoding method for a video signal. Such a decoding method can be performed by the system 1000 described in FIG. 10, and more precisely, can be implemented by the processor 1010. In at least one embodiment, in step 1200, the processor 1010 accesses a signal (e.g., received at an input interface or read from a media support) and extracts and analyzes high-level syntax elements to determine tools selected by the coding device. In step 1210, these tools are used to perform decoding, generating a decoded image that can, for example, be provided to or displayed on a device.
[0093] References to "one embodiment," or "an embodiment," or "one implementation," or "an embodiment," and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment," or "in an embodiment," or "in one embodiment," or "in an embodiment," and other variations thereof, in various places throughout this specification do not necessarily all refer to the same embodiment.
[0094] Additionally, this application or its claims may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0095] Additionally, this application or its claims may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, predicting information, or estimating information.
[0096] Additionally, the application or its claims may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory or optical media storage). Furthermore, "receiving" generally involves in some way an operation such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0097] For example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selection of only the first listed option (A), or selection of only the second listed option (B), or selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass selection of only the first listed option (A), or selection of only the second listed option (B), or selection of only the third listed option (C), or selection of only the first and second listed options (A and B), or selection of only the first and third listed options (A and C), or selection of only the second and third listed options (B and C), or selection of all three options (A, B, and C). This can be expanded to include as many items as are listed, as would be readily apparent to one skilled in this and related arts.
[0098] As will be apparent to those skilled in the art, implementations may generate various signals formatted to carry information that may be, for example, stored or transmitted. Information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
[0099] In variants of at least the first, second, third, fourth, fifth, sixth, seventh, and eighth aspects of an embodiment, the parameter representing the type of partitioning includes a first value of quadtree and binary tree partitioning, a second value of quadtree and binary tree partitioning + ternary tree partitioning, a third value of quadtree and binary tree partitioning + asymmetric binary tree partitioning, and a fourth value of quadtree and binary tree partitioning + ternary tree partitioning + asymmetric binary tree partitioning.
[0100] In a variant of at least the first, second, third, fourth, fifth, sixth, seventh, and eighth aspects of an embodiment, the parameters further include a size of a block to which the sample adaptive offset is applied with respect to the coded slice.
[0101] In a variant of at least the first, second, third, fourth, fifth, sixth, seventh, and eighth aspects of an embodiment, the type of intra prediction mode includes at least one of an intra bi-prediction mode and a multi-reference intra prediction mode.
Claims
1. 1. A video signal formatted to include information according to a video coding standard, the video signal including a bitstream having video content and high-level syntax information, the high-level syntax information comprising: a parameter indicating a type of intra-prediction mode; A parameter that indicates the type of partitioning; a parameter describing the block size adaptation of the sample adaptive offset loop filter; a parameter representing an illumination compensation mode for the bi-predictive block; A video signal including at least one parameter from the above.
2. The parameter representing the type of partitioning is: a first value for quadtree and binary tree partitioning; a second value of quadtree and binary tree partitioning plus ternary tree partitioning; a third value of quad-tree and binary tree partitioning plus asymmetric binary tree partitioning; a fourth value of quadtree and binary tree partitioning + ternary tree partitioning + asymmetric binary tree partitioning; 2. The video signal of claim 1, comprising:
3. 3. A video signal according to claim 1 or 2, further comprising a size of the block to which the sample adaptive offset is applied with respect to the coded slice.
4. 4. A video signal according to any one of claims 1 to 3, wherein the types of intra prediction modes include at least one of intra bi-prediction modes and multi-reference intra prediction modes.
5. 1. A storage medium having coded video signal data, the video signal data being coded and including a bitstream having video content and high level syntax information, the high level syntax information comprising: a parameter indicating a type of intra-prediction mode; A parameter that indicates the type of partitioning; a parameter describing the block size adaptation of the sample adaptive offset loop filter; a parameter representing an illumination compensation mode for the bi-predictive block; A storage medium containing at least one parameter from
6. 1. A video encoder for encoding picture data of at least one block in a picture, the encoding comprising: a parameter indicating a type of intra-prediction mode; A parameter that indicates the type of partitioning; a parameter describing the block size adaptation of the sample adaptive offset loop filter; a parameter representing an illumination compensation mode for the bi-predictive block; and 10. An apparatus comprising: a video encoder, wherein the parameters are inserted into a high-level syntax element of the coded picture data.
7. coding picture data of at least one block in a picture, comprising: a parameter indicating a type of intra-prediction mode; A parameter that indicates the type of partitioning; a parameter describing the block size adaptation of the sample adaptive offset loop filter; a parameter representing an illumination compensation mode for the bi-predictive block; and encoding the data using at least one parameter from inserting said parameters into a high-level syntax element of said coded picture data; 1. A method in a video encoder, comprising:
8. 1. A video decoder for decoding picture data of at least one block in a picture, wherein the decoding comprises: a parameter indicating a type of intra-prediction mode; A parameter that indicates the type of partitioning; a parameter describing the block size adaptation of the sample adaptive offset loop filter; a parameter representing an illumination compensation mode for the bi-predictive block; and 10. An apparatus comprising: a video decoder, wherein the parameters are obtained from high-level syntax elements of coded picture data.
9. obtaining parameters from high-level syntax elements of the coded picture data; decoding picture data for at least one block in a picture, a parameter indicating a type of intra-prediction mode; A parameter that indicates the type of partitioning; a parameter describing the block size adaptation of the sample adaptive offset loop filter; a parameter representing an illumination compensation mode for the bi-predictive block; and decoding the encoded data using at least one parameter from 11. A method in a video decoder, comprising:
10. The parameter representing the type of partitioning is: a second value of quadtree and binary tree partitioning plus ternary tree partitioning; a first value for quadtree and binary tree partitioning; a third value of quad-tree and binary tree partitioning plus asymmetric binary tree partitioning; a fourth value of quadtree and binary tree partitioning + ternary tree partitioning + asymmetric binary tree partitioning; 10. The method of claim 9, comprising:
11. The method of claim 9 or 10, further comprising the size of the block to which the sample adaptive offset is applied with respect to the coded slice.
12. The method of any one of claims 9 to 11, wherein the types of intra prediction modes include at least one of intra bi-prediction modes and multi-reference intra prediction modes.
13. A non-transitory computer readable medium containing data content generated by the method or apparatus of any one of claims 6 to 12.
14. A computer program comprising program code instructions executable by a processor for performing the steps of the method according to at least one of claims 7 or 9 to 12.
15. A computer program product stored on a non-transitory computer readable medium and comprising program code instructions executable by a processor for performing the steps of the method according to at least one of claims 7 or 9-12.
Citation Information
Patent Citations
Image encoding / decoding method and device
JP2019535211A
Image encoding / decoding method and device, and recording medium in which bitstream is stored
WO2018016823A1
Moving-image decoder, moving-image decoding method, moving-image encoder, moving-image encoding method, and computer readable recording medium
WO2018055910A1
Image processing device and image processing method
WO2018070267A1