Methods and devices for fine control of image encoding and decoding processes

By introducing aggregation high-level syntax elements into the bitstream portion, the problem of imprecise control of encoding tools and features in existing technologies is solved, enabling fine control and efficiency improvement in the video encoding process.

CN115769587BActive Publication Date: 2026-07-17INTERDIGITAL CE PATENT HOLDINGS SAS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2021-05-10
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing video coding schemes struggle to achieve fine-grained control over coding tools and features, making it difficult to accurately control the encoding and decoding processes through configuration file definitions.

Method used

By introducing high-level aggregation syntax elements in the bitstream section, explicit indications are made of whether encoding tools or features, such as multi-type trees, scaling matrices, long-term reference pictures, and maximum transform unit size, to be allowed, thus enabling fine-grained control over the encoding process.

Benefits of technology

It enables fine-grained control over the video encoding process, improving the efficiency and accuracy of encoding and decoding while reducing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769587B_ABST
    Figure CN115769587B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for decoding, the method comprising: obtaining (601) an encoded video stream including a bitstream portion, the bitstream portion aggregating high-level syntax elements, at least one of the syntax elements providing information indicating whether an encoding tool or feature corresponding to the high-level syntax element is permitted in the encoded video stream; and determining (603) from the high-level syntax elements included in the bitstream portion whether an encoding tool or feature is permitted to decode the encoded video stream, wherein the encoding tool or feature is at least one of a multi-type tree, a scaling matrix, a long-term reference picture, a maximum transform unit size equal to a predetermined highest possible maximum transform unit size, or a weighted prediction.
Need to check novelty before this filing date? Find Prior Art

Description

1. Technical Field

[0001] At least one embodiment of the present invention generally relates to methods and apparatus for image encoding and decoding, and more specifically, to methods for restricting the use of at least one encoding tool or feature. 2. Background Technology

[0002] To achieve high compression efficiency, video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in the video content. During encoding, images of the video content are divided into sample blocks (i.e., pixels), and these blocks are then partitioned into one or more sub-blocks, hereinafter referred to as original sub-blocks. Intra-frame or inter-frame prediction is then applied to each sub-block to utilize intra-frame or inter-frame image correlations. Regardless of the prediction method used (intra-frame or inter-frame), a predicted sub-block is determined for each original sub-block. The sub-blocks representing the difference between the original sub-blocks and the predicted sub-blocks (typically denoted as prediction error sub-blocks, prediction residual sub-blocks, or simply residual blocks) are then transformed, quantized, and entropy-coded to generate the encoded video stream. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to the transform, quantization, and entropy coding.

[0003] Compared to earlier video compression methods such as MPEG-1 (ISO / CEI-11172), MPEG-2 (ISO / CEI 13818-2), or MPEG-4 / AVC (ISO / CEI 14496-10), the complexity of video compression methods has increased significantly. In fact, many new coding tools have emerged, or existing coding tools have been improved upon in previous generations of video compression standards (for example, in the international standard called Universal Video Coding (VVC), being developed by a joint collaborative group of ITU-T and ISO / IEC experts known as the Joint Video Experts Group (JVET), or in the standard HEVC (ISO / IEC 23008-2 – MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265)).

[0004] Not all possible encoding tools / features need to be activated during the encoding process. The activation / deactivation of some encoding tools can be controlled, for example, using high-level syntax elements such as constraint flags. Constraint flags are used to define configuration files / sub-configuration files where certain encoding tools / features are deactivated. Most encoding tools / features are associated with constraint flags. However, it can be noted that constraint flags are missing for multiple tools / features. These missing constraint flags make it difficult to provide fine-grained control over the encoding and decoding process when defining configuration files.

[0005] The goal is to propose a solution that allows for the simple definition of configuration files / sub-configuration files that enable fine-grained control over the coding process. 3. Summary of the Invention

[0006] In a first aspect, one or more embodiments of the present invention provide a method for decoding, the method comprising: obtaining an encoded video stream comprising a bitstream portion comprising aggregated high-level syntax elements, at least one of the syntax elements providing information indicating whether an encoding tool or feature corresponding to the high-level syntax elements is permitted in the encoded video stream; and determining from the high-level syntax elements included in the bitstream portion whether an encoding tool or feature is permitted to decode the encoded video stream, wherein the encoding tool or feature is at least one of a multi-type tree, a scaling matrix, a long-term reference picture, a maximum transform unit size equal to a predetermined highest possible maximum transform unit size, or a weighted prediction.

[0007] In a second aspect, one or more embodiments of the present invention provide a method for encoding, the method comprising: obtaining a video sequence to be encoded and a set of encoding constraints; and setting values ​​of high-level syntax elements in a bitstream portion comprising aggregated high-level syntax elements according to data representing the set of encoding constraints, at least one of the syntax elements providing information indicating whether encoding tools or features corresponding to the high-level syntax elements are permitted to be used to encode the video sequence, wherein the encoding tools or features are at least one of a multi-type tree, a scaling matrix, a long-term reference picture, a maximum transform unit size equal to a predetermined highest possible maximum transform unit size, or a weighted prediction.

[0008] In a third aspect, one or more embodiments of the present invention provide an apparatus for decoding, the apparatus comprising: means for obtaining an encoded video stream comprising a bitstream portion including aggregated high-level syntax elements, at least one of the syntax elements providing information indicating whether an encoding tool or feature corresponding to the high-level syntax elements is permitted in the encoded video stream; and means for determining from the high-level syntax elements included in the bitstream portion whether an encoding tool or feature is permitted to decode the encoded video stream, wherein the encoding tool or feature is at least one of a multi-type tree, a scaling matrix, a long-term reference picture, a maximum transform unit size equal to a predetermined highest possible maximum transform unit size, or a weighted prediction.

[0009] In a fourth aspect, one or more embodiments of the present invention provide an apparatus for encoding, the apparatus comprising: means for obtaining a video sequence to be encoded and a set of encoding constraints; and means for setting values ​​of high-level syntax elements in a bitstream portion of aggregated high-level syntax elements according to data representing the set of encoding constraints, at least one of the syntax elements providing information indicating whether encoding of the video sequence is permitted using an encoding tool or feature corresponding to the high-level syntax element, wherein the encoding tool or feature is at least one of a multi-type tree, a scaling matrix, a long-term reference picture, a maximum transform unit size equal to a predetermined highest possible maximum transform unit size, or a weighted prediction.

[0010] In a fifth aspect, one or more embodiments of the present invention provide an apparatus comprising the device described in the third or fourth aspect.

[0011] In a sixth aspect, one or more embodiments of the present invention provide a signal comprising data representing a portion of a bitstream of aggregated high-level syntax elements, at least one of the syntax elements providing information indicating whether an encoding tool or feature corresponding to the high-level syntax element is permitted in the encoded video stream; wherein the encoding tool or feature is at least one of a multi-type tree, a scaling matrix, a long-term reference picture, a maximum transform unit size equal to a predetermined highest possible maximum transform unit size, or a weighted prediction.

[0012] In a seventh aspect, one or more embodiments of the present invention provide a computer program comprising program code instructions for implementing the method according to the first or second aspect.

[0013] In an eighth aspect, one or more embodiments of the present invention provide an information storage medium that stores program code instructions for implementing the method according to the first or second aspect. 4. Description of the attached drawings

[0014] Figure 1 An example of the partitions that the pixel images of the original video have gone through is shown;

[0015] Figure 2 The method for encoding a video stream, performed by the encoding module, is illustrated schematically.

[0016] Figure 3 A method for decoding encoded video streams (i.e., bitstreams) is illustrated schematically.

[0017] Figure 4AAn example of a hardware architecture for a processing module capable of implementing an encoding or decoding module is schematically shown, in which various aspects and implementation schemes are implemented;

[0018] Figure 4B A block diagram of an example system is shown, in which various aspects and implementation schemes are implemented;

[0019] Figure 5 A method for signaling the activation of certain encoding tools / features during the encoding process is schematically depicted; and,

[0020] Figure 6 A method for determining the activation of a tool during the decoding process is illustrated schematically. 5. Detailed Implementation

[0021] In the following description, some implementations use tools developed in the context of VVC or HEVC. However, these implementations are not limited to video encoding / decoding methods corresponding to VVC or HEVC, and are applicable to other video encoding / decoding methods, as well as image encoding / decoding methods in which certain encoding tools / features can be activated / deactivated.

[0022] about Figure 1 , Figure 2 and Figure 3 We describe a video compression method. This method uses many encoding tools / features. As mentioned above, the activation / deactivation of some encoding tools can be controlled using high-level syntax elements such as constraint flags. Constraint flags are aggregated in a bitstream section called `general_constraint_info`. Each constraint flag provides information indicating whether the corresponding tool is allowed in the encoded video stream. For example, the following constraint flags are defined in the bitstream section `general_constraint_info`:

[0023]

[0024]

[0025] Table Tab1

[0026] For most encoding tools / features, define constraint flags to disable them. For example, the constraint flag no_alf_constraint_flag specifies that ALF (Adaptive Loop Filtering) is disabled.

[0027] It can be noted that constraint flags are missing for several tools / features. The relevant tools / features are:

[0028] • Multi-type tree (MTT);

[0029] • Maximum transform unit size;

[0030] • Zoom list;

[0031] • Long-term reference image prediction;

[0032] • Weighted forecasting.

[0033] These tools are described in more detail below.

[0034] Figure 1 An example of the partitioning experienced by pixel sample 11 of the original video 10 is shown. Here, the sample is considered to consist of three components: one luminance component and two chrominance components. In this case, the sample corresponds to a pixel. However, the following embodiments are applicable to images composed of samples including another number of components (e.g., where the sample includes a grayscale sample of one component), or images composed of samples including three color components and a transparency component and / or a depth component.

[0035] The image is divided into multiple coded entities. First, as... Figure 1 As indicated by reference numeral 13, the image is divided into a grid of blocks called coding tree units (CTUs). A CTU consists of N×N luminance sample blocks and two corresponding chrominance sample blocks. N is typically a power of two, for example, a maximum of "128". Next, the image is divided into one or more groups of CTUs. For example, the image may be divided into one or more tile rows and tile columns, where a tile is a sequence of CTUs covering a rectangular area of ​​the image. In some cases, a tile may be divided into one or more bricks, each brick consisting of at least one row of CTUs within the tile. Above the concepts of tiles and bricks, there is another coding entity called a slice, which may contain at least one tile of the image or at least one brick of a tile.

[0036] exist Figure 1 In the example, as indicated by reference numeral 12, image 11 is divided into three slices S1, S2 and S3, each slice comprising multiple tiles (not shown).

[0037] like Figure 1As indicated by reference numeral 14, a CTU can be partitioned into a hierarchical tree of one or more sub-blocks called coding units (CUs). The CTU is the root (i.e., parent node) of the hierarchical tree and can be partitioned into multiple CUs (i.e., child nodes). If each CU is not further partitioned into smaller CUs, then each CU becomes a leaf of the hierarchical tree; or if each CU is further partitioned into smaller CUs (i.e., child nodes), then each CU becomes a parent node of the smaller CUs. Several types of hierarchical trees can be applied, including, for example, quadtrees, binary trees, and ternary trees. In a quadtree, a CTU (or CU) can be partitioned into four equal-sized square CUs (i.e., it can be the parent node of four equal-sized square CUs). In a binary tree, a CTU (or CU) can be partitioned horizontally or vertically into two equal-sized rectangular CUs. In a ternary tree, a CTU (or CU) can be partitioned horizontally or vertically into three rectangular CUs. For example, a CU with a height of N and a width of M is vertically (or horizontally) divided into a first CU with a height of N (or N / 4) and a width of M / 4 (or M), a second CU with a height of N (or N / 2) and a width of M / 2 (or M), and a third CU with a height of N (or N / 4) and a width of M / 4 (or M).

[0038] exist Figure 1 In the example, firstly, CTU 14 is partitioned into "4" square CUs using quadtree partitioning. The top-left CU is a leaf of the hierarchical tree because it is not further partitioned, meaning it is not the parent node of any other CU. The top-right CU is further partitioned into "4" smaller square CUs using quadtree partitioning again. The bottom-right CU is vertically partitioned into "2" rectangular CUs using binary tree partitioning. The bottom-left CU is vertically partitioned into "3" rectangular CUs using ternary tree partitioning.

[0039] The combination of binary and ternary trees is called a Multi-Type Tree (MTT). MTT is a relatively new encoding tool, but it does not define a single high-level syntax such as a Sequence Parameter Set (SPS) level flag or constraint flag. For MTT, three SPS level syntax elements are defined as follows:

[0040]

[0041] Table TAB2

[0042] The semantics of these two SPS-level syntax elements are as follows:

[0043] `sps_max_mtt_hierarchy_depth_intra_slice_luma` specifies the default maximum hierarchical depth of coding units generated by multi-type tree partitioning of quad-leaf trees in slices of the reference SPS with a `sh_slice_type` equal to "2" (I). When `sps_partition_constraints_override_enabled_flag` is equal to "1", the default maximum hierarchical depth can be overridden by `ph_max_mtt_hierarchy_depth_intra_slice_luma`, which exists in the `PH` of the reference SPS. The value of `sps_max_mtt_hierarchy_depth_intra_slice_luma` should be in the range of "0" to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive.

[0044] `sps_max_mtt_hierarchy_depth_inter_slice` specifies the default maximum hierarchical depth of coding units generated by multi-type tree segmentation of quad-leaf trees in slices of the reference SPS with a `sh_slice_type` equal to "0" (B) or "1" (P). When `sps_partition_constraints_override_enabled_flag` is equal to "1", the default maximum hierarchical depth can be overridden by `ph_max_mtt_hierarchy_depth_inter_slice`, which exists in the `PH` of the reference SPS. The value of `sps_max_mtt_hierarchy_depth_inter_slice` should be in the range of "0" to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive.

[0045] `sps_max_mtt_hierarchy_depth_intra_slice_chroma` specifies the default maximum hierarchical depth of chroma coding units generated by multi-type tree partitioning of chroma quadtree leaves with a `treeType` equal to `DUAL_TREE_CHROMA` in slices of the reference SPS with a `sh_slice_type` equal to "2" (I). When `sps_partition_constraints_override_enabled_flag` equals "1", the default maximum hierarchical depth can be overridden by `ph_max_mtt_hierarchy_depth_chroma` present in the `PH` of the reference SPS. The value of `sps_max_mtt_hierarchy_depth_intra_slice_chroma` should be in the range of "0" to 2*(CtbLog2SizeY - MinCbLog2SizeY), inclusive. When not present, the value of `sps_max_mtt_hierarchy_depth_intra_slice_chroma` is inferred to be equal to "0".

[0046] To completely disable MTT, the three SPS level syntax elements sps_max_mtt_hierarchy_depth_intra_slice_luma, sps_max_mtt_hierarchy_depth_inter_slice, and sps_max_mtt_hierarchy_depth_intra_slice_chroma should be zero. A simple way to disable MTT is expected.

[0047] During image encoding, partitioning is adaptive, with each CTU being partitioned to optimize the compression efficiency of the CTU criteria.

[0048] In some compression methods, the concepts of prediction units (PU) and transform units (TU) emerge. In this case, the coding entity used for prediction (i.e., PU) and the coding entity used for transform (i.e., TU) can be sub-partitions of the CU. For example, as... Figure 1 This indicates that a CU of size 2N×2N can be divided into PUs 1411 of size N×2N or size 2N×N. Additionally, the CU can be divided into “4” TUs 1412 of size N×N or “16” TUs of size (N / 2)×(N / 2).

[0049] In some implementations, the maximum transform unit (TU) size is defined as, for example, equal to "64" or "32". For certain configurations, it is important to limit the maximum transform size to "32" to reduce overall complexity. Regarding the maximum transform unit size, the following SPS-level flag, `sps_max_luma_transform_size_64_flag`, is defined. Its semantics are:

[0050] A value of "1" for `sps_max_luma_transform_size_64_flag` specifies a maximum transform size of "64" in the luminance sample. A value of "0" for `sps_max_luma_transform_size_64_flag` specifies a maximum transform size of "32" in the luminance sample. When it does not exist, the value of `sps_max_luma_transform_size_64_flag` is inferred to be "0".

[0051] The flag `sps_max_luma_transform_size_64_flag` fixes the highest possible maximum TU size to "64". However, the highest possible maximum TU size can be fixed to other values, such as "128" or "256".

[0052] In this application, the terms "block" or "image block" or "subblock" may be used to refer to any of CTU, CU, PU, ​​and TU. Additionally, the terms "block" or "image block" may be used to refer to macroblocks, partitions, and subblocks as specified in MPEG-4 / AVC or other video coding standards, and more generally to an array of samples of numerous sizes.

[0053] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture”, “subpicture”, “slice” and “frame” are used interchangeably.

[0054] Figure 2 A method for encoding a video stream, performed by an encoding module, is illustrated schematically. Variations of this encoding method are envisioned, but for clarity, they are described below. Figure 2 The method used for encoding is given, but not all expected variants are described.

[0055] The encoding of the current original image 201 during step 202 begins with a partition of the current original image 201, as per the relevant information. Figure 1 As described. Therefore, the current image is partitioned into 201 blocks: CTU, CU, PU, ​​TU, etc. For each block, the coding module determines the coding mode between intra-frame prediction and inter-frame prediction.

[0056] The intra-frame prediction, represented by step 203, includes predicting samples of the current block from a prediction block according to an intra-frame prediction method. This prediction block is derived from samples of a reconstructed block located near the causal relationship of the current block to be encoded. The result of the intra-frame prediction is a prediction direction indicating which samples from nearby blocks are used, and a residual block obtained by calculating the difference between the current block and the prediction block.

[0057] Inter-frame prediction involves predicting samples for the current block from sample blocks (called reference blocks) of images preceding or following the current image (called the reference image). Two types of reference images are defined: "Short-Term Reference Picture (STRP)" and "Long-Term Reference Picture (LTRP)". Both STRP and LTRP can be used as reference images for the current block. The corresponding SPS level flag, called `sps_long_term_ref_pics_flag`, is defined with the following semantics:

[0058] A value of "0" for sps_long_term_ref_pics_flag indicates that no LTRP is used for inter-frame prediction of any coded pictures in CLVS. A value of "1" for sps_long_term_ref_pics_flag indicates that LTRP can be used for inter-frame prediction of one or more coded pictures in CLVS.

[0059] During the encoding of the current block according to the inter-frame prediction method, the motion estimation step 204 determines the block in the reference image that is closest to the current block based on a similarity criterion. During step 204, a motion vector indicating the location of the reference block in the reference image is determined. This motion vector is used during the motion compensation step 205, during which a residual block is calculated in the form of the difference between the current block and the reference block.

[0060] In the first video compression standard, the aforementioned one-way inter-frame prediction mode was the only available inter-frame mode. As video compression standards have evolved, the family of inter-frame modes has grown significantly and now includes many different inter-frame modes. One example of a tool included in the family of inter-frame modes is weighted prediction. Weighted prediction (WP) is a coding tool that allows for efficient coding of video content with attenuation. WP allows signaling weighting parameters (weights and offsets) to each reference image in each of the reference image lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference images are applied. Two SPS level flags control weighted prediction as follows:

[0061]

[0062] Table TAB3

[0063] The meaning of these symbols is:

[0064] A value of "1" for `sps_weighted_pred_flag` specifies that weighted predictions can be applied to P slices of the reference SPS. A value of "0" for `sps_weighted_pred_flag` specifies that weighted predictions should not be applied to the P slices of the reference SPS.

[0065] A value of "1" for `sps_weighted_bipred_flag` specifies that explicit weighted predictions can be applied to B slices of the reference SPS. A value of "0" for `sps_weighted_bipred_flag` specifies that explicit weighted predictions are not applied to the B slices of the reference SPS.

[0066] To disable weighted prediction, both flags must be set to zero. A more convenient way to control the activation / deactivation of weighted prediction is desirable.

[0067] During the selection step 206, the coding module selects the prediction mode that optimizes compression performance from the tested prediction modes (intra-frame prediction mode, inter-frame prediction mode) according to the rate / distortion criterion (i.e., RDO criterion).

[0068] When a prediction mode is selected, the residual block is transformed during step 207 and quantized during step 209. During quantization, in the transform domain, the transform coefficients are weighted by a scaling matrix in addition to the quantization parameters. The scaling matrix is ​​a coding tool that allows some frequencies to be supported at the expense of others. Generally, low frequencies are advantageous. Some video compression methods allow the application of user-defined scaling matrices (also known as scaling lists) instead of the default scaling matrix. In this case, the parameters of the scaling matrix need to be emitted to the decoder. Note that this feature is not needed in many coding scenarios. In fact, in most cases, all transform coefficients are treated equally. Regarding the scaling matrix, the SPS level flag `sps_explicit_scaling_list_enabled_flag` is defined to disable the scaling matrix. Its semantics are:

[0069] A value of "1" for `sps_explicit_scaling_list_enabled_flag` specifies that an explicit scaling list, signaled in the scaling list APS, is used during scaling of transform coefficients when decoding of slices is enabled for CLVS (Coding Layer Video Sequence). A value of "0" for `sps_explicit_scaling_list_enabled_flag` specifies that an explicit scaling list is used during scaling of transform coefficients when decoding of slices is disabled for CLVS.

[0070] It should be noted that the encoding module can skip the transformation and apply quantization directly to the untransformed residual signal.

[0071] When the current block is encoded according to the intra-frame prediction mode, during step 210, the prediction direction and the transformed and quantized residual block are encoded by the entropy encoder.

[0072] When encoding the current block according to the inter-frame prediction mode, motion data associated with the inter-frame prediction mode is encoded in step 208.

[0073] Generally speaking, two modes can be used to encode motion data, namely AMVP (Adaptive Motion Vector Prediction) mode and merge mode.

[0074] The AMVP model basically includes signaling a reference image for predicting the current block, a motion vector prediction index, and a motion vector difference (also known as a motion vector residual).

[0075] The merging pattern includes a signal indicating the index of some motion data collected in a list of motion data prediction values. This list consists of either "5" or "7" candidates and is constructed in the same way on both the decoder and encoder sides. Therefore, the merging pattern aims to derive some motion data taken from the merging list. The merging list typically contains motion data associated with some spatially and temporally adjacent blocks, which are available in their reconstructed state when processing the current block.

[0076] Once predicted, the motion information is then encoded by an entropy encoder along with the transformed and quantized residual blocks during step 210. It should be noted that the encoding module can bypass the transform and quantization; that is, entropy encoding can be applied to the residuals without applying the transform or quantization process. The result of the entropy encoding is inserted into the encoded video stream (i.e., the bitstream) 211.

[0077] It should be noted that entropy encoders can be implemented in the form of context-adaptive binary arithmetic encoders (CABAC). CABAC encodes binary symbols, which maintains low complexity and allows for probabilistic modeling of the more frequently used bits of any symbol.

[0078] After quantization step 209, the current block is reconstructed so that the pixels corresponding to that block are available for future prediction. This reconstruction stage is also called the prediction loop. Therefore, inverse quantization is applied to the transformed and quantized residual block during step 212, and inverse transform is applied during step 213. The prediction block of the current block is reconstructed based on the prediction mode used for the current block obtained during step 214. If the current block is encoded according to an inter-frame prediction mode, the encoding module applies motion compensation to a reference block using the motion information of the current block during step 216, where appropriate. If the current block is encoded according to an intra-frame prediction mode, the reference block of the current block is reconstructed during step 215 using the prediction direction corresponding to the current block. The reference block and the reconstructed residual block are added together to obtain the reconstructed current block.

[0079] Following reconstruction, during step 217, an in-loop post-filter designed to reduce coding artifacts is applied to the reconstruction block. This post-filter is called an in-loop post-filter because it occurs in the prediction loop to obtain the same reference image at the encoder as at the decoder, thereby avoiding drift between the encoding and decoding processes. Examples of in-loop post-filters include deblocking filtering, SAO (Sample Adaptive Shift) filtering, and adaptive loop filtering (ALF) with block-based filter adaptation.

[0080] During entropy coding step 210, parameters representing the activation or deactivation of the in-loop deblocking filter and the characteristics of the in-loop deblocking filter when activated are introduced into the encoded video stream 211.

[0081] When reconstructing a block, the block is inserted into the reconstructed image stored in the decoded image buffer (DPB) 219 during step 218. The reconstructed image thus stored can then be used as a reference image for other images to be encoded.

[0082] Figure 3 The schematic depiction is used for the purpose of... Figure 2 The described method is for decoding an encoded video stream (i.e., a bitstream) 211. The decoding method is performed by a decoding module. Variations of this decoding method are contemplated, but for clarity, the following description is preferred. Figure 3 The method used for decoding is described, but not all expected variations are described.

[0083] Decoding is performed block by block. For the current block, it begins with entropy decoding during step 310. Entropy decoding allows the acquisition of the prediction pattern for the current block.

[0084] If the current block has already been encoded according to the intra-prediction mode, entropy decoding allows for the acquisition of information representing the intra-prediction direction and the residual block.

[0085] If the current block has already been encoded according to the inter-frame prediction mode, entropy decoding allows for the acquisition of information representing motion data and residual blocks. Where appropriate, during step 308, motion data is reconstructed for the current block according to either the AMVP mode or the merge mode. In the merge mode, the motion data obtained through entropy decoding includes indices from a candidate list of motion vector prediction values. The decoding module applies the same process as the encoding module to reconstruct the candidate lists for both the regular merge mode and the sub-block merge mode. Using the reconstructed lists and indices, the decoding module is able to retrieve motion vectors for predicting the block's motion vectors.

[0086] The decoding method includes steps 312, 313, 315, 316, and 317, which are identical in all respects to steps 212, 213, 215, 216, and 217 of the encoding method. At the encoding module level, step 214 includes a mode selection process that evaluates each mode according to a rate distortion criterion and selects the optimal mode, while step 314 only includes reading information representing the selected mode from bitstream 211. In step 318, the decoded block is saved in the decoded image and the decoded image is stored in DPB 319. When the decoding module decodes a given image, the image stored in DPB 319 is the same as the image stored in DPB 219 by the encoding module during the encoding of the given image. The decoded image can also be output by the decoding module for, for example, display.

[0087] Figure 4A The illustration schematically shows examples of hardware architectures for processing module 40, modified according to different aspects and implementation schemes, capable of implementing either an encoding module or a decoding module, which can respectively implement... Figure 2 Methods for encoding and Figure 3The method for decoding. As a non-limiting example, the processing module 40 includes the following items connected by a communication bus 405: a processor or CPU (central processing unit) 400 containing one or more microprocessors, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture; random access memory (RAM) 401; read-only memory (ROM) 402; a storage unit 403, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drive and / or optical disk drive, or storage media reader, such as SD (Secure Digital) card reader and / or hard disk drive (HDD) and / or network accessible storage device; at least one communication interface 404 for exchanging data with other modules, devices, or equipment. The communication interface 404 may include, but is not limited to, a transceiver configured to transmit and receive data through a communication channel. Communication interface 404 may include, but is not limited to, a modem or network card.

[0088] If processing module 40 implements a decoding module, then communication interface 404 enables, for example, processing module 40 to receive encoded video streams and provide decoded video streams. If processing module 40 implements an encoding module, then communication interface 404 enables, for example, processing module 40 to receive raw image data to be encoded and provide encoded video streams.

[0089] Processor 400 is capable of executing instructions loaded into RAM 401 from ROM 402, external memory (not shown), storage media, or a communication network. When processing module 40 is powered on, processor 400 is capable of reading instructions from RAM 401 and executing those instructions. These instructions form a computer program that enables, for example, processor 400 to implement... Figure 3 The described decoding method or about Figure 2 The encoding method described herein includes the following aspects and implementation schemes as described in this document.

[0090] All or part of the algorithms and steps of the encoding or decoding method may be implemented in software by executing a set of instructions by a programmable machine such as a DSP (Digital Signal Processor) or a microcontroller, or in hardware by a machine or dedicated component such as an FPGA (Field Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).

[0091] Figure 4BA block diagram illustrating an example of a system 4 in which various aspects and embodiments are implemented is shown. System 4 may be embodied as a device including the various components described below and configured to perform one or more aspects and embodiments described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 4 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, system 4 includes a processing module 40 implementing a decoding module or an encoding module. However, in another embodiment, system 4 may include a first processing module 40 implementing a decoding module and a second processing module 40 implementing an encoding module, or a single processing module 40 implementing both a decoding module and an encoding module. In various embodiments, system 40 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 4 is configured to implement one or more aspects of the aspects described in this document.

[0092] System 4 includes at least one processing module 40, which is capable of implementing one or both of an encoding module or a decoding module.

[0093] Inputs to processing module 40 may be provided by various input modules as shown in box 42. Such input modules include, but are not limited to: (i) a radio frequency (RF) module that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) a component (COMP) input module (or a set of COMP input modules); (iii) a universal serial bus (USB) input module; and / or (iv) a high-definition multimedia interface (HDMI) input module. Figure 4B Other examples not shown include composite video.

[0094] In various embodiments, the input module of block 42 has associated corresponding input processing elements as known in the art. For example, the RF module may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) re-band-limiting the signal to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF module of various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting received signals to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to baseband. In one set-top box implementation, the RF module and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various implementations rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various implementations, the RF module includes an antenna.

[0095] Additionally, the USB and / or HDMI modules may include corresponding interface processors for connecting System 4 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processing module 40. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processing module 40. Demodulation, error correction, and demultiplexing streams are provided to processing module 40.

[0096] Various components of System 4 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected using suitable connection arrangements (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards) and data can be transferred between these components. For example, in System 4, processing module 40 is interconnected with other components of System 4 via bus 405.

[0097] The communication interface 404 of the processing module 40 allows the system 4 to communicate over the communication channel 41. For example, the communication channel 41 can be implemented in a wired and / or wireless medium.

[0098] In various implementations, a wireless network such as Wi-Fi, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), is used to stream or otherwise provide data to system 4. In these implementations, the Wi-Fi signal is received via a communication channel 41 and a communication interface 404 suitable for Wi-Fi communication. The communication channel 41 in these implementations is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other cloud-based communications. Other implementations use a set-top box that delivers data via an HDMI connection to input block 42 to provide streaming data to system 4. Still other implementations use an RF connection to input block 42 to provide streaming data to system 4. As mentioned above, various implementations provide data in a non-streaming manner. Additionally, various implementations use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0099] System 4 can provide output signals to various output devices, including a display 46, a speaker 47, and other peripheral devices 48. The display 46 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 46 can be used in televisions, tablets, laptops, cellular phones (mobile phones), or other devices. The display 46 can also be integrated with other components (e.g., as in a smartphone) or standalone (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 46 include one or more of a standalone digital video disc (or digital universal disc, both terms being DVR), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 48 that provide functionality based on the output of System 4. For example, an optical disc player performs the function of playing the output of System 4.

[0100] In various embodiments, control signals are transmitted between system 4 and display 46, speaker 47, or other peripheral devices 48 using signaling protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices are communicatively coupled to system 4 via dedicated connections through corresponding interfaces 43, 44, and 45. Alternatively, output devices can be connected to system 4 via communication interface 404 using communication channel 41. Display 46 and speaker 47 can be integrated into a single unit with other components of system 4 in electronic devices such as televisions. In various embodiments, display interface 43 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0101] For example, if the RF portion of input 42 is part of a separate set-top box, then display 46 and speaker 47 may optionally be separate from one or more other components. In various embodiments where display 46 and speaker 47 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0102] Various specific implementations involve decoding. As used in this application, "decoding" may encompass all or part of a process performed, for example, on a received encoded video stream, to produce a final output suitable for display. In various implementations, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and prediction. In various implementations, such a process also includes, or alternatively includes, processes performed by a decoder of the various specific implementations or embodiments described in this application, such as determining whether an MTT is activated, a scaling matrix, a long-term reference picture, a maximum TU size equal to "32", or weighted prediction.

[0103] Whether the phrase “decoding process” specifically refers to a subset of operations or broadly refers to a wider decoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0104] Various specific implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass, for example, all or part of the process performed on an input video sequence to produce an encoded video stream. In various implementations, such processes include one or more processes typically performed by an encoder, such as partitioning, prediction, transform, quantization, in-loop post-filtering, and entropy coding. In various implementations, such processes also include, or alternatively include, processes performed by an encoder of the various specific implementations or embodiments described herein, such as for activating / deactivating MTT, scaling matrices, long-term reference pictures, a maximum TU size equal to “32,” or weighted prediction.

[0105] Whether the phrase “encoding process” specifically refers to a subset of operations or broadly refers to a wider encoding process will be clear based on the specific context of the description and is believed to be well understood by those skilled in the art.

[0106] It should be noted that the names of syntax elements, tags, containers, and encoding tools used in this article are descriptive terms. Therefore, they do not preclude the use of other syntax element, tag, container, or encoding tool names.

[0107] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0108] Various implementation schemes refer to rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered. Rate distortion optimization is generally formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different approaches exist to solve the rate distortion optimization problem. For example, these methods may be based on extensive testing of all encoding options (including all considered modes or encoding parameter values) and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to reduce encoding complexity, particularly for calculating approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed residual signal. A hybrid of these two approaches can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.

[0109] The specific embodiments and aspects described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even if discussed only in the context of a single form of specific embodiment (e.g., discussed only as a method), specific embodiments of the discussed features may be implemented in other forms (e.g., apparatus or program). Apparatus may be implemented, for example, in suitable hardware, software, and firmware. Methods may be implemented, for example, in a processor, which generally refers to a processing device.

[0110] The processing device includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0111] The reference to "an implementation scheme" or "implementation scheme" or "a specific implementation" or "specific implementation," and other variations thereof, means that the specific features, structures, characteristics, etc., described in connection with the implementation scheme are included in at least one implementation scheme. Therefore, the appearance of the phrase "in an implementation scheme" or "in an implementation scheme" or "in a specific implementation" or "in a specific implementation," and any other variations appearing throughout this application, do not necessarily refer to the same implementation scheme.

[0112] In addition, this application may involve "determining" various types of information. Determining information may include, for example, estimated information, calculated information, predicted information, information inferred from other information, information retrieved from memory, or information obtained, for example, from another device, module, or user, one or more of these.

[0113] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, inferring information, or estimating information, or one or more of these.

[0114] Furthermore, this application may relate to "receiving" various types of information. Like "access," "receiving" is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or more. Moreover, "receiving" typically involves one or more of the following during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, inferring information, or estimating information.

[0115] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” “at least one of A and B,” and “one or more of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one,” “one or more” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As a further example, in the cases of “A, B, and / or C,” “at least one of A, B, and C,” and “one or more of A, B, and C,” such phrases are intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to as many of the listed items as possible.

[0116] Moreover, as used herein, the term "signaling" refers to (among other things) instructing the corresponding decoder to do something. For example, in some implementations, the encoder signals constraint flags indicating whether the MTT, scaling matrix, long-term reference picture, maximum TU size equal to 32, or weighted prediction is activated. Thus, in one implementation, the same parameters are used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters and others, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. Bit savings are achieved in various implementations by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various implementations, information is signaled to the corresponding decoder using one or more syntax elements, tags, etc. Although the verb form of the term "signal" has been used above, the term "signal" can also be used as a noun herein.

[0117] It will be apparent to those skilled in the art that the embodiments can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the embodiments. For example, a signal may be formatted to carry an encoded video stream of the embodiment. Such signals may be formatted as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the encoded video stream and modulating a carrier wave using the encoded video stream. The information carried by the signal may be, for example, analog or digital information. It is known that signals can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0118] Figure 5 A method for signaling the activation of certain encoding tools / features during the encoding process is illustrated schematically. Figure 5 Methods such as using Figure 2 The described method is performed before encoding the first image of the video sequence.

[0119] In the implementation, the bitstream section `general_constraint_info` is modified as follows to include new constraint flags for activating / deactivating the MTT, scaling matrix, long-term reference image, maximum TU size equal to "64", or weighted prediction:

[0120]

[0121]

[0122]

[0123] Table TAB4

[0124] In Table TAB4, the added constraint symbols are indicated in bold.

[0125] The meanings of these symbols are as follows:

[0126] A `no_mtt_constraint_flag` value of "1" specifies that `sps_max_mtt_hierarchy_depth_intra_slice_luma`, `sps_max_mtt_hierarchy_depth_inter_slice`, and `sps_max_mtt_hierarchy_depth_intra_slice_chroma` should be equal to "0". A `no_mtt_constraint_flag` value of "0" does not impose this constraint. In other words, if `no_mtt_constraint_flag` is equal to "1", MTT is disabled.

[0127] A `max_luma_transform_size_32_constraint_flag` equal to "1" specifies that `sps_max_luma_transform_size_64_flag` should be equal to "0". A `max_luma_transform_size_32_constraint_flag` equal to "0" does not impose this constraint. In other words, `max_luma_transform_size_32_constraint_flag` equal to "1" has a maximum TU size of "32" and does not allow TUs of size "64".

[0128] A `no_scaling_list_constraint_flag` value of "1" specifies that `sps_explicit_scaling_list_enabled_flag` should be equal to "0". A `no_scaling_list_constraint_flag` value of "0" does not impose this constraint. In other words, when `no_scaling_list_constraint_flag` is equal to "1", the scaling matrix is ​​disabled, and the use of non-default scaling matrices is disabled.

[0129] A `no_long_term_ref_pic_constraint_flag` value of "1" specifies that `sps_long_term_ref_pics_flag` should be equal to "0". A `no_long_term_ref_pic_constraint_flag` value of "0" does not impose this constraint. In other words, when `no_long_term_ref_pic_constraint_flag` is equal to "1", no LTRP is used for inter-frame prediction. When `intra_only_constraint_flag` is equal to "1", the value of `no_long_term_ref_pic_constraint_flag` should be equal to "1".

[0130] A `no_weighted_pred_constraint_flag` value of "1" specifies that `sps_weighted_pred_flag` and `sps_weighted_bipred_flag` should be equal to "0". A `no_weighted_pred_constraint_flag` value of 0 does not impose this constraint. In other words, when `no_weighted_pred_constraint_flag` is equal to "1", weighted prediction is disabled. When `intra_only_constraint_flag` is equal to "1", the value of `no_weighted_pred_constraint_flag` should be equal to "1".

[0131] Back Figure 5 In the method, in step 501, processing module 40 obtains the video sequence to be encoded. During step 501, processing module 40 also receives data representing a profile / sub-profile or, for example, a set of encoding constraints fixed by the user.

[0132] In step 502, processing module 40 sets the values ​​of constraint flags in the bitstream portion general_constraint_info. These constraint flags are set based on data representing a profile / sub-profile or a set of encoding constraints, or default values. For example, if MTT is disabled (specifically, maximum transform unit size is "32", scaling matrix usage is disabled, LTRP usage is disabled, and weighted prediction is disabled), then processing module 40 sets the value of constraint flag no_mtt_constraint_flag (specifically, max_luma_transform_size_32_constraint_flag, no_scaling_list_constraint_flag, no_long_term_ref_pic_constraint_flag, and no_weighted_pred_constraint_flag) to "1". If MTT activation is allowed at the general_constraint_info level (specifically, the maximum transform unit size "64" is allowed at the general_constraint_info level, scaling matrices are allowed at the general_constraint_info level, LTRP is allowed at the general_constraint_info level, and weighted prediction is allowed at the general_constraint_info level), then the processing module 40 sets the constraint flag no_mtt_constraint_flag (specifically, max_luma_transform_size_32_constraint_flag, no_scaling_list_constraint_flag, no_long_term_ref_pic_constraint_flag, and no_weighted_pred_constraint_flag) to "0".

[0133] Figure 6 A method for determining the activation of a tool during the decoding process is illustrated schematically. Figure 6 Methods such as receiving the encoded video stream and using Figure 3 The described method is executed before decoding the first image of the encoded video stream. The received encoded video stream includes the bitstream portion general_constraint_info.

[0134] In step 601, the processing module 40 obtains the encoded video stream including the bitstream portion general_constraint_info.

[0135] In step 602, the processing module 40 parses the bitstream portion general_constraint_info.

[0136] In step 603, processing module 40 determines whether MTT, scaling matrix, LTRP, maximum transform unit size "64", or weighted prediction is allowed from the constraint flags contained in the bitstream portion general_constraint_info. To do this, processing module 40 determines whether the constraint flags no_mtt_constraint_flag, no_scaling_list_constraint_flag, max_luma_transform_size_32_constraint_flag, no_long_term_ref_pic_constraint_flag, or no_weighted_pred_constraint_flag exist in the bitstream portion general_constraint_info, and if they exist, determines the values ​​of these flags. If no_mtt_constraint_flag (which, respectively, no_scaling_list_constraint_flag, max_luma_transform_size_32_constraint_flag, no_long_term_ref_pic_constraint_flag, and no_weighted_pred_constraint_flag) equals "1", then processing module 40 determines in step 605 that MTT (which, respectively, is not the default scaling matrix, the maximum TU size equal to "64", LTRP, and weighted prediction) is not allowed in the encoded video stream. In this case, decoding of the encoded video stream is performed without using unauthorized tools / features.

[0137] If no_mtt_constraint_flag (and respectively, no_scaling_list_constraint_flag, max_luma_transform_size_32_constraint_flag, no_long_term_ref_pic_constraint_flag, and no_weighted_pred_constraint_flag), the processing module 40 determines in step 604 that MTT (and respectively, scaling matrix, maximum TU size equal to 64, LTRP, and weighted prediction) is allowed at the general_constraint_info level in the bitstream portion. In this case, if this use is not otherwise prevented in the encoded video stream, the decoding of the encoded video stream can use the allowed tools / features.

[0138] Additionally, note that in the implementation scheme, if the intra_only_constraint_flag of the general constraint flags syntax of the VVC specification is equal to "1", then the newly introduced constraint flags no_long_term_ref_pic_constraint_flag and no_weighted_pred_constraint_flag are also equal to "1".

[0139] In fact, in the VVC specification, an intra_only_constraint_flag that is equal to "1" specifies that sh_slice_type should be equal to I(Intra). An intra_only_constraint_flag that is equal to "0" will not impose this constraint.

[0140] In fact, these two proposed constraint flags relate to the coding of inter-frame blocks, and therefore to the coding of inter-frame images. Therefore, they are not related to the VVC-coded bitstream in which intra_only_constraint_flag equals 1.

[0141] Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim classes and types:

[0142] • Includes a bitstream or signal that transmits a syntax for information generated according to any of the embodiments described;

[0143] • Inserting syntax elements into the signaling allows the decoder to adapt the decoding process in a manner corresponding to that used by the encoder;

[0144] • Creating and / or transmitting and / or receiving and / or decoding bitstreams or signals comprising one or more of the syntax elements or variations thereof;

[0145] • Creation and / or transmission and / or reception and / or decoding according to any one of the embodiments described;

[0146] • The method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any one of the embodiments described above;

[0147] • An adaptive television, set-top box, mobile phone, tablet computer or other electronic device that performs an encoding or decoding process according to any of the described implementation schemes;

[0148] • A television, set-top box, mobile phone, tablet computer or other electronic device that performs an adaptive encoding or decoding process according to any of the described implementation schemes and displays the resulting image (e.g., using a monitor, screen or other type of display);

[0149] • An adaptive television, set-top box, cellular phone, tablet computer, or other electronic device that selects (e.g., using a tuner) a channel to receive a signal including an encoded image and performs a decoding process according to any of the described embodiments;

[0150] • An adaptive television, set-top box, cellular phone, tablet or other electronic device that receives signals including encoded images over the air (e.g., using an antenna) and performs a decoding process according to any of the embodiments described.

Claims

1. A method for decoding, the method comprising: Obtain video data including general constraint information; A first general constraint information level syntax element is read from the general constraint information section, wherein a first value of the first general constraint information syntax element applies a third value to at least one sequence parameter set level syntax element, the third value indicating that the multi-type tree tool combining binary trees and ternary trees is not allowed; and a second value of the first general constraint information level syntax element indicates that no constraint is applied to the at least one sequence parameter set level syntax element. as well as The video data is decoded using the multi-type tree tool based on the values ​​of the at least one sequence parameter set level syntax element.

2. A method for encoding, the method comprising: Obtain the video sequence to be encoded in the video data and a set of encoding constraints; The signaling notification includes a general constraint information portion of a first general constraint information level syntax element, wherein a first value of the first general constraint information syntax element imposes a third value on at least one sequence parameter set level syntax element, the third value indicating that the multi-type tree tool combining binary trees and ternary trees is not allowed; and a second value of the first general constraint information level syntax element indicates that no constraint is imposed on the at least one sequence parameter set level syntax element. as well as The video data is encoded using the multi-type tree tool based on the values ​​of the at least one sequence parameter set level syntax element.

3. A device for decoding, the device comprising electronic circuitry, the electronic circuitry being configured... Set for: Obtain video data including general constraint information; Read a first general constraint information level syntax element from the general constraint information section, wherein a first value of the first general constraint information syntax element applies a third value to at least one sequence parameter set level syntax element, the third value indicating that multi-type tree tools combining binary and ternary trees are not allowed; and a second value of the first general constraint information level syntax element indicates that no constraint is applied to the at least one sequence parameter set level syntax element; and The video data is decoded using the multi-type tree tool based on the values ​​of the at least one sequence parameter set level syntax element.

4. A device for encoding, the device comprising electronic circuitry, the electronic circuitry being configured... Set for: Obtain the video sequence to be encoded in the video data and a set of encoding constraints; The signal notification includes a general constraint information portion of a first general constraint information level syntax element, wherein a first value of the first general constraint information syntax element imposes a third value on at least one sequence parameter set level syntax element, the third value indicating that the multi-type tree tool combining binary and ternary trees is not allowed; and a second value of the first general constraint information level syntax element indicates that no constraint is imposed on the at least one sequence parameter set level syntax element; and The video data is encoded using the multi-type tree tool based on the values ​​of the at least one sequence parameter set level syntax element.

5. A computer program product comprising, when executed by a processor, program code instructions that cause the processor to execute the method of claim 1.

6. A computer program product comprising, when executed by a processor, program code instructions that cause the processor to execute the method of claim 2.

7. A computer-readable medium having program code thereon that, when executed by one or more processors, causes the processors to perform the method of claim 1.

8. A computer-readable medium having program code thereon that, when executed by one or more processors, causes the processors to perform the method of claim 2.