Encoder, Decoder, and Corresponding Method
Patent Information
- Application Number
- KR1020227013372
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-24
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2040-09-24
Smart Images

Figure 112022042792990-PCT00032_ABST
Abstract
Description
Technology Field
[0001] This application claims priority from application No. PCT / CN2019 / 107594 filed on September 24, 2019. The disclosure of the aforementioned patent application is hereby incorporated by reference in its entirety.
[0002] The embodiments of the present application (disclosure) generally relate to the field of picture processing, and more specifically to inter-layer prediction. Background Technology
[0003] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcasting, digital TV, video transmission over the Internet and mobile networks, real-time interactive applications such as video chatting, video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems, and security applications for camcorders.
[0004] Even for relatively short videos, the amount of video data required to depict them can be substantial, which can cause difficulties when data needs to be streamed or transmitted over communication networks with limited bandwidth. Therefore, video data is typically compressed before being transmitted over modern communication networks. Since memory resources may be limited, the size of the video can also be an issue when it is stored on a storage device. Video compression devices often reduce the amount of data required to represent a digital video image by using software and / or hardware at the source to code the video data before transmission or storage. The compressed data is then received at the destination by a video decompression device that decodes the video data. With limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little to no sacrifice in picture quality are desirable.
[0005] Embodiments of the present application provide an apparatus and method for encoding and decoding according to an independent claim.
[0006] The aforementioned task and other tasks are achieved by the subject matter of the independent claim. Additional forms of implementation are evident from the dependent claim, the detailed description, and the drawings.
[0007] Certain embodiments are outlined in the appended independent claims, and other embodiments are outlined in the dependent claims.
[0008] According to a first aspect, the present invention relates to a method for decoding a coded video bitstream. The method is performed by a decoding device. The method comprises: a step of obtaining a sequence parameter set (SPS) level syntax element from a bitstream—specificating that an SPS level syntax element equal to a preset value indicates that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than the preset value indicates that the SPS references the VPS—; and a step of obtaining an inter-layer enabled syntax element that, when an SPS level syntax element is greater than the preset value, indicates whether one or more inter-layer reference pictures (ILRP) can be used for inter-prediction of one or more coded pictures. It includes a step of predicting one or more coded pictures based on the values of enable syntax elements between layers.
[0009] In inter-layer prediction, the coded picture and its reference picture belong to different layers, and different layers may correspond to different resolutions, with lower spatial resolution being used as a reference for higher spatial resolution. The inter-layer enable syntax element specifies whether inter-layer prediction is enabled; therefore, when the inter-layer enable syntax element specifies that inter-layer prediction is disabled, the syntax element associated with inter-layer prediction does not need to be signaled, which may result in a reduction in bitrate. Furthermore, the fact that the Video Parameter Set (VPS) is for multiple layers and no VPS is referenced by the SPS means that multiple layers are not required when decoding the picture associated with the SPS, i.e., only a single layer will be used when decoding the picture associated with the SPS. Inter-layer prediction cannot be performed when there is only a single layer; therefore, not signaling the inter-layer enable syntax element when no VPS is referenced by the SPS will further reduce the bitrate.
[0010] Here, a bitstream is a sequence of bits that forms one or more coded video sequences (CVS).
[0011] The coded video sequence (CVS) here is a sequence of AUs.
[0012] Here, the coded layer video sequence (CLVS) is a sequence of PUs with the same nuh_layer_id value.
[0013] Here, the access unit (AU) is a set of PUs belonging to different layers and containing coded pictures associated with the same time point for output from the DPB.
[0014] Here, a picture unit (PU) is a set of NAL units that are associated with each other according to a specified classification rule, have a consecutive decoding order, and contain exactly one coded picture.
[0015] Here, the Inter-Layer Reference Picture (ILRP) is a picture within the same AU as the current picture, and nuh_layer_id is smaller than the nuh_layer_id of the current picture.
[0016] Here, SPS is a syntax structure containing syntax elements that apply to zero or more total CLVS.
[0017] In a possible implementation of the method according to this first aspect, the VPS includes a syntax element describing inter-layer prediction information of a layer in a coded video sequence (CVS), and the SPS includes an SPS level syntax element and an inter-layer enable syntax element, wherein the CVS includes one or more ILRPs and one or more coded pictures.
[0018] Here, when the VPS is referenced by the SPS, the VPS includes a syntax element describing inter-layer prediction information of the layer to which one or more ILRPs and one or more coded pictures belong.
[0019] In a possible embodiment of the method according to such a first aspect or any preceding implementation of the first aspect, the step of predicting one or more coded pictures based on the value of an inter-layer enable syntax element comprises: predicting one or more coded pictures by referencing one or more ILRPs when it becomes possible to use the value of an inter-layer enable syntax element specifying one or more inter-layer reference pictures (ILRPs) for inter-predicting one or more coded pictures, wherein one or more ILRPs are obtained based on inter-layer prediction information contained in a VPS referenced by an SPS.
[0020] In a possible implementation of the method according to such a first aspect or any preceding implementation of the first aspect, the coded picture and the ILRP of the coded picture belong to different layers.
[0021] In a possible implementation of the method according to such a first aspect or any preceding implementation of the first aspect, an SPS level syntax element identical to a preset value further specifies that the coded video sequence (CVS) contains only one layer of coded pictures.
[0022] In a possible implementation form of the method according to such a first aspect or any preceding implementation of the first aspect, the preset value is 0.
[0023] In a possible embodiment of the method according to such a first aspect or any preceding implementation of the first aspect, the step of predicting one or more coded pictures based on the value of an inter-layer enable syntax element comprises: predicting one or more coded pictures without referencing any ILRP when the value of an inter-layer enable syntax element specifying one or more ILRPs is not used for inter-predicting of one or more coded pictures.
[0024] According to a second aspect, the present invention relates to a method for encoding a coded video bitstream. The method is performed by an encoding device. The method comprises: a step of encoding a sequence parameter set (SPS) level syntax element into a bitstream — specifying that an SPS level syntax element equal to a preset value is not referenced by any video parameter set (VPS) by the SPS, and specifying that an SPS level syntax element greater than the preset value is referenced by the SPS by the VPS —; and, when an SPS level syntax element is greater than the preset value, a step of encoding an inter-layer enable syntax element into a bitstream, wherein the inter-layer enable syntax element specifies whether one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures.
[0025] In a possible implementation of the method according to this second aspect, the VPS includes a syntax element describing inter-layer prediction information of layers in a coded video sequence (CVS), and the SPS includes an SPS level syntax element and an inter-layer enable syntax element, wherein the CVS includes one or more ILRPs and one or more coded pictures.
[0026] In a possible implementation of the method according to such a second aspect or any preceding implementation of the second aspect, the coded picture and the ILRP of the coded picture belong to different layers.
[0027] In a possible implementation of the method according to such a second aspect or any preceding implementation of the second aspect, an SPS level syntax element identical to a preset value further specifies that the coded video sequence (CVS) contains only one layer of coded pictures.
[0028] In a possible implementation form of the method according to such a second aspect or any preceding implementation of the second aspect, the preset value is 0.
[0029] In a possible embodiment of the method according to such a second aspect or any preceding implementation of the second aspect, the step of encoding an inter-layer enable syntax element into a bitstream comprises: a step of encoding an inter-layer enable syntax element into a bitstream that specifies that one or more ILRPs can be used for inter-predicting of one or more coded pictures based on a determination that one or more ILRPs can be used for inter-predicting of one or more coded pictures.
[0030] In a possible embodiment of the method according to such a second aspect or any preceding implementation of the second aspect, the step of encoding inter-layer enable syntax elements into a bitstream comprises: the step of encoding inter-layer enable syntax elements into a bitstream that specify, based on a determination that one or more ILRPs are not used for inter-predicting of one or more coded pictures, that one or more ILRPs are not used for inter-predicting of one or more coded pictures.
[0031] According to a third aspect, the present invention relates to a decoder for decoding a coded video bitstream, wherein the decoder comprises: an acquisition unit configured to acquire a sequence parameter set (SPS) level syntax element from the bitstream, wherein an SPS level syntax element equal to a preset value specifies that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than the preset value specifies that the SPS references the VPS; the acquisition unit is further configured to acquire an inter-layer enable syntax element that specifies whether, when the SPS level syntax element is greater than the preset value, one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures; and a prediction unit configured to predict one or more coded pictures based on the value of the inter-layer enable syntax element.
[0032] In a possible implementation of the method according to this third aspect, the VPS includes a syntax element describing inter-layer prediction information of layers in a coded video sequence (CVS), and the SPS includes an SPS level syntax element and an inter-layer enable syntax element, wherein the CVS includes one or more ILRPs and one or more coded pictures.
[0033] In a possible implementation of the method according to such a third aspect or any preceding implementation of the third aspect, the prediction unit is configured to predict one or more coded pictures by referencing one or more ILRPs when the value of an inter-layer enable syntax element specifying one or more inter-layer reference pictures (ILRPs) becomes available for inter-prediction of one or more coded pictures, and the one or more ILRPs are obtained based on inter-layer prediction information contained in a VPS referenced by an SPS.
[0034] In such a third aspect or a possible implementation of the method according to any preceding implementation of the third aspect, the coded picture and the ILRP of the coded picture belong to different layers.
[0035] In a possible implementation of the method according to such a third aspect or any preceding implementation of the third aspect, an SPS level syntax element identical to a preset value further specifies that the coded video sequence (CVS) contains only one layer of coded pictures.
[0036] In a possible implementation form of the method according to such a third aspect or any preceding implementation of the third aspect, the preset value is 0.
[0037] In a possible implementation of the method according to such a third aspect or any preceding implementation of the third aspect, the prediction unit is configured to predict one or more coded pictures without referencing any ILRP when the value of an inter-layer enable syntax element specifying one or more ILRPs is not used for inter-prediction of one or more coded pictures.
[0038] According to a fourth aspect, the present invention relates to an encoder for encoding a coded video bitstream, wherein the encoder comprises: a first encoding unit configured to encode a sequence parameter set (SPS) level syntax element into a bitstream—specificating that an SPS level syntax element equal to a preset value is not referenced by any video parameter set (VPS) by the SPS, and an SPS level syntax element greater than a preset value is referenced by the SPS by the VPS—; and a second encoding unit configured to encode an inter-layer enable syntax element into a bitstream when the SPS level syntax element is greater than a preset value, wherein the inter-layer enable syntax element specifies whether one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures.
[0039] In a possible implementation of the method according to this fourth aspect, the encoder further includes a determination unit configured to determine whether an SPS level syntax element is greater than a preset value.
[0040] In a possible embodiment of the method according to such a fourth aspect or any preceding implementation of the fourth aspect, the VPS includes a syntax element describing inter-layer prediction information of a layer in a coded video sequence (CVS), and the SPS includes an SPS level syntax element and an inter-layer enable syntax element, wherein the CVS includes one or more ILRPs and one or more coded pictures.
[0041] In such a fourth aspect or a possible implementation of the method according to any preceding implementation of the fourth aspect, the coded picture and the ILRP of the coded picture belong to different layers.
[0042] In a possible implementation of the method according to such a fourth aspect or any preceding implementation of the fourth aspect, an SPS level syntax element identical to a preset value further specifies that the coded video sequence (CVS) contains only one layer of coded pictures.
[0043] In a possible implementation form of the method according to such a fourth aspect or any preceding implementation of the fourth aspect, the preset value is 0.
[0044] In a possible implementation of the method according to such a fourth aspect or any preceding implementation of the fourth aspect, the second encoding unit is configured to encode a layer-to-layer enable syntax element into a bitstream that specifies that one or more ILRPs can be used for inter-predicting of one or more coded pictures, based on a determination that one or more ILRPs can be used for inter-predicting of one or more coded pictures.
[0045] In a possible implementation of the method according to such a fourth aspect or any preceding implementation of the fourth aspect, the second encoding unit is configured to encode a layer-to-layer enable syntax element into a bitstream that specifies that one or more ILRPs are not used for inter-predicting of one or more coded pictures, based on a determination that one or more ILRPs are not used for inter-predicting of one or more coded pictures.
[0046] In a possible embodiment of the method according to such a fourth aspect or any preceding implementation of the fourth aspect, the encoder further comprises a determination unit configured to determine whether one or more ILRPs can be used for inter-predicting of one or more coded pictures.
[0047] The method according to the first aspect of the present invention may be performed by an apparatus according to the third aspect of the present invention. Additional features and embodiments of the method according to the third aspect of the present invention correspond to features and embodiments of the apparatus according to the first aspect of the present invention.
[0048] The method according to the second aspect of the present invention may be performed by the apparatus according to the fourth aspect of the present invention. Additional features and embodiments of the method according to the fourth aspect of the present invention correspond to features and embodiments of the apparatus according to the second aspect of the present invention.
[0049] The method according to the second aspect can be extended to an implementation corresponding to the implementation of the first device according to the first aspect. Therefore, the implementation of this method includes the feature(s) of the implementation corresponding to the first device.
[0050] The advantage of the method according to the second aspect is the same as the advantage of the corresponding implementation form of the first device according to the first aspect.
[0051] According to a fifth aspect, the present invention relates to an apparatus for decoding a video stream, the apparatus comprising a processor and a memory. The memory stores instructions that cause the processor to perform a method according to a first aspect.
[0052] According to the sixth aspect, the present invention relates to an apparatus for encoding a video stream, the apparatus comprising a processor and a memory. The memory stores instructions that cause the processor to perform a method according to the second aspect.
[0053] According to the seventh aspect, a computer-readable storage medium is proposed that stores instructions that, when executed, cause one or more processors to code video data. The instructions cause one or more processors to perform a method according to the first or second aspect, or any possible embodiment of the first or second aspect.
[0054] According to the eighth aspect, the present invention relates to a computer program comprising program code for performing a method according to the first or second aspect, or any possible embodiment of the first or second aspect, when executed on a computer.
[0055] According to the ninth aspect, the present invention relates to a non-transient storage medium comprising an encoded bitstream decoded by an image decoding device, wherein the bitstream is generated by dividing a frame of a video signal or image signal into a plurality of blocks and comprises a plurality of syntax elements, wherein the plurality of syntax elements comprises an inter-layer enable syntax element that specifies whether one or more inter-layer reference pictures (ILRPs) can be used for inter-predicting of one or more coded pictures, subject to the condition that an SPS level syntax element is greater than a preset value, wherein an SPS level syntax element equal to the preset value specifies that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than the preset value specifies that the SPS references the VPS.
[0056] Details of one or more embodiments are presented in the attached drawings and detailed description below. Other features, problems, and advantages will be apparent from the description, drawings, and claims.
[0057] Furthermore, the following embodiments are provided.
[0058] In one embodiment, a method for decoding a coded video bitstream is provided, the method being:
[0059] A step of parsing a first syntax element that specifies whether the layer with index i uses cross-layer prediction — i is an integer, and i is greater than 0 —;
[0060] A step of parsing a second syntax element that specifies whether the layer with index j is a direct reference layer to the layer with index i when the first condition is satisfied — where j is an integer, and j is less than i and greater than or equal to 0, and the first condition includes the first syntax element specifying that the layer with index i can use inter-layer prediction and i is greater than a preset value (e.g., 1) —;
[0061] It includes the step of predicting a picture of a hierarchy having index i based on the value of a second syntax element.
[0062] In one embodiment, this method is:
[0063] When condition 2 is satisfied, the method further includes the step of predicting a picture of a layer having index i using a layer having index j as a direct reference layer for a layer having index i, wherein j is an integer and j is less than i and greater than or equal to 0, and the second condition includes a syntax element specifying that the layer having index i can use inter-layer prediction and i is equal to a preset value.
[0064] In one embodiment, this method is:
[0065] 2. When condition 2 is satisfied, the method further includes the step of determining that the value of the second syntax element specifies that the layer with index j is a direct reference layer to the layer with index i.
[0066] In one embodiment, a picture of a layer having index i includes a picture within the layer having index i or a picture associated with the layer having index i.
[0067] In one embodiment, a method for decoding a coded video bitstream is provided, the method being:
[0068] A step of parsing a syntax element that specifies whether the layer with dex i uses cross-layer prediction — i is an integer, and i is greater than 0 —;
[0069] When the condition is satisfied, the method includes the step of predicting a picture of the layer having index i using the layer having index j as a direct reference layer for the layer having index i, wherein j is an integer and is equal to i-1, and the condition includes a syntax element specifying that the layer having index i can use inter-layer prediction.
[0070] In one embodiment, a picture of a layer having index i includes a picture within the layer having index i or a picture associated with the layer having index i.
[0071] In one embodiment, a method for decoding a coded video bitstream is provided, the method being:
[0072] A step of parsing a syntax element that specifies whether at least one long-term reference picture (LTRP) is used for inter-prediction of any coded picture in a coded video sequence (CVS) — each picture of at least one LTRP is marked as "used for long-term reference," but inter-level reference pictures (ILRP) are not marked —;
[0073] It includes a step of predicting one or more coded pictures in CVS based on the value of a tax element.
[0074] In one embodiment, a method for decoding a coded video bitstream is provided, the method being:
[0075] A step to determine whether the condition is satisfied — the condition includes the layer index of the current layer being greater than a preset value —;
[0076] A step of parsing a first syntax element that specifies whether at least one inter-layer reference picture (ILRP) is used for inter-prediction of any coded picture in a coded video sequence (CVS) when a condition is satisfied;
[0077] 1. Includes the step of predicting one or more coded pictures in CVS based on the value of a syntax element.
[0078] In one embodiment, the preset value is 0.
[0079] In one embodiment, the condition further includes that the second syntax element (e.g., sps_video_parameter_set_id) is greater than 0.
[0080] In one embodiment, a method for decoding a coded video bitstream is provided, the method being:
[0081] A step for determining whether a condition is satisfied — the condition includes the fact that the hierarchy index of the current hierarchy is greater than a preset value and the current entry within the referenced picture list structure is an ILRP entry —;
[0082] When the condition is satisfied, a step of parsing a syntax element that specifies an index for a list of direct dependent layers of the current layer;
[0083] It includes the step of predicting one or more coded pictures in CVS based on the reference picture list structure of the current entry in which the ILRP is obtained, using an index for the list of adjacent dependent layers.
[0084] In one embodiment, the preset value is 1.
[0085] In one embodiment, an encoder (20) is provided that includes a processing circuit for executing a method according to any of the prior embodiments.
[0086] In one embodiment, a decoder (30) is provided that includes a processing circuit for executing a method according to any of the prior embodiments.
[0087] In one embodiment, a computer program product is provided that includes program code for performing a method according to any of the prior embodiments when executed on a computer or processor.
[0088] In one embodiment, a decoder is provided, and the decoder is:
[0089] A processor of my caliber; and
[0090] It includes a non-transient computer-readable storage medium coupled to a processor and storing a program for execution by the processor, and the program is configured to execute a method according to any of the prior embodiments when executed by the processor.
[0091] In one embodiment, an encoder is provided, and the encoder is:
[0092] A processor of my caliber; and
[0093] It includes a non-transient computer-readable storage medium coupled to a processor and storing a program for execution by the processor, and the encoder is configured to execute a method according to any of the prior embodiments when the program is executed by the processor.
[0094] In one embodiment, a non-transient computer-readable medium is provided, which transmits program code that, when executed by a computer device, causes the computer device to perform a method of any of the prior embodiments. Brief explanation of the drawing
[0095] Next, embodiments of the present invention are described in more detail with reference to the attached diagrams and drawings. FIG. 1a is a block diagram showing an example of a video coding system configured to implement an embodiment of the present invention. FIG. 1b is a block diagram showing another example of a video coding system configured to implement an embodiment of the present invention. FIG. 2 is a block diagram showing an example of a video encoder configured to implement an embodiment of the present invention. FIG. 3 is a block diagram showing an exemplary structure of a video decoder configured to implement an embodiment of the present invention. FIG. 4 is a block diagram illustrating an example of an encoding device or a decoding device. Figure 5 is a block diagram illustrating another example of an encoding device or a decoding device. Figure 6 is a block diagram illustrating scalable coding by two layers. FIG. 7 is a block diagram illustrating an exemplary structure of a content supply system (3100) that realizes a content delivery service. FIG. 8 is a block diagram illustrating the structure of an example of a terminal device. FIG. 9 is a flowchart of a decoding method according to one embodiment. FIG. 10 is a flowchart of an encoding method according to one embodiment. FIG. 11 is a schematic diagram of an encoder according to one embodiment. FIG. 12 is a schematic diagram of a decoder according to one embodiment. Next, unless explicitly stated otherwise, the same reference numeral will refer to the same or at least functionally equivalent features. Specific details for implementing the invention
[0096] In the following description, reference is made to the accompanying drawings, which illustrate specific aspects of embodiments of the invention or specific aspects in which embodiments of the invention may be used, and which form part of the content of the present disclosure. It is understood that embodiments of the invention may be used in other aspects and may include structural or logical modifications not illustrated in the drawings. Accordingly, the following detailed description is not to be construed as limiting, and the scope of the invention is defined by the appended claims.
[0097] For example, it is understood that the disclosure relating to the method described may also hold true for the corresponding device or system configured to perform the method, and vice versa. On the other hand, for example, if a particular device is described based on one or more units, such as functional units, the corresponding method may include one or more units, such as functional units, to perform the steps of the one or more method described (e.g., one unit performing one or more steps, or multiple units each performing one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a particular device is described based on one or more units, such as functional units, the corresponding method may include one step to perform the function of one or more units (e.g., one step performing the function of one or more units, or multiple steps each performing the function of one or more of the multiple units), even if such one or more steps are not explicitly described or illustrated in the drawings. Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with one another unless specifically stated otherwise.
[0098] Video coding typically refers to the processing of a sequence of pictures that form a video or a video sequence. In the field of video coding, the terms "frame" or "image" may be used as synonyms instead of "picture." Video coding (or coding in general) comprises two parts: video encoding and video decoding. Video encoding is performed at the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed at the destination side and typically involves reverse processing by comparing with an encoder to reconstruct the video picture. An embodiment referring to the "coding" of a video picture (or picture in general) will be understood to relate to the "encoding" or "decoding" of the video picture or each video sequence. The combination of the encoding and decoding parts is also referred to as a CODEC (Coding and Decoding).
[0099] In the case of lossless video coding, the original video picture can be reconstructed; that is, the reconstructed video picture has the same quality as the original video picture (assuming there is no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, additional compression is performed, for example through quantization, to reduce the amount of data representing video pictures that cannot be perfectly reconstructed by the decoder; that is, the quality of the reconstructed video picture is lower or worse than that of the original video picture.
[0100] Several video coding standards belong to the group of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the sample domain with 2D transformation coding to apply quantization in the transformation domain). Each picture in a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. That is, in the encoder, the video is typically processed at the block (video block) level—i.e., encoded—by generating prediction blocks using, for example, spatial (intra-picture) prediction and / or temporal (inter-picture) prediction, obtaining residual blocks by subtracting the prediction blocks from the current blocks (blocks currently processed / to be processed), transforming the residual blocks, and quantizing the residual blocks in the transformation domain to reduce (compress) the amount of data to be transmitted, whereas in the decoder, inverse processing is applied to the encoded or compressed blocks compared with the encoder to reconstruct the current blocks for representation. Furthermore, the encoder duplicates the decoder processing loop so that both the encoder and decoder will generate the same predictions (e.g., intra-predictions and inter-predictions) and / or reconstructions to process subsequent blocks, i.e., to code.
[0101] Next, embodiments of a video coding system (10), a video encoder (20), and a video decoder (30) are described based on FIGS. 1 to 3.
[0102] FIG. 1a is a schematic block diagram illustrating an exemplary coding system (10) capable of utilizing the technology of the present application, such as a video coding system (10) (or abbreviated as coding system (10)). The video encoder (20) (or abbreviated as encoder (20)) and video decoder (30) (or abbreviated as decoder (30)) of the video coding system (10) represent examples of devices that may be configured to perform the technology according to various examples described in the present application.
[0103] As illustrated in FIG. 1a, the coding system (10) includes a source device (12) configured to provide encoded picture data (21) to a destination device (14) for decoding encoded picture data (21), for example.
[0104] The source device (12) includes an encoder (20) and may additionally, i.e. optionally, include a picture source (16), a preprocessor (or postprocessing unit) (18), such as a picture preprocessor (18), and a communication interface or communication unit (22).
[0105] The picture source (16) may include, or be any kind of picture capture device, e.g., a camera for capturing real-world pictures, and / or any kind of picture generation device, e.g., a computer graphics processor for generating computer-animated pictures, or any other kind of device for acquiring and / or providing real-world pictures, computer-animated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may be any kind of memory or storage for storing any of the aforementioned pictures.
[0106] In distinction from the processing performed by the preprocessor (18) and the preprocessing unit (18), the picture or picture data (17) may also be referred to as the raw picture or raw picture data (17).
[0107] The preprocessor (18) is configured to receive (raw) picture data (17) and to perform preprocessing on the picture data (17) to obtain a preprocessed picture (19) or preprocessed picture data (19). The preprocessing performed by the preprocessor (18) may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It may be understood that the preprocessing unit (18) may be an optional component.
[0108] A video encoder (20) is configured to receive preprocessed picture data (19) and provide encoded picture data (21) (additional details will be described below, for example, based on FIG. 2).
[0109] The communication interface (22) of the source device (12) may be configured to receive encoded picture data (21) and to transmit the encoded picture data (21) (or any additionally processed version thereof) to another device, such as a destination device (14) or any other device, for storage or direct reconstruction via the communication channel (13).
[0110] The destination device (14) includes a decoder (30) (e.g., a video decoder (30)) and may additionally, i.e. optionally, include a communication interface or communication unit (28), a post-processor (32) (or post-processing unit (32)) and a display device (34).
[0111] The communication interface (28) of the destination device (14) is configured to receive encoded picture data (21) (or any additionally processed version thereof) from, for example, a direct source device (12) or from any other source, for example, a storage device, for example, an encoded picture data storage device, and to provide the encoded picture data (21) to the decoder (30).
[0112] The communication interface (22) and the communication interface (28) may be configured to transmit or receive encoded picture data (21) or encoded data (13) through a direct communication link between the source device (12) and the destination device (14), such as a direct wired or wireless connection, or through any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any combination thereof.
[0113] The communication interface (22) may be configured to package, for example, the encoded picture data (21) into a suitable format, such as a packet, and / or process the encoded picture data using any kind of transmission encoding or processing for transmission through a communication link or communication network.
[0114] A communication interface (28) forming a corresponding part of a communication interface (22) may be configured, for example, to receive transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or de-packaging to obtain encoded picture data (21).
[0115] Both the communication interface (22) and the communication interface (28) may be configured as a unidirectional communication interface or a bidirectional communication interface, as indicated by the arrow for the communication channel (13) of FIG. 1a pointing from the source device (12) to the destination device (14), and may be configured to send and receive messages, for example, to establish a connection, and to exchange any other information related to the communication link and / or data transmission, for example, the transmission of encoded picture data.
[0116] The decoder (30) is configured to receive encoded picture data (21) and provide decoded picture data (31) or decoded picture (31) (additional details will be described below, for example, based on FIG. 3 or FIG. 5).
[0117] The post-processor (32) of the destination device (14) is configured to post-process decoded picture data (31), e.g., decoded picture (31), (also referred to as reconstructed picture data) to obtain post-processed picture data (33), e.g., post-processed picture data (33). The post-processing performed by the post-processing unit (32) may include, e.g., color format conversion (e.g., from YCbCr to RGB), color correction, trimming or resampling, or any other processing, e.g., preparing the decoded picture data (31) for display by, e.g., a display device (34).
[0118] The display device (34) of the destination device (14) is configured to receive post-processed picture data (33) to display a picture to, for example, a user or viewer. The display device (34) may be any type of display for displaying the reconstructed picture, such as an integrated or external display or monitor, or may include such. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.
[0119] Although FIG. 1a illustrates the source device (12) and the destination device (14) as separate devices, embodiments of the device may also include both the source device (12) and the destination device (14), or both devices or both functions of the source device (12) or the corresponding function and the destination device (14) or the corresponding function. In these embodiments, the source device (12) or the corresponding function and the destination device (14) or the corresponding function may be implemented using the same hardware and / or software, and / or by separate hardware and / or software, or by any combination thereof.
[0120] As is obvious to those skilled in the art based on the description, the existence of different units or functions and the exact division of functions within the source device (12) and / or destination device (14) shown in FIG. 1a may vary depending on the actual device and application.
[0121] An encoder (20) (e.g., a video encoder (20)) or a decoder (30) (e.g., a video decoder (30)) or both the encoder (20) and the decoder (30) may be implemented through the processing circuit shown in FIG. 1b, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated to video coding, or any combination thereof. The encoder (20) may be implemented through the processing circuit (46) to implement various modules as discussed for the encoder (20) of FIG. 2 and / or any other encoder system or subsystem described herein. The decoder (30) may be implemented via a processing circuit (46) to implement various modules as discussed for the decoder (30) of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations as discussed later. As illustrated in FIG. 5, if the technology is partially implemented in software, the device may store instructions for the software in a suitable non-transient computer-readable storage medium and execute instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder (20) and the video decoder (30) may be integrated as part of a combined encoder / decoder (CODEC) in a single device, for example, as illustrated in FIG. 1b.
[0122] The source device (12) and the destination device (14) may include any of the broad range of devices, such as any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content delivery server), a broadcast receiver device, a broadcast transmitter device, etc., and may not use an operating system or may use any type of operating system. In some cases, the source device (12) and the destination device (14) may be equipped for wireless communication. Thus, the source device (12) and the destination device (14) may be wireless communication devices.
[0123] In some cases, the video coding system exemplified in FIG. 1a ( 10 ) is merely an example, and the technology of the present application may be applied to video coding setups (e.g., video encoding or video decoding) that do not necessarily involve any data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory or streamed over a network. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve data from memory to decode it. In some examples, encoding and decoding are performed by a device that does not communicate with each other but simply encodes data into memory and / or retrieves data from memory to decode it.
[0124] For convenience of explanation, embodiments of the present invention are described herein with reference to, for example, reference software for High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC), and next-generation video coding standards developed by the ITU-T Video Coding Experts Group (VCEG) and the Joint Collaboration Team on Video Coding (JCT-VC) of the ISO / IEC Motion Picture Experts Group (MPEG). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.
[0125] Encoder and Encoding Method
[0126] FIG. 2 illustrates a schematic block diagram of an exemplary video encoder (20) configured to implement the technology of the present application. In the example of FIG. 2, the video encoder (20) includes an input (201) (or input interface (201)), a residual calculation unit (204), a transform processing unit (206), a quantization unit (208), an inverse quantization unit (210) and an inverse transform processing unit (212), a reconstruction unit (214), a loop filter unit (220), a decoded picture buffer (DPB) (230), a mode selection unit (260), an entropy encoding unit (270), and an output (272) (or output interface (272)). The mode selection unit (260) may include an inter-prediction unit (244), an intra-prediction unit (254), and a partitioning unit (262). The inter prediction unit (244) may include a motion estimation unit (not shown) and a motion compensation unit. The video encoder (20) shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder according to a hybrid video codec.
[0127] While the residual calculation unit (204), the transformation processing unit (206), the quantization unit (208), and the mode selection unit (260) may be described as forming the forward signal path of the encoder (20), the inverse quantization unit (210), the inverse transformation processing unit (212), the reconstruction unit (214), the buffer (216), the loop filter (220), the decoded picture buffer (DPB) (230), the inter prediction unit (244), and the intra prediction unit (254) may be described as forming the reverse signal path of the video encoder (20), and the reverse signal path of the video encoder (20) corresponds to the signal path of the decoder (see video decoder (30) in FIG. 3). The inverse quantization unit (210), inverse transformation processing unit (212), reconstruction unit (214), loop filter (220), decoded picture buffer (DPB) (230), inter prediction unit (244) and intra prediction unit (254) are also referred to as forming the "built-in decoder" of the video encoder (20).
[0128] Picture and picture partitioning (picture and block)
[0129] The encoder (20) may be configured to receive, for example, a picture (17) (or picture data (17)), such as a picture of a sequence of pictures forming a video or video sequence, through an input (201). The received picture or picture data may also be a preprocessed picture (19) (or preprocessed picture data (19)). For brevity, the following description refers to the picture (17). The picture (17) may also be referred to as the current picture or (particularly, in video coding to distinguish the current picture from another picture, such as the same video sequence, i.e., a previously encoded and / or decoded picture of a video sequence that also includes the current picture) the picture to be coded.
[0130] A (digital) picture is a two-dimensional array or matrix of samples having intensity values, or can be regarded as such a two-dimensional array or matrix. Samples within the array may also be referred to as pixels (picture elements in shorthand form) or pixels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are used; that is, the picture may be represented by an array of three samples or may contain an array of three samples. In an RGB format or color space, the picture contains corresponding arrays of red, green, and blue samples. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, such as YCbCr, which includes a luminance component designated as Y (sometimes L is also used instead) and two chrominance components designated as Cb and Cr. The luminance (or abbreviated as luma) component (Y) represents brightness or gray level intensity (e.g., in a grayscale picture), while the two chrominance (or abbreviated as chroma) components (Cb, Cr) represent color or color information components. Accordingly, a picture in the YCbCr format contains an array of luminance samples of luminance sample values (Y) and two arrays of chrominance samples of chrominance values (Cb, Cr). A picture in the RGB format can be converted or switched to the YCbCr format, and vice versa, and the process is also known as color conversion or switching. If the picture is monochromatic, the picture may contain only an array of luminance samples. Accordingly, the picture may be, for example, an array of luminance samples in a monochromatic format, or an array of luminance samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0131] An embodiment of the video encoder (20) may include a picture partitioning unit (not shown in FIG. 2) configured to partition a picture (17) into a plurality of (typically non-overlapping) picture blocks (203). These blocks may also be referred to as root blocks, macro blocks (H.264 / AVC) or coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size and a corresponding grid defining that block size for all pictures in the video sequence, or to change the block size between pictures or a subset or group of pictures and partition each picture into a corresponding block.
[0132] In an additional embodiment, the video encoder may be configured to directly receive blocks (203) of a picture (17), such as one, several, or all blocks forming the picture (17). The picture blocks (203) may also be referred to as the current picture blocks or the picture blocks to be coded.
[0133] Like the picture (17), the picture block (203) is smaller in dimension than the picture (17) but is a two-dimensional array or matrix of samples having intensity values (sample values), or can be considered as such a two-dimensional array or matrix. That is, the block (203) may include, for example, one array of samples (e.g., a luminance array for a monochrome picture (17), or a luminance or chroma array for a color picture) or three arrays of samples (e.g., a luminance array and two chroma arrays for a color picture (17)) or any other number and / or type of array, depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of the block (203) defines the size of the block (203). Accordingly, the block may be, for example, an M×N (M-column x N-row) array of samples or an M×N array of transformation factors.
[0134] An embodiment of the video encoder (20) as illustrated in FIG. 2 can be configured to encode a picture (17) in blocks, for example, encoding and prediction are performed for each block (203).
[0135] An embodiment of the video encoder (20) as illustrated in FIG. 2 may be further configured to partition and / or encode a picture by using slices (also referred to as video slices), wherein the picture may be partitioned into one or more slices (typically not overlapping) or encoded using one or more slices, and each slice may include one or more blocks (e.g., CTU) or one or more groups of blocks (e.g., tiles (H.265 / HEVC and VVC) or bricks (VVC)).
[0136] An embodiment of the video encoder (20) as illustrated in FIG. 2 may be further configured to partition and / or encode a picture by using a slice / tile group (also referred to as a video tile group) and / or a tile (also referred to as a video tile), wherein the picture may be partitioned into one or more slice / tile groups (typically not overlapping) or encoded using one or more slice / tile groups, and each slice / tile group may include, for example, one or more blocks (e.g., CTU) or one or more tiles, and each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTU), for example, complete or partial blocks.
[0137] Residual calculation
[0138] The residual calculation unit (204) is configured to obtain the residual block (205) in the sample domain by calculating the residual block (205) (also referred to as residual (205)) based on the picture block (203) and the prediction block (265) (additional details regarding the prediction block (265) will be provided later), for example, by subtracting the sample value of the prediction block (265) from the sample value of the picture block (203) in sample units (pixel units).
[0139] conversion
[0140] The transformation processing unit (206) is configured to obtain transformation coefficients (207) in the transformation domain by applying a transformation, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values of the residual block (205). The transformation coefficients (207) may also be referred to as transformation residual coefficients and may represent the residual block (205) in the transformation domain.
[0141] The conversion processing unit (206) may be configured to apply integer approximations of DCT / DST, such as the conversion specified for H.265 / HEVC. Compared to the orthogonal DCT conversion, these integer approximations are typically scaled by specific factors. To preserve the norm of the residual blocks processed by the forward and inverse conversions, additional scaling factors are applied as part of the conversion process. The scaling factors are typically selected based on specific constraints, such as a scaling factor that is a power of 2 for the shift operation, the bit depth of the conversion factor, and a trade-off between accuracy and implementation cost. For example, a specific scaling factor may be specified for the inverse conversion by, for example, the inverse conversion processing unit (212) (and the corresponding inverse conversion by, for example, the inverse conversion processing unit (312) of the video decoder (30), and accordingly, a corresponding scaling factor may be specified for the forward conversion by, for example, the conversion processing unit (206) of the encoder (20).
[0142] An embodiment of the video encoder (20) (each conversion processing unit (206)) may be configured to output conversion parameters, such as a type of conversion or conversions encoded or compressed through the direct or entropy encoding unit (270), so that, for example, the video decoder (30) may receive and use conversion parameters for decoding.
[0143] Quantization
[0144] The quantization unit (208) may be configured to obtain quantized coefficients (209) by quantizing the transformation coefficients (207) by, for example, applying scalar quantization or vector quantization. The quantized coefficients (209) may also be referred to as quantized transformation coefficients (209) or quantized residual coefficients (209).
[0145] The quantization process can reduce the bit depth associated with some or all of the conversion factors (207). For example, during quantization, an n-bit conversion factor can be rounded down to an m-bit conversion factor, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, different scaling can be applied to achieve finer or more approximate quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to more approximate quantization. The applicable quantization step size can be indicated by the quantization parameter (QP). The quantization parameter can be, for example, an index for a predefined set of applicable quantization step sizes. For example, a small quantization parameter can correspond to fine quantization (small quantization step size) and a large quantization parameter can correspond to approximate quantization (large quantization step size), or vice versa. Quantization may involve division by the quantization step size, and, for example, the corresponding and / or inverse quantization by the inverse quantization unit (210) may involve multiplication by the quantization step size. An embodiment according to some standards, such as HEVC, may be configured to determine the quantization step size using quantization parameters. Generally, the quantization step size may be calculated based on the quantization parameters using a fixed-point approximation of an expression involving division. Additional scaling factors may be introduced in the quantization and inverse quantization to restore the norm of the residual block, and the norm of the residual block may be modified due to the scaling used in the fixed-point approximation of the expression for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and inverse quantization may be combined.Alternatively, a custom quantization table is used and can be signaled from the encoder to the decoder, for example, as a bitstream. Quantization is a lossy operation, and here, the loss increases with increasing quantization step size.
[0146] An embodiment of the video encoder (20) (each quantization unit (208)) can be configured to output quantization parameters (QP) encoded, for example, directly or through an entropy encoding unit (270), so that, for example, the video decoder (30) can receive and apply the quantization parameters for decoding.
[0147] Inverse quantization
[0148] The inverse quantization unit (210) is configured to obtain an inverse quantized coefficient (211) by applying the inverse of the quantization method applied by the quantization unit (208) to the quantized coefficient, for example, based on or using the same quantization step size as the quantization unit (208). The inverse quantized coefficient (211) may also be referred to as an inverse quantized residual coefficient (211) and corresponds to the transformation coefficient (207) — which is not typically the same as the transformation coefficient due to losses caused by quantization.
[0149] Inverse transformation
[0150] The inverse transform processing unit (212) is configured to obtain a reconstructed residual block (213) (or the corresponding inverse quantized coefficient (213)) in the sample domain by applying an inverse transform of the transform applied by the transform processing unit (206), such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or another inverse transform. The reconstructed residual block (213) may also be referred to as a transform block (213).
[0151] Reconstruction
[0152] The reconstruction unit (214) (e.g., an adder or summer (214)) is configured to obtain a reconstruction block (215) in the sample domain by adding the transformation block (213) (i.e., the reconstruction residual block (213)) to the prediction block (265) by adding, for example, the sample values of the reconstruction residual block (213) and the prediction block (265) — on a sample basis.
[0153] Filtering
[0154] A loop filter unit (220) (or short "loop filter" (220)) is configured to filter a reconstructed block (215) to obtain a filtered block (221), or generally, to filter a reconstructed sample to obtain a filtered sample value. The loop filter unit is configured, for example, to smooth pixel transitions or to improve video quality in other ways. The loop filter unit (220) may include one or more loop filters, such as a de-blocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, the loop filter unit (220) may include a de-blocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be the de-blocking filter, the SAO, and the ALF. In another example, a process called luma mapping with chroma scaling (LMCS) (i.e., adaptive in-loop reshaper) is added. This process is performed before block separation. In another example, the block separation filter process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. In FIG. 2, the loop filter unit (220) is shown as an in-loop filter, but in other configurations, the loop filter unit (220) may be implemented as a post-loop filter. The filtered block (221) may also be referred to as a filtered reconstructed block (221).
[0155] An embodiment of the video encoder (20) (each loop filter unit (220)) can be configured to output loop filter parameters (e.g., SAO filter parameters or ALF filter parameters or LMCS parameters) encoded, for example, directly or through the entropy encoding unit (270), so that, for example, the decoder (30) can receive and apply the same loop filter parameters or each loop filter for decoding.
[0156] Decoded picture buffer
[0157] The decoded picture buffer (DPB) (230) may be a reference picture for encoding video data by the video encoder (20), or generally a memory that stores reference picture data. The DPB (230) may be formed by any memory device among various memory devices, such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) (230) may be configured to store one or more filtered blocks (221). The decoded picture buffer (230) may be further configured to store other previously filtered blocks of the same current picture or different pictures, such as previously reconstructed pictures, such as previously reconstructed and filtered blocks (221), and may provide, for example, for inter-predicting, the complete previously reconstructed, i.e., decoded picture (and corresponding reference blocks and samples) and / or partially reconstructed current picture (and corresponding reference blocks and samples). The decoded picture buffer (DPB) (230) may also be configured to store, for example, one or more unfiltered reconstructed blocks (215) or, generally, unfiltered reconstructed samples, or any other additionally processed version of a reconstructed block or sample, if the reconstructed block (215) is not filtered by the loop filter unit (220).
[0158] Mode Selection (Partitioning and Prediction)
[0159] The mode selection unit (260) comprises a partitioning unit (262), an inter prediction unit (244), and an intra prediction unit (254), and is configured to receive or acquire original picture data, e.g., original block (203) (current block (203) of the current picture (17)), and reconstructed picture data, e.g., from one or more previously decoded pictures of the same (current) picture, e.g., a decoded picture buffer (230) or other buffer (e.g., a line buffer not shown). The reconstructed picture data is used as reference picture data for prediction, e.g., inter prediction or intra prediction, to acquire a prediction block (265) or a predictor (265).
[0160] The mode selection unit (260) may be configured to determine or select a partitioning for a current block prediction mode (which does not include partitioning) and a prediction mode (e.g., intra or inter prediction mode), and to generate a corresponding prediction block (265) used for calculating the residual block (205) and for reconstructing the reconstructed block (215).
[0161] An embodiment of the mode selection unit (260) may be configured to select a partitioning and prediction mode that provides the best match, or that is, the minimum residual (the minimum residual means better compression for transmission or storage), or the minimum signaling overhead (the minimum signaling overhead means better compression for transmission or storage), or both, or balances both (e.g., from those supported by the mode selection unit (260) or available to the mode selection unit (260). The mode selection unit (260) may be configured to determine the partitioning and prediction mode based on rate distortion optimization (RDO), that is, to select a prediction mode that provides the minimum rate distortion. In this context, terms such as "best," "minimum," and "optimal" do not necessarily mean the overall "best," "minimum," or "optimal," but may also refer to the satisfaction of termination or selection criteria, such as values exceeding or falling below a threshold, or other constraints that potentially lead to a "suboptimal choice" while reducing complexity and processing time.
[0162] In other words, the partitioning unit (262) may be configured to partition a picture from a video sequence into a sequence of coding tree units (CTUs), and further partition the CTU (203) into smaller block partitions or subblocks (which again form blocks) using, for example, quad tree partitioning (QT), binary partitioning (BT), or ternary tree partitioning (TT) or any combination thereof, and to perform prediction for each block partition or subblock, for example, the mode selection includes the selection of a tree structure of the partitioned block (203), and the prediction mode is applied to each block partition or subblock.
[0163] Next, partitioning (e.g., by the partitioning unit (260)) and prediction processing (by the inter prediction unit (244) and the intra prediction unit (254)) performed by the exemplary video encoder (20) will be described in more detail.
[0164] Partitioning
[0165] The partitioning unit (262) may be configured to partition a picture from a video sequence into a sequence of coding tree units (CTUs), and the partitioning unit (262) may partition (or divide) the coding tree units (CTUs) (203) into smaller partitions, for example, smaller blocks of square or rectangular size. For a picture having a three-sample array, the CTU consists of an N×N block of luminance samples along with two corresponding blocks of chroma samples. The maximum allowable size of the luminance block in the CTU is specified as 128×128 in the universal video coding (VVC) under development, but this may be specified as a value other than 128×128 in the future, for example, 256×256. The CTUs of the picture may be clustered / grouped as slice / tile groups, tiles, or bricks. A tile covers a rectangular area of the picture, and a tile may be divided into one or more bricks. A brick consists of multiple rows of CTUs within a tile. A tile that is not partitioned into multiple bricks may be referred to as a brick. However, a brick is a true subset of tiles and is not referred to as a tile. There are two modes of tile groups supported in VVC: raster-scan slice / tile group mode and rectangular slice mode. In raster-scan tile group mode, a slice / tile group contains a sequence of tiles from a tile raster scan of the picture. In rectangular slice mode, a slice contains multiple bricks of the picture that collectively form a rectangular area of the picture. The bricks within a rectangular slice are the order of the slice's brick raster scan. These smaller blocks (which may also be referred to as subblocks) can be further partitioned into much smaller partitions.This is also referred to as tree partitioning or hierarchical tree partitioning, where, for example, a root block at root tree level 0 (hierarchy level 0, depth 0) can be recursively partitioned, for example, into two or more blocks of the next sub-tree level, for example, nodes of tree level 1 (hierarchy level 1, depth 1), and these blocks can be partitioned again into two or more blocks of the next sub-level, for example, tree level 2 (hierarchy level 2, depth 2), until partitioning is terminated, for example, because a termination criterion is satisfied, for example, because the maximum tree depth or minimum block size is reached. Additionally, blocks that are not partitioned are also referred to as leaf blocks or leaf nodes of the tree. A tree that uses partitioning into 2 partitions is referred to as a binary tree (BT), a tree that uses partitioning into 3 partitions is referred to as a ternary tree (TT), and a tree that uses partitioning into 4 partitions is referred to as a quad tree (QT).
[0166] For example, a coding tree unit (CTU) may be or include a CTB of luminance samples, two corresponding CTBs of chroma samples of a picture having a three-sample array, or a CTB of samples of a picture or monochrome picture coded using three distinct color planes and syntax structures used to code the samples. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value N such that the division of components into CTBs is partitioned. A coding unit (CU) may be or include a coding block of luminance samples, two corresponding coding blocks of chroma samples of a picture having a three-sample array, or a coding block of samples of a picture or monochrome picture coded using three distinct color planes and syntax structures used to code the samples. Correspondingly, a coding block (CB) may be an M×N block of samples for some values M and N such that the division of CTBs into coding blocks is partitioned.
[0167] In an embodiment, for example according to HEVC, a coding tree unit (CTU) can be partitioned into CUs by using a quadtree structure denoted as a coding tree. A decision on whether to code a picture region is made at the leaf CU level by using inter-picture (temporal) or intra-picture (spatial) prediction. Each leaf CU can be further partitioned into one, two, or four PUs based on a PU partitioning type. Within a single PU, the same prediction process is applied, and relevant information is transmitted to the decoder on a PU-by-PU basis. After obtaining residual blocks by applying the prediction process based on the PU partitioning type, the leaf CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.
[0168] In an embodiment, according to the latest video coding standard currently under development, referred to as Universal Video Coding (VVC), a combined quad tree nested multitype tree using binary and ternary split segmentation structures is used, for example, to partition coding tree units. In the coding tree structure within a coding tree unit, the CU may have a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned by a quaternary tree. Subsequently, the quaternary tree leaf nodes may be further partitioned by a multitype tree structure. The multitype tree structure includes four partitioning types: vertical binary split (SPLIT_BT_VER), horizontal binary split (SPLIT_BT_HOR), vertical ternary split (SPLIT_TT_VER), and horizontal ternary split (SPLIT_TT_HOR). Multitype tree leaf nodes are referred to as Coding Units (CUs), and if the CU is not too large for the maximum transformation length, this segmentation is used for prediction and transformation processing without any additional partitioning. This means that in most cases, CUs, PUs, and TUs have the same block size in a quadtree with an nested multitype tree coding block structure. An exception occurs when the supported maximum transformation length is smaller than the width or height of the CU's coding component. VVC develops a unique signaling mechanism for partitioning information in a quadtree with an nested multitype tree coding tree structure. In the signaling mechanism, the Coding Tree Unit (CTU) is treated as the root of the quaternary tree and is first partitioned by the quaternary tree structure. Subsequently, each quaternary tree leaf node (if it is large enough to allow this) is further partitioned by the multitype tree structure.In a multi-type tree structure, a first flag (mtt_split_cu_flag) is signaled to indicate whether a node is further partitioned; when a node is further partitioned, a second flag (mtt_split_cu_vertical_flag) is signaled to indicate the direction of the split, and then a third flag (mtt_split_cu_binary_flag) is signaled to indicate whether the split is a binary split or a ternary split. Based on the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of the CU can be derived by a decoder based on a predefined rule or table. It should be noted that for certain designs, such as the 64×64 luminance block and 32×32 chroma pipelining design in a VVC hardware decoder, TT splitting is prohibited when the width or height of the luminance coding block is greater than 64, as shown in Fig. 6. TT splitting is also prohibited when the width or height of the chroma coding block is greater than 32. The pipelining design will split the picture into virtual pipeline data units (VPDUs), which are defined as non-overlapping units in the picture. In a hardware decoder, successive VPDUs are processed simultaneously by multiple pipeline stages. Since the VPDU size is approximately proportional to the buffer size in most pipeline stages, it is important to keep the VPDU size small. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) and binary tree (BT) partitioning can lead to an increase in the VPDU size.
[0169] In addition, it should be noted that when part of a tree node block extends beyond the bottom or right picture boundary, the tree node block is forcibly split until all samples of each coded CU are located inside the picture boundary.
[0170] For example, an intra-subpartition (ISP) tool can divide a predicted block of Luma Intra vertically or horizontally into two or four subpartitions depending on the block size.
[0171] For example, the mode selection unit (260) of the video encoder (20) may be configured to perform any combination of the partitioning techniques described herein.
[0172] As previously described, the video encoder (20) is configured to determine or select the best or optimal prediction mode from a set of (e.g., predetermined) prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0173] Intra prediction
[0174] A set of intra prediction modes may include 35 different intra prediction modes, such as non-directional modes like DC (or average) mode and planar mode, or directional modes like those defined in HEVC, for example, or 67 different intra prediction modes, such as non-directional modes like DC (or average) mode and planar mode, or directional modes like those defined in VVC. For example, several conventional angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks, for example, as defined in VVC. As another example, to avoid division operations for DC prediction, only the longer side is used to calculate the average for non-square blocks. And the result of intra prediction in planar mode may be further modified by a position-dependent intra prediction combination (PDPC) method.
[0175] The intra prediction unit (254) is configured to generate an intra prediction block (265) according to the intra prediction mode of the set of intra prediction modes using a reconstructed sample of neighbor blocks of the same current picture.
[0176] The intra prediction unit (254) (or generally the mode selection unit (260)) is further configured to output intra prediction parameters (or generally information indicating the intra prediction mode selected for the block) to the entropy encoding unit (270) in the form of syntax elements (266) so as to be included in the encoded picture data (21), so that, for example, the video decoder (30) can receive and use the prediction parameters for decoding.
[0177] Inter-prediction (including inter-layer prediction)
[0178] A set (or possible) inter-prediction mode depends on an available reference picture (i.e., a previously at least partially decoded picture stored in the DPB (230)) and other inter-prediction parameters, such as whether the entire reference picture is used to search for the best matching reference block or only a part of the reference picture, such as a search window area around the area of the current block, and / or whether pixel interpolation is applied, such as half / semi-pel, quarter-pel, and / or 1 / 16 pixel interpolation.
[0179] In addition to the prediction modes above, skip mode, direct mode, and / or other inter-prediction modes may be applied.
[0180] For example, in extended merge prediction, the merge candidate list of such a mode is constructed by including the following five types of candidates in order: spatial MVP from spatial neighbor CUs, temporal MVP from collated CUs, history-based MVP from FIFO tables, pairwise average MVP, and zero MV. Additionally, decoder-side motion vector refinement (DMVR) based on bidirectional matching may be applied to increase the accuracy of the MV in the merge mode. Merge mode with MVD (MVD) is derived from merge mode based on motion vector difference. To determine whether MMVD mode is used for a CU, the MMVD flag is signaled immediately after transmitting the skip flag and merge flag. Furthermore, an adaptive motion vector resolution (AMVR) method at the CU level may be applied. AMVR allows the MVD of a CU to be coded with different precision. Depending on the prediction mode for the current CU, the MVD of the current CU can be adaptively selected. When the CU is coded in merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. A weighted average of the inter and intra prediction signals is performed to obtain the CIIP prediction. Affine motion compensation prediction, the affine motion field of the block is described by motion information of two control points (4-parameters) or three control point motion vectors (6-parameters).Subblock-based temporal motion vector prediction (SbTMVP) is similar to temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vectors of sub-CUs within the current CU. Bi-directional optical flow (BDOF), previously referred to as BIO, is a simpler version that requires significantly less computation, particularly in terms of the number of multiplications and the size of the multiplier. In the triangular partition mode, the CU is evenly divided into two triangular-shaped partitions using diagonal or anti-diagonal partitioning. Furthermore, the bidirectional prediction mode is extended beyond a simple average to enable a weighted average of the two prediction signals.
[0181] The inter-prediction unit (244) may include a motion estimation (ME) unit and a motion compensation (MC) unit (both not shown in FIG. 2). The motion estimation unit is configured to receive or acquire a picture block (203) (the current picture block (203) of the current picture (17)) and a decoded picture (231) for motion estimation, or at least one or more previously reconstructed blocks, e.g., reconstructed blocks, e.g., one or more other / different previously decoded pictures (231). For example, a video sequence may include a current picture and a previously decoded picture (231), or, in other words, the current picture and the previously decoded picture (231) may form a sequence of pictures forming a video sequence or be part of it.
[0182] The encoder (20) may be configured to select a reference block from among a plurality of reference blocks of different pictures, for example, the same or a plurality of different pictures, and to provide an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block and / or a reference picture (or reference picture index) as an inter-prediction parameter for the motion estimation unit. This offset is also called a motion vector (MV).
[0183] The motion compensation unit is configured to acquire, for example, receive inter-prediction parameters and to acquire inter-prediction blocks (265) by performing inter-prediction based on or using the inter-prediction parameters. Motion compensation performed by the motion compensation unit may involve fetching or generating prediction blocks based on motion / block vectors determined by motion estimation, and possibly performing interpolation with subpixel precision. Interpolation filtering may potentially increase the number of candidate prediction blocks that can be used to code picture blocks by generating additional pixel samples from known pixel samples. Upon receiving a motion vector for the PU of the current picture block, the motion compensation unit may find the location of the prediction block pointed to by the motion vector in one of the reference picture lists.
[0184] The motion compensation unit may also generate a block and a syntax element associated with the video slice for use by the video decoder (30) when decoding a picture block of the video slice. In addition to or alternative to the slice and individual syntax elements, a tile group and / or a tile and individual syntax elements may be generated or used.
[0185] Entropy Coding
[0186] The entropy encoding unit (270) is configured to apply or bypass (uncompress) an entropy encoding algorithm or method (e.g., variable length coding (VLC) method, context adaptive VLC (CAVLC) method, arithmetic coding method, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy encoding method or technique) to quantized residual coefficients (209), inter-predictor parameters, intra-predictor parameters, loop filter parameters, and / or other syntax elements in order to obtain encoded picture data (21) that can be output through an output (272) in the form of, for example, an encoded bitstream (21), so that, for example, video The decoder (30) can receive and use parameters for decoding. The encoded bitstream (21) can be transmitted to the video decoder (30) or stored in memory for subsequent transmission or retrieve by the video decoder (30).
[0187] Other structural variations of the video encoder (20) may be used to encode the video stream. For example, a non-transform-based encoder (20) may quantize the residual signal directly without a transform processing unit (206) for a specific block or frame. In other implementations, the encoder (20) may have a quantization unit (208) and an inverse quantization unit (210) combined into a single unit.
[0188] Decoder and Decoding Method
[0189] FIG. 3 illustrates an example of a video decoder (30) configured to implement the technology of the present application. The video decoder (30) is configured to receive encoded picture data (21) (e.g., encoded bitstream (21)), which is encoded by, for example, an encoder (20), and to obtain a decoded picture (331). The encoded picture data or bitstream includes information for decoding data representing the encoded picture data, such as a picture block (and / or tile group or tile) of an encoded video slice and associated syntax elements.
[0190] In the example of FIG. 3, the decoder (30) includes an entropy decoding unit (304), an inverse quantization unit (310), an inverse transformation processing unit (312), a reconstruction unit (314) (e.g., a summer (314)), a loop filter (320), a decoded picture buffer (DPB) (330), a mode application unit (360), an inter prediction unit (344), and an intra prediction unit (354). The inter prediction unit (344) may be a motion compensation unit or may include a motion compensation unit. In some examples, the video decoder (30) may perform a decoding pass that is generally opposite to the encoding pass described for the video encoder (100) from FIG. 2.
[0191] As described in relation to the encoder (20), the inverse quantization unit (210), the inverse transformation processing unit (212), the reconstruction unit (214), the loop filter (220), the decoded picture buffer (DPB) (230), the inter prediction unit (344), and the intra prediction unit (354) are also referred to as forming the "built-in decoder" of the video encoder (20). Accordingly, the inverse quantization unit (310) may be functionally identical to the inverse quantization unit (110), the inverse transformation processing unit (312) may be functionally identical to the inverse transformation processing unit (212), the reconstruction unit (314) may be functionally identical to the reconstruction unit (214), the loop filter (320) may be functionally identical to the loop filter (220), and the decoded picture buffer (330) may be functionally identical to the decoded picture buffer (230). Accordingly, the description provided for each unit and function of the video encoder (20) is applied to each unit and function of the video decoder (30).
[0192] Entropy decoding
[0193] The entropy decoding unit (304) is configured to parse the bitstream (21) (or generally the encoded picture data (21)) and, for example, perform entropy decoding on the encoded picture data (21) to obtain, for example, any or all parameters among the quantized coefficients (309) and / or decoded coding parameters (not shown in FIG. 3), such as inter-prediction parameters (e.g., reference picture index and motion vector), intra-prediction parameters (e.g., intra-prediction mode or index), transformation parameters, quantization parameters, loop filter parameters and / or other syntax elements. The entropy decoding unit (304) may be configured to apply a decoding algorithm or method corresponding to the encoding method described in relation to the entropy encoding unit (270) of the encoder (20). The entropy decoding unit (304) may be further configured to provide inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode application unit (360), and to provide other parameters to other units of the decoder (30). The video decoder (30) may receive syntax elements at the video slice level and / or video block level. In addition to or alternative to slices and individual syntax elements, tile groups and / or tiles and individual syntax elements may be received and / or used.
[0194] Inverse quantization
[0195] The dequantization unit (310) may be configured to receive quantization parameters (QP) (or information generally related to dequantization) and quantized coefficients from encoded picture data (21) (e.g., by parsing and / or decoding by the entropy decoding unit (304)) and to apply dequantization to the decoded quantized coefficients (309) based on the quantization parameters to obtain dequantized coefficients (311), which may also be referred to as transform coefficients (311). The dequantization process may include using quantization parameters determined by the video encoder (20) for each video block of the video slice (or tile or group of tiles) to determine the degree of quantization and, likewise, the degree of dequantization to be applied.
[0196] Inverse transformation
[0197] The inverse transformation processing unit (312) may be configured to receive an inverse quantized coefficient (311), also referred to as a transformation coefficient (311), and to apply a transformation to the inverse quantized coefficient (311) to obtain a reconstructed residual block (213) in the sample domain. The reconstructed residual block (213) may also be referred to as a transformation block (313). The transformation may be an inverse transformation, such as an inverse DCT, an inverse DST, an inverse integer transformation, or a conceptually similar inverse transformation process. The inverse transformation processing unit (312) may be further configured to receive transformation parameters or corresponding information from the encoded picture data (21) (e.g., by parsing and / or decoding by the entropy decoding unit (304)) to determine the transformation to be applied to the inverse quantized coefficient (311).
[0198] Reconstruction
[0199] The reconstruction unit (314) (e.g., an adder or summer (314)) may be configured to obtain a reconstruction block (315) in the sample domain by adding the reconstruction residual block (313) to the prediction block (365), for example, by adding the sample value of the reconstruction residual block (313) and the sample value of the prediction block (365).
[0200] Filtering
[0201] A loop filter unit (320) (inside or after the coding loop) is configured to filter the reconstructed block (315) to obtain a filtered block (321), for example, to smooth pixel transitions or improve video quality in other ways. The loop filter unit (320) may include one or more loop filters, such as a block separation filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination thereof. For example, the loop filter unit (220) may include a block separation filter, an SAO filter, and an ALF filter. The order of the filtering process may be block separation filter, SAO, and ALF. In another example, a process called luminance mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop reshaper) is added. This process is performed before block separation. In another example, the block separation filter process may also be applied to internal subblock edges, such as affine subblock edges, ATMVP subblock edges, subblock transformation (SBT) edges, and intra-subpartition (ISP) edges. In FIG. 3, the loop filter unit (320) is shown as an in-loop filter, but in other configurations, the loop filter unit (320) may be implemented as a post-loop filter.
[0202] Decoded picture buffer
[0203] Subsequently, the decoded video block (321) of the picture is stored in the decoded picture buffer (330), and the decoded picture buffer (330) stores the decoded picture (331) as a reference picture for subsequent motion compensation for another picture and / or for each output display.
[0204] The decoder (30) is configured to output, for example through an output (312), the decoded picture (311) to be presented or shown to the user.
[0205] prediction
[0206] The inter prediction unit (344) may be identical to the inter prediction unit (244) (in particular, motion compensation unit), and the intra prediction unit (354) may be identical to the inter prediction unit (254) in terms of function, and performs partitioning or partitioning determination and prediction based on each information or partitioning and / or prediction parameter received from the encoded picture data (21) (e.g., by parsing and / or decoding by the entropy decoding unit (304)). The mode application unit (360) may be configured to obtain a prediction block (365) by performing a prediction (intra or inter prediction, which may include inter-layer prediction) on each block based on the reconstructed picture, block, or each sample (filtered or unfiltered).
[0207] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit (354) of the mode application unit (360) is configured to generate a prediction block (365) for a picture block of the current video slice based on the intra-prediction mode and data signaled from a previously decoded block of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, the inter-prediction unit (344) of the mode application unit (360) (e.g., motion compensation unit) is configured to generate a prediction block (365) for a video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit (304). In the case of inter-prediction, the prediction block may be generated from one of the reference pictures within one of the reference picture lists. The video decoder (30) may configure reference frame lists, List 0 and List 1, using default configuration techniques based on the reference pictures stored in the DPB (330). For embodiments using a tile group (e.g., video tile group) and / or a tile (e.g., video tile) in addition to or alternatively to a slice (e.g., video slice), or may be applied in the same or similar manner by the embodiments, for example, a video may be coded using I, P, or B tile groups and / or tiles.
[0208] The mode application unit (360) is configured to determine prediction information for a video block of a current video slice by parsing motion vectors or related information and other syntax elements, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, the mode application unit (360) uses some of the received syntax elements to determine a prediction mode used to code a video block of a video slice (e.g., intra or inter prediction), an inter prediction slice type (e.g., B slice, P slice, or GPB slice), configuration information for one or more of a reference picture list for the slice, motion vectors for each inter-encoded video block of the slice, an inter prediction state for each inter-encoded video block of the slice, and other information for decoding a video block of the current video slice. For an embodiment using a tile group (e.g., video tile group) and / or a tile (e.g., video tile) in addition to or alternatively to a slice (e.g., video slice), or may be applied in the same or similar manner by the embodiment, for example, a video may be coded using I, P, or B tile groups and / or tiles.
[0209] An embodiment of the video decoder (30) as illustrated in FIG. 3 may be configured to partition and / or decode a picture by using slices (also referred to as video slices), wherein the picture may be partitioned into one or more slices (typically not overlapping) or decoded using one or more slices, and each slice may include one or more blocks (e.g., CTU) or one or more groups of blocks (e.g., tiles (H.265 / HEVC and VVC) or bricks (VVC)).
[0210] An embodiment of the video decoder (30) as illustrated in FIG. 3 may be configured to partition and / or decode a picture by using a slice / tile group (also referred to as a video tile group) and / or a tile (also referred to as a video tile), wherein the picture may be partitioned into one or more slice / tile groups (typically not overlapping) or decoded using one or more slice / tile groups, and each slice / tile group may include, for example, one or more blocks (e.g., CTU) or one or more tiles, and each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTU), for example, complete or partial blocks.
[0211] Other variations of the video decoder (30) may be used to decode the encoded picture data (21). For example, the decoder (30) may generate an output video stream without a loop filtering unit (320). For example, a non-transform-based decoder (30) may directly dequantize the residual signal for a specific block or frame without an inverse transformation processing unit (312). In other implementations, the video decoder (30) may have a dequantization unit (310) and an inverse transformation processing unit (312) combined into a single unit.
[0212] It should be understood that in the encoder (20) and decoder (30), the processing result of the current stage can be output to the next stage after further processing. For example, after interpolation filtering, motion vector derivation, or loop filtering, additional action, such as clipping or shifting, can be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0213] It should be noted that additional motions may be applied to the derived motion vector of the current block (including, but not limited to, control point motion vectors in affine mode, affine, planar, ATMVP mode, temporal motion vectors, etc.). For example, the value of the motion vector is limited to a predefined range according to its representation bit. If the representation bit of the motion vector is bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1, where "^" signifies exponentiation. For example, if bitDepth is equal to 16, the range is -32768 to 32767; if bitDepth is 18, the range is -131072 to 131071. For example, the values of the derived motion vectors (e.g., the MVs of four 4x4 subblocks within a single 8x8 block) are restricted such that the maximum difference between the integer parts of the four 4x4 subblock MVs is N pixels or less, such as 1 pixel or less. Here, two methods for restricting motion vectors according to bitDepth are provided.
[0214] FIG. 4 is a schematic diagram of a video coding device (400) according to one embodiment of the present disclosure. The video coding device (400) is suitable for implementing the disclosed embodiment as described in the present specification. In one embodiment, the video coding device (400) may be a decoder, such as the decoder (30) of FIG. 1a, or an encoder, such as the video encoder (20) of FIG. 1a.
[0215] A video coding device (400) includes an entry port (410) (or input port (410)) and a receiver unit (RX) (420) for receiving data; a processor, logic unit, or central processing unit (CPU) (430) for processing data; a transmitter unit (Tx) (440) and an exit port (450) (or output port (450)) for transmitting data; and a memory (460) for storing data. The video coding device (400) may also include an optical-to-electrical (OE) component and an electrical-to-optical (EO) component coupled to the entry port (410), receiver unit (420), transmitter unit (440), and exit port (450) for the entry or entry of an optical or electrical signal.
[0216] The processor (430) is implemented by hardware and software. The processor (430) may be implemented as one or more CPU chips, cores (e.g., as multi-core processors), FPGAs, ASICs, and DSPs. The processor (430) communicates with an entry port (410), a receiver unit (420), a transmitter unit (440), an exit port (450), and memory (460). The processor (430) includes a coding module (470). The coding module (470) implements the disclosed embodiments described above. For example, the coding module (470) implements a process or prepares or provides various coding operations. Thus, the inclusion of the coding module (470) provides substantial improvements to the functionality of the video coding device (400) and influences the transition of the video coding device (400) to other states. Alternatively, the coding module (470) may be implemented as instructions stored in memory (460) and executed by the processor (430).
[0217] The memory (460) may include one or more disks, tape drives, and solid-state drives and may be used as an overflow data storage device for storing such program when the program is selected for execution, and for storing instructions and data read during program execution. The memory (460) may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0218] FIG. 5 is a simplified block diagram of a device (500) that can be used as either or both of the source device (12) and the destination device (14) of FIG. 1, according to an exemplary embodiment.
[0219] The processor (502) within the device (500) may also be referred to as a central processing unit. Alternatively, the processor (502) may be any other type of device, or a number of devices capable of manipulating or processing information that is currently existing or will be developed in the future. The disclosed implementation may be carried out with a single processor, e.g., processor (502), as illustrated, but advantages in speed and efficiency may be achieved by using more than one processor.
[0220] The memory (504) of the device (500) may be a read-only memory (ROM) device or a random access memory (RAM) device in the implementation. Any other suitable type of storage device may be used as the memory (504). The memory (504) may contain code and data (506) accessed by the processor (502) using the bus (512). The memory (504) may further include an operating system (508) and an application program (510), the application program (510) may include at least one program that enables the processor (502) to perform the method described herein. For example, the application program (510) may include applications 1 through N, which further include a video coding application that performs the method described herein.
[0221] The device (500) may also include one or more output devices, such as a display (518). The display (518) may be, for example, a touch-sensitive display that combines the display with a touch-sensitive element operable to detect touch input. The display (518) may be coupled to the processor (502) via a bus (512).
[0222] Although illustrated as a single bus in this specification, the bus (512) of the device (500) may be composed of multiple buses. Additionally, the secondary storage (514) may be directly coupled to other components of the device (500) or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, the device (500) may be implemented in a wide variety of configurations.
[0223] Scalable coding
[0224] Scalable coding including quality scalable (PSNR scalable), spatial scalable, etc. For example, as shown in FIG. 6, a sequence can be down-sampled to a low spatial resolution version. Both the low spatial resolution version and the original spatial resolution (high spatial resolution) version will be encoded. And generally, the low spatial resolution will be coded first, and this will be used as a reference for the later coded high spatial resolution.
[0225] To describe the layer information (number, dependencies, output), the Video Parameter Set (VPS) is defined as follows:
[0226]
[0227] vps_max_layer_minus1 + 1 specifies the maximum allowed number of layers in each CVS referencing the VPS.
[0228] Same as 1 vps_all_independent_layer_flag vps_all_independent_layer_flag specifies that all layers within the CVS are coded independently without using inter-layer prediction. vps_all_independent_layer_flag, such as 0, specifies that one or more layers within the CVS may use inter-layer prediction. The value of vps_all_independent_layer_flag is inferred to be 1 if it does not exist. When vps_all_independent_layer_flag is 1, the value of vps_independent_layer_flag[i] is inferred to be 1. When vps_all_independent_layers_flag is 0, the value of vps_independent_layer_flag
[0000] is inferred to be 1.
[0229] vps_layer_id[ i ] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[ m ] will be less than vps_layer_id[ n ].
[0230] Same as 1 vps_independent_layer_flag [ i ] specifies that the layer with index i does not use cross-layer prediction. vps_independent_layer_flag[ i ], such as 0, specifies that the layer with index i can use cross-layer prediction, and vps_layer_dependency_flag[ i ] exists in the VPS.
[0231] Like 0 vps_direct_dependency_flag [ i ][ j ] specifies that the layer with index j is not a direct reference layer to the layer with index i. vps_direct_dependency_flag[ i ][ j ], such as 1, specifies that the layer with index j is a direct reference layer to the layer with index i. If vps_direct_dependency_flag[ i ][ j ] does not exist for i and j within the range from 0 to vps_max_layer_minus1, it is inferred to be 0.
[0232] The variable DirectDependentLayerIdx[ i ][ j ], which specifies the j-th direct dependent layer of the i-th layer, is derived as follows:
[0233]
[0234] The variable GeneralLayerIdx[i], which specifies the layer index of the layer where nuh_layer_id is the same as vps_layer_id[i], is derived as follows:
[0235]
[0236] A brief explanation is as follows:
[0237] vps_max_layer_minus1 + 1 means the number of layers.
[0238] vps_all_independent_layer_flag indicates whether all layers are coded independently.
[0239] vps_layer_id[ i ] indicates the layer ID of the i-th layer.
[0240] vps_independent_layer_flag[ i ] indicates whether the i-th layer is coded independently.
[0241] vps_direct_dependency_flag[ i ][ j ] indicates whether the j-th layer is used for references to the i-th layer.
[0242] Here, the syntax elements vps_independent_layer_flag[ i ] and vps_direct_dependency_flag[ i ][ j ] are inter-layer prediction information of the layers, i and j are layer identifiers, and different layers correspond to different layer identifiers.
[0243] DPB management and reference picture marking.
[0244] To manage such reference pictures during the decoding process, the decoded pictures need to be retained in the Decoding Picture Buffer (DPB) for reference use for subsequent picture decoding. To indicate such pictures, their picture order count (POC) information needs to be signaled directly or indirectly in the slice header. Typically, there are two reference picture lists, list0 and list1. Additionally, reference picture indices need to be included to signal pictures within the lists. In the case of unidirectional prediction, reference pictures are fetched from one reference picture list, and in the case of bidirectional prediction, reference pictures are fetched from two reference picture lists.
[0245] All reference pictures are stored in the DPB. All pictures within the DPB are marked as "Used for long-term reference," "Used for short-term reference," or "Not used for reference," and only one is marked for each of the three states. If a picture is marked as "Not used for reference," it will no longer be used for reference. If a picture also does not need to be stored for output, it can be removed from the DPB. The state of a reference picture can be signaled from the slice header or derived from the slice header information.
[0246] A new reference picture management method called the RPL (reference picture list) method has been proposed. The RPL proposes an entire set or sets of reference pictures for the currently coded picture, and the reference pictures within the set are used for decoding the current picture or future (later or subsequent) pictures. Therefore, the RPL reflects picture information in the DPB, and even if a reference picture is not used for reference to the current picture but is to be used for reference to a subsequent picture, this needs to be stored in the RPL.
[0247] After the picture is reconstructed, it will be stored in the DPB and marked as "Used for short-term reference" by default. DPB management operations will begin after parsing the RPL information within the slice header.
[0248] Reference Picture List Configuration.
[0249] Reference picture information can be signaled via the slice header. Additionally, some RPL candidates may exist in the sequence parameter set (SPS); in this case, the slice header may include an RPL index to obtain the necessary RPL information without signaling the entire RPL syntax structure. Alternatively, the entire RPL syntax structure may be signaled in the slice header.
[0250] Introduction of the RPL method.
[0251] To save cost bits for RPL signaling, some RPL candidates may exist in the SPS. A picture can use the RPL index (ref_pic_list_idx[ i ]) to retrieve its RPL information from the SPS. RPL candidates are signaled as follows:
[0252]
[0253] Semantics are as follows:
[0254] Same as 1 rpl1_same_as_rpl0_flag specifies that the syntax structures num_ref_pic_lists_in_sps
[0001] and ref_pic_list_struct( 1, rplsIdx ) do not exist, and the following applies:
[0255] - The value of num_ref_pic_lists_in_sps
[0001] is inferred to be the same as the value of num_ref_pic_lists_in_sps
[0000] .
[0256] - The value of each syntax element in ref_pic_list_struct( 1, rplsIdx ) is inferred to be the same as the value of the corresponding syntax element in ref_pic_list_struct( 0, rplsIdx ) for rplsIdx in the range of 0 to num_ref_pic_lists_in_sps
[0000] - 1.
[0257] num_ref_pic_lists_in_sps [ i ] specifies the number of ref_pic_list_struct( listIdx, rplsIdx ) syntax structures with i in listIdx included in SPS. The value of num_ref_pic_lists_in_sps[ i ] will be in the range from 0 to 64.
[0258] In addition to obtaining RPL information based on the RPL index from the SPS, RPL information can be signaled in the slice header.
[0259]
[0260] Same as 1 ref_pic_list_sps_flag [ i ] specifies that the reference picture list (i) of the current slice is derived based on one of the ref_pic_list_struct( listIdx, rplsIdx ) syntax structures in the SPS where listIdx is i. ref_pic_list_sps_flag[ i ], such as 0, specifies that the reference picture list (i) of the current slice is derived based on the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure where listIdx is i, which is directly included in the slice header of the current picture.
[0261] When ref_pic_list_sps_flag[ i ] does not exist, the following applies:
[0262] - If num_ref_pic_lists_in_sps[ i ] is equal to 0, the value of ref_pic_list_sps_flag[ i ] is inferred to be equal to 0.
[0263] - Otherwise (num_ref_pic_lists_in_sps[ i ] is greater than 0), if rpl1_idx_present_flag is equal to 0, the value of ref_pic_list_sps_flag
[0001] is inferred to be the same as ref_pic_list_sps_flag
[0000] .
[0264] - Otherwise, the value of ref_pic_list_sps_flag[ i ] is inferred to be the same as pps_ref_pic_list_sps_idc[ i ] - 1.
[0265] ref_pic_list_idx[ i ] specifies an index of the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure, where listIdx is equal to i, which is used to derive the reference picture list (i) of the current picture, in the list of the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure included in the SPS, where listIdx is equal to i. The syntax element ref_pic_list_idx[ i ] is represented by the Ceil( Log2( num_ref_pic_lists_in_sps[ i ] ) ) bit. The value of ref_pic_list_idx[ i ] is inferred to be equal to 0 if it does not exist. The value of ref_pic_list_idx[ i ] will be in the range from 0 to num_ref_pic_lists_in_sps[ i ] - 1. When ref_pic_list_sps_flag[ i ] is equal to 1 and num_ref_pic_lists_in_sps[ i ] is equal to 1, the value of ref_pic_list_idx[ i ] is inferred to be equal to 0. When ref_pic_list_sps_flag[ i ] is equal to 1 and rpl1_idx_present_flag is equal to 0, the value of ref_pic_list_idx
[0001] is inferred to be equal to ref_pic_list_idx
[0000] .
[0266] The variable RplsIdx[ i ] is derived as follows:
[0267]
[0268] slice_poc_lsb_lt [ i ][ j ] specifies the value of MaxPicOrderCntLsb modulo of the picture order count of the j-th LTRP entry in the i-th reference picture list. The length of the slice_poc_lsb_lt[ i ][ j ] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.
[0269] The variable PocLsbLt[ i ][ j ] is derived as follows:
[0270]
[0271] Same as 1 delta_poc_msb_present_flag [ i ][ j ] indicates that delta_poc_msb_cycle_lt[ i ][ j ] exists. delta_poc_msb_present_flag[ i ][ j ] equal to 0 indicates that delta_poc_msb_cycle_lt[ i ][ j ] does not exist.
[0272] Set prevTid0Pic to the previous picture that has the same nuh_layer_id as the current picture in decoding order, has a TemporalId equal to 0, and is not a RASL or RADL picture. Set setOfPrevPocVals to the set consisting of the following:
[0273] -PicOrderCntVal in prevTid0Pic;
[0274] - PicOrderCntVal of each picture referenced by an entry in RefPicList
[0000] or RefPicList
[0001] of prevTid0Pic and having the same nuh_layer_id as the current picture
[0275] - In the decoding order, the PicOrderCntVal of each picture following prevTid0Pic has the same nuh_layer_id as the current picture and precedes the current picture in the decoding order.
[0276] When there is more than one modulo MaxPicOrderCntLsb value in setOfPrevPocVals that is equal to PocLsbLt[ i ][ j ], the value of delta_poc_msb_present_flag[ i ][ j ] will be equal to 1.
[0277] delta_poc_msb_cycle_lt[ i ][ j ] specifies the value of the variable FullPocLt[ i ][ j ] as follows:
[0278]
[0279] The value of delta_poc_msb_cycle_lt[ i ][ j ] ranges from 0 to 2 (32 - log2_max_pic_order_cnt_lsb_minus4 - 4) It will be in the range up to. The value of delta_poc_msb_cycle_lt[ i ][ j ] is inferred to be 0 if it does not exist.
[0280] The syntax structure of RPL is as follows:
[0281]
[0282] num_ref_entries [ listIdx ][ rplsIdx ] specifies the number of entries in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure. The value of num_ref_entries[ listIdx ][ rplsIdx ] will be in the range from 0 to sps_max_dec_pic_buffering_minus1 + 14.
[0283] Like 0 ltrp_in_slice_header_flag [ listIdx ][ rplsIdx ] specifies that the POC LSB of an LTRP entry within the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure exists in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure. ltrp_in_slice_header_flag[ listIdx ][ rplsIdx ], as in 1, specifies that the POC LSB of an LTRP entry within the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure does not exist in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure.
[0284] Same as 1 inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] specifies that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an ILRP entry. inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ], which is equal to 0, specifies that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is not an ILRP entry. The value of inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is inferred to be equal to 0 if it does not exist.
[0285] Same as 1 st_ref_pic_flag [ listIdx ][ rplsIdx ][ i ] specifies that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an STRP entry. st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ], which is equal to 0, specifies that the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure is an LTRP entry. When inter_layer_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is equal to 0 and st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] does not exist, the value of st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is inferred to be equal to 1.
[0286] The variable NumLtrpEntries[ listIdx ][ rplsIdx ] is derived as follows:
[0287]
[0288] abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] specifies the value of the variable AbsDeltaPocSt[ listIdx ][ rplsIdx ][ i ] as follows:
[0289]
[0290] (7-121)
[0291] The value of abs_delta_poc_st[ listIdx ][ rplsIdx ][ i ] ranges from 0 to 2 15 - It will be in the range up to 1.
[0292] Same as 1 strp_entry_sign_flag [ listIdx ][ rplsIdx ][ i ] specifies that the i-th entry in the syntax structure ref_pic_list_struct( listIdx, rplsIdx ) has a value greater than or equal to 0. strp_entry_sign_flag[ listIdx ][ rplsIdx ][ i ], equal to 0, specifies that the i-th entry in the syntax structure ref_pic_list_struct( listIdx, rplsIdx ) has a value less than 0. The value of strp_entry_sign_flag[ listIdx ][ rplsIdx ][ i ] is inferred to be equal to 1 if it does not exist.
[0293] The list DeltaPocValSt[ listIdx ][ rplsIdx ] is derived as follows:
[0294]
[0295] rpls_poc_lsb_lt[ listIdx ][ rplsIdx ][ i ] specifies the value of MaxPicOrderCntLsb as the picture order count modulo of the picture referenced by the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure. The length of the rpls_poc_lsb_lt[ listIdx ][ rplsIdx ][ i ] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.
[0296] A general description of the RPL structure.
[0297] For each list, an RPL structure exists. First, num_ref_entries[ listIdx istrplsIdx ] is signaled to indicate the number of referenced pictures in the list. ltrp_in_slice_header_flag[ listIdx istrplsIdx ] is used to indicate whether the Least Significant Bit (LSB) information is signaled in the slice header. If the current referenced picture is not an inter-layer referenced picture, st_ref_pic_flag[ listIdx ][ rplsIdx ][ i ] is used to indicate whether that picture is a long-term referenced picture. If that picture is a short-term referenced picture, POC information (abs_delta_poc_st, strp_entry_sign_flag) is signaled. If ltrp_in_slice_header_flag[tlistIdx istrplsIdx ] is 0, rpls_poc_lsb_lt[plistIdx istrplsIdx plsj++ ] is used to derive the LSB information of the current reference picture. The MSB (Most Significant Bit) can be derived directly or based on the information in the slice header (delta_poc_msb_present_flag[ei ][ j ], delta_poc_msb_cycle_lt[ei ][ j ]).
[0298] Decoding process for constructing a reference picture list
[0299] This process is called at the start of the decoding process for each slice of the non-IDR picture.
[0300] Reference pictures are addressed via reference indices. A reference index is an index for a reference picture list. When decoding slice I, the reference picture list is not used for decoding the slice data. When decoding slice P, only reference picture list 0 (i.e., RefPicList
[0000] ) is used for decoding the slice data. When decoding slice B, both reference picture list 0 and reference picture list 1 (i.e., RefPicList
[0001] ) are used for decoding the slice data.
[0301] At the start of the decoding process for each slice of a non-IDR picture, reference picture lists RefPicList
[0000] and RefPicList
[0001] are derived. The reference picture lists are used for marking the reference pictures as specified in Clause 8.3.3 or for decoding the slice data.
[0302] Note 1 - In the case of the I slice of a non-IDR picture that is not the first slice of the picture in Fig. 1, RefPicList
[0000] and RefPicList
[0001] may be derived for bitstream conformity check purposes, but their derivation is not required for the decoding of the current picture or the picture following the current picture in the decoding order. In the case of the P slice that is not the first slice of the picture, RefPicList
[0001] may be derived for bitstream conformity check purposes, but their derivation is not required for the decoding of the current picture or the picture following the current picture in the decoding order.
[0303] The reference picture lists RefPicList
[0000] and RefPicList
[0001] are configured as follows:
[0304]
[0305] After the RPL is configured, where refPicLayerId is the ILRP layer identifier and PicOrderCntVal is the POC value, the marking process is as follows:
[0306] Decoding process for reference picture marking
[0307] This process is called once per picture, after the decoding of the slice header and the decoding process for constructing the list of reference pictures for the slice as specified in Clause 8.3.2, but before the decoding of the slice data. This process may cause one or more reference pictures in the DPB to be marked as "not used for reference" or "used for long-term reference".
[0308] Decoded pictures within the DPB may be marked as "not for reference," "for short reference," or "for long reference," but at any given moment during the operation of the decoding process, they may be marked as only one of these three. Assigning one of these markings to a picture implicitly removes the other of these markings, where applicable. When a picture is referred to as being marked as "for reference," this collectively means pictures marked as "for short reference" or "for long reference" (but not both).
[0309] STRP and ILRP are identified by their nuh_layer_id and PicOrderCntVal values. LTRP is identified by their nuh_layer_id value and Log2(MaxLtPicOrderCntLsb) of their PicOrderCntVal value.
[0310] If the current picture is a CLVSS picture, (if any) all referenced pictures in the current DPB with the same nuh_layer_id as the current picture are marked as "not used for reference".
[0311] Otherwise, the following applies:
[0312] - For each LTRP entry in RefPicList
[0000] or RefPicList
[0001] , if the referenced picture is an STRP with the same nuh_layer_id as the current picture, the picture is marked as "used for long-term reference".
[0313] - In DPB, each referenced picture with the same nuh_layer_id as the current picture that is not referenced by any entry in RefPicList
[0000] or RefPicList
[0001] is marked as "not used for reference".
[0314] - For each ILRP entry in RefPicList
[0000] or RefPicList
[0001] , the referenced picture is marked as "Used for long-term reference".
[0315] Note that here, the ILRP (inter-layer reference picture) is marked as "used for long-term reference."
[0316] There are two syntaxes for inter-layer reference information in SPS.
[0317]
[0318] sps_video_parameter_set_id When greater than 0, it specifies the value of vps_video_parameter_set_id for the VPS referenced by SPS. When sps_video_parameter_set_id is equal to 0, SPS does not reference a VPS, and no VPS is referenced when decoding each CVS that references SPS.
[0319] Like 0 long_term_ref_pics_flag isSpecifies that LTRP is not used for inter prediction of any coded picture in CVS. A long_term_ref_pics_flag such as 1 specifies that LTRP may be used for inter prediction of one or more coded pictures in CVS.
[0320] Like 0 inter_layer_ref_pics_present_flag specifies that ILRP is not used for inter-layer ref_pics_present_flag in CVS. An inter_layer_ref_pics_flag equal to 1 specifies that ILRP may be used for inter-layer ref_pics_present_flag in CVS. When sps_video_parameter_set_id is equal to 0, the value of inter_layer_ref_pics_present_flag is inferred to be equal to 0.
[0321] A brief explanation is as follows:
[0322] long_term_ref_pics_flag is used to indicate whether LTRP can be used in the decoding process.
[0323] inter_layer_ref_pics_present_flag is used to indicate whether ILRP can be used in the decoding process.
[0324] Therefore, when inter_layer_ref_pics_present_flag is equal to 1, there may be an ILRP used in the decoding process, which is marked as "used for long-term reference". In this case, there is a long_term_ref_pics_flag, such as LTRP used in the decoding process, or even 0. Thus, there is a discrepancy with the semantics of long_term_ref_pics_flag.
[0325] In existing methods, some syntax elements for inter-layer reference information are always signaled without considering the index of the current layer. The present invention proposes adding some conditions to the syntax elements to improve signaling efficiency.
[0326] Since long_term_ref_pics_flag is used only to control the parsing of ltrp_in_slice_header_flag and st_ref_pic_flag, the semantic is modified to control the parsing of flags parsed in the RPL.
[0327] Syntax elements for cross-layer reference information are signaled by considering the index of the current layer. If the information can be derived by the index of the current layer, the information does not need to be signaled.
[0328] Since long_term_ref_pics_flag is used only to control the parsing of ltrp_in_slice_header_flag and st_ref_pic_flag, the semantic is modified to control the parsing of flags parsed in the RPL.
[0329] Syntax elements for cross-layer reference information are signaled by considering the index of the current layer. If the information can be derived by the index of the current layer, the information does not need to be signaled.
[0330] First embodiment (modifying the semantics of long_term_ref_pics_flag to eliminate discrepancies between LTRP and ILRP)
[0331] Since long_term_ref_pics_flag is used solely to control the parsing of ltrp_in_slice_header_flag and st_ref_pic_flag, the semantics are modified as follows:
[0332] Same as 1 long_term_ref_pics_flag isSpecifies that ltrp_in_slice_header_flag and st_ref_pic_flag do not exist in the syntax structure ref_pic_list_struct( listIdx, rplsIdx ). long_term_ref_pics_flag, such as 0, specifies that these syntax elements do not exist in the syntax structure ref_pic_list_struct( listIdx, rplsIdx ).
[0333] Additionally, the semantic can be modified to exclude ILRP as follows:
[0334] Like 0 long_term_ref_pics_flag is Specifies that LTRP is not used for inter-layer reference picture prediction in CVS. A long_term_ref_pics_flag such as 1 specifies that LTRP may be used for inter-layer reference picture prediction in CVS. Here, LTRP does not include ILRP (inter-layer reference picture).
[0335] 2nd embodiment
[0336] Note here that when i is equal to 1, this means that layer1 needs to reference another layer. Only layer0 can be the reference layer, so vps_direct_dependency_flag[ i ][ j ] does not need to be signaled. Only when i is greater than 1 does vps_direct_dependency_flag[ i ][ j ] need to be signaled.
[0337]
[0338] Like 0 vps_direct_dependency_flag[ i ][ j ] specifies that the layer with index j is not a direct reference layer to the layer with index i. vps_direct_dependency_flag[ i ][ j ] equal to 1 specifies that the layer with index j is a direct reference layer to the layer with index i. When vps_direct_dependency_flag[ i ][ j ] does not exist for i and j within the range from 0 to vps_max_layer_minus1, if i is equal to 1 and vps_independent_layer_flag[ i ] is equal to 0, the value of vps_direct_dependency_flag[ i ][ j ] is inferred to be equal to 1, otherwise it is inferred to be equal to 0.
[0339] Third embodiment
[0340] Note that if sps_video_parameter_set_id (SPS level syntax element) is equal to 0, it means that there are no multiple layers, and therefore there is no need to signal inter_layer_ref_pics_flag (inter-layer enabled syntax element), and the flag is 0 by default.
[0341]
[0342] Like 0 inter_layer_ref_pics_present_flag specifies that ILRP is not used for inter prediction of any coded picture in CVS. An inter_layer_ref_pics_flag, such as 1, specifies that ILRP may be used for inter prediction of one or more coded pictures in CVS. When inter_layer_ref_pics_flag does not exist, it is inferred to be 0.
[0343] Note here that if GeneralLayerIdx[ nuh_layer_id ] is equal to 0, the current layer is the 0th layer, and it cannot reference any other layer. Therefore, there is no need to signal inter_layer_ref_pics_present_flag, and the value is 0 by default.
[0344]
[0345] Like 0 inter_layer_ref_pics_present_flag specifies that ILRP is not used for inter prediction of any coded picture in CVS. An inter_layer_ref_pics_flag, such as 1, specifies that ILRP may be used for inter prediction of one or more coded pictures in CVS. When inter_layer_ref_pics_flag does not exist, it is inferred to be 0.
[0346] If you code both of the two cases mentioned above, another application example is shown below:
[0347]
[0348] Like 0 inter_layer_ref_pics_present_flag specifies that ILRP is not used for inter prediction of any coded picture in CVS. An inter_layer_ref_pics_flag, such as 1, specifies that ILRP may be used for inter prediction of one or more coded pictures in CVS. When inter_layer_ref_pics_flag does not exist, it is inferred to be 0.
[0349] Fourth embodiment (to improve coding efficiency by eliminating redundancy information signaling, inter-layer reference information is signaled by considering the index of the current layer).
[0350] Note here that if GeneralLayerIdx[ nuh_layer_id ] is equal to 1, the current layer is layer1, which can only reference layer0, while layer0's ilrp_idc must be 0. Therefore, in this case, there is no need to signal ilrp_idc.
[0351]
[0352] ilrp_idc [ listIdx ][ rplsIdx ][ i ] specifies the index of the ILRP of the i-th entry in the ref_pic_list_struct( listIdx, rplsIdx ) syntax structure for the list of direct dependencies in the list of direct dependencies. The value of ilrp_idc[ listIdx ][ rplsIdx ][ i ] will be in the range from 0 to GeneralLayerIdx[ nuh_layer_id ] - 1. When GeneralLayerIdx[ nuh_layer_id ] is equal to 1, the value of ilrp_idc[ listIdx ][ rplsIdx ][ i ] is inferred to be equal to 0.
[0353] Fifth embodiment
[0354] It is noted here that some or all of the embodiments of Examples 1 to 4 can be combined to form a new embodiment.
[0355] For example, Example 1 + Example 2 + Example 3 + Example 4, or Example 2 + Example 3 + Example 4, or other combinations.
[0356] The following is a description of the encoding method as well as the decoding method as illustrated in the above-mentioned embodiment, and the application of the system using it.
[0357] FIG. 7 is a block diagram illustrating a content supply system (3100) for realizing a content distribution service. The content supply system (3100) includes a capture device (3102) and a terminal device (3106), and optionally includes a display (3126). The capture device (3102) communicates with the terminal device (3106) via a communication link (3104). The communication link may include the communication channel (13) described above. The communication link (3104) includes, but is not limited to, WiFi, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of any kind thereof.
[0358] The capture device (3102) generates data and can encode the data by an encoding method as illustrated in the above embodiment. Alternatively, the capture device (3102) can distribute the data to a streaming server (not illustrated in the drawing), and the server encodes the data and transmits the encoded data to a terminal device (3106). The capture device (3102) includes, but is not limited to, a camera, a smartphone or pad, a computer or laptop, a video conferencing system, a PDA, a vehicle-mounted device, or any combination of these. For example, the capture device (3102) may include a source device (12) as described above. When the data includes video, the video encoder (20) included in the capture device (3102) can actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device (3102) can actually perform audio encoding processing. In some actual scenarios, the capture device (3102) distributes the encoded video and audio data by multiplexing them together. In other actual scenarios, for example in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. The capture device (3102) distributes the encoded audio data and encoded video data individually to the terminal device (3106).
[0359] In the content delivery system (3100), the terminal device (310) receives and plays encoded data. The terminal device (3106) may be a device capable of receiving and restoring data, such as a smartphone or pad (3108), a computer or laptop (3110), a network video recorder (NVR) / digital video recorder (DVR) (3112), a TV (3114), a set-top box (STB) (3116), a video conferencing system (3118), a video surveillance system (3120), a personal digital assistant (PDA) (3122), a vehicle-mounted device (3124), or any combination of these, and may be capable of decoding the encoded data mentioned above. For example, the terminal device (3106) may include a destination device (14) as described above. When the encoded data contains video, the video decoder (30) included in the terminal device is prioritized to perform video decoding. When the encoded data contains audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0360] In the case of a terminal device having a display, for example, a smartphone or pad (3108), a computer or laptop (3110), a network video recorder (NVR) / digital video recorder (DVR) (3112), a TV (3114), a personal digital assistant (PDA) (3122), or a vehicle-mounted device (3124), the terminal device can supply decoded data to its display. In the case of a terminal device not equipped with a display, for example, an STB (3116), a video conferencing system (3118), or a video surveillance system (3120), an external display (3126) is contacted internally to receive and display the decoded data.
[0361] When each device of this system performs encoding or decoding, a picture encoding device or a picture decoding device such as that shown in the above-mentioned embodiment may be used.
[0362] FIG. 8 is a diagram illustrating an example structure of a terminal device (3106). After the terminal device (3106) receives a stream from a capture device (3102), a protocol processing unit (3202) analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, a Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any combination of any kind thereof.
[0363] After the protocol processing unit (3202) processes the stream, a stream file is generated. The file is output to the demultiplexing unit (3204). The demultiplexing unit (3204) can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some actual scenarios, for example in a video conferencing system, the encoded audio data and encoded video data are not multiplexed. In this situation, the encoded data is transmitted to the video decoder (3206) and audio decoder (3208) without passing through the demultiplexing unit (3204).
[0364] Through demultiplexing processing, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder (3206), including a video decoder (30) as described in the above-mentioned embodiment, decodes the video (ES) by a decoding method as illustrated in the above-mentioned embodiment to generate a video frame and feeds this data to a synchronization unit (3212). An audio decoder (3208) decodes the audio ES to generate an audio frame and feeds this data to the synchronization unit (3212). Alternatively, the video frame may be stored in a buffer (not shown in FIG. 8) before being fed to the synchronization unit (3212). Similarly, the audio frame may be stored in a buffer (not shown in FIG. 8) before being fed to the synchronization unit (3212).
[0365] The synchronization unit (3212) synchronizes video frames and audio frames and supplies video / audio to a video / audio display (3214). For example, the synchronization unit (3212) synchronizes the presentation of video and audio information. The information may be coded in syntax using timestamps regarding the presentation of coded audio and visual data and timestamps regarding the delivery of the data stream itself.
[0366] If subtitles are included in the stream, a subtitle decoder (3210) decodes the subtitles, synchronizes them with video frames and audio frames, and supplies the video / audio / subtitles to a video / audio / subtitle display (3216).
[0367] The present invention is not limited to the system mentioned above, and the picture encoding device or picture decoding device in the embodiment mentioned above may be integrated into other systems, for example, automotive systems.
[0368] mathematical operators
[0369] The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations, such as exponentiation and division of real values, are defined. Numbering and counting conventions generally start from 0, for example, "first" is equivalent to the 0th, and "second" is equivalent to the 1st.
[0370] arithmetic operators
[0371] The following arithmetic operators are defined as follows:
[0372] + addition
[0373] - Subtraction (as a two-factor operator) or negation (as a unary prefix operator)
[0374] Multiplication involving matrix multiplication
[0375] x y Exponentialization. Specifies x as a power of y. In other contexts, such notation is used for superscripts not intended for interpretation as exponentialization.
[0376] / Integer division by truncation of the result toward 0. For example, 7 / 4 and -7 / -4 are truncated to 1, and -7 / 4 and 7 / 4 are truncated to -1.
[0377] ÷ is used to denote division in mathematical expressions where truncation or rounding is not intended.
[0378] It is used to denote division in mathematical expressions where truncation or rounding is not intended.
[0379] The sum of f(i) where i takes all integer values from x to y, including y.
[0380] x % y modulus. The remainder of dividing x by y, defined only for integers x and y such that x >= 0 and y > 0.
[0381] logical operators
[0382] The following logical operators are defined as follows:
[0383] x && y Boolean logical "and" of x and y
[0384] Boolean logic of x || yx and y "or"
[0385] ! Boolean logic "not"
[0386] x ? y : z If x is TRUE or not 0, evaluate to the value of y; otherwise, evaluate to the value of z.
[0387] relational operators
[0388] The following relational operators are defined as follows:
[0389] Excess
[0390] >= more
[0391] <
[0392] <= Below
[0393] == Equals
[0394] != not equal to
[0395] When a relational operator is applied to a syntax element or variable to which a "na" (not applicable) value is assigned, the "na" value is treated as a distinct value for the syntax element or variable. The "na" value is considered not to be the same as any other value.
[0396] bit-wise operators
[0397] The following bitwise operators are defined as follows.
[0398] & bitwise "and". When operating on integer arguments, the operation is performed on the two's complement representation of the integer value. When operating on binary arguments containing fewer bits than the other argument, the shorter argument is extended by adding a more significant bit equal to 0.
[0399] | Bitwise "or". When operating on integer arguments, the operation is performed on the two's complement representation of the integer value. When operating on binary arguments containing fewer bits than other arguments, the shorter argument is extended by adding a more significant bit equal to 0.
[0400] ^ Bitwise "exclusive or". When operating on an integer argument, the operation is performed on the two's complement representation of the integer value. When operating on a binary argument containing fewer bits than the other argument, the shorter argument is extended by adding a more significant bit equal to 0.
[0401] Arithmetic right shift of the two's complement integer representation of x by the binary digit x >> yy. This function is defined only for non-negative integer values of y. As a result of the right shift, the bit shifted to the most significant bit (MSB) has the same value as the MSB of x before the shift operation.
[0402] Arithmetic left shift of the two's complement integer representation of x by the binary digit x << yy. This function is defined only for non-negative integer values of y. As a result of the left shift, the bit shifted to the least significant bit (LSB) has a value equal to 0.
[0403] assignment operator
[0404] The following arithmetic operators are defined as follows:
[0405] = assignment operator
[0406] ++ increment, that is x ++ isx = x Equivalent to + 1; when used in an array index, it evaluates to the value of the variable before the increment operation.
[0407] -- decrease, i.e., x-- Is x = x - Equivalent to 1; when used in an array index, it evaluates to the value of the variable before the decrement operation.
[0408] += an increment of a specified amount, that is, x += 3 is equal to x = x + 3, and x += (-3) is equal to x = x + (-3).
[0409] -= a specified amount of decrease, that is, x -= 3 is equivalent to x = x - 3, and x -= (-3) is equivalent to x = x - (-3).
[0410] Range notation
[0411] The following notation is used to specify a range of values:
[0412] x = y...zx takes integer values from y to z, where x, y, and z are integers and z is greater than y.
[0413] mathematical functions
[0414] The following mathematical function is defined:
[0415]
[0416] Asin( x ) is a trigonometric inverse sine function that operates on an argument x in the range of -1.0 to 1.0, with output values within the range of -π÷2 to π÷2 in radians.
[0417] Atan(x) is a trigonometric inverse tangent function that operates on argument x; the output value is within the range of -π÷2 to π÷2 in radians.
[0418]
[0419] Ceil( x ) The smallest integer greater than or equal to x.
[0420] Clip1 Y ( x ) = Clip3( 0, ( 1 << BitDepth Y ) - 1, x )
[0421] Clip1 C ( x ) = Clip3( 0, ( 1 << BitDepth C ) - 1, x )
[0422]
[0423] Cos( x ) is a trigonometric cosine function that operates on an argument x in radians.
[0424] Floor( x ) The largest integer less than or equal to x.
[0425]
[0426] Ln( x ) natural logarithm of x (logarithm with base e, where e is the natural logarithm base constant 2.718 281 828…).
[0427] Log2( x ) The logarithm of x with base 2.
[0428] Log10( x ) The logarithm of x with base 10.
[0429]
[0430]
[0431] Round(x) = Sign(x) * Floor(Abs(x) + 0.5)
[0432]
[0433] Sin(x) is a trigonometric sine function that operates on an argument x in radians.
[0434] Sqrt( x ) =
[0435] Swap( x, y ) = ( y, x )
[0436] Tan(x) is a trigonometric tangent function that operates on an argument x in radians.
[0437] Operation precedence
[0438] When the precedence of expressions is not explicitly indicated by the use of parentheses, the following rules apply:
[0439] - A higher priority operation is evaluated before an arbitrary operation of lower priority.
[0440] Operations of the same priority are evaluated sequentially from left to right.
[0441] The table below specifies the priority of operations from highest to lowest; higher positions in the table indicate higher priority.
[0442] For such operators also used in the C programming language, the precedence used in this specification is the same as that used in the C programming language.
[0443] Table: Operation priority from the highest (at the top of the table) to the lowest (at the bottom of the table)
[0444] Operation (using operands x, y, and z) "x++", "x--" "!x", "-x" (as unary prefix operators) x y "x * y", "x / y", "x ÷ y", , "x % y" "x + y", "x - y" (as a 2-factor operator), "x << y", "x >> y" "x < y", "x <= y", "x > y", "x >= y" "x == y", "x != y" "x & y" "x | y" "x && y" "x || y" "x ? y : z" "x‥y" "x = y", "x += y", "x -= y"
[0445] Text description of logical operations
[0446] In the text, a statement of a logical operation to be mathematically explained in the following form:
[0447] if( condition 0 )
[0448] Command 0
[0449] else if( condition 1 )
[0450] Imperative statement 1
[0451] …
[0452] else / * Useful comments on the remaining conditions * /
[0453] Command n
[0454] It can be explained in the following way:
[0455] … is as follows / … the following applies:
[0456] - If condition is 0, statement 0
[0457] - Otherwise, if condition 1, statement 1
[0458] - …
[0459] - Otherwise (a useful remark on the remaining conditions), statement n
[0460] In the text, the "…if. otherwise, …if. otherwise, …" statements are introduced into "…is as follows" or "…the following applies," immediately followed by "…if." The final condition of "…if. otherwise, …if. otherwise, …" is always "otherwise, …". The interleaved "…if. otherwise, …if. otherwise, …" statements can be identified by matching "…is as follows" or "…the following applies" with the final "otherwise, …".
[0461] In the text, a statement of a logical operation to be mathematically explained in the following form:
[0462] if( condition 0a && condition 0b )
[0463] Command 0
[0464] else if( condition 1a || condition 1b )
[0465] Imperative statement 1
[0466] …
[0467] else
[0468] Command n
[0469] It can be explained in the following way:
[0470] … is as follows / … the following applies:
[0471] - If all of the following conditions are true, statement 0:
[0472] - Condition 0a
[0473] - Condition 0b
[0474] - Otherwise, if one or more of the following conditions are true, statement 1:
[0475] - Condition 1a
[0476] - Condition 1b
[0477] - …
[0478] - Otherwise, statement n
[0479] In the text, a statement of a logical operation to be mathematically explained in the following form:
[0480] if( condition 0 )
[0481] Command 0
[0482] if( condition 1 )
[0483] Imperative statement 1
[0484] It can be explained in the following way:
[0485] When condition 0, statement 0
[0486] When condition 1, statement 1.
[0487] Although embodiments of the present invention have been described primarily based on video coding, it should be noted that the coding system (10), encoder (20) and decoder (30) (and corresponding system (10)) and other embodiments described herein may also be configured for still image processing or coding, that is, for the processing or coding of individual pictures independent of any preceding or consecutive pictures, as in video coding. Generally, when picture processing coding is limited to a single picture (17), only the inter-prediction unit (244 (encoder), 344 (decoder)) may not be available. All other functions of the video encoder (20) and video decoder (30) (also referred to as tools or techniques) can be used for still image processing, such as residual calculation (204 / 304), transformation (206), quantization (208), inverse quantization (210 / 310), (inverse) transformation (212 / 312), partitioning (262 / 362), intra prediction (254 / 354) and / or loop filtering (220, 320), and entropy coding (270) and entropy decoding (304).
[0488] For example, embodiments of the encoder (20) and decoder (30), and the functions described herein with reference to, for example, the encoder (20) and decoder (30), may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on a computer-readable medium or transmitted through a communication medium as one or more instructions or codes and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a type of medium such as a data storage medium, or a communication medium including any medium that enables the transfer of a computer program from one place to another according to a communication protocol, for example. In this way, a computer-readable medium may generally correspond to (1) a non-transient type of computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for the implementation of the technology described herein. A computer program product may include a computer-readable medium.
[0489] In particular, a method for decoding a coded video bitstream, implemented in a decoder as exemplified in FIG. 9, is provided, the method comprising the following steps: S901, obtaining a sequence parameter set (SPS) level syntax element from the bitstream ― specifying that an SPS level syntax element equal to a preset value indicates that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than the preset value indicates that the SPS references the VPS ―. S902, obtaining an inter-layer enable syntax element that, when the SPS level syntax element is greater than the preset value, indicates whether one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures; and S903, predicting one or more coded pictures based on the value of the inter-layer enable syntax element.
[0490] Similarly, a method for encoding a video bitstream containing coded data, implemented in an encoder as exemplified in FIG. 10, is provided, the method comprising the following steps: S1001, encoding a sequence parameter set (SPS) level syntax element into the bitstream ― SPS level syntax element equal to a preset value specifies that no video parameter set (VPS) is referenced by the SPS, and SPS level syntax element greater than a preset value specifies that the SPS references the VPS ―; S1003, encoding an inter-layer enable syntax element into the bitstream when the SPS level syntax element is greater than a preset value ― The inter-layer enable syntax element specifies whether one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures ―.
[0491] Additionally, this method may further include a step of determining whether S1002, an SPS level syntax element is greater than a preset value.
[0492] FIG. 11 illustrates a decoder (1100) configured to decode a video bitstream containing coded data for a plurality of pictures. The decoder (1100) according to the illustrated example comprises: an acquisition unit (1110) configured to acquire a sequence parameter set (SPS) level syntax element from the bitstream—where an SPS level syntax element equal to a preset value specifies that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than the preset value specifies that the SPS references the VPS; and the acquisition unit (1110) is further configured to acquire an inter-layer enable syntax element that specifies whether, when an SPS level syntax element is greater than the preset value, one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures. It includes a prediction unit (1120) configured to predict one or more coded pictures based on the values of enable syntax elements between layers.
[0493] Here, the acquisition unit may be an entropy decoding unit (304). The prediction unit (1120) may be an inter-prediction unit (344). The decoder (1100) may be a destination device (14), a decoder (30), a device (500), a video decoder (3206), or a terminal device (3106).
[0494] Similarly, an encoder (1200) configured to encode a video bitstream containing coded data for a plurality of pictures as exemplified in FIG. 12 is provided. The encoder (1200) comprises: a first encoding unit (1210) configured to encode a sequence parameter set (SPS) level syntax element into a bitstream—where an SPS level syntax element equal to a preset value specifies that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than a preset value specifies that the SPS references the VPS—; and a second encoding unit (1220) configured to encode an inter-layer enable syntax element into a bitstream when the SPS level syntax element is greater than a preset value, wherein the inter-layer enable syntax element specifies whether one or more inter-layer reference pictures (ILRP) can be used for inter-predicting of one or more coded pictures.
[0495] In a possible implementation of the method according to this fourth aspect, the encoder further includes a determination unit configured to determine whether an SPS level syntax element is greater than a preset value.
[0496] The first encoding unit (1210) and the second encoding unit (1220) may be an entropy encoding unit (270). The determination unit may be a mode selection unit (260). The encoder (1200) may be a source device (12), an encoder (20), or a device (500).
[0497] By way of example, not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disc storage or other magnetic storage devices, flash memory, or any other medium accessible by a computer that can be used to store desired program code in the form of instructions or data structures. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, and microwave is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals, or other transient media, but instead relate to non-transient types of storage media. Discs (disk and disc) as used herein include compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, wherein a disc typically reproduces data magnetically, while a disc reproduces data optically by a laser. A combination of the above should also be included within the scope of computer-readable media.
[0498] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the term “processor” as used herein may mean any of the aforementioned structures or any other structures suitable for the implementation of the technology described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be integrated into a combined codec. Furthermore, this technology may be fully implemented by one or more circuits or logic elements.
[0499] The technology of the present disclosure may be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). While various components, modules, or units are described in the present disclosure to highlight the functional aspects of a device configured to perform the disclosed technology, they do not necessarily require realization by different hardware units. Rather, as previously described, various units may be combined into a codec hardware unit, or provided by a set of interacting hardware units comprising one or more processors as previously described, together with appropriate software and / or firmware.
Claims
Claim 1 A method for decoding a coded video bitstream, comprising: a step of obtaining a sequence parameter set (SPS) level syntax element from the bitstream — specifying that if the SPS level syntax element is equal to a preset value, no video parameter set (VPS) is referenced by the SPS, and if the SPS level syntax element is greater than the preset value, the SPS references the VPS —; a step of deriving a variable GeneralLayerIdx[i] that specifies the layer index of a layer having a nuh_layer_id equal to vps_layer_id[i] — where vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer —; a step of obtaining an inter-layer enable syntax element from the bitstream when the SPS level syntax element is greater than the preset value — the inter-layer enable syntax element A method for decoding a coded video bitstream comprising the steps of: determining whether one or more inter-layer reference pictures (ILRP) can be used for inter-prediction of one or more coded pictures; when the SPS level syntax element is equal to the preset value, the inter-layer enable syntax element does not exist in the bitstream, and the value of the inter-layer enable syntax element is inferred to be equal to 0, and when GeneralLayerIdx[nuh_layer_id] is equal to 0, the value of the inter-layer enable syntax element is equal to 0 by default; and predicting one or more coded pictures based on the value of the inter-layer enable syntax element. Claim 2 A method for decoding a coded video bitstream according to claim 1, wherein the VPS includes a syntax element describing inter-layer prediction information of a layer in a coded video sequence (CVS), the SPS includes the SPS level syntax element and the inter-layer enable syntax element, and the CVS includes the one or more ILRPs and the one or more coded pictures. Claim 3 In paragraph 2, the step of predicting one or more coded pictures based on the value of the inter-layer enable syntax element comprises: a step of predicting one or more coded pictures by referencing the one or more ILRPs when it becomes possible to use the value of the inter-layer enable syntax element specifying the one or more inter-layer reference pictures (ILRPs) for inter-predicting one or more coded pictures, wherein the one or more ILRPs are obtained based on inter-layer prediction information included in the VPS referenced by the SPS, a method for decoding a coded video bitstream. Claim 4 A method for decoding a coded video bitstream according to claim 1, wherein the coded picture and the ILRP of the coded picture belong to different layers. Claim 5 A method for decoding a coded video bitstream according to claim 1, wherein an SPS level syntax element identical to a preset value further specifies that the coded video sequence (CVS) contains only one layer of coded picture. Claim 6 A method for decoding a coded video bitstream according to claim 1, wherein the preset value is 0. Claim 7 A method for decoding a coded video bitstream according to claim 1, wherein the step of predicting one or more coded pictures based on the value of the inter-layer enable syntax element comprises: predicting one or more coded pictures without referencing any ILRP when the value of the inter-layer enable syntax element specifying the one or more ILRPs is not used for inter-predicting of one or more coded pictures. Claim 8 A method for encoding a coded video bitstream, comprising the step of encoding a sequence parameter set (SPS) level syntax element into the bitstream — specifying that an SPS level syntax element equal to a preset value indicates that no video parameter set (VPS) is referenced by the SPS, and an SPS level syntax element greater than the preset value indicates that the SPS references the VPS —; when the SPS level syntax element is greater than the preset value, the step of encoding an inter-layer enable syntax element into the bitstream — the inter-layer enable syntax element specifies whether one or more inter-layer reference pictures (ILRPs) can be used for inter-predicting of one or more coded pictures; and when the SPS level syntax element is equal to the preset value, the inter-layer enable syntax element does not exist in the bitstream, and the value of the inter-layer enable syntax element is inferred to be equal to 0 —; A method for encoding a coded video bitstream, comprising the step of deriving a variable GeneralLayerIdx[i] that specifies the layer index of a layer having the same nuh_layer_id as vps_layer_id[i] — wherein vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer — and when GeneralLayerIdx[nuh_layer_id] is equal to 0, the value of the inter-layer enable syntax element is equal to 0 by default. Claim 9 A method for encoding a coded video bitstream according to claim 8, wherein the VPS includes a syntax element describing inter-layer prediction information of a layer in a coded video sequence (CVS), the SPS includes the SPS level syntax element and the inter-layer enable syntax element, and the CVS includes the one or more ILRPs and the one or more coded pictures. Claim 10 A method for encoding a coded video bitstream, wherein, in paragraph 8, the coded picture and the ILRP of the coded picture belong to different layers. Claim 11 In claim 8, a method for encoding a coded video bitstream, wherein an SPS level syntax element identical to a preset value further specifies that the coded video sequence (CVS) contains only one layer of coded picture. Claim 12 A method for encoding a coded video bitstream, wherein the preset value is 0, in claim 8. Claim 13 A method for encoding a coded video bitstream, wherein, in claim 8, the step of encoding the inter-layer enable syntax element into the bitstream comprises: the step of encoding the inter-layer enable syntax element into the bitstream, which specifies that the one or more ILRPs can be used for inter-predicting of one or more coded pictures based on a determination that the one or more ILRPs can be used for inter-predicting of one or more coded pictures. Claim 14 A method for encoding a coded video bitstream, wherein, in claim 8, the step of encoding the inter-layer enable syntax element into the bitstream comprises: the step of encoding the inter-layer enable syntax element into the bitstream, which specifies that the one or more ILRPs are not used for inter-predicting of one or more coded pictures based on a determination that the one or more ILRPs are not used for inter-predicting of one or more coded pictures. Claim 15 A decoding device comprising a processing circuit for executing a method according to any one of claims 1 to 7. Claim 16 An encoding device comprising a processing circuit for executing a method according to any one of claims 8 to 14. Claim 17 A computer program stored on a computer-readable storage medium, comprising program code that causes the computer to perform a method according to any one of claims 1 to 14 when executed by a computer. Claim 18 A decoding device comprising: one or more processors as a decoding device; and a non-transient computer-readable storage medium coupled to said processors and storing a program for execution by said processors, wherein said program is configured to execute a method according to any one of claims 1 through 7 when executed by said processors, a decoding device. Claim 19 An encoding device comprising: one or more processors as an encoding device; and a non-transient computer-readable storage medium coupled to said processors and storing a program for execution by said processors, wherein said programming configures said encoding device to execute a method according to any one of claims 8 through 14 when executed by said processors. Claim 20 A non-transient computer-readable storage medium containing program code that, when executed by a computer device, causes the computer device to perform the method of any one of claims 1 to 14. Claim 21 A non-transient storage medium comprising an encoded bitstream decoded by an image decoding device, wherein the bitstream comprises a plurality of syntax elements, the plurality of syntax elements comprises a sequence parameter set (SPS)-level syntax element, wherein the SPS-level syntax element being equal to a preset value specifies that no video parameter set (VPS) is referenced by the SPS, and the SPS-level syntax element being greater than the preset value specifies that the SPS references the VPS; wherein when the SPS-level syntax element is greater than the preset value, the plurality of syntax elements further comprise an inter-layer enable syntax element specifying whether one or more inter-layer reference pictures (ILRPs) can be used for inter-predicting of one or more coded pictures; wherein when the SPS-level syntax element is equal to the preset value, the inter-layer enable syntax element does not exist in the bitstream, and the layer A non-transient storage medium in which the value of an inter-layer enable syntax element is inferred to be equal to 0, and when the variable GeneralLayerIdx[nuh_layer_id] is equal to 0, the value of the inter-layer enable syntax element is equal to 0 by default, GeneralLayerIdx[i] is derived and specifies the layer index of the layer having a nuh_layer_id equal to vps_layer_id[i], wherein vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. Claim 22 delete Claim 23 delete Claim 24 delete