Avoiding SPS errors in sub-bitstream extraction
By correctly identifying and retaining the sequence parameter set (SPS) in the video encoding sub-bitstream extraction process, the hierarchical decoding error caused by SPS errors in the prior art is solved, and encoding efficiency and stability are improved.
Patent Information
- Application Number
- JP2024008764
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-07
- Filing Date
- 2024-01-24
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2040-10-06
AI Technical Summary
In the process of extracting sub-bitstream, existing video encoding technologies are prone to hierarchical decoding errors due to mistaken deletion of sequence parameter sets (SPS), which affects video quality and encoding efficiency.
During the sub-bitstream extraction process, the sequence parameter set (SPS) is correctly identified and retained based on whether the encoding hierarchy uses intersection layer prediction, and set independent layer flags if necessary to avoid accidentally deleting SPS.
It effectively avoids hierarchical decoding errors caused by SPS error deletion, improves the stability and encoding efficiency of video encoding, and reduces resource consumption of encoder and decoder.
Smart Images

Figure 0007675231000010 
Figure 0007675231000011 
Figure 0007675231000012
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims priority to U.S. Provisional Patent Application No. 62 / 911,808, entitled "Scalability in Video Coding," filed on October 7, 2019 by Ye-Kui Wang, which is incorporated herein by reference.
[0002] This disclosure relates generally to video coding, and more particularly to sub-bitstream A mechanism to prevent errors when extraction is performed on a multi-layer bitstream. It is related to ism. [Background technology]
[0003] The amount of video data required to render even a relatively short video is substantially large. As a result, data is streamed or bandwidth capacity is limited. Complications can arise when transmitted in any other manner over a communications network. Therefore, in today's telecommunications networks, video data is compressed before being transmitted. The size of the video may be limited, and memory resources may be limited. This can be problematic when the video is stored on a storage device. Compression devices use software and / or hardware at the source to transmit or It codes the video data prior to storage, thereby representing a digital video image. The compressed data is then used to render the video. The video data is received by a video decompression device that decodes it. Due to limited resources and ever-increasing demands for video quality, Improved compression and decompression techniques that improve compression ratios without sacrificing much or anything are desirable. It is nice. Summary of the Invention [Means for solving the problem]
[0004] In one embodiment, the present disclosure includes a method implemented by a decoder, the method comprising: The decoder obtains a coded layer video sequence (CLVS) for the layer, and A bitstream containing a sequence parameter set (SPS) referenced by CLVS When the layer does not use inter-layer prediction, the SPS receives the nuh_layer of the CLVS. A Network Abstraction Layer (NAL) unit header layer identifier (nuh_layer_i) equal to the _id value. d) receiving, by a decoder, a coding signal from the CLVS based on the SPS having a value; and decoding the coded picture to generate a decoded picture.
[0005] Some video coding systems code a video sequence into layers of pictures. Pictures in different layers have different characteristics. The coder can transmit different layers to the decoder depending on the constraints on the decoder side. To perform this function, the encoder compresses all layers into a single bitstream. On demand, the encoder can encode A sub-bitstream extraction process can be performed to remove irrelevant information. The result is an extracted bitmap that contains only the data in the layer required by the decoder. The description of how the layers are related is given in the video parameters. Simulcast layers can be included in a dataset (VPS). Simulcast layers refer to other layers. Simulcast layers are layers that are set to display without being When the sub-bitstream is transmitted to the You can safely remove the VPS as it is not needed to decode the trailer. Some variables in other parameter sets may refer to the VPS. Removing the VPS when the cast layer is transmitted can improve coding efficiency. Furthermore, failure to correctly identify the SPS may result in As a result, the SPS may be accidentally removed along with the VPS. This is because the SPS is missing. This can be problematic as in some cases the layers cannot be decoded correctly at the decoder. Examples of the present invention avoid errors when the VPS is removed during sub-bitstream extraction. Specifically, the sub-bitstream extraction process includes a mechanism for Remove NAL units based on r_ids. SPS is a layer for which CLVS does not use inter-layer prediction. has a nuh_layer_id equal to the nuh_layer_id of the CLVS referencing the SPS when A layer that does not use inter-layer prediction is a simulcast layer. Therefore, the SPS will have the same nuh_layer_id as the layer when the VPS was removed. In this case, the SPS is incorrectly removed by the sub-bitstream extraction process. The use of inter-layer prediction is enabled by the VPS independent layer flag (vps_independent_layer_fl ag). However, the vps_independent_layer_flag is signaled by Therefore, when an SPS does not refer to a VPS, vps_independent_la vps_i for the current layer, represented as yer_flag[GeneralLayerIdx[nuh_layer_id]] Independent_layer_flag is inferred to be 1. The SPS is inferred from the SPS VPS identifier (sps_video_param When vps_independent_layer_id is set to 0, the VPS is not referenced. When flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1 (e.g., when the VPS does not exist) ), the decoder and / or the hypothetical reference decoder (HRD) must be aware that the current layer / CLVS is inter-layer predicted. By adopting this line of reasoning, we can infer that SP S is supported even when the VPS and corresponding parameters are removed from the bitstream. Include the appropriate nuh_layer_id to avoid extraction by the bitstream extraction process As a result, the functionality of the encoder and decoder is improved. ,The coding efficiency is improved by reducing the unnecessary,data from a bitstream that contains only the simulcast layer. This is enhanced by successfully removing the VPS, which allows both the encoder and the decoder use of processor, memory, and / or network signaling resources in Reduce the dose.
[0006] Optionally, in any of the preceding aspects, another implementation of the aspect further comprises: Specifies that x[nuh_layer_id] is equal to the current layer index.
[0007] Optionally, in any of the aforementioned aspects, another implementation of the aspect further comprises: It is stipulated that t_layer_flag specifies whether the corresponding layer uses inter-layer prediction. Determine.
[0008] Optionally, in any of the aforementioned aspects, another implementation of the aspect further comprises: When t_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the layer is inter-layer predicted. It is stipulated that the following shall not be used.
[0009] Optionally, in any of the aforementioned aspects, another implementation of the aspect is to provide a method for detecting a plurality of SPS-based retrievals by the SPS. sps_video_parameter_set_id, which specifies the identifier (ID) value for the VPS referenced by When sps_video_parameter_set_id is equal to 0, vps_independent_layer_flag[Genera Specifies that lLayerIdx[nuh_layer_id]] is inferred to be equal to 1.
[0010] Optionally, in any of the aforementioned aspects, another implementation of the aspect further comprises: Specifies that when meter_set_id is equal to 0, the SPS does not reference a VPS.
[0011] Optionally, in any of the above aspects, another implementation of the aspect is to provide a method for CLVS to use the same nuh_ It specifies that the layer_id is a sequence of coded pictures having the same layer_id value.
[0012] In one embodiment, the present disclosure includes a method implemented by an encoder, The encoder encodes the CLVS for the layer in the bitstream. The encoder then encodes the SPS referenced by the CLVS into the bitstream. When a layer does not use inter-layer prediction, SPS uses nuh_layer of CLVS. nuh_layer_id value equal to the nuh_layer_id value of the and storing the bitstream for communication by the reader to a decoder. .
[0013] Some video coding systems code a video sequence into layers of pictures. Pictures in different layers have different characteristics. The coder can transmit different layers to the decoder depending on the constraints on the decoder side. To perform this function, the encoder compresses all layers into a single bitstream. On demand, the encoder can encode A sub-bitstream extraction process can be performed to remove irrelevant information. The result is an extracted bitmap that contains only the data in the layer required by the decoder. The description of how the layers are related is given in the video parameters. Simulcast layers can be included in a dataset (VPS). Simulcast layers refer to other layers. Simulcast layers are layers that are set to display without being When the sub-bitstream is transmitted to the You can safely remove the VPS as it is not needed to decode the trailer. Some variables in other parameter sets may refer to the VPS. Removing the VPS when the cast layer is transmitted can improve coding efficiency. Furthermore, failure to correctly identify the SPS may result in As a result, the SPS may be accidentally removed along with the VPS. This is because the SPS is missing. This can be problematic as in some cases the layers cannot be decoded correctly at the decoder. Examples of the present invention avoid errors when the VPS is removed during sub-bitstream extraction. Specifically, the sub-bitstream extraction process includes a mechanism for Remove NAL units based on r_ids. SPS is a layer for which CLVS does not use inter-layer prediction. has a nuh_layer_id equal to the nuh_layer_id of the CLVS referencing the SPS when A layer that does not use inter-layer prediction is a simulcast layer. Therefore, the SPS will have the same nuh_layer_id as the layer when the VPS was removed. In this case, the SPS is incorrectly removed by the sub-bitstream extraction process. The use of inter-layer prediction is signaled by vps_independent_layer_flag. However, the vps_independent_layer_flag is removed when the VPS is removed. Therefore, when an SPS does not refer to a VPS, vps_independent_layer_flag[GeneralLayer vps_independent_layer_fla for the current layer, represented as [rIdx[nuh_layer_id]] g is inferred to be 1. SPS is inferred to be VP Does not reference S. In addition, vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] When is equal to 1 (e.g., when the VPS is not present), the decoder and / or HRD are currently It can be inferred that the layer / CLVS does not use inter-layer prediction. By adopting the above logic, the SPS can extract the VPS and corresponding parameters from the bitstream. to avoid extraction by the sub-bitstream extraction process even when The result is that the encoder and decoder can be constrained to include the appropriate nuh_layer_id for In addition, the coding efficiency is improved by using only the simulcast layer. This is enhanced by successfully removing unnecessary VPS from the bitstream, This reduces the processor, memory, and / or network overhead in both the encoder and decoder. This reduces the usage of network signaling resources.
[0014] Optionally, in any of the preceding aspects, another implementation of the aspect further comprises: Specifies that x[nuh_layer_id] is equal to the current layer index.
[0015] Optionally, in any of the aforementioned aspects, another implementation of the aspect further comprises: It is stipulated that t_layer_flag specifies whether the corresponding layer uses inter-layer prediction. Determine.
[0016] Optionally, in any of the aforementioned aspects, another implementation of the aspect further comprises: When t_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the layer is inter-layer predicted. It is stipulated that the following shall not be used.
[0017] Optionally, in any of the aforementioned aspects, another implementation of the aspect is to provide a method for detecting a plurality of SPS-based retrievals by the SPS. sps_video_parameter_set_id, which specifies the ID value for the VPS referenced by sps_v When video_parameter_set_id is equal to 0, vps_independent_layer_flag[GeneralLayerIdx Specifies that [nuh_layer_id] is inferred to be equal to 1.
[0018] Optionally, in any of the aforementioned aspects, another implementation of the aspect further comprises: Specifies that when meter_set_id is equal to 0, the SPS does not reference a VPS.
[0019] Optionally, in any of the above aspects, another implementation of the aspect is to provide a method for CLVS to use the same nuh_ It specifies that the layer_id is a sequence of coded pictures having the same layer_id value.
[0020] In one embodiment, the present disclosure includes a video coding device, which includes a processor. a processor coupled to the processor; a memory coupled to the processor; and a transmitter coupled to the processor, the receiver, the memory, and the transmitter according to the above-mentioned aspects. The present invention is configured to perform any one of the above methods.
[0021] In one embodiment, the present disclosure provides a method for using a video coding device. A non-transitory computer-readable medium containing a computer program product, The program product, when executed by a processor, A method for performing the method of any of the above aspects, comprising: The computer-executable instructions include:
[0022] In one embodiment, the present disclosure includes a decoder that performs CLVS and CL A receiving means for receiving a bitstream including an SPS referenced by a VS, the receiving means comprising: PS is nuh_layer_id, which is equal to the nuh_layer_id value of CLVS when the layer does not use inter-layer prediction. A receiving means for receiving a coded picture from a CLVS based on an SPS, the receiving means having an id value and decoding the coded picture from the CLVS based on the SPS. a decoding means for decoding the video sequence to generate a decoded picture; a transfer means for transferring the decoded picture for display as part of a sequence; Equipped with.
[0023] Some video coding systems code a video sequence into layers of pictures. Pictures in different layers have different characteristics. The coder can transmit different layers to the decoder depending on the constraints on the decoder side. To perform this function, the encoder compresses all layers into a single bitstream. On demand, the encoder can encode A sub-bitstream extraction process can be performed to remove irrelevant information. The result is an extracted bitmap that contains only the data in the layer required by the decoder. The description of how the layers are related is given in the video parameters. Simulcast layers can be included in a dataset (VPS). Simulcast layers refer to other layers. Simulcast layers are layers that are set to display without being When the sub-bitstream is transmitted to the You can safely remove the VPS as it is not needed to decode the trailer. Some variables in other parameter sets may refer to the VPS. Removing the VPS when the cast layer is transmitted can improve coding efficiency. Furthermore, failure to correctly identify the SPS may result in As a result, the SPS may be accidentally removed along with the VPS. This is because the SPS is missing. This can be problematic as in some cases the layers cannot be decoded correctly at the decoder. Examples of the present invention avoid errors when the VPS is removed during sub-bitstream extraction. Specifically, the sub-bitstream extraction process includes a mechanism for Remove NAL units based on r_ids. SPS is a layer for which CLVS does not use inter-layer prediction. has a nuh_layer_id equal to the nuh_layer_id of the CLVS referencing the SPS when A layer that does not use inter-layer prediction is a simulcast layer. Therefore, the SPS will have the same nuh_layer_id as the layer when the VPS was removed. In this case, the SPS is incorrectly removed by the sub-bitstream extraction process. The use of inter-layer prediction is signaled by vps_independent_layer_flag. However, the vps_independent_layer_flag is removed when the VPS is removed. Therefore, when an SPS does not refer to a VPS, vps_independent_layer_flag[GeneralLayer vps_independent_layer_fla for the current layer, represented as [rIdx[nuh_layer_id]] g is inferred to be 1. SPS is inferred to be VP Does not reference S. In addition, vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] When is equal to 1 (e.g., when the VPS is not present), the decoder and / or HRD are currently It can be inferred that the layer / CLVS does not use inter-layer prediction. By adopting the above logic, the SPS can extract the VPS and corresponding parameters from the bitstream. to avoid extraction by the sub-bitstream extraction process even when The result is that the encoder and decoder can be constrained to include the appropriate nuh_layer_id for In addition, the coding efficiency is improved by using only the simulcast layer. This is enhanced by successfully removing unnecessary VPS from the bitstream, This reduces the processor, memory, and / or network overhead in both the encoder and decoder. This reduces the usage of network signaling resources.
[0024] Optionally, in any of the aforementioned aspects, another implementation of the aspect is a decoder comprising: It is further configured to perform the method of any of the aspects.
[0025] In one embodiment, the present disclosure includes an encoder that The bitstream contains the SPS referenced by the CLVS. An encoding means for encoding, the SPS comprising: When the CLVS nuh_layer_id is not present, the encoding is constrained to have a nuh_layer_id value equal to the CLVS nuh_layer_id value. and storage means for storing the bitstream for communication to a decoder. It is equipped with.
[0026] Some video coding systems code a video sequence into layers of pictures. Pictures in different layers have different characteristics. The coder can transmit different layers to the decoder depending on the constraints on the decoder side. To perform this function, the encoder compresses all layers into a single bitstream. On demand, the encoder can encode A sub-bitstream extraction process can be performed to remove irrelevant information. The result is an extracted bitmap that contains only the data in the layer required by the decoder. The description of how the layers are related is given in the video parameters. Simulcast layers can be included in a dataset (VPS). Simulcast layers refer to other layers. Simulcast layers are layers that are set to display without being When the sub-bitstream is transmitted to the You can safely remove the VPS as it is not needed to decode the trailer. Some variables in other parameter sets may refer to the VPS. Removing the VPS when the cast layer is transmitted can improve coding efficiency. Furthermore, failure to correctly identify the SPS may result in As a result, the SPS may be accidentally removed along with the VPS. This is because the SPS is missing. This can be problematic as in some cases the layers cannot be decoded correctly at the decoder. Examples of the present invention avoid errors when the VPS is removed during sub-bitstream extraction. Specifically, the sub-bitstream extraction process includes a mechanism for Remove NAL units based on r_ids. SPS is a layer for which CLVS does not use inter-layer prediction. has a nuh_layer_id equal to the nuh_layer_id of the CLVS referencing the SPS when A layer that does not use inter-layer prediction is a simulcast layer. Therefore, the SPS will have the same nuh_layer_id as the layer when the VPS was removed. In this case, the SPS is incorrectly removed by the sub-bitstream extraction process. The use of inter-layer prediction is signaled by vps_independent_layer_flag. However, the vps_independent_layer_flag is removed when the VPS is removed. Therefore, when an SPS does not refer to a VPS, vps_independent_layer_flag[GeneralLayer vps_independent_layer_fla for the current layer, represented as [rIdx[nuh_layer_id]] g is inferred to be 1. SPS is inferred to be VP Does not reference S. In addition, vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] When is equal to 1 (e.g., when the VPS is not present), the decoder and / or HRD are currently It can be inferred that the layer / CLVS does not use inter-layer prediction. By adopting the above logic, the SPS can extract the VPS and corresponding parameters from the bitstream. to avoid extraction by the sub-bitstream extraction process even when The result is that the encoder and decoder can be constrained to include the appropriate nuh_layer_id for In addition, the coding efficiency is improved by using only the simulcast layer. This is enhanced by successfully removing unnecessary VPS from the bitstream, This reduces the processor, memory, and / or network overhead in both the encoder and decoder. This reduces the usage of network signaling resources.
[0027] Optionally, in any of the aforementioned aspects, another implementation of the aspect is further characterized in that the encoder comprises: It is further configured to perform the method of any of the above aspects.
[0028] For clarity, any one of the above-mentioned embodiments may be combined with any other of the above-mentioned embodiments. may be combined with any one or more of the above to form new embodiments within the scope of the present disclosure. It may be formed.
[0029] These and other features will become more apparent from the following detailed description taken in conjunction with the accompanying drawings and claims. will be understood.
[0030] In order that this disclosure may be more fully understood, reference may be made to the accompanying drawings, in which like numerals represent like parts and in which: For further details, reference is made to the following brief description. [Brief description of the drawings]
[0031] [Figure 1] 1 is a flowchart of an exemplary method for coding a video signal. [Diagram 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding. [Diagram 3] 1 is a schematic diagram illustrating an example video encoder. [Figure 4] 1 is a schematic diagram illustrating an exemplary video decoder. [Diagram 5] 1 is a schematic diagram illustrating an exemplary hypothetical reference decoder (HRD). [Figure 6] 1 is a schematic diagram illustrating an example multi-layer video sequence configured for inter-layer prediction. [Figure 7] FIG. 2 is a schematic diagram illustrating an exemplary bitstream. [Figure 8] 1 is a schematic diagram of an example video coding device. [Figure 9] 1 is a flowchart of an example method for encoding a multi-layer video sequence into a bitstream to support preserving a sequence parameter set (SPS) upon extraction of a sub-bitstream for a simulcast layer. [Figure 10] 1 is a flow chart of an example method for decoding a video sequence from a bitstream including a simulcast layer extracted from a multi-layer bitstream in which SPS is preserved during sub-bitstream extraction. [Figure 11]1 is a schematic diagram of an example system for coding a multi-layered video sequence into a bitstream to support preserving SPS during sub-bitstream extraction for a simulcast layer. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0032] First, exemplary implementations of one or more embodiments are presented below, but the disclosed The systems and / or methods described herein are not limited to any currently known or existing technology. It should be understood that the present disclosure may be implemented using a variety of techniques. The following examples, including the exemplary designs and implementations illustrated and described herein, are provided to The present disclosure should in no way be limited to the illustrative implementations, diagrams, and techniques shown, but includes all equivalents. It may be amended within the scope of the appended claims together with its scope.
[0033] The following terms are defined as follows, unless used in the present specification to the contrary: Specifically, the following definitions are intended to further clarify the present disclosure. However, terms may be explained differently in different contexts. Thus, the following definitions should be considered supplementary to the terms provided herein for such terms. The above definitions should not be construed as limiting any other definitions of the descriptions provided herein.
[0034] A bitstream is a video data stream that is compressed for transmission between an encoder and a decoder. An encoder uses an encoding process to generate a sequence of bits that contains the A device configured to compress video data into a bitstream using a decoded video signal. The decoder uses a decoding process to convert the video data into a bitstream for display. A picture is a frame or a subframe of a picture. The arrays of luma samples and / or chroma samples that make up the filter. The encoded or decoded pictures are shown in order to clarify the explanation. A coded picture may be referred to as the current picture. A video codec with a specific value of the NAL unit header layer identifier (nuh_layer_id) in The Video Streaming Layer (VCL) and Network Abstraction Layer (NAL) units are included in the entire picture. A coded representation of a picture that contains a coding tree unit (CTU). A decoded picture is a coded picture that is generated by applying a decoding process to a coded picture. A NAL unit is a picture generated by emulation if desired. A raw byte sequence payload that is an indication of the type of data, interspersed with anti-malware bytes. A VCL NAL unit is a syntax structure that contains data in the form of a picture. A NAL unit that is coded to contain video data, such as a coded slice. Non-VCL NAL units are used to decode video data, perform conformance checking, and Syntax and / or parameters that support the execution of a task or other action. A NAL unit that contains non-video data. A layer is indicated by a layer Id (identifier). Specified characteristics such as common resolution, frame rate, image size, etc. A set of VCL NAL units that share a common VCL NAL unit (such as a The NAL unit header layer identifier (nuh_layer_id) identifies the layer that contains the NAL unit. This is a syntax element that specifies the identifier.
[0035] A hypothetical reference decoder (HRD) decodes the bitstream generated by the encoding process. On the encoder to check the variability of the stream and verify compliance with the specified constraints The bitstream conformance test is performed on the encoded If the bitstream complies with a standard such as Versatile Video Coding (VVC), The Video Parameter Set (VPS) is a test to determine whether a video A sequence parameter set (S PS) is applied to zero or more entire coded layer video sequences (CLVS). The nuh_layer_id is a syntax structure that contains syntax elements that are used to 4 is a sequence of coded pictures having SPS video parameter set identification. The child (sps_video_parameter_set_id) is a syntax element that specifies the identifier (ID) of the VPS reference by the SPS. The general layer index (GeneralLayerIdx[i]) is the layer index of the corresponding layer i. It is a derived variable that specifies the index, hence it has a layer ID of nuh_layer_id. The current layer has the index specified by GeneralLayerIdx[nuh_layer_id]. The current layer index corresponds to the layer being encoded or decoded. The layer index to be used. VPS independent layer flag (vps_independent_layer_flag[i]) is a syntax element that specifies whether the corresponding layer i uses inter-layer prediction. Therefore, vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is Specifies whether the current layer uses inter-layer prediction. Inter-layer prediction is a method for predicting the quality of a layer by using different layers. The current layer's current reference picture is based on a reference picture from the previous layer (e.g., the same access unit). Access Unit is a mechanism for coding blocks of sample values of a picture. An AU is a set of coded time units in different layers that are all associated with the same output time. A set of pictures that are encoded in the VPS parameter set. is a syntax element that provides an ID for the VPS for reference by other syntax elements / constructs. A coded video sequence is a sequence of one or more coded A decoded video sequence is a set of one or more decoded pictures. This is a set of pre-populated pictures.
[0036] The acronyms used in this specification are: Access Unit (AU), Coding Tree Lock (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coding Coded Layered Video Sequence (CLVS), Coded Layered Video Sequence Start (CLVSS), Coded Video Sequence (CVS), Coded Video Sequence Start (CVSS), Joint Video Experts Team (JVET), Hypothetical Reference Decoder (HRD), Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Output Layer Set (OLS), Operation Point (OP), Picture Order Count (POC), Random Access Point (RAP), Low Byte Sequence Payload (RBSP), Sequence Payload Parameter Set (SPS), Video Parameter Set (VPS), Versatile Video Codec The first is VVC.
[0037] Many video compression techniques reduce the size of video files with minimal data loss. For example, video compression techniques can be employed to reduce the spatial (e.g., picture
[0023] In the video sequence, a prediction method is provided for predicting the image quality of a video sequence by performing intra-picture (intra-picture) prediction and / or temporal (e.g., inter-picture) prediction. The method may include reducing or eliminating data redundancy in a block-based For video coding, a video slice (e.g., a video picture, or a video A video block (part of a video picture) may be partitioned into several video blocks, which block, coding tree block (CTB), coding tree unit (CTU), They are sometimes called Piping Units (CUs), and / or Coding Nodes. Video blocks in an intra-coded (I) slice of a picture are The pixel is coded using spatial prediction with respect to reference samples in neighboring blocks. Video in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a channel A block can be spatially predicted with respect to reference samples in neighboring blocks in the same picture, or Coding is performed by employing temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be , which may be referred to as a reference frame and / or a reference image. The result is a prediction block that represents the original image block. It represents the pixel difference between the inter-coded block and the predicted block. A block is made up of a motion vector pointing to a block of reference samples that form the prediction block, and and the residual data indicating the difference between the coded block and the predicted block. Intra-coded blocks are coded according to the intra-coding mode and For further compression, the residual data is encoded according to the pixel These results in residual transform coefficients, which can be converted from the vector domain to the transform domain. The quantized transform coefficients may be first arranged in a two-dimensional array. The processed transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. Video compression can be applied to achieve even greater compression. Shrinkage techniques are described in more detail below.
[0038] To ensure that the encoded video can be accurately decoded, the video is It is encoded and decoded according to the corresponding video coding standard. The standard is the International Telecommunications Union (ITU) Standardization Sector (ITU-T) H.261, the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU- T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H. Advanced Video Coding (AVC), also known as H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding, also known as ITU-T H.265 or MPEG-H Part 2 AVC includes Scalable Video Coding (SVC), Multiview Video Coding (MVC), and 3D Video Coding (DVC). Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and Three-Dimensional (3D)A HEVC includes extensions such as Scalable HEVC (SHVC) and Multiview HEVC. It includes extensions such as Multi-Voltage-Hevc (MV-HEVC), and 3D HEVC (3D-HEVC). The Japan Video Experts Team (JVET) is a leading provider of Versatile Video Coding (VVC) The video coding standard called VVC has been developed. Yes, it is included in the Working Draft (WD).
[0039] Some video coding systems code a video sequence into layers of pictures. Pictures in different layers have different characteristics. The coder can transmit different layers to the decoder depending on the constraints on the decoder side. To perform this function, the encoder compresses all layers into a single bitstream. On demand, the encoder can encode A sub-bitstream extraction process can be performed to remove irrelevant information. The result is an extracted bitmap that contains only the data in the layer required by the decoder. The description of how the layers are related is given in the video parameters. Simulcast layers can be included in a dataset (VPS). Simulcast layers refer to other layers. Simulcast layers are layers that are set to display without being When the sub-bitstream is transmitted to the You can safely remove the VPS as it is not needed to decode the trailer. Some variables in other parameter sets may refer to the VPS. Removing the VPS when the cast layer is transmitted can improve coding efficiency. , which may result in errors. Furthermore, Failure to properly identify the SPS may result in the incorrect removal of the SPS along with the VPS. This means that if the SPS is missing, the layer cannot be decoded correctly at the decoder. This can be problematic.
[0040] Disclosed herein is a method for extracting a sub-bitstream in which the VPS is removed during extraction of the sub-bitstream. This is a mechanism to avoid errors when transferring sub-bitstreams. The extraction process uses network abstraction based on NAL unit layer identifiers (nuh_layer_ids). Remove layer (NAL) units. SPS is included in layers where CLVS does not use inter-layer prediction. When the SPS is generated, the nuh_layer_i d. Layers that do not use inter-layer prediction are constrained to have nuh_layer_id equal to , the simulcast layer. Therefore, the SPS is the layer when the VPS is removed. In this way, the SPS can The use of inter-layer prediction is enabled by the VPS independent layer flag (vps_ However, vps_independent _layer_flag is removed when the VPS is removed. Therefore, when the SPS does not reference the VPS, , currently represented as vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] vps_independent_layer_flag for this layer is inferred to be 1. SPS is an SPS VPS When the identifier (sps_video_parameter_set_id) is set to 0, it does not reference the VPS. When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1 (for example, For example, when no VPS is present), the decoder and / or the hypothetical reference decoder (HRD) We can deduce that CLVS / Earth / CLVS does not use inter-layer prediction. By adopting SPS, the VPS and corresponding parameters are removed from the bitstream. To avoid extraction by the sub-bitstream extraction process even when The nuh_layer_id can be constrained to include the appropriate nuh_layer_id. In addition, the coding efficiency is improved for video that contains only the simulcast layer. This is enhanced by successfully removing unnecessary VPS from the hotstream, The processor, memory, and / or network in both the encoder and decoder Reduces signaling resource usage.
[0041] FIG. 1 is a flow chart of an exemplary operational method 100 of coding a video signal. Specifically, the video signal is encoded in an encoder. The process compresses the video signal by using various mechanisms and creates a video file. Reduce file size. Smaller file sizes result in lower associated bandwidth overhead. This allows the compressed video file to be transmitted to the user while reducing the The decoder decodes the compressed video file and converts it to the original video for display to the end user. The decoding process generally involves the decoder reconstructing a consistent video signal. Mirroring the encoding process to allow the audio signal to be reconstructed do.
[0042] In step 101, a video signal is input to an encoder. For example, the video signal may be It may be an uncompressed video file stored in memory. The file is captured by a video capture device, such as a video camera, and Video files can be encoded to support live streaming of , may contain both audio and video components. A video frame contains a sequence of image frames which, when viewed in sequence, give the visual impression of motion. A beam is made up of light, referred to herein as luma components (or luma samples), and chrominance. A pixel contains a representation of a color, called a color component (or color sample). In an example, the frames may also include depth values to support three-dimensional displays.
[0043] In step 103, the video is segmented into a number of blocks. Subdividing the pixels in each frame into square and / or rectangular blocks for compression For example, High Efficiency Video Coding (HEVC) (combined with H.265 and MPEG-H Part 2) In a multi-frame (also known as a multi-frame), a frame is first scaled to a predefined size (e.g., The image is divided into coding tree units (CTUs), which are blocks of 64 pixels x 64 pixels. A CTU contains both luma and chroma samples. The coding tree is , splitting the CTU into several blocks and then supporting further encoding It may be employed to recursively subdivide the block until a configuration is achieved. The luma component of the image may be subdivided until each block contains relatively homogenous illumination values. The chroma components of the frame are then subdivided until each block contains relatively homogenous color values. Therefore, the segmentation mechanism may vary depending on the content of the video frame. do.
[0044] In step 105, various compression mechanisms are employed to compress the image segmented in step 103. For example, inter-prediction and / or intra-prediction may be employed. Inter prediction is based on the tendency for objects in a common scene to appear in consecutive frames. It is designed to take advantage of the fact that the object has a certain orientation within its reference frame. Blocks depicting an object need not be repeated in adjacent frames. Specifically, objects such as tables remain in a consistent position across multiple frames. Thus, once a table is written, adjacent frames can be used as reference frames. Matching objects across multiple frames A pattern matching mechanism can be employed to detect object movement and color. Objects that move across multiple frames can be expressed by moving the camera, etc. As a specific example, a video may show a car moving across the screen over several frames. A motion vector can be used to describe such a movement. A vector is a vector that goes from the coordinates of an object in a frame to the coordinates of an object in a reference frame. is a 2-dimensional vector that provides an offset to the current Image blocks in a frame are offset from their corresponding blocks in a reference frame. The motion vectors can then be encoded as a set of
[0045] Intra prediction encodes blocks within a common frame. Take advantage of the fact that ma and chroma components tend to cluster within a frame. For example, a patch of green in a tree segment is positioned adjacent to a similar patch of green. Intra prediction tends to use multiple directional prediction modes (for example, 33 types in HEVC). Use the directional modes (such as 3D), Planar, and Direct Current (DC) modes. These directional modes are indicates that the block is similar / identical to the samples in the neighboring block in the corresponding direction. Planar mode is when a series of blocks along a row / column (e.g. a plane) is split into two parts, one near the end of the row and the other near the edge of the row. It is shown that the pixel values can be interpolated based on neighboring blocks. A relatively constant gradient is used to show smooth transitions of light / color across rows / columns. DC mode is used for boundary smoothing and the blocks are related to the angular direction of the directional prediction mode. Similar / same as the average value associated with the samples of all neighboring blocks Therefore, the intra-predicted blocks are not based on the actual values but on the various related prediction models. In addition, an image block can be represented as an inter-prediction block. Image blocks can be represented as motion vector values rather than actual values. However, the prediction block may not accurately represent the image block in some cases. Any differences are stored in the residual block. To further compress the file, A transform may be applied to the residual block.
[0046] Various filtering techniques may be applied in step 107. In HEVC, the filters are: It is applied according to the in-loop filtering scheme described above. The prediction of the block vector can result in blocky images at the decoder. The prediction scheme of the source encodes a block and then uses it later as a reference block. The in-loop filtering scheme may be , noise suppression filter, deblocking filter, adaptive loop filter, and sample Iteratively apply adaptive offset (SAO) filters to a block / frame. These filters are , reducing such blocking artifacts and thus improving the quality of the encoded file. Moreover, these filters can accurately reconstruct the The artifacts are reduced by reducing the number of reconstructed images based on the reconstructed reference block. It is less likely to introduce further artifacts in subsequent blocks that are encoded. become.
[0047] After the video signal is segmented, compressed, and filtered, the resulting digital The data is encoded into a bitstream in step 109. and supports proper video signal reconstruction at the decoder. The present invention includes any signaling data that is desirable for the purpose of determining whether a particular Various circuits for sending segment data, prediction data, residual blocks, and coding instructions to the decoder The bitstream may contain flags for transmission to a decoder on demand. The bitstream can also be broadcast to multiple decoders. The creation of the bitstream may be , is an iterative process. Thus, steps 101, 103, 105, 107, and 109 are The frame and block of the image may be sequentially and / or simultaneously performed. The order shown is presented for clarity and ease of explanation and is not intended to be used with video code. The present invention is not intended to limit the coding process to any particular order.
[0048] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme to The decoder converts the stream of data into the corresponding syntax and video data. In step 111, the syntax data from the bitstream is used to determine the The partitioning should be consistent with the block partitioning results in step 103. Next, entropy encoding / decoding as employed in step 111 is performed. The encoder determines the spatial location of values in the input image. The compression process includes selecting a block partitioning scheme from several possible choices based on the In signaling the correct choice, many bins are As used herein, a bin is a set of two variables that are treated as variables. Entropy coding is a method for coding a number of bits at a time (e.g., a bit value that can vary depending on the context). The encoder discards any options that are clearly infeasible for a particular case. , leaving a set of acceptable options. Then, for each acceptable option, Each option is assigned a codeword. The length of the codeword depends on the allowable options. Based on the number of options (e.g., one bin for 2 options, two for 3-4 options, etc.) The encoder then encodes the codeword for the selected option. In this scheme, the codeword is the potential A small set of acceptable options, as opposed to a unique selection from a large set. of the codeword, since it is desirable to uniquely indicate a selection from the subset. The decoder then selects the acceptable options in a similar manner to the encoder. The set of allowable options is then used to decode the selection. By determining The selected selection can be determined.
[0049] In step 113, the decoder performs block decoding. The decoder then performs an inverse transform to generate the residual block. The residual block and the corresponding prediction block are used to reconstruct an image block according to the The predicted block is an intra-predicted block as generated by the encoder in step 105. The reconstructed image block may include both inter-predicted and inter-predicted blocks. , within a frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax for step 113 is also as described above. can be signaled in the bitstream via entropy coding.
[0050] In step 115, filtering is repeated in the encoder in a manner similar to step 107. It is performed on frames of the constructed video signal. For example, noise suppression filters, deblocking filters, etc. The blocking filter, adaptive loop filter, and SAO filter are used to can be applied to the frame to remove artifacts. After the frame has been filtered, The video signal is output to a display in step 117 for viewing by an end user. obtain.
[0051] FIG. 2 illustrates an exemplary coding and decoding scheme for video coding. 2 is a schematic diagram of a codec system 200. Specifically, the codec system 200 operates in accordance with a first method. 00 implementation. It is generalized to depict components employed in both coders and decoders. The codec system 200 is described with respect to steps 101 and 103 of the method of operation 100. 2, a video signal is received and segmented as shown in FIG. 2, resulting in a segmented video signal 201. The codec system 200 then repeats steps 105, 107, and 108 of the method 100. When operating as an encoder as described with respect to FIG. 9, the segmented video signal 2 The codec system 200 compresses the encoded bitstream. When operating as a datum, steps 111, 113, 115, and 117 of the method of operation 100 are described. A codec system generates an output video signal from a bitstream as specified in 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component a compensation component 219, a motion estimation component 221, a scaling and inverse transformation component components 229, filter control analysis components 227, in-loop filter components 225 , the decoded picture buffer component 223, and the header formatting It includes a coding and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are connected as shown. In FIG. 2, the black lines represent the ends. The dashed lines indicate the movement of data to be coded / decoded, while the dashed lines indicate the movement of other components. The components of the codec system 200 are all The decoder may be a component of the codec system 200. For example, the decoder may include an intra-picture prediction component. 217, motion compensation component 219, scaling and inverse transform component 229, loop The filter component 225 and the decoded picture buffer component 2 23. These components are now described.
[0052] The segmented video signal 201 is divided into several blocks of pixels by a coding tree. A coding tree is a captured video sequence partitioned into blocks. Various division modes are used to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. A lock may be referred to as a node of the coding tree. A larger parent node has a higher The number of times a node is subdivided depends on the number of nodes in the coding tree. The partitioned blocks may in some cases be divided into coding units. For example, a CU may be included in a CT containing a luma block, a red-difference chroma (Cr) block, and a blue-difference chroma (Cb) block The partitioning mode may be a subpart of U. The partitioning mode may vary depending on the partitioning mode adopted for the node. A binary tree ( The segmented video signal 201 may include a ternary tree (BT), a ternary tree (TT), and a quad tree (QT). A general coder control component 211 and a transform scaling and quantization component 2 13, the intra-picture estimation component 215, the filter control analysis component 227, and the motion The result is forwarded to the estimation component 221.
[0053] The general coder control component 211 controls the bitstream according to application constraints. and configured to make decisions relating to coding of images of a video sequence into a frame. For example, the general coder control component 211 may provide a trade-off between bit rate / bit ratio and reconstruction quality. Such decisions are based on storage space / bandwidth optimization. This can be done based on bandwidth availability and image resolution requirements. Port 211 also mitigates buffer underrun and overrun problems. To manage these issues, we use a The general coder control component 211 controls the segmentation, prediction, and frame rate of the other components. For example, the general coder control component 211 manages the filtering of dynamically increase compression complexity to increase bandwidth usage, or reduce resolution and bandwidth It is acceptable to reduce the compression complexity in order to reduce bandwidth usage. The control component 211 controls the other components of the codec system 200 to generate the video. The general coder control component balances the quality of the signal reconstruction with the bit rate issue. The component 211 creates control data that controls the operation of other components. The data is also forwarded to the Header Formatting and CABAC component 231 for bitstreaming. It is encoded in the stream and signals parameters for decoding at the decoder. Ring it.
[0054] The segmented video signal 201 also includes a motion estimation component 221 for inter prediction. and transmitted to the motion compensation component 219. A slice may be divided into multiple video blocks. and a motion compensation component 219 for performing temporal prediction in one or more reference frames. Inter-predictive coding of a received video block for one or more blocks of The codec system 200 performs multiple coding passes to For example, an appropriate coding mode for each block of video data may be selected.
[0055] The motion estimation component 221 and the motion compensation component 219 are highly integrated. Although these are illustrated separately for conceptual purposes, the motion estimation component 221 The motion estimation is performed by generating motion vectors, which estimate the motion for a video block. A motion vector is, for example, a coding vector for a predictive block. The predicted block may indicate the displacement of the coded object in terms of pixel differences. A predicted block is a block that is known to be a good match for the block to be matched. Such pixel differences can be calculated using the sum of absolute differences (SAD) , Sum of Squared Differences (SSD), or other difference metrics. A number of pre-coded objects, including the Loading Tree Block (CTB) and CU, are included. For example, a CTU may be split into a CTB, which may then be split into a CB to be included in a CU. A CU may be divided into a prediction unit (PU) that contains prediction data and / or a change to the CU. The motion estimation component can be encoded as a transform unit (TU) that contains the transformed residual data. Component 221 uses rate-distortion analysis as part of the rate-distortion optimization process. Thus, the motion vectors, PU and TU are generated. For example, the motion estimation component 22 1. Multiple reference blocks, multiple motion vectors, etc. for the current block / frame The method may determine and select the reference block, motion vector, etc. that has the best rate-distortion performance. The best rate-distortion performance is determined by the quality of the video reconstruction (e.g., the amount of data lost due to compression) and Balance both coding efficiency (i.e., size of the final encoding) .
[0056] In some examples, the codec system 200 may include a decoded picture buffer controller. The values for the sub-integer pixel positions of the reference picture stored in the component 223 are For example, the video codec system 200 may calculate a 1 / 4 pixel position of the reference picture. The values of the fractional pixel positions, 1 / 8 pixel positions, or other fractional pixel positions may be interpolated. The motion estimation component 221 performs motion search for full pixel and fractional pixel positions. The motion estimation component may perform a search and output motion vectors with fractional pixel accuracy. The counter 221 calculates the image by comparing the position of the PU with the position of the prediction block of the reference picture. Calculate the motion vectors for the PUs of the video blocks in the inter-coded slices. The motion estimation component 221 outputs the encoding and Header formatting for the call and motion and motion data to CABAC component 231 The calculated motion vector is output as the data.
[0057] The motion compensation performed by the motion compensation component 219 is 21 to fetch or generate a prediction block based on the motion vector determined by Again, the motion estimation component 221 and the motion compensation component The components 219 may be functionally integrated in some examples. Upon receiving the motion vectors for the PUs of the block, the motion compensation component 219 The residual video block may then be coded by identifying the location of the predictive block to which the residual video block points. The pixel values of the predicted block are subtracted from the pixel values of the current video block being mapped. , pixel difference values. In general, the motion estimation component The motion estimation component 221 performs motion estimation on the luma components, and the motion compensation component 219 performs motion estimation on the chroma components. For both the luminance and luma components, the motion vectors calculated based on the luma component are used. The prediction block and the residual block are subjected to the transform scaling and quantization component 213. will be forwarded to.
[0058] The segmented video signal 201 also includes an intra-picture estimation component 215 and a picture The prediction component 217 is sent to the motion estimation component 221 and the motion compensation component 222. As well as the intra picture estimation component 215 and the intra picture prediction component 219, Components 217 may be highly integrated, but are illustrated separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 are described above. As shown in FIG. 1, the motion estimation component 221 and the motion compensation component 222 are As an alternative to the inter prediction performed by component 219, In particular, the intra picture estimation component 215 determines the intra prediction mode to use to encode the current block. In some examples, the intra picture estimation component 215 may include a plurality of tested intra predictions. Select an appropriate intra-prediction mode to encode the current block from the mode The selected intra prediction mode is then added to the header format for encoding. The call is forwarded to the call processing and CABAC component 231.
[0059] For example, the intra picture estimation component 215 may be implemented using various tested intra prediction modes. Rate-distortion analysis is used to calculate rate-distortion values for each mode, and the Select the intra prediction mode with the best rate-distortion performance. In general, encoded blocks are encoded to generate encoded blocks. the amount of distortion (or error) between the original unencoded block and the encoded Determine the bit rate (e.g., number of bits) used to generate the completed block. The intra picture estimation component 215 determines which intra prediction mode is used for the block. for various encoded blocks to determine which gives the best rate-distortion value. The ratio is calculated from the distortion and rate. In addition, the intra-picture estimation component 21 5 uses Depth Modeling Mode (DMM) based on Rate-Distortion Optimization (RDO) to generate depth maps. The image may be configured to code a depth block of
[0060] The intra picture prediction component 217, when implemented in an encoder, performs intra picture estimation. The prediction block is generated based on the selected intra-prediction mode determined by the component 215. When implemented in a decoder, it generates a residual block from the The residual block may be read from the prediction The residual block contains the difference in values between the measured block and the original block. The residual block is then The image is forwarded to the filtering and quantization component 213. 15 and the intra-picture prediction component 217 operate on both the luma and chroma components. obtain.
[0061] The transform scaling and quantization component 213 further compresses the residual block. The transform scaling and quantization component 213 is configured to perform discrete cosine The residual block is a transform, such as the discrete cosine transform (DCT), the discrete sine transform (DST), or a conceptually similar transform. The wavelet transform is applied to the block to generate a video block containing residual transform coefficient values. Integer transforms, subband transforms, or other types of transforms could also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 converts the transformed It is also configured to scale the residual information. Such scaling may be used to scale the residual information for different peripheries. Applying a scale factor to the residual information so that the wavenumber information is quantized to different granularities This can affect the final visual quality of the reconstructed video. The filtering and quantization component 213 applies a conversion factor to further reduce the bit rate. The quantization process is also configured to quantize the quantization numbers associated with some or all of the coefficients. The degree of quantization can be adjusted by adjusting the quantization parameter. In some examples, the transformation scaling and quantization components The processor 213 may then perform a scan of the matrix containing the quantized transform coefficients. The conversion factor is forwarded to the header formatting and CABAC component 231, which converts the bit Encoded into the stream.
[0062] The scaling and inverse transform component 229 performs transformation to support motion estimation. Apply the inverse operation of the scaling and quantization component 213. The inverse transform component 229 may, for example, convert a block to a prediction block for another current block. Reconstruct the residual block in the pixel domain for later use as a possible reference block. Apply inverse scaling, transformation, and / or quantization to the motion estimation component. The motion estimation component 221 and / or the motion compensation component 219 estimate the motion of subsequent blocks / frames. By adding the residual block back to the corresponding predicted block for use in determining A filter is applied to the reconstructed reference block, which It reduces artifacts created during scaling, quantization, and conversion. Such artifacts would otherwise occur when subsequent blocks are predicted inaccurately. It may also cause predictions (and create additional artifacts).
[0063] The filter control analysis component 227 and the in-loop filter component 225 Apply a filter to the difference block and / or the reconstructed image block. For example, The transformed residual block from the scaler and inverse transform component 229 is The corresponding prediction block from the prediction component 217 and / or the motion compensation component 219 The filter can then be combined with the block to reconstruct the original image block. In some examples, the filter may be applied to the residual block instead. As with the other components in Figure 2, the filter control analysis component The in-loop filter component 225 and the in-loop filter component 227 are highly integrated and The reconstructed reference block may be implemented in a 3D representation, but is depicted separately for conceptual purposes. The filters applied to a particular spatial region are determined by how such filters are applied to the Contains several parameters to adjust how the filter is applied. The filter 227 analyzes the reconstructed reference block to determine if such a filter should be applied. Such data is stored in the encoding. Filter control data for header formatting and CABAC component 2 31. The in-loop filter component 225 determines based on the filter control data Then, we apply such a filter. The filter can be a deblocking filter, a noise suppression filter, etc. Such filters may include a filter, a SAO filter, and an adaptive loop filter. can be in the spatial / pixel domain (e.g., on a reconstructed pixel block) or It can be applied in the frequency domain.
[0064] When acting as an encoder, the filtered reconstructed image block, the residual The block and / or the predicted block are subsequently estimated by motion estimation as described above. The decoded picture is stored in the decoded picture buffer component 223 for use. When acting as a decoded picture buffer, the decoded picture buffer component 223 stores the reconstructed The filtered and reconstructed blocks are stored and displayed as part of the output video signal. The decoded picture buffer component 223 transfers the predicted block, residual block, and Any memory capable of storing the residual blocks and / or reconstructed image blocks. The device may be a remote device.
[0065] The header formatting and CABAC component 231 is Receives data from various components of the The header is then encoded into a coded bitstream for transmission. The DaFormatting and CABAC components 231 are general control data and filters Generate various headers to encode the control data, such as the control data, and , prediction data including intra prediction data and motion data, as well as quantized transform coefficients Any residual data in the form of data is encoded into the bitstream. The bitstream is the desired bitstream for a decoder to reconstruct the original segmented video signal 201. Such information includes all the necessary information, such as the intra-prediction mode index table (also called codeword mapping table), encoding for various blocks Defining the rendering context, indicating the most likely intra-prediction mode, and indicating partition information Such data can be encoded using entropy coding. For example, the information can be encoded using Context Adaptive Variable Length Coding (CAVLC). ), CABAC, Syntax-Based Context-Adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioned Entropy (PIPE) coding or another entropy coding technique is used. According to the entropy coding, the code can be encoded by using The encoded bitstream is then transmitted to another device (e.g., a video decoder) or , or may be archived for later transmission or removal.
[0066] 3 is a block diagram illustrating an example video encoder 300. 300 implements the encoding functionality and / or operates the codec system 200. may be employed to implement steps 101, 103, 105, 107, and / or 109 of method 100. The encoder 300 segments an input video signal, resulting in a segmented video signal 301. 201. The segmented video signal 202 is then A video signal 301 is compressed into a bitstream by components of the encoder 300. The data is then encoded.
[0067] Specifically, the partitioned video signal 301 is subjected to an intra-picture prediction control for intra prediction. The intra-picture prediction component 317 performs intra-picture estimation. 215 and intra-picture prediction component 217. The segmented video signal 301 may also include a decoded picture buffer component. The motion compensation component 321 transfers the motion to the motion compensation component 322 for inter prediction based on a reference block in the motion compensation component 323. The motion compensation component 321 is a combination of the motion estimation component 221 and the motion compensation component 322. The intra-picture prediction component may be substantially similar to the intra-picture prediction component 219. The prediction block and the residual block from the motion compensation component 321 are The block is forwarded to the transform and quantization component 313 for transformation and quantization. The transform and quantization component 313 is a transform scaling and quantization component. The transformed and quantized residual block may be substantially similar to that of the input 213. The block and the corresponding prediction block (along with associated control data) are stored in the bitstream. The resulting image is forwarded to the entropy coding component 331 for coding into a sigma-based image. The entropy coding component 331 handles header formatting and CABAC It may be substantially similar to component 231.
[0068] The transformed and quantized residual block and / or the corresponding prediction block are The motion compensation component 321 converts the input image data into a reference block and converts and quantizes the converted image data into a reference block. The inverse transform and quantization component 329 forwards the inverse transform and quantization component 313. The scaling and quantization component 329 is substantially the same as the scaling and inverse transform component 229. The in-loop filter in the in-loop filter component 325 may be similar. The data is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 is connected to the filter control analysis component 227 and the loop The in-loop filter may be substantially similar to the in-loop filter component 225. The in-loop filter component 325 is described with respect to the in-loop filter component 225. The filtered block is then subjected to a motion compensation process. Decoded picture buffer for use as a reference block by component 321 The decoded picture is stored in the decoded picture buffer component 323. , as being substantially similar to the decoded picture buffer component 223 good.
[0069] FIG. 4 is a block diagram illustrating an exemplary video decoder 400. implements the decoding functionality of the codec system 200 and / or the operating method 100 The decoder 4 may be employed to implement steps 111, 113, 115, and / or 117 of the 00 receives the bitstream from, for example, the encoder 300 and prepares the A reconstructed output video signal is generated based on the bitstream to obtain a reconstructed output video signal.
[0070] The bitstream is received by an entropy decoding component 433. The Entropy Decoding Component 433 supports CAVLC, CABAC, SBAC, and PIPE codes. Entropy decoding, such as decoding or other entropy coding techniques For example, the entropy decoding component is configured to implement a Component 433 is additional data encoded as a codeword in the bitstream. Header information may be employed to provide context for interpreting the decoded data. The information includes general control data, filter control data, classification information, motion data, prediction data, and any other necessary steps for decoding the video signal, such as the quantized transform coefficients from the residual block. The quantized transform coefficients are then inverse transformed for reconstruction into the residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantize component 329.
[0071] The reconstructed residual block and / or the predicted block are then processed based on an intra prediction operation. The image blocks are then forwarded to the intra-picture prediction component 417 for reconstruction into image blocks. The intra picture prediction component 417 is a combination of the intra picture estimation component 215 and the intra picture prediction. It may be similar to the intra-picture prediction component 217. Component 417 employs a prediction mode to identify the location of a reference block within a frame. , and apply the residual block to the result to reconstruct the intra-predicted image block. Only intra-predicted image blocks and / or residual blocks and corresponding inter-predicted The data is stored in the decoded picture buffer component 223 and the in-loop frame buffer component 224, respectively. The in-loop filter component may be substantially similar to the filter component 225. The decoded picture is then transferred to the decoded picture buffer component 423 via the decoded picture buffer component 425. The in-loop filter component 425 filters the reconstructed image block, the residual block, and and / or predictive blocks, and such information is added to the decoded picture. The decoded picture is stored in the buffer component 423. The reconstructed image block from the input 423 is processed by the motion compensation component 4 for inter prediction. 21. The motion compensation component 421 is connected to the motion estimation component 221 and / or or motion compensation component 219. Specifically, The motion compensation component 421 calculates the motion from the reference block to generate a prediction block. The vector is taken and the residual block is applied to the result to reconstruct the image block. The resulting reconstructed block also includes an in-loop filter component 425. The decoded picture may be forwarded to the decoded picture buffer component 423 via the The picture buffer component 423 can be reconfigured within a frame via partition information. , and continues to store additional reconstructed image blocks. Such frames are The sequence may be placed into a sequence that is then displayed as a reconstructed output video signal. The image is output to Ray.
[0072] 5 is a schematic diagram illustrating an example HRD 500. The HRD 500 includes a codec system 200 and and / or may be employed in an encoder such as encoder 300. The bit stream created in step 109 is then input to a decoder 400 or the like. The bit may be checked before being forwarded to the decoder. The stream is continuously transferred through the HRD500 as the bitstream is encoded. It is possible that a portion of the bitstream fails to conform to the associated constraints. If such a failure occurs, the HRD500 will notify the encoder of such a failure and tell the encoder to use a different mechanism. This allows the algorithm to re-encode the corresponding section of the bitstream. do.
[0073] The HRD 500 includes a virtual stream scheduler (HSS) 541. The HSS 541 is a virtual distribution mechanism. The virtual delivery mechanism is a component that is configured to run the HR Regarding the timing and data flow of the bit stream 551 input to the D500, Used to check the conformance of a stream or a decoder. For example, HSS541 receives a bitstream 551 output from the encoder, and In a specific example, HSS541 may manage the conformance testing process. It controls the rate at which pre-filled pictures move through the HRD 500 and the bitstream 551 It can be verified that it does not contain non-conforming data.
[0074] The HSS 541 can transfer the bitstream 551 to the CPB 543 at a predefined rate. The HRD 500 may manage data in a decoding unit (DU) 553. The DU 553 may , an access unit (AU) or a subset of AUs and associated non-video code A Video Streaming Layer (VCL) and a Network Abstraction Layer (NAL) unit are called AUs. AUs contain one or more pictures associated with an output time. For example, an AU may be a single A single picture in a one-layer bitstream and a multi-layer bitstream Each picture in an AU corresponds to a corresponding VCL NAL unit. Thus, the DU 553 may be divided into slices contained in one or more pictures. , one or more slices of a picture, or a combination thereof. The parameters used to decode a picture and / or slice are non-V Therefore, the DU 553 decodes the VCL NAL unit within the DU 553. CPB54 contains non-VCL NAL units that contain data necessary to support loading 3 is a first-in, first-out buffer in the HRD 500. The CPB 543 stores the video in decoding order. CPB543 is used for bitstream conformance verification. The video data for the image processing is stored.
[0075] The CPB 543 forwards the DU 553 to the decoding process component 545. The processing component 545 is a component that complies with the VVC standard. For example, The decoding process component 545 performs the decoding process employed by the end user. The decoding process component 545 may emulate the decoder 400. Decode DU553 at a rate that can be achieved by the end user's decoder. The Handling Process Component 545 prevents the CPB 543 from overflowing (or overflowing the buffer). If the DU553 cannot be decoded fast enough to prevent underruns, the bitstream 551 is non-conforming and should be re-encoded.
[0076] The decoding process component 545 decodes the DU553 and outputs the decoded DU55 The decoded DU 555 contains the decoded picture. The decoded picture buffer component 2 is forwarded to the DPB 547. 23, 323, and / or 423. In order to support the above, a mark for use as a reference picture 556 obtained from the decoded DU 555 is The marked pictures are added to the decoding process to support further decoding. The DPB 547 returns the decoded video sequence to the process component 545. The picture 557 is output as a bitstream 551 by the encoder. A reconstructed picture generally mirrors the picture encoded in .
[0077] The picture 557 is forwarded to the output cropping component 549. The component 549 is configured to apply an adaptive cropping window to the picture 557. This results in the output cropped picture 559. The completed picture 559 is the fully reconstructed picture. The coded picture 559 is visible to the end user after decoding the bitstream 551. Therefore, the encoder outputs a cropped picture. You can review the 559 to ensure that the encoding is satisfactory. Cut.
[0078] The HRD 500 is initialized based on the HRD parameters in the bitstream 551. For example, The HRD 500 shall read the HRD parameters from the VPS, SPS, and / or SEI messages. The HRD 500 may then determine the bit rate based on the information in such HRD parameters. As a specific example, the HRD 500 may perform a conformance test operation on the HRD parameter stream 551. One or more CPB delivery schedules may be determined from the meter. To and / or from memory locations such as B and / or DPB This specifies the timing for the delivery of the video data. Therefore, the CPB delivery schedule The rule specifies the timing for delivery of AUs, DU553s, and / or pictures to / from CPB543s. The HRD500 specifies the DPB delivery schedule in the DPB547, which is similar to the CPB delivery schedule. Note that rules may be employed.
[0079] Video is further encoded for use by decoders with different levels of hardware capability. They are coded into different layers and / or OLS for various network conditions. The CPB delivery schedule will be chosen to reflect these issues. Therefore, the upper layer sub-bitstreams are optimized for optimal hardware and network conditions. Therefore, the upper layer must use the large amount of memory in the CPB 543 and the DPB 547. Receive one or more CPB delivery schedules that employ short delays for the transfer of DU553 to the Similarly, the lower layer sub-bitstreams can be trusted using limited decoder hardware. capacity and / or poor network conditions. The YA has a small amount of memory in CPB543 and a longer delay for the transfer of DU553 towards DPB547. Then, the OLS, layer, Each sublayer, or combination thereof, is tested according to a corresponding delivery schedule; The resulting sub-bitstream meets the expected conditions for the sub-bitstream. Therefore, it is possible to ensure that the bitstream can be correctly decoded under certain conditions. The HRD parameters in the Ream 551 indicate the CPB delivery schedule, and the HRD 500 also determines the CPB delivery schedule. Determine the schedule and sync the CPB delivery schedule to the corresponding OLS, layer, and / or sub The image may contain sufficient data to allow correlation to a breaker.
[0080] FIG. 6 illustrates an example multi-layer video sequence configured to perform inter-layer prediction 621. 6 is a schematic diagram illustrating a multi-layer video sequence 600. The multi-layer video sequence 600 may be implemented, for example, by a method 100, a codec system 200 and / or an encoder such as encoder 300 may be used. and a codec system 200 and / or a decoder such as decoder 400. Further, the multi-layer video sequence 600 can be decoded by the HRD 500, etc. The multi-layer video sequence 600 can be checked for conformance by the HRD of Figure 1 shows an example application for layers in a coded video sequence. The multi-layer video sequence 600 includes layer N 631 and layer The present invention is directed to a video sequence that employs multiple layers, such as a video sequence of ...
[0081] In one example, the multi-layer video sequence 600 may employ inter-layer prediction 621. Inter-layer prediction 621 is a method for predicting the quality of a picture in a different layer by comparing pictures 611, 612, 613, and 614 with picture 615. , 616, 617, and 618. In the illustrated example, pictures 611, 612 , 613, and 614 are part of layer N+1 632, and pictures 615, 616, 617, and 618 is part of layer N 631. Layers such as layer N 631 and / or layer N+1 632 , all related to similar values of characteristics such as similar size, quality, resolution, signal-to-noise ratio, and capacity. A layer is a group of pictures that are attached to a VCL NAL unit and associated A VCL NAL unit can be formally defined as the set of non-VCL NAL units that represent a picture. A coded slice of a video stream containing video data, such as a coded slice of a video stream. Non-VCL NAL units are used to decode video data, conformance checks, and Syntax and / or parameters that support running a check or other action. Any NAL unit that contains non-video data.
[0082] In the illustrated example, layer N+1 632 is associated with a larger image size than layer N 631. Thus, pictures 611, 612, 613, and 614 in layer N+1 632 are In this example, a picture in layer N 631 that is larger than pictures 615, 616, 617, and 618 is size (e.g., larger height and width, and therefore more samples) However, such a picture may be separated into layers N+1 632 and N+3 by other characteristics. Only two layers, layer N+1 632 and layer N 631, are shown in the figure. Although shown, the set of pictures may be separated into any number of layers based on relevant characteristics. Layer N+1 632 and Layer N 631 may also be indicated by a Layer Id. is the item of data associated with the picture and the layer in which the picture is shown Therefore, each picture 611-618 is associated with a corresponding layer ID. 631, which determines which layer N+1 632 or layer N 631 is the corresponding picture. For example, the layer Id may indicate whether the NAL unit header layer identifier (nuh_layer_ id), which is used to identify the NAL unit (e.g., slices and is a syntax element that specifies the identifier of a layer that contains Layers associated with lower quality / bitstream size, such as layer N 631. Layers are generally assigned lower layer IDs and are referred to as lower layers. Layers associated with higher quality / bitstream size, such as layer N+1 632 is typically assigned a higher layer Id and is referred to as the upper layer.
[0083] The pictures 611-618 of the different layers 631-632 are arranged to be displayed in alternative ways. As a specific example, the decoder may adjust the current display time if a smaller picture is desired. Picture 615 may be decoded and displayed, or the decoder may select a larger picture if desired. If desired, picture 611 may be decoded and displayed at the current display time. Pictures 611 to 614 in layer N+1 632 are the lower layer (regardless of the difference in picture size). Each of the pictures 615-618 in layer N 631 contains substantially the same image data as the corresponding pictures 615-618 in layer N 631. In general, picture 611 contains substantially the same image data as picture 615, and picture 612 contains substantially the same image data as picture 615. It contains substantially the same image data as picture 616, and so on.
[0084] Pictures 611-618 refer to other pictures 611-618 in the same layer N 631 or N+1 632. A picture can be coded with reference to another picture in the same layer. The result is an inter prediction 623. The inter prediction 623 is indicated by a solid arrow. For example, picture 613 is a composite of pictures 611, 612, and 613 in layer N+1 632. and / or employing inter-prediction 623 using one or two of 614 as a reference. A picture may be coded by one-way inter prediction. and / or two pictures are referenced for bidirectional inter prediction For example, picture 617 may be a subset of pictures 615, 616, and / or 618 in layer N 631. By using one or two of them as a reference, we can adopt inter-prediction 623. One picture may be referenced for unidirectional inter prediction. and / or two pictures are referenced for bidirectional inter prediction. However, when performing inter prediction 623, a picture is used as a reference to another picture in the same layer. When used, a picture may be referred to as a reference picture. For example, picture 612 is A reference picture used to code the picture 613 according to inter prediction 623 Inter prediction 623 may be considered as intra-layer prediction in a multi-layer context. Therefore, inter prediction 623 is a method for predicting whether a reference picture and a current picture are the same. The indicated sample in a reference picture different from the current picture if they are in the same layer. is a mechanism for coding samples of the current picture by referencing .
[0085] Pictures 611 to 618 can be arranged to refer to other pictures 611 to 618 in different layers. This process is known as inter-layer prediction 621 and is represented by the dashed line The inter-layer prediction 621 is indicated by the arrows. The indicated reference pictures are in different layers and therefore have different layer IDs. A mechanism for coding samples of the current picture by referencing samples in For example, a picture in lower layer N 631 has a corresponding picture in higher layer N+1 632. A corresponding picture may be used as a reference picture for coding the corresponding picture. Thus, the picture 611 is coded by referring to the picture 615 according to the inter-layer prediction 621. In such a case, the picture 615 is used as an inter-layer reference picture. The inter-layer reference picture is a reference picture used for inter-layer prediction 621. In most cases, inter-layer prediction 621 is performed based on the assumption that a current picture, such as picture 611, Use only inter-layer reference pictures that are in the same AU and are in a lower layer, such as picture 615 If multiple layers (e.g., more than two) are available, In this case, the inter-layer prediction 621 uses multiple inter-layer reference pictures at a level lower than the current picture. Based on the received signal, the current picture can be encoded / decoded.
[0086] The video encoder may implement many different combinations of inter prediction 623 and inter-layer prediction 621. and / or permutation of the multi-layer video to encode the pictures 611-618. For example, picture 615 may be a sequence of pictures that are predicted using intra prediction. Pictures 616 to 618 can then be coded using picture 615 as a reference picture. In addition, the pixel may be coded according to inter prediction 623 by using the pixel as a pixel. The picture 611 is an inter-layer reference picture by using the picture 615. Pictures 612-614 may then be coded according to the prediction 621. 623 as a reference picture. Therefore, a reference picture is a single layer reference for different coding mechanisms. It can act as both a reference picture and an inter-layer reference picture. By coding the enhancement layer N+1 632 picture based on the has a much lower coding efficiency than inter prediction 623 and inter-layer prediction 621. Therefore, it is possible to avoid adopting intra prediction. The loading inefficiency can be limited to the smallest / lowest quality picture, and therefore , can be limited to coding a minimum amount of video data. and / or pictures used as inter-layer reference pictures are added to the reference picture list This may be indicated in an entry in a reference picture list included in the structure.
[0087] Layers, such as Layer N+1 632 and Layer N 631, are included in the output layer set (OLS). Note that OLS is a one or more layer based method where at least one layer is the output layer. is a set of layers. For example, layer N 631 is included in the first OLS, and layer N Both layer N-1 632 and layer 633 may be included in the second OLS. Allows different OLS to be sent to different decoders depending on conditions. For example, The sub-bitstream extraction process performs the master bitstream extraction before the target OLS is sent to the decoder. It is possible to remove data that is not related to the target OLS from the multi-layer video sequence 600. Therefore, the encoded copy of the multi-layer video sequence 600 can be The OLS is stored in the coder (or corresponding content server) and is then used by various OLSs when requested. and transmitted to different decoders.
[0088] A simulcast layer is a layer that does not employ inter-layer prediction 621. For example, Layer N+1 632 is coded by referring to layer N 631 based on inter-layer prediction 621. However, layer 631 is coded by referencing other layers. Therefore, layer 631 is a simulcast layer. A scalable video sequence, such as the multi-layer video sequence 600, is generally , a base layer and one or more enhancements that enhance some characteristic of the base layer. In FIG. 6, layer N 631 is the base layer. The ears are typically coded as simulcast layers. Figure 6 shows an example. It is possible, but not limited to, that video sequences with multiple layers have dependencies Note also that many different combinations / permutations of A stream may include any number of layers, and any number of such layers may be simulcast. For example, the inter-layer prediction 621 can be omitted entirely, and in In this case, all layers are simulcast layers. An application displays two or more output layers, hence the multiview Applications typically use two or more simulcast layers It may include a base layer and may include an enhancement layer corresponding to each base layer.
[0089] Simulcast layers are processed differently than layers that use inter-layer prediction 621. For example, when coding a layer that uses inter-layer prediction 621, The decoder indicates the number of layers and even the dependencies between layers to support decoding. However, for the simulcast layer, such information is omitted. For example, the configuration of Layer N+1 632 and Layer N 631 is described in more detail below. However, layer N 631 does not have such information. Therefore, the VPS can be decoded as , can be removed from the corresponding bitstream. If any parameters remaining in the system refer to a VPS, this may cause errors. In addition, each layer can be described by an SPS. The SPS is a multi-layer bitstream to prevent the SPS from being accidentally removed along with the VPS. These and other issues are discussed in more detail below. do.
[0090] FIG. 7 is a schematic diagram illustrating an example bitstream 700. The stream 700 is generated by the codec system 200 and / or the decoder 400 according to the method 100. The codec system 200 and / or the encoder 300 generate the encoded data for decoding. Furthermore, the bitstream 700 may include the multi-layer video sequence 600. In addition, the bitstream 700 may be used to control the operation of an HRD, such as the HRD 500. Based on such parameters, the HRD 500 determines the decoding The bitstream 700 is checked for conformance to the standard before being transmitted to a decoder for can be checked.
[0091] The bitstream 700 includes a VPS 711, one or more SPSs 713, and a number of picture parameters. The VPS 711 includes a video processing set (PPS) 715, a number of slice headers 717, and image data 720. For example, VPS 711 contains data related to the entire bitstream 700. It may contain data related to OLS, layers, and / or sublayers used in the SPS 713 is a pre-coded video sequence included in bitstream 700. Each layer contains sequence data that is common to all pictures in the or a plurality of coded video sequences, each of which may include The video sequence may refer to SPS 713 for corresponding parameters. The parameters include picture sizing, bit depth, coding tool parameters, It can include bit rate limits, etc. While each sequence refers to SPS713, In some cases a single SPS713 may contain data for multiple sequences. Note that the PPS 715 contains parameters that apply to the entire picture. Each picture in a video sequence may reference a PPS 715. However, in some cases a single PPS715 may contain data for multiple pictures. Note that multiple similar pictures can be represented as similar parameters. In such a case, a single PPS 715 can be coded to contain all such similar pictures. The PPS 715 may contain data for slices, quantization, etc. in the corresponding picture. It can show available coding tools for parameters, offsets, etc.
[0092] The slice header 717 contains parameters specific to each slice in the picture. Thus, there may be one slice header 717 for each slice in a video sequence. The header 717 stores slice type information, a picture order count (POC), a reference picture list, a prediction These may include measurement weights, tile entry points, or deblocking parameters. In some examples, the bitstream 700 may also include a picture header, which is a syntax construct that contains parameters that apply to all slices in a single picture. For this reason, the picture header and slice header Da 717 can be used interchangeably in some contexts. For example, some The meter determines whether such parameters are common to all slices in the picture. The slice header 717 may be moved between the picture header 717 and the slice header 717 depending on the
[0093] The image data 720 may be encoded according to inter prediction, inter-layer prediction, and / or intra prediction. The encoded video data, as well as the corresponding transformed and quantized residual data For example, image data 720 includes layers 723 and 724, pictures 725 and 726, and Layers 723 and 724 may include slices 727 and 728. 2, etc. VCL NAL unit 741 shares the same frame rate, image size, etc. For example, layer 723 has the same nuh_layer_id 732. Similarly, layers 724 may include a set of pictures 725 that share the same nuh_layout The layers 723 and 724 may include a set 726 of pictures that share an er_id 732. They may be similar in nature but contain different content. For example, Layers 723 and 724 include Layer N 631 and Layer N+1 632 from FIG. Therefore, the coded bitstream 700 may include multiple layers 723 and 724. Only two layers 723 and 724 are shown for clarity of illustration. However, any number of layers 723 and 724 may be included in the bitstream 700.
[0094] nuh_layer_id 732 identifies layers 723 and / or 724 that contain at least one NAL unit. This is a syntax element that specifies an identifier for a layer. For example, the first layer, known as the base layer, Low quality layers are created using the lowest nuh_layer_id 732 with a higher value for higher quality layers. The lower layer may then include a nuh_layer_id 732 value. The layer with the smaller value is layer 723 or 724, and the upper layer is the layer with the larger value of nuh_layer_id 732. The data for layers 723 and 724 is stored in nuh_layer_ For example, the parameter set and video data are correlated based on the id 732. nuh_layer, which corresponds to the lowest layer 723 or 724 that contains the appropriate parameter set / video data _id 732 value. Therefore, the set of VCL NAL units 741 may be When a set of units 741 all have a particular value of nuh_layer_id 732, layers 723 and and / or part of 724.
[0095] Pictures from the set of pictures 725 and 726 make up a frame or its fields. An array of luma samples and / or chroma samples that form a pixel. Pictures from the sets of pictures 725 and 726 can be output for display or A coded picture that is used to support the coding of other pictures Pictures from sets of pictures 725 and 726 are substantially similar, but The set of pictures 725 is contained in layer 723, and the set of pictures 726 is contained in layer 724. Each of the pictures from the sets of pictures 725 and 726 is divided into one or more slices. Slices 727 / 728 contain a single NAL unit, such as VCL NAL unit 741. An integer number of complete tiles or (for example, A coding tree unit (CTU) row may be defined as an integer number of consecutive complete coding tree unit (CTU) rows. Slice 727 and slice 728 are included in picture 725 and layer 723, 7 is substantially similar to picture 726 and layer 724, except that slice 728 is included in picture 726 and layer 724. Slices 727 / 728 are further divided into CTUs and / or coding tree blocks (CTBs). A CTU is a block of predefined size that can be partitioned by the coding tree. A CTB is a group of samples. A CTB is a subset of a CTU and is the sum of the luma or cT CTU / CTB is divided into coding blocks based on the coding tree. The coding block is then encoded / decoded according to a prediction mechanism. It can be decoded.
[0096] Coded layer video sequence (CLVS) 743 and CLVS 744 are coded picture 725 and a sequence of pictures 725, respectively, that have the same nuh_layer_id 732 value. For example, CLVS 743 and / or 744 may be a sequence of pictures 725 and / or 726 contained in a single layer 723 and / or 724, respectively. Thus, CLVS 743 includes all of the pictures 725 in layer 723, and CLVS 744 includes all of the pictures 726 in layer 724, respectively. Each of CLVS 743 and 744 references a corresponding SPS 713. Depending on the example, CLVS 743 and 744 Refer to the same SPS713 or CLVS743 and 744 may each refer to a different SPS 713.
[0097] The bitstream 700 may be coded as a sequence of NAL units. A unit is a representation of the video data and / or the context for the syntax it supports. A NAL unit can be a VCL NAL unit 741 or a non-VCL NAL unit 742. The VCL NAL unit 741 contains image data 720 and associated slices. A NAL unit is a NAL unit coded to contain video data, such as JPEG 2000 (SEQ ID NO: 717). By way of example, each slice 727, 728 and associated slice header 717 may be a single V CL NAL units 741. Non-VCL NAL units 742 decode video data. Syntax and support for coding, performing conformance checks, or other actions NAL units that contain non-video data such as video and / or parameters. For example, non-VC L NAL unit 742 can be a VPS711, SPS713, PPS715, picture header, or other supported Therefore, the bitstream 700 may contain a sequence of V CL NAL unit 741 and non-VCL NAL unit 742. Each NAL unit is 732, which is used by an encoder or decoder to determine which layer 723 or 724 corresponds to This allows determining whether a NAL unit contains
[0098] A bitstream 700 containing multiple layers 723 and 724 is encoded and For example, the decoder may store layers 723, 724, and / or an OLS including multiple layers 723 and 724 may be required. In this example, layer 723 is a base layer and layer 724 is an enhancement layer. Additional layers may also be employed in the bitstream 700. The content server receives the layers 723 and 724 required to decode the requested output layer. and / or 724 should be sent to the decoder. When used for size, decoders that require the largest picture size will use layers 723 and 724. A minimum picture size may be required. Decoders that require intermediate picture sizes can only receive layer 723. may receive layer 723 and other intermediate layers, but does not receive the top layer 724, Therefore, the entire bitstream cannot be received. The same approach can be used for frame rate, pixel count, etc. It may also be used for other layer characteristics such as texture resolution.
[0099] The sub-bitstream extraction process 729 extracts sub-bitstreams from the bitstream 700. The subframe 701 is extracted and employed to support the functionality described above. Bitstream 701 is a stream of NAL units (e.g., non-VCL NAL units) from bitstream 700. The sub-bitstream is a subset of the VCL NAL unit 742 and the VCL NAL unit 741. A layer 701 may contain data relating to one or more layers, but may not be related to other layers. In the illustrated example, sub-bitstream 701 may not include data related to layer 7. It contains data related to layer 23, but does not contain data related to layer 724. The bitstream 701 includes an SPS 713, a PPS 715, a slice header 717, and layers 723, CLV S 743, picture 725, and slice 727. The frame extraction process 729 removes NAL units based on nuh_layer_id 732. For example, VCL NAL units 741 and non-VCL NAL units 74 associated only with the upper layer 724 2 contains a higher nuh_layer_id 732 value and therefore By removing all NAL units with Each NAL unit is then extracted to support the sub-bitstream extraction process 729. For this purpose, the nuh_layer_id 732 must be less than or equal to the nuh_layer_id 732 of the lowest layer that contains a NAL unit. The bitstream 700 and the sub-bitstream 701 each generally include a bit Note that it may be referred to as a stream.
[0100] In the illustrated example, the sub-bitstream 701 is a simulcast layer (e.g. As pointed out above, the simulcast layer includes the It is an optional layer that does not use inter-layer prediction. VPS711 records the configuration of layers 723 and 724. However, this data is not included in the Simulcast data, such as layer 723. Therefore, the sub-bitstream extraction process is not required to decode the sub-layer. Process 729 supports increased coding efficiency when extracting simulcast layers. Remove VPS711 to support some video coding systems. In particular, some parameters of the SPS713 refer to the VPS711. When the VPS711 is removed, the decoder and / or HRD may not use such parameters. It is not possible to resolve such parameters because the data referenced by them no longer exists. As a result, it may not be possible to perform conformance tests for the simulcast layer in the HRD. Alternatively, this may result in an error when executing the When the simulcast layer is transmitted to Furthermore, if each SPS713 cannot be correctly identified, the result is a sub-bit error. The stream extraction process 729 removes the VPS 711 for the simulcast layer. This can result in improper removal of SPS713 when
[0101] This disclosure addresses these errors. Specifically, the SPS713 is a ray-based Refer to SPS713 when Y723 does not use inter-layer prediction (is a simulcast layer). SPS713 is constrained to contain the same nuh_layer_id 732 as CLVS743, which it C when layer 724, including 44, uses inter-layer prediction (is not a simulcast layer) CLVS744's nuh_layer_id 732 may be equal to or less than the nuh_layer_id 732 of LVS744. Both SPS 713 and 743 may optionally refer to the same SPS 713. contains the same nuh_layer_id 732 as CLVS743 / layer 723, so it is a simulcast layer (for example For example, layer 723) is not removed by the sub-bitstream extraction process 729 .
[0102] Further, the SPS 713 includes an sps_video_parameter_set_id 731. t_id 731 is a syntax element that specifies the ID of the VPS 711 referenced by the SPS 713 . Specifically, VPS 711 includes vps_video_parameter_set_id 735, which is used by other syntax This is a syntax element that provides the ID of the VPS711 referenced by a VPS711 property element / structure. When 1 is present, sps_video_parameter_set_id 731 is set to vps_video_parameter_set_id 73 It is set to a value of 5. However, when SPS713 is used for the simulcast layer , sps_video_parameter_set_id 731 is set to 0. In other words, sps_video When _parameter_set_id 731 is greater than 0, for VPS711 referenced by SPS713 , vps_video_parameter_set_id 735 is specified. sps_video_parameter_set_id 731 is specified. When equal to 0, SPS713 does not reference VPS711 and any coded No VPS711 is referenced when decoding a multi-layer video sequence. This is done by using different layers (e.g. one SPS for the simulcast layer and one SPS for the non-simulcast layer). Use a separate SPS713 for the multicast layer) or a separate SPS for the sub-bitstream. During the stream extraction process 729, the value of sps_video_parameter_set_id 731 is changed. In this way, sps_video_parameter_set_id 731 can be used to During the system extraction process 729, the VPS 711 is removed, but it incorrectly references an ID that is not available. Never.
[0103] In addition, various variables derived by the HRD and / or the decoder are also used as parameters for the VPS711. Therefore, such variables refer to the sps_video_parameter_set_id 731 set to 0. This is the default for multi-layer bitstreams. While such variables are working correctly, the VPS711 is not aware of them for the simulcast layer. ensure that the decoded values can be properly resolved to usable values when extracted. The reader and / or the HRD may receive the bitstream 700 and / or the sub-bitstream 701. GeneralLayerIdx[i] can be derived based on the corresponding A derived variable that specifies the index of layer i. Therefore, GeneralLayerIdx[i] is , and include the current layer's nuh_layer_id732 as layer i in GeneralLayerIdx[i]. This can be employed to determine the layer index of the current layer by: The general layer index (GeneralLayerIdx[nuh_layer_id]) corresponding to nuh_layer_id Therefore, GeneralLayerIdx[nuh_layer_id] is the ID for the corresponding layer. This process is also applicable to non-simulcast layer 724, such as layer 725. It works fine for the cast layer, but gives an error for the simulcast layer 723. Therefore, sps_video_parameter_set_id 731 is set to 0 (simultaneous GeneralLayerIdx[nuh_layer_id] is set to 0 when the layer is a general-cast layer. and / or 0.
[0104] As another example, VPS 711 includes the VPS independent layer flag (vps_independent_layer_flag) 733. vps_independent_layer_flag 733 indicates whether the corresponding Specifies whether the layer being processed uses inter-layer prediction. layer_flag[GeneralLayerIdx[nuh_layer_id]] is the index GeneralLayerIdx[nuh_layer er_id] specifies whether the current layer uses inter-layer prediction. Therefore, VPS711 with vps_independent_layer_flag733 indicates that layer 723 sent to the decoder When it is a simulcast layer, it is not transmitted to the decoder. Therefore, the reference is not an error. However, the simulcast layers use inter-layer prediction. Therefore, vps_independent_layer_flag733 for the simulcast layer is , 1, which means that inter-layer prediction is performed for the corresponding layer 723. Therefore, vps_independent_layer_flag[GeneralLayerIdx[n uh_layer_id]] is the ID for the current layer when sps_video_parameter_set_id is set to 0. It is inferred to be set to 1 to indicate that inter-layer prediction is not used. In this way, errors are detected in the bit stream prior to transmission of a simulcast layer, such as layer 723. This is avoided when the VPS is removed from the stream. As a result, the encoder and decoder Furthermore, the coding efficiency is improved by using only the simulcast layer. This is enhanced by successfully removing unnecessary VPS from the bitstream, , both in the encoder and the decoder, the processor, memory, and / or network This reduces the usage of signaling resources.
[0105] The above information will then be explained in more detail below in this specification. Layered video coding is also known as scalable video coding or This is also called scalable video coding. The scalability is supported by using multi-layer coding technology. A multi-layer bitstream may include a base layer (BL) and one or more An example of scalability is spatial scalability. ,Quality / Signal-to-Noise Ratio (SNR) Scalability, Multi-view Scalability, Frame rate scalability, etc. When multi-layer coding techniques are used, A picture or part of it is coded without the use of reference pictures (int It is coded by referring to a reference picture in the same layer (inference prediction). prediction), and / or by referencing reference pictures in other layers. The reference picture used for inter-layer prediction of the current picture can be The picture in the different layers is called an inter-layer reference picture (ILRP). A multi-layer coding approach for spatial scalability with different resolutions An example is illustrated.
[0106] Some video coding families provide profiles for single-layer coding. Provides support for scalability in profiles separated from files. Scalable Video Coding (SVC) is a method for the scalability of video in terms of spatial, temporal and quality. Advanced Video Coding (AVC) scalable encoding provides support for For SVC, the flag is set for each macroblock (MB) in the EL picture. The EL MB is then signaled using the same location block from the lower layer. Prediction from co-located blocks is used to estimate texture, motion vectors, and In an SVC implementation, the design may include The unmodified AVC implementation cannot be directly reused for SVC EL macroblock syntax and The syntax and decoding process is different from the AVC syntax and decoding process. do.
[0107] Scalable HEVC (SHVC) provides support for spatial and quality scalability. Multiview HEVC (MV-HEVC) is an extension of HEVC that provides multiview scalability. 3D HEVC (3D-HEVC) is an extension of HEVC that provides support for 3D video coding. The HEVC extension provides support for 3D video coding that is highly advanced and efficient. Temporal scalability is an integral part of the single-layer HEVC codec. In the multi-layer extension of HEVC, the decoded pictures used for inter-layer prediction may be included. Such pictures come only from the same AU and are treated as long-term reference pictures (LTRPs). The picture is added to the reference picture list along with other temporal reference pictures in the current layer. Inter-layer prediction (ILP) is performed at the prediction unit (PU) level. The reference index is then adjusted to refer to an inter-layer reference picture in the reference picture list. Spatial scalability is achieved by setting the ILRP When a reference picture has a different spatial resolution than the current picture being Resample a picture or part of it. Reference picture resampling is a method to resample a picture. This can be implemented either at the bell or coding block level.
[0108] VVC may also support layered video coding. A VVC bitstream is , can contain multiple layers. The layers are all considered independent of each other. For example, each layer may be coded without using inter-layer prediction. In some cases, the layer is also called a simulcast layer. Some of the layers are coded using ILP. A flag in the VPS indicates that the layer is simulcast. It can indicate whether a layer is a layer or not, and whether some layer uses ILP. When several layers use ILP, the layer dependencies between the layers are also signaled in the VPS. Unlike SHVC and MV-HEVC, VVC cannot specify OLS. OLS is a layer Contains the specified set, and one or more layers in the set of layers are treated as output layers. The output layer is the output OLS layer. In an implementation, when a layer is a simulcast layer, only one layer In some implementations of VVC, any When a layer uses ILP, the entire bitstream including all layers is decoded. In addition, a particular layer among the layers is specified as the output layer. The output layers can be the top layer only, all layers, or the top layer + the indicated layers. The set of lower layers may be denoted as
[0109] The above aspects raise several scalability-related issues. Scalability design in the system is based on layer-specific profiles, tiers, and levels. (PTL), as well as layer-specific Coded Picture Buffer (CPB) operations. Signaling efficiency should be improved. Sequence-level HRD for sublayers The efficiency of parameter signaling should be improved. DPB parameter signaling is This needs to be improved. Some designs have single-layer bitstreams referencing VPSs. The value range of num_ref_entries[][] in such a design is invalid. This can cause unexpected errors to the decoder. The decoding process involves sub-bitstream extraction, which places a burden on the decoder implementation. The general decoding process for such a design is Such a configuration cannot work for scalable bitstreams that contain multiple layers. The derivation of the value of the variable NoOutputOfPriorPicsFlag in the design is picture-based, and as such In such a design, AU-based is not possible. nesting_ols_flag is equal to 1, the nesting SEI message is directly connected to the OLS, not to a layer of OLS. Non-scalable nested SEI messages should be simplified to apply , payloadType is 0 (buffering period), 1 (picture timing), or 130 (decode When this is equal to 0th OLS (0th OLS input information), it can be specified to apply only to the 0th OLS.
[0110] In general, this disclosure describes various approaches for scalability in video coding. The techniques are based on VVC. However, these techniques , and also applies to layered video coding based on other video codec specifications. One or more of the above problems may be solved as follows. Specifically, this disclosure provides: A method for improved scalability support in video coding is included.
[0111] Below are some examples of different definitions. OP is determined by the OLS index and the highest value of TemporalId. The output layer may be a temporal subset of the OLS, identified by the An OLS may be a set of layers, and one or more of the layers may be The layer of the model is specified to be the output layer. The OLS layer index is the The subbits may be the index of the layer in the OLS into a list of layers that can be The stream extraction process uses the target OLS index and the target highest TemporalId NAL units in the bitstream that do not belong to the target set, as determined by It may be a specified process of removing from the bitstream, A frame contains NAL units in the bitstream that belong to a target set.
[0112] An example video parameter set RBSP syntax is as follows:
[0113] [Table 1A]
[0114] [Table 1B]
[0115] An exemplary sequence parameter set RBSP syntax is as follows:
[0116] [Table 2A]
[0117] [Table 2B]
[0118] An exemplary DPB parameter syntax is as follows:
[0119] [Table 3]
[0120] An exemplary general HRD parameter syntax is as follows:
[0121] [Table 4]
[0122] An exemplary OLD HRD parameter syntax is as follows:
[0123] [Table 5]
[0124] An example sub-layer HRD parameter syntax is as follows:
[0125] [Table 6]
[0126] Exemplary video parameter set RBSP semantics are: vps_max_la yers_minus1+1 specifies the maximum number of layers allowed in each CVS that references a VPS. _layers_minus1+1 is the number of temporal sublayers that may exist in each CVS that references the VPS. vps_max_sub_layers_minus1 is a value in the range 0 to 6. vps_all_layers_same_num_sub_layers_flag equal to 1 means that the number of temporal sublayers is the same. v equal to 0 specifies that the VPS is the same for all layers in each CVS that references it. ps_all_layers_same_num_sub_layers_flag specifies that the layers in each CVS that references the VPS are the same. Specifies whether the image may or may not have any number of temporal sub-layers. When vps_all_layers_same_num_sub_layers_flag is set, the value of vps_all_layers_same_num_sub_layers_flag can be inferred to be equal to 1. vps_all_independent_layers_flag is equal to 0, all layers in the CVS use inter-layer prediction. vps_all_ind equal to 0. dependent_layers_flag indicates whether one or more of the layers in the CVS may use inter-layer prediction. When not present, the value of vps_all_independent_layers_flag is equal to 1. When vps_all_independent_layers_flag is equal to 1, vps_independent The value of ent_layer_flag[i] is inferred to be equal to 1. When is equal to 0, the value of vps_independent_layer_flag[0] is inferred to be equal to 1.
[0127] vps_direct_dependency_flag[i][j] equal to 0 means that the layer with index j is vp equal to 1 specifies that the layer with index i is not a direct reference layer. s_direct_dependency_flag[i][j] indicates that the layer with index j has vps_direct_dependency_flag Specifies that the layer is a direct reference layer to the layer you want to add. [i][j] does not exist for i and j in the range 0 to vps_max_layers_minus1 The flag is inferred to be equal to 0. It specifies the jth directly subordinate layer of the ith layer. The variable DirectDependentLayerIdx[i][j], and the layer with layer index j are Variable LayerUse that specifies whether the layer is used as a reference layer by any other layer. dAsRefLayerFlag[j] may be derived as follows: for(i=0; i<=vps_max_layers_minus1; i++) LayerUsedAsRefLayerFlag[j]=0 for(i=1; i <vps_max_layers_minus1; i++) if(!vps_independent_layer_flag[i]) for(j=i-1, k=0; j>=0; j--) if(vps_direct_dependency_flag[i][j]) { DirectDependentLayerIdx[i][k++]=j LayerUsedAsRefLayerFlag[j]=1 }
[0128] A variable that specifies the layer index of the layer whose nuh_layer_id is equal to vps_layer_id[i]. The number GeneralLayerIdx[i] may be derived as follows: for(i=0; i<=vps_max_layers_minus1; i++) GeneralLayerIdx[vps_layer_id[i]]=i
[0129] each_layer_is_an_ols_flag equal to 1 indicates that each output layer set contains exactly one layer. Each layer in the bitstream itself is a single contained layer, and only one output layer is allowed. Specifies that the output layer set is an ols layer. vps_max_layers_minus specifies that the output layer set may contain multiple layers. If 1 is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. If not, when vps_all_independent_layers_flag is equal to 0, each_layer_is_an_ols_f The value of lag is inferred to be equal to 0.
[0130] ols_mode_idc equal to 0 means that the total number of OLSs specified by the VPS is vps_max_layers_minus1 +1, the i-th OLS contains layers with layer indices from 0 to i, and each OL Specifies that only the top layer in the OLS for S is output. ols_mode_id equal to 1 c is the total number of OLSs specified by the VPS, equal to vps_max_layers_minus1+1, and the i-th OLS contains layers with layer indices from 0 to i, and for each OLS, Specifies that layers of are to be output. ols_mode_idc equal to 2 specifies that the The total number of OLSs that have been used is explicitly signaled, and for each OLS, the top layer and ols_mode_i Specifies that an explicitly signaled set of the highest layer is output. The value of dc can range from 0 to 2. If vps_all_independent_layers_flag is equal to 1, When each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is assumed to be equal to 2. num_output_layer_sets_minus1+1 is the number of layers generated by the VPS when ols_mode_idc is equal to 2. Specifies the total number of OLSs specified.
[0131] The variable TotalNumOlss, which specifies the total number of OLSs specified by the VPS, is derived as follows: It is possible. if(vps_max_layers_minus1==0) TotalNumOlss=1 else if(each_layer_is_an_ols_flag || ols_mode_idc==0 || ols_mode_idc==1) TotalNumOlss=vps_max_layers_minus1+1 else if(ols_mode_idc==2) TotalNumOlss=num_output_layer_sets_minus1+1
[0132] layer_included_flag[i][j] is the flag for the jth layer (for example, nuh_layer_id is vps_layer_i d[j]) is included in the ith OLS when ols_mode_idc is equal to 2. layer_included_flag[i][j] equal to 1 indicates that the jth layer is included in the ith OLS. layer_included_flag[i][j] equal to 0 specifies that the jth layer is included in the ith Specifies that layers are not included in the OLS. The variable NumLayers specifies the number of layers in the i-th OLS. InOls[i] and the variable LayerIdIn specifying the nuh_layer_id value of the jth layer in the ith OLS Ols[i][j] can be derived as follows: NumLayersInOls[0] = 1 LayerIdInOls[0][0]=vps_layer_id[0] for(i=1, i <TotalNumOlss; i++) { if(each_layer_is_an_ols_flag) { NumLayersInOls[i] = 1 LayerIdInOls[i][0]=vps_layer_id[i] } else if(ols_mode_idc==0 | | ols_mode_idc==1) { NumLayersInOls[i]=i+1 for(j=0; j <NumLayersInOls[i]; j++) LayerIdInOls[i][j]=vps_layer_id[j] } else if(ols_mode_idc==2) { for(k=0, j=0; k<=vps_max_layers_minus1; k++) if(layer_included_flag[i][k]) LayerIdInOls[i][j++]=vps_layer_id[k] NumLayersInOls[i] = j } }
[0133] nuh_layer_id specifies the OLS layer index of the layer equal to LayerIdInOls[i][j]. The variable OlsLayeIdx[i][j] can be derived as follows: for(i=0, i <TotalNumOlss; i++) for j=0; j <NumLayersInOls[i]; j++) OlsLayeIdx[i][LayerIdInOls[i][j]]=j
[0134] The bottom layer in each OLS is assumed to be an independent layer. In other words, the layers range from 0 to TotalNum For each i in the range up to Olss-1, vps_independent_layer_flag[GeneralLayerIdx[Layer The value of rIdInOls[i][0]] shall be equal to 1. Each layer shall have at least In other words, the range from 0 to vps_max_layers_minus1. A specific value of nuhLayerId that is equal to one of vps_layer_id[k] for k in the range For each layer, there is at least one pair of values of i and j, where i runs from 0 to T otalNumOlss-1, j is in the range of NumLayersInOls[i]-1, and LayerIdInOls[ The value of [i][j] is equal to nuhLayerId. Any layer in the OLS can be the output layer of the OLS or the It shall be a reference layer (directly or indirectly) of the output layer.
[0135] vps_output_layer_flag[i][j] is the jth output layer in the ith OLS when ols_mode_idc is equal to 2. Specifies whether the ith layer is output. vps_output_layer_flag[i] equal to 1 specifies whether the ith layer is output. vps_output_layer_ equal to 0 specifies that the jth layer in the jth OLS is to be output. flag[i] specifies that the j-th layer in the i-th OLS is not output. When pendent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, vps_ou The value of tput_layer_flag[i] is inferred to be equal to 1. A value of 1 indicates the jth layer in the ith OLS. a value of 0 means that the jth layer in the ith OLS is not output. The designated variable OutputLayerFlag[i][j] can be derived as follows: for(i=0, i <TotalNumOlss; i++) { OutputLayerFlag[i][NumLayersInOls[i]-1]=1 for(j=0; j <NumLayersInOls[i]-1; j++) if(ols_mode_idc[i]==0) OutputLayerFlag[i][j]=0 else if(ols_mode_idc[i]==1) OutputLayerFlag[i][j]=1 else if(ols_mode_idc[i]==2) OutputLayerFlag[i][j]=vps_output_layer_flag[i][j] }
[0136] The 0th OLS is the lowest layer (for example, the layer whose nuh_layer_id is equal to vps_layer_id[0]). Note that for the 0th OLS, only the included layers are output. vps_num_ptls specifies the number of VPS profile_tier_level() syntax structures. pt_present_flag[i] equal to 1 indicates the profile, tier, and general constraint information. Specifies that exists in the i-th profile_tier_level() syntax structure in the VPS. pt_present_flag[i] equal to 0 indicates that profile, tier, and general constraint information is , specifies that it does not exist in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 0. If pt_present_flag[i] is equal to 0, When the profile is created, the profile for the i-th profile_tier_level() syntax structure in the VPS is created. The tier and general constraint information are stored in the (i-1)th profile_tier_level() syntax in the VPS. It is inferred that the structure is identical to that of the box structure.
[0137] ptl_max_temporal_id[i] is the value of the i-th profile_tier_level() syntax in the VPS. ptl_max_temporal Specifies the TemporalId of the highest sublayer representation present in the text box structure. The value of _id[i] must be in the range of 0 to vps_max_sub_layers_minus1. When x_sub_layers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is assumed to be equal to 0. vps_max_sub_layers_minus1 is greater than 0 and vps_all_layers_same_num_sub_la When yers_flag is equal to 1, the value of ptl_max_temporal_id[i] is equal to vps_max_sub_layers_minus It is inferred to be equal to 1. vps_ptl_byte_alignment_zero_bit shall be inferred to be equal to 0.
[0138] ols_ptl_idx[i] is the value of the profile_tier_level() syntax structure applied to the ith OLS. Specifies an index into the list of profile_tier_level() syntax structures within the VPS. When present, the value of ols_ptl_idx[i] must be in the range from 0 to vps_num_ptls-1. When NumLayersInOls[i] is equal to 1, the profile_tier_l applied to the i-th OLS is The evel() syntax structure exists in the SPS referenced by the layer in the i-th OLS. vps_num_dpb_params specifies the number of dpb_parameters() syntax structures in the VPS. The value of ps_num_dpb_params must be in the range of 0 to 16. If it is not present, , the value of vps_num_dpb_params can be inferred to be equal to 0. same_dpb_size_outp equal to 1 ut_or_nonoutput_flag is the value of the layer_nonoutput_dpb_params_idx[i] syntax element in the VPS. same_dpb_size_output_or_nonoutput_flag equal to 0 specifies that the ayer_nonoutput_dpb_params_idx[i] syntax element may or may not exist in the VPS. vps_sub_layer_dpb_params_present_flag specifies that the dpb_params in the VPS cannot be present. max_dec_pic_buffering_minus1[ ], max_num_reord in the parameters() syntax structure Controls the presence of the er_pics[ ] and max_latency_increase_plus1[ ] syntax elements When not present, vps_sub_dpb_params_info_present_flag is set to 0. is inferred to be equal to
[0139] dpb_size_only_flag[i] equal to 1 overrides max_num_reorder_pics[ ] and max_latency_inc The rease_plus1[ ] syntax element is the VPS of the i-th dpb_parameters() syntax structure. dpb_size_only_flag[i] equal to 1 specifies that the The r_pics[ ] and max_latency_increase_plus1[ ] syntax elements are the i-th dpb_paramet ers() syntax structure that can be present in that VPS. The maximum number of DPB parameters that can be present in the i-th dpb_parameters() syntax structure in the VPS. Specifies the TemporalId of the upper sublayer representation. The value of dpb_max_temporal_id[i] ranges from 0 to vps vps_max_sub_layers_minus1 must be in the range of 0. , the value of dpb_max_temporal_id[i] can be inferred to be equal to 0. _layers_minus1 is greater than 0 and vps_all_layers_same_num_sub_layers_flag is equal to 1 When layer_output_dpb_params_idx[i] is the ith layer when it is an output layer in OLS. The dpb_parameters() syntax in the VPS applies to the dpb_parameters() syntax structure. Specifies an index into the list of dpb_p structures. The value of arams_idx[i] shall be in the range of 0 to vps_num_dpb_params-1.
[0140] If vps_independent_layer_flag[i] is equal to 1, the i-th layer is used as the output layer. The dpb_parameters() syntax structure applied to the layer is the SPS referenced by the layer. The dpb_parameters() syntax structure exists in the vps_independent() structure. nt_layer_flag[i] is equal to 1), the following applies: vps_num_dpb_params is equal to 1 When the bitstream is The requirement for frame conformance is that the value of layer_output_dpb_params_idx[i] must be greater than or equal to dpb_size_only_flag[lay er_output_dpb_params_idx[i]] may be set to a value equal to 0.
[0141] layer_nonoutput_dpb_params_idx[i] is the i-th layer is a non-output layer in OLS The dpb_parameters() syntax structure that applies to the i-th layer when Specifies an index into the list of parameters() syntax structures. If present, The value of layer_nonoutput_dpb_params_idx[i] must be in the range from 0 to vps_num_dpb_params-1. If same_dpb_size_output_or_nonoutput_flag is equal to 1, the following applies: If vps_independent_layer_flag[i] is equal to 1, the i-th layer is a non-output layer. The dpb_parameters() syntax structure that applies to the ith layer when the layer is The dpb_parameters() syntax structure exists in the SPS referenced by the Otherwise (vps_independent_layer_flag[i] is equal to 1), layer_nonoutput_dpb_pa The value of rams_idx[i] is inferred to be equal to layer_output_dpb_params_idx[i]. If the same_dpb_size_output_or_nonoutput_flag is not set (same_dpb_size_output_or_nonoutput_flag is equal to 0), vps_num_dpb_param When s is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.
[0142] general_hrd_params_present_flag equal to 1 specifies the syntax element num_units_in_tick and time_scale and the syntax structure general_hrd_parameters() are part of the SPS RBSP syntax. general_hrd_params_present_flag equal to 0 specifies that the , the syntax elements num_units_in_tick and time_scale, and the syntax structure generator Specifies that al_hrd_parameters() is not present in the SPS RBSP syntax structure. nits_in_tick is the number of ticks that occur for one increment of the clock tick counter (called a clock tick). num_units is the number of time units for a clock running at the corresponding frequency time_scalehertz (Hz). _in_tick must be greater than 0. Clock ticks are in seconds, and num_units_i Equal to n_tick divided by time_scale. For example, if the picture rate of a video signal is 25 Hz, time_scale may be equal to 27,000,000 and num_units_in_tick may be equal to 1 ,080,000, so that a clock tick is equal to 0.04 seconds. time_scale is the number of time units that pass in one second. For example, a 27MHz clock A time coordinate system that measures time using clocks has a time_scale of 27,000,000. The value of cale must be greater than 0.
[0143] vps_extension_flag equal to 0 means that the VPS RBSP syntax structure does not contain the vps_extension_data_fl Specifies that the ag syntax element is not present. vps_extension_flag equal to 1 If the vps_extension_data_flag syntax element is present in the VPS RBSP syntax structure, vps_extension_data_flag can have any value. The presence and value of a_flag may not affect decoder conformance to a profile. A decoder that supports this feature may ignore all vps_extension_data_flag syntax elements.
[0144] An exemplary sequence parameter set RBSP semantics is as follows: SPS RB The SP must be contained in at least one access unit with TemporalId equal to 0 or it must be outside the access unit. provided through external means and available to the decoding process before being referenced Suppose that an SPS NAL unit that contains an SPS RBSP is a PPS NAL unit that references the SPS NAL unit. The sps_seq_p in CVS shall have a nuh_layer_id equal to the lowest nuh_layer_id value in the All SPS NAL units with a particular value of parameter_set_id have the same content. When sps_decoding_parameter_set_id is greater than 0, it is referenced by the SPS. Specifies the value of dps_decoding_parameter_set_id for the DPS to be When _set_id is equal to 0, the SPS does not reference the DPS, and decodes each CLVS that references the SPS. When the DPS is not referenced, the value of sps_decoding_parameter_set_id is shall be the same in all SPSs referenced by the coded picture do.
[0145] When sps_video_parameter_set_id is greater than 0, it is used for the VPS referenced by the SPS. Specify the value of vps_video_parameter_set_id to be used. When the SPS is not referenced, the VPS cannot be referenced, and the VPS cannot be referenced when decoding each CLVS that references the SPS. and the value of GeneralLayerIdx[nuh_layer_id] shall be inferred to be equal to 0. The value of s_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1. When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1 , an SPS referenced by a CLVS with a particular nuh_layer_id value nuhLayerId is The layer shall have a nuh_layer_id equal to
[0146] sps_max_sub_layers_minus1+1 may exist in each CLVS that references the SPS. Specifies the maximum number of temporal sublayers. The value of sps_max_sub_layers_minus1 ranges from 0 to vps_ma x_sub_layers_minus1. sps_reserved_zero_4bits is the sps_reserved_zero_4bits shall be equal to 0 in the bitstream in which it is used. The value may be reserved.
[0147] sps_ptl_dpb_present_flag equal to 1 indicates that the profile_tier_level() syntax structure and Specifies that the dpb_parameters() and dpb_parameters() syntax structures are present in the SPS. Equal to 0. sps_ptl_dpb_present_flag is also supported by the profile_tier_level() syntax structure and the dpb_parameter sps_ptl_dpb_present Specifies that the rs() syntax structure is not present in the SPS. The value of _flag shall be equal to vps_independent_layer_flag[nuh_layer_id]. If dependent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the variable MaxDecPic BuffMinus1 is the max_dec_pic_buffer in the dpb_parameters() syntax structure in the SPS. It is set to be equal to ng_minus1[sps_max_sub_layers_minus1], otherwise , MaxDecPicBuffMinus1 is layer_nonoutput_dpb_params_idx[GeneralLayerIdx[n max_dec_pic_buffering_minus of the dpb_parameters() syntax structure of the [uh_layer_id]]th gdr_enabled_f equal to 1. lag specifies that GDR pictures may be present in the CLVS that references the SPS. gdr_en equal to 0 abled_flag specifies that no GDR pictures exist in the CLVS that references the SPS.
[0148] sps_sub_layer_dpb_params_flag is the flag for the dpb_parameters() syntax in the SPS. x_dec_pic_buffering_minus1[i], max_num_reorder_pics[i], and max_latency_increa Used to control the presence of the se_plus1[i] syntax element. , sps_sub_dpb_params_info_present_flag is inferred to be equal to 0. long_t equal to 0 erm_ref_pics_flag specifies whether LTRP should be used for inter prediction of any coded picture in CLVS. long_term_ref_pics_flag equal to 1 specifies that one or more or specifies that LTRP may be used for inter prediction of multiple previously coded pictures. do.
[0149] Example general profile, tier, and level semantics are as follows: The profile_tier_level() syntax structure specifies the level information and, optionally, the profile Profile, tier, subprofile, and general constraint information (shown as PT information) When the profile_tier_level() syntax structure is included in the DPS, OlsInSco pe is an OLS that includes all layers in the entire bitstream that references the DPS. When the ofile_tier_level() syntax structure is included in the VPS, OlsInScope is The profile_tier_level() syntax structure is one or more OLSs specified by the SP When contained in S, OlsInScope is an independent layer, the layer that references the SPS. This is an OLS that includes only the lowest layer.
[0150] general_profile_idc indicates the profile that OlsInScope matches. ag specifies the tier context for the interpretation of general_level_idc. les specifies the number of general_sub_profile_idc[i] syntax elements. file_idc[i] indicates the ith registered interoperability metadata. vel_idc indicates the level at which OlsInScope is compatible. The higher the value of general_level_idc, the higher the level. Note that the DPS signaled for OlsInScope indicates that the The maximum level that can be signaled in the SPS for the CVS contained within OlsInScope. When OlsInScope matches multiple profiles, the l_profile_idc is the preferred decoded result or It should also be noted that the profile that provides the bitstream identification should be Please note that the profile_tier_level() syntax structure is included in the DPS and the CVS of OlsInScope When fitting different profiles, general_profile_idc and level_idc are used in OlsInScope The profile and level for the decoder that can decode the Please also note that
[0151] sub_layer_level_present_flag[i] equal to 1 indicates that the level information has TemporalId equal to i. Present in the profile_tier_level() syntax structure for the sublayer representation A sub_layer_level_present_flag[i] equal to 0 indicates that the level information is equal to i. The profile_tier_level() syntax structure for the sublayer representation with the given localId. Specifies that no alignment is present. ptl_alignment_zero_bits shall be equal to 0. Syntax The semantics of the sub_layer_level_idc[i] element is different from the specification of an inference for a non-existent value. Apart from that, we have the same syntax element general_level_idc, but with a TemporalId equal to i. This applies to the sublayer representations that have
[0152] Exemplary DPB parameter semantics are: dpb_parameters(maxSubLa yersMinus1, subLayerInfoFlag) syntax structure, DPB size, maximum picture order change It provides information about the number of CVSs, the number of CVSs, and the maximum wait time for each CVS. When a tax structure is included in the VPS, the dpb_parameters() syntax structure is applied. The LS is specified by the VPS. If the dpb_parameters() syntax structure is included in the SPS, When using the dpb_parameters() syntax structure, it is assumed that the structure is an independent layer. It is applied to OLS that includes only layers that are the lowest layer among the layers in the model.
[0153] max_dec_pic_buffering_minus1[i]+1 is the buffer size for each CLVS in CVS when Htid is equal to i. The maximum required decoded picture buffer in the picture storage buffer unit. Specifies the size. The value of max_dec_pic_buffering_minus1[i] ranges from 0 to MaxDpbSize-1. When i is greater than 0, max_dec_pic_buffering_minus1[i] is within the range m ax_dec_pic_buffering_minus1[i-1] or more. subLayerInfoFlag is equal to 0. Due to this, max_dec_pic_buffering_minus1[i] ranges from 0 to maxSubLayersMinus1-1. If max_dec_pic_buffering_minus1[i] does not exist for i in the range of Inferred to be equal to dec_pic_buffering_minus1[maxSubLayersMinus1].
[0154] max_num_reorder_pics[i] is the number of reordered pics for each CLVS in the CVS when Htid is equal to i. A CL that can precede any picture in decoding order and be followed by that picture in output order. Specifies the maximum number of pictures allowed in a VS. The value of max_num_reorder_pics[i] ranges from 0 to max_dec_ pic_buffering_minus1[i]. When i is greater than 0, max_num _reorder_pics[i] shall be equal to or greater than max_num_reorder_pics[i-1]. Due to lag equal to 0, max_num_reorder_pics[i] goes from 0 to maxSubLayersMinus1 If not present for i in the range of i to -1, then max_num_reorder_pics[i] is It is inferred to be equal to reorder_pics[maxSubLayersMinus1].
[0155] max_latency_increase_plus1[i] not equal to 0 calculates the value of MaxLatencyPictures[i] For each CLVS in CVS, it is used to find any A CLVS picture that precedes a picture in output order and can be followed in decoding order by that picture. Specifies the maximum number of pictures in a frame when max_latency_increase_plus1[i] is not equal to 0. , the value of MaxLatencyPictures[i] may be specified as follows: MaxLatencyPictures[i]=max_num_reorder_pics[i]+max_latency_increase_plus1[i]-1 When max_latency_increase_plus1[i] is equal to 0, the corresponding limit is not expressed.
[0156] The value of max_latency_increase_plus1[i] should be in the range of 0 to 2-2. max_latency_increase_plus1[i] is set from 0 to 0 due to ubLayerInfoFlag being equal to 0. max_latency_increment does not exist for i in the range maxSubLayersMinus1-1. ase_plus1[i] is inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1]. do.
[0157] Exemplary general HDR parameter semantics are: The eters() syntax structure specifies the HRD parameters used in the HRD calculation. rd_params_minus1+1 is the ols_hrd_parameters that are present in the general_hrd_parameters() syntax structure. Specifies the number of parameters() syntax structures. The value of num_ols_hrd_params_minus1 can be 0 or When TotalNumOlss is greater than 1, num_ols_hrd_par The value of ams_minus1 is inferred to be equal to 0. hrd_cpb_cnt_minus1+1 is the bitstream of the CVS. Specifies the number of alternate CPBs in the stream. The value of hrd_cpb_cnt_minus1 must be in the range 0 to 31. hrd_max_temporal_id[i] is the number of layers whose HRD parameters are the i-th layer_level. Specifies the TemporalId of the top-level sublayer representation contained in the l_hrd_parameters() syntax structure. The value of hrd_max_temporal_id[i] is in the range from 0 to vps_max_sub_layers_minus1. When vps_max_sub_layers_minus1 is equal to 0, hrd_max_temporal_id[ The value of ols_hrd_idx[i] is inferred to be equal to 0. ols_hrd_idx[i] is the ols_hrd Specifies the index of the _parameters() syntax structure. The value of ols_hrd_idx[[i] is 0. If not present, o The value of ls_hrd_idx[[i] is inferred to be equal to 0.
[0158] Exemplary reference picture list structure semantics are: ref_pic_list The _struct(listIdx, rplsIdx) syntax structure exists in the SPS or in the slice header. Whether the syntax structure is included in the slice header or the SPS may be determined. If present in the slice header, the following applies: ref_pic_list_structure The t(listIdx, rplsIdx) syntax structure is the list of the current picture (the picture that contains the slice). Specifies the reference picture list listIdx, otherwise (if present in the SPS), ref_pic_ The list_struct(listIdx, rplsIdx) syntax structure defines a list of reference pictures for the listIdx. specifies candidates for the current pixel in the semantics specified in the rest of this section. The term "char" refers to the ref_pic_list_struct(listIdx, rplsIdx) syntax in the SPS. One or more ref_pic_list_idx[listIdx] equals an index into the list of pic structures. Each picture in the CVS has a slice count and references the SPS. stIdx][rplsIdx] is an entry in the ref_pic_list_struct(listIdx, rplsIdx) syntax structure. The value of num_ref_entries[listIdx][rplsIdx] ranges from 0 to MaxDecPicBuffM It shall be within the range of inus1+14.
[0159] An exemplary general decoding process is as follows: The output of this process is the bitstream BitstreamToDecode. The decoding process is performed according to the specified profile and All decoders conforming to a level will receive bitstreams conforming to that profile and level. Invokes the decoding process associated with that profile for the stream. to produce a numerically identical cropped decoded output picture when The same as that formed by the process described herein. Any decoding process that produces a cropped decoded output picture may Decoding process (with correct output order or timing as specified) Meets service requirements.
[0160] For each IRAP AU in the bitstream, the following applies: Either it is the first AU in decoding order or each picture is instantaneous decoding refresh (IDR) pictures or each picture after the end of the sequence NAL unit in decoding order If this is the first picture in the layer following the Otherwise, the variable HandleCraAsCvsStartFlag is set to If set to a value, HandleCraAsCvsStartFlag is provided by an external mechanism. The NoIncorrectPicOutputFlag is set equal to the value specified by HandleCraAsCvsStartFlag. g. Otherwise, HandleCraAsCvsStartFlag and NoIn correctPicOutputFlag is set to both equal 0.
[0161] For each Gradual Decoding Refresh (GDR) AU in the bitstream, the following applies: The AU is either the first AU in the bitstream in decoding order, or is the first picture of the layer following the end of the sequence NAL unit in decoding order. If so, the variable NoIncorrectPicOutputFlag is set equal to 1. If so, some external mechanism sets the variable HandleGdrAsCvsStartFlag to the value for the AU. If HandleGdrAsCvsStartFlag is available to be started by an external mechanism, The NoIncorrectPicOutputFlag is set equal to the value provided by HandleGdrAsC vsStartFlag is set equal to HandleGdrAsCvsStartFlag otherwise and NoIncorrectPicOutputFlag are both set equal to 0. For both RGB and GDR pictures, the above operations identify the CVS in the bitstream. Decoding is performed for each coded picture in BitstreamToDecode. are called repeatedly for each memory in decoding order.
[0162] An exemplary decoding process for reference picture list construction is as follows: This process starts the decoding process for each slice of a non-IDR picture. A reference picture is addressed through a reference index. The reference index is an index into the reference picture list. When the slice data is decoded, the reference picture list is not used. When decoding a slice, only reference picture list 0 (e.g., RefPicList[0]) is Used when decoding slice data. When decoding B slices, the reference pixel Both reference picture list 0 and reference picture list 1 (for example, RefPicList[1]) are included in the slice data. This is used to decode the data.
[0163] The following constraints are applied to bitstream conformance: For num_ref_entries[i][RplsIdx[i]], num_ref_entries[i][RplsIdx[i]] is not less than NumRefIdxActive[i]. Each active entry in RefPicList[0] or RefPicList[1] references The picture is assumed to exist in the DPB and has a TemporalId less than or equal to the current picture's TemporalId. It shall have an Id, referenced by each entry in RefPicList[0] or RefPicList[1]. The picture to be referenced shall not be the current picture and shall have non_reference_picture equal to 0. _flag. RefPicList[0] or RefPicList[1] of a slice of a picture Short-term reference picture (STRP) entries in a slice and in different slices of the same slice or picture The long-term reference picture (LTRP) entries in Rice's RefPicList[0] or RefPicList[1] are the same. RefPicList[0] or RefPicList[1] should not refer to the same picture. The difference between the PicOrderCntVal of the entry and the PicOrderCntVal of the picture referenced by that entry. There should be no LTRP entries with a value greater than 224.
[0164] Set setOfRefPics to all the pictures in RefPicList[0] that have the same nuh_layer_id as the current picture. All entries in RefPicList[1] that have the same nuh_layer_id as the current picture Let setOfRefPics be the set of unique pictures referenced by all entries in the The number of pictures shall be less than or equal to MaxDecPicBuffMinus1, and setOfRefPics shall be the number of pictures. The same is true for all slices of the current picture. When it is a layer access (STSA) picture, RefPicList[0] or RefPicList[1] contains the current There shall be no active entry with a TemporalId equal to the TemporalId of the current picture. The current picture is assigned a TemporalId that is equal to the TemporalId of the current picture in decoding order. If the picture is a picture following an STSA picture with the specified Id, it shall precede the STSA picture in decoding order. The current pico- list is included as an active entry in RefPicList[0] or RefPicList[1]. There shall be no picture with a TemporalId equal to the TemporalId of the current picture.
[0165] Each inter-layer reference pic in RefPicList[0] or RefPicList[1] of the slice of the current picture The picture referenced by the Interleave Picture (ILRP) entry has the same access unit as the current picture. The slice of the current picture is assumed to be in RefPicList[0] or RefPicList[ The pictures referenced by each ILRP entry in [1] shall be present in the DPB and shall The RefP of the slice shall have a nuh_layer_id smaller than the nuh_layer_id of the picture. Each ILRP entry in icList[0] or RefPicList[1] should be an active entry. be.
[0166] An exemplary HRD specification is as follows: HRD specifies bitstream and decoder conformance. This is used to check the integrity of the bitstream, denoted as the entireBitstream. Bitstream Conformance is used to check the conformance of a bitstream, referred to as the whole A set of bitstream conformance tests is used. The set of bitstream conformance tests is specified by the VPS. The purpose of this test is to test the suitability of each OP for each OLS specified in the test.
[0167] For each test, the following steps are applied in the order listed, then The process described after these steps in section 3.3.1 follows. The operation point under test has the OLS index opOlsIdx and the highest Tempora It is selected by selecting the target OLS with the lId value opTid. The value of opOlsIdx is , in the range from 0 to TotalNumOlss-1. The value of opTid is in the range from 0 to vps_max_sub_layers_mi Each pair of selected values of opOlsIdx and opTid is in the range of the entireBitstore. Invoke the sub-bitstream extraction process with am, opOlsIdx, and opTid as inputs The sub-bitstreams output by the The nuh_layer_id is equal to the nuh_layer_id of LayerIdInOls[opOlsIdx] in BitstreamToDecode. There is at least one VCL NAL unit with the ayer_id value. There is at least one VCL NAL unit whose mporalId is equal to opTid.
[0168] The layers in targetOp include all layers in the entireBitstream and opTid is equal to enti If the TemporalId value is equal to or greater than the highest TemporalId value among all NAL units in the reBitstream, If so, BitstreamToDecode is set to be equal to the entireBitstream. If so, BitstreamToDecode takes the entireBitstream, opOlsIdx, and opTid as inputs. It is set to be output by invoking the bitstream extraction process. The values of rgetOlsIdx and Htid are equal to the opOlsIdx and opTid of the targetOp, respectively. The value of ScIdx is selected. The selected ScIdx is set in the range from 0 to hrd_cpb_cnt_minus The buffering period SEI message applicable to TargetOlsIdx shall be in the range of 1 to 1. Messages (present in the TargetLayerBitstream or available through an external mechanism) The access unit in BitstreamToDecode associated with is the HRD initialization point. and is referred to as access unit 0 for each layer of the target OLS.
[0169] The subsequent step is to use the OLS layer index TargetOlsLayerIdx in the target OLS. The ols_hrd_parameters() syntax applicable to BitstreamToDecode is applied to each layer that has The syntax structure and the sub_layer_hrd_parameters() syntax structure are selected as follows: The ols_hrd_idx[TargetOlsIdx] number in the VPS (or provided through an external mechanism) The first ols_hrd_parameters() syntax structure is selected. In the rs() syntax construct, if BitstreamToDecode is a Type I bitstream, The sub_layer_hrd_paramet that comes immediately after the condition "if(general_vcl_hrd_params_present_flag)" The ers(Htid) syntax structure is selected and the variable NalHrdModeFlag is set equal to 0. Otherwise (if BitstreamToDecode is a type II bitstream), Condition "if(general_vcl_hrd_params_present_flag)" (in this case, the variable NalHrdModeFlag is 0 is set to equal to ) or the condition "if (general_nal_hrd_params_present_flag) " (in this case the variable NalHrdModeFlag is set equal to 1) The sub_layer_hrd_parameters(Htid) syntax structure that comes in is selected. When ode is a Type II bitstream and NalHrdModeFlag is equal to 0, the filler data All non-VCL NAL units except for the data NAL units, as well as NAL units from the NAL unit stream All leading_zero_8bits, zero_byte, and start_code_pref that form the byte stream The ix_one_3bytes and trailing_zero_8bits syntax elements, when present, The remaining bitstream is assigned to BitstreamToDecode. can be.
[0170] When decoding_unit_hrd_params_present_flag is equal to 1, the CPB is level (in this case the variable DecodingUnitHrdFlag is set equal to 0) or At the coding unit level (in this case the variable DecodingUnitHrdFlag is set equal to 1) Otherwise, it is scheduled to run on either the ingUnitHrdFlag is set equal to 0, and the CPB operates at the access unit level. The system is scheduled to
[0171] For each access unit in BitstreamToDecode, starting with access unit 0, The buffering period SEI message associated with the access unit and applied to the TargetOlsIdx. message (present in BitstreamToDecode or available through an external mechanism) The picture that is associated with the access unit and applied to TargetOlsIdx is selected. Chatiming period SEI message (present in BitstreamToDecode or external mechanism) is selected, DecodingUnitHrdFlag is equal to 1, and dec When oding_unit_cpb_params_in_pic_timing_sei_flag is equal to 0, The decoding unit associated with the TargetOlsIdx is the decoding unit that is applied to the TargetOlsIdx. BitstreamInfo SEI message (present in BitstreamToDecode or by an external mechanism) available through the program) is selected.
[0172] Each suitability test involves a combination of one option in each of the above steps. When there are multiple options for a step, only one is required for any particular conformance test. Only this option is selected. All possible combinations of all steps are tested against the compatibility test. For each operation point under test, The number of exponential bitstream conformance tests is equal to n0*n1*n2*n3, where n0, n1, n2, and n The value of 3 is specified as follows: n1 is equal to hrd_cpb_cnt_minus1+1. n1 is the number of buffers The access unit in BitstreamToDecode associated with the ring period SEI message n2 is the number of bits in the bitstream. n2 is derived as follows: If the bitstream is a Type II bitstream, n0 is equal to 1. Otherwise (if BitstreamToDecode is a Type II bitstream, n0 is equal to 2 if the decoding_unicast stream is a If t_hrd_params_present_flag is equal to 0, then n3 is equal to 1. Otherwise, n3 is equal to 2. equal.
[0173] The HRD consists of a bitstream extractor (optionally present), a coded picture buffer, The decoder conceptually includes a CPB, a concurrent decoding process, and a sub-DPB for each layer. It includes a coded picture buffer (DPB), and output cropping. For stream conformance testing, the CPB size (in bits) is CpbSize[Htid][ScIdx]. , DPB parameters for each layer: max_dec_pic_buffering_minus1[Htid], max_num_reord er_pics[Htid] and MaxLatencyPictures[Htid] indicate whether the layer is an independent layer. and whether the layer is an output layer of the target OLS. These are found in or derived from the dpb_parameters() syntax structure used.
[0174] The HRD can operate as follows: HDR is initialized with decoding unit 0, CPB and Both the DPB and each sub-DPB of the DPB are set to empty (the fill amount of the sub-DPB for each sub-DPB is equal to 0). After initialization, the HRD receives the SEI message for the following buffering period. It cannot be reinitialized by a message according to its specified arrival schedule. The data associated with the decoding unit flowing into each CPB is The HSS is a time scheduler that is associated with each decoding unit. The data that has been removed is decoded instantaneously at the CPB removal time of the decoding unit. Each decoded picture is instantly decoded by the process. Each decoded picture is placed in the DPB. The decoded picture is not needed for inter prediction reference anymore. and are removed from the DPB when they are no longer needed for output.
[0175] An exemplary operation of the decoded picture buffer is as follows: Applied independently to each set of selected Decoded Picture Buffer (DPB) parameters The decoded picture buffer conceptually contains sub-DPBs, each of which may contain one Includes a picture storage buffer for storing decoded pictures of the two layers. Each picture storage buffer is marked for reference use or for later output. As described herein, the decoded pictures may include those that are retained for The process is applied sequentially, in order of increasing nuh_layer_id values of layers in the OLS, with the lowest These are applied independently for each layer, starting from the layer. When this process is applied, only the sub-DPB for the particular layer is affected. In their process description, a DPB refers to a sub-DPB for a particular layer, and the The layer that is currently being used is called the current layer.
[0176] In the operation of the output timing DPB, if PicOutputFlag in the same access unit is equal to 1, The new decoded pictures are sequentially ordered by the ascending nuh_layer_id value of the decoded pictures. The picture n and the current picture are output as the access point for a particular value of nuh_layer_id. Let n be the coded or decoded picture of access unit n, where n is a non-negative The removal of a picture from the DPB before the current picture is decoded is done as follows: Removal of a picture from the DPB before the current picture is decoded (but not after the current picture is decoded). The first decoding of access unit n (after parsing the slice header of the first slice of This occurs virtually instantaneously at the CPB removal time of the next It proceeds as follows.
[0177] The decoding process for reference picture list construction is invoked, and the reference pictures The decoding process for the marking is invoked. If the current AU is not AU0, When the loaded video sequence start (CVSS) AU, the next ordered step The variable NoOutputOfPriorPicsFlag is set for the decoder under test as follows: The pic_width_max_in_lum derived for any picture in the current AU. a_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_pl ane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8 or max_dec_pic_buffer The value of ing_minus1[Htid] is the sum of the pi c_width_in_luma_samples, pic_height_in_luma_samples, chroma_format_idc, separate _colour_plane_flag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or max_de If the value of c_pic_buffering_minus1[Htid] is different, NoOutputOfPriorPicsFlag is set to no_outp It shall be set to 1 by the decoder under test, regardless of the value of ut_of_prior_pics_flag. NoOutputOfPriorPicsFlag should be set to no_output_of_prior_pics_flag. may be preferable under these conditions, but the decoder under test In this case, you are allowed to set NoOutputOfPriorPicsFlag to 1. Otherwise, No OutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag. do.
[0178] The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied, and as a result, all Picture storage buffers are emptied without the output of the pictures they contain, and the DPB is filled. The amount is set equal to 0. For any picture k in the DPB, both of the following conditions If is true, then all such pictures k in the DPB are removed from the DPB. The picture k may be marked as unused for reference or the picture k may have a PictureOutput equal to 0. tFlag or the DPB output time is the first decoded time of the current picture n. m) and DpbOut putTime[k] is less than or equal to DuCpbRemovalTime[m]. For each picture removed from the DPB, The DPB fullness is decremented by one.
[0179] The operation of the output order DPB may be as follows: These processes The above-mentioned parameters may be applied independently to each set of decoded picture buffer (DPB) parameters. The decoded picture buffer conceptually contains sub-DPBs, each of which represents one layer. The picture storage buffer includes a picture storage buffer for storing the decoded pictures of the picture Each of the data storage buffers is marked for reference use or for future output. Contains the decoded picture stored in the DPB. A process is then invoked to output and remove the current decoded pixel. The process of marking and storing the kucha is called, followed by adding These processes are called bumping processes. It is applied to each layer independently, starting with the increasing order of the nuh_layer_id values of the layers in the OLS. When these processes are applied to a particular layer, Only the sub-DPBs that are
[0180] In the operation of the output order DPB, the same access unit is used as in the operation of the output timing DPB. Any decoded pictures in the unit with PicOutputFlag equal to 1 are also The nth picture and the current picture are output consecutively in ascending order of nuh_layer_id value. The coded or decoded picture of access unit n for a particular value of ayer_id. The output and removal of pictures from the DPB are It is explained as follows:
[0181] Before decoding the current picture (but not before the slice header of the first slice of the current picture) The output and removal of pictures from the DPB (after parsing the header) is done by the action that contains the current picture. Virtually instantaneous when the first decoding unit of the access unit is removed from the CPB. The decoding process for constructing a reference picture list is called The decoding process for the reference picture marking is then invoked. If the current AU is a CVSS AU that is not AU0, the following ordered steps are applied: The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows: pic_width_max_in_luma_samples, pic_h eight_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_de pth_luma_minus8, bit_depth_chroma_minus8 or max_dec_pic_buffering_minus1[Htid] The pic_width_in_luma_ samples, pic_height_in_luma_samples, chroma_format_idc, separate_colour_plane_fl ag, bit_depth_luma_minus8, bit_depth_chroma_minus8, or max_dec_pic_buffering_ If the value of minus1[Htid] is different, NoOutputOfPriorPicsFlag is set to no_output_of_prior_pics May be set to 1 by the decoder under test regardless of the value of _flag.
[0182] Set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag. Although it may be preferable under these conditions to use You are allowed to set NoOutputOfPriorPicsFlag to 1 before PriorPicsFlag is set equal to no_output_of_prior_pics_flag. The value of the variable NoOutputOfPriorPicsFlag derived for the target decoder is H Applies to RD. If NoOutputOfPriorPicsFlag is equal to 1, all pixels in the DPB The picture storage buffers are emptied without outputting the pictures they contain, and the DPB fullness is set equal to 0. Otherwise (if NoOutputOfPriorPicsFlag is equal to 0) all pictures marked as unused for reference, including pictures marked as not needed for output, and pictures marked as unused for reference. All picture storage buffers are emptied (without output) and all non-empty pictures in the DPB are The cache storage buffer is emptied by repeatedly calling bumping, and the DP The B sufficiency is set equal to zero.
[0183] Otherwise (if the current picture is not a CLVSS picture), All picture storage buffers, including pictures marked as unused for reference For each picture storage buffer emptied, a DPB is filled. The foot volume is decremented by 1. When one or more of the following conditions are true, For each additional Picture Storage Buffer that is emptied, a DPB refill is performed until neither of the previous two is true. The bumping process is then called again, with the foot volume further decremented by 1. The number of pictures in the DPB marked as needing reordering is greater than max_num_reorder_pics[Htid]. max_latency_increase_plus1[Htid] is not equal to 0 and the associated variable PicLa At least one output marked as required whose latencyCount is greater than or equal to MaxLatencyPictures[Htid] At most one picture is in the DPB. The number of pictures in the DPB is limited to max_dec_pic_buffering_min It is greater than or equal to us1[Htid]+1.
[0184] In one example, the additional bumping can occur as follows: Then, the last decoding unit of the access unit n that contains the current picture is removed from the CPB. This can happen virtually instantly when the current picture is removed. When the current picture has a byte, each picture in the DPB that is marked as needed for output and follows the current picture in output order is For each model, the associated variable PicLatencyCount is equal to PicLatencyCount+1. The following also applies: If the current decoded picture is Pic equal to 1, If the current decoded picture has trueOutputFlag, it is marked as needed for output. , the associated variable PicLatencyCount is set equal to 0. If not (the current decoded picture has PictureOutputFlag equal to 0), The current decoded picture is marked as not needed for output.
[0185] The bumping process is performed when one or more of the following conditions are true: It is called repeatedly until none of them are true. DPs marked as required in the output The number of pictures in B is greater than max_num_reorder_pics[Htid]. e_plus1[Htid] is not equal to 0 and the associated variable PicLatencyCount is equal to MaxLatencyP At least one picture marked as needed for output that is greater than or equal to ictures[Htid] is included in the DPB It's inside.
[0186] The bumping process involves the following ordered steps: One or more pictures are marked as required for output. Each of these pictures is selected as having the minimum value of nuh_ Cropped using an adaptive cropping window for the picture, ordered by layer_id. The image is then cropped, the cropped picture is output, and the picture is marked as not needed for output. One of the pictures that was cropped and output, including the picture marked as unused. Each picture storage buffer that was The amount is decremented by 1. Any two For two pictures picA and picB, if picA is output earlier than picB, The value of picOrderCntVal of picB is less than the value of PicOrderCntVal of picB.
[0187] An exemplary sub-bitstream extraction process is as follows: The inputs are the bitstream inBitstream, the target OLS index targetOlsIdx, and The output of this process is the sub-bitstream. The bitstream conformance requirement is that any input bitstream For a frame, the bitstream, an index into the list of OLSs specified by the VPS. and tIdTarget equal to any value in the range from 0 to 6. The output from this process is an output sub-bitstream that satisfies the following conditions: The output sub-bitstream may be a conforming bitstream. The system has a nuh_layer_id equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx]. The output sub-bitstream contains at least one VCL NAL unit that matches the tIdTarget. Contains at least one VCL NAL unit with equal TemporalId. A stream contains one or more coded slice NALs with TemporalId equal to 0. A coded slice NAL unit that contains a unit but has nuh_layer_id equal to 0 It is not necessary to have a
[0188] The output sub-bitstream OutBitstream is derived as follows: outBitstream is set to be identical to the bitstream inBitstream. All NAL units with TemporalId greater than et are removed from outBitstream. All NAL units with nuh_layer_id that are not included in the list LayerIdInOls[targetOlsIdx] Bits are removed from the outBitstream. Scalable Contains the nesting SEI message, and the value of i is in the range from 0 to nesting_num_olss_minus1. If there is no SEI NAL unit, NestingOlsIdx[i] is equal to targetOlsIdx, and all SEI NAL units are When targetOlsIdx is greater than 0, the buffering period is set to 0, and the 0 (Decoding unit information) or payloadType equal to 130 (Decoding unit information). All SEI NAL units, including non-scalable nested SEI messages, are sent to the outBitstream is removed from
[0189] An example scalable nesting SEI message syntax is as follows:
[0190] [Table 7]
[0191] Exemplary general SEI payload semantics are as follows: Non-scalableness On the applicable layer or OLS of the non-scalable network SEI message, the following applies: For the stream SEI message, payloadType is 0 (buffering period) or 1 (picture timing 130 (decoding unit information), or 130 (decoding unit information), non-scalable nested SE The I message applies only to the 0th OLS. For non-scalable nested SEI messages, When payloadType is equal to any value in VclAssociatedSeiList, the non-scalable The Bullet SEI message is not sent until a VCL NAL unit contains an SEI message. applies only to the layer with nuh_layer_id equal to the nuh_layer_id of the NAL unit .
[0192] The bitstream conformance requirements impose the following restrictions on the values of nuh_layer_id in SEI NAL units: A non-scalable nested SEI message may be applied with 0 (buffering pay equal to 1 (decoding period), 1 (picture timing), or 130 (decoding unit information) loadType, SEI NAL units that contain non-scalable nested SEI messages are vps_layer_id[0] shall have nuh_layer_id equal to vps_layer_id[0]. When a message has a payloadType equal to any value in VclAssociatedSeiList , a SEI NAL unit that contains a non-scalable nested SEI message is has a nuh_layer_id equal to the value of nuh_layer_id of the VCL NAL unit to which it is associated The SEI NAL unit containing the scalable nesting SEI message shall be The lowest value of nuh_layer_id of all layers to which the Bull Nest SEI message applies (scaling nesting_ols_flag in the nesting SEI message is equal to 0) or scalable The minimum value of nuh_layer_id of all layers in the OLS to which the re-nested SEI message applies ( nuh (when nesting_ols_flag in the scalable nesting SEI message is equal to 1) It should have a _layer_id.
[0193] Exemplary scalable nesting SEI message semantics are as follows: Scalable nesting SEI messages can be used to nest SEI messages with a specific OLS or Scalable nesting SEI messages provide a mechanism for associating a specific layer with a specific A message contains one or more SEI messages. Scalable nesting SEI messages The SEI message contained in is also called a scalable nested SEI message. The stream conformance requirement is to nest SEI messages within a scalable nesting SEI message. The following restrictions on inclusion may apply:
[0194] pa equal to 132 (decoded picture hash) or 133 (scalable nesting) A SEI message with yloadType is included in a scalable nesting SEI message. The scalable nesting SEI message may not have a buffering period. When including picture timing or decoding unit information SEI messages, The scalable nesting SEI message has the following values: 0 (buffering period), 1 (picture timing payloadType not equal to 130 (Decoding Unit Information) It does not include other SEI messages.
[0195] The bitstream conformance requirements apply to SEIs, including scalable nesting SEI messages. The following restrictions may apply to values of nal_unit_type of a NAL unit: The scalable nesting SEI message is 0 (buffering period), 1 (picture timing 130 (Decoding Unit Information), 145 (Subordinate RAP Indication), or 168 (Frame Filtering) When the scalability is achieved by including an SEI message with payloadType equal to The SEI NAL unit that contains the routing SEI message has a nal_unit_type equal to PREFIX_SEI_NUT. e should be present.
[0196] nesting_ols_flag equal to 1 indicates that the scalable nested SEI message applies to a particular OLS. nesting_ols_flag equal to 0 specifies that scalable nested SEI messages are to be Specifies that a particular message applies to a particular layer. The following restrictions may apply to the value of ing_ols_flag: The I message can be 0 (buffering period), 1 (picture timing), or 130 (decode When the SEI message contains a payloadType equal to the nesting_o The value of ls_flag shall be equal to 1. The scalable nesting SEI message shall When a message contains an SEI message with a payloadType equal to a value in sociatedSeiList, nesting The value of _ols_flag shall be equal to 0. nesting_num_olss_minus1+1 is the scalable nesting Specifies the number of OLSs to which the nesting SEI message applies. The value of nesting_num_olss_minus1 is 0 The range is from to TotalNumOlss-1.
[0197] nesting_ols_idx_delta_minus1[i] is the scalable delta when nesting_ols_flag is equal to 1. A variable Nestin that specifies the OLS index of the ith OLS to which the ReNest SEI message applies. Used to derive gOlsIdx[i]. The value of nesting_ols_idx_delta_minus1[i] must be between 0 and The variable NestingOlsIdx[i] should be in the range from 0 to TotalNumOlss-2. It can be derived. if(i==0) NestingOlsIdx[i]=nesting_ols_idx_delta_minus1[i] else NestingOlsIdx[i]=NestingOlsIdx[i-1]+nesting_ols_idx_delta_minus1[i]+1
[0198] nesting_all_layers_flag equal to 1 indicates that the scalable nested SEI message will be It applies to all layers with nuh_layer_id equal to or greater than the nuh_layer_id of the NAL unit. nesting_all_layers_flag equal to 0 specifies that scalable nesting SEI messages The message is sent to all layers with nuh_layer_id equal to or greater than the nuh_layer_id of the current SEI NAL unit. nesting_num_layers_minus1+ 1 specifies the number of layers to which the scalable nesting SEI message applies. The value of m_layers_minus1 ranges from 0 to vps_max_layers_minus1-GeneralLayerIdx[nuh_layer_id] and nuh_layer_id is in the range of nuh_layer_id of the current SEI NAL unit. nesting_layer_id[i] is the scalar when nesting_all_layers_flag is equal to 0. Specifies the nuh_layer_id value of the i-th layer to which the nested SEI message applies. The value of ing_layer_id[i] must be greater than nuh_layer_id, and nuh_layer_id must be the current SEIN This is the nuh_layer_id of the AL unit.
[0199] When nesting_ols_flag is equal to 0, the scalable nested SEI message is applied to the The NestingNumLayers variable specifies the number of layers, and the scalable nesting SEI message Specifies the list of layer nuh_layer_id values to apply to, in the range 0 to NestingNumLayers-1 The list NestingLayerId[i] for i in the layer can be derived as follows, where nuh_layer_id is currently The nuh_layer_id of the SEI NAL unit. if(nesting_all_layers_flag) { NestingNumLayers=vps_max_layers_minus1+1-GeneralLayerIdx[nuh_layer_id] for(i=0; i <NestingNumLayers; i++) NestingLayerId[i]=vps_layer_id[GeneralLayerIdx[nuh_layer_id]+i] } else { NestingNumLayers=nesting_num_layers_minus1+1 for(i=0; i <NestingNumLayers; i++) NestingLayerId[i]=(i==0) ? nuh_layer_id: nesting_layer_id[i] }
[0200] nesting_num_seis_minus1+1 specifies the number of scalable nested SEI messages. The value of sting_num_seis_minus1 must be in the range of 0 to 63. Let it be equal to 0.
[0201] 8 is a schematic diagram of an example video coding device 800. The device 800 is adapted to implement the disclosed examples / embodiments as described herein. The video coding device 800 is suitable for use in upstream and / or downstream a downstream port 820 including a transmitter and / or receiver for communicating data downstream; The video codec includes a video port 850, and / or a transceiver unit (Tx / Rx) 810. The operating device 800 may include a logic unit and / or a central processing unit for processing data. The video processing system includes a processor 830 including a CPU and a memory 832 for storing data. The coding device 800 may communicate with the device via an electrical, optical, or wireless communication network. 8. An electrical connector coupled to the upstream port 850 and / or the downstream port 820 for data communication. Optical-Electrical (OE) components, Electrical-Optical (EO) components, and / or wireless Video coding device 800 may also include a data communications component. Input and / or output (I / O) devices 860 for transmitting data and receiving data from a user The I / O device 860 may also include a display for displaying video data, a The I / O device 860 may include an output device such as a speaker for outputting data. input devices such as keyboards, mice, trackballs, and / or output devices The device may also include a corresponding interface for interactively manipulating the chair.
[0202] The processor 830 is implemented in hardware and software. 30 may include one or more CPU chips, cores (e.g., as a multi-core processor), filters, etc. Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), and Digital The processor 830 may be implemented as a digital signal processor (DSP). , Tx / Rx 810, upstream port 850, and memory 832. The processor 830 The coding module 814 includes a coding module 814 for implementing the methods 100, 900, and 1000. The present invention is directed to implementing the disclosed embodiments described herein, such as The multi-layer video sequence 600, the bitstream 700, and / or the sub-bitstreams The coding module 814 may employ the above-described Any other method / mechanism for coding may also be implemented. The deck system 200, the encoder 300, the decoder 400, and / or the HRD 500 may be implemented in a For example, the coding module 814 may include a simulcast layer and no VPS. Employed to encode, extract, and / or decode bitstreams containing Additionally, the coding module 814 may include, as part of the sub-bitstream extraction, Various syntax elements and syntax changes are made to avoid errors based on the reference to the VPS that is extracted. It may be employed to set and / or infer variables and / or The loading module 814 performs n when the layer including the CLVS is a simulcast layer. It is adopted to constrain uh_layer_id to be the same as nuh_layer_id in CLVS, Therefore, the coding module 814 can prevent erroneous extraction of the SPS. Implementing mechanisms to address one or more of the problems described above Thus, the coding module 814 may be configured to code the video data. It is intended to provide additional functionality and / or coding efficiency when coding Therefore, the coding module 814 may include a coding device 800. The functionality of the video coding device 800 is improved, and furthermore, the video coding technology is Furthermore, the coding module 814 is a video coding device. Alternatively, the coding module 814 may convert the memory 800 to a different state. 832 and executed by the processor 830 (e.g., The present invention may be implemented as a computer program product stored in a
[0203] Memory 832 can be a disk, tape drive, solid state drive, read-only drive, -Memory (ROM), Random Access Memory (RAM), Flash Memory, Ternary Associative Memory (TCA M), static random access memory (SRAM), or one or more memory types. The memory 832 stores programs when such programs are selected for execution. , to store instructions and data that are read during program execution. It can be used as a bar flow data storage device.
[0204] FIG. 9 shows the sub-bitstream extraction process 729 for the simulcast layer. To support holding SPS713, bitstreams such as Bitstream 700 9 is a flow chart of an example method 900 for encoding a multi-layer video sequence into a stream. The method 900 includes, when performing the method 100, the codec system 200, the encoder 300, and the 00, and / or an encoder such as video coding device 800. Additionally, the method 900 may operate on the HRD 500, thus allowing multi-layer video Conformance tests may be performed on the sequence 600 and / or its extracted layers. .
[0205] The method 900 includes an encoder receiving a video sequence and, based on, for example, user input, Deciding to encode the video sequence into a multi-layer bitstream In step 901, the encoder calculates the CLVS of the coded picture. The coded picture is encoded into a layer in the bitstream. For example, an encoder may store the pictures in a video sequence in a set of units. Encode the channel as one or more CLVS into one or more layers, and then layer / CLVS can be encoded into a multi-layer bitstream. A stream contains one or more layers and one or more CLVSs. A layer is a set of layers that are A set of VCL NAL units with a given header Id and associated non-VCL NAL units is In addition, CLVS may include a sequence of coded pictures that have the same nuh_layer_id value. As a concrete example, a VCL NAL unit is identified by nuh_layer_id. Specifically, a set of VCL NAL units may be associated with a layer. A set of elements are part of a layer when they all have a particular value of nuh_layer_id. , video data of encoded pictures, as well as coding data for such pictures. It includes a set of VCL NAL units that contain any parameter sets used to Such parameters may be obtained from the VPS, SPS, PPS, picture header, slice header, or or other parameter set or syntax structure. A plurality of layers may be output layers, and therefore one or more of the CLVSs contained in the layer may be The layer that is not the output layer is called a reference layer, and is used to reconstruct the output layer. However, such a supported layer / C LVS is not intended to be output by a decoder. In this way, the encoder It is possible to encode various layer / CLV combinations for transmission to the decoder when The layer / CLVS can be configured based on network conditions, hardware capabilities, and / or user preferences. This allows the decoder to obtain different representations of the video sequence depending on the In the present example, at least one of the layers may be transmitted using inter-layer prediction. Therefore, at least one CLVS is a simulcast layer that does not It is included in the cast layer.
[0206] In step 903, an encoder may encode the SPS into a bitstream. An SPS is referenced by one or more CLVSs. Specifically, an SPS is a record that contains a CLVS. When the layer does not use inter-layer prediction, it has a nuh_layer_id value equal to the nuh_layer_id value of CLVS. In addition, SPS uses nuh_lay of CLVS when layers including CLVS use inter-layer prediction. In this way, the SPS can identify any It has the same nuh_layer_id as the simulcast layer. In addition, the SPS is has a nuh_layer_id that is less than or equal to the nuh_layer_id of any non-simulcast layer. In this way, an SPS can be referenced by multiple CLVS / layers. However, an SPS is a sub-bit When stream extraction is applied to the CLVS / layer, which is a simulcast layer, the sub- are not removed by stream extraction.
[0207] The SPS will use the ID for the VPS referenced by the SPS, such as vps_independent_layer_flag. It can be encoded to include a sps_video_parameter_set_id that specifies the value. , sps_video_parameter_set_id is the SPS Specifies the value of vps_video_parameter_set_id for the VPS referenced by the SPS. Each codec that does not reference S and references SPS when sps_video_parameter_set_id is equal to 0 The VPS is not referenced at all when decoding a pre-mapped layer video sequence. Therefore, sps_video_parameter_set_id is Simulcast The SPS referenced by the coded layer video sequence contained in the It is set to 0 and / or inferred to be 0 when retrieved from the
[0208] The HRD is employed to perform conformance testing and therefore conformance to standards such as VVC. The HRD can be used to check the bitstream for sps_video_par When ameter_set_id is equal to 0, set GeneralLayerIdx[nuh_layer_id] to 0 It can be set and / or inferred to be equal to 0. id] is equal to the current layer index for the corresponding layer, and therefore the current It indicates the layer index of the current layer for the simulcast layer. The layer index is set to / inferred to 0. Additionally, the HRD is vps_independent_layer_flag[GeneralLayerIdx[nuh_ layer_id]] is equal to 1. Specifically, vps_independent_layer_flag is , specifies whether the corresponding layer uses inter-layer prediction. r_flag[i] is included in the VPS and is set to 0 to indicate that the i-th layer uses inter-layer prediction. or set to 1 to indicate that the i-th layer does not use inter-layer prediction. Therefore, vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] The current layer with index GeneralLayerIdx[nuh_layer_id] is an inter-layer predicted Therefore, a layer can use vps_independent_layer_flag[G Inter-layer prediction is not used when the current layer is When a layer is a simullayer, the VPS is omitted and no inter-layer prediction is employed. Therefore, the inference for a value of 1 when sps_video_parameter_set_id is equal to 0 is simulcast The cast layer is extracted during bitstream extraction for the simulcast layer. ,We make sure that it works properly in HRD and during decoding,while avoiding references to VPS. Therefore, this reasoning is based on the assumption that the VPS is eliminated for the simulcast layer. to prevent sub-bitstream extraction errors that would otherwise occur when Then, HRD will set SPS, sps_video_parameter_set_id, GeneralLayerIdx[nuh_la yer_id], and / or vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] Based on the coded picture from the VCL NAL unit in the simulcast layer / CLVS Therefore, the HRD can decode the multi-level picture. Simulcast layer of the ear bitstream is VP to simulcast layer Fits into the bitstream without unexpected errors caused by the omission of S It is possible to verify whether
[0209] In step 905, the encoder prepares a video stream for communication to the decoder when requested. The encoder can store the sub-bitstreams. to get the simulcast layer and then bitstream / sub when desired. The bitstream can also be transmitted to an encoder.
[0210] FIG. 10 illustrates a block diagram of a multi-layer bitstream extracted from a multi-layer bitstream, such as bitstream 700. From a bitstream, such as a sub-bitstream 701 that contains a simulcast layer 1 is a flowchart of an exemplary method 1000 for decoding a video sequence from SPS713; is retained during the sub-bitstream extraction process 729. When executed, the codec system 200, the decoder 400, and / or the video codec The method 1000 may be employed by a decoder, such as the HRD 500. A multi-layer video sequence 600, which has been checked for conformance by the HRD, or may be employed on the extracted layer.
[0211] The method 1000 may be implemented by a decoder that generates a multi-layer bitstream, for example as a result of the method 900. Contains the coded video sequence of the simulcast layer extracted from the video stream. In step 1001, the decoder starts receiving the bitstream. The layer receives a bitstream containing the CLVS in the layer. The layer is Simulcast extracted from multi-layer bitstream by content server A simulcast layer may be a set of coded pictures. For example, the bitstream contains a CLVS that contains each coded picture. One or more associated with a simulcast layer as identified by layer_id. A coded picture is a set of VCL NAL units. The layer must be able to read VCL NAL units and associated non-VCL NAL units that have the same layer Id. For example, a layer may include a set of video data of an encoded picture. and any parameter set used to code such a picture. Therefore, the set of VCL NAL units Denotes that a set of VCL NAL units all have a particular value of nuh_layer_id for a layer. In addition, CLVS will not use the nuh_layer_id value for all coded pictures that have the same nuh_layer_id value. The simulcast layer is also the output layer and employs inter-layer prediction. Therefore, CLVS also does not employ inter-layer prediction.
[0212] The bitstream also contains the SPS referenced by the CLVS. The SPS is the layer that contains the CLVS. When inter-layer prediction is not used, it has a nuh_layer_id value equal to the CLVS nuh_layer_id value. Additionally, SPS uses CLVS's nuh_layer_id when layers including CLVS use inter-layer prediction. In this way, the SPS can It has the same nuh_layer_id as the multicast layer. In addition, the SPS is has a nuh_layer_id that is less than or equal to the nuh_layer_id of the non-simulcast layer. An SPS can be referenced by multiple CLVS / layers. However, an SPS is a sub-bitstream. When the stream extraction is applied to the CLVS / layer, which is a simulcast layer, the encoder / It is not removed by the sub-bitstream extraction at the content server.
[0213] The SPS will use the ID for the VPS referenced by the SPS, such as vps_independent_layer_flag. It contains the sps_video_parameter_set_id that specifies the value. t_id is the video ID of the VPS referenced by the SPS when sps_video_parameter_set_id is greater than 0. Specify the value of ps_video_parameter_set_id. Furthermore, SPS does not refer to VPS, and sps_video_p Each coded layer video sequence that references an SPS when parameter_set_id is equal to 0 The VPS is not referenced at all when decoding sequences. er_set_id is the codec that contains the sps_video_parameter_set_id in the simulcast layer. Set to 0 when retrieved from the SPS referenced by the rendered layer video sequence. and / or inferred to be 0. When the system contains only simulcast layers including CLVS that does not employ inter-layer prediction, VP Does not contain S.
[0214] In step 1003, the decoder decodes the picture from the CLVS based on the SPS to obtain a decoded picture. For example, the decoder can generate a pre-coded picture using the sps_video_parameter When r_set_id is equal to 0, set GeneralLayerIdx[nuh_layer_id] equal to 0 and / or can be inferred to be equal to 0. is equal to the current layer index for the corresponding layer, and therefore the current layer index. It indicates the layer index and hence the current layer index for the simulcast layer. The index is set to 0 / inferred to 0. In addition, the decoder vps_independent_layer_flag[GeneralLayerIdx[nuh when _parameter_set_id is equal to 0 _layer_id]] is equal to 1. Specifically, vps_independent_layer_flag is , specifies whether the corresponding layer uses inter-layer prediction. r_flag[i] is included in the VPS and is set to 0 to indicate that the i-th layer uses inter-layer prediction. or set to 1 to indicate that the i-th layer does not use inter-layer prediction. Therefore, vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] The current layer with index GeneralLayerIdx[nuh_layer_id] is an inter-layer predicted Therefore, a layer can use vps_independent_layer_flag[G Inter-layer prediction is not used when the current layer is When the layer is simullayer, the VPS is not included in the received bitstream and the layer Inter prediction is not employed. Therefore, the value of 1 when sps_video_parameter_set_id is equal to 0 The inference of the value is done by the simulcast layer by This works fine during decoding, avoiding references to the VPS that are extracted during frame extraction. Therefore, this reasoning ensures that the VPS is not excluded from the simulcast layer. The sub-bitstream extraction errors that would otherwise occur when The decoder then checks the SPS, sps_video_parameter_set_id, GeneralLayer rIdx[nuh_layer_id], and / or vps_independent_layer_flag[GeneralLayerIdx[nuh_ Coding from VCL NAL units in the simulcast layer / CLVS based on the layer_id The decoder may then decode the coded picture to generate a decoded picture. , in step 1005, the video signal is decoded for display as part of the decoded video sequence. The completed picture can be transferred.
[0215] FIG. 11 shows the sub-bitstream extraction process 729 for the simulcast layer. Supports bit-for-bit multi-layer video sequences to hold SPS713 FIG. 11 is a schematic diagram of an exemplary system 1100 for coding into the stream 700. The system 1100 includes a codec system 200, an encoder 300, a decoder 400, and / or a video The encoding device 800 may be implemented by an encoder and a decoder. In addition, the system 1100 employs the HRD 500 to generate a multi-layer video sequence 600, 7. Performing conformance tests on stream 700 and / or sub-bitstream 701 Additionally, the system 1100 may include the methods 100, 900, and / or 1000. This can be adopted when implementing
[0216] The system 1100 includes a video encoder 1102. The video encoder 1102 converts the CLVS into a video signal. an encoding module 1105 for encoding the layer in the bit stream; The encoding module 1105 further encodes the CLVS-encoded sigma-based ... and encoding an SPS referenced by the layer, the SPS being a layer that performs inter-layer prediction. When unused it is constrained to have a nuh_layer_id value equal to the CLVS nuh_layer_id value. The video encoder 1102 stores the bitstream for communication to a decoder. The video encoder 1102 further includes a storage module 1106. The video encoder further comprises a transmission module 1107 for transmitting the bitstream. The coder 1102 may be further configured to perform any of the steps of the method 900.
[0217] The system 1100 also includes a video decoder 1110. The video decoder 1110 can be used to decode CLs in a layer. Receive module for receiving bitstreams containing SPS referenced by VS and CLVS When a layer does not use inter-layer prediction, the SPS uses the nuh_layer_id of the CLVS. The video decoder 1110 receives the code from the CLVS based on the SPS. A decoding program for decoding the loaded picture to generate a decoded picture. The video decoder 1110 further comprises a decoding module 1113. A transfer module 11 that transfers the decoded picture for display as part of the sequence. 15. The video decoder 1110 may be configured to perform any of the steps of the method 1000. It may be further configured as follows.
[0218] The first component is the line between the first component and the second component, When there is no intervening component, except for a race or another medium, The first component is directly coupled to the second component. There are no intervening components other than wires, traces, or another medium between the components. At some point, it is indirectly connected to a second component. The phrase and its variations refer to both direct and indirect binding. In addition, the word "about" means a range including ±10% of the subsequent numerical value, unless otherwise specified. It means the surrounding area.
[0219] The steps of the exemplary methods described herein are not necessarily illustrated. It should also be understood that the steps of such methods need not be performed in order. The above should be understood to be merely exemplary. In methods that conform to the principles of the present invention, additional steps may be included in such methods, and certain steps may be included in the methods that conform to the principles of the present invention. Groups may be omitted or combined.
[0220] Although several embodiments are provided in this disclosure, the disclosed systems and The method may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. It can be understood that the above-mentioned example is illustrative and not restrictive. It should be understood, however, that the present invention is not to be limited to the details set forth in the specification. For example, the various elements or components may be combined in another system or or may be combined, or some features may be omitted or implemented. Not obtained.
[0221] In addition, although described and illustrated in various embodiments as being discrete or separate, The techniques, systems, subsystems, and methods described herein are not intended to be limiting without departing from the scope of the present disclosure. Combined or integrated with other systems, components, techniques, or methods Other examples of modifications, substitutions, and alterations may be ascertained by one of ordinary skill in the art and may be incorporated herein by reference. This can be done without departing from the spirit and scope of the disclosure herein. [Explanation of symbols]
[0222] 100 How it works 200 Coding and Decoding (Codec) Systems 201 Segmented Video Signal 211 General Coder Control Component 213 Transform Scaling and Quantization Components 215 In-Picture Estimation Component 217 Intra-Picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling and Inverse Transformation Components 231 Header Formatting and Context-Adaptive Binary Arithmetic Coding (CABAC) )component 300 Video Encoder 301 Segmented Video Signal 313 Transform and Quantize Components 317 In-Picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Component 400 Video Decoder 417 In-Picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Component 500HRD 541 Virtual Stream Scheduler (HSS) 543 CPB 545 Decoding Process Components 547 DPB 549 Output Cropping Component 551 Bitstream 553 Decoding Unit (DU) 555 Decoded DUs 556 Reference Pictures 557 Pictures 559 Output Cropped Picture 600 Multi-layer Video Sequences 611, 612, 613, 614 Pictures 615, 616, 617, 618 Pictures 621 Inter-layer Prediction 623 Inter Prediction 631 Layer N 632 Layer N+1 700 bitstream 701 Sub-Bitstream 711 VPS 713 SPS 715 Picture Parameter Set (PPS) 717 Slice Header 720 Image data 723, 724 Layer 725, 726 Pictures Set of 726 pictures 727, 728 slices 729 Sub-Bitstream Extraction Process 731 sps_video_parameter_set_id 732 nuh_layer_id 733 VPS independent layer flag (vps_independent_layer_flag) 735 vps_video_parameter_set_id 741 VCL NAL Units 742 non-VCL NAL units 743 Coded Layered Video Sequence (CLVS) 744 CLVS 800 Video Coding Device 810 Transceiver Unit (Tx / Rx) 814 Coding Module 820 Downstream Port 830 Processor 832 Memory 850 Upstream Port 860 Input and / or Output (I / O) Devices 900 ways 1000 ways 1100 System 1102 Video Encoder 1105 Encoding Module 1106 Memory Module 1107 Transmission Module 1110 Video Decoder 1111 Receiver Module 1113 Decoding Module
Claims
1. A method implemented by an encoder, comprising: encoding, by the encoder, a coded layered video sequence (CLVS) for the layer in a bitstream; encoding, by the encoder, into the bitstream a sequence parameter set (SPS) referenced by the CLVS, the SPS being constrained to have a network abstraction layer (NAL) unit header layer identifier (nuh_layer_id) value equal to the CLVS nuh_layer_id value when the layer does not use inter-layer prediction, the SPS including a video parameter set (VPS) identifier (sps_video_parameter_set_id) that specifies an identifier (ID) value for a video parameter set (VPS) referenced by the SPS, and a general layer index (GeneralLayerIdx[nuh_layer_id]) corresponding to nuh_layer_id is set to 0 when the sps_video_parameter_set_id is equal to 0; storing, by the encoder, the bitstream for communication to a decoder.
2. The method of claim 1 , wherein GeneralLayerIdx[nuh_layer_id] is equal to the current layer index.
3. The method of claim 1 or 2, wherein the VPS referenced by the SPS has a VPS independent layer flag (vps_independent_layer_flag) that specifies whether the corresponding layer uses inter-layer prediction.
4. The method of claim 3, wherein the layer does not use inter-layer prediction when the vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1.
5. The method of claim 4 , wherein when sps_video_parameter_set_id is equal to 0, the vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.
6. The method of claim 1 , wherein the SPS does not reference a VPS when the sps_video_parameter_set_id is equal to 0.
7. The method of claim 1 , wherein the CLVS is a sequence of coded pictures having the same nuh_layer_id value.
8. Encoding a coded layered video sequence (CLVS) for the layer in the bitstream; encoding means for encoding, into the bitstream, a sequence parameter set (SPS) referenced by the CLVS, the SPS being constrained to have a network abstraction layer (NAL) unit header layer identifier (nuh_layer_id) value equal to the CLVS nuh_layer_id value when the layer does not use inter-layer prediction, the SPS including a video parameter set (VPS) identifier (sps_video_parameter_set_id) that specifies an identifier (ID) value for a video parameter set (VPS) referenced by the SPS, and a general layer index (GeneralLayerIdx[nuh_layer_id]) corresponding to nuh_layer_id is set to 0 when the sps_video_parameter_set_id is equal to 0; storage means for storing said bitstream for communication to a decoder; An encoder comprising:
9. An encoder according to claim 8, further configured to perform the method according to any one of claims 1 to 7.
10. A device for storing a bitstream, comprising: at least one receiver configured to receive the bitstream; at least one memory configured to store the bitstream; The bitstream comprises: a coded layer video sequence (CLVS) for a layer and a sequence parameter set (SPS) referenced by the CLVS, the SPS having a network abstraction layer (NAL) unit header layer identifier (nuh_layer_id) value equal to a nuh_layer_id value of the CLVS when the layer does not use inter-layer prediction, the SPS including an SPS VPS identifier (sps_video_parameter_set_id) that specifies an identifier (ID) value for a video parameter set (VPS) referenced by the SPS; A device, wherein when the sps_video_parameter_set_id is equal to 0, the general layer index (GeneralLayerIdx[nuh_layer_id]) corresponding to the nuh_layer_id is equal to 0.
11. A method for storing a bitstream, comprising the steps of: receiving said bitstream by at least one receiver; storing the bitstream by at least one memory; The bitstream comprises: a coded layer video sequence (CLVS) for a layer and a sequence parameter set (SPS) referenced by the CLVS, the SPS having a network abstraction layer (NAL) unit header layer identifier (nuh_layer_id) value equal to a nuh_layer_id value of the CLVS when the layer does not use inter-layer prediction, the SPS including an SPS VPS identifier (sps_video_parameter_set_id) that specifies an identifier (ID) value for a video parameter set (VPS) referenced by the SPS; A method in which when the sps_video_parameter_set_id is equal to 0, the general layer index (GeneralLayerIdx[nuh_layer_id]) corresponding to the nuh_layer_id is equal to 0.