Sub-bitstream extraction
By introducing image sequence counting and sub-image sub-bitstream extraction techniques from the VVC standard, the problems of low bitstream consistency and decoding efficiency in multi-layer video and 360° immersive media are solved, achieving efficient image management and random access, and improving encoding and decoding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2021-09-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing video codec standards struggle to efficiently manage image sequence counting and supplementary enhancement information when handling multi-layered video and 360° immersive media, resulting in low bitstream consistency and decoding efficiency.
It adopts the Picture Order Count (POC) and sub-picture sub-bitstream extraction technology from the VVC standard, and notifies the most significant bit and least significant bit of the POC through signaling. It allows mixing different types of pictures and supports the processing of mixed NAL unit types and scalable nested SEI messages within sub-pictures, ensuring bitstream consistency and decoding efficiency.
It improves decoding efficiency and bitstream consistency for multi-layer video and 360° immersive media, supports viewport adaptation and efficient random access, and reduces transmission bit rate.
Smart Images

Figure CN116325753B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is filed to promptly claim priority and benefit from International Patent Application No. PCT / CN2020 / 117596, filed on September 25, 2020. The entire disclosure of the aforementioned application is incorporated herein by reference as part of the disclosure. Technical Field
[0003] This patent document relates to the generation, storage, and consumption of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to process video or image representations according to file formats.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream comprising multiple layers of video, wherein the bitstream includes multiple Supplemental Enhancement Information (SEI) messages associated with a specific output layer set (OLS) or a specific layer's access unit (AU) or decoding unit (DU), including multiple SEI messages of a message type different from a scalable nested type based on a format rule, and wherein the format rule specifies that each of the multiple SEI messages has the same SEI payload content because it is associated with a specific OLS or a specific layer's AU or DU.
[0007] In another example, a different video processing method is disclosed. This method includes: performing a conversion between a video including the current block and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies constraints on the bitstream due to the exclusion of syntax fields from the bitstream, the syntax fields indicating the sequence count of images associated with the current picture including the current block or the currently accessed unit including the current block.
[0008] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a video bitstream, wherein the bitstream conforms to a sub-bitstream extraction process sequence defined by rules, which specify that (a) the sub-bitstream extraction process excludes the removal of all Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) units that satisfy certain conditions from the output bitstream, or (b) the sub-bitstream extraction process sequence includes replacing a parameter set with a replacement parameter set before removing SEINAL units from the output bitstream.
[0009] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0010] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0011] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0012] In another example, a computer-readable medium on which a bitstream is stored is disclosed. The bitstream is generated or processed using the methods described in this document.
[0013] These and other features are described throughout this document. Attached Figure Description
[0014] Figure 1 The image is displayed as being divided into 18 slices, 24 strips, and 24 sub-images.
[0015] Figure 2 This demonstrates a typical viewport-dependent 360° video transmission scheme based on sub-pictures.
[0016] Figure 3 This shows an example of extracting a sub-image from a bitstream containing two sub-images and four stripes.
[0017] Figure 4 This is a block diagram of an example video processing system.
[0018] Figure 5 This is a block diagram of a video processing device.
[0019] Figure 6 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0020] Figure 7 This is a block diagram illustrating an encoder according to some embodiments of the disclosed technology.
[0021] Figure 8 This is a block diagram illustrating a decoder according to some embodiments of the disclosed technology.
[0022] Figure 9-11 This is a flowchart of an example method for video processing according to some embodiments of the disclosed technology. Detailed Implementation
[0023] The use of chapter headings in this document is for ease of understanding and does not limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0024] 1. Preliminary Discussion
[0025] This document relates to video codec techniques. Specifically, it covers Picture Order Counting (POC), Supplemental Enhancement Information (SEI), and subpicture sub-bitstream extraction. These ideas can be applied individually or in various combinations to any standard or non-standard video codec, such as the recently completed Multi-Functional Video Codec (VVC).
[0026] 2. Abbreviations
[0027] ACT Adaptive Color Transformation
[0028] ALF Adaptive Loop Filter
[0029] AMVR Adaptive Motion Vector Accuracy
[0030] APS Adaptive Parameter Set
[0031] AU Access Unit
[0032] AUD Access Unit Delimiter
[0033] AVC Advanced Video Codec (Rec.ITU-T H.264|ISO / IEC 14496-10)
[0034] B Two-way prediction
[0035] BCW CU-level bidirectional weighted prediction
[0036] BDOF bidirectional optical flow
[0037] BDPCM (Block-based Incremental Pulse Codec Modulation)
[0038] BP buffer period
[0039] CABAC Context-Based Adaptive Binary Arithmetic Encoding and Decoding
[0040] CB codec block
[0041] CBR Constant Bit Rate
[0042] CCALF Cross-Component Adaptive Loop Filter
[0043] CLVS codec layer video sequence
[0044] CLVSS codec layer video sequence start
[0045] CPB image buffer
[0046] CRA Clean Random Access
[0047] CRC Cyclic Redundancy Check
[0048] CTB codec tree block
[0049] CTU (Codec Tree Unit)
[0050] CU encoding / decoding unit
[0051] CVS codec video sequence
[0052] CVSS codec video sequence begins
[0053] DPB Decoding Image Buffer
[0054] DCI decoding capability information
[0055] DRAP relies on random access points
[0056] DU decoding unit
[0057] DUI Decoding Unit Information
[0058] EG index Golomb
[0059] EGk k-order exponent Golomb
[0060] EOB bitstream end
[0061] EOS sequence ends
[0062] FD padding data
[0063] FIFO (First In First Out)
[0064] FL fixed length
[0065] GBR Green, Blue and Red
[0066] GCI General Constraint Information
[0067] GDR Gradual Decoding and Refresh
[0068] GPM Geometric Segmentation Mode
[0069] HEVC High-Efficiency Video Codec (Rec.ITU-T H.265|ISO / IEC 23008-2)
[0070] HRD Assumption Reference Decoder
[0071] HSS Assumption Stream Scheduler
[0072] I within the frame
[0073] IBC Intra-block Copy
[0074] IDR Instant Decoding and Refresh
[0075] ILRP interlayer reference image
[0076] IRAP Intra-Frame Random Access Point
[0077] LFNST Low-Frequency Inseparable Transform
[0078] LPS Minimum Possible Symbols
[0079] LSB (Least Significant Bit)
[0080] LTRP Long-Term Reference Image
[0081] LMCS Luminance Mapping and Chroma Scaling
[0082] MIP (Matrix-Based Intra-Frame Prediction)
[0083] MPS Most Probable Symbol
[0084] MSB Most significant bit
[0085] MTS Multiple Transformation Selection
[0086] MVP Motion Vector Prediction
[0087] NAL Network Abstraction Layer
[0088] OLS Output Layer Set
[0089] OP operation point
[0090] OPI Operation Point Information
[0091] P prediction
[0092] PH image header
[0093] POC Image Sequential Counting
[0094] PPS Image Parameter Set
[0095] PROF refines predictions using optical flow.
[0096] PT Image Time Sequence
[0097] PU Image Unit
[0098] QP Quantization Parameters
[0099] RADL Random Access Decodable Preamble (Image)
[0100] RASL Random Access Skip Preamble (Image)
[0101] RBSP raw byte sequence payload
[0102] RGB red, green and blue
[0103] RPL Reference Image List
[0104] SAO Sample Adaptive Offset
[0105] SAR sample aspect ratio
[0106] SEI Supplemental Enhancement Information
[0107] SH strip header
[0108] SLI sub-image level information
[0109] SODB data bit string
[0110] SPS Sequence Parameter Set
[0111] STRP Short-Term Reference Image
[0112] STSA Stepwise Temporal Sublayer Access
[0113] TR Cutoff Rice
[0114] TU Transformer
[0115] VBR Variable Bit Rate
[0116] VCL (Video Codec Layer)
[0117] VPS Video Parameter Set
[0118] VSEI General Supplemental Enhancement Information (Rec.ITU-T H.274|ISO / IEC 23002-7)
[0119] VUI Video Availability Information
[0120] VVC (Video Codec) is a multi-functional video codec (Rec.ITU-T H.266|ISO / IEC 23090-3).
[0121] 3. Overview of Video Encoding and Decoding
[0122] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed the MPEG-1 and MPEG-4 Visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). When the Multifunctional Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce the bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which concluded on July 1, 2020.
[0123] The Multi-Functional Video Coding (VVC) standard (ITU-TH.266|ISO / IEC23090-3) and the related Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-TH.274|ISO / IEC23002-7) are designed for the widest range of applications, including traditional applications such as television broadcasting, video conferencing, or playback of stored media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple codec video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360° immersive media.
[0124] 3.1. Image Sequence Counting in HEVC and VVC (POC)
[0125] In HEVC and VVC, POC is basically used as an image ID to identify images in many parts of the decoding process, including DPB management, a part of which is reference image management.
[0126] For the newly introduced PH, in VVC, the information of the least significant bit (LSB) of the POC, which is used to derive the POC value and has the same value for all stripes of the picture, is signaled in the PH, as opposed to HEVC where it is signaled in the SH. VVC also allows the signaling notification of the cyclic value of the most significant bit (MSB) of the POC in the PH, so that the POC value can be derived without tracking the POC MSB, depending on the POC information of the earlier decoded picture. For example, this allows mixing IRAP and non-IRAP pictures within an AU in a multi-layer bitstream. Another difference between POC signaling notification in HEVC and VVC is that in HEVC, the POC LSB is not signaled for IDR pictures, which showed some disadvantages during the later development of multi-layer extensions of HEVC to enable mixing IDR and non-IDR pictures within an AU. Therefore, in VVC, for each picture including the IDR picture, the POC LSB information is signaled. The signaling notification of POC LSB information for IDR images also makes it easier to merge IDR and non-IDR images from different bitstreams into a single image, because otherwise, processing the POC LSB in the merged image would require some complex design.
[0127] The decoding process for VVC's Proof of Concept (POC) is defined as follows:
[0128] 8.3.1 Decoding process of image sequential counting
[0129] The output of this process is PicOrderCntVal, which is the image order count of the current image.
[0130] Each encoded / decoded image is associated with an image order count variable, denoted as PicOrderCntVal.
[0131] Assume that the variable currLayerIdx is set to be equal to GeneralLayerIdx[nuh_layer_id].
[0132] PicOrderCntVal is exported as follows:
[0133] – If vps_independent_layer_flag[currLayerIdx] equals 0 and there is an image picA in the current AU, its nuh_layer_id equals layerIdA such that GeneralLayerIdx[layerIdA] is in the list ReferenceLayerIdx[currLayerIdx], PicOrderCntVal is exported as equal to PicOrderCntValpicA, and the value of ph_pic_order_cnt_lsb should be the same in all VCLNAL cells of the current AU.
[0134] —Otherwise, the PicOrderCntVal of the current image is derived in accordance with the rest of this subsection.
[0135] When ph_poc_msb_cycle_val does not exist and the current image is not a CLVSS image, the variables prevPicOrderCntLsb and prevPicOrderCntMsb are exported as follows:
[0136] – Let prevTid0Pic be the previous image in the decoding order, whose nuh_layer_id is equal to the nuh_layer_id of the current image, whose TemporalId and ph_non_ref_pic_flag are both equal to 0, and which is not a RASL or RADL image.
[0137] Note 1 – In a sub-bitstream consisting only of intra-pictures, extracted from a single-layer bitstream and used for intra-picture-only effects playback, prevTid0Pic is the preceding intra-picture in decoding order. To ensure correct POC export, the encoder can choose to include ph_poc_msb_cycle_val for intra-pictures, or set the value of sps_log2_max_pic_order_cnt_lsb_minus4 large enough that the POC difference between the current picture and prevTid0Pic is less than MaxPicOrderCntLsb / 2.
[0138] Note 2 – When vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0 for any i value in the layer index and j is equal to the layer index of the current layer, the prevTid0Pic in some OLS extracted sub-bitstreams will be the previous IRAP or GDR picture in the current layer whose ph_recovery_poc_cnt is equal to 0 in the decoding order. To ensure that such sub-bitstreams are consistent bitstreams, the codec can choose to include ph_poc_msb_cycle_val for each IRAP or GDR picture, where ph_recovery_poc_cnt is equal to 0, or set the value of sps_log2_max_pic_order_cnt_lsb_minus4 large enough that the POC difference between the current picture and prevTid0Pic is less than MaxPicOrderCntLsb / 2.
[0139] The variable prevPicOrderCntLsb is set to equal ph_pic_order_cnt_lsb of prevTid0Pic.
[0140] The variable `prevPicOrderCntMsb` is set to be equal to the `PicOrderCntMsb` of `prevTid0Pic`. The variable `PicOrderCntMsb` for the current image is exported as follows:
[0141] – If ph_poc_msb_cycle_val exists, then set PicOrderCntMsb to be equal to ph_poc_msb_cycle_val * MaxPicOrderCntLsb.
[0142] Otherwise (ph_poc_msb_cycle_val does not exist), if the current image is a CLVSS image, PicOrderCntMsb is set to 0.
[0143] Otherwise, PicOrderCntMsb will export as follows:
[0144]
[0145] PicOrderCntVal exports the following:
[0146] PicOrderCntVal = PicOrderCntMsb + ph_pic_order_cnt_lsb (197)
[0147] Note 3 – All CLVSS images where ph_poc_msb_cycle_val does not exist have a PicOrderCntVal equal to ph_pic_order_cnt_lsb, because for those images PicOrderCntMsb is set to 0.
[0148] The value of PicOrderCntVal should be in the range of -2. 31 to 2 31 The range is -1 (inclusive).
[0149] In a CVS, the PicOrderCntVal values of any two codec images with the same nuh_layer_id value should not be the same.
[0150] All images in any given AU should have the same PicOrderCntVal value.
[0151] The function PicOrderCnt(picX) is defined as follows:
[0152] PicOrderCnt( picX ) = PicOrderCntVal of the picture picX (198)
[0153] The function DiffPicOrderCnt(picA, picB) is defined as follows:
[0154] DiffPicOrderCnt( picA, picB ) = PicOrderCnt( picA ) - PicOrderCnt(picB ) (199)
[0155] The bitstream should not contain values that cause the DiffPicOrderCnt(picA,picB) used during decoding to be outside the range of -2. 15 to 2 15 Data within the range of -1 (inclusive).
[0156] Note 4 – Let X be the current image, and Y and Z be two other images in the same CVS. When both DiffPicOrderCnt(X,Y) and DiffPicOrderCnt(X,Z) are positive or both are negative, Y and Z are considered to be in the same output order direction as X.
[0157] 3.2.VUI and SEI messages
[0158] The VUI is a syntax structure sent as part of the SPS (and possibly in the HEVC VPS). The VUI carries information that does not affect the standard decoding process, but may be important for the correct rendering of the encoded and decoded video.
[0159] SEI assists in processes related to decoding, display, or other purposes. Like VUI, SEI does not affect the specification decoding process. SEI is carried in SEI messages. Decoder support for SEI messages is optional. However, SEI messages do affect bitstream consistency (e.g., if the syntax of SEI messages in the bitstream does not conform to the specification, the bitstream is not conforming to the specification), and some SEI messages are required in the HRD specification.
[0160] The VUI syntax structures and most SEI messages used with VVC are not specified in the VVC specification, but rather in the VSEI specification. The SEI messages required for HRD conformance testing are specified in the VVC specification. VVC v1 defines five SEI messages related to HRD conformance testing, and VSEI v1 specifies 20 additional SEI messages. The SEI messages carried in the VSEI specification do not directly affect the behavior of the conformance decoder and are defined to allow them to be used in a codec-agnostic manner, thus allowing VSEI to be used with other video codec standards besides VVC in the future. The VSEI specification does not specifically mention VVC syntax element names, but rather refers to variables whose values are set in the VVC specification.
[0161] Compared to HEVC, VVC's VUI syntax structure focuses only on information related to the correct rendering of the image and does not include any timing information or bitstream limitation indications. In VVC, the VUI is signaled in the SPS, which includes a length field preceding the VUI syntax structure to signal the length of the VUI payload (in bytes). This allows the decoder to easily skip information and, more importantly, allows for convenient future VUI syntax extensions by adding new syntax elements directly to the end of the VUI syntax structure in a manner similar to SEI message syntax scaling.
[0162] The VUI syntax structure contains the following information:
[0163] • The content is interwoven or progressive;
[0164] • Does the content include frame-encapsulated stereoscopic video or projected omnidirectional video?
[0165] • Aspect ratio of the sample points;
[0166] • Is the content suitable for overscan display?
[0167] • Color description, including color primary colors, matrix, and transmission characteristics, is particularly important for signaling communication of Ultra High Definition (UHD) and High Definition (HD) color spaces as well as High Dynamic Range (HDR);
[0168] • Chromaticity position relative to luminance (signaling notification of progressive content has been clarified compared to HEVC).
[0169] When SPS does not contain any VUI, the information is considered unspecified, and if the content of the bitstream is intended for presentation on a display, the information must be communicated externally or specified by the application.
[0170] Table 1 lists all SEI messages specified for VVC v1, along with the specification containing their syntax and semantics. Of the 20 SEI messages specified in the VSEI specification, many are inherited from HEVC (e.g., the padding payload and two user data SEI messages). Some SEI messages are essential for the proper processing or rendering of encoded video content. This is the case, for example, for SEI messages related to primary display color magnitude, content light level information, or alternative transport characteristics, which are particularly relevant to HDR content. Other examples include SEI messages for isorectangular projection, spherical rotation, region packing, or omnidirectional viewports, which are related to signaling notification and processing of 360° video content.
[0171] Table 1: SEI Message List in VVC v1
[0172]
[0173]
[0174] The new SEI messages specified for VVC v1 include frame field information SEI messages, sample aspect ratio information SEI messages, and sub-picture level information SEI messages.
[0175] The Frame Field Information (SEI) message contains information indicating how the associated picture should be displayed (e.g., field parity or frame repetition period), the source scan type of the associated picture, and whether the associated picture is a copy of a previous picture. In previous video codec standards, this information was typically signaled along with the timing information of the associated picture in the Picture Timing SEI message. However, it has been observed that Frame Field Information and Timing Information are two different types of information and are not necessarily signaled together. A typical example involves signaling the timing information at the system level but signaling the Frame Field Information within the bitstream. Therefore, it was decided to remove the Frame Field Information from the Picture Timing SEI message and instead signal it in a dedicated SEI message. This change also makes it possible to modify the syntax of the Frame Field Information to convey more and clearer instructions to the display, such as pairing fields together or specifying more values for frame repetition.
[0176] Sample Aspect Ratio (SEI) messages can signal different sample aspect ratios for different images within the same sequence, while the corresponding information contained in the VUI applies to the entire sequence. This can be relevant when using reference image resampling features with scaling factors that cause different images in the same sequence to have different sample aspect ratios.
[0177] The Sub-Picture Level Information (SEI) message provides level information for a sequence of sub-pictures.
[0178] 3.3. Image Segmentation and Sub-images in VVC
[0179] In VVC, an image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the image. The CTUs within a slice are scanned in raster scan order within that slice.
[0180] A strip consists of an integer number of consecutive complete CTU lines within an integer number of complete slices or images.
[0181] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a series of complete slices within a sheet raster scan of an image. In rectangular stripe mode, a stripe contains multiple complete slices that together form a rectangular area of the image, or multiple consecutive complete CTU rows of a single slice that together form a rectangular area of the image. Slices within a rectangular stripe are scanned in sheet raster scan order within the rectangular area corresponding to that stripe.
[0182] A sub-image contains one or more stripes that collectively cover a rectangular area of the image.
[0183] 3.3.1. The concept and function of sub-images
[0184] In VVC, each sub-image consists of one or more complete rectangular strips that collectively cover a rectangular area of the image, for example, as shown below. Figure 1 As shown. Sub-images can be defined as extractable (i.e., independent of other sub-images of the same image and earlier images in the decoding order) or non-extractable. Regardless of whether a sub-image is extractable, the encoder can control whether to apply loop filtering (including deblocking, SAO, and ALF) separately across sub-image boundaries for each sub-image.
[0185] Functionally, sub-images are similar to the Motion Constrained Piece Set (MCTS) of HEVC. For use cases such as viewport-dependent 360° video streaming optimization and region of interest (ROI) applications, they allow for the independent encoding, decoding, and extraction of rectangular subsets of the encoded image sequence.
[0186] In 360° video streaming (also known as omnidirectional video), at any given moment, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user, and the user can rotate their head at any time to change their viewing direction, thus changing the current viewport. While it is desirable to have at least some lower-quality representations of areas not covered by the current viewport available at the client end, ready to be presented to the user in case they suddenly change their viewing direction to anywhere on the sphere, the high-quality representation of the omnidirectional video is only needed for the current viewport presented to the user at any given moment. Dividing the high-quality representation of the entire omnidirectional video into sub-images of appropriate granularity enables... Figure 1 The optimization shown has 12 high-precision sub-images on the left-hand side and the remaining 12 lower-precision sub-images of the omnidirectional video on the right-hand side.
[0187] Another typical sub-image-based viewport-dependent 360° video transmission scheme is as follows: Figure 2 As shown, only the higher-precision representation of the full video consists of sub-pictures, while the lower-precision representation of the full video does not use sub-pictures and can be encoded and decoded using a lower-frequency RAP compared to the higher-precision representation. The client receives the lower-precision full video, while for the higher-precision video, the client only receives and decodes the sub-pictures covering the current viewport.
[0188] 3.3.2. Difference between sub-images and MCTS
[0189] There are several important design differences between subpictures and MCTS. First, the subpicture feature in VVC allows the motion vectors of the codec block to point outside the subpicture, even though the subpicture can be extracted by applying sample padding to the subpicture boundary, similar to at the picture boundary. Second, the selection and derivation of motion vectors introduce additional variations during motion vector refinement on the decoder side in merge mode and VVC. This allows for higher encoding / decoding efficiency compared to the non-canonical motion constraints applied on the codec side in MCTS. Third, when extracting one or more extractable subpictures from a picture sequence to create a subbitstream as a conforming bitstream, it is not necessary to rewrite the SH (and PH NAL units, if present). In HEVC MCTS-based subbitstream extraction, the SH needs to be rewritten. Note that in both HEVC MCTS extraction and VVC subpicture extraction, the SPS and PPS need to be rewritten. However, there are typically only a few parameter sets in a bitstream, and each picture has at least one stripe, so rewriting the SH can be a significant burden for the application system. Fourth, the stripes of different subpictures within a picture allow for different NAL unit types. This is a characteristic commonly referred to as the hybrid NAL unit type or hybrid subpicture type within a picture, which will be discussed in more detail below. Fifth, VVC specifies HRD and level definitions for subpicture sequences, thus allowing the codec to guarantee the consistency of the sub-bitstream for each extractable subpicture sequence.
[0190] 3.3.3. Mixed Sub-image Types within an Image
[0191] In AVC and HEVC, all VCL NAL units within a picture must have the same NAL unit type. VVC introduces the option to mix subpicks within a picture that have some different VCL NAL unit types, thus providing support for random access not only at the picture level but also at the subpick level. In VVC, VCL NAL units within a subpick still need to have the same NAL unit type.
[0192] The ability to randomly access sub-images from IRAP is beneficial for 360° video applications. In similar... Figure 2 In the viewport-dependent 360° video transmission scheme shown, the content of spatially adjacent viewports largely overlaps. That is, during a viewport orientation change, only a small portion of sub-pictures within the viewport are replaced by new sub-pictures, while the majority of sub-pictures remain in the viewport. The newly introduced sub-picture sequence must begin with an IRAP stripe; however, by allowing inter-frame prediction of the remaining sub-pictures during viewport changes, a significant reduction in the overall transmission bit rate can be achieved.
[0193] The PPS of the image reference provides an indication of whether an image contains only one type of NAL unit or more than one type (i.e., using a flag called pps_mixed_nalu_types_in_pic_flag). An image can consist of sub-images containing IRAP stripes and sub-images that also contain trailing stripes. Several other combinations of different NAL unit types within an image are also allowed, including leading image stripes of NAL unit types RASL and RADL. This allows merging sub-image sequences with open-GOP and close-GOP codec structures extracted from different bitstreams into a single bitstream.
[0194] 3.3.4. Sub-image layout and ID signaling notification
[0195] The layout of sub-images in VVC is signaled in SPS and therefore remains unchanged within CLVS. Each sub-image is signaled by the position of its top-left CTU and its width and height (in terms of the number of CTUs), thus ensuring that sub-images cover the rectangular area of the image at the CTU granularity. The order in which sub-images are signaled in SPS determines the index of each sub-image within the image.
[0196] To extract and merge sub-picture sequences without rewriting the SH or PH, the stripe addressing scheme in VVC is based on the sub-picture ID and the sub-picture-specific stripe index to associate a stripe with a sub-picture. In the SH, signaling informs the sub-picture ID and sub-picture-level stripe index of the sub-picture containing the stripe. Note that the value of the sub-picture ID for a specific sub-picture may differ from the value of its sub-picture index. The mapping between the two is either signaled in the SPS or PPS (but not both) or implicitly inferred. If present, the sub-picture ID mapping needs to be rewritten or added when rewriting the SPS and PPS during the sub-picture sub-bitstream extraction process. Together, the sub-picture ID and sub-picture-level stripe index indicate to the decoder the exact location of the first decoding CTU of the stripe within the DPB slot of the decoded picture. After sub-bitstream extraction, the sub-picture ID remains unchanged, while the sub-picture index may change. Even if the raster scan CTU address of the first CTU in a stripe of a sub-picture has changed compared to the value in the original bitstream, the unchanged values of the sub-picture ID and sub-picture level stripe index in the corresponding SH will still correctly determine the position of each CTU in the decoded image of the extracted sub-bitstream. Figure 3 The example demonstrates how to extract sub-images using sub-image IDs, sub-image indexes, and sub-image level stripe indices, with the example containing two sub-images and four stripes.
[0197] Similar to sub-image extraction, sub-image signaling notification allows multiple sub-images from different bitstreams to be merged into a single bitstream by simply rewriting the SPS and PPS, provided that the different bitstreams are generated in a coordinated manner (e.g., using different sub-image IDs, but otherwise primarily aligned with SPS, PPS, and PH parameters, such as CTU size, chroma format, encoding / decoding tools, etc.).
[0198] Although subpicks and stripes are signaled independently in the SPS and PPS respectively, there are inherent mutual constraints between the subpick and stripe layouts to form a consistent bitstream. First, the existence of subpicks requires the use of rectangular stripes and prohibits raster scan stripes. Second, the stripes of a given subpick should be consecutive NAL units in decoding order, meaning that the subpick layout constrains the order of the NAL units of the codec stripes within the bitstream.
[0199] 3.4. General Sub-Bitstream Extraction Process in VVC
[0200] Similar to HEVC, the VVC specification includes a sub-bitstream extraction process, allowing the extraction of sub-bitstreams corresponding to a specific operation point (i.e., the OLS and the included temporal sublayers). In HEVC, however, the extraction process is part of the decoding process; the decoder will need to discard NAL units that are not relevant to the operation point when they appear in the bitstream. The VVC design assumes that the bitstream fed to the decoder does not contain NAL units that do not belong to the indicated operation point; that is, when necessary, NAL units that are not relevant to the operation point are discarded by the extractor, which is not part of the decoder. A significant difference compared to HEVC is that in VVC, the handling of scalable nested HRDSEI messages is standardized. For example, the extracted bitstream carries the correct HRD timing parameters for the target operation point. This process involves removing the original BP SEI message and the DUI SEI message (if any) when the target operation point does not include all layers in the bitstream, and inserting the appropriate SEI message originally included in the scalable nested SEI message. However, PT SEI messages have special handling; the above operations are only required when there is no indication that each PT SEI message applies to all OLSs.
[0201] The specification for the general sub-bitstream extraction process in VVC is as follows:
[0202] C.6 General Sub-Bitstream Extraction Process
[0203] The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the highest target TemporalId value tIdTarget.
[0204] The output of this process is the sub-bitstream outBitstream.
[0205] An OLS with an OLS index targetOlsIdx is called a target OLS.
[0206] The requirement for bitstream consistency of the input bitstream is that any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0207] – The output sub-bitstream is the output of the process specified in this sub-term, where the bitstream targetOlsIdx is equal to the index of the OLS list specified by the VPS, and the input tIdTarget is equal to any value in the range of vps_ptl_max_tid[vps_ols_ptl_idx[targetOlsIdx]] (inclusive).
[0208] – The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each nuh_layer_id value in LayerIdInOls[targetOlsIdx].
[0209] – The output sub-bitstream contains at least one VCLNAL unit whose TemporalId is equal to tIdTarget.
[0210] Note – A consistent bitstream contains one or more striped NAL units encoded with TemporalId equal to 0, but does not necessarily contain striped NAL units encoded with nuh_layer_id equal to 0. The output sub-bitstream OutBitstream is derived by applying the following ordered steps:
[0211] 1. The bitstream outBitstream is set to be the same as the bitstream inBitstream.
[0212] 2. Remove all NAL cells from outBitstream whose TemporalId is greater than tIdTarget.
[0213] 3. Remove all NAL units with nuh_layer_id from outBitstream that are not included in the list LayerIdInOls[targetOlsIdx], are not DCI, OPI, VPS, AUD, or EOB NAL units, and are not SEI NAL units containing non-scalable nested SEI messages with a payloadType equal to 0, 1, 130, or 203.
[0214] 4. Remove all APS and VCL NAL units, along with their associated non-VCL NAL units, from outBitstream if all of the following conditions are true: nal_unit_type equals PH_NUT or FD_NUT, or nal_unit_type equals SUFFIX_SEI_NUT or PREFIX_SEI_NUT, and SEI messages containing a payload type not equal to any of 0 (BP), 1 (PT), 130 (DUI), and 203 (SLI):
[0215] –nal_unit_type equals APS_NUT, TRAIL_NUT, STSA_NUT, RADL_NUT, or RASL_NUT, or nal_unit_type equals GDR_NUT and the associated ph_recovery_poc_cnt is greater than 0.
[0216] –TemporalId is greater than or equal to NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]].
[0217] 5. When all VCL NAL units of AU are removed through steps 2, 3 or 4 above, and AUD or OPINAL units exist in AU, remove the AUD or OPINAL units from outBitstream.
[0218] 6. For each OPI NAL unit in outBitstream, set opi_htid_info_present_flag to 1, set opi_ols_info_present_flag to 1, set opi_htid_plus1 to tIdTarget+1, and set opi_ols_idx to targetOlsIdx.
[0219] 7. When AUD exists in an AU in outBitstream, and the AU is an IRAP or GDR AU, set the aud_irap_or_gdr_flag of AUD to 1.
[0220] 8. Remove all SEI NAL units from outBitstream that contain scalable nested SEI messages with sn_ols_flag equal to 1 and no i value in the range from 0 to sn_num_olss_minus1 (inclusive) such that NestingOlsIdx[i] equals targetOlsIdx.
[0221] 9. Remove all SEI NAL cells from outBitstream that contain scalable nested SEI messages with sn_ols_flag equal to 0 and whose values in the list NestingLayerId are not equal to the values in the list LayerIdInOls[targetOlsIdx].
[0222] 10. When LayerIdInOls[targetOlsIdx] does not include all values of nuh_layer_id in all VCL NAL units of the bitstream, the following apply in the order listed:
[0223] a. Remove all SEI NAL units from outBitstream that contain non-scalable nested SEI messages with payloadType equal to 0 (BP), 130 (DUI), or 203 (SLI).
[0224] b. When general_same_pic_timing_in_all_ols_flag equals 0, remove all SEI NAL cells from outBitstream that contain non-scalable nested SEI messages with payloadType equal to 1 (PT).
[0225] c. When outBitstream contains SEI NAL unit seiNalUnitA, which contains scalable nested SEI messages with sn_ols_flag equal to 1 and sn_subpic_flag equal to 0 applied to the target ols, or when NumLayersInOls[targetOlsIdx] equals 1 and outBitstream contains SEI NAL unit seiNalUnitA, which contains scalable nested SEI messages with sn_ols_flag equal to 0 and sn_subpic_flag equal to 0 applied to the layers in outBitstream, generate a new SEI NAL unit seiNalUnitB, include it in the PU containing the seinal unit, immediately following seiNalUnitA, extract scalable nested SEI messages from the scalable nested SEI messages and include them directly in seiNalUnitB (as non-scalable nested SEI messages), and remove seiNalUnitA from outBitstream.
[0226] 3.5. Sub-image Bitstream Extraction Process in VVC
[0227] VVC allows for HRD conformance testing on each independently encoded and decoded sub-picture. Specifically, it allows the extraction of bitstream portions associated with each sub-picture to form a valid bitstream, and testing the conformance of this bitstream to the HRD model. Conformance testing requires defining a complete HRD model for such sub-picture bitstreams, and VVC allows carrying necessary information, in addition to other HRD parameters, through a new SEI message called Subpicture Level Information (SLI) SEI message.
[0228] The SLI SEI message provides level information for the subpicture sequence, which is needed to derive the CPB size and bitrate values of the HRD model describing the decoder's processing of the subpicture bitstream. While the original bitstream consisting of multiple subpictures can adhere to constraints defined by a specific level indicated by the parameter set of the bitstream (e.g., level 5.1 for 4K at 60Hz), the subpicture sub-bitstreams of that specific bitstream can correspond to lower levels (e.g., level 3 for 720p at 60Hz). Furthermore, VVC allows the level of a specific subpicture sub-bitstream to be expressed by a score of the reference level, which allows for finer-grained level signaling notification than in HEVC, and this information also provides guidance to systems merging multiple subpictures into a single joint bitstream regarding how much each subpicture sub-bitstream will contribute to the level constraints of the merged bitstream. Additional attributes, such as bitstreams exhibiting a constant bitrate, or, in the case of multi-layer bitstreams, the level contribution of layers where subpicture segmentation is not applied, can also be signaled in the SLI SEI message, thus enabling the derivation of a complete HRD model for each individual subpicture sequence in such scenarios. Another part of the effort to allow conforming sub-picture sub-bitstreams is to apply many consistency-related constraints that have both picture ranges and sub-picture ranges, such as the minimum compression ratio or bin bit ratio for VCL NAL cells belonging to a single sub-picture.
[0229] In HEVC, sub-bitstream extraction and conformance testing of the Code-Code-Side Region (MCTS) require the use of MCTS Extraction Information Set (SEI) messages to carry the parameter set of this sub-bitstream in a nested manner.
[0230] VVC adds a new sub-picture sub-bitstream extraction process. By removing unnecessary NAL units related to other sub-pictures and proactively rewriting relevant parts of the parameter set to accurately reflect the attributes of the sub-picture sub-bitstream, it allows the generation of a consistent bitstream from independently encoded and decoded sub-pictures. For example, this process includes rewriting the level indicators and HRD parameters in the VPS and SPS, as well as the picture size, segmentation information, consistency window offset, scaling window offset, and virtual boundary position in appropriate parts of each parameter set. The sub-picture sub-bitstream extraction process rewrites some information in the bitstream's parameter set based on the information provided in the SLI SEI message.
[0231] The specification for the VVC sub-image sub-bitstream extraction process is as follows:
[0232] C.7 Sub-image Sub-bit Stream Extraction Process
[0233] The input to this process is a list of bitstream inBitstream, target OLS index targetOlsIdx, target highest TemporalId value tIdTarget, and target subpic index value subpicIdxTarget[i], where i ranges from 0 to NumLayersInOls[targetOlsIdx]1 (inclusive).
[0234] The output of this process is the sub-bit stream outBitstream.
[0235] An OLS with an OLS exponent targetOlsIdx is called a target OLS. Within the layers of a target OLS, graphs with a reference SPS whose sps_num_subpics_minus1 is greater than 0 are called multiSubpicLayers.
[0236] The requirement for bitstream consistency of the input bitstream is that any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0237] – The output sub-bitstream is the output of the process specified in this sub-clause, where the bitstream targetOlsIdx is equal to the index of the list of OLS specified by the VPS, tIdTarget is equal to any value in the range of 0 to vps_max_sublayers_minus1 (inclusive), and a list of i from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive) is taken as input, satisfying the following condition:
[0238] The value of –subpicIdxTarget[i] is equal to a value in the range of 0 to sps_num_subpics_minus1 (inclusive), such that sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] equals 1, where sps_num_subpics_minus1 and sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] are found in the SPS referenced by the layer whose nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i], or inferred based on that SPS.
[0239] Note 1 – When the layer with nuh_layer_id equals LayerIdInOls[targetOlsIdx][i] has a sps_num_subpics_minus1 of 0, the value of subpicIdxTarget[i] is equal to 0.
[0240] – For any two distinct integer values of m and n, when sps_num_subpics_minus1 is greater than 0, for two layers whose nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][m] and LayerIdInOls[targetOlsIdx][n] respectively, subpicIdxTarget[m] is equal to subpicIdxTarget[n].
[0241] – The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each nuh_layer_id value in the list LayerIdInOls[targetOlsIdx].
[0242] – The output sub-bitstream contains at least one VCL NAL unit whose TemporalId is equal to tIdTarget.
[0243] Note 2 – A consistent bitstream contains one or more striped NAL units with TemporalId equal to 0, but does not necessarily contain striped NAL units with nuh_layer_id equal to 0.
[0244] –For each i in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), the output sub-bitstream contains at least one VCL NAL unit where nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is equal to SubpicIdVal[subpicIdxTarget[i]].
[0245] The output sub-bitstream, outBitstream, is exported through the following sequential steps:
[0246] 1. The sub-bitstream extraction procedure specified in Appendix C.6 is called with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the output of the procedure is assigned to outBitstream.
[0247] 2. For each i value in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), remove from outBitstream all VCL NAL units, their associated padding data NAL units, and their associated SEI NAL units containing the padding payload SEI message, where nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is not equal to SubpicIdVal[subpicIdxTarget[i]].
[0248] 3. When an SLI SEI message applied to the target OLS exists and the sli_cbr_constraint_flag of the SLI SEI message is equal to 0, remove all NAL units with nal_unit_type equal to FD_NUT and SEI NAL units containing padding payload SEI messages.
[0249] 4. Remove all SEI NAL cells from outBitstream that contain scalable nested SEI messages with sn_subpic_flag equal to 1, and ensure that none of the sn_subpic_id[j] values from j to sn_num_subpics_minus1 (inclusive) are equal to any SubpicIdVal[subpicIdxTarget[i]] value of any layer in multiSubpicLayers.
[0250] 5. When at least one VCL NAL cell has been removed via step 2, remove all SEI NAL cells from outBitstream that contain scalable nested SEI messages with sn_subpic_flag equal to 0.
[0251] 6. If some external device not specified in this document can be used to provide a replacement parameter set for the sub-bitstream outBitstream, then replace all parameter sets with the replacement parameter set. Otherwise, apply the following ordered steps:
[0252] a. For any layer in multiSubpicLayers, the variable spIdx is set to the value of subpicIdxTarget[i].
[0253] b. When the SLI SEI message applied to the target OLS exists, for k in the range of 0 to tIdTarget–1 (inclusive), the values of general_level_idc and sublayer_level_idc[k] in the vps_ols_ptl_idx[targetOlsIdx]-th entry of the profile_tier_level() syntax structure list in all referenced VPSs (when they exist) and the values of general_level_idc and sublayer_level_idc[k] in the profile_tier_level() syntax structure in all referenced SPSs (when NumLayersInOls[targetOlsIdx] equals 1) are set to equal to SubpicLevelIdc[spIdx][tIdTarget] and SubpicLevelIdc[spIdx][k], respectively, which is derived from Equation 1621 for the spIdx-th subpicture sequence.
[0254] c. When an SLI SEI message applied to the target OLS exists, for k in the range 0 to tIdTarget (inclusive), let spLvIdx be set equal to SubpicLevelIdx[spIdx][k], where SubpicLevelIdx[spIdx][k] is derived from Equation 1621 for the spIdx-th subpic sequence. When the VCL HRD parameter or NAL... When the HRD parameter exists, for k in the range of 0 to tIdTarget (inclusive), set the corresponding values of cpb_size_value_minus1[k][j] and bit_rate_value_minus1[k][j] of the j-th CPB in the vps_ols_timing_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] ols_timing_hrd_parameters() syntax structure in all referenced VPSs (if they exist), and also set the corresponding values of cpb_size_value_minus1[k][j] and bit_rate_value_minus1[k][j] of the j-th CPB in the ols_timing_hrd_parameters() syntax structure in all referenced SPSs. The corresponding values of [j] (when NumLayersInOls[targetOlsIdx] equals 1) are made to correspond to SubpicCpbSizeVcl[spLvIdx][spIdx][k] and SubpicCpbSizeNal[spLvIdx][spIdx][k] derived from equations 1617 and 1618, and to correspond to SubpicBitrateVcl[spLvIdx][spIdx][k] and SubpicBitrateNal[spLvIdx][spIdx][k] derived from equations 1619 and 1620, respectively, where j is in the range from 0 to hrd_cpb_cnt_minus1 (inclusive), and i is in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive).
[0255] d. For each layer in multiSubpicLayers, the following ordered steps apply to rewriting the SPS and PPS referenced by the images in that layer:
[0256] i. The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are exported as follows:
[0257]
[0258] ii. Set the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS to be equal to subpicWidthInLumaSamples and subpicHeightInLumaSamples, respectively.
[0259] iii. Set the value of pps_num_subpics_minus1 in all referenced PPS to 0.
[0260] iv. Set the values of the syntax elements sps_subpic_ctu_top_left_x[spIdx] and sps_subpic_ctu_top_left_y[spIdx] (if they exist) in all referenced SPS to 0.
[0261] v. For each j that is not equal to spIdx, remove all referenced syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j], and sps_subpic_id[j] (if they exist) from the SPS.
[0262] vi. When spIdx is greater than 0 and the referenced SPS's sps_subpic_id_mapping_explicitly_signalled_flag is equal to 0, set the values of sps_subpic_id_mapping_explicitly_signalled_flag and sps_subpic_id_mapping_present_flag to 1, and add sps_subpic_id[0] equal to spIdx to the SPS.
[0263] vii. For each j that is not equal to spIdx, remove the syntax element pps_subpic_id[j] from all referenced PPS (if it exists).
[0264] viii. Set the syntax elements in all referenced PPS to signal slices and stripes to remove all slice rows, slice columns, and stripes that are not related to the sub-picture whose sub-picture index is equal to spIdx.
[0265] The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are exported as follows:
[0266]
[0267]
[0268] The values of sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, sps_pic_height_max_in_luma_samples, sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in these equations are derived from the original SPS values before the rewrite.
[0269] Note 3 – For the images in the layers of multiSubpicLayers in both the input and output bitstreams, the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are equal to pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples, respectively. Therefore, in these equations, sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples can be replaced with pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples, respectively.
[0270] x. Set the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in all referenced SPS to be equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively.
[0271] The variables subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset are exported as follows:
[0272]
[0273] In these equations, the values of sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, and sps_pic_height_max_in_luma_samples are from the original SPS before rewriting, and the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset are from the original PPS before rewriting.
[0274] xii. Set the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in all referenced PPS NAL cells to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset, respectively.
[0275] xiii. The variables numVerVbs, subpicVbx[i], numHorVbs, and subpicVby[i] are derived as follows:
[0276]
[0277] The values of sps_num_ver_virtual_boundaries, sps_virtual_boundary_pos_x_minus1[i], sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_num_hor_virtual_boundaries, sps_virtual_boundary_pos_y_minus1[i], sps_subpic_ctu_top_left_y[spIdx], and sps_subpic_height_minus1[spIdx] in these equations are derived from the original SPS values before the rewrite.
[0278] xiv. For i in the range 0 to numVerVbs-1 (inclusive) and j in the range 0 to numHorVbs-1 (inclusive), when sps_virtual_boundaries_present_flag equals 1, set the values of the syntax elements sps_num_ver_virtual_boundaries, sps_virtual_boundary_pos_x_minus1[i], sps_num_hor_virtual_boundaries, and sps_virtual_boundary_pos_y_minus1[j] in all referenced SPSs to numVerVbs, subpicVbx[i]–1, numHorVbs, and subpicVby[j]–1, respectively. Virtual boundaries outside the extracted subpicks are removed. When both numVerVbs and numHorVbs are equal to 0, set the value of sps_virtual_boundaries_enabled_flag in all referenced SPSs to 0, and remove the syntax elements sps_virtual_boundaries_present_flag, sps_num_ver_virtual_boundaries, sps_virtual_boundary_pos_x_minus1[i], sps_num_hor_virtual_boundaries, and sps_virtual_boundary_pos_y_minus1[i].
[0279] e. When an SLI SEI message applicable to the target OLS exists, the following applies:
[0280] i. If sli_cbr_constraint_flag equals 1, set cbr_flag[tIdTarget][j] of the j-th CPB in the vps_ols_timing_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] ols_timing_hrd_parameters() syntax structure in all referenced VPSs to equal 1, and when NumLayersInOls[targetOlsIdx] equals 1, set cbr_flag[tIdTarget][j] of the cbr_flag[tIdTarget][j] in the ols_timing_hrd_parameters() syntax structure in all referenced SPSs to equal 1.
[0281] ii. Otherwise, (if sli_cbr_constraint_flag equals 0), set cbr_flag[tIdTarget][j] to 0. In both cases, j is in the range of 0 to hrd_cpb_cnt_minus1 (inclusive).
[0282] 7. When at least one VCL NAL cell has been removed in step 2, the following will be applied in the order listed:
[0283] a. Remove all SEI NAL units from outBitstream that contain non-scalable nested SEI messages with payloadType equal to 0 (BP), 130 (DUI), 203 (SLI), or 132 (decoded image hash).
[0284] b. When general_same_pic_timing_in_all_ols_flag equals 0, remove all SEI NAL cells from outBitstream that contain non-scalable nested SEI messages with payloadType equal to 1 (PT).
[0285] c. When outBitstream contains a SEI NAL unit seiNalUnitA, which contains scalable nested SEI messages with sn_ols_flag equal to 1 and sn_subpic_flag equal to 1 applied to the target ols and subpicks in outBitstream, or when NumLayersInOls[targetOlsIdx] equals 1 and outBitstream contains a SEI NAL unit seiNalUnitA, which contains scalable nested SEI messages with sn_ols_flag equal to 0 and sn_subpic_flag equal to 1 applied to the layers and subpicks in outBitstream, generate a new SEI NAL unit seiNalUnitB, include it in the PU containing the seinal unit, immediately following seiNalUnitA, extract scalable nested SEI messages from the scalable nested SEI messages and include them directly in seiNalUnitB (as non-scalable nested SEI messages), and remove seiNalUnitA from outBitstream.
[0286] 4. Examples of technical problems solved by the disclosed technical solutions
[0287] The design of VVC's POC, SEI, and sub-image sub-bitstream extraction has the following problems:
[0288] 1) Clause C.4 of the VVC specification includes the following constraints related to POC:
[0289] Set currPicLayerId to equal the nuh_layer_id of the current image.
[0290] For each current image, set the variables maxPicOrderCnt and minPicOrderCnt to the maximum and minimum values of PicOrderCntVal, respectively, for the following images, where nuh_layer_id equals currPicLayerId:
[0291] – Current image.
[0292] – The preceding image in the decoding order has both TemporalId and ph_non_ref_pic_flag equal to 0, and is not a RASL or RADL image.
[0293] – The STRP referenced by all entries in RefPicList[0] and all entries in RefPicList[1] of the current image.
[0294] – For all images n with PictureOutputFlag equal to 1, AuCpbRemovalTime[n] is less than AuCpbRemovalTime[currPic] and DpbOutputTime[n] is greater than or equal to AuCpbRemovalTime[currPic], where currPic is the current image.
[0295] For each current image that is not a CLVSS image, the value of maxPicOrderCnt-minPicOrderCnt should be less than MaxPicOrderCntLsb / 2.
[0296] If the second bullet point above is relaxed, the above constraints do not allow the following bitstreams to be consistent bitstreams:
[0297] a. The original bitstream is a single-layer bitstream, and the sub-bitstream `outBitstream` is extracted from the bitstream by removing all non-intra-encoded images. `outBitstream` must contain at least one image `picA` such that the difference between `picA` and `prevTid0Pic`'s POC value (i.e., `TemporalId` and `ph_non_ref_pic_flag` are both 0, and it is not the preceding image in the same layer of a RASL or RADL image in decoding order) is greater than `MaxPicOrderCntLsb / 2`. Based on these constraints, this bitstream `outBitstream` is not a consistent bitstream because it violates the constraints.
[0298] However, if picA has the POC MSB value signaled in PH, its POC can be correctly derived. Therefore, if the second bullet point is changed as follows, the bitstream outBitstream may become a consistent bitstream:
[0299] – When the current image does not have ph_poc_msb_cycle_val, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0300] b. The original bitstream is a multi-layer bitstream. The layer with layer index j is the direct reference layer of another layer in the bitstream with layer index layerB, and vps_max_tid_il_ref_pics_plus1[i][j] equals 0. Then, in the extracted bitstream outBitstream, according to the general sub-bitstream extraction process, for an OLS where layer j is the output layer but layer i is not, there are only GDR or IRAP images with ph_recovery_poc_cnt equal to 0 in layer i. At this time, as above, if there is at least one image picA in outBitstream at layer i, such that the difference in POC values between image picA and image prevTid0Pic (i.e., the TemporalId and ph_non_ref_pic_flag of the previous image in the same layer according to the decoding order are both equal to 0, and the image is not a RASL or RADL image) is greater than MaxPicOrderCntLsb / 2. According to the above constraints, this bitstream outBitstream is not a consistent bitstream because it violates the constraints. Since if the original bitstream is a consistent bitstream, then the extracted bitstream, such as outBitstream, needs to be a consistent bitstream as well; therefore, the original bitstream is not a consistent bitstream either.
[0301] However, if picA has a POC MSB value signaled in PH, its POC can be correctly derived. Similarly, if the second bulleted item above is changed as follows, both the bitstream outBitstream and the original bitstream in this example may become consistent bitstreams:
[0302] – When the current image does not have ph_poc_msb_cycle_val, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0303] c. Similar to the raw bitstream and outBitstream in examples a and b above, but containing multiple layers, and for image picA, there is another image picB in the same AU, and this image belongs to the reference layer containing picA. In other words, the POC value of picA will be derived to be equal to the POC value of picB. In this case, even if the difference in POC between picA and prevTid0Pic is greater than MaxPicOrderCntLsb / 2, the POC value of picA can still be correctly derived, as long as the POC value of picB can be correctly derived.
[0304] Therefore, if the second bulleted item above is changed as follows, the bitstream outBitstream in this type of example may become a consistent bitstream:
[0305] – When there is no image belonging to the reference layer in the current AU, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0306] 2) In the general sub-bitstream extraction process, step 9 removes all SEI NAL units from the outBitstream that contain scalable nested SEI messages with sn_ols_flag equal to 0 and whose list NestingLayerId has no value equal to the value in the list LayerIdInOls[targetOlsIdx]. However, this step is actually unnecessary because: 1) For SEI NAL units containing scalable nested SEI messages with sn_ols_flag equal to 0, NestingLayerId will always include the layer ID of the SEI NAL unit. 2) If the layer ID of the SEI NAL unit is not in the list LayerIdInOls[targetOlsIdx], it may have already been removed in step 3.
[0307] 3) During the subpicture sub-bitstream extraction process, after at least one VCL NAL unit has been removed in step 2, step 5 removes all SEINAL units from the outBitstream that contain scalable nested SEI messages with sn_subpic_flag equal to 0, including nested SLI SEI messages. However, step 6 may require nested SLI SEI messages (if present) to rewrite the parameter set.
[0308] 4) Nested and non-nested HRD-related SEI messages are allowed to coexist and be applied to the same OLS. Similarly, nested and non-nested non-HRD-related SEI messages are also allowed to coexist and be applied to the same layer. However, there is a lack of constraint requiring that the content of such SEI messages for a specific payload type be identical.
[0309] 5. Examples of technical solutions
[0310] To address the aforementioned and other issues, a methodology summarized below is disclosed. These items should be considered as examples for explaining general concepts, rather than interpreted in a narrow way. Furthermore, these items can be used individually or in combination in any way.
[0311] 1) To resolve issue 1, it is recommended to change the second bulleted item in the POC-related constraints described in issue 1 to one of the following:
[0312] a. When the current image does not have ph_poc_msb_cycle_val and there is no image in the current AU belonging to the reference layer of the current layer, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0313] i. Alternatively, replace the phrase "There is no image belonging to the reference layer of the current layer in the current AU" with "The PocFromIlrpFlag of the current image is equal to 0", and specify PocFromIlrpFlag as follows:
[0314] Let the variable currLayerIdx equal GeneralLayerIdx[nuh_layer_id]. If vps_independent_layer_flag[currLayerIdx] equals 0, and the current AU contains an image picA with nuh_layer_id equal to layerIdA, such that GeneralLayerIdx[layerIdA] is in the list ReferenceLayerIdx[currLayerIdx], then set the variable PocFromIlrpFlag to 1. Otherwise, PocFromIlrpFlag is set to 0.
[0315] b. When the current image does not have ph_poc_msb_cycle_val, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0316] c. When there is no image belonging to the reference layer in the current AU, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0317] i. Alternatively, replace "There is no image in the current AU that belongs to the reference layer of the current layer" with "The PocFromIlrpFlag of the current image is equal to 0", and specify PocFromIlrpFlag as above.
[0318] d. When ph_poc_msb_cycle_val does not exist in the current image or the current AU, no image belongs to the reference layer of the current layer, the TemporalId and ph_non_ref_pic_flag of the previous image in the decoding order are both equal to 0, and the image is not a RASL or RADL image.
[0319] i. Alternatively, replace the phrase "There is no image in the current AU that belongs to the reference layer of the current layer" with "The PocFromIlrpFlag of the current image is equal to 0", and specify PocFromIlrpFlag as above.
[0320] 2) To solve problem 2, remove step 9 of the general sub-bit stream extraction process.
[0321] 3) To solve problem 3, step 5 of the sub-bitstream extraction process of the sub-image is moved after the entire step 6.
[0322] 4) To address problem 4, the following constraint is proposed: When there are multiple SEI messages associated with a specific AU or DU and applied to a specific OLS or layer with a specific payloadType value not equal to 133, the SEI messages will have the same SEI payload content, regardless of whether some or all of these SEI messages are scalable nested.
[0323] Please note that SEI messages with payloadType equal to 133 are scalable nested SEI messages.
[0324] 6. Example
[0325] Below are some example embodiments of the invention summarized in Section 5 above, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-S2001-vH. Most of the relevant added or modified sections are shown in bold, italic, and underlined fonts, for example, To indicate an addition, some deleted parts are indicated by bold, italic, or double brackets, for example, feature This indicates a deletion. There may be other changes that are essentially editable, and therefore not highlighted.
[0326] 6.1. First Embodiment
[0327] This example applies to projects 1.ai, 2, 3, and 4.
[0328] 8.3.1 Decoding process of image sequential counting
[0329] The output of this process is PicOrderCntVal, which is the image order count of the current image.
[0330] Each encoded / decoded image is associated with an image order count variable, denoted as PicOrderCntVal.
[0331] Set the variable currLayerIdx to be equal to GeneralLayerIdx[nuh_layer_id].
[0332] PicOrderCntVal is exported as follows:
[0333] – If vps_independent_layer_flag[currLayerIdx] equals 0 and there is an image picA in the current AU with nuh_layer_id equal to layerIdA, such that GeneralLayerIdx[layerIdA] is in the list ReferenceLayerIdx[currLayerIdx], then PicOrderCntVal is exported as PicOrderCntVal equal to picA. Furthermore, the value of ph_pic_order_cnt_lsb should be the same in all VCL NAL units of the current AU.
[0334] -otherwise, And the current image's PicOrderCntVal is exported as specified in the remainder of this sub-clause.
[0335] When ph_poc_msb_cycle_val does not exist and the current image is not a CLVSS image, the variables prevPicOrderCntLsb and prevPicOrderCntMsb are exported as follows:
[0336] – Let prevTid0Pic be the image whose nuh_layer_id is equal to the nuh_layer_id of the current image in the decoding order, and whose TemporalId and ph_non_ref_pic_flag are both equal to 0, and which is not a RASL or RADL image.
[0337] Note 1 – In a sub-bitstream consisting only of intra-frame pictures, extracted from a single-layer bitstream and used for intra-frame picture-only special effects playback, prevTid0Pic is The previous intra-frame image in the decoding order. To ensure correct POC export, the encoder can be selected as... The intra-frame image contains ph_poc_msb_cycle_val, or the value of sps_log2_max_pic_order_cnt_lsb_minus4 is set large enough that the POC difference between the current image and prevTid0Pic is less than MaxPicOrderCntLsb / 2.
[0338] Note 2 – When vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0 for any i value in the layer index and j is equal to the layer index of the current layer, the prevTid0Pic in some OLS extracted sub-bitstreams will be the previous IRAP or GDR picture in the current layer whose ph_recovery_poc_cnt is equal to 0 in the decoding order. To ensure that such sub-bitstreams are consistent bitstreams, the codec can choose to include ph_poc_msb_cycle_val for each IRAP or GDR picture, where ph_recovery_poc_cnt is equal to 0, or set the value of sps_log2_max_pic_order_cnt_lsb_minus4 large enough that the POC difference between the current picture and prevTid0Pic is less than MaxPicOrderCntLsb / 2.
[0339] The variable prevPicOrderCntLsb is set to equal ph_pic_order_cnt_lsb of prevTid0Pic.
[0340] The variable prevPicOrderCntMsb is set to be equal to the PicOrderCntMsb of prevTid0Pic. ...
[0342] C.4 Bitstream Consistency
[0343] Bitstreams of encoded and decoded data conforming to this specification shall meet all the requirements specified in this sub-clause.
[0344] Bitstreams should be constructed in accordance with the syntax, semantics, and constraints specified in this specification, excluding this appendix.
[0345] The first encoded / decoded image in the bitstream should be an IRAP image (i.e., an IDR image or a CRA image) or a GDR image.
[0346] The bitstream is tested by HRD to determine whether it conforms to the provisions of subclause C.1.
[0347] Set currPicLayerId to equal the nuh_layer_id of the current image.
[0348] For each current image, set the variables maxPicOrderCnt and minPicOrderCnt to the maximum and minimum values of PicOrderCntVal, respectively, for the following images, where nuh_layer_id equals currPicLayerId:
[0349] – Current image.
[0350] – The preceding image in the decoding order has both TemporalId and ph_non_ref_pic_flag equal to 0, and is not a RASL or RADL image.
[0351] – The STRP referenced by all entries in RefPicList[0] and all entries in RefPicList[1] of the current image.
[0352] – For all images n with PictureOutputFlag equal to 1, AuCpbRemovalTime[n] is less than AuCpbRemovalTime[currPic] and DpbOutputTime[n] is greater than or equal to AuCpbRemovalTime[currPic], where currPic is the current image.
[0353] Each bitstream consistency test should meet all of the following conditions: ...
[0355] 8. For each current image that is not a CLVSS image, the value of maxPicOrderCnt-minPicOrderCnt should be less than MaxPicOrderCntLsb / 2. ...
[0357] C.6 General Sub-Bitstream Extraction Process ...
[0359] The bitstream outBitstream is set to be the same as the bitstream inBitstream.
[0360] 1. The bitstream outBitstream is set to be the same as the bitstream inBitstream.
[0361] 2. Remove all NAL cells from outBitstream whose TemporalId is greater than tIdTarget.
[0362] 3. Remove all NAL units with nuh_layer_id from outBitstream that are not included in the list LayerIdInOls[targetOlsIdx], are not DCI, OPI, VPS, AUD, or EOB NAL units, and are not SEI NAL units containing non-scalable nested SEI messages with a payloadType equal to 0, 1, 130, or 203.
[0363] 4. Remove all APS and VCL NAL units, along with their associated non-VCL NAL units, from outBitstream if all of the following conditions are true: nal_unit_type equals PH_NUT or FD_NUT, or nal_unit_type equals SUFFIX_SEI_NUT or PREFIX_SEI_NUT, and SEI messages containing a payload type not equal to any of 0 (BP), 1 (PT), 130 (DUI), and 203 (SLI):
[0364] –nal_unit_type equals APS_NUT, TRAIL_NUT, STSA_NUT, RADL_NUT, or RASL_NUT, or nal_unit_type equals GDR_NUT and the associated ph_recovery_poc_cnt is greater than 0.
[0365] –TemporalId is greater than or equal to NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]].
[0366] 5. When all VCL NAL units of AU are removed through steps 2, 3 or 4 above, and AUD or OPINAL units exist in AU, remove the AUD or OPINAL units from outBitstream.
[0367] 6. For each OPI NAL unit in outBitstream, set opi_htid_info_present_flag to 1, set opi_ols_info_present_flag to 1, set opi_htid_plus1 to tIdTarget+1, and set opi_ols_idx to targetOlsIdx.
[0368] 7. When AUD exists in an AU in outBitstream, and the AU is an IRAP or GDR AU, set the aud_irap_or_gdr_flag of AUD to 1.
[0369] 8. Remove all SEI NAL units from outBitstream that contain scalable nested SEI messages with sn_ols_flag equal to 1 and no i value in the range from 0 to sn_num_olss_minus1 (inclusive) such that NestingOlsIdx[i] equals targetOlsIdx.
[0370] Remove all SEI NAL cells from outBitstream that contain scalable nested SEI messages with sn_ols_flag equal to 0 and whose values in the list NestingLayerId are not equal to the values in the list LayerIdInOls[targetOlsIdx].
[0371] 9. When LayerIdInOls[targetOlsIdx] does not include all values of nuh_layer_id in all VCL NAL units of the bitstream, the following apply in the order listed: ...
[0373] C.7 Sub-image Sub-bitstream Extraction Process ...
[0375] The output sub-bitstream, outBitstream, is exported through the following sequential steps:
[0376] 1. The sub-bitstream extraction procedure specified in Appendix C.6 is called with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the output of the procedure is assigned to outBitstream.
[0377] 2. For each i value in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive of end value 1), remove from outBitstream all VCL NAL units, their associated padding data NAL units, and their associated SEI NAL units containing the padding payload SEI message, where nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is not equal to SubpicIdVal[subpicIdxTarget[i]].
[0378] 3. When an SLI SEI message applied to the target OLS exists and the sli_cbr_constraint_flag of the SLI SEI message is equal to 0, remove all NAL units with nal_unit_type equal to FD_NUT and SEI NAL units containing padding payload SEI messages.
[0379] 4. Remove all SEI NAL cells from outBitstream that contain scalable nested SEI messages with sn_subpic_flag equal to 1, and ensure that none of the sn_subpic_id[j] values from j to sn_num_subpics_minus1 (inclusive) are equal to any SubpicIdVal[subpicIdxTarget[i]] value of any layer in multiSubpicLayers.
[0380] When at least one VCL NAL cell is removed via step 2, all SEI NAL cells containing scalable nested SEI messages with sn_subpic_flag equal to 0 are removed from outBitstream.
[0381] 5. If some external device not specified in this document can be used to provide a replacement parameter set for the sub-bitstream outBitstream, then replace all parameter sets with the replacement parameter set. Otherwise, apply the following ordered steps: ...
[0383]
[0384] 7. When at least one VCL NAL cell has been removed in step 2, the following will be applied in the order listed: ...
[0386] Figure 4 This is a block diagram of an example video processing system 4000 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0387] System 4000 may include an encoding / decoding component 4004 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 4004 can reduce the average bit rate of the video from input 4002 to the output of encoding / decoding component 4004 to produce an encoded / decoded representation of the video. Therefore, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 4004 can be stored or transmitted via connected communication, as represented by component 4006. The stored or communicated bitstream (or encoded / decoded) representation of the video received at input 4002 can be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it should be understood that the encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations will be inverted by the decoder to retrieve the encoded / decoded results.
[0388] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.
[0389] Figure 5 This is a block diagram of a video processing apparatus 5000. Apparatus 5000 can be used to implement the methods described herein (such as...). Figure 9-11 One or more of the methods shown herein. The device 5000 can be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 5000 may include one or more processors 5002, one or more memories 5004, and video processing hardware 5006. The processors 5002(s) may be configured to implement one or more methods described herein. The memories 5004(s) may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 5006 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing hardware 5006 may be at least partially included in the processor 5002, such as a graphics coprocessor.
[0390] Figure 6 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.
[0391] like Figure 6As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the destination device 120 may be referred to as a video decoding device.
[0392] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0393] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems that generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax elements. I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0394] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0395] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to connect to an external display device.
[0396] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVM), and other current and / or other standards.
[0397] Figure 7 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 6 The video encoder 114 in the system 100 shown in the figure.
[0398] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 7 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0399] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0400] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0401] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes... Figure 7 The examples are shown separately.
[0402] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0403] The mode selection unit 203 can, for example, select one of the intra-frame or inter-frame encoding / decoding modes based on the error result, and provide the obtained intra-frame or inter-frame encoded / decoded blocks to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the encoded / decoded blocks for use as reference images. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. The mode selection unit 203 can also select the precision of the motion vector (e.g., sub-pixel or integer pixel precision) for the blocks in the inter-frame prediction case.
[0404] To perform inter-frame prediction for the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information of the image from buffer 213 (rather than the image associated with the current video block) and decoded samples.
[0405] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, the different operations performed depend on whether the current video block is in an I-strip, a P-strip, or a B-strip.
[0406] In some examples, motion estimation unit 204 can perform unidirectional prediction of the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference image in list 0 or list 1 contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0407] In other examples, motion estimation unit 204 can perform bidirectional prediction of the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference images in list 0 or list 1 contain the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0408] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding process.
[0409] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.
[0410] In one example, the motion estimation unit 204 may indicate in the syntax structure associated with the current video block that the current video block has the same motion information value as another video block.
[0411] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicating video block. Video decoder 300 can use the motion vector of the indicating video block and the motion vector difference to determine the motion vector of the current video block.
[0412] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling notification.
[0413] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0414] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0415] In other examples, such as in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform a subtraction operation.
[0416] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0417] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0418] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0419] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0420] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0421] Figure 8 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 6 The video decoder 114 in the system 100 shown in the figure.
[0422] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 8 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0423] exist Figure 8 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 7 The decoding process is the overall inversion of the encoding process described.
[0424] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and merge modes.
[0425] The motion compensation unit 302 can generate motion compensation blocks, possibly based on interpolation filters. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.
[0426] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation values of a sub-integer number of pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.
[0427] The motion compensation unit 302 can use some syntactic information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0428] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0429] The reconstruction unit 306 can sum the residual blocks using the corresponding prediction blocks generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on the display device.
[0430] Figure 9-11 It shows that it can be used in, for example Figure 4-8 The embodiments shown are example methods for implementing the above-described technical solution.
[0431] Figure 9A flowchart illustrating an example method 900 for video processing is shown. Method 900 includes, at operation 910, performing a conversion between video and a bitstream comprising multiple layers of video, the bitstream also including multiple Supplemental Enhancement Information (SEI) messages associated with a specific output layer set (OLS) or a specific layer's access unit (AU) or decoding unit (DU), the multiple SEI messages being of a message type different from the scalable nested type, based on a format rule that specifies that each of the multiple SEI messages has the same SEI payload content due to their association with a specific OLS or a specific layer's AU or DU.
[0432] Figure 10 A flowchart of an example method 1000 for video processing is shown. Method 1000 includes: in operation 1010, performing a conversion between a video including the current block and a bitstream of the video, the bitstream conforming to a format rule that specifies constraints on the bitstream due to the exclusion of syntax fields from the bitstream, the syntax fields indicating the picture sequence count (POC) associated with the current picture including the current block or the current access unit (AU) including the current block.
[0433] Figure 11 A flowchart of an example method 1100 for video processing is shown. Method 1100 includes: in operation 1110, performing a conversion between a video and a bitstream of the video, the bitstream conforming to a sub-bitstream extraction process in a sequence defined by rules, the rules specifying that (a) the sub-bitstream extraction process excludes the operation of removing all Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) units that meet certain conditions from the output bitstream, or (b) the sequence of the sub-bitstream extraction process includes replacing a parameter set with a replacement parameter set before removing SEI NAL units from the output bitstream.
[0434] The following solutions show example implementations of the techniques discussed in the previous chapter (e.g., Items 1-4).
[0435] The following provides a list of preferred solutions for some embodiments.
[0436] A1. A method for processing video data, comprising:
[0437] Perform conversion between video and bitstreams containing multiple layers of video.
[0438] The bitstream includes multiple supplemental enhancement information (SEI) messages associated with a specific output layer set (OLS) or a specific layer's access unit (AU) or decoding unit (DU).
[0439] This includes multiple SEI messages of different message types than scalable nested types, based on format rules, and
[0440] The format rules specify that multiple SEI messages may have the same SEI payload content because they are associated with a specific OLS or a specific layer's AU or DU.
[0441] A2. According to the method described in solution A1, the value of the payload type of the scalable nested SEI message is equal to 133.
[0442] A3. The method according to solution A1, wherein the plurality of SEI messages includes at least one SEI message contained in a scalable nested SEI message and at least one SEI message not contained in a scalable nested SEI message.
[0443] A4. A method for processing video data, comprising:
[0444] Perform the conversion between the current block of video and the video bitstream.
[0445] The bitstream conforms to the format rules, and
[0446] The format rule specifies that, due to the constraint on the bitstream by excluding the syntax field from the bitstream, the syntax field indicates the picture sequence count (POC) associated with the current picture including the current block or the current access unit (AU) including the current block.
[0447] A5. The method described in solution A4, wherein the syntax field indicating the POC associated with the current image and the current AU is excluded, and wherein the constraint specifies that there is no image belonging to the reference layer of the current layer in the current AU.
[0448] A6. The method described in solution A4 or A5, wherein the constraint specifies that the preceding image in the decoding order is neither a Random Access Skip Preamble (RASL) image nor a Random Access Decodeable Preamble (RADL) image.
[0449] A7. According to the method described in solution A4, wherein the constraint stipulates that, since the current AU excludes images belonging to the reference layer of the current layer in the current AU, the preceding image in the decoding order is neither a Random Access Skip Preamble (RASL) image nor a Random Access Decodeable Preamble (RADL) image.
[0450] A8. The method described in solution A6 or A7, wherein the temporal identifier of the previous image is equal to 0, and wherein the temporal identifier is TemporalId.
[0451] A9. The method according to any one of solutions A4 to A8, wherein the syntax field is ph_poc_msb_cycle_val.
[0452] A10. The method according to any one of solutions A1 to A9, wherein the conversion includes decoding video from a bitstream.
[0453] A11. The method according to any one of solutions A1 to A9, wherein the conversion includes encoding the video into a bitstream.
[0454] A12. A method for storing a bitstream representing video to a computer-readable recording medium, comprising:
[0455] Generate a bitstream from the video according to one or more of the methods described in solutions A1 to A9; and
[0456] Storing bitstreams in computer-readable recording media.
[0457] A13. A video processing apparatus, comprising a processor configured to implement the method described in any one or more of solutions A1 to A12.
[0458] A14. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to perform the method according to one or more of solutions A1 to A12.
[0459] A15. A computer-readable medium storing a bit stream generated according to any one or more of solutions A1 to A12.
[0460] A16. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions A1 to A12.
[0461] The following is another list of preferred solutions for some embodiments.
[0462] B1. A method for processing video data, comprising:
[0463] Perform conversion between video and video bitstream.
[0464] The bitstream conforms to the order of the sub-bitstream extraction process defined by the rule that (a) the sub-bitstream extraction process excludes the removal of all Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) units that meet the conditions from the output bitstream, or (b) the order of the sub-bitstream extraction process includes replacing the parameter set with a replacement parameter set before removing the SEI NAL units from the output bitstream.
[0465] B2. The method according to solution B1, wherein the condition specifies that at least one of the SEI NAL units includes a scalable nested SEI message, the scalable nested SEI message including (a) a flag equal to zero and (b) a first identifier having a value different from the value of a second identifier in the list.
[0466] B3. The method according to solution B2, wherein a flag indicates whether a scalable nested SEI message is applicable to a particular set of output layers, wherein a first identifier is a nested layer identifier, and wherein a second identifier is a layer identifier of a layer in the set of output layers.
[0467] B4. The method described in solution B3, wherein the flag is sn_ols_flag.
[0468] B5. The method described in solution B3, wherein the first identifier is NestingLayerId and the second identifier is LayerIdInOls[targetOlsIdx].
[0469] B6. The method according to solution B1, wherein at least one of the SEI NAL units includes a scalable nested SEI message.
[0470] B7. The method according to solution B6, wherein the SEI NAL unit is removed from the output bitstream due to the removal of at least one video codec layer (VCL) NAL unit from the output bitstream.
[0471] B8. The method described in solution B6 or B7, wherein the scalable nested SEI message includes a flag equal to zero.
[0472] B9. The method described in solution B8, wherein the flag is sn_subpic_flag.
[0473] B10. The method according to any one of solutions B1 to B9, wherein the conversion includes decoding video from the bitstream.
[0474] B11. The method according to any one of solutions B1 to B9, wherein the conversion includes encoding the video into a bitstream.
[0475] B12. A method for storing a bitstream representing video to a computer-readable recording medium, comprising:
[0476] Generate a bitstream from the video according to one or more of the methods described in solutions B1 to B9; and
[0477] Storing bitstreams in computer-readable recording media.
[0478] B13. A video processing apparatus, comprising a processor configured to implement the method described in any one or more of solutions B1 to B12.
[0479] B14. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to implement one or more of the methods described in solutions B1 to B12.
[0480] B15. A computer-readable medium for storing a bit stream generated according to any one or more of solutions B1 to B12.
[0481] B16. A video processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions B1 to B12.
[0482] The following is yet another list of preferred solutions for some embodiments.
[0483] P1. A method for video processing, comprising: performing a conversion between a video comprising one or more layers and a codec representation of the video, the one or more layers comprising one or more pictures, wherein the codec representation conforms to a format rule, wherein the format rule specifies constraints on the codec representation in the absence of a syntax field indicating a picture order count associated with the current picture in the currently accessed unit.
[0484] P2. According to the method described in solution P1, the format rules specify the constraint that the preceding image in the decoding order is either randomly accessed skipping the preamble or randomly accessing a decodeable preamble image.
[0485] P3. According to the method described in solution P1, the format rules specify constraints based on the condition that there is no image belonging to the reference layer of the current layer in the current access unit.
[0486] P4. A method for video processing, comprising: performing a conversion between a video comprising one or more layers and a codec representation of the video, the one or more layers comprising one or more pictures, wherein the codec representation conforms to a sub-bitstream extraction process defined by rules.
[0487] P5. According to the method described in solution P4, the rule specifies that the sub-bitstream extraction process omits the step of removing Scalable Nested Supplemental Enhancement Information (SEI) messages from the output bitstream of all SEI network abstraction layer units based on conditions.
[0488] P6. According to the method described in solution P4, wherein the rule specifies that the order of the sub-bitstream extraction process includes replacing the parameter set with a replacement parameter set before removing the Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) units containing scalable nested SEI messages from the output bitstream.
[0489] P7. A video processing method, comprising: performing a conversion between a video comprising one or more layers and a codec representation of the video, the one or more layers comprising one or more pictures, wherein the codec representation conforms to a format rule, wherein the format rule specifies that in the presence of multiple Supplemental Enhancement Information (SEI) messages, the multiple SEI messages having a specific payload type value not equal to 133, the multiple SEI messages being associated with a specific access unit or decoding unit and applied to a specific set of output layers or layers, the multiple SEI messages having the same SEI payload content.
[0490] P8. According to the approach in solution P7, some or all of the multiple SEI messages are scalable and nested.
[0491] P9. The method according to any one of solutions P1 to P8, wherein the transformation includes generating a codec representation from the video.
[0492] P10. The method according to any one of solutions P1 to P8, wherein the conversion includes decoding the encoding / decoding representation to generate video.
[0493] P11. A video decoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions P1 to P10.
[0494] P12. A video encoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions P1 to P10.
[0495] P13. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method of any one of solutions P1 to P10.
[0496] P14. A computer-readable medium storing a codec representation generated according to any one of solutions P1 to P10.
[0497] In the solution described in this paper, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described in this paper, the decoder can use the format rules to parse the syntax elements in the codec representation, determining the presence or absence of syntax elements according to the format rules to generate the decoded video.
[0498] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-occurring or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the error residuals from the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0499] The disclosures and other schemes, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, containing the structures disclosed in this document and their equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, i.e., one or more computer program instruction modules for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a complex influencing machine-readable propagating signals, or combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0500] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0501] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be performed by special-purpose logic circuitry (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), and the apparatus can be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).
[0502] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magneto-optical, magneto-optical, or optical disc), or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0503] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0504] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all the operations shown, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0505] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.
Claims
1. A method for processing video data, comprising: Perform the conversion between the video and the video bitstream. Wherein, the bitstream conforms to the order of the sub-bitstream extraction process defined by the rules, the rules specifying that (a) the sub-bitstream extraction process excludes the operation of removing all supplementary enhancement information SEI network abstraction layer NAL units that meet the conditions from the output bitstream, or (b) the order of the sub-bitstream extraction process includes: replacing the parameter set with a replacement parameter set before removing the SEI NAL unit from the output bitstream.
2. The method according to claim 1, wherein, The condition specifies that the SEI NAL unit includes scalable nested SEI messages, which include (a) a flag equal to zero and (b) no value in a first list of the first identifier equal to a value in a second list of the second identifier.
3. The method according to claim 2, wherein, The flag indicates whether the scalable nested SEI message is applicable to a specific set of output layers, wherein the first identifier is a nested layer identifier, and wherein the second identifier is a layer identifier of a layer in the set of output layers.
4. The method according to claim 3, wherein, The flag is sn_ols_flag.
5. The method according to claim 3, wherein, The first identifier is NestingLayerId, and the second identifier is LayerIdInOls[targetOlsIdx].
6. The method according to claim 1, wherein, At least one of the SEI NAL units includes scalable nested SEI messages.
7. The method according to claim 6, wherein, The SEI NAL unit is removed from the output bitstream due to the removal of at least one VCL NAL unit from the output bitstream.
8. The method according to claim 6, wherein, The scalable nested SEI message includes a flag equal to zero.
9. The method according to claim 8, wherein, The flag is sn_subpic_flag.
10. The method according to any one of claims 1 to 9, wherein, The conversion includes decoding the video from the bitstream.
11. The method according to any one of claims 1 to 9, wherein, The conversion includes encoding the video into the bitstream.
12. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Perform the conversion between the video and the video bitstream. Wherein, the bitstream conforms to the order of the sub-bitstream extraction process defined by the rules, the rules specifying that (a) the sub-bitstream extraction process excludes the operation of removing all supplementary enhancement information SEI network abstraction layer NAL units that meet the conditions from the output bitstream, or (b) the order of the sub-bitstream extraction process includes: replacing the parameter set with a replacement parameter set before removing the SEI NAL unit from the output bitstream.
13. The apparatus according to claim 12, wherein, The condition specifies that the SEI NAL unit includes scalable nested SEI messages, which include (a) a flag equal to zero and (b) no value in a first list of the first identifier equal to a value in a second list of the second identifier. Wherein, the flag indicates whether the scalable nested SEI message is applicable to a specific set of output layers, wherein the first identifier is a nested layer identifier, and wherein the second identifier is a layer identifier of a layer in the set of output layers, and The flag is sn_ols_flag, and the first identifier is NestingLayerId, and the second identifier is LayerIdInOls[targetOlsIdx].
14. The apparatus according to claim 12, wherein, At least one of the SEI NAL units includes scalable nested SEI messages. Specifically, the removal of the SEI NAL unit from the output bitstream is caused by the removal of at least one Video Codec Layer (VCL) NAL unit from the output bitstream. The scalable nested SEI message includes a flag equal to zero, and the flag is sn_subpic_flag.
15. A non-transitory computer-readable storage medium for storing instructions that cause a processor to perform the following operations: Perform the conversion between the video and the video bitstream. in, The bitstream conforms to the order of the sub-bitstream extraction process defined by rules, which stipulate that (a) the sub-bitstream extraction process excludes the operation of removing all SEI network abstraction layer NAL units that meet the conditions from the output bitstream, or (b) the order of the sub-bitstream extraction process includes: replacing the parameter set with a replacement parameter set before removing the SEI NAL units from the output bitstream.
16. The non-transitory computer-readable storage medium according to claim 15, wherein, The condition specifies that the SEI NAL unit includes scalable nested SEI messages, which include (a) a flag equal to zero and (b) no value in a first list of the first identifier equal to a value in a second list of the second identifier. Wherein, the flag indicates whether the scalable nested SEI message is applicable to a specific set of output layers, wherein the first identifier is a nested layer identifier, and wherein the second identifier is a layer identifier of a layer in the set of output layers, and The flag is sn_ols_flag, and the first identifier is NestingLayerId, and the second identifier is LayerIdInOls[targetOlsIdx].
17. The non-transitory computer-readable storage medium according to claim 15, wherein, At least one of the SEI NAL units includes scalable nested SEI messages. Specifically, the removal of the SEI NAL unit from the output bitstream is caused by the removal of at least one Video Codec Layer (VCL) NAL unit from the output bitstream. The scalable nested SEI message includes a flag equal to zero, and the flag is sn_subpic_flag.
18. A non-transitory computer-readable recording medium storing a computer program / instructions and a bitstream thereon, wherein the computer program / instructions, when executed by a processor, implement a method for generating a video bitstream by a video processing apparatus, wherein... The method includes: The bitstream of the video is generated. Wherein, the bitstream conforms to the order of the sub-bitstream extraction process defined by the rules, the rules specifying that (a) the sub-bitstream extraction process excludes the operation of removing all supplementary enhancement information SEI network abstraction layer NAL units that meet the conditions from the output bitstream, or (b) the order of the sub-bitstream extraction process includes: replacing the parameter set with a replacement parameter set before removing the SEI NAL unit from the output bitstream.
19. The non-transitory computer-readable recording medium according to claim 18, wherein, The condition specifies that at least one of the SEI NAL units includes a scalable nested SEI message, the scalable nested SEI message including (a) a flag equal to zero and (b) no value in a first list of the first identifier equal to a value in a second list of the second identifier. Wherein, the flag indicates whether the scalable nested SEI message is applicable to a specific set of output layers, wherein the first identifier is a nested layer identifier, and wherein the second identifier is a layer identifier of a layer in the set of output layers, and The flag is sn_ols_flag, and the first identifier is NestingLayerId, and the second identifier is LayerIdInOls[targetOlsIdx].
20. The non-transitory computer-readable recording medium according to claim 18, wherein, At least one of the SEI NAL units includes scalable nested SEI messages. Specifically, the removal of the SEI NAL unit from the output bitstream is caused by the removal of at least one Video Codec Layer (VCL) NAL unit from the output bitstream. The scalable nested SEI message includes a flag equal to zero, and the flag is sn_subpic_flag.
21. A method for storing a video bitstream, comprising: The bitstream that generates the video, and The bitstream is stored in a non-transitory computer-readable storage medium. Wherein, the bitstream conforms to the order of the sub-bitstream extraction process defined by the rules, the rules specifying that (a) the sub-bitstream extraction process excludes the operation of removing all supplementary enhancement information SEI network abstraction layer NAL units that meet the conditions from the output bitstream, or (b) the order of the sub-bitstream extraction process includes: replacing the parameter set with a replacement parameter set before removing the SEI NAL unit from the output bitstream.
22. A video processing apparatus comprising a processor configured to implement the method of any one of claims 1 to 11.
23. A computer-readable medium having instructions stored thereon, which, when executed, cause a processor to perform the method of any one of claims 1 to 11.