Sub-bit stream extraction
The VVC sub-bitstream extraction process addresses the challenges of SEI handling and bitstream conformance in video coding standards by efficiently extracting and merging sub-bitstreams, reducing bitrate and enhancing decoding efficiency for viewport-dependent 360° video streaming and multi-layer applications.
Patent Information
- Application Number
- JP2025092282
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-15
AI Technical Summary
Existing video coding standards face challenges in efficiently handling supplemental enhancement information (SEI) messages and sub-bitstream extraction processes, particularly in supporting viewport-dependent 360° video streaming and multi-layer bitstreams, leading to complex bitstream conformance and high transmission bitrate.
Implementing a sub-bitstream extraction process in VVC that excludes the operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units and allows for replacing parameter sets prior to removing SEI NAL units, while ensuring HRD conformance and supporting sub-picture level information for efficient extraction and merging of sub-bitstreams.
Enables efficient extraction and merging of sub-bitstreams with reduced bitrate and improved decoding efficiency, supporting viewport-dependent 360° video streaming and multi-layer bitstreams by simplifying the bitstream conformance process.
Smart Images

Figure 2025120244000001_ABST
Abstract
Description
[Technical Field]
[0001] This application is a divisional application of Patent Application No. 2023-518401, which is based on International Patent Application No. PCT / CN2021 / 120531 filed on September 26, 2021, which claims priority to and the benefit of International Patent Application No. PCT / CN2020 / 117596 filed on September 25, 2020. All of the above patent applications are incorporated herein by reference in their entirety.
[0002] This patent document relates to the creation, storage and consumption of digital audio-video media information in a file format. [Background technology]
[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks, and as the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands for digital video usage are expected to continue to increase. Summary of the Invention
[0004] This document discloses techniques that can be used by video encoders and decoders to process coded representations of videos or images according to a file format.
[0005] In one example aspect, a video processing method is disclosed, comprising: performing conversion between a video and a bitstream of the video having multiple layers, the bitstream having multiple supplemental enhancement information (SEI) messages associated with a particular output layer set (OLS) or access units (AUs) or decoding units (DUs) of a particular layer, the multiple SEI messages having a message type different from a scalable nesting type based on a formatting rule, the formatting rule specifying that each of the multiple SEI messages has the same SEI payload content due to the multiple SEI messages being associated with the particular OLS or the AUs or DUs of the particular layer.
[0006] In another example aspect, another video processing method is disclosed, comprising: performing a conversion between a video having a current block and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying constraints on the bitstream due to syntax fields being excluded from the bitstream, the syntax field indicating a picture order count associated with a current picture having the current block or a current access unit having the current block.
[0007] In yet another example aspect, another video processing method is disclosed, comprising: performing a conversion between a video and a bitstream of the video, the bitstream conforming to an ordering of a sub-bitstream extraction process defined by rules, the rules specifying that (a) the sub-bitstream extraction process excludes an operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from an output bitstream, or (b) the ordering of the sub-bitstream extraction process includes replacing a parameter set with a replacement parameter set prior to removing the SEI NAL units from the output bitstream.
[0008] In yet another exemplary embodiment, a video encoder apparatus is disclosed, the video encoder comprising a processor configured to implement the above-described method.
[0009] In yet another example embodiment, a video decoder device is disclosed, the video decoder having a processor configured to implement the above-described method.
[0010] In yet another exemplary embodiment, a computer readable medium having stored thereon code, in the form of processor executable code, embodying one of the methods described herein, is disclosed.
[0011] In yet another exemplary embodiment, a computer-readable medium having stored thereon a bitstream is disclosed, the bitstream being generated or processed using the methods described herein.
[0012] These and other features are described throughout this document. [Brief explanation of the drawings]
[0013] [Figure 1] It shows a picture divided into 18 tiles, 24 slices, and 24 sub-pictures. [Figure 2] 1 illustrates a typical sub-picture-based viewport-dependent 360° video delivery scheme. [Figure 3] 1 shows an example of extracting one subpicture from a bitstream containing two subpictures and four slices. [Figure 4] FIG. 1 is a block diagram of an example of a video processing system. [Figure 5] FIG. 1 is a block diagram of a video processing device. [Figure 6] FIG. 1 is a block diagram illustrating a video coding system in accordance with some embodiments of the present disclosure. [Figure 7]FIG. 1 is a block diagram illustrating an encoder in accordance with some embodiments of the disclosed technology. [Figure 8] FIG. 2 is a block diagram illustrating a decoder in accordance with some embodiments of the disclosed technology. [Figure 9] 1 is a flowchart of an example method for video processing in accordance with some embodiments of the disclosed technology. [Figure 10] 1 is a flowchart of an example method for video processing in accordance with some embodiments of the disclosed technology. [Figure 11] 1 is a flowchart of an example method for video processing in accordance with some embodiments of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION
[0014] Section headings are used in this document for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section alone. Also, H.266 terminology is used in some descriptions solely for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein may also be applicable to other video codec protocols and designs.
[0015] 1. Opening remarks This document relates to video file formats. Specifically, this document relates to Picture Order Count (POC), Supplemental Enhancement Information (SEI), and Subpicture Sub-bitstream Extraction. The ideas can be applied individually or in various combinations to any video coding standard, such as the recently established Versatile Video Coding (VVC), or to non-standard video codecs.
[0016] 2. Abbreviations ACT adaptive color transform ALF adaptive loop filter AMVR adaptive motion vector resolution APS adaptation parameter set AU access unit AUD access unit delimiter AVC advanced video coding (Recommendation ITU-T H.264 | ISO / IEC 14496-10) B bi-predictive BCW bi-prediction with CU-level weights BDOF bi-directional optical flow BDPCM block-based delta pulse code modulation BP buffering period CABAC context-based adaptive binary arithmetic coding CB coding block CBR constant bit rate CCALF cross-component adaptive loop filter CLVS coded layer video sequence CLVSS coded layer video sequence start CPB coded picture buffer CRA clean random access CRC cyclic redundancy check CTB coding tree block CTU coding tree unit CU coding unit CVS coded video sequence CVSS coded video sequence start DPB decoded picture buffer DCI decoding capability information DRAP dependent random access point DU decoding unit DUI decoding unit information EG exponential-Golomb EGk k-th order exponential-Golomb EOB end of bitstream EOS end of sequence FD filler data (filler (dummy) data) FIFO first-in, first-out FL fixed-length GBR green, blue, and red GCI general constraints information GDR gradual decoding refresh GPM geometric partitioning mode HEVC high efficiency video coding (Recommendation ITU-T H.265 | ISO / IEC 23008-2) HRD hypothetical reference decoder HSS hypothetical stream scheduler I intra IBC intra block copy IDR instantaneous decoding refresh ILRP inter-layer reference picture IRAP intra random access point LFNST low frequency non-separable transform LPS least probable symbol LSB least significant bit LTRP long-term reference picture LMCS luma mapping with chroma scaling MIP matrix-based intra prediction MPS most probable symbol MSB most significant bit MTS multiple transform selection MVP motion vector prediction NAL network abstraction layer OLS output layer set OP operation point OPI operating point information P predictive PH picture header POC picture order count PPS picture parameter set PROF prediction refinement with optical flow PT picture timing PU picture unit QP quantization parameter RADL random access decodable leading (picture) RASL random access skipped leading (Random Access Skip Leading (Picture) RBSP raw byte sequence payload RGB red, green, and blue RPL reference picture list SAO sample adaptive offset SAR sample aspect ratio SEI supplemental enhancement information SH slice header SLI subpicture level information SODB string of data bits SPS sequence parameter set STRP short-term reference picture STSA step-wise temporal sublayer access TR truncated rice TU transform unit VBR variable bit rate VCL video coding layer VPS video parameter set VSEI versatile supplemental enhancement information (Recommendation ITU-T H.274 | ISO / IEC 23002-7) VUI video usability information VVC versatile video coding (Recommendation ITU-T H.266 | ISO / IEC 23090-3)
[0017] 3. Video Coding Review Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC, and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture that utilizes transform coding in addition to temporal prediction. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, numerous new methods have been adopted by JVET and incorporated into reference software named the Joint Exploration Model (JEM). JVET was later renamed the Joint Video Experts Team (JVET) when the VVC project was officially launched. VVC is a new coding standard established by JVET at its 19th meeting, which ended on July 1, 2020, that aims to achieve a 50% bitrate reduction compared to HEVC.
[0018] The VVC standard (ITU-T H.266 | ISO / IEC 23090-3) and the related VSEI standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for use in the widest range of applications, including both traditional uses such as television broadcasting, videoconferencing, or playback from storage media, and newer and more advanced uses such as adaptive bitrate streaming, video region extraction, content synthesis and fusion from multiple coded video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive 360° immersive media.
[0019] 3.1. Picture Order Count (POC) in HEVC and VVC In HEVC and VVC, the POC is essentially used as a picture ID to identify the picture in many parts of the decoding process, including DPB management, part of which is reference picture management.
[0020] With the newly introduced PH, VVC signals the POC least significant bit (LSB) information, which is used to derive the POC value and has the same value for all slices of a picture, in the PH, in contrast to HEVC, where it is signaled in the SH. VVC also enables signaling of the POC most significant bit (MSB) cycle value in the PH, allowing derivation of the POC value without tracking the POC MSB, which relies on the POC information of previously decoded pictures. This enables, for example, mixing IRAP and non-IRAP pictures within an AU of a multi-layer bitstream. A further difference in POC signaling between HEVC and VVC is that HEVC does not signal the POC LSB for IDR pictures, which was found to present some drawbacks during the development of later multi-layer extensions of HEVC to enable mixing IDR and non-IDR pictures within an AU. Therefore, in VVC, POC LSB information is signaled for each picture, including IDR pictures. Signaling POC LSB information for IDR pictures also makes it easier to support merging IDR and non-IDR pictures from different bitstreams into one picture, because otherwise, handling POC LSBs in a merged picture would require complex design.
[0021] The POC decoding process in VVC is specified as follows: 8.3.1 Decoding process for picture order count The output of this process is PicOrderCntVal, the picture order count of the current picture. Each coded picture is associated with a picture order count variable, denoted PicOrderCntVal. The variable currLayerIdx shall be set equal to GeneralLayerIdx[nuh_layer_id]. PicOrderCntVal is derived as follows: - if vps_independent_layer_flag[currLayerIdx] is equal to 0 and there exists a picture picA in the current AU with nuh_layer_id equal to layerIdA and GeneralLayerIdx[layerIdA] is in the list ReferenceLayerIdx[currLayerIdx], then PicOrderCntVal is derived to be equal to the PicOrderCntVal of picA and the value of ph_pic_order_cnt_lsb shall be the same in all VCL NAL units of the current AU. - Otherwise, the PicOrderCntVal of the current picture is derived as specified in the remainder of this subclause. When ph_poc_msb_cycle_val is not present and the current picture is not a CLVSS picture, the variables prevPicOrderCntLsb and prevPicOrderCntMsb are derived as follows: - Let prevTid0Pic be the previous picture in decoding order that has nuh_layer_id equal to the nuh_layer_id of the current picture, has TemporalId and ph_non_ref_pic_flag both equal to 0, and is not a RASL or RADL picture: NOTE 1 - In a sub-bitstream consisting only of intra pictures extracted from a single-layer bitstream and used in intra-picture-only trick-play, prevTid0Pic is the preceding intra picture in decoding order. To ensure correct POC derivation, the encoder may choose either to include ph_poc_msb_cycle_val for intra pictures or to set the value of sps_log2_max_pic_order_cnt_lsb_minus4 to be large enough so that the POC difference between the current picture and prevTid0Pic is less than MaxPicOrderCntLsb / 2; NOTE 2 - When vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0 for any value of i in the layer index and j equal to the layer index of the current layer, prevTid0Pic in the extracted sub-bitstream for some of the OLS is the preceding IRAP or GDR picture, in decoding order, in the current layer with ph_recovery_poc_cnt equal to 0. To ensure that such sub-bitstreams are conforming bitstreams, an encoder may choose either to include ph_poc_msb_cycle_val for each IRAP or GDR picture with ph_recovery_poc_cnt equal to 0, or to set the value of sps_log2_max_pic_order_cnt_lsb_minus4 large enough so that the POC difference between the current picture and prevTid0Pic is less than MaxPicOrderCntLsb / 2. - The variable prevPicOrderCntLsb is set equal to the ph_pic_order_cnt_lsb of prevTid0Pic. - The variable prevPicOrderCntMsb is set equal to the PicOrderCntMsb of prevTid0Pic. The variable PicOrderCntMsb for the current picture is derived as follows: if((ph_pic_order_cnt_lsb <prevPicOrderCntLsb)&& ((prevPicOrderCntLsb-ph_pic_order_cnt_lsb)>=(MaxPicOrderCntLsb / 2))) PicOrderCntMsb=prevPicOrderCntMsb+MaxPicOrderCntLsb (196) elseif((ph_pic_order_cnt_lsb>prevPicOrderCntLsb)&& ((ph_pic_order_cnt_lsb-prevPicOrderCntLsb)>(MaxPicOrderCntLsb / 2))) PicOrderCntMsb=prevPicOrderCntMsb-MaxPicOrderCntLsb else PicOrderCntMsb=prevPicOrderCntMsb PicOrderCntVal is derived as follows: PicOrderCntVal=PicOrderCntMsb+ph_pic_order_cnt_lsb (197) NOTE 3 - All CLVSS pictures for which ph_poc_msb_cycle_val is not present have PicOrderCntVal equal to ph_pic_order_cnt_lsb, since PicOrderCntMsb is set equal to 0 for those pictures. The value of PicOrderCntVal is -2 inclusive. 31 From 2 31 It should be in the range of -1. Within one CVS, the PicOrderCntVal values for any two coded pictures with the same value of nuh_layer_id shall not be the same. All pictures within any particular AU shall have the same value of PicOrderCntVal. The function PicOrderCnt(picX) is defined as follows: PicOrderCnt(picX) = PicOrderCntVal of picture picX (198) The function DiffPicOrderCnt(picA,picB) is defined as follows: DiffPicOrderCnt(picA,picB)=PicOrderCnt(picA)-PicOrderCnt(picB) (199) The bitstream is -2 inclusive for the value of DiffPicOrderCnt(picA,picB) used in the decoding process. 15 From 2 15 Data that results in a value outside the range -1 shall not be included: NOTE 4 - Let X be the current picture, and Y and Z be two other pictures in the same CVS. Y and Z are considered to be in the same output forward direction from X if DiffPicOrderCnt(X,Y) and DiffPicOrderCnt(X,Z) are both positive or both negative.
[0022] 3.2. VUI and SEI Messages VUI is a syntax structure transmitted as part of the SPS (and possibly also in the VPS in HEVC). VUI carries information that does not affect the standard decoding process but may be important for the proper rendering of the coded video.
[0023] The SEI supports processes related to decoding, display, or other purposes. Like the VUI, the SEI does not affect the standard decoding process. The SEI is carried within SEI messages. Decoder support for SEI messages is optional. However, SEI messages affect bitstream conformance (e.g., if the syntax of an SEI message in a bitstream does not conform to the specification, the bitstream will not conform), and some SEI messages are required by the HRD specification.
[0024] The VUI syntax structure and most SEI messages used with VVC are not specified in the VVC specification, but rather in the VSEI specification. SEI messages required for HRD conformance testing are specified in the VVC specification. VVC v1 defines five SEI messages related to HRD conformance testing, and VSEI v1 specifies 20 additional SEI messages. The SEI messages carried in the VSEI specification do not directly affect compliant decoder behavior and are defined to be used in a coding format-independent manner, allowing VSEI to be used with other video coding standards in addition to VVC in the future. Rather than specifically referencing VVC syntax element names, the VSEI specification references variables whose values are set within the VVC specification.
[0025] Compared to HEVC, the VUI syntax structure of VVC focuses only on information related to the proper rendering of pictures and does not include timing information or bitstream constraint indications. In VVC, the VUI is signaled within the SPS, which includes a length field before the VUI syntax structure to signal the length of the VUI payload in bytes. This allows decoders to easily jump between information and, more importantly, allows for convenient future VUI syntax extensions by adding new syntax elements directly to the end of the VUI syntax structure, in a manner similar to SEI message syntax extensions.
[0026] The VUI syntax structure includes the following information: · The content is interlaced or progressive; Whether the content contains frame-packed stereoscopic or projected omnidirectional video; Sample aspect ratio; · Whether the content is suitable for overscan display; Color descriptions, including primaries, matrices, and transfer characteristics, which are particularly important for being able to signal Ultra-High Definition (UHD) versus High-Definition (HD) color spaces as well as High Dynamic Range (HDR); Chroma position relative to luma (signaling clarified for progressive content compared to HEVC).
[0027] When an SPS does not contain any VUI, the information is considered unspecified and must be conveyed via external means or specified by the application when the bitstream's contents are intended for rendering on a display.
[0028] Table 1 lists all SEI messages specified in VVC v1 and their specifications, including their syntax and semantics. Of the 20 SEI messages specified in the VVC v1 specification, many (e.g., filler payload and both user data SEI messages) are inherited from HEVC. Some SEI messages are essential for the correct processing or rendering of coded video content. This is true, for example, for mastering display color volume, content light level information, or alternative transfer characteristic SEI messages that are particularly relevant to HDR content. Other examples include equirectangular projection, spherical rotation, region-wise packing, or omnidirectional viewport SEI messages, which are relevant to signaling and processing 360° video content. [Table 1] TIFF2025120244000003.tif215162
[0029] New SEI messages defined for VVC v1 include a frame-field information SEI message, a sample aspect ratio information SEI message, and a sub-picture level information SEI message.
[0030] The frame-field information SEI message contains information indicating how the associated picture should be displayed (e.g., field parity or frame repetition period), the source scan type of the associated picture, and whether the associated picture is a duplicate of a previous picture. This information, along with the timing information of the associated picture, used to be signaled in the picture timing SEI message in previous video coding standards. However, it has been noticed that frame-field information and timing information are two different types of information that are not necessarily signaled together. One typical example is to signal timing information at the system level, but signal frame-field information within the bitstream. Therefore, it was decided to remove frame-field information from the picture timing SEI message and signal it in a dedicated SEI message instead. This change also modifies the syntax of the frame-field information to allow for more explicit additional instructions to be conveyed to the display, such as field pairing or more values for frame repetition.
[0031] The sample aspect ratio SEI message allows signaling different sample aspect ratios for different pictures in the same sequence, while the corresponding information contained in the VUI applies to the entire sequence, which may be relevant when using reference picture resampling functions in conjunction with scaling factors that cause different pictures of the same sequence to have different sample aspect ratios.
[0032] The sub-picture level information SEI message provides information at the level of the sub-picture sequence.
[0033] 3.3. Picture Partitioning and Subpictures in VVC In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that cover a rectangular area of the picture. The CTUs within a tile are scanned in raster scan order within that tile.
[0034] A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture.
[0035] Two modes of slicing are supported: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a sequence of complete tiles in a tile raster scan of the picture. In rectangular slice mode, a slice contains either several complete tiles that collectively form a rectangular area of the picture, or several contiguous complete CTU rows of a tile that collectively form a rectangular area of the picture. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular area corresponding to the slice.
[0036] A subpicture contains one or more slices that collectively cover a rectangular area of the picture.
[0037] 3.3.1. Subpicture Concept and Function In VVC, each sub-picture consists of one or more complete rectangular slices that collectively cover a rectangular area of a picture, as shown, for example, in Figure 1. A sub-picture is either defined as extractable (i.e., coded independently of other sub-pictures of the same picture and of pictures that precede it in decoding order) or non-extractable. Whether a sub-picture is extractable or not, the encoder can control, for each sub-picture individually, whether in-loop filtering (including deblocking, SAO, and ALF) is applied across sub-picture boundaries.
[0038] Functionally, sub-pictures are similar to motion constrained tile sets (MCTS) in HEVC: they both allow independent coding and extraction of rectangular subsets of a sequence of pictures to be coded for use cases such as viewport-dependent 360° video streaming optimization and region of interest (ROI) applications.
[0039] In streaming 360° video, also known as omnidirectional video, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is rendered to the user at any particular moment, and the user can turn their head at any time to change their viewing direction and, therefore, their current viewport. While it is desirable to have at least a lower-quality representation of the areas not covered by the current viewport available at the client and ready to be rendered to the user in case the user suddenly changes their looking direction somewhere on the sphere, a high-quality representation of the omnidirectional video is only needed for the current viewport being rendered to the user at any given moment. Dividing the high-quality representation of the entire omnidirectional video into sub-pictures of appropriate granularity allows for an optimization such as that shown in Figure 1, with 12 high-resolution sub-pictures on the left and the remaining 12 sub-pictures of the omnidirectional video at lower resolution on the right.
[0040] Another exemplary sub-picture-based viewport-dependent 360° video delivery scheme is shown in Figure 2, where only the higher resolution full video representation is composed of sub-pictures, and the lower resolution full video representation can be coded with fewer RAPs than the higher resolution representation without using sub-pictures. The client receives the lower resolution full video, while for the higher resolution video, the client receives and decodes only the sub-pictures that currently cover the viewport.
[0041] 3.3.2. Differences between Subpicture and MCTS There are several important design differences between subpictures and MCTS. First, the subpicture feature in VVC allows motion vectors of coding blocks to point outside of a subpicture, even if the subpicture is extractable, by applying sample padding at subpicture boundaries as well as at picture boundaries. Second, additional modifications are introduced for motion vector selection and derivation in merge mode and the decoder-side motion vector refinement process of VVC. This allows for higher coding efficiency compared to non-standard motion constraints applied at the encoder side for MCTS. Third, when extracting one or more extractable subpictures from a sequence of pictures to create a conforming subbitstream, rewriting of the SH (and PH NAL units, if present) is not required. Subbitstream extraction based on the HEVC MCTS requires rewriting of the SH. Note that both HEVC MCTS extraction and VVC subpicture extraction require rewriting of the SPS and PPS. However, typically, only a few parameter sets exist in a bitstream, while each picture has at least one slice, so rewriting the SH can be a significant burden for application systems. Fourth, slices of different subpictures within a picture are allowed to have different NAL unit types. This is a feature often referred to as mixed NAL unit types or mixed subpicture types within a picture, as will be described in more detail below. Fifth, VVC specifies HRD and level definitions for subpicture sequences, so that the conformance of each sub-bitstream of an extractable subpicture sequence can be guaranteed by the encoder.
[0042] 3.3.3. Mixing Subpicture Types Within a Picture In AVC and HEVC, all VCL NAL units within a picture are required to have the same NAL unit type. VVC introduces the option to mix sub-pictures with specific different VCL NAL unit types within a picture, thus providing support for random access not only at the picture level but also at the sub-picture level. In VVC, VCL-NAL units within a sub-picture are still required to have the same NAL unit type.
[0043] The random access capability from IRAP subpictures is beneficial for 360° video applications. In a viewport-dependent 360° video distribution scheme similar to that shown in Figure 2, the contents of spatially adjacent viewports largely overlap, i.e., only a portion of the subpictures within a viewport are replaced by new subpictures during a viewport orientation change, while the majority of subpictures remain within the viewport. Although the subpicture sequence newly introduced to a viewport must start with an IRAP slice, a significant reduction in the overall transmission bitrate can be achieved when the remaining subpictures are allowed to perform inter-prediction during a viewport change.
[0044] An indication of whether a picture contains only one type of NAL unit or more than one type is provided in the PPS referenced by that picture (i.e., using a flag called pps_mixed_nalu_types_in_pic_flag). A picture can simultaneously have sub-pictures containing IRAP slices and sub-pictures containing trailing slices. A few other combinations of different NAL unit types within a picture are also allowed, including leading picture slices of NAL unit types RASL and RADL, which makes it possible to merge sub-picture sequences with open GOP and closed GOP coding structures extracted from different bitstreams into one bitstream.
[0045] 3.3.4. Subpicture Layout and ID Signaling The layout of subpictures in VVC is signaled in the SPS and is therefore constant in CLVS. Each subpicture is signaled by the location of its top-left CTU and its width and height in CTU numbers, thus ensuring that the subpicture covers a rectangular area of the picture at CTU granularity. The order in which the subpictures are signaled in the SPS determines the index of each subpicture within the picture.
[0046] To enable extraction and merging of subpicture sequences without rewriting the SH or PH, the slice addressing scheme in VVC is based on a subpicture ID and a subpicture-specific slice index to associate slices with subpictures. In SH, the subpicture ID and subpicture-level slice index of the subpicture containing the slice are signaled. Note that the value of the subpicture ID of a particular subpicture can be different from the value of its subpicture index. The mapping between these two is either signaled within the SPS or PPS (but never both), or implicitly inferred. If present, the subpicture ID mapping needs to be rewritten or added when rewriting the SPS and PPS in the subpicture sub-bitstream extraction process. Both the subpicture ID and the subpicture-level slice index indicate to the decoder the exact location of the first decoded CTU of a slice within the DPB slot of the decoded picture. After sub-bitstream extraction, the subpicture ID of a subpicture remains unchanged, but the subpicture index may change. Even if the raster scan CTU address of the first CTU in a slice in a subpicture changes compared to its value in the original bitstream, the unchanged values of the subpicture ID and subpicture level slice index in each SH still accurately determine the location of each CTU in the decoded picture of the extracted subbitstream. Figure 3 illustrates the use of the subpicture ID, subpicture index, and subpicture level slice index to enable subpicture extraction using an example including two subpictures and four slices.
[0047] Similar to sub-picture extraction, signaling about sub-pictures allows merging several sub-pictures from different bitstreams into one bitstream by simply rewriting the SPS and PPS, provided that the different bitstreams are generated cooperatively (e.g., using distinct sub-picture IDs, but otherwise using mostly aligned SPS, PPS, and PH parameters, e.g., CTU size, chroma format, coding tool, etc.).
[0048] Although sub-pictures and slices are signaled independently within the SPS and PPS, respectively, there are inherent inter-constraints between the sub-picture layout and the slice layout in order to form a conforming bitstream. First, the presence of sub-pictures requires the use of rectangular slices and prohibits raster-scan slices. Second, the slices of a given sub-picture are assumed to be consecutive NAL units in decoding order, which means that the sub-picture layout constrains the order of coded slice NAL units in the bitstream.
[0049] 3.4. General Sub-Bitstream Extraction Process in VVC Similar to HEVC, the VVC specification includes a sub-bitstream extraction process that allows for the extraction of sub-bitstreams corresponding to specific operation points (i.e., the OLS and included temporal sublayers). While in HEVC the extraction process is part of the decoding process, i.e., the decoder needs to discard NAL units that are not associated with the operation point when present in the bitstream, the VVC design assumes that the bitstream provided to the decoder does not contain NAL units that do not belong to the indicated operation point, i.e., NAL units that are not associated with the operation point are discarded, when necessary, by an extractor that is not part of the decoder. A notable difference compared to HEVC is that in VVC, the handling of scalable nested HRD SEI messages is standardized, e.g., the extracted bitstream carries the correct HRD timing parameters for the target operation point. The process involves removing the original BP SEI message and, if any, the DUI SEI message, and inserting the appropriate SEI message that was originally included in the scalable nesting SEI message when the target operation point does not include all layers in the bitstream. However, PT SEI messages have special handling and the above actions are only required if each PT SEI message does not indicate that it applies to all OLSs.
[0050] The general sub-bitstream extraction process specification in VVC is as follows: C.6 General sub-bitstream extraction process The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target highest TemporalId value tIDTarget. The output of this process is the sub-bitstream outBitstream. The OLS with OLS index targetOlsIdx is called the target OLS. The requirement for a bitstream to be conformant with respect to an input bitstream is that any output sub-bitstream that satisfies all of the following conditions is a conforming bitstream: - The output sub-bitstream is the output of the process specified in this subclause, which took as input the bitstream, targetOlsIdx equal to an index into the list of OLSs specified by the VPS, and tIDTarget equal to any value in the range from 0 to vps_ptl_max_tid[vps_ols_ptl_idx[targetOlsIdx]], inclusive. - The output sub-bitstream contains at least one VCL NAL unit with nuh_layer_id equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx]. - the output sub-bitstream contains at least one VCL NAL unit with TemporalId equal to tIDTarget: NOTE − A conforming bitstream contains one or more coded slice NAL units with TemporalId equal to 0, but need not contain any coded slice NAL units with nuh_layer_id equal to 0. The output sub-bitstream OutBitstream is derived by applying the following ordered steps: 1. The bitstream outBitstream is set to be identical to the bitstream inBitstream. 2. Remove all NAL units with TemporalId greater than tIDTarget from outBitstream. 3. Remove from outBitstream all NAL units that have a nuh_layer_id that is not included in the list LayerIdInOls[targetOlsIdx], and that are not DCI, OPI, VPS, AUD, or EOB NAL units, and that are not SEI NAL units containing non-scalable nested SEI messages with payloads of PayloadType equal to 0, 1, 130, or 203. 4. Remove from outBitstream all APS and VCL NAL units for which all of the following conditions are true, and their associated non-VCL NAL units that had a nal_unit_type equal to PH_NUT or FD_NUT, or that contained an SEI message with a PayloadType equal to SUFFIX_SEI_NUT or PREFIX_SEI_NUT and not equal to any of 0 (BP), 1 (PT), 130 (DUI), and 203 (SLI): - nal_unit_type is equal to APS_NUT, TRAIL_NUT, STSA_NUT, RADL_NUT, or RASL_NUT, or nal_unit_type is equal to GDR_NUT and the associated ph_recovery_poc_cnt is greater than 0; - TemporalId is greater than or equal to NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]]. 5. When all VCL NAL units of an AU have been removed by steps 2, 3, or 4 above and an AUD or OPI NAL unit is present in the AU, remove the AUD or OPI NAL unit from outBitstream. 6. For each OPI NAL unit in outBitstream, set opi_htid_info_present_flag equal to 1, set opi_ols_info_present_flag equal to 1, set setopi_htid_plus1 equal to tIdTarget+1, and set opi_ols_idx equal to targetOlsIdx. 7. When an AUD is present within an AU in outBitstream and the AU becomes an IRAP or GDR AU, set the AUD's aud_irap_or_gdr_flag equal to 1. 8. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages with sn_ols_flag equal to 1 and no value of i in the range from 0 to sn_num_olss_minus1, inclusive, such that NestingOlsIdx[i] is equal to targetOlsIdx. 9. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages that have sn_ols_flag equal to 0 and that do not have a value in the list NestingLayerId equal to a value in the list LayerIdInOls[targetOlsIdx]. 10. When LayerIdInOls[targetOlsIdx] does not contain all values of nuh_layer_id in all VCL NAL units in bitstream inBitstream, the following apply in the order listed: a. Remove all SEI NAL units, including non-scalable nested SEI messages, with payloadType equal to 0 (BP), 130 (DUI), or 203 (SLI), from outBitstream; b. When general_same_pic_timing_in_all_ols_flag is equal to 0, remove all SEI NAL units, including non-scalable nested SEI messages, with payloadType equal to 1(PT) from outBitstream; c. When outBitstream contains an SEI NAL unit seiNalUnitA containing a scalable nesting SEI message with sn_ols_flag equal to 1 or sn_subpic_flag equal to 0 that applies to the target OLS, or when NumLayersInOls[targetOlsIdx] is equal to 1 and outBitstream contains an SEI NAL unit seiNalUnitA containing a scalable nesting SEI message with sn_ols_flag equal to 0 and sn_subpic_flag equal to 0 that applies to a layer in outBitstream, generate a new SEI NAL unit seiNalUnitB and include it in the PU that contains seiNalUnitA immediately after seiNalUnitA, extract the scalable nested SEI messages from the scalable nesting SEI message and include them directly in seiNalUnitB (as non-scalable nested SEI messages), and remove seiNalUnitA from outBitstream.
[0051] 3.5. Subpicture Sub-bitstream Extraction Process in VVC VVC allows for HRD conformance testing of each independently coded sub-picture, i.e., the bitstream portion associated with each sub-picture can be extracted to form a valid bitstream, and the bitstream can then be tested for conformance to the HRD model. Conformance testing requires the definition of a complete HRD model for such sub-picture sub-bitstreams, and VVC allows for the necessary information to be carried, in addition to other HRD parameters, by a new SEI message called the Sub-picture Level Information (SLI) SEI message.
[0052] The SLI SEI message provides level information for subpicture sequences, which is required to derive the CPB size and bitrate values of the HRD model that describes the processing of subpicture bitstreams by a decoder. An original bitstream consisting of multiple subpictures may adhere to the limits defined by a particular level indicated in the bitstream's parameter set (e.g., level 5.1 for 4K at 60 Hz), while a subpicture sub-bitstream of that particular bitstream may correspond to a lower level (e.g., level 3 for 720p at 60 Hz). Furthermore, VVC allows the level of a particular subpicture sub-bitstream to be expressed as a fraction of a reference level, which allows for level signaling with finer granularity than HEVC, and this information also provides systems that merge multiple subpictures into a single joint bitstream with guidance on how much each subpicture sub-bitstream contributes to the level limit of the merged bitstream. Additional characteristics, such as, for example, the level contributions of constant bitrate bitstreams or layers for which sub-picture partitioning does not apply in the case of multi-layer bitstreams, can also be signaled in the SLI SEI message, thus allowing the derivation of a complete HRD model for each individual sub-picture sequence, even for such scenarios. A further part of the effort to enable conforming sub-picture sub-bitstreams is the application of a number of conformance-related constraints with picture scope as well as sub-picture scope to the VCL NAL units belonging to individual sub-pictures, such as constraints on minimum compression ratios or bin-to-bit ratios.
[0053] In HEVC, sub-bitstream extraction and conformance testing of independently coded regions (i.e., MCTS) requires that parameter sets for such sub-bitstreams be carried in a nested form by the MCTS extraction information set SEI message.
[0054] VVC adds a new subpicture sub-bitstream extraction process that enables the generation of conforming bitstreams from independently coded subpictures by removing unnecessary NAL units associated with other subpictures and actively rewriting the relevant portions of the parameter sets to correctly reflect the characteristics of the subpicture sub-bitstream. For example, this process involves rewriting the level indicator and HRD parameters in the VPS and SPS, and the picture size, partitioning information, adaptation window offset, scaling window offset, and virtual boundary position in the appropriate sections of the respective parameter sets. The subpicture sub-bitstream extraction process rewrites some of the information in the bitstream's parameter set based on information provided in the SLI SEI message.
[0055] The specifications for the sub-picture sub-bitstream extraction process are as follows: C.7 Subpicture Sub-Bitstream Extraction Process The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, the target highest TemporalId value tIDTarget, and a list of target subpicIdxTarget[i] for i from 0 to NumLayersInOls[targetOlsIdx]-1, inclusive. The output of this process is the sub-bitstream outBitstream. The OLS with OLS index targetOlsIdx is referred to as the target OLS. Layers in the target OLS with a referenced SPS of sps_num_subpics_minus1 greater than 0 are referred to as multiSubpicLayers. The requirement for a bitstream to be conformant with respect to an input bitstream is that any output sub-bitstream that satisfies all of the following conditions is a conforming bitstream: - The output sub-bitstream is the output of the process specified in this subclause, which takes as input the bitstream, targetOlsIdx equal to an index into the list of OLSs specified by the VPS, tIDTarget equal to any value in the range 0 to vps_max_sublayers_minus1, inclusive, and a list subpicIdxTarget[i], for i between 0 and NumLayersInOls[targetOlsIdx]-1, inclusive, such that: - the value of subpicIdxTarget[i] is equal to a value in the range 0 to sps_num_subpics_minus1, inclusive, such that sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] is equal to 1, and sps_num_subpics_minus1 and sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] are found in or estimated based on the SPS referenced by the layer with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i]: NOTE 1 - When sps_num_subpics_minus1 for a layer with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] is equal to 0, the value of subpicIdxTarget[i] is equal to 0; - For any two distinct integer values of m and n, subpicIdxTarget[m] is equal to subpicIdxTarget[n] when sps_num_subpics_minus1 is greater than 0 for both layers with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][m] and LayerIdInOls[targetOlsIdx][n], respectively. - The output sub-bitstream contains at least one VCL NAL unit with nuh_layer_id equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx]. - the output sub-bitstream contains at least one VCL NAL unit with TemporalId equal to tIDTarget: NOTE 2 - A conforming bitstream contains one or more coded slice NAL units with TemporalId equal to 0, but need not contain any coded slice NAL units with nuh_layer_id equal to 0. - The output sub-bitstream contains at least one VCL NAL unit with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id equal to SubpicIdVal[subpicIdxTarget[i]], for each i in the range from 0 to NumLayersInOls[targetOlsIdx]-1, inclusive. The output sub-bitstream outBitstream is derived by the following ordered steps: 1. The sub-bitstream extraction process specified in Appendix C.6 is invoked with inBitstream, targetOlsIdx, and tIDTarget as inputs, and the output of the process is assigned to outBitstream. 2. For each value of i in the range from 0 to NumLayersInOls[targetOlsIdx]-1, inclusive, remove from outBitstream all VCL NAL units with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id not equal to SubpicIdVal[subpicIdxTarget[i]], their associated filler data NAL units, and their associated SEI NAL units containing filler payload SEI messages. 3. When there is an SLI SEI message that applies to the target OLS and the sli_cbr_constraint_flag of that SLI SEI message is equal to 0, remove all NAL units with nal_unit_type equal to FD_NUT and SEI NAL units that contain filler payload SEI messages. 4. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages that have sn_subpic_flag equal to 1 and none of the sn_subpic_id[j] values for j between 0 and sn_num_subpics_minus1, inclusive, equals any of the SubpicIdVal[subpicIdxTarget[i]] values for any layer in multiSubpicLayers. 5. When at least one VCL NAL unit has been removed by step 2, remove all SEI NAL units, including scalable nesting SEI messages with sn_subpic_flag equal to 0, from outBitstream. 6. If some external means not specified in this document is available to provide a replacement parameter set for the sub-bitstream outBitstream, replace all parameter sets with the replacement parameter set. Otherwise, the following ordered steps apply: a. The variable spIdx is set equal to the value of subpicIdxTarget[i] for one of the layers in multiSubpicLayers; b. when an SLI SEI message that applies to the target OLS exists, in the vps_ols_ptl_idx[targetOlsIdx]th entry in the list of profile_tier_level() syntax structures in all referenced VPSs, when present, and in the profile_tier_level() syntax structures in all referenced SPSs when NumLayersInOls[targetOlsIdx] is equal to 1, set the values of general_level_idc and sublayer_level_idc[k] for k in the range from 0 to tIDTarget-1, inclusive, to SubpicLevelIdc[spIdx][tIDTarget] and SubpicLevelIdc[spIdx][k], respectively, derived by equation 1621 for the spIdxth subpicture sequence; c. When there is an SLI SEI message that applies to the target OLS, let spLvIdx be set equal to SubpicLevelIdx[spIdx][k], for k in the range 0 to tIDTarget, inclusive, where SubpicLevelIdx[spIdx][k] is derived by equation 1621 for the spIdx-th subpicture sequence. VCL HRD parameters or NAL When an HRD parameter exists, for k in the range 0 to tIDTarget, inclusive, the cpb_size_value_minus1[k][j] and bit_rate_value_minus1[k][j] of the jth CPB in the vps_ols_timing_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]]th ols_timing_hrd_parameters() syntax structure in all referenced VPSs, when present, and in the ols_timing_hrd_parameters() syntax structure in all referenced SPSs, when NumLayersInOls[targetOlsIdx] is equal to 1. 1618], where j ranges from 0 to hrd_cpb_cnt_minus1, inclusive, and i ranges from 0 to NumLayersInOls[targetOlsIdx]-1, inclusive; d. For each layer in the multiSubpicLayers, the following ordered steps are applied to rewrite the SPS and PPS referenced by pictures in that layer: i. The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are derived as follows: subpicWidthInLumaSamples=Min((sps_subpic_ctu_top_left_x[spIdx]+ sps_subpic_width_minus1[spIdx]+1)*CtbSizeY, pps_pic_width_in_luma_samples)- sps_subpic_ctu_top_left_x[spIdx]*CtbSizeY (1600) subpicHeightInLumaSamples=Min((sps_subpic_ctu_top_left_y[spIdx]+ sps_subpic_height_minus1[spIdx]+1)*CtbSizeY, pps_pic_height_in_luma_samples)- sps_subpic_ctu_top_left_y[spIdx]*CtbSizeY (1601) ii. Set the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPSs and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPSs equal to subpicWidthInLumaSamples and subpicHeightInLumaSamples, respectively; iii. Set the values of sps_num_subpics_minus1 in all referenced SPSs and pps_num_subpics_minus1 in all referenced PPSs equal to 0; iv. When present, set the values of the syntax elements sps_subpic_ctu_top_left_x[spIdx] and sps_subpic_ctu_top_left_y[spIdx] in all referenced SPSs equal to 0; v. For each j not equal to spIdx, remove the syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j], and sps_subpic_id[j] in all referenced SPSs, when present; vi. When spIdx is greater than 0 and sps_subpic_id_mapping_explicitly_signalled_flag of the referenced SPS is equal to 0, set the values of sps_subpic_id_mapping_explicitly_signalled_flag and sps_subpic_id_mapping_present_flag both equal to 1, and add sps_subpic_id[0] equal to spIdx to the SPS; vii. For each j not equal to spIdx, remove the syntax element pps_subpic_id[j] in all referenced PPSs, when present; viii. configure syntax elements in all referenced PPSs for tile and slice signaling to remove all tile rows, tile columns, and slices that are not associated with a subpicture with a subpicture index equal to spIdx; ix. The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are derived as follows: subpicConfWinLeftOffset=sps_subpic_ctu_top_left_x[spIdx]==0? Sps_conf_win_left_offset:0 (1602) subpicConfWinRightOffset=(sps_subpic_ctu_top_left_x[spIdx]+ sps_subpic_width_minus1[spIdx]+1)*CtbSizeY>= (1603) sps_pic_width_max_in_luma_samples?sps_conf_win_right_offset:0 subpicConfWinTopOffset=sps_subpic_ctu_top_left_y[spIdx]==0? sps_conf_win_top_offset:0 (1604) subpicConfWinBottomOffset=(sps_subpic_ctu_top_left_y[spIdx]+ sps_subpic_height_minus1[spIdx]+1)*CtbSizeY>= (1605) sps_pic_height_max_in_luma_samples?sps_conf_win_bottom_offset:0 where the values of sps_subpic_ctu_top_left_x[spIdx]x[], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, sps_height_max_in_luma_samples, sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, sps_conf_win_bottom_offset in these formulas are the values from the original SPS before they were rewritten: NOTE 3 - For pictures in layers of multiSubpicLayers in both the input bitstream and the output bitstream, the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are equal to pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples, respectively. Therefore, in these formulas, sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples may be replaced by pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples, respectively; x. Set the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in all referenced SPSs equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively; xi. The variables subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset are derived as follows: subpicScalWinLeftOffset = pps_scaling_win_left_offset - (1606) sps_subpic_ctu_top_left_x[spIdx] * CtbSizeY / SubWidthC rightSubpicBd = (sps_subpic_ctu_top_left_x[spIdx] + sps_subpic_width_minus1[spIdx] + 1) * CtbSizeY subpicScalWinRightOffset = (rightSubpicBd >= sps_pic_width_max_in_luma_samples)? (1607) pps_scaling_win_right_offset : pps_scaling_win_right_offset - (sps_pic_width_max_in_luma_samples - rightSubpicBd) / SubWidthC subpicScalWinTopOffset = pps_scaling_win_top_offset - (1608) sps_subpic_ctu_top_left_y[spIdx] * CtbSizeY / SubHeightC botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] + sps_subpic_height_minus1[spIdx] + 1) * CtbSizeY subpicScalWinBotOffset = (botSubpicBd >= sps_pic_height_max_in_luma_samples)? (1609) pps_scaling_win_bottom_offset:pps_scaling_win_bottom_offset- (sps_pic_height_max_in_luma_samples-botSubpicBd) / SubHeightC where the values of sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, and sps_pic_height_max_in_luma_samples in these formulas are the values from the original SPS before they were rewritten, and the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in these formulas are the values from the original PPS before they were rewritten; xii. setting the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in all referenced PPS NAL units equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset, respectively; xiii. The variables numVerVbs, subpicVbx[i], numHorVbs, and subpicVby[i] are derived as follows: numVerVbs=0; subpicX=sps_subpic_ctu_top_left_x[spIdx] for(i=0;i<sps_num_ver_virtual_boundaries;i++){ vbX=sps_virtual_boundary_pos_x_minus1[i]+1 if(vbX>(subpicX*CtbSizeY / 8)&&vbX<Min((subpicX+ (1610) sps_subpic_width_minus1[spIdx]+1)*CtbSizeY / 8, pps_pic_width_in_luma_samples / 8)) subpicVbx[numVerVbs++]=vbX-subpicX*CtbSizeY / 8 } numHorVbs=0; subpicY=sps_subpic_ctu_top_left_y[spIdx] for(i=0;i<sps_num_hor_virtual_boundaries;i++){ vbY=sps_virtual_boundary_pos_y_minus1[i]+1 if(vbY>(subpicY*CtbSizeY / 8)&&vbY<Min((subpicY+ (1611) sps_subpic_height_minus1[spIdx]+1)*CtbSizeY / 8, pps_pic_height_in_luma_samples / 8)) subpicVby[numHorVbs++]=vbY-subpicY*CtbSizeY / 8 } where the values of sps_num_ver_virtual_boundaries, sps_virtual_boundary_pos_x_minus1[i], sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_num_hor_virtual_boundaries, sps_virtual_boundary_pos_y_minus1[i], sps_subpic_ctu_top_left_y[spIdx], and sps_subpic_height_minus1[spIdx] in these equations are the values from the original SPS before they were rewritten; xiv. When sps_virtual_boundaries_present_flag is equal to 1, for i in the range 0 to numVerVbs-1, inclusive, and j in the range 0 to numHorVbs-1, inclusive, set the values of the sps_num_ver_virtual_boundaries, sps_virtual_boundary_pos_x_minus1[i], sps_num_hor_virtual_boundaries, and sps_virtual_boundary_pos_y_minus1[j] syntax elements in all reference SPSs to be equal to numVerVbs, subpicVbx[i]-1, numHorVbs, and subpicVby[j]-1, respectively. Virtual boundaries outside the extracted subpictures are removed. When both numVerVbs and numHorVbs are equal to 0, set the value of sps_virtual_boundaries_enabled_flag in all referenced SPSs to be equal to 0 and remove the syntax elements sps_virtual_boundaries_present_flag, sps_num_ver_virtual_boundaries, sps_virtual_boundary_pos_x_minus1[i], sps_num_hor_virtual_boundaries, and sps_virtual_boundary_pos_y_minus1[i]; e. When there is an SLI SEI message that applies to the target OLS, the following applies: i. If sli_cbr_constraint_flag is equal to 1, set cbr_flag[tIDTarget][j] of the jth CPB to 1 in the vps_ols_timing_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]]th ols_timing_hrd_parameters() syntax structure in all referenced VPSs, and in the ols_timing_hrd_parameters() syntax structure in all referenced SPSs when NumLayersInOls[targetOlsIdx] is equal to 1; ii. Otherwise (sli_cbr_constraint_flag is equal to 0), set cbr_flag[tIDTarget][j] equal to 0, where j is in the range 0 to hrd_cpb_cnt_minus1, inclusive; 7. When at least one VCL NAL unit has been removed by step 2, the following apply in the order listed: a. Remove all SEI NAL units, including non-scalable nested SEI messages, with payloadType equal to 0 (BP), 130 (DUI), 203 (SLI), or 132 (decoded picture hash) from outBitstream; b. When general_same_pic_timing_in_all_ols_flag is equal to 0, remove all SEI NAL units, including non-scalable nested SEI messages, with payloadType equal to 1(PT) from outBitstream; c. When outBitstream contains an SEI NAL unit seiNalUnitA containing a scalable nesting SEI message with sn_ols_flag equal to 1 or sn_subpic_flag equal to 1 that applies to the target OLS and subpicture in outBitstream, or when NumLayersInOls[targetOlsIdx] is equal to 1 and outBitstream contains an SEI NAL unit seiNalUnitA containing a scalable nesting SEI message with sn_ols_flag equal to 0 and sn_subpic_flag equal to 1 that applies to the layer and subpicture in outBitstream, a new SEI Generate NAL unit seiNalUnitB and include it in the PU that contains seiNalUnitA immediately after seiNalUnitA, extract the scalable nested SEI messages from the scalable nesting SEI message and include them directly in seiNalUnitB (as non-scalable nested SEI messages), and remove seiNalUnitA from the outBitstream.
[0056] 4. Examples of technical problems solved by the disclosed technical solutions The design of POC, SEI, and sub-picture sub-bitstream extraction in VVC has the following problems: 1) Section C.4 of the VVC specification contains the following constraints related to POC: Let currPicLayerId equal to the nuh_layer_id of the current picture; For each current picture, the variables maxPicOrderCnt and minPicOrderCnt shall be set equal to the maximum and minimum, respectively, of the PicOrderCntVal values of the following pictures with nuh_layer_id equal to currPicLayerId: - Current Picture; - a leading picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture; - STRP referenced by all entries in RefPicList[0] and all entries in RefPicList[1] of the current picture; - all pictures n with current picture currPic having PictureOutputFlag equal to 1, AuCpbRemovalTime[n] less than AuCpbRemovalTime[currPic], and DpbOutputTime[n] greater than or equal to AuCpbRemovalTime[currPic]; For each current picture that is not a CLVSS picture, the value of maxPicOrderCnt-minPicOrderCnt shall be less than MaxPicOrderCntLsb / 2; The above constraints do not allow the following bitstreams, which would be conforming bitstreams if the second bullet above were relaxed: a. The original bitstream is a single-layer bitstream, and a sub-bitstream outBitstream is extracted from the bitstream by removing all non-intra-coded pictures. In outBitstream, there is at least one picture picA such that the difference in POC value between picture picA and picture prevTid0Pic, i.e., the preceding picture in decoding order in the same layer that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture, is greater than MaxPicOrderCntLsb / 2. According to the above constraint, this bitstream outBitstream is not a conforming bitstream because it violates the constraint; However, if picA has the POC MSB value signaled in PH, its POC can be correctly derived. Therefore, the bitstream outBitstream can be a conforming bitstream if the second item is changed as follows: - a preceding picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture when ph_poc_msb_cycle_val is not present for the current picture; b. The original bitstream is a multi-layer bitstream, a layer with layer index j is a direct reference layer of another layer with layer index layerB in the bitstream, and vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0. Then, in the extracted bitstream outBitstream, according to the general sub-bitstream extraction process, for an OLS where layer j is an output layer but layer i is not, only GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0 exist in layer i. In this case, similarly to above, in outBitstream, there exists at least one picture picA in layer i such that the difference in POC value between picture picA and picture prevTid0Pic, i.e., the preceding picture in the same layer in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture, is greater than MaxPicOrderCntLsb / 2. According to the above constraint, this bitstream outBitstream is not a conforming bitstream because it violates the constraint: an extracted bitstream such as outBitstream is required to be a conforming bitstream if the original bitstream was a conforming bitstream, and therefore the original bitstream is also not a conforming bitstream; However, if picA has the POC MSB value signaled in the PH, its POC can be correctly derived. Again, both the bitstream outBitstream and the original bitstream in this example can be conforming bitstreams if the second item is changed as follows: - a preceding picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture when ph_poc_msb_cycle_val is not present for the current picture; c. The original bitstream and outBitstream are similar to those in examples a and b above, but contain multiple layers. For a picture picA, there is another picture picB in the same AU, which belongs to the reference layer of the layer containing picA. In other words, the POC value of picA will be derived equal to the POC value of picB. In this case, even if the picture POC difference between picA and prevTid0Pic is greater than MaxPicOrderCntLsb / 2, the POC value of picA can still be derived correctly as long as the POC value of picB can be derived correctly; Thus, the bitstream outBitstream in this example could be a conforming bitstream if the second bullet above was changed to: - When there is no picture in the current AU that belongs to a reference layer of the current layer, a leading picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture. 2) In the general sub-bitstream extraction process, step 9 removes from outBitstream all SEI NAL units containing scalable nesting SEI messages that have sn_ols_flag equal to 0 and that do not have a value in the list NestingLayerId equal to a value in the list LayerIdInOls[targetOlsIdx]. However, this step is not actually necessary because: 1) For SEI NAL units containing scalable nesting SEI messages that have sn_ols_flag equal to 0, NestingLayerId always contains the layer ID of the SEI NAL unit; and 2) if an SEI NAL unit has a layer ID that is not in the list LayerIdInOls[targetOlsIdx], it will have already been removed by step 3. 3) In the subpicture sub-bitstream extraction process, when at least one VCL NAL unit has been removed by step 2, step 5 removes from outBitstream all SEI NAL units containing scalable nesting SEI messages with sn_subpic_flag equal to 0, including nested SLI SEI messages. However, nested SLI SEI messages, when present, may be needed by step 6 for parameter set rewriting. 4) Both nested and non-nested HRD-related SEI messages are allowed to exist and apply to the same OLS. Similarly, both nested and non-nested non-HRD-related SEI messages are allowed to exist and apply to the same layer. However, there is no constraint requiring the content of such SEI messages of a particular payload type to be the same.
[0057] 5. Examples of technical solutions To solve the above problems and others, the following summarized methods are disclosed. These items should be considered as examples to illustrate the overall concept and should not be construed narrowly. Also, these items can be applied individually or in any combination: 1) To solve Problem 1, it is proposed to change the second item in the POC-related constraints described in Problem 1 to one of the following: a. When ph_poc_msb_cycle_val is not present for the current picture and there is no picture belonging to the reference layer of the current layer in the current AU, a leading picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture: i. Alternatively, the phrase "there is no picture belonging to the reference layer of the current layer in the current AU" is replaced by "PocFromIlrpFlag of the current picture is equal to 0", and PocFromIlrpFlag is defined as follows: The variable currLayerIdx is set equal to GeneralLayerIdx[nuh_layer_id]. If vps_independent_layer_flag[currLayerIdx] is equal to 0 and there is a picture picA in the current AU with nuh_layer_id equal to layerIdA whose GeneralLayerIdx[layerIdA] is in the list ReferenceLayerIdx[currLayerIdx], then the variable PocFromIlrpFlag is set equal to 1. Otherwise, PocFromIlrpFlag is set equal to 0; b. When ph_poc_msb_cycle_val is not present for the current picture, the preceding picture in decoding order has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture; c. When there is no picture in the current AU that belongs to the reference layer of the current layer, the preceding picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture: i. Alternatively, the phrase "there is no picture belonging to the reference layer of the current layer in the current AU" is replaced by "PocFromIlrpFlag of the current picture is equal to 0", and PocFromIlrpFlag is defined as above; d. When ph_poc_msb_cycle_val does not exist for the current picture, or when there is no picture belonging to the reference layer of the current layer in the current AU, the preceding picture in decoding order that has TemporalId and ph_non_ref_pic_flag both equal to 0 and is not a RASL or RADL picture: i. Alternatively, the phrase "there is no picture belonging to the reference layer of the current layer in the current AU" is replaced by "PocFromIlrpFlag of the current picture is equal to 0", and PocFromIlrpFlag is defined as above. 2) To solve problem 2, we eliminate step 9 of the general sub-bitstream extraction process. 3) To solve problem 3, step 5 of the sub-picture sub-bitstream extraction process is moved to after step 6 entirely. 4) To solve problem 4, the following constraint is proposed: When there are multiple SEI messages with a particular value of payloadType not equal to 133 associated with a particular AU or DU and applied to a particular OLS or layer, the SEI messages shall have the same SEI payload content, regardless of whether some or all of these SEI messages are scalably nested. Note that an SEI message with payloadType equal to 133 is a scalable nesting SEI message.
[0058] 6. Implementation Below are some example embodiments of some of the inventive aspects summarized in Section 5 that can be applied to the VVC specification. The text subject to change is based on the latest VVC text in JVET-S2001-vH. The most relevant parts that have been added or changed are shown in bold, italic, and underlined font, e.g., "feature" in bold, italic, and underlined indicates an addition, and some parts that have been deleted are surrounded by double brackets in bold, italic, e.g., "feature" surrounded by "bold, italic, [[]]" indicates a deletion. There may be some changes that are not highlighted because they are editorial in nature. 6.1. First embodiment This embodiment relates to items 1.ai, 2, 3, and 4. (outside 1) TIFF2025120244000004.tif219170(outside 2) TIFF2025120244000005.tif154168(outside 3) TIFF2025120244000006.tif229170(outside 4) TIFF2025120244000007.tif187168
[0059] 4 is a block diagram illustrating an example of a video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 that receives video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0060] System 4000 may include a coding component 4004 that may implement various coding or encoding methods described herein. Coding component 4004 may reduce the average bitrate of video from input 4002 to the output of coding component 4004, generating a coded representation of the video. Coding techniques are therefore sometimes referred to as video compression techniques or video transcoding techniques. The output of coding component 4004 may be stored or transmitted via a communication connection, as represented by component 4006. The stored or communicated bitstream (or coded) representation of the video received at input 4002 may be used by component 4008 to generate pixel values or displayable video that are sent to display interface 910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Also, while certain video processing operations may be referred to as “coding” operations or tools, it is understood that coding tools or operations are used in an encoder, and corresponding decoding tools or operations that reverse the results of the coding are performed in a decoder.
[0061] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be embodied in a variety of electronic devices, such as, for example, mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0062] FIG. 5 is a block diagram of a video processing device 5000. The device 5000 may be used to implement one or more of the methods described herein (e.g., the methods shown in FIGS. 9-11). The device 5000 may be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 5000 may include one or more processors 5002, one or more memories 5004, and video processing hardware 5006. The processor(s) 5002 may be configured to execute one or more methods described herein. The memory(s) 5004 may be used to store data and code used to execute the methods and techniques described herein. The video processing hardware 5006 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing hardware 5006 may be at least partially included in the processor 5002, such as a graphics coprocessor.
[0063] FIG. 6 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0064] 6, video coding system 100 may include source device 110 and destination device 120. Source device 110 generates encoded video data and may be referred to as a video encoder. Destination device 120 can decode the encoded video data generated by source device 110 and may be referred to as a video decoder.
[0065] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface .
[0066] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a series of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted via the I / O interface 116 directly over the network 130a to the destination device 120. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0067] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0068] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to interface with an external display device.
[0069] Video encoder 114 and video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0070] FIG. 7 is a block diagram illustrating an example of a video encoder 200, which may be video encoder 114 in system 100 shown in FIG.
[0071] Video encoder 200 may be configured to perform any or all of the techniques described in this disclosure. In the example of FIG. 7, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0072] The functional components of the video encoder 200 may include a division unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0073] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0074] Also, some components, such as the motion estimation unit 204 and the motion compensation unit 205, although shown separately in the example of FIG. 5 for illustrative purposes, may be highly integrated.
[0075] Division unit 201 may divide a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support a variety of video block sizes.
[0076] The mode select unit 203 may select one of a plurality of coding modes, intra or inter, based on, for example, an error result, and provide the resulting intra- or inter-coded block to a residual generation unit 207, which generates residual block data, and to a reconstruction unit 212, which reconstructs a coding block for use as a reference picture. In some examples, the mode select unit 203 may select a combination of intra and inter predication (CIIP) mode, in which prediction is based on an inter prediction signal and an intra prediction signal. The mode select unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision) in the case of inter prediction.
[0077] To perform inter prediction on a current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a prediction video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0078] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice, for example.
[0079] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 or list 1 for reference video blocks for the current video block. Motion estimation unit 204 may then generate a reference index that points to the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector that indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Based on the reference video block indicated by the motion information of the current video block, motion compensation unit 205 may generate a prediction video block for the current block.
[0080] In another example, motion estimation unit 204 may perform bidirectional prediction on the current video block, where motion estimation unit 204 may search reference pictures in list 0 for a reference video block for the current video block and may also search reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 204 may then generate reference indices that point to the reference pictures in lists 0 and 1 that contain the reference video blocks, and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. Motion estimation unit 204 may output the reference indices and the motion vector for the current video block as motion information for the current video block. Based on the reference video blocks indicated by the motion information for the current video block, motion compensation unit 205 may generate a prediction video block for the current block.
[0081] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.
[0082] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block by reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0083] In one example, motion estimation unit 204 may point to a value within a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0084] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) within a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the pointed-to video block. Video decoder 300 may use the motion vector of the pointed-to video block and the motion vector difference to determine the motion vector of the current video block.
[0085] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0086] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks within the same picture. The predictive data for the current video block may include a predictive video block and various syntax elements.
[0087] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.
[0088] In other examples, for example, in skip mode, residual data for the current video block may not exist and residual generation unit 207 may not perform a subtraction operation.
[0089] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0090] After transform processing unit 208 generates the transform coefficient video block for the current video block, quantization unit 209 may quantize the transform coefficient video block for the current video block based on one or more quantization parameter (QP) values for the current video block.
[0091] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 202 to generate a reconstructed video block for the current block that is stored in buffer 213.
[0092] After reconstruction unit 212 reconstructs the video blocks, a loop filtering operation may be performed to reduce video blocking artifacts in the video blocks.
[0093] An entropy encoding unit 214 may receive data from other functional components of the video encoder 200. Once the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0094] FIG. 8 is a block diagram illustrating an example of a video decoder 300, which may be video decoder 124 in system 100 shown in FIG.
[0095] Video decoder 300 may be configured to perform any or all of the techniques described in this disclosure. In the example of FIG. 8, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0096] 8, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. Video decoder 300 may, in some examples, perform a decoding pass that is generally inverse to the encoding pass described with respect to video encoder 200 (e.g., FIG. 7).
[0097] An entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-encoded video data, and from the entropy-decoded video data, a motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.
[0098] The motion compensation unit 302 may optionally perform interpolation based on an interpolation filter to generate the motion-compensated blocks. An identifier for the interpolation filter used with sub-pixel precision may be included in the syntax element.
[0099] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of the reference block using the interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 according to received syntax information and generate the predictive block using the interpolation filters.
[0100] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is divided, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0101] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using an intra prediction mode, e.g., received in the bitstream. The inverse quantization unit 303 inverse quantizes, e.g., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0102] A reconstruction unit 306 may add the residual block with a corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for later motion compensation / intra prediction and also generates a decoded video for presentation on a display device.
[0103] 9-11 illustrate example ways in which the technical solutions described above in the embodiments illustrated in, for example, FIGS. 4-8 can be implemented.
[0104] 9 shows a flowchart of an example method 900 for video processing. The method 900 includes, at operation 910, performing a conversion between a video and a bitstream of the video having multiple layers, the bitstream further having multiple supplemental enhancement information (SEI) messages associated with a particular output layer set (OLS) or access units (AUs) or decoding units (DUs) of a particular layer, the multiple SEI messages having a message type different from a scalable nesting type based on a format rule, the format rule specifying that each of the multiple SEI messages has the same SEI payload content due to the multiple SEI messages being associated with the particular OLS or AUs or DUs of the particular layer.
[0105] 10 shows a flowchart of an example method 1000 for video processing. The method 1000 includes, at operation 1010, performing a conversion between a video having a current block and a bitstream of the video, the bitstream conforming to format rules that specify constraints on the bitstream due to syntax fields being excluded from the bitstream, the syntax fields indicating a Picture Order Count (POC) associated with a current picture having the current block or a current access unit (AU) having the current block.
[0106] 11 shows a flowchart of an example method 1100 for video processing. Method 1100 includes, at act 1110, performing a conversion between video and a bitstream of the video, where the bitstream conforms to a sub-bitstream extraction process order defined by a rule, where the rule specifies that (a) the sub-bitstream extraction process excludes an operation that removes all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from the output bitstream, or (b) the sub-bitstream extraction process order includes replacing a parameter set with a replacement parameter set prior to removing SEI NAL units from the output bitstream.
[0107] The following solutions represent example implementations of the techniques described in the preceding sections (eg, items 1 through 4).
[0108] Next, we provide a list of solutions that are preferred by some embodiments.
[0109] A1. A method for processing video data, comprising: performing a conversion between video and a bitstream of the video having multiple layers, the bitstream having multiple supplemental enhancement information (SEI) messages associated with a particular output layer set (OLS) or access units (AUs) or decoding units (DUs) of a particular layer, the multiple SEI messages having a message type different from a scalable nesting type based on a format rule, the format rule specifying that each of the multiple SEI messages has the same SEI payload content due to the multiple SEI messages being associated with the AUs or DUs of the particular OLS or the particular layer.
[0110] A2. The method of solution A1, wherein the value of the payload type of the scalable nesting SEI message is equal to 133.
[0111] A3. The method of Solution A1, wherein the plurality of SEI messages comprises at least one SEI message included in a scalable nesting SEI message and at least one SEI message not included in a scalable nesting SEI message.
[0112] A4. A method for processing video data, comprising the step of performing a conversion between a video having a current block and a bitstream of the video, wherein the bitstream complies with format rules, the format rules specifying constraints on the bitstream due to syntax fields being excluded from the bitstream, and the syntax fields indicating a picture order count (POC) associated with a current picture having the current block or a current access unit (AU) having the current block.
[0113] A5. A method of solution A4, wherein the syntax fields indicating the POC associated with the current picture and the current AU are excluded, and the constraint specifies that there are no pictures belonging to reference layers of the current layer within the current AU.
[0114] A6. The method of solution A4 or A5, wherein the constraint specifies that the leading picture in decoding order is neither a random access skip-reading (RASL) picture nor a random access decodable-reading (RADL) picture.
[0115] A7. A method of solution A4, wherein the constraint specifies that the preceding picture in decoding order is neither a random access skip reading (RASL) picture nor a random access decodable reading (RADL) picture due to the current AU excluding pictures belonging to a reference layer of the current layer within the current AU.
[0116] A8. The method of solution A6 or A7, wherein the temporal identifier of the preceding picture is equal to zero and the temporal identifier is TemporalId.
[0117] A9. Any of solutions A4 to A8, wherein the syntax field is ph_poc_msb_cycle_val.
[0118] A10. The method of any of Solutions A1 to A9, wherein the converting comprises decoding the video from the bitstream.
[0119] A11. The method of any one of Solutions A1 to A9, wherein the converting comprises encoding the video into the bitstream.
[0120] A12. A method for storing a bitstream representing an image on a computer-readable recording medium, comprising the steps of generating the bitstream from the image according to a method described in any one or more of Solutions A1 to A9, and storing the bitstream on the computer-readable recording medium.
[0121] A13. A video processing device comprising a processor configured to perform the method according to any one or more of solutions A1 to A12.
[0122] A14. A computer readable medium having stored thereon instructions which, when executed, cause a processor to perform the method described in one or more of solutions A1 to A12.
[0123] A15. A computer-readable medium storing a bitstream generated according to any one or more of solutions A1 to A12.
[0124] A16. A video processing device for storing a bitstream, the video processing device being configured to perform the method according to any one or more of solutions A1 to A12.
[0125] Next, we provide another list of solutions that are preferred by some embodiments.
[0126] B1. A method for processing video data, comprising: performing a conversion between video and a bitstream of the video, wherein the bitstream conforms to an ordering of a sub-bitstream extraction process defined by rules, the rules specifying that (a) the sub-bitstream extraction process excludes the operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from the output bitstream; or (b) the ordering of the sub-bitstream extraction process includes replacing a parameter set with a replacement parameter set prior to removing the SEI NAL units from the output bitstream.
[0127] B2. The method of Solution B1, wherein the condition specifies that at least one of the SEI NAL units has a scalable nesting SEI message having (a) a flag equal to zero and (b) a first identifier having a value different from the value of a second identifier in the list.
[0128] B3. The method of solution B2, wherein the flag indicates whether the scalable nesting SEI message applies to a particular output layer set, the first identifier is a nesting layer identifier, and the second identifier is a layer identifier of a layer within the output layer set.
[0129] B4. The method of solution B3, wherein the flag is sn_ols_flag.
[0130] B5. The method of solution B3, wherein the first identifier is a NestingLayerId and the second identifier is a LayerIdInOls[targetOlsIdx].
[0131] B6. The method of Solution B1, wherein at least one of the SEI NAL units includes a scalable nesting SEI message.
[0132] B7. The method of Solution B6, wherein the SEI NAL unit is removed from the output bitstream due to at least one video coding layer (VCL) NAL unit being removed from the output bitstream.
[0133] B8. The method of solution B6 or B7, wherein the scalable nesting SEI message has a flag equal to 0.
[0134] B9. The method of solution B8, wherein the flag is sn_subpic_flag.
[0135] B10. The method of any of Solutions B1 to B9, wherein said converting comprises decoding said video from said bitstream.
[0136] B11. The method of any of Solutions B1 to B9, wherein the converting comprises encoding the video into the bitstream.
[0137] B12. A method for storing a bitstream representing a video on a computer-readable recording medium, comprising the steps of generating the bitstream from the video according to a method described in any one or more of solutions B1 to B9, and storing the bitstream on the computer-readable recording medium.
[0138] B13. A video processing device comprising a processor configured to perform the method according to any one or more of solutions B1 to B12.
[0139] B14. A computer-readable medium having stored thereon instructions which, when executed, cause a processor to perform the method described in one or more of solutions B1 to B12.
[0140] B15. A computer-readable medium storing a bitstream generated according to any one or more of solutions B1 to B12.
[0141] B16. A video processing device for storing a bitstream, the video processing device being configured to perform the method according to any one or more of solutions B1 to B12.
[0142] Next, we provide yet another list of solutions that are preferred by some embodiments.
[0143] P1. A method of video processing, comprising: performing a conversion between video having one or more layers with one or more pictures and a coded representation of the video, the coded representation conforming to format rules, the format rules specifying constraints on the coded representation in the absence of a syntax field indicating a picture order count associated with a current picture in a current access unit.
[0144] P2. The method of solution P1, wherein the format rule specifies a constraint that the leading picture in decoding order is neither a random access skip-leading picture nor a random access decodable leading picture.
[0145] P3. The method of solution P1, wherein the format rule specifies that the constraint is based on the condition that there are no pictures belonging to a reference layer of the current layer in the current access unit.
[0146] P4. A method of video processing, comprising a step of performing a conversion between a video having one or more layers with one or more pictures and a coded representation of said video, said coded representation following an order of sub-bitstream extraction processes defined by rules.
[0147] P5. The method of solution P4, wherein the rule specifies that the sub-bitstream extraction process omits, for all SEI Network Abstraction Layer units, the step of removing scalable nesting supplemental enhancement information (SEI) messages from the output bitstream in accordance with a condition.
[0148] P6. The method of solution P4, wherein the rule specifies that the order of the sub-bitstream extraction process includes replacing a parameter set with a replacement parameter set prior to removing a supplemental enhancement information (SEI) Network Abstraction Layer (NAL) unit containing a scalable nesting SEI message in the output bitstream.
[0149] P7. A method of video processing, comprising: performing a conversion between video having one or more layers with one or more pictures and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies that if there are multiple supplemental enhancement information (SEI) messages with a particular value of payload type not equal to 133 associated with a particular access unit or decoding unit and applied to a particular output layer set or layer, the multiple SEI messages have the same SEI payload content.
[0150] P8. The method of solution P7, wherein some or all of the plurality of SEI messages are scalably nested.
[0151] P9. The method of any of solutions P1 to P8, wherein the transforming comprises generating a coded representation from the video.
[0152] P10. The method of any of solutions P1 to P8, wherein said converting comprises decoding said coded representation to generate said video.
[0153] P11. A video decoding device comprising a processor configured to perform the method according to one or more of solutions P1 to P10.
[0154] P12. A video coding device comprising a processor configured to perform the method according to one or more of solutions P1 to P10.
[0155] P13. A computer program product storing computer code which, when executed by a processor, causes the processor to carry out the method described in any of solutions P1 to P10.
[0156] P14. A computer-readable medium storing a coded representation generated according to any of solutions P1 to P10.
[0157] In the solutions described herein, an encoder can comply with the formatting rules by generating a coded representation according to the formatting rules, and a decoder can use the formatting rules to parse syntax elements in the coded representation with knowledge of the presence and absence of syntax elements according to the formatting rules to generate decoded video.
[0158] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied in converting a pixel representation of a video to a corresponding bitstream representation or vice versa. The bitstream representation of a current video block may correspond to bits located together in the bitstream or bits scattered in multiple different locations, e.g., as specified by syntax. For example, a macroblock may be encoded using bits in a header and other fields in the bitstream, with error residual values being transformed and coded. Also, in the transform, a decoder may parse the bitstream with knowledge that fields may or may not be present based on a decision, as described in the solution above. Similarly, an encoder may determine whether to include or not include certain syntax fields and generate a coded representation by including or excluding those syntax fields from the coded representation accordingly.
[0159] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document, including the structures disclosed herein and their structural equivalents, can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in combinations of one or more of these. The disclosed and other embodiments can be implemented as one or more computer program products, e.g., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter producing a machine-readable propagated signal, or a combination of one or more of these. The term "data processing apparatus" encompasses any apparatus, device, and machine that processes data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations of these. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver device.
[0160] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program at hand, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers, either collocated or distributed across multiple locations and interconnected by a communications network.
[0161] The processes and logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. These processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0162] Processors suitable for the execution of a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or is operatively coupled to receive data from or transfer data to the mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0163] While this patent document contains numerous details, these should not be construed as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operating in a particular combination, and even initially claimed as such, in some cases one or more features from a claimed combination may be removed from the combination, or the claimed combination may be subject to subcombinations or variations of the subcombination.
[0164] Similarly, although the figures may depict operations in a particular order, this should not be understood as requiring that those operations be performed in the particular order or sequence shown, or that all of the operations shown be performed, to achieve desired results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0165] Only a few implementations and examples have been described, and other implementations, extensions and variations may be made based on what is described and illustrated in this patent document.
Claims
1. 1. A method for processing video data, comprising: performing a conversion between the video and a bitstream of the video; and the bitstream is processed according to a sub-bitstream extraction process order, the sub-bitstream extraction process excluding an operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from an output bitstream, or the sub-bitstream extraction process order comprises replacing a parameter set with a replacement parameter set prior to removing the SEI NAL units from the output bitstream. method.
2. 2. The method of claim 1 , wherein the condition specifies that the SEI NAL unit has a scalable nesting SEI message that (a) has a flag equal to zero, and (b) no value in a first list of first identifiers is equal to a value in a second list of second identifiers.
3. 3. The method of claim 2, wherein the flag indicates whether the scalable nested SEI message applies to a particular output layer set, the first identifier is a nesting layer identifier, and the second identifier is a layer identifier of a layer within the output layer set.
4. The method of claim 2 or 3, wherein the flag is sn_ols_flag.
5. The method of claim 2 , wherein the first identifier is a NestingLayerId and the second identifier is a LayerIdInOls[targetOlsIdx].
6. The method of claim 1 , wherein at least one of the SEI NAL units comprises a scalable nesting SEI message.
7. 7. The method of claim 1, wherein the SEI NAL unit is removed from the output bitstream due to at least one video coding layer (VCL) NAL unit being removed from the output bitstream.
8. The method of claim 6 or 7, wherein the scalable nesting SEI message has a flag equal to 0.
9. The method of claim 8 , wherein the flag is sn_subpic_flag.
10. The method of claim 1 , wherein the converting comprises decoding the video from the bitstream.
11. The method of any of claims 1 to 9, wherein the converting comprises encoding the video into the bitstream.
12. 1. An apparatus for processing video data comprising a processor and a non-transitory memory having instructions that, upon execution by the processor, cause the processor to: performing a conversion between the video and a bitstream of the video; the bitstream is processed according to a sub-bitstream extraction process order, the sub-bitstream extraction process excluding an operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from an output bitstream, or the sub-bitstream extraction process order comprises replacing a parameter set with a replacement parameter set prior to removing the SEI NAL units from the output bitstream. Device.
13. A non-transitory computer-readable storage medium having stored thereon instructions that cause a processor to: performing a conversion between the video and a bitstream of the video; (b) the bitstream is processed according to a sub-bitstream extraction process order, the sub-bitstream extraction process excluding an operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from an output bitstream, or (b) the sub-bitstream extraction process order includes replacing a parameter set with a replacement parameter set prior to removing the SEI NAL units from the output bitstream. A computer-readable storage medium.
14. 1. A method for storing a video bitstream, comprising: generating the bitstream of the video; storing the bitstream on a non-transitory computer-readable recording medium; and the bitstream is processed according to a sub-bitstream extraction process order, the sub-bitstream extraction process excluding an operation of removing all satisfying supplemental enhancement information (SEI) network abstraction layer (NAL) units from an output bitstream, or the sub-bitstream extraction process order comprises replacing a parameter set with a replacement parameter set prior to removing the SEI NAL units from the output bitstream. method.