Signaling of assistance information

CN116671111BActive Publication Date: 2026-08-18DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180066926.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-29
Filing Date
2021-09-29
Publication Date
2026-08-18
Estimated Expiration
2041-09-29

Smart Images

  • Figure CN116671111B_ABST
    Figure CN116671111B_ABST
Patent Text Reader

Abstract

Systems, methods, and apparatus, including computer programs encoded on a computer-readable medium, for encoding, decoding, or transcoding digital video are described. One example method of processing video data includes performing a conversion between video and a bitstream of the video according to format rules, wherein the format rules specify that a supplemental enhancement information field included in the bitstream indicates whether the bitstream includes one or more video layers that represent auxiliary information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on International Patent Application No. PCT / CN2021 / 121513, filed on September 29, 2021, which claims priority and interest in International Patent Application No. PCT / CN2020 / 118711, filed on September 29, 2020. All of the aforementioned patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] This patent relates to digital video encoding and decoding technologies, including video encoding, transcoding, or decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This paper discloses techniques that can be used by video encoders and decoders to process video or image representations according to document formats.

[0006] In one example aspect, a method for processing video data is disclosed. The method includes: performing a conversion between video and video bitstreams according to format rules, wherein the format rules specify supplemental enhancement information fields or video availability information syntax structures included in the bitstream that indicate whether the bitstream includes a multi-view bitstream, in which multiple views are encoded and decoded in multiple video layers.

[0007] In another example, a method for processing video data is disclosed. The method includes performing a conversion between video and video bitstreams according to format rules, wherein the format rules specify that supplementary enhancement information fields included in the bitstream indicate whether the bitstream includes one or more video layers representing auxiliary information.

[0008] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video comprising video images and a codec representation of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that fields included in the codec representation indicate that the video is a multi-view video.

[0009] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising video images and a codec representation of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies fields included in the codec representation indicating that the video is encoded or decoded into multiple video layers in the codec representation.

[0010] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0011] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0012] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0013] In another example, a computer-readable medium on which a bitstream is stored is disclosed. The bitstream is generated or processed using the methods described in this document.

[0014] These and other features are described in this paper. Attached Figure Description

[0015] Figure 1 This is a block diagram of an example video processing system;

[0016] Figure 2 A block diagram of a video processing device;

[0017] Figure 3 A flowchart of an example method for video processing;

[0018] Figure 4 This is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure;

[0019] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure;

[0020] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure;

[0021] Figure 7 Here is an example of a bitstream with two OLS, where OLS2 has vps_max_tid_il_ref_pics_plus1[1][0] equal to 0; and

[0022] Figures 8 to 9 A flowchart of an example method for video processing. Detailed Implementation

[0023] Chapter headings are used in this document for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, edit changes to text relative to the current draft of the VVC specification are indicated by strikethrough indicating canceled text and highlighting indicating added text (including bold and italics).

[0024] 1. Introduction

[0025] This article relates to video codec technology. Specifically, it relates to signaling notification of scalability size information for multi-function video codec (VVC) video bitstreams. These ideas can be applied individually or in various combinations to any video codec standard or non-standard video codec, such as the recently completed VVC.

[0026] 2. Abbreviations

[0027] ACT Adaptive Color Transformation

[0028] ALF Adaptive Loop Filter

[0029] AMVR Adaptive Motion Vector Resolution

[0030] APS Adaptive Parameter Set

[0031] AU Access Unit

[0032] AUD access unit separator

[0033] AVC Advanced Video Codec (Rec. ITU-T H.264 | ISO / IEC 14496-10)

[0034] B Two-way prediction

[0035] BCW features bidirectional prediction with CU-level weights.

[0036] BDOF bidirectional optical flow

[0037] BDPCM is based on block-based incremental pulse coding and decoding modulation.

[0038] BP buffer period

[0039] CABAC Context-Based Adaptive Binary Arithmetic Encoding and Decoding

[0040] CB codec block

[0041] CBR constant bit rate

[0042] CCALF Cross-Component Adaptive Loop Filter

[0043] CLVS codec layer video sequence

[0044] CLVSS codec layer video sequence begins

[0045] CPB image buffer

[0046] CRA Fully Random Access

[0047] CRC Cyclic Redundancy Check

[0048] CTB codec tree block

[0049] CTU encoding / decoding tree unit

[0050] CU encoding / decoding unit

[0051] CVS encoded video sequence

[0052] CVSS codec video sequence start

[0053] DPB Decode Image Buffer

[0054] DCI decoding capability information

[0055] DRAP depends on random access point

[0056] DU decoding unit

[0057] DUI Decoding Unit Information

[0058] EG index Golomb

[0059] EGkk-order exponent Golomb

[0060] EOB bitstream end

[0061] End of EOS sequence

[0062] FD filler data

[0063] FIFO (First In First Out)

[0064] FL fixed length

[0065] GBR green, blue and red

[0066] GCI General Constraint Information

[0067] GDR is being gradually decoded and refreshed.

[0068] GPM geometric segmentation mode

[0069] HEVC High-Efficiency Video Codec (Rec. ITU-T H.265 | ISO / IEC 23008-2)

[0070] HRD Hypothetical Reference Decoder

[0071] HSS Hypothetical Flow Scheduler

[0072] Within I-frame

[0073] IBC Intra-Block Copying

[0074] IDR Instant Decoding and Refresh

[0075] ILRP interlayer reference image

[0076] IRAP Intra-Frame Random Access Point

[0077] LFNST Low-Frequency Inseparable Transform

[0078] LPS least likely symbol

[0079] LSB Least Significant Bit

[0080] LTRP Long-Term Reference Image

[0081] LMCS has luminance mapping with chroma scaling

[0082] MIP-based intra-frame prediction

[0083] MPS most likely symbol

[0084] MSB most significant bit

[0085] MTS Multiple Transformation Selection

[0086] MVP motion vector prediction

[0087] NAL Network Abstraction Layer

[0088] OLS Output Layer Set

[0089] OP operation point

[0090] OPI Operation Point Information

[0091] P prediction

[0092] PH image header

[0093] POC image sequential counting

[0094] PPS Image Parameter Set

[0095] PROF refines the prediction using optical flow.

[0096] PT image timer

[0097] PU image unit

[0098] QP quantization parameters

[0099] RADL random access decodeable front-end (image)

[0100] RASL random access skips the preprocessor (image)

[0101] RBSP raw byte sequence payload

[0102] RGB red, green and blue

[0103] RPL Reference Image List

[0104] SAO Sample Adaptive Migration

[0105] SAR sample aspect ratio

[0106] SEI Supplemental Enhancement Information

[0107] SH strip header

[0108] SLI sub-picture level information

[0109] SODB data bit string

[0110] SPS sequence parameter set

[0111] STRP Short-Term Reference Image

[0112] STSA is gradually being integrated into the time-domain sublayer.

[0113] TR cuts off Rice

[0114] TU Transformer

[0115] VBR Variable Bit Rate

[0116] VCL video codec layer

[0117] VPS Video Parameter Set

[0118] VSEI General Supplemental Enhancement Information (Rec. ITU-T H.274 | ISO / IEC 23002-7)

[0119] VUI Video Availability Information

[0120] VVC (Video Codec) (Rec. ITU-T H.266 | ISO / IEC 23090-3)

[0121] 3. Preliminary Discussion

[0122] 3.1. Video codec standards

[0123] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed the MPEG-1 and MPEG-4 Visual standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). When the Multifunctional Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce the bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which concluded on July 1, 2020.

[0124] The Multi-Functional Video Coding (VVC) standard (ITU-TH.266|ISO / IEC 23090-3) and the related Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-TH.274|ISO / IEC 23002-7) are designed for the widest range of applications, including traditional applications such as television broadcasting, video conferencing, or playback of stored media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple codec video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360° immersive media.

[0125] 3.2. Video-based point cloud compression (V-PCC)

[0126] ISO / IEC 23090-5, Information technology—Encoding and decoding representation of immersive media—Part 5: Video codecs based on visual volumetric video (V3C) and video-based point cloud compression (V-PCC), or V-PCC for short, is a standard that specifies the encoding and decoding representation of point cloud signals. The V-PCC standard is another recently completed standard.

[0127] V-PCC specifies data types, such as occupancy, geometry, texture attributes, material attributes, transparency attributes, reflection attributes, and normal attributes, which can be encoded and decoded using specific video codecs (such as VVC, HEVC, AVC, etc.).

[0128] 3.3. Temporal Scalability Support in VVC

[0129] VVC includes temporal scalability support similar to that in HEVC. This support includes signaling notification of the temporal ID in the NAL unit header, restrictions that images of specific temporal sublayers cannot be used as inter-frame prediction references by images of lower temporal sublayers, the sub-bitstream extraction process, and the requirement that the output of each sub-bitstream extraction from appropriate inputs must conform to the bitstream. Media-Aware Network Elements (MANEs) can use the temporal ID in the NAL unit header for stream adaptation purposes based on temporal scalability.

[0130] 3.4. Image resolution changes within a VVC sequence

[0131] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC allows changing the image resolution within a sequence at locations where IRAP images are not encoded; IRAP images are always intra-frame encoded and decoded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling the reference image used for inter-frame prediction when the reference image has a different resolution than the current image being decoded.

[0132] To allow for the reuse of motion compensation modules in existing implementations, the scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). The horizontal and vertical scaling ratios are derived from the image's width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.

[0133] RPR allows for resolution changes without requiring the encoding and decoding of IRAP images, which can lead to instantaneous bitrate spikes in streaming or video conferencing scenarios, for example, in response to changes in network conditions. RPR can also be used in applications requiring scaling of the entire video area or specific regions of interest. Negative scaling window offsets are allowed to support a wider range of scaling-based applications. Negative scaling window offsets also enable the extraction of sub-image sequences from multiple bitstreams while maintaining the same scaling window for the extracted sub-bitstream as in the original bitstream.

[0134] Unlike the spatial scalability in HEVC's scalable extension, where image resampling and motion compensation are applied in two different stages, RPR in VVC is performed at the block level as part of the same process, where the derivation of sample locations and motion vector scaling is performed during motion compensation.

[0135] In efforts to limit the complexity of the implementation, image resolution changes are not permitted within the CLVS when each image in the CLVS has multiple sub-images. Furthermore, when RPR is used between the current image and the reference image, decoder-side motion vector thinning, bidirectional optical flow, and prediction thinning utilizing optical flow are not applied. The juxtaposed images used to derive temporal motion vector candidates are also restricted to having the same image size, scaling window offset, and CTU size as the current image.

[0136] To support RPR, VVC differs from HEVC in several other aspects of its design. First, the image resolution and corresponding consistency and scaling windows are signaled in the PPS, not the SPS, where the maximum image resolution and corresponding consistency window are signaled. In applications, the maximum image resolution with the corresponding consistency window offset in the SPS can be used as the expected or desired image output size after cropping. Second, for a single-layer bitstream, each image storage (the slot in the DPB used to store one decoded image) occupies the buffer size required to store the decoded image with the maximum image resolution.

[0137] 3.5. Multi-tier Scalability Support in VVC

[0138] The RPR (Reference Prediction Process) in the VVC core design enables inter-frame prediction from reference images of different sizes, allowing VVC to easily support multi-layered bitstreams with varying resolutions, such as two layers with standard definition and high definition resolution respectively. This functionality can be integrated into the VVC decoder without requiring any additional signal processing level codecs, as the necessary upsampling capabilities to support spatial scalability can be provided by reusing the RPR upsampling filter. However, additional high-level syntax design is required to support bitstream scalability.

[0139] VVC supports scalability, but only within multi-layer profiles. Unlike scalability support in any earlier video codec standards (including extensions to AVC and HEVC), VVC scalability is designed to be as friendly as possible to single-layer decoder implementations. The decoding capability of a multi-layer bitstream is specified as if the bitstream contained only a single layer. For example, decoding capability for a DPB size can be specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, a decoder designed for a single-layer bitstream can decode multi-layer bitstreams without significant modifications.

[0140] Compared to the multi-layered extensions of AVC and HEVC, HLS offers significant simplification at the expense of some flexibility. For example, 1) the IRAP AU needs to include a picture of every layer present in CVS, avoiding the need to specify a layer-by-layer decoding process, and 2) a simpler design including POC signaling notifications in VVC, instead of a complex POC reset mechanism, to ensure that the exported POC values ​​are identical for all pictures in the AU.

[0141] Similar to HEVC, information about layers and layer dependencies is included in the VPS. Information about the OLS is provided for signaling notifications about which layers are included in the OLS, which layers are output, and other information such as PTL and HRD parameters associated with each OLS. Similar to HEVC, there are three operating modes: output all layers, output only the highest layer, or output layers with specific indications in a custom output mode.

[0142] There are some differences in the OLS design between VVC and HEVC. First, in HEVC, the layer set is signaled, and then the OLS is signaled based on the layer set; for each OLS, the output layer is signaled. The design in HEVC allows a layer to belong to an OLS that is neither an output layer nor a layer required for decoding an output layer. In VVC, the design requires that any layer in the OLS is either an output layer or a layer required for decoding an output layer. Therefore, in VVC, the OLS is signaled by indicating the output layer of the OLS, and other layers belonging to the OLS are derived only by the layer dependencies indicated in the VPS. Furthermore, VVC requires that each layer be contained in at least one OLS.

[0143] Another difference in the VVC OLS design is that, unlike HEVC, where the OLS consists of all NAL units belonging to the set of identified layers mapped to the OLS, VVC excludes some NAL units belonging to non-output layers mapped to the OLS. More specifically, the VVC OLS consists of the set of layers mapped to the OLS, where non-output layers only include IRAP or GDR pictures (ph_recovery_poc_cnt equals 0) or pictures from sublayers used for inter-layer prediction. This allows indicating the optimal level value for the multi-layer bitstream, considering only all “necessary” pictures from all sublayers within the layer forming the OLS, where “necessary” here means required for output or decoding. Figure 7 An example of a two-layer bitstream with vps_max_tid_il_ref_pics_plus1[1][0] equal to 0 is shown, i.e., when extracting OLS2, only the sub-bitstream of the IRAP picture from layer L0 is retained.

[0144] Considering that allowing different RAP periodicity at different layers is beneficial in some scenarios, similar to AVC and HEVC, where AUs are allowed to have layers with unaligned RAPs, the Access Unit Delimiter (AUD) is extended to include a flag indicating whether the AU is an IRAP AU or a GDR AU, compared to HEVC, to more quickly identify RAPs in multi-layer bitstreams. Furthermore, when the VPS indicates multiple layers, the AUD is mandatory for such IRAP or GDR AUs. However, for single-layer bitstreams indicated by the VPS or bitstreams not referencing the VPS, as in HEVC, the AUD is entirely optional, since in this case, the RAP can be easily detected from the NAL unit type and corresponding parameter set of the first stripe in the AU.

[0145] In order to enable multiple layers to share SPS, PPS and APS, and to ensure that the bitstream extraction process does not discard the parameter set required for the decoding process, the VCL NAL unit of the first layer can refer to the SPS, PPS or APS with the same or lower layer ID value, as long as all OLS including the first layer also include the layer identified by the lower layer ID value.

[0146] 3.6.VUI and SEI messages

[0147] The VUI is a syntax structure sent as part of the SPS (and possibly in the HEVC VPS). The VUI carries information that does not affect the standard decoding process, but may be important for the correct rendering of the encoded and decoded video.

[0148] SEI assists in processes related to decoding, display, or other purposes. Like VUI, SEI does not affect standard decoding processes. SEI is carried within SEI messages. Decoder support for SEI messages is optional. However, SEI messages do affect bitstream consistency (e.g., if the syntax of SEI messages in the bitstream does not conform to the specification, the bitstream is not conforming to the specification), and some SEI messages are required in the HRD specification.

[0149] The VUI syntax structures and most SEI messages used with VVC are not specified in the VVC specification, but rather in the VSEI specification. The SEI messages required for HRD conformance testing are specified in the VVC specification. VVC v1 defines five SEI messages related to HRD conformance testing, and VSEI v1 specifies 20 additional SEI messages. The SEI messages carried in the VSEI specification do not directly affect the behavior of the conformance decoder and are defined to allow them to be used in a codec-agnostic manner, thus allowing VSEI to be used with other video codec standards besides VVC in the future. The VSEI specification does not specifically mention the names of VVC syntax elements, but rather refers to variables whose values ​​are set in the VVC specification.

[0150] Compared to HEVC, VVC's VUI syntax structure focuses only on information related to the correct rendering of the image and does not include any timing information or bitstream limit indications. In VVC, the VUI is signaled in the SPS, which includes a length field before the VUI syntax structure to signal the length of the VUI payload (in bytes). This allows the decoder to easily skip information and, more importantly, allows for convenient future VUI syntax extensions by adding new syntax elements directly to the end of the VUI syntax structure in a manner similar to SEI message syntax extensions.

[0151] The VUI syntax structure contains the following information:

[0152] • The content is interwoven or progressive;

[0153] • Does the content include frame-encapsulated stereoscopic video or projected omnidirectional video?

[0154] • Aspect ratio of the sample points;

[0155] • Is the content suitable for overscan display?

[0156] • Color description, including primary colors, matrix, and transmission characteristics, is particularly important for signaling communication of Ultra High Definition (UHD) and High Definition (HD) color spaces as well as High Dynamic Range (HDR);

[0157] • Chromaticity position relative to luminance (signaling notification of progressive content has been clarified compared to HEVC).

[0158] When SPS does not contain any VUI, the information is considered unspecified, and if the content of the bitstream is intended for rendering on a display, the information must be communicated externally or specified by the application.

[0159] Table 1 lists all SEI messages specified for VVC v1, along with the specification containing their syntax and semantics. Of the 20 SEI messages defined in the VSEI specification, many are inherited from HEVC (e.g., the padding payload and two user data SEI messages). Some SEI messages are essential for the proper processing or rendering of encoded video content. This is the case, for example, for SEI messages related to primary display color magnitude, content light level information, or alternative transport characteristics, which are particularly relevant to HDR content. Other examples include SEI messages for isometric projection, spherical rotation, area packing, or omnidirectional viewports, which are related to signaling notification and processing of 360° video content.

[0160] Table 1: SEI Message List in VVC v1

[0161]

[0162]

[0163] The new SEI messages specified for VVC v1 include frame field information SEI messages, sample aspect ratio information SEI messages, and sub-picture level information SEI messages.

[0164] The Frame Field Information (SEI) message contains information indicating how the associated picture should be displayed (e.g., field parity or frame repetition period), the source scan type of the associated picture, and whether the associated picture is a copy of a previous picture. In previous video codec standards, this information was typically signaled along with the timing information of the associated picture in the Picture Timing SEI message. However, it was observed that Frame Field Information and timing information are two different types of information and are not necessarily signaled together. A typical example involves signaling timing information at the system level but signaling Frame Field Information within the bitstream. Therefore, it was decided to remove Frame Field Information from the Picture Timing SEI message and instead signal it in a dedicated SEI message. This change also made it possible to modify the syntax of the Frame Field Information to convey more and clearer instructions to the display, such as pairing fields together or specifying more values ​​for frame repetition.

[0165] Sample Aspect Ratio (SEI) messages can signal different sample aspect ratios for different images within the same sequence, while the corresponding information contained in the VUI applies to the entire sequence. This can be relevant when using reference image resampling features with scaling factors that cause different images in the same sequence to have different sample aspect ratios.

[0166] The Sub-Picture Level Information (SEI) message provides level information for a sequence of sub-pictures.

[0167] 4. The technical problems solved by the publicly disclosed technical solutions

[0168] VVC supports multi-layer scalability. However, given a VVC multi-layer bitstream, it is unknown whether the OLS bitstream is a multi-view bitstream or simply a bitstream composed of multiple layers with SNR and / or spatial scalability. Furthermore, given a VVC multi-layer bitstream, it is unknown whether one or more layers represent auxiliary information such as alpha, depth, etc., and if so, which layers represent what.

[0169] 5. List of technical solutions

[0170] To address the aforementioned problems, the following summarized methods are disclosed. Inventions should be considered as examples for interpreting general concepts, and not interpreted in a narrow sense. Furthermore, these inventions can be applied individually or in any combination.

[0171] 1) Information indicating whether a VVC video bitstream is a multi-view bitstream is signaled in the VVC video bitstream.

[0172] a. In one example, this information is signaled in an SEI message, for example, named a scalability size SEI message.

[0173] i. In one example, the Scalable Size SEI message provides information about the bitstreamInScope, which is defined as a sequence of AUs that, in decoding order, includes the AU containing the current Scalable Size SEI message, followed by zero or more AUs, including all subsequent AUs but excluding any subsequent AUs containing the Scalable Size SEI message.

[0174] ii. In one example, the SEI message includes a flag indicating whether the bitstream can be a multi-view bitstream.

[0175] iii. In one example, the SEI message indicates the view ID for each layer.

[0176] 1. In one example, the SEI message includes a flag indicating whether the view ID is signaled for each layer.

[0177] 2. In one example, the length (in bits) of the view ID for each layer is signaled in the SEI message.

[0178] b. In one example, this information is signaled as part of the VUI.

[0179] 2) Information indicating whether the VVC video bitstream includes one or more layers representing auxiliary information is signaled in the VVC video bitstream.

[0180] a. In one example, this information is signaled in an SEI message, for example, named a scalability size SEI message.

[0181] i. In one example, the Scalable Size SEI message provides information about the bitstreamInScope, which is defined as a sequence of AUs that, in decoding order, includes the AU containing the current Scalable Size SEI message, followed by zero or more AUs, including all subsequent AUs but excluding any subsequent AUs containing the Scalable Size SEI message.

[0182] ii. In one example, the SEI message includes a flag indicating whether the bitstream may contain auxiliary information carried by one or more layers.

[0183] iii. In one example, the SEI message indicates the auxiliary ID for each layer.

[0184] 1. In one example, the SEI message includes a flag indicating whether the auxiliary ID is signaled for each layer.

[0185] 2. In one example, the value of the auxiliary ID, such as 0, indicates that the layer does not contain an auxiliary image.

[0186] 3. In one example, the value of the auxiliary ID, such as 1, indicates that the type of auxiliary information is alpha.

[0187] 4. In one example, the value of the auxiliary ID, such as 2, indicates that the type of auxiliary information is depth.

[0188] 5. In one example, the value of the auxiliary ID, such as 3, indicates that the type of auxiliary information is occupied, for example, as specified in V-PCC.

[0189] 6. In one example, the value of the auxiliary ID, such as 4, indicates that the type of auxiliary information is geometric, such as that specified in V-PCC.

[0190] 7. In one example, the value of the auxiliary ID, such as 5, indicates that the type of auxiliary information is an attribute, such as that specified in V-PCC.

[0191] 8. In one example, the value of the auxiliary ID, such as 6, indicates that the type of auxiliary information is a texture attribute, such as that specified in V-PCC.

[0192] 9. In one example, the value of the auxiliary ID, such as 7, indicates that the type of auxiliary information is a material property, such as that specified in V-PCC.

[0193] 10. In one example, the value of the auxiliary ID, such as 8, indicates that the type of auxiliary information is a transparency attribute, such as that specified in V-PCC.

[0194] 11. In one example, the value of the auxiliary ID, such as 9, indicates that the type of auxiliary information is a reflectivity attribute, such as that specified in V-PCC.

[0195] 12. In one example, the value of the auxiliary ID, such as 10, indicates that the type of auxiliary information is a normal attribute, such as that specified in V-PCC.

[0196] b. In one example, this information is signaled as part of the VUI.

[0197] 6. Example

[0198] Below are some example embodiments of this invention summarized in Chapter 5 above, which can be applied to the VVC and VSEI specifications.

[0199] 6.1 First Embodiment

[0200] This embodiment applies to Project 1, 1.a and all its sub-projects 2, 2.a, 2.ai, 2.a.ii, 2.a.iii, 2.a.iii.1, 2.a.iii.2, 2.a.iii.3 and 2.a.iii.4.

[0201] 6.1.1. Scalability Size SEI Message Syntax

[0202]

[0203] 6.1.2. Scalability Size SEI Message Semantics

[0204] The Scalability Size SEI message provides scalability size information for each layer in bitstreamInScope (defined below), such as 1) the view ID of each layer when bitstreamInScope may be a multi-view bitstream; and 2) the auxiliary ID of each layer when there may be auxiliary information (such as depth or alpha) carried by one or more layers in bitstreamInScope.

[0205] bitstreamInScope is an AU sequence in decoding order. This AU sequence includes the AU containing the current scalability size SEI message, followed by zero or more AUs, including all subsequent AUs, but excluding any subsequent AUs containing the scalability size SEI message.

[0206] The increment of sd_max_layers_minus1 indicates the maximum number of layers in bitstreamInScope.

[0207] A value of 1 for `sd_multiview_info_flag` indicates that `bitstreamInScope` may be a multiview bitstream, and that the `sd_view_id_val[]` syntax element exists in the Scalable Size SEI message. A value of 0 for `sd_multiview_flag` indicates that `bitstreamInScope` is not a multiview bitstream, and that the `sd_view_id_val[]` syntax element does not exist in the Scalable Size SEI message.

[0208] A value of 1 for `sd_auxilary_info_flag` indicates that auxiliary information may exist, carried by one or more layers in `bitstreamInScope`, and that the `sd_aux_id[]` syntax element exists in the Scalable Size SEI message. A value of 0 for `sd_auxilary_info_flag` indicates that no auxiliary information exists, carried by one or more layers in `bitstreamInScope`, and that the `sd_aux_id[]` syntax element does not exist in the Scalable Size SEI message.

[0209] sd_view_id_len specifies the length (in bits) of the sd_view_id_val[i] syntax element.

[0210] `sd_view_id_val[i]` specifies the view ID of the i-th level in `bitstreamInScope`. The length of the `sd_view_id_val[i]` syntax element is `sd_view_id_len` bits. If it does not exist, the value of `sd_view_id_val[i]` is inferred to be equal to 0.

[0211] A value of 0 for sd_aux_id[i] indicates that the i-th layer in bitstreamInScope does not contain an auxiliary image. A value greater than 0 for sd_aux_id[i] indicates the type of auxiliary image in the i-th layer of bitstreamInScope, as specified in Table 2.

[0212] Table 2 – Mapping of sd_aux_id[i] to the type of auxiliary image

[0213]

[0214] Note 1 – The interpretation of auxiliary images associated with sd_aux_id in the range of 128 to 159 (inclusive) is specified in a manner other than the sd_aux_id value.

[0215] For a bitstream conforming to this version of the specification, sd_aux_id[i] should be in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive). Although the value of sd_aux_id[i] should be in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive), in this version of the specification, the decoder should allow the value of sd_aux_id[i] to be in the range of 0 to 255 (inclusive).

[0216] Figure 1 This is a block diagram of an example video processing system 1900 that can implement the various techniques disclosed herein. Various implementations may include some or all of the components in system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[0217] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described herein. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via connected communication, as represented by component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or displayable video that is sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations for retrieving the codec results are performed by the decoder.

[0218] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described herein can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0219] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more of the methods described herein. Apparatus 3600 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors 3602(s) may be configured to implement one or more methods described herein. The memories 3604(s) may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing hardware 3606 may be at least partially included in the processor 3602, such as a graphics coprocessor.

[0220] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0221] like Figure 4 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0222] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0223] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems that generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 includes a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0224] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0225] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to connect to an external display device.

[0226] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or other standards.

[0227] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 4 The video encoder 114 in the system 100 shown in the figure.

[0228] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0229] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0230] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0231] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes... Figure 5 The examples are shown separately.

[0232] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0233] The mode selection unit 203 can, for example, select one of the intra-frame or inter-frame encoding / decoding modes based on the error result, and provide the obtained intra-frame or inter-frame encoded / decoded blocks to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the encoded blocks for use as reference images. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. The mode selection unit 203 can also select the resolution of the motion vector (e.g., sub-pixel or integer pixel precision) for the blocks in the inter-frame prediction case.

[0234] To perform inter-frame prediction for the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information of the image from buffer 213 (rather than the image associated with the current video block) and decoded samples.

[0235] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, the different operations performed depend on whether the current video block is in an I-strip, a P-strip, or a B-strip.

[0236] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference image in list 0 or list 1 contains the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0237] In other examples, motion estimation unit 204 can perform bidirectional prediction for the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 204 can then generate a reference index indicating that the reference images in list 0 or list 1 contain the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0238] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding process.

[0239] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.

[0240] In one example, the motion estimation unit 204 may indicate in the syntax structure associated with the current video block that the current video block has the same motion information value as another video block.

[0241] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicating video block. Video decoder 300 can use the motion vector of the indicating video block and the motion vector difference to determine the motion vector of the current video block.

[0242] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling notification.

[0243] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0244] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0245] In other examples, such as in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform a subtraction operation.

[0246] Transform unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0247] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0248] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.

[0249] After the video block is reconstructed in reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0250] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0251] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 4 The video decoder 124 in the system 100 shown in the figure.

[0252] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0253] exist Figure 6 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 5 The decoding process is the overall inversion of the encoding process described.

[0254] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and merge modes.

[0255] The motion compensation unit 302 can generate motion compensation blocks, possibly based on interpolation filters. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0256] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation values ​​of a sub-integer number of pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0257] The motion compensation unit 302 can use some syntactic information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0258] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies the inverse transform.

[0259] The reconstruction unit 306 can sum the residual blocks using the corresponding prediction blocks generated by the motion compensation unit 302 or the intra-frame prediction unit 303 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on the display device.

[0260] The following provides a list of preferred solutions for some embodiments.

[0261] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 1).

[0262] 1. A method for video processing (e.g., Figure 3 The method described in 700 includes performing (702) a conversion between a video containing video images and a video codec representation, wherein the bitstream conforms to a format rule, wherein the format rule specifies that fields included in the codec representation indicate that the video is a multi-view video.

[0263] 2. The method according to Solution 1, wherein the field is included in the supplementary enhancement information portion of the codec representation.

[0264] 3. The method according to Solution 1, wherein the field is included in the video availability information section of the codec representation.

[0265] The following solutions illustrate example implementations of the techniques discussed in the previous chapter (e.g., Project 2).

[0266] 4. A video processing method, comprising: performing a conversion between a video including video images and a codec representation of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that fields included in the codec representation indicate that the video is encoded or decoded in codec representations of multiple video layers.

[0267] 5. The method according to Solution 4, wherein the field is included in the supplementary enhancement information portion of the codec representation.

[0268] 6. The method according to Solution 4, wherein the field is included in the video availability information section of the codec representation.

[0269] 7. The method according to any one of solutions 1-6, wherein the conversion includes generating a codec representation from the video.

[0270] 8. The method according to any one of solutions 1-6, wherein the conversion includes decoding the encoding / decoding representation to generate video.

[0271] 9. A video decoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions 1 to 8.

[0272] 10. A video encoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions 1 to 8.

[0273] 11. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 8.

[0274] 12. A computer-readable medium storing a codec representation generated according to any one of solutions 1 to 8.

[0275] 13. The methods, apparatus or systems described in this document.

[0276] In the solution described herein, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described herein, the decoder uses the format rules to parse the syntax elements in the codec representation and determines the presence or absence of these syntax elements to generate the decoded video.

[0277] Figure 8 This is a flowchart of an example method for video processing. Operation 802 includes: performing a conversion between video and video bitstreams according to format rules, wherein the format rules specify supplemental enhancement information fields or video availability information syntax structures included in the bitstream that indicate whether the bitstream includes a multi-view bitstream, in which multiple views are encoded and decoded in multiple video layers.

[0278] In some embodiments, the formatting rules specify that the supplemental enhancement information field is included in the scalability size information of the supplemental enhancement information message in the bitstream. In some embodiments, the formatting rules specify that the supplemental enhancement information message includes a first flag indicating whether the bitstream is a multi-view bitstream. In some embodiments, the formatting rules specify that the supplemental enhancement information message includes a view identifier for each of the multiple video layers of the bitstream. In some embodiments, the formatting rules specify the bit length of the view identifier for each video layer included in the supplemental enhancement information message.

[0279] In some embodiments, the format rules specify that the supplemental enhancement information message includes a second flag indicating whether a view identifier is included in the bitstream of each video layer. In some embodiments, the format rules specify that the scalability size information in the supplemental enhancement information message provides information related to the sequence of access units, which, in decoding order, include access units containing second scalability size information from a second supplemental enhancement information message, followed by zero or more access units, including all subsequent access units but excluding any subsequent access units containing third scalability size information in a third supplemental enhancement information message. In some embodiments, the format rules specify that the supplemental enhancement information field is included in a video availability information syntax structure in the bitstream. In some embodiments, the bitstream is a multi-functional video codec bitstream. In some embodiments, performing a conversion includes encoding video into a bitstream. In some embodiments, performing a conversion includes generating a bitstream from video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments, performing a conversion includes decoding video from the bitstream.

[0280] Figure 9 This is a flowchart of an example method for video processing. Operation 902 includes: performing a conversion between video and video bitstream according to format rules, wherein the format rules specify that supplementary enhancement information fields included in the bitstream indicate whether the bitstream includes one or more video layers representing auxiliary information.

[0281] In some embodiments, the formatting rules specify that the supplementary enhancement information field is included in the scalability size information of the supplementary enhancement information message in the bitstream. In some embodiments, the formatting rules specify that the supplementary enhancement information message includes a first flag indicating whether the bitstream contains auxiliary information for one or more video layers. In some embodiments, the formatting rules specify that the supplementary enhancement information message includes an auxiliary identifier for each of the multiple video layers in the bitstream. In some embodiments, the formatting rules specify that a first value for the auxiliary identifier of a video layer indicates that the video layer does not include auxiliary pictures.

[0282] In some embodiments, the format rule specifies a second value for the auxiliary identifier of the video layer indicating that the type of auxiliary information for the video layer is α. In some embodiments, the format rule specifies a third value for the auxiliary identifier of the video layer indicating that the type of auxiliary information for the video layer is depth. In some embodiments, the format rule specifies that the supplementary enhancement information message includes a second flag indicating whether the auxiliary identifier is included in the bitstream of each video layer. In some embodiments, the format rule specifies a fourth value for the auxiliary identifier of the video layer indicating that the type of auxiliary information for the video layer is occupancy.

[0283] In some embodiments, the formatting rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is geometry. In some embodiments, the formatting rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is an attribute. In some embodiments, the formatting rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is a texture attribute. In some embodiments, the formatting rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is a material attribute. In some embodiments, the formatting rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is a transparency attribute.

[0284] In some embodiments, the format rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is a reflection attribute. In some embodiments, the format rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary information of the video layer is a normal attribute. In some embodiments, the format rule specifies that the scalability size information in the supplementary enhancement information message provides information related to the sequence of access units, which, in decoding order, include access units containing the second scalability size information in the second supplementary enhancement information message, followed by zero or more access units, including all subsequent access units but excluding any subsequent access units containing the third scalability size information in the third supplementary enhancement information message.

[0285] In some embodiments, the format rules specify that supplemental enhancement information fields are included in the video availability information within the bitstream. In some embodiments, the video is a multi-function video codec video. In some embodiments, performing the conversion includes encoding the video into a bitstream. In some embodiments, performing the conversion includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments, performing the conversion includes decoding the video from the bitstream.

[0286] In some embodiments, a video decoding apparatus includes a processor configured to implement the methods described in one or more of the techniques described herein. In some embodiments, a video encoding apparatus includes a processor configured to implement the methods described in one or more of the techniques described herein. In some embodiments, a computer program product has computer instructions stored thereon that, when executed by a processor, cause the processor to implement the methods described herein. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to any of the methods described herein.

[0287] In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause a processor to implement any of the methods described herein. In some embodiments, a method for generating a bitstream includes: generating a bitstream of video according to any of the methods described herein, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, an apparatus, or a bitstream generated according to the disclosed method or system described in this document.

[0288] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-occurring or scattered at different positions within the bitstream. For example, a macroblock can be encoded based on the error residuals from the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0289] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on the decision or determination, the conversion from video blocks to a bitstream representation of the video will be performed using that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the decision or determination.

[0290] Some embodiments of the technology disclosed herein include deciding or determining to disable video processing tools or modes. In one example, when video processing tools or modes are disabled, the encoder will not use the tools or modes in the conversion of video blocks to a bitstream representation of the video. In another example, when video processing tools or modes are disabled, the decoder will process the bitstream using knowledge that the bitstream has not yet been modified based on the video processing tools or modes that have been decided or determined to be disabled.

[0291] The disclosures and other schemes, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, containing the structures disclosed in this document and their equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products encoded on a computer-readable medium, i.e., one or more computer program instruction modules for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a complex influencing machine-readable propagating signals, or combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0292] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0293] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be performed by special-purpose logic circuitry (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), and the apparatus can be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).

[0294] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magneto-optical, magneto-optical, or optical disc) for storing data, or operatively coupled to receive data from or transfer data to a mass storage device (e.g., magneto-optical, magneto-optical, or optical disc), or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0295] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0296] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all the operations shown, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0297] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.

Claims

1. A method for processing video data, comprising: Perform the conversion between the video and the video bitstream according to the format rules. The format rules specified in the format rules indicate whether the supplemental enhancement information field included in the bitstream is allowed to include one or more video layers representing auxiliary information; The format rules further specify that the supplementary enhancement information field is included in the scalability size information of the supplementary enhancement information message in the bitstream.

2. The method according to claim 1, wherein, The formatting rules specify that the supplemental enhancement information message includes an auxiliary identifier for each video layer in one or more video layers of the bitstream.

3. The method according to claim 2, wherein, The format rule specifies that the first value of the auxiliary identifier for the video layer indicates that the video layer does not include auxiliary images.

4. The method according to claim 3, wherein, The first value is equal to 0.

5. The method according to claim 2, wherein, The format rule specifies that the second value of the auxiliary identifier of the video layer indicates that the type of the auxiliary image of the video layer is α-plane.

6. The method according to claim 5, wherein, The second value is equal to 1.

7. The method according to claim 2, wherein, The format rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of auxiliary image of the video layer is a depth image.

8. The method according to claim 7, wherein, The third value is equal to 2.

9. The method according to any one of claims 1 to 8, wherein, The video is a multi-functional video codec.

10. The method according to any one of claims 1 to 8, wherein, Performing the conversion includes encoding the video into the bitstream.

11. The method according to any one of claims 1 to 8, wherein, Performing the conversion includes decoding the video from the bitstream.

12. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Perform the conversion between the video and the video bitstream according to the format rules. The format rules specified in the format rules indicate whether the supplemental enhancement information field included in the bitstream is allowed to include one or more video layers representing auxiliary information; The format rules further specify that the supplementary enhancement information field is included in the scalability size information of the supplementary enhancement information message in the bitstream.

13. The apparatus according to claim 12, wherein, The format rules specify that the supplemental enhancement information message includes an auxiliary identifier for each video layer of one or more video layers of the bitstream.

14. The apparatus according to claim 13, wherein, The format rule specifies that the first value of the auxiliary identifier for the video layer indicates that the video layer does not include auxiliary images, wherein the first value is equal to 0. The format rule specifies that the second value of the auxiliary identifier for the video layer indicates that the type of the auxiliary image for the video layer is α-plane, wherein the second value is equal to 1, and The format rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of the auxiliary picture of the video layer is a depth picture, wherein the third value is equal to 2.

15. The apparatus according to claim 12, wherein, The video is a multi-functional video codec.

16. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: Perform the conversion between the video and the video bitstream according to the format rules. The format rules specified in the format rules indicate whether the supplemental enhancement information field included in the bitstream is allowed to include one or more video layers representing auxiliary information; in, The format rules also specify that the supplemental enhancement information field is included in the scalability size information of the supplemental enhancement information message in the bitstream.

17. The non-transitory computer-readable storage medium according to claim 16, wherein, The format rules specify that the supplemental enhancement information message includes an auxiliary identifier for each video layer of one or more video layers of the bitstream. The format rule specifies that the first value of the auxiliary identifier for the video layer indicates that the video layer does not include auxiliary images, wherein the first value is equal to 0. The format rule specifies that the second value of the auxiliary identifier for the video layer indicates that the type of the auxiliary image for the video layer is α-plane, wherein the second value is equal to 1. The format rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of the auxiliary image of the video layer is a depth image, wherein the third value is equal to 2, and The video is a multi-functional video codec.

18. A non-transitory computer-readable recording medium having stored thereon instructions and a bitstream of video, wherein, when executed by a video processing apparatus, the instructions implement a method for generating the bitstream of the video, wherein, The method includes: The bitstream of the video is generated according to the format rules. The format rule specifies that the supplementary enhancement information field included in the bitstream indicates whether the bitstream is allowed to include one or more video layers representing auxiliary information; wherein the format rule further specifies that the supplementary enhancement information field is included in the scalability size information of the supplementary enhancement information message in the bitstream.

19. The non-transitory computer-readable recording medium according to claim 18, wherein, The format rules specify that the supplemental enhancement information message includes an auxiliary identifier for each video layer of one or more video layers of the bitstream. The format rule specifies that the first value of the auxiliary identifier for the video layer indicates that the video layer does not include auxiliary images, wherein the first value is equal to 0. The format rule specifies that the second value of the auxiliary identifier for the video layer indicates that the type of the auxiliary image for the video layer is α-plane, wherein the second value is equal to 1. The format rule specifies that the third value of the auxiliary identifier of the video layer indicates that the type of the auxiliary image of the video layer is a depth image, wherein the third value is equal to 2, and The video is a multi-functional video codec.

20. A method for storing a bitstream of video, comprising: A method is executed to generate the bitstream of the video. The bitstream is stored in a non-transitory computer-readable recording medium. in, The method includes: The bitstream of the video is generated according to a format rule, wherein the format rule specifies that the supplementary enhancement information field included in the bitstream indicates whether the bitstream is allowed to include one or more video layers representing auxiliary information; The format rules further specify that the supplementary enhancement information field is included in the scalability size information of the supplementary enhancement information message in the bitstream.

Citation Information

Patent Citations

  • Method and apparatus for coding video into bitstream carrying region-based post processing parameters into sei nesting message

    CN109479148A

  • Video encoding and decoding method, apparatus and system

    US20150172693A1