Signaling of Multi-View Information
By signaling information through SEI messages or VUI fields, the proposed solution addresses the challenge of processing multi-view and auxiliary information layers in VVC bitstreams, enhancing the efficiency and accuracy of video coding techniques.
Patent Information
- Application Number
- JP2023518729
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-29
- Filing Date
- 2021-09-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-09-29
AI Technical Summary
Existing digital video coding techniques struggle to efficiently process and differentiate between multi-view bitstreams and bitstreams with multiple layers representing auxiliary information, such as alpha or depth, within the VVC standard.
The proposed solution involves signaling information through SEI messages or VUI fields to indicate whether a VVC bitstream is multi-view or contains auxiliary information, using flags and identifiers to specify the nature of each layer.
This approach enables accurate identification and processing of multi-view and auxiliary information layers within VVC bitstreams, improving the efficiency and effectiveness of video encoding and decoding processes.
Smart Images

Figure 0007690577000005 
Figure 0007690577000006 
Figure 0007690577000007
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This This application claims priority and the benefit thereof to International Patent Application No. PCT / CN2020 / 118711, filed on September 29, 2020 is based on International Patent Application No. PCT / CN2021 / 121512, filed on September 29, 2021, claiming and is . Above hereby All patent applications incorporated their full texts are by reference incorporated herein by reference in its entirety.
[0002] [Technical Field] This patent specification relates to digital video coding techniques including video encoding, transcoding, or decoding.
Background Art
[0003] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of user devices capable of receiving and displaying video increases, the bandwidth demand for digital video utilization is expected to continue to grow.
Summary of the Invention
[0004] This specification discloses techniques that can be used by a video encoder and a decoder to process a coded representation of a video or an image according to a file format.
[0005] In an exemplary aspect, a method for processing video data is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to format rules, where the format rules are determined by a supplemental enhancement information field included in the bitstream or a video user - ability information syntax structure, such that it is indicated whether the bitstream has a multi - view bitstream in which a plurality of videos are coded in a plurality of video layers.
[0006] In another exemplary aspect, a method for processing video data is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to formatting rules, where the formatting rules are defined such that a supplemental enhancement information field included in the bitstream indicates whether the bitstream has one or more video layers representing auxiliary information.
[0007] In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video including video pictures and a coded representation of the video, where the coded representation complies with formatting rules that are defined such that a field included in the coded representation indicates that the video is a multi-view video.
[0008] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video including video pictures and a coded representation of the video, where the coded representation complies with formatting rules that are defined such that a field included in the coded representation indicates that the video is coded in a coded representation in multiple video layers.
[0009] In yet another exemplary aspect, a video encoder device is disclosed. The video encoder has a processor configured to implement the above method.
[0010] In yet another exemplary aspect, a video decoder device is disclosed. The video decoder has a processor configured to implement the above method.
[0011] In yet another exemplary aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0012] In yet another example aspect, a computer-readable medium storing a bitstream is disclosed. The bitstream is generated or processed using the methods described herein.
[0013] These and other features are described throughout this specification.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0015] Section headings are used in this specification for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to only that section. Further, the term H.266 is used in some descriptions only for ease of understanding and not to limit the scope of the disclosed techniques. As such, the techniques described herein are applicable to other video codec protocols and designs. In this specification, editorial changes are indicated to the text with strike-through for text deletion and highlighting for text addition with respect to the current draft of the VVC specification.
[0016] [1. Introduction] This specification relates to video coding techniques. Specifically, it relates to the signaling of scalability dimension information of a Versatile Video Coding (VVC) video bitstream. The ideas may be applied, individually or in various combinations, to any video coding standard or non-standard video codec, such as the recently finalized VVC.
[0017] [2. Acronyms] ACT Adaptive Colour Transform ALF Adaptive Loop Filter AMVR Adaptive Motion Vector Resolution APS Adaptation Parameter Set AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding(Rec.ITU-T H.265|ISO / IEC 14496-10) B Bi-predictive BCW Bi-prediction with CU-level Weights BDOF Bi-Directional Optical Flow BDPCM Block-Based Delta Pulse Code Modulation BP Buffering Period CABAC Context-based Adaptive Binary Arithmetic Coding CB Coding Block CBR Constant Bit Rate CCALF Cross-Component Adaptive Loop Filter CLVS Coded Layer Video Sequence CLVSS Coded Layer Video Sequence Start CPB Coded Picture Buffer CRA Clean Random Access CRC Cyclic Redundancy Check CTB Coding Tree Block CTU Coding Tree Unit CU Coding Unit CVS Coded Video Sequence CVSS Coded Video Sequence Start DCI Decoding Capability Information DPB Decoded Picture Buffer DRAP Dependent Random Access Point DU Decoding Unit DUI Decoding Unit Information EG Exponential-Golomb EGk k-th order Exponential-Golomb EOB End Of Bitstream EOS End Of Sequence FD Filler Data FIFO First-In, First-Out FL Fixed-Length GBR Green, Blue, and Red GCI General Constraints Information GDR Gradual Decoding Refresh GPM Geometric Partitioning Mode HEVC High Efficiency Video Coding(Rec.ITU-T H.265|ISO / IEC 23008-2) HRD Hypothetical Reference Decoder HSS Hypothetical Stream Scheduler I Intra IBC Intra Block Copy IDR Instantaneous Decoding Refresh ILRP Inter-Layer Reference Picture IRAP Intra Random Access Point LFNST Low Frequency Non-Separable Transform LPS Least Probable Symbol LSB Least Significant Bit LTRP Long-Term Reference Picture LMCS Luma Mapping with Chroma Scaling MIP Matrix-based Intra Prediction MPS Most Probable Symbol MSB Most Significant Bit MTS Multiple Transform Selection MVP Motion Vector Prediction NAL Network Abstraction Layer OLS Output Layer Set OP Operation Point OPI Operating Point Information P Predictive PH Picture Header POC Picture Order Count PPS Picture Parameter Set PROF Prediction Refinement with Optical Flow PT Picture Timing PU Picture Unit QP Quantization Parameter RADL Random Access Decodable Leading (picture) RASL Random Access Skipped Leading (picture) RBSP Raw Byte Sequence Payload RGB Red, Green, and Blue RPL Reference Picture List SAO Sample Adaptive Offset SAR Sample Aspect Ratio SEI Supplemental Enhancement Information SH Slice Header SLI Subpicture Level Information SODB String Of Data Bits SPS Sequence Parameter Set STRP Short-Term Reference Picture STSA Step-wise Temporal Sublayer Access TR Truncated Rice TU Transform Unit VBR Variable Bit Rate VCL Video Coding Layer VPS Video Parameter Set VSEI Versatile Supplemental Enhancement Information(Rec.ITU-T H.274|ISO / IEC 23002-7) VUI Video Usability Information VVC Versatile Video Coding(Rec.ITU-T H.266|ISO / IEC 23090-3)
[0018] [3. Initial Discussion] [3.1. Video Coding Standard] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and the two organizations jointly created the H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standard specifications. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes time prediction plus transform coding. To explore future video coding technologies beyond HEVC, the JVET (Joint Video Exploration Team) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by the JVET and placed in the reference software named JEM (Joint Exploration Model). Later, when the VVC (Versatile Video Coding) project officially started, the JVET was renamed the JVET (Joint Video Experts Team). VVC is the new coding standard specification that was finally agreed upon at the 19th JVET, which ended on June 1, 2020, and aims to reduce the bit rate by 50% compared to HEVC.
[0019] The VVC (Versatile Video Coding) standard specification (ITU-T H.266 | ISO / IEC 23090-3) and the related VSEI (Versatile Supplemental Enhancement Information) standard specification (ITU-T H.274 | ISO / IEC 23002-7) are designed for use in a maximally wide range of applications, including both conventional uses such as television broadcasting, video conferencing, or playback from storage media, and newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple coded video bitstreams, multi-view video, scalable layer coding, and viewport-adaptive 360° immersive video.
[0020] [3.2. Video based Point Cloud Compression (V-PCC)] ISO / IEC 23090-5, Information technology - Coded Representation of Immersive Media - Part 5: Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC), which is also abbreviated as V-PCC, is a standard specification that defines the coded representation of point cloud signals. The V-PCC standard specification is another recently finalized standard specification.
[0021] V-PCC defines data types such as occupancy, geometry, texture attributes, material attributes, transparency attributes, reflectance attributes, and standard attributes that can be coded using specific video codecs such as VVC, HEVC, AVC, etc.
[0022] [3.3. Support for Temporal Scalability in VVC] VVC includes similar support for temporal scalability as seen in HEVC. Such support includes signaling of the temporal ID in the NAL unit header, a restriction that pictures of a particular temporal sublayer cannot be used as inter-prediction references by pictures of lower temporal sublayers, the sub-bitstream extraction process, and the requirement that each sub-bitstream extraction output of appropriate input must be a conforming bitstream. MANE (Media-Aware Network Element(s)) can utilize the temporal ID within the NAL unit header for stream adaptation based on temporal scalability.
[0023] [3.4. Change of Picture Resolution within a Sequence in VVC] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC enables the change of picture resolution within a sequence at certain positions without encoding an IRAP picture that is always intra-coded. This feature is sometimes called Reference Picture Resampling (RPR) because it requires resampling of the reference picture when the reference picture used for inter-prediction has a different resolution from the current picture being decoded.
[0024] To enable reusing the motion compensation module of existing implementations, the scaling ratio is restricted to be 1 / 2 or more (2x downsampling from the reference picture to the current picture) and 8 or less (8x upsampling). The horizontal and vertical scaling ratios are derived based on the picture widths and heights specified for the reference picture and the current picture, as well as the left, right, up, and down scaling offsets.
[0025] Without the need to code IRAP pictures that cause instantaneous bitrate spikes in streaming or video conferencing scenarios to adapt to changes in network conditions, for example, RPR enables resolution changes. RPR can also be used in adaptive scenarios where zooming of the entire video area or an area of interest is required. The scaling window offset can be negative to support applications based on a wider range of zooming. A negative scaling window offset also enables the extraction of a sub-picture sequence from a multi-layer bitstream while maintaining the same scaling window as seen in the original bitstream for the extracted sub-bitstream.
[0026] Unlike the spatial scalability in the scalable extension of HEVC where picture resampling and motion compensation are applied in two different stages, RPR in VVC is executed as part of the same process at the block level where the derivation of sample positions and motion vector scaling are performed during motion compensation.
[0027] Aiming to limit implementation complexity, changes in picture resolution within CLVS are not permitted if the pictures within CLVS have multiple sub-pictures per picture. Furthermore, decoder side motion vector refinement, bi-directional optical flow, and prediction refinement with optical flow are not applied when RPR is used between the current picture and the reference picture. Collocated pictures for the derivation of temporal motion vector candidates are also restricted to have the same picture size, scaling window offset, and CTU size as the current picture.
[0028] Regarding the support of RPR, other aspects of the VVC design have been made different from HEVC. First, the picture resolution and the corresponding conformance and scaling window are signaled in the PPS rather than in the SPS, while in the SPS, the maximum picture resolution and the corresponding conformance window are signaled. In applications, the maximum picture resolution with the corresponding conformance window offset in the SPS can be used as the intended or desired picture output size after cropping. Second, in the case of a single-layer bitstream, each picture store (slot in the DPB for storing one decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.
[0029] [3.5. Support for Multilayer Scalability in VVC] By having the ability to perform inter prediction from reference pictures of different sizes than the current picture using RPR in the VVC core design, VVC can easily support a bitstream that includes multiple layers of different resolutions, for example, two layers each having standard definition resolution and high definition resolution. In the VVC decoder, such a function can be incorporated without the need for any additional signal processing level coding tools, in that the upsampling function required for the support of spatial scalability can be provided by reusing the RPR upsampling filter. Nevertheless, additional high-level syntax design is required to enable the scalability support of the bitstream.
[0030] Scalability is supported in VVC but is only included in the multi-layer profile. Different from the scalability support in any future video coding standard including the extensions of AVC and HEVC, the design of VVC scalability is made to be as suitable as possible for single-layer decoder design. The decoding capabilities for multi-layer bitstreams are defined as if there were only a single layer in the bitstream. For example, decoding capabilities such as the DPB size are defined in a way that does not depend on the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not require significant changes to be able to decode a multi-layer bitstream.
[0031] Compared with the design of the multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, 1) IRAP AUs are required to include pictures for each layer present in the CVS, eliminating the need to define a decoding process that starts at the layer level, and 2) instead of a complex POC re-setting mechanism, a much simpler design for POC signaling is included in VVC, ensuring that the derived POC values are the same for all pictures within an AU.
[0032] Similar to HEVC, information regarding layers and layer dependencies is included in the VPS. OLS information is provided for signaling which layers are included in the OLS, which layers are output, and other information such as the PTL and HRD parameters associated with each OLS. Similar to HEVC, there are three operating modes to output all layers, only the top layer, or any specific specified layer in a custom output mode.
[0033] There are several differences between the OLS design in VVC and the OLS design in HEVC. First, in HEVC, a layer set is signaled, then the OLS is signaled based on the layer set, and for each OLS, the output layer is signaled. With the HEVC design, a layer can belong to an OLS that is neither the output layer nor a layer necessary to decode the output layer. In VVC, the design requires that any layer within an OLS is either the output layer or a layer necessary to decode the output layer. Thus, in VVC, the OLS is signaled by indicating the output layer of the OLS, and then the other layers belonging to the OLS are easily derived by the layer dependencies indicated in the VPS. Further, VVC requires that each layer be included in at least one OLS.
[0034] Another difference in the VVC OLS design is that, in contrast to HEVC where the OLS is composed of all NAL units belonging to the identified set of layers mapped to the OLS, VVC can exclude some NAL units belonging to non-output layers mapped to the OLS. More specifically, the VVC OLS consists of a set of layers mapped to the OLS that includes non-output layers containing only pictures from IRAP or GDR pictures with ph_recovery_poc_cnt equal to 0 or pictures from sub-layers used for inter-layer prediction. This enables showing the optimal level values of the multi-layer bitstream considering only the "necessary" pictures of all sub-layers within the layers forming the OLS. Here, "necessary" means necessary for output or decoding. Figure 7 shows an example of a two-layer bitstream with vps_max_tid_il_ref_pics_plus1[1][0] equal to 0, i.e., a sub-bitstream where only the IRAP pictures from layer L0 are retained when OLS2 is extracted.
[0035] Considering several scenarios where it is beneficial to allow different RAP periodicities in different layers, similar to AVC and HEVC, an AU is allowed to have layers that contain unaligned RAPs. For faster identification of RAPs within a multi-layer bitstream, i.e., AUs that have RAPs in all layers, the Access Unit Delimiter (AUD) is extended compared to HEVC, which has a flag indicating whether the AU is an IRAP AU or a GDR AU. Further, the AUD is mandated to be present in such IRAP or GDR AUs when the VPS indicates multiple layers. However, in the case of a single-layer bitstream indicated by the VPS or a bitstream that does not refer to the VPS, the AUD is completely optional, as in HEVC. This is because in this case, the RAP can be easily detected from the NAL unit type of the first slice within the AU and each parameter set.
[0036] To enable sharing of SPS, PPS, and APS by multiple layers, and at the same time to ensure that the bitstream extraction process does not waste the parameter sets required in the decoding process, the VCL NAL units of the first layer can refer to SPS, PPS, or APS with the same or lower layer ID values as long as all OLSs including that first layer also include layers identified by lower layer ID values.
[0037] [3.6.VUI and SEI Messages] VUI is a syntax structure that is transmitted as part of the SPS (and, optionally, also in the HEVC VPS). VUI carries information that may be important for the proper rendering of the coded video although it does not affect the canonical decoding process.
[0038] The SEI supports processes related to decoding, display, or other purposes. Similar to the VUI, the SEI does not affect the canonical decoding process. The SEI is carried in SEI messages. Decoder support for SEI messages is optional. However, SEI messages affect bitstream compliance (e.g., if the syntax of SEI messages in the bitstream does not conform to the specification, the bitstream is non-compliant), and some SEI messages are required by the HRD specification.
[0039] The VUI syntax structure and most SEI messages used in VVC are not defined in the VVC specification, but rather in the VSEI specification. The SEI messages required for the HRD compliance test are defined in the VVC specification. VVC v1 defines five SEI messages related to the HRD compliance test, and VSEI v1 specifies 20 additional SEI messages. SEI messages carried in the VSEI specification do not directly affect the compliant decoder behavior and are defined in a way that they can be used in a way not tied to the coding format, thus allowing VSEI to be used in the future by other video coding standards in addition to VVC. Instead of specifically referring to VVC syntax element names, the VSEI specification refers to variables whose values are set within the VVC specification.
[0040] Compared to HEVC, the VUI syntax structure of VVC focuses only on information related to the proper rendering of pictures and does not include any timing information or bitstream restriction instructions. In VVC, the VUI is signaled within the SPS, and the SPS includes a length field before the VUI syntax structure to notify the length of the VUI payload in bytes. This allows the decoder to easily skip over the information, and more importantly, enables a convenient future VUI syntax structure by directly adding new syntax elements at the end of the VUI syntax structure in a similar way to SEI message syntax extensions.
[0041] The VUI syntax structure includes the following information: ● Content that is interlaced or progressive; ● Whether the content includes frame-packed stereoscopic video or projected omnidirectional video; ● Sample aspect ratio: ● Whether the content is suitable for overscan display; ● Color description including primary colors, matrices, and transfer characteristics, which is particularly important for signaling the color space and high dynamic range (HDR) of ultra-high definition (UHD) versus high definition (HD) to enable signaling; ● The position of chroma (chroma) compared to luminance (luma) (for progressive content compared to HEVC, the signaling is clarified).
[0042] When the SPS does not include any VUI, the information is considered not specified and must be specified by the application or conveyed by external means if the bitstream is intended for rendering on a display.
[0043] Table 1 lists all the SEI messages defined for VVC v1 and their specifications including syntax and semantics. Of the 20 SEI messages defined in the VSEI specification, many are inherited from HEVC (e.g., filler payload and both user data SEI messages). Some SEI messages are essential for the accurate processing or rendering of the coded video content. This applies, for example, to the master display color volume, content light level information, or alternative transfer characteristic SEI messages particularly relevant to HDR content. Other examples include the orthographic projection, spherical rotation, per-region packing, omnidirectional viewport SEI messages, etc., which are related to the signaling and processing of 360° video content.
Table 1
[0044] The new SEI messages defined for VVC v1 include frame field SEI messages, sample aspect ratio information SEI messages, and sub-picture level information SEI messages.
[0045] The frame field SEI message includes information indicating how the associated picture should be displayed (field parity or frame repetition period), the scan type of the associated picture, and whether the associated picture is a duplicate of the previous picture. This information, together with the timing information of the associated image, was notified in the picture timing SEI message in previous video coding standard specifications. However, it has been observed that the frame field information and the timing information are two different types of information that are not necessarily signaled together. A typical example consists of signaling the timing information at the system level while signaling the frame field information within the bitstream. Therefore, it was decided to remove the frame field information from the picture timing SEI message and signal it instead within a dedicated SEI message. This change also made it possible to change the syntax of the frame field information to convey more explicit instructions to the display, such as field pairing and more values for frame repetition.
[0046] The sample aspect ratio SEI message makes it possible to signal different sample aspect ratios for different pictures within the same sequence, while the corresponding information included in the VUI applies to the entire sequence. It may be relevant when using the reference picture resampling function with scaling factors that give different pictures within the same sequence different sample aspect ratios.
[0047] The sub-picture level information SEI message provides the level information of the sub-picture sequence.
[0048] [4. Technical problems to be solved by the disclosed technical solution] VVC supports multi-layer scalability. However, given a VVC multi-layer bitstream, it is unknown whether the OLS bitstream is a multi-view bitstream or simply a bitstream consisting of multiple layers with SNR and / or spatial scalability. Furthermore, given a VVC multi-layer bitstream, it is unknown whether there are one or more layers representing auxiliary information such as alpha, depth, etc., and if so, which layer represents what.
[0049] [5. List of technical solutions] To solve the above problems, the methods summarized below are disclosed. The inventions should be regarded as examples for explaining the overview and should not be interpreted in a narrow sense. Furthermore, these inventions may be applied individually or combined in any way. 1) Information indicating whether the VVC video bitstream is a multi-view bitstream is signaled in the VVC video bitstream. a. In one example, the information is signaled by an SEI message (e.g., called a scalability dimension SEI message). i. In one example, the scalability dimension SEI message provides information on the bitstream bitstreamInScope. The bitstream bitstreamInScope is defined as a sequence of AUs in decoding order that have zero or more AUs following the current scalability dimension SEI message, up to and including any subsequent AU that contains the scalability dimension SEI message, but not including that AU itself. ii. In one example, the SEI message includes a flag indicating whether the bitstream can be a multi-view bitstream. iii. In one example, the SEI message indicates the view ID of each layer. 1. In one example, the SEI message includes a flag indicating whether the view ID is signaled layer by layer. 2. In one example, the length in bits of the view ID for each layer is signaled in the SEI message. b. In one example, the information is signaled as part of the VUI. 2) Information indicating whether the VVC video bitstream includes one or more layers representing auxiliary information is signaled in the VVC video bitstream. a. In one example, the information is signaled in an SEI message (e.g., a scalability dimension SEI message). i. In one example, the scalability dimension SEI message provides information on the bitstream bitstreamInScope. The bitstream bitstreamInScope is defined as a sequence of zero or more AUs that follow the AU containing the current scalability dimension SEI message up to and not including any subsequent AU that contains a scalability dimension SEI message. ii. In one example, the SEI message includes a flag indicating whether the bitstream can contain auxiliary information carried by one or more layers. iii. In one example, the SEI message indicates the auxiliary ID of each layer. 1. In one example, the SEI message includes a flag indicating whether the auxiliary ID is signaled layer by layer. 2. In one example, the value of the auxiliary ID (e.g., 0) indicates that the layer does not include an auxiliary picture. 3. In one example, the value of the auxiliary ID (e.g., 1) indicates that the type of the auxiliary information is alpha. 4. In one example, the value of the auxiliary ID (e.g., 2) indicates that the type of the auxiliary information is depth. 5. In one example, the value of the auxiliary ID (e.g., 3) indicates that the type of the auxiliary information is occupancy (e.g., as defined in V-PCC). 6. In one example, the value of the auxiliary ID (e.g., 4) indicates that the type of the auxiliary information is geometry (e.g., as defined in V-PCC). 7. In one example, the value of the auxiliary ID (e.g., 5) indicates that the type of the auxiliary information is attribute (e.g., as defined in V-PCC). 8. In one example, the value of the auxiliary ID (e.g., 6) indicates that the type of the auxiliary information is texture attribute (e.g., as defined in V-PCC). 9. In one example, the value of the auxiliary ID (e.g., 7) indicates that the type of the auxiliary information is material attribute (e.g., as defined in V-PCC). 10. In one example, the value of the auxiliary ID (e.g., 8) indicates that the type of the auxiliary information is transparency attribute (e.g., as defined in V-PCC). 11. In one example, the value of the auxiliary ID (e.g., 9) indicates that the type of the auxiliary information is reflectance attribute (e.g., as defined in V-PCC). 12. In one example, the value of the auxiliary ID (e.g., 10) indicates that the type of the auxiliary information is standard attribute (e.g., as defined in V-PCC). b. In one example, the information is signaled as part of the VUI.
[0050] [6. Embodiment] The following are some exemplary embodiments of some aspects of the present invention summarized in Section 5 above that can be applied to the VVC specification and the VSEI specification.
[0051] [6.1. First Embodiment] This example relates to items 1, 1.a, and all of its sub-items, 2, 2.a, 2.a.i, 2.a.ii, 2.a.iii, 2.a.iii.1, 2.a.iii.2, 2.a.iii.3, and 2.a.iii.4.
[0052] [6.1.1. Syntax of the Scalability Dimension SEI Message] [Table 2]
[0053] [6.1.2. Semantics of the Scalability Dimension SEI Message] The Scalability Dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), which is, 1) the view ID of each layer if bitstreamInScope can be a multi-view bitstream, and 2) the auxiliary ID of each layer if there can be auxiliary information (e.g., depth or alpha) carried by one or more layers in bitstreamInScope. bitstreamInScope is a contiguous sequence of AUs in decoding order that has zero or more AUs following the current AU containing the Scalability Dimension SEI message, up to and not including any subsequent AU containing the Scalability Dimension SEI message. One plus sd_max_layers_minus1 indicates the maximum number of layers within bitstreamInScope. An sd_multiview_info_flag equal to 1 indicates that bitstreamInScope can be a multi-view bitstream and that the sd_view_id_val[] syntax element is present in the Scalability Dimension SEI message. An sd_multiview_info_flag equal to 0 indicates that bitstreamInScope is not a multi-view bitstream and that the sd_view_id_val[] syntax element is not present in the Scalability Dimension SEI message. The sd_auxilary_info_flag equal to 1 indicates that there may be auxiliary information carried by one or more layers within bitstreamInScope, and the sd_aux_id[] syntax element exists in the scalability dimension SEI message. The sd_auxilary_info_flag equal to 0 indicates that there is no auxiliary information carried by one or more layers within bitstreamInScope, and the sd_aux_id[] syntax element does not exist in the scalability dimension SEI message. sd_view_id_len specifies the length in bits of the sd_view_id_val[i] syntax element. sd_view_id_val[i] specifies the view ID of the i-th layer within bitstreamInScope. The length of the sd_view_id_val[i] syntax element is sd_view_id_len bits. If it does not exist, the value of sd_view_id_val[i] is presumed to be equal to 0. The sd_aux_id[i] equal to 0 indicates that the i-th layer within bitstreamInScope does not contain an auxiliary picture. The sd_aux_id[i] greater than 0 indicates the type of the auxiliary picture in the i-th layer within bitstreamInScope as specified in Table 2 below.
Table 3
[0054] FIG. 1 is a block diagram showing an exemplary video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input unit 1902 that receives video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. The input unit 1902 may correspond to a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular networks.
[0055] System 1900 may include a coding component 1904 that can implement various coding or encoding methods described herein. The coding component 1904 may reduce the average bitrate of the video from the input unit 1902 to the output unit of the coding component 1904 so as to generate a coded representation of the video. Coding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the coding component 1904 may be stored or transmitted via a connected communication as represented by the component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input unit 1902 may be used by a component 1908 that generates a displayable video that is sent to the pixel values or the display interface 1910. The process of generating a video that a user can view from the bitstream representation is sometimes referred to as video decompression. Further, while certain video processing operations are referred to as "coding" operations or tools, it will be understood that such coding tools or operations are used in an encoder and the corresponding decoding tools or operations that reverse the results of the coding are to be performed by a decoder.
[0056] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI (registered trademark)) or a Displayport (registered trademark), etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI (Peripheral Component Interconnect), IDE (Integrated Drive Electronics) interface, etc. The techniques described herein may be embodied in various electronic devices such as a cellular phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0057] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 can be used to implement one or more of the methods described herein. The apparatus 3600 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor 3602 may be configured to implement one or more of the methods described herein. The memory (memories) 3604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described herein in a hardware circuit. In some embodiments, the video processing hardware 3606 may be at least partially included in the processor 3602, such as a graphics coprocessor.
[0058] Figure 4 is a block diagram representing an exemplary video coding system 100 that can utilize the techniques of the present disclosure.
[0059] As shown in Figure 4, the video coding system 100 may include a source device 110 and a destination device 120. The source device 110 may generate encoded video data and may be referred to as a video encoding device. The destination device 120 may be able to decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0060] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0061] Video source 112 may include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system that generates video data, or a combination of such sources. The video data may have one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits that forms a coded representation of the video data. The bitstream may also include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly through network 130a to destination device 120 via I / O interface 116. The encoded video data may also be stored in storage medium / server 130b for access by destination device 120.
[0062] Destination device 120 may include I / O interface 126, video decoder 124, and display device 122.
[0063] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain the encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to the user. Display device 122 may be integrated with destination device 120 or may be configured to interface with an external display device and be outside of destination device 120.
[0064] Video encoder 114 and video decoder 124 may operate according to video compression standards such as the HEVC (High Efficiency Video Coding) standard, the VVC (Versatile Video Coding) standard, and other current and / or further standard specifications.
[0065] FIG. 5 is a block diagram representing an example of a video encoder 200, which may be the video encoder 114 of the system 100 shown in FIG. 4.
[0066] The video encoder 200 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 5, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0067] The functional components of the video encoder 200 may include a partitioning unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0068] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode where at least one reference picture is the picture in which the current video block is located.
[0069] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but are shown separately in the example of FIG. 5 for the sake of explanation.
[0070] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0071] The mode selection unit 203 may select, for example, one of the intra or inter coding modes based on the error result, and supply the resulting intra or inter coded block to the residual generation unit 207 that generates residual block data, and to the reconstruction unit 212 that reconstructs the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select an Intra-Inter Composite Prediction (CIIP) mode in which the prediction is based on both an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution (e.g., sub-pixel or integer pixel accuracy) for the motion vector of the block in the case of inter prediction.
[0072] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block of the current video block based on the motion information and the decoded samples of the picture from the buffer 213 other than the picture associated with the current video block.
[0073] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for the current video block, for example, according to whether the current video block is an I slice, a P slice, or a B slice.
[0074] In some examples, the motion estimation unit 204 may perform uni-directional prediction for the current video block. The motion estimation unit 204 may search for a reference video block for the current video block from the reference pictures in list 0 or list 1. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that includes the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0075] In other examples, the motion estimation unit 204 may perform bi-directional prediction for the current video block. The motion estimation unit 204 may search for a reference video block for the current video block from the reference pictures in list 0, and may also search for another reference video block for the current video block from the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in lists 0 and 1 that include the reference video blocks, and a motion vector indicating the spatial displacement between those reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0076] In some examples, the motion estimation unit 204 may output a full set of motion information for the decoder's decoding process.
[0077] In some examples, the motion estimation unit 204 may not output a full set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of other video blocks. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0078] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 indicating that the current video block has the same motion information as another video block in the syntax structure associated with the current video block.
[0079] In other examples, the motion estimation unit 204 may identify another video block and a Motion Vector Difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0080] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 are Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0081] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks within the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0082] The residual generation unit 207 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block (e.g., indicated by a negative sign). The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples within the current video block.
[0083] In other examples, for example, in skip mode, for the current video block, there may be no residual data for the current video block, and the residual generation unit 207 may not need to perform a subtraction operation.
[0084] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0085] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0086] The inverse quantization unit 210 and the inverse transform unit 211 may each apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 and generate a reconstructed video block related to the current block for storage in the buffer 213.
[0087] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0088] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0089] FIG. 6 is a block diagram illustrating an example of a video decoder 300, which may be the video decoder 124 of the system 100 shown in FIG. 4.
[0090] The video decoder 300 may be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 6, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0091] In the example of FIG. 6, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 may perform a decoding path generally inverse to the encoding path described with respect to video encoder 200 (FIG. 5).
[0092] Entropy decoding unit 301 may extract the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 may decode the entropy-coded video data, and from the entropy-decoded video data, motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 may determine such information, for example, by performing AMVP and merge mode.
[0093] Motion compensation unit 302 may optionally perform interpolation based on an interpolation filter to generate a motion-compensated block. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.
[0094] Motion compensation unit 302 may use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolation values for sub-integer pixels of the reference block. Motion compensation unit 302 may determine the interpolation filter used by video encoder 200 according to the received syntax information and use that interpolation filter to generate a prediction block.
[0095] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, the partition information that describes how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence.
[0096] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes, i.e., dequantizes, the quantized video block coefficients supplied in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0097] The reconstruction unit 306 may add the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303 to the residual block to form a decoded block. Optionally, a deblocking filter may also be applied to the decoded block to filter it in order to remove blocking artifacts. The decoded video blocks are then stored in the buffer 307, which supplies reference blocks for subsequent motion compensation / intra prediction and further generates the decoded video for presentation on a display device.
[0098] A list of solutions preferred by some embodiments is given next.
[0099] The following shows exemplary embodiments of the techniques discussed in the previous section (e.g., item 1).
[0100] Solution 1. A method of video processing (e.g., method 700 represented in FIG. 3), comprising: executing a step (702) of converting between a video including video pictures and a coded representation of the video; wherein the coded representation complies with format rules; the format rules define that a field indicating that the video is a multi-view video is included in the coded representation; Method.
[0101] Solution 2. the field is included in a supplementary enhancement information part of the coded representation; The method of Solution 1.
[0102] Solution 3. the field is included in a video user usability information part of the coded representation; The method of Solution 1.
[0103] The following shows exemplary embodiments of the technology discussed in the previous section (e.g., item 2).
[0104] Solution 4. A method of video processing, comprising: executing a step of converting between a video including video pictures and a coded representation of the video; wherein the coded representation complies with format rules; the format rules define that a field indicating that the video is coded in the coded representation in a plurality of video layers is included in the coded representation; Method.
[0105] Solution 5. the field is included in a supplementary enhancement information part of the coded representation; The method of Solution 4.
[0106] Solution 6. The field is included in the video user usability information part of the coded representation, the method of Solution 4.
[0107] Solution 7. The conversion has generating the coded representation from the video, the method of any one of Solutions 1 to 6.
[0108] Solution 8. The conversion has decoding the coded representation to generate the video, the method of any one of Solutions 1 to 6.
[0109] Solution 9. A video decoding device having a processor configured to implement the method described in one or more of Solutions 1 to 8.
[0110] Solution 10. A video encoding device having a processor configured to implement the method described in one or more of Solutions 1 to 8.
[0111] Solution 11. A computer program product storing computer code that, when executed by a processor, causes the processor to implement the method described in any one of Solutions 1 to 8.
[0112] Solution 12. A computer-readable medium storing a coded representation generated according to any one of Solutions 1 to 8.
[0113] The method, device, or system described herein.
[0114] In the solution described in this specification, the encoder may comply with the formatting rules by generating the coded representation in accordance with the formatting rules. In the solution described in this specification, the decoder may use the formatting rules to parse the syntax elements in the coded representation after knowing the presence or absence of the syntax elements that comply with the formatting rules in order to generate the decoded video.
[0115] FIG. 8 is a flowchart of an exemplary method of video processing. Operation 802 includes performing a conversion between a video and a bitstream of the video in accordance with formatting rules, where the formatting rules are defined such that the bitstream indicates whether the bitstream has a multi-view bitstream in which multiple views are coded in multiple video layers by virtue of a supplemental enhancement information field included in the bitstream or a video user capability information syntax structure.
[0116] In some embodiments, the formatting rules are defined such that the supplemental enhancement information field is included in the scalability dimension information in the supplemental enhancement information message in the bitstream. In some embodiments, the formatting rules are defined such that the supplemental enhancement information message includes a first flag indicating whether the bitstream is a multi-view bitstream. In some embodiments, the formatting rules are defined such that the supplemental enhancement information message includes the view identifier of each video layer among the multiple video layers of the bitstream. In some embodiments, the formatting rules are defined such that the supplemental enhancement information message includes the bit length of the view identifier of each video layer.
[0117] In some embodiments, the formatting rule defines that the supplemental enhancement information message includes a second flag indicating whether the view identifier is included in the bitstream for each video layer. In some embodiments, the formatting rule provides information regarding a series of access units that, in decoding order, includes an access unit including second scalability dimension information in a second supplemental enhancement information message, where zero or more access units follow that access unit and include subsequent access units up to and including a subsequent access unit including third scalability dimension information in a third supplemental enhancement information message, where the scalability dimension information in the supplemental enhancement information message is included in the subsequent access unit. In some embodiments, the formatting rule defines that the supplemental enhancement information field is included in a video usability information syntax structure in the bitstream. In some embodiments, the bitstream is a versatile video coding bitstream. In some embodiments, the step of performing the conversion includes encoding the video into a bitstream. In some embodiments, the step of performing the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments, the step of performing the conversion includes decoding the video from the bitstream.
[0118] FIG. 9 is a flowchart of an exemplary method of video processing. Operation 902 includes performing a conversion between a video and a bitstream of the video according to a formatting rule, the formatting rule defining that the bitstream includes a supplemental enhancement information field indicating whether the bitstream has one or more video layers representing auxiliary information.
[0119] In some embodiments, the formatting rule defines that the supplementary enhancement information field is included in the scalability dimension information within the supplementary enhancement information message in the bitstream. In some embodiments, the formatting rule defines that the supplementary enhancement information message includes a first flag indicating whether the bitstream includes auxiliary information for one or more video layers. In some embodiments, the formatting rule defines that the supplementary enhancement information message includes an auxiliary identifier for each of a plurality of video layers in the bitstream. In some embodiments, the formatting rule defines that a first value of the auxiliary identifier of a video layer indicates that the video layer does not include an auxiliary picture.
[0120] In some embodiments, the formatting rule defines that a second value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of the video layer is alpha. In some embodiments, the formatting rule defines that a third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of the video layer is depth. In some embodiments, the formatting rule defines that the supplementary enhancement information message includes a second flag indicating whether the auxiliary identifier is included in the bitstream for each video layer. In some embodiments, the formatting rule defines that a fourth value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of the video layer is occupancy.
[0121] In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is geometry. In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is an attribute. In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is a texture attribute. In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is a material attribute. In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is a transparency attribute.
[0122] In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is a reflectance attribute. In some embodiments, the formatting rule is defined such that the third value of the auxiliary identifier of a video layer indicates that the type of the auxiliary information of that video layer is a normal attribute. In some embodiments, the formatting rule provides information regarding a series of access units that, in decoding order, includes an access unit that includes second scalability dimension information in a second supplementary enhancement information message, followed by zero or more access units that include third scalability dimension information in a third supplementary enhancement information message up to, but not including, a subsequent access unit that includes third scalability dimension information in a subsequent access unit that includes the third scalability dimension information in the third supplementary enhancement information message.
[0123] In some embodiments, the formatting rule defines that the supplementary enhancement information field is included in the video user capability information in the bitstream. In some embodiments, the video is a versatile video coding video. In some embodiments, the step of performing the conversion includes encoding the video into a bitstream. In some embodiments, the step of performing the conversion includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments, the step of performing the conversion includes decoding the video from the bitstream.
[0124] In some embodiments, the video decoding device has a processor configured to implement one or more of the methods recited in this patent document. In some embodiments, the video encoding device has a processor configured to implement one or more of the methods recited in this patent document. In some embodiments, the computer program product stores computer instructions that cause the processor to implement the techniques recited in this patent document when executed by the processor. In some embodiments, the non-transitory computer-readable storage medium stores a bitstream generated according to any one of the methods recited in this patent document.
[0125] In some embodiments, the non-transitory computer-readable storage medium stores instructions for a processor to perform any of the methods recited in this patent document. In some embodiments, the method of bitstream generation includes generating a video bitstream according to any of the methods recited in this patent document, and storing the bitstream in a computer-readable program medium. In some embodiments, the method, apparatus, bitstream generated according to the disclosed method, or system recited in this patent document.
[0126] In this patent document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may correspond to bits that are either at the same position within the bitstream or spread out at different locations, as defined by the syntax, for example. For example, a macroblock may be encoded using the transformed and coded error residual values, and further, bits in the headers and other fields within the bitstream. Further, during the conversion, the decoder may parse the bitstream based on a determination that there may or may not be some fields, as described in the above solutions. Similarly, the encoder may determine whether a particular syntax field should or should not be included, and accordingly, generate a coded representation by including or excluding the syntax field in the coded representation.
[0127] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in processing video blocks, but based on the use of the tool or mode, it is not necessarily required to change the resulting bitstream. That is, the conversion from a video block to a video bitstream representation will use that tool or mode when the video processing tool or mode is enabled based on a decision or determination. In other examples, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been changed based on that video processing tool or mode. That is, the conversion from a video bitstream representation to a video block will be performed using a video processing tool or mode enabled based on a decision or determination.
[0128] Some embodiments of the disclosed technology involve making a decision or determination to disable a video processing tool or mode. In an example, when a video processing tool or mode is disabled, the encoder does not use that tool or mode in the conversion from a video block to a video bitstream representation. In other examples, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been changed using a video processing tool or mode disabled based on a decision or determination.
[0129] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” includes, by way of example, all apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer programs in question, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., an electrical, optical, or electromagnetic signal generated by a machine, that is generated to encode information for transmission to an appropriate receiver device.
[0130] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code), or in portions of files that hold other programs or data (e.g., one or more scripts stored in a markup language document). A computer program can be deployed to be executed on one computer, or on one location, or on multiple computers distributed across multiple locations and interconnected by a communication network.
[0131] The processes and logic flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data to generate output. The processes and logic flows can also be executed by dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus can be implemented as such.
[0132] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or be operatively coupled to receive data from or transfer data to one or more such mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks, including all forms of non-volatile memory, media, and memory devices. The processor and the memory may be enhanced or incorporated in a dedicated logic circuit.
[0133] This specification includes many details, but they are to be construed not as limitations on the scope of any subject or of anything that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. The particular features described herein in connection with separate embodiments may be implemented in combination in a single embodiment. Conversely, the various features described in connection with a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may be described above as acting in certain combinations and even initially claimed as such, but one or more features from a claimed combination may in some cases be excisable from that combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0134] Similarly, operations are represented in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in that particular order or in a sequential order to achieve the desired result, or that all of the operations shown be performed. Further, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all embodiments.
[0135] Only a few implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
Claim 1 A method for processing video data, comprising: executing a conversion between a video and a bitstream of the video according to format rules; wherein the format rules are defined such that the bitstream is allowed to have a multi-view bitstream in which a plurality of views are coded in a plurality of video layers, as indicated by a supplementary enhancement information (SEI) field included in the bitstream; wherein the format rules are defined such that the supplementary enhancement information (SEI) field is included in scalability dimension information within a supplementary enhancement information (SEI) message in the bitstream; A method. Claim 2 wherein the format rules are defined such that the supplementary enhancement information (SEI) message includes a view identifier for each video layer of the plurality of video layers of the bitstream; The method according to claim 1. Claim 3 wherein the format rules are defined such that the supplementary enhancement information (SEI) message includes the length of the bits of the view identifier for each video layer; The method according to claim 2. Claim 4 wherein the format rules are defined such that the supplementary enhancement information (SEI) message includes a flag indicating whether the view identifier is included in the bitstream for each video layer; The method according to claim 2. Claim 5 wherein the bitstream is a versatile video coding bitstream; The method according to any one of claims 1 to 4. Claim 6 wherein the step of executing the conversion includes encoding the video into the bitstream; The method according to any one of claims 1 to 5. Claim 7 wherein the step of executing the conversion includes decoding the video from the bitstream; The method according to any one of claims 1 to 5. Claim 8 A processor and a non-transitory memory including instructions, wherein when the instructions are executed by the processor, the processor is caused to execute a conversion between a video and a bitstream of the video according to format rules, The format rule is defined such that the supplementary enhancement information (SEI) field included in the bitstream indicates whether the bitstream is allowed to have a multi-view bitstream in which a plurality of views are coded in a plurality of video layers. The format rule is defined such that the supplementary enhancement information (SEI) field is included in scalability dimension information in a supplementary enhancement information (SEI) message in the bitstream. Device. Claim 9 Store in a processor instructions for causing the processor to perform a conversion between video and a bitstream of the video according to a format rule. The format rule is defined such that the supplementary enhancement information (SEI) field included in the bitstream indicates whether the bitstream is allowed to have a multi-view bitstream in which a plurality of views are coded in a plurality of video layers. The format rule is defined such that the supplementary enhancement information (SEI) field is included in scalability dimension information in a supplementary enhancement information (SEI) message in the bitstream. Non-transitory computer-readable storage medium. Claim 10 A method for storing a bitstream of video, comprising: generating the bitstream of the video according to a format rule; and storing the bitstream in a non-transitory computer-readable recording medium. The format rule is defined such that the supplementary enhancement information (SEI) field included in the bitstream indicates whether the bitstream is allowed to have a multi-view bitstream in which a plurality of views are coded in a plurality of video layers. The format rule is defined such that the supplementary enhancement information (SEI) field is included in scalability dimension information in a supplementary enhancement information (SEI) message in the bitstream. Method.
Citation Information
Patent Citations
Video encoding system and method of operating the video encoding system
JP2015508580A
Signaling of View ID bit depth within parameter set
JP2016528801A