Dependence on random access point indication in video bitstreams
By modifying the semantics of the DRAP Indicator SEI message and introducing the Type 2DRAP Indicator SEI message, the complexity of decoding random access points in multi-layer bitstreams is solved, achieving more efficient decoding and streaming.
Patent Information
- Application Number
- CN202111142837.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-29
- Filing Date
- 2021-09-28
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Existing video encoding and decoding technologies have difficulties in effectively parsing DRAP indication SEI messages in multi-layer bitstreams. This leads to the decoder needing to parse and deduce a large amount of information to determine the type of random access point and the reference image, increasing system complexity and latency.
By modifying the semantics of the DRAP Indicator SEI message to make it applicable to multi-layer bitstreams, and introducing the Type 2 DRAP Indicator SEI message to explicitly define the random access point identifier and reference image list, the decoder can correctly decode images that depend on random access points and their subsequent images without decoding associated IRAP images.
It simplifies the decoding process of multi-layer bitstreams, reduces system complexity and latency, and improves decoding efficiency and streaming flexibility.
Smart Images

Figure CN114339245B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is made to timely claim priority to and benefit of U.S. Provisional Patent Application No. 63 / 084,953, filed September 29, 2020, under the applicable patent laws and / or rules conforming to the Paris Convention. The entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this application for all purposes made in accordance with the law. TECHNICAL FIELD
[0003] This patent document relates to digital video coding techniques, including video encoding, transcoding, or decoding. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, it is expected that the bandwidth demand for digital video usage will continue to grow. SUMMARY
[0005] This document discloses techniques that can be used by video encoders and decoders to process coded representations of video or images according to a file format.
[0006] In one example aspect, a method of processing visual media data is disclosed. The method includes performing a conversion between the visual media data and a bitstream of the visual media data including multiple layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in a decoding order and an output order without having to decode other pictures in the layer other than an intra random access point (IRAP) picture associated with the DRAP picture.
[0007] In another example aspect, another method of processing visual media data is disclosed. The method includes performing a conversion between the visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies whether and how a second type of supplemental enhancement information (SEI) message is included in the bitstream different from a first type of SEI message, and wherein the first type of SEI message and the second type of SEI message indicate a first type of dependent random access point (DRAP) picture and a second type of DRAP picture, respectively.
[0008] In another example aspect, another method of processing visual media data is disclosed. The method includes performing a conversion between visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies that a supplemental enhancement information (SEI) message referring to a dependent random access point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating a number of intra random access point (IRAP) pictures or DRAP pictures within a same coded layer video sequence (CLVS) as the DRAP picture.
[0009] In yet another example aspect, a video processing apparatus is disclosed. The video processing apparatus includes a processor configured to implement the above-described method.
[0010] In yet another example aspect, a method of storing visual media data into a file comprising one or more bitstreams is disclosed. The method corresponds to the above-described method, and further includes storing the one or more bitstreams to a non-transitory computer-readable recording medium.
[0011] In yet another example aspect, a computer-readable medium storing a bitstream is disclosed. The bitstream is generated according to the above-described method.
[0012] In yet another example aspect, a video processing apparatus storing a bitstream is disclosed, wherein the video processing apparatus is configured to implement the above-described method.
[0013] In yet another example aspect, a computer-readable medium is disclosed, on which a bitstream conforms to a file format generated according to the above-described method.
[0014] These, additional, and other aspects will be described throughout the present document. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 is a block diagram of an example video processing system.
[0016] Figure 2 is a block diagram of a video processing apparatus.
[0017] Figure 3 is a flowchart of an example method of video processing.
[0018] Figure 4 is a block diagram illustrating a video coding system according to some embodiments of the present disclosure.
[0019] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0020] Figure 6is a block diagram illustrating a decoder according to some embodiments of the disclosure.
[0021] Figures 7 to 9 is a flowchart of an example method of processing visual media data based on some embodiments of the disclosed technology. DETAILED DESCRIPTION
[0022] For ease of understanding, section headings are used in this document, and the teachings and embodiments disclosed in each section are not meant to be limited to only that section. Also, the use of H.266 terminology in some descriptions is merely for ease of understanding, and is not meant to limit the scope of the disclosed technology. As such, the technology described herein is applicable to other video codec protocols and designs as well. In this document, editorial changes to text relative to the current draft of the VVC specification are shown by strikeout for deleted text and highlight (including boldface italics) for added text.
[0023] 1. Preliminary Discussion
[0024] This document relates to video coding technology. In particular, this document relates to support for cross-random access point (RAP) references in video coding based on supplemental enhancement information (SEI) messages. These ideas can be applied to any video coding standard or non-standard video codec, such as the recently finalized Versatile Video Coding (VVC), individually or in various combinations.
[0025] 2. Abbreviations
[0026] ACT adaptive color transform
[0027] ALF adaptive loop filter
[0028] AMVR adaptive motion vector resolution
[0029] APS adaptive parameter set
[0030] AU access unit
[0031] AUD access unit delimiter
[0032] AVC advanced video coding (Rec. ITU-T H.264 | ISO / IEC 14496-10)
[0033] B bi-prediction
[0034] BCW bi-prediction with CU-level weights
[0035] BDOF bi-directional optical flow
[0036] BDPCM block-based delta pulse code modulation
[0037] BP buffering period
[0038] CABAC context-based adaptive binary arithmetic coding
[0039] CB coded block
[0040] CBR constant bit rate
[0041] CCALF cross-component adaptive loop filter
[0042] CLVS coded layer video sequence
[0043] CLVSS coded layer video sequence start
[0044] CPB coded picture buffer
[0045] CRA clean random access
[0046] CRC cyclic redundancy check
[0047] CTB coded tree block
[0048] CTU coded tree unit
[0049] CU coding unit
[0050] CVS coded video sequence
[0051] CVSS coded video sequence start
[0052] DPB decoded picture buffer
[0053] DCI decoding capability information
[0054] DRAP dependent random access point
[0055] DU decoding unit
[0056] DUI decoding unit information
[0057] EG exponential Golomb
[0058] EGk k-th order exponential Golomb
[0059] EOB end of bitstream
[0060] EOS end of sequence
[0061] FD filler data
[0062] FIFO first in, first out
[0063] FL fixed length
[0064] GBR green, blue, and red
[0065] GCI general constraint information
[0066] GDR gradual decoding refresh
[0067] GPM geometric partition mode
[0068] HEVC high efficiency video coding (Rec. ITU-T H.265 | ISO / IEC 23008-2)
[0069] HRD hypothetical reference decoder
[0070] HSS hypothetical stream scheduler
[0071] I intra
[0072] IBC intra block copy
[0073] IDR instantaneous decoding refresh
[0074] ILRP inter-layer reference picture
[0075] IRAP intra random access point
[0076] LFNST low-frequency non-separable transform
[0077] LPS least probable symbol
[0078] LSB least significant bit
[0079] LTRP long-term reference picture
[0080] LMCS luma mapping with chroma scaling
[0081] MIP matrix-based intra prediction
[0082] MPS most probable symbol
[0083] MSB most significant bit
[0084] MTS multiple transform selection
[0085] MVP motion vector prediction
[0086] NAL network abstraction layer
[0087] OLS output layer set
[0088] OP operation point
[0089] OPI operation point information
[0090] P prediction
[0091] PH picture header
[0092] POC picture order count
[0093] PPS picture parameter set
[0094] PROF prediction refinement with optical flow
[0095] PT picture timing
[0096] PU picture unit
[0097] QP quantization parameter
[0098] RADL random access decodable leading (picture)
[0099] RAP random access point
[0100] RASL random access skipped leading (picture)
[0101] RBSP raw byte sequence payload
[0102] RGB red, green, and blue
[0103] RPL reference picture list
[0104] SAO sample adaptive offset
[0105] SAR sample aspect ratio
[0106] SEI supplemental enhancement information
[0107] SH slice header
[0108] SLI subpicture level information
[0109] SODB data bit string
[0110] SPS sequence parameter set
[0111] STRP short-term reference picture
[0112] STSA stepping time duration sub-layer access
[0113] TR truncated rice
[0114] TU transform unit
[0115] VBR variable bit rate
[0116] VCL video coding layer
[0117] VPS video parameter set
[0118] VSEI Versatile Supplementary Enhancement Information (Rec. ITU-T H.274 | ISO / IEC 23002-7)
[0119] VUI Video Usability Information
[0120] VVC Versatile Video Coding (Rec. ITU-T H.266 | ISO / IEC 23090-3)
[0121] 3. Introduction to video coding
[0122] 3.1. Video coding standards
[0123] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction plus transform coding is exploited. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially started, the JVET was later renamed as Joint Video Expert Team (JVET). VVC is the new coding standard finalized by the JVET at its 19th meeting, which ended on 1st July 2020, with the goal of 50% bitrate reduction compared to HEVC.
[0124] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcast, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding and viewport-adaptive 360° immersive media.
[0125] 3.2. Picture Order Count (POC) in HEVC and VVC
[0126] In HEVC and VVC, POC is basically used as a picture ID for pictures in many parts of the decoding process, including DPB management, where a part is reference picture management.
[0127] For newly introduced PH in VVC, the information of POC least significant bit (LSB) is signaled in PH, unlike in HEVC where it is signaled in SH, which is used to derive POC values and has the same value for all slices of a picture. VVC also allows signaling of POC most significant bit (MSB) period value in PH to enable derivation of POC values without tracking POC MSB, which relies on POC information of earlier decoded pictures. This allows, for example, mixing of IRAP and non-IRAP within an AU in multi-layer bitstreams. An additional difference between POC signaling in HEVC and VVC is that in HEVC, no POC LSB is signaled for IDR pictures, which showed some drawbacks during the later development of multi-layer extension of HEVC in order to enable mixing of IDR and non-IDR pictures within an AU. Therefore, in VVC, POC LSB information is signaled for every picture, including IDR pictures. The signaling of POC LSB information for IDR pictures also makes it easier to support merging of IDR pictures and non-IDR pictures from different bitstreams into one picture, otherwise handling POC LSB in the merged picture would require some complex design.
[0128] 3.3. Random access and its support in HEVC and VVC
[0129] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, search in local playback and streaming, and stream adaptation in streaming, a bitstream needs to include frequent random access points, which are usually intra-coded pictures, but can also be inter-coded pictures (e.g., in the case of gradual decoding refresh).
[0130] HEVC signals intra random access point (IRAP) pictures through NAL unit types in the NAL unit header. Three types of IRAP pictures are supported, namely instantaneous decoder refresh (IDR), clean random access (CRA), and broken link access (BLA) pictures. IDR pictures constrain the inter prediction structure to not reference any picture before the current group of pictures (GOP), traditionally referred to as a closed-GOP random access point. CRA pictures have less restriction by allowing certain pictures to reference pictures before the current GOP, in the case of random access where all pictures are discarded. CRA pictures are traditionally referred to as an open-GOP random access point. BLA pictures typically result from splicing of two bitstreams or a part thereof at a CRA picture, for example during stream switching. To enable the system to better use IRAP pictures, a total of six different NAL units are defined to signal properties of IRAP pictures that can be used to better match the stream access point types defined in the ISO base media file format (ISOBMFF), which is used to support dynamic adaptive streaming over HTTP (DASH).
[0131] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or another type without associated RADL pictures) and one type of CRA picture. These are essentially the same as in HEVC. VVC does not include the BLA picture type in HEVC for two main reasons: i) the basic functionality of BLA pictures can be achieved by a CRA picture plus an end of sequence NAL unit, whose presence indicates that the following pictures start a new CVS in a single-layer bitstream; ii) during the development of VVC, it was desired to specify fewer NAL unit types than HEVC, which is indicated by the use of five bits instead of six bits for the NAL unit type field in the NAL unit header.
[0132] Another key difference between VVC and HEVC in terms of random access support is that the support of GDR in VVC is in a more normative way. In GDR, the decoding of the bitstream can start from an inter-coded picture, and although not the entire picture region can be correctly decoded at the beginning, the entire picture region will be correct after a number of pictures. AVC and HEVC also support GDR, and the signaling of GDR random access points and recovery points uses the recovery point SEI message. In VVC, a new NAL unit type is specified for indicating GDR pictures, and the recovery point is signaled in the picture header syntax structure. It is allowed that a CVS and a bitstream can start from a GDR picture. This means that it is allowed that the entire bitstream contains only inter-coded pictures, without a single intra-coded picture. The main benefit of specifying GDR support in this way is that it provides a consistent behavior for GDR. GDR enables an encoder to smooth the bit rate of the bitstream by distributing intra-coded slices or blocks in multiple pictures, instead of intra-coding the entire picture, thus significantly reducing the end-to-end delay, which is considered more important today than before, as wireless display, online gaming, drone-based applications become more popular.
[0133] Another GDR related feature in VVC is the virtual boundary signaling. On the pictures between a GDR picture and its recovery point, the boundary between the refreshed region (i.e., the correctly decoded region) and the un-refreshed region can be signaled as a virtual boundary, and when signaled, no in-loop filtering across the boundary will be applied, thus no decoding mismatch of some samples at or near the boundary will occur. This can be useful when the application determines to display the correctly decoded region during the GDR process.
[0134] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.
[0135] 3.4. VUI and SEI messages
[0136] VUI is a syntax structure sent as part of SPS (and possibly in VPS of HEVC as well). The information carried by VUI does not affect the normative decoding process, but the information can be important for correctly rendering the coded video.
[0137] SEI assists processes related to decoding, display, or other purposes. Like VUI, SEI does not affect the normative decoding process either. SEI is carried in SEI messages. Decoder support of SEI messages is optional. However, SEI messages do affect bitstream conformance (e.g., a bitstream is not conforming if the syntax of SEI messages in the bitstream does not conform to the specification), and some SEI messages are required in the HRD specification.
[0138] The VUI syntax structure and most SEI messages used with VVC are not specified in the VVC specification but in the VSEI specification. The SEI information needed for HRD conformance testing is specified in the VVC specification. VVC vl defines 5 SEI messages related to HRD conformance testing and VSEI vl specifies 20 additional SEI messages. The SEI messages carried in the VSEI specification do not directly impact the conforming decoder behavior and have been defined so they can be used in a codec format agnostic way, allowing VSEI to be used in the future with other video coding standards than VVC. The VSEI specification does not refer to VVC syntax element names but to variables whose values are set in the VVC specification.
[0139] In contrast to HEVC, the VVC VUI syntax structure only concerns information related to the correct rendering of the pictures and does not contain any timing information or bitstream restriction indication. In VVC, the VUI is signaled in the SPS which contains a length field before the VUI syntax structure signaling the length of the VUI payload in bytes. This enables the decoder to easily skip the information and more importantly, allows for convenient future VUI syntax extension by directly adding new syntax elements at the end of the VUI syntax structure in a similar way as SEI message syntax extension.
[0140] The VUI syntax structure contains the following information:
[0141] • Whether the content is interlaced or progressive;
[0142] • Whether the content contains frame-packed stereoscopic video or projected omnidirectional video;
[0143] • Sample aspect ratio;
[0144] • Whether the content is suitable for over scan display;
[0145] • Color description including the primary chromaticities, the matrix coefficients and the transfer characteristics, which is particularly important for being able to signal Ultra High Definition (UHD) and High Definition (HD) color spaces as well as High Dynamic Range (HDR);
[0146] • Chromaticity location compared to luminance (clarified signaling for progressive content compared to HEVC).
[0147] When the SPS does not contain any VUI, this information is considered unspecified and must be conveyed by external means or specified by the application if the content of the bitstream is intended to be presented on a display.
[0148] Table 1 lists all SEI messages specified for VVC vl, as well as the specification containing their syntax and semantics. Out of the 20 SEI messages specified in the VSEI specification, many are inherited from HEVC (e.g., padding payload and user data SEI messages). Some SEI messages are essential for the correct handling or presentation of coded video content. This is the case, for example, for the primary display color volume, content light level information, or alternative transfer characteristics SEI messages that are particularly relevant for HDR content. Other examples include equirectangular projection, sphere rotation, region packing, or omnidirectional viewport SEI messages that are relevant for the signaling and handling of 360° video content.
[0149] Table 1: List of SEI messages in VVC vl
[0150]
[0151]
[0152] The new SEI messages specified for VVC vl include the frame-field information SEI message, the sample aspect ratio information SEI message, and the subpicture level information SEI message.
[0153] The frame-field information SEI message contains information indicating how the associated picture should be displayed (e.g., field parity or frame repeat period), the source scan type of the associated picture, and whether the associated picture is a copy of a previous picture. In previous video coding standards, this information can be signaled together with the timing information of the associated picture in the picture timing SEI message. However, it is observed that the frame-field information and the timing information are two different kinds of information, which do not necessarily need to be signaled together. A typical example is to signal the timing information at the system level, but to signal the frame-field information in the bitstream. Therefore, it is decided to remove the frame-field information from the picture timing SEI message and instead signal it in a dedicated SEI message. This change also enables the syntax of the frame-field information to be modified to convey additional and more explicit instructions to the display, such as pairing fields together, or more values for frame repetition.
[0154] The sample aspect ratio SEI message enables different sample aspect ratios to be signaled for different pictures within the same sequence, whereas the corresponding information contained in the VUI applies to the entire sequence. This can lead to different sample aspect ratios for different pictures of the same sequence when the reference picture resampling functionality with scaling factors is used.
[0155] The subpicture level information SEI message provides level information for a subpicture sequence.
[0156] 3.5. Cross-RAP reference
[0157] A video coding method based on Cross-RAP Reference (CRR), also called External Decoded Refresh (EDR), is proposed in JVET-M0360, JVET-N0119, JVET-O0149 and JVET-P0114.
[0158] The basic idea of this video coding method is as follows. Instead of coding random access points as Intra coded IRAP pictures (except the first picture in the bitstream), they are coded using inter prediction to overcome the unavailability of early pictures if random access points are coded as IRAP pictures. The trick is to provide a limited number of early pictures, usually representing different scenes of the video content, by a separate video bitstream, which can be called external device. Such early pictures are called external pictures. Thus, each external picture can be used for inter prediction reference by pictures across random access points. The coding efficiency gain comes from coding random access points as inter predicted pictures and having more available reference pictures for pictures following the EDR pictures in decoding order.
[0159] As described below, the bitstream coded with this video coding method can be used in ISOBMFF and DASH based applications.
[0160] DASH content preparation operation
[0161] 1) The video content is encoded into one or more representations, each having a specific spatial resolution, temporal resolution and quality.
[0162] 2) Each specific representation of the video content is represented by a main stream, which can also be represented by an external stream. The main stream contains coded pictures, which can or can not contain EDR pictures. When at least one EDR picture is included in the main stream, the external stream also exists and contains external pictures. When no EDR picture is included in the main stream, the external stream does not exist.
[0163] 3) Each main stream is carried in a Main Stream Representation (MSR). Each EDR picture of the MSR is the first picture of a Segment.
[0164] 4) Each external stream (if exists) is carried in an External Stream Representation (ESR).
[0165] 5) For each segment in an MSR that starts with an EDR picture, there is a segment in the corresponding ESR with the same segment start time derived from the MPD, carrying the external pictures needed to decode that EDR picture and the subsequent pictures in decoding order in the bitstream carried in the MSR.
[0166] 6) The MSRs for the same video content are included in one Adaptation Set (AS). The ESRs for the same video content are included in one AS.
[0167] DASH streaming operation
[0168] 1) The client obtains the MPD for a DASH media presentation, parses the MPD, selects an MSR, and determines the start presentation time at which to consume the content.
[0169] 2) The client requests segments of the MSR, starting with the segment that includes the picture with the presentation time equal to (or sufficiently close to) the start presentation time.
[0170] a. If the first picture in the start segment is an EDR picture, the corresponding segment in the associated ESR (with the same segment start time derived from the MPD) is also requested, preferably before the MSR segment is requested. Otherwise, no segment of the associated ESR is requested.
[0171] 3) When switching to a different MSR, the client requests segments of the switched-to MSR starting with the first segment whose segment start time is greater than the start time of the last requested segment from the switched-from MSR.
[0172] a. If the first picture in the start segment of the switched-to MSR is an EDR picture, the corresponding segment in the associated ESR is also requested, preferably before the MSR segment is requested. Otherwise, no segment of the associated ESR is requested.
[0173] 4) When operating continuously on the same MSR (after decoding the start segment following a seek or stream switch operation), no segment of the associated ESR needs to be requested, including when requesting any segment that starts with an EDR picture.
[0174] 3.6. DRAP indication SEI message
[0175] The VSEI specification includes a DRAP indication SEI message, as follows:
[0176]
[0177] A picture associated with a dependent random access point (DRAP) indication SEI message is referred to as a DRAP picture.
[0178] The presence of the DRAP indication SEI message indicates that the constraints on picture order and picture referencing specified in this clause apply. These constraints can enable decoders to correctly decode the DRAP picture and pictures following it in decoding order and output order without the need to decode any pictures other than the associated IRAP picture of the DRAP picture.
[0179] The presence of the DRAP indication SEI message indicates that the following constraints all apply:
[0180] - the DRAP picture is a trailing picture;
[0181] - the temporal sub-layer identifier of the DRAP picture is equal to 0;
[0182] - the DRAP picture does not include any picture in active entries of its reference picture lists other than the associated IRAP picture of the DRAP picture;
[0183] - any picture following the DRAP picture in decoding order and output order does not include in active entries of its reference picture lists any picture preceding the DRAP picture in decoding order or output order, except the associated IRAP picture of the DRAP picture.
[0184] 4. Technical problems solved by the disclosed technical solutions
[0185] The function of the DRAP indication SEI message can be considered as a subset of the CRR method. For simplicity, pictures associated with the DRAP indication SEI message are referred to as type 1 DRAP pictures.
[0186] From an encoding perspective, while the CRR method proposed in JVET-P0114 or earlier JVET contributions is not adopted by VVC, an encoder can still encode a video bitstream in such a way that some pictures only rely on the associated IRAP picture for inter prediction reference (such as type 1 DRAP pictures indicated by the DRAP SEI message), and some other pictures (for example, referred to as type 2 DRAP pictures) only rely on some pictures in a picture set consisting of the associated IRAP picture and some other (type 1 or type 2) DRAP pictures.
[0187] However, given a VVC bitstream, it is not known whether such Type 2 DRAP pictures exist in the bitstream. Moreover, even when it is known that such Type 2 DRAP pictures exist in the bitstream, in order to compose media files from ISOBMFF and DASH media representations based on such VVC bitstreams to enable CRR or EDR streaming operations, the file and DASH media representation editors would need to parse and derive a large amount of information, including POC values and valid entries in the reference picture lists, to determine whether a particular picture is a Type 2 DRAP picture, and if so, which earlier IRAP or DRAP picture is needed to randomly access from that particular picture, so that appropriate sets of pictures can be included in separate, time-synchronized file tracks and DASH representations.
[0188] There is also a problem that the semantics of the DRAP indication SEI message only apply to single-layer bitstreams.
[0189] 5. List of solutions
[0190] To address the above problems, among others, methods summarized as follows are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow way. Moreover, these items can be applied individually or in any combination.
[0191] 1) In one example, the semantics of the DRAP indication SEI message are changed so that this SEI message can apply to multi-layer bitstreams, i.e., the semantics enable the decoder to correctly decode the DRAP picture (i.e., the picture associated with the DRAP indication SEI message) and the pictures in the same layer and following the DRAP in decoding order and output order, without the need to decode any other picture in the same layer other than the associated IRAP picture of the DRAP picture.
[0192] a. For example, it is required that the DRAP picture does not include in its reference picture list valid entries any picture in the same layer other than the associated IRAP picture of the DRAP picture.
[0193] b. In one example, it is required that any picture in the same layer and following the DRAP picture in decoding order and output order does not include in its reference picture list valid entries any picture in the same layer and preceding the DRAP picture in decoding order or output order, other than the associated IRAP picture of the DRAP picture.
[0194] 2) In one example, the RAP picture ID for the DRAP picture is signaled in the DRAP indication SEI message to specify the identifier of the RAP picture, which can be an IRAP picture or a DRAP picture.
[0195] a. In one example, a presence flag indicating whether a RAP picture ID is present in the DRAP indication is signaled, and when the flag is equal to a specific value, e.g., 1, the RAP picture ID is signaled in the DRAP indication SEI message, and when the flag is equal to another value, e.g., 0, the RAP picture ID is not signaled in the DRAP indication SEI message.
[0196] 3) In one example, a DRAP picture associated with a DRAP indication SEI message allows referring to an associated IRAP picture or a preceding picture in decoding order, which is a GDR picture with ph_recovery_poc_cnt equal to 0, for inter prediction reference.
[0197] 4) In one example, a new SEI message, e.g., named Type 2 DRAP indication SEI message, and each picture associated with the new SEI message is referred to as a special type of picture, e.g., Type 2 DRAP picture.
[0198] 5) In one example, it is specified that Type 1 DRAP pictures (associated with a DRAP indication SEI message) and Type 2 DRAP pictures (associated with a Type 2 DRAP indication SEI message) are collectively referred to as DRAP pictures.
[0199] 6) In one example, a Type 2 DRAP indication SEI message includes a RAP picture ID, e.g., denoted as RapPicId, to specify an identifier of a RAP picture, which can be an IRAP picture or a DRAP picture, and a syntax element, e.g., t2drap_num_ref_rap_pics_minus1, indicating a number of IRAP or DRAP pictures within the same CLVS as the Type 2 DRAP picture and can be included in the active entries of the reference picture lists of the Type 2 DRAP picture.
[0200] a. In one example, the syntax element, e.g., t2drap_num_ref_rap_pics_minus1, indicating the number is coded as u(3) using 3 bits.
[0201] b. Optionally, the syntax element, e.g., t2drap_num_ref_rap_pics_minus1, indicating the number is coded as ue(v).
[0202] 7) In one example, for the RAP picture ID of a DRAP picture, in the DRAP indication SEI message or the Type 2 DRAP indication SEI message, one or more of the following methods apply:
[0203] a. In one example, the syntax element signaling the RAP picture ID is coded using 16 bits as u(16).
[0204] i. Alternatively, the syntax element signaling the RAP picture ID is coded using ue(v).
[0205] b. In one example, instead of signaling the RAP picture ID in the DRAP indication SEI message, the POC value of the DRAP picture is signaled, e.g., using se(v) or i(32).
[0206] i. Optionally, the POC delta relative to the POC value of the associated IRAP picture is signaled, e.g., using ue(v) or u(16).
[0207] 8) In one example, it is specified that each IRAP or DRAP picture that is an IRAP or a DRAP is associated with a RAP picture ID, RapPicId.
[0208] a. In one example, it is specified that the value of RapPicId for an IRAP picture is inferred to be equal to 0.
[0209] b. In one example, it is specified that the values of RapPicId for any two IRAP or DRAP pictures within a CLVS shall be different.
[0210] c. Furthermore, in one example, the values of RapPicId for IRAP and DRAP pictures within a CLVS will increase with the increasing decoding order of the IRAP or DRAP picture.
[0211] d. Furthermore, in one example, within the same CLVS, the RapPicId of a DRAP picture shall be greater than the RapPicId of an IRAP or DRAP picture that precedes it in decoding order by 1.
[0212] 9) In one example, the type 2 DRAP indication SEI message further includes a list of RAP picture IDs, one list for each of the IRAP or DRAP pictures that are within the same CLVS as the type 2 DRAP picture and can be included in the active entries of the reference picture lists of the type 2 DRAP picture.
[0213] a. In one example, each of the RAP picture IDs in the list is coded as the same as the RAP picture ID of the DRAP picture associated with the type 2 DRAP indication SEI message.
[0214] b. Optionally, the values of the list of RAP picture IDs are required to be increasing in increasing order of the values of the list indices i, and the value of RapPicId for the i-th DRAP pic is coded using ue(v) and the increment between 1) the value of RapPicId for the (i-1)-th DRAP or IRAP pic (when i is greater than 0) or 2) 0 (when i is equal to 0).
[0215] c. Optionally, each of the list of RAP picture IDs is coded to represent the POC value of the RAP picture, e.g., coded as se(v) or i(32).
[0216] d. Optionally, each of the list of RAP picture IDs is coded to represent the POC increment relative to the POC value of the associated IRAP picture, e.g., signaled using ue(v), u(16).
[0217] e. Optionally, each of the list of RAP picture IDs is coded to represent the POC increment between the POC value of the current picture and 1) the POC value of the (i-1)-th DRAP or IRAP pic (when i is greater than 0) or 2) the POC value of the IRAP picture (when i is equal to 0), e.g., using ue(v) or u(16).
[0218] f. Optionally, in addition, it is required that for any two values of list index values i and j, for the list of RAP picture IDs, the i-th IRAP or DRAP picture shall precede the j-th IRAP or DRAP picture in decoding order when i is less than j.
[0219] 6. Embodiments
[0220] Below are some example embodiments of the inventive aspects summarized in Section 5 above, which can be applied to the VSEI specification. The changed text is based on the latest VSEI text in JVET-S2007-v7. Most of the relevant parts that have been added or modified are highlighted in bold italics, and some deleted parts are marked by double brackets (e.g., [[a]] means the letter “a” is deleted). Some other changes that can be of editorial nature are not highlighted.
[0221] 6.1. First embodiment
[0222] This embodiment is a change to the existing DRAP indication SEI message.
[0223] 6.1.1. Dependent random access point indication SEI message syntax
[0224]
[0225] 6.1.2. Dependent random access point indication SEI message semantics
[0226] A picture associated with a random access point (RAP) indication SEI message is referred to as a type 1 RAP picture.
[0227] A type 1 RAP picture and a type 2 RAP picture (associated with a type 2 RAP indication SEI message) are collectively referred to as a RAP picture.
[0228] The presence of a RAP indication SEI message indicates that the constraints specified in this subclause on picture order and picture referencing apply. These constraints can enable a decoder to correctly decode a type 1 RAP picture and pictures in the same layer and following the type 1 RAP picture in decoding order and output order without the need to decode any other picture in the same layer except the associated IRAP picture of the type 1 RAP picture.
[0229] The constraints indicated by the presence of a RAP indication SEI message are as follows, all of which apply:
[0230] - The type 1 RAP picture is a trailing picture.
[0231] - The type 1 RAP picture has a temporal sub-layer identifier equal to 0.
[0232] - The type 1 RAP picture does not include, in the same layer, any picture in the active entries of its reference picture lists except the associated IRAP picture of the type 1 RAP picture.
[0233] - Any picture in the same layer and following the type 1 RAP picture in decoding order and output order does not include, in the active entries of its reference picture lists, any picture in the same layer and preceding the type 1 RAP picture in decoding order or output order except the associated IRAP picture of the type 1 RAP picture.
[0234] drap_rap_id_in_clvs specifies the RAP picture ID of the type 1 Drap picture, denoted as RapPicId.
[0235] Each IRAP picture or DRAP picture is associated with a RapPicId as IRAP or DRAP, respectively. The value of RapPicId for an IRAP picture is inferred to be equal to 0. The values of RapPicId for any two IRAP pictures or DRAP pictures within a CLVS shall be different.
[0236] 6.2. Second embodiment
[0237] This embodiment is for the new type 2 RAP indication SEI message.
[0238] 6.2.1. Type 2 DRAP indication SEI message syntax
[0239]
[0240]
[0241] 6.2.2. Type 2 DRAP indication SEI message semantics
[0242] A picture associated with a type 2 DRAP indication SEI message is referred to as a type 2 DRAP picture.
[0243] Type 1 DRAP pictures (associated with a DRAP indication SEI message) and type 2 DRAP pictures are collectively referred to as DRAP pictures.
[0244] The presence of a type 2 DRAP indication SEI message indicates that the constraints specified in this subclause on picture order and picture referencing apply. These constraints can enable a decoder to correctly decode a type 2 DRAP picture and pictures in the same layer and following the type 2 DRAP picture in decoding order and output order without the need to decode any other picture in the same layer, except for the picture list referenceablePictures, which consists of the list of IRAP pictures or DRAP pictures in decoding order within the same CLVS that are identified by the t2drap_ref_rap_id[ i ] syntax elements.
[0245] The constraints indicated by the presence of a type 2 DRAP indication SEI message are as follows, all of which apply:
[0246] - the type 2 DRAP picture is a trailing picture;
[0247] - the type 2 DRAP picture has a temporal sub-layer identifier equal to 0;
[0248] - the type 2 DRAP picture does not include in the valid entries of its reference picture lists any picture in the same layer, except for referenceablePictures;
[0249] - any picture in the same layer and following the type 2 DRAP picture in decoding order and output order does not include in the valid entries of its reference picture lists any picture in the same layer and preceding the type 2 DRAP picture in decoding order or output order, except for referenceablePictures;
[0250] - Any picture in the list referenceablePictures does not include in the active entries of its reference picture list any picture that is in the same layer and is not a picture that is at an earlier position in the list referenceablePictures.
[0251] NOTE - Thus, the first picture in referenceablePictures, even if it is a DRAP picture and not an IRAP picture, does not include in the active entries of its reference picture list any picture that is in the same layer.
[0252] t2drap_rap_id_in_clvs specifies the RAP picture identifier of the type 2 DRAP picture, denoted as RapPicId.
[0253] Each IRAP picture or DRAP picture, as IRAP or DRAP, is associated with a RapPicId. The value of RapPicId for an IRAP picture is inferred to be equal to 0. The values of RapPicId for any two IRAP pictures or DRAP pictures within a CLVS shall be different.
[0254] In bitstreams conforming to this version of this Specification, t2drap_reserved_zero_13bits shall be equal to 0. Other values of t2drap_reserved_zero_13bits are reserved for future use by ITU-T | ISO / IEC. Decoders shall ignore the value of t2drap_reserved_zero_13bits.
[0255] t2drap_num_ref_rap_pics_minus1 plus 1 specifies the number of IRAP pictures or DRAP pictures that are within the same CLVS as the type 2 DRAP picture and that can be included in the active entries of the reference picture list of the type 2 DRAP picture.
[0256] t2drap_ref_rap_id[ i ] specifies the RapPicId of the i-th IRAP picture or DRAP picture that is within the same CLVS as the type 2 DRAP picture and that can be included in the active entries of the reference picture list of the type 2 DRAP picture.
[0257] Figure 1is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 1900. The system 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be in a compressed or encoded format. The input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0258] The system 1900 can include a codec component 1904 that can implement various coding or encoding methods described in this document. The codec component 1904 can reduce the average bitrate of video from the input 1902 to an output of the codec component 1904 to produce a coded representation of the video. The coding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 can be stored, or transmitted via a communication connection as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 can be used by component 1908 to generate pixel values or a displayable video to a display interface 1910. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0259] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI), or Display port, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0260] Figure 2is a block diagram of a video processing device 3600. The device 3600 can be used to implement one or more of the methods described herein. The device 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in the present document. The memory(ies) 604 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement, in hardware circuitry, some of the techniques described in the present document. In some embodiments, the video processing hardware 3606 can be included at least in part in the processor 3602 (e.g., a graphics co-processor).
[0261] Figure 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0262] As shown in Figure 4 , the video coding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, where the source device 110 can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the destination device 120 can be referred to as a video decoding device.
[0263] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0264] The video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 by the I / O interface 116 through the network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by the destination device 120.
[0265] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0266] The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the destination device 120, or can be external to the destination device 120 configured to interface with the external display device.
[0267] The video encoder 114 and the video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[0268] Figure 5 is a block diagram illustrating an example of a video encoder 200 that can be Figure 4 the video encoder 114 in the system 100 shown.
[0269] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 5 examples, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0270] The functional components of the video encoder 200 can include a partitioning unit 201, a prediction unit 202 (which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0271] In other examples, the video encoder 200 can include more, less, or different functional components. In examples, the prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0272] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately for illustrative purposes. Figure 5 in examples.
[0273] The partitioning unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0274] The mode selection unit 203 can select one of the coding modes (e.g., intra or inter) based on the error results and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra and inter prediction modes (CIIP) in which the prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0275] To perform inter prediction for a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 to the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0276] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0277] In some examples, the motion estimation unit 204 can perform single prediction for a current video block, and the motion estimation unit 204 can search for a reference picture of list 0 or list 1 for a reference video block of the current video block. The motion estimation unit 204 can then generate a reference index indicating the reference picture in list 0 or list 1, which contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0278] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0, and can also search for another reference video block for the current video block in list 1. The motion estimation unit 204 can then generate a reference index that indicates the reference picture in list 0 and list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the reference video block and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0279] In some examples, the motion estimation unit 204 can output a full set of motion information for the current video for use in decoding processing at the decoder.
[0280] In some examples, the motion estimation unit 204 can not output a full set of motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0281] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.
[0282] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0283] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0284] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0285] Residual generation unit 207 can generate residual data for a current video block by subtracting (e.g., indicated by the minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0286] In other examples, such as in a skip mode, there can be no residual data for the current video block for the current video block, and residual generation unit 207 can not perform the subtraction operation.
[0287] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0288] After transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0289] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0290] After reconstruction unit 212 reconstructs the video block, loop filtering operations can be performed to reduce video block artifacts in the video block.
[0291] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0292] Figure 6 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 4 the video decoder 114 in the system 100 shown.
[0293] The video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 In examples, the video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0294] In Figure 6 In the example of FIG. 3, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. Video decoder 300 may, in some examples, perform a decoding process generally reciprocal to the encoding process described with respect to video encoder 200. Figure 5
[0295] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data and, from the entropy decoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes.
[0296] Motion compensation unit 302 can generate a motion compensated block, and can perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used at sub-pixel precision can be included in a syntax element.
[0297] Motion compensation unit 302 can use an interpolation filter as used by video encoder 200 during encoding of the video block to calculate the interpolation of sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.
[0298] Motion compensation unit 302 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence.
[0299] Intra-prediction unit 303 can use intra-prediction modes, e.g., received in the bitstream, to form a prediction block from spatially neighboring blocks. Inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transformation unit 303 applies an inverse transform.
[0300] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates the decoded video for presentation on the display device.
[0301] The following is a list of preferred solutions for some embodiments.
[0302] The first set of solutions is provided below, which illustrates example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0303] 1. A video processing method (e.g., Figure 3 The method 700 described herein includes: performing (702) a conversion between a video comprising multiple layers and a codec representation of the video, wherein the codec representation is organized according to a format rule; wherein the format rule specifies that supplementary enhancement information (SEI) is included in the codec representation, wherein the SEI information carries information sufficient to enable the decoder to decode dependent random access point (DRAP) pictures and / or decode pictures in a layer in the order of decoding and output, without needing to decode other pictures in that layer, except for intra-frame random access pictures (IRAP) of DRAP pictures.
[0304] 2. According to the method of Solution 1, DRAP images are excluded from the list of reference images for any image in that layer, except for IRAP images.
[0305] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).
[0306] 3. A video processing method, comprising: performing a conversion between a multi-layered video and a codec representation of the video, wherein the codec representation is organized according to a format rule; wherein the format rule specifies that a Supplemental Enhancement Information (SEI) message is included in a codec representation of a Random Access Point (RAP) image, wherein the SEI message includes an identifier of the RAP image.
[0307] 4. According to the method of Solution 3, where RAP is an intra-frame random access image.
[0308] 5. Based on the approach of Solution 3, where RAP is based on Random Access Image (DRAP).
[0309] The following solutions illustrate preferred example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0310] 6. The method according to solution 5, wherein the DRAP picture is allowed to refer to an associated intra random access picture or a previous picture in decoding order that is a gradual decoding refresh picture.
[0311] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., items 4-6).
[0312] 7. A video processing method comprising performing a conversion between a video comprising multiple layers and a coded representation of the video, wherein the coded representation is organized according to a format rule; wherein the format rule specifies whether and how a type-2 supplemental enhancement information (SEI) message referring to a dependent random access picture (DRAP) is included in the coded representation.
[0313] 8. The method according to solution 7, wherein the format rule specifies that the type-2 SEI message and each picture associated with the message are treated as a special type of picture.
[0314] 9. The method according to solution 7, wherein the format rule specifies that the type-2 SEI message includes an identifier of a random access picture (RAP) referred to as a type-2 RAP picture and a syntax element indicating a number of pictures in the same coded video layer as the random access picture, such that the pictures are included in a valid reference picture list of the type-2 RAP picture.
[0315] 10. The method according to any of solutions 1-9, wherein the conversion comprises generating the coded representation from the video.
[0316] 11. The method according to any of solutions 1-9, wherein the conversion comprises decoding the coded representation to generate the video.
[0317] 12. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 11.
[0318] 13. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 11.
[0319] 14. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement a method recited in any of solutions 1 to 11.
[0320] 15. A computer readable medium storing a coded representation generated according to any of solutions 1 to 11.
[0321] 16. The method, apparatus or system described in the present document.
[0322] The second set of preferred solutions provides example embodiments of the techniques discussed in the previous section (e.g., items 1, 1.a, 1.b, 2, 2.a, 3).
[0323] 1. A method of processing visual media data (e.g., method 710 as shown in Figure 7 FIG. 7), comprising performing (712) a conversion between a visual media data and a bitstream of the visual media data comprising multiple layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in decoding order and output order without having to decode other pictures in the layer except for an intra random access point (IRAP) picture associated with the DRAP picture.
[0324] 2. The method according to solution 1, wherein the DRAP picture excludes pictures in the layer from being valid entries in a reference picture list of the DRAP picture except for the IRAP picture.
[0325] 3. The method according to solution 1, wherein a first picture included in the layer and following the DRAP picture in decoding order and output order excludes a second picture included in the layer and preceding the DRAP picture in decoding order and output order from being valid entries in a reference picture list of the first picture except for the IRAP picture.
[0326] 4. The method according to solution 1, wherein the format rule further specifies that the SEI message includes an identifier of a random access point (RAP) picture.
[0327] 5. The method according to solution 4, wherein the RAP picture is an IRAP picture or a DRAP picture.
[0328] 6. The method according to solution 4, wherein the format rule further specifies that a presence flag indicating a presence of the identifier of the RAP picture in the SEI message is included in the bitstream.
[0329] 7. The method according to solution 6, wherein the presence flag having a value equal to a first value indicates that the identifier of the RAP picture is present in the SEI message.
[0330] 8. The method according to solution 6, wherein the presence flag having a value equal to a second value indicates that the identifier of the RAP picture is omitted from the SEI message.
[0331] 9. The method of solution 1, wherein the allowed DRAP picture refers to an IRAP picture or a previous picture in decoding order that is a gradual decoding refresh (GDR) picture whose decoded picture output timing is equal to 0.
[0332] 10. The method of any of solutions 1 to 9, wherein the bitstream is a general video coding bitstream.
[0333] 11. The method of any of solutions 1 to 10, wherein performing the conversion comprises generating the bitstream from the visual media data.
[0334] 12. The method of any of solutions 1 to 10, wherein performing the conversion comprises reconstructing the visual media data from the bitstream.
[0335] 13. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a conversion between the visual media data and a bitstream of the visual media data comprising multiple layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in decoding order and output order without having to decode other pictures in the layer except for an intra random access point (IRAP) picture associated with the DRAP picture.
[0336] 14. The apparatus of solution 13, wherein the DRAP picture excludes pictures in the layer from being valid entries in a reference picture list of the DRAP picture except for the IRAP picture.
[0337] 15. The apparatus of solution 13, wherein a first picture included in the layer and following the DRAP picture in decoding order and output order excludes a second picture included in the layer and preceding the DRAP picture in decoding order and output order from being valid entries in a reference picture list of the first picture except for the IRAP picture.
[0338] 16. The apparatus of solution 13, wherein the bitstream is a general video coding bitstream.
[0339] 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a conversion between a visual media data and a bitstream comprising the visual media data according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in decoding order and output order without having to decode other pictures in the layer except for an intra random access point (IRAP) picture associated with the DRAP picture.
[0340] 18. The non-transitory computer-readable storage medium of solution 17, wherein the bitstream is a Versatile Video Coding bitstream.
[0341] 19. A non-transitory computer-readable storage medium storing a bitstream of visual media data generated by a method performed by a visual media data processing apparatus, wherein the method comprises determining that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in decoding order and output order without having to decode other pictures in the layer except for an intra random access point (IRAP) picture associated with the DRAP picture; and generating the bitstream based on the determination.
[0342] 20. The non-transitory computer-readable storage medium of solution 19, wherein the bitstream is a Versatile Video Coding bitstream.
[0343] 21. A visual media data processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 12.
[0344] 22. A method of storing a bitstream of visual media data comprising a method recited in any one of solutions 1 to 12, further comprising storing the bitstream to a non-transitory computer-readable storage medium.
[0345] 23. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of solutions 1 to 12.
[0346] 24. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0347] 25. A visual media data processing apparatus that stores a bitstream, wherein the visual media data processing apparatus is configured to implement a method recited in any one or more of solutions 1 to 12.
[0348] 26. A computer readable medium having a bitstream thereon complying with a format rule recited in any one of solutions 1 to 12.
[0349] The third set of solutions provides preferred example implementations of the techniques discussed in the previous section (e.g., item 4 to item 8).
[0350] 1. A method of processing visual media data (e.g., method 800 as shown in Figure 8 FIG. 8), comprising performing 802 a conversion between visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies whether and how a second type of supplemental enhancement information (SEI) message is included in the bitstream differently from a first type of SEI message, and wherein the first type of SEI message and the second type of SEI message are indicative of a first type of dependent random access point (DRAP) picture and a second type of DRAP picture, respectively.
[0351] 2. The method of solution 1, wherein the format rule further specifies that the second type of SEI message includes a random access point (RAP) picture identifier.
[0352] 3. The method of solution 1, wherein a random access point (RAP) picture identifier is included in the bitstream for the first type of DRAP picture or the second type of DRAP picture.
[0353] 4. The method of solution 3, wherein the RAP picture identifier is coded as u(16), u(16) being an unsigned integer using 16 bits, or the RAP picture identifier is coded as ue(v), ue(v) being an unsigned integer using an exponential Golomb code.
[0354] 5. The method of solution 1, wherein the format rule further specifies that the first type of SEI message or the second type of SEI message includes information about a picture order count (POC) value of the first type of DRAP picture or the second type of DRAP picture.
[0355] 6. The method of solution 1, wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier.
[0356] 7. The method according to solution 6, wherein the format rule further specifies that the value of the RAP picture identifier of an IRAP picture is inferred to be equal to 0.
[0357] 8. The method according to solution 6, wherein the format rule further specifies that the values of the RAP picture identifiers of any two IRAP or DRAP pictures within a coded layer video sequence (CLVS) are different from each other.
[0358] 9. The method according to solution 6, wherein the format rule further specifies that the values of the RAP picture identifiers of IRAP or DRAP pictures within a coded layer video sequence (CLVS) are increasing in the decoding order of the IRAP or DRAP pictures.
[0359] 10. The method according to solution 6, wherein the format rule further specifies that the value of the RAP picture identifier of a DRAP picture is one greater than the value of the previous IRAP or DRAP picture in the decoding order within a coded layer video sequence (CLVS).
[0360] 11. The method according to any of solutions 1 to 10, wherein performing the conversion comprises generating the bitstream from the visual media data.
[0361] 12. The method according to any of solutions 1 to 10, wherein performing the conversion comprises reconstructing the visual media data from the bitstream.
[0362] 13. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a conversion between visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies whether and how a second type of supplemental enhancement information (SEI) message is included in the bitstream differently from a first type of SEI message, and wherein the first and second types of SEI messages indicate first and second types of dependent random access point (DRAP) pictures, respectively.
[0363] 14. The apparatus according to solution 13, wherein the format rule further specifies that the second type of SEI message includes a random access point (RAP) picture identifier.
[0364] 15. The apparatus according to solution 13, wherein, for the first type of DRAP picture or the second type of DRAP picture, a random access point (RAP) picture identifier is included in the bitstream, the RAP picture identifier is coded as u(16), u(16) being an unsigned integer using 16 bits, or the RAP picture identifier is coded as ue(v), ue(v) being an unsigned integer using an exponential Golomb code.
[0365] 16. The apparatus according to solution 13, wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier, and the format rule further specifies that a value of the RAP picture identifier for the IRAP picture is inferred to be equal to 0.
[0366] 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a conversion between visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies whether and how a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream, and wherein the first type of SEI message and the second type of SEI message are indicative of a first type of dependent random access point (DRAP) picture and a second type of DRAP picture, respectively.
[0367] 18. The non-transitory computer-readable storage medium according to solution 17, wherein the format rule further specifies that the second type of SEI message includes a random access point (RAP) picture identifier, and wherein, for the first type of DRAP picture or the second type of DRAP picture, a random access point (RAP) picture identifier is included in the bitstream, the RAP picture identifier is coded as u(16), which is an unsigned integer using 16 bits, or as ue(v), which is an unsigned integer using an exponential Golomb code, and wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier, and the format rule further specifies that a value of the RAP picture identifier for the IRAP picture is inferred to be equal to 0.
[0368] 19. A non-transitory computer-readable storage medium storing a bitstream of visual media data generated by a method performed by a visual media data processing apparatus, wherein the method comprises determining whether and how a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream; and generating the bitstream based on the determination.
[0369] 20. The non-transitory computer-readable storage medium according to solution 19, wherein the format rule further specifies that the second type of SEI message includes a random access point (RAP) picture identifier, and wherein, for the first type of DRAP picture or the second type of DRAP picture, the random access point (RAP) picture identifier is included in the bitstream, the RAP picture identifier is coded as u(16), which is an unsigned integer using 16 bits, or ue(v), which is an unsigned integer using an exponential Golomb code, and wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier, and the format rule further specifies that the value of the RAP picture identifier for the IRAP picture is inferred to be equal to 0.
[0370] 21. A visual media data processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 12.
[0371] 22. A method of storing a bitstream of visual media data, comprising a method recited in any one of solutions 1 to 12, further comprising storing the bitstream to a non-transitory computer-readable storage medium.
[0372] 23. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of solutions 1 to 12.
[0373] 24. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0374] 25. A visual media data processing apparatus storing a bitstream, wherein the video processing apparatus is configured to implement a method recited in any one or more of solutions 1 to 12.
[0375] 26. A computer-readable medium having a bitstream thereon conforming to a format rule recited in any one of solutions 1 to 12.
[0376] The fourth set of solutions provides preferred example implementations of the techniques discussed in the previous section (e.g., item 6 and item 9).
[0377] 1. A method of processing visual media data (e.g., as in Figure 9The illustrated method 900) comprises performing 902 a conversion between a visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies that a supplemental enhancement information (SEI) message referring to a dependent random access point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating a number of intra random access point (IRAP) pictures or dependent random access point (DRAP) pictures within a same coded layer video sequence (CLVS) as the DRAP picture.
[0378] 2. The method according to solution 1, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of the DRAP picture.
[0379] 3. The method according to solution 1, wherein the syntax element is coded as u(3), u(3) being an unsigned integer using 3 bits, or the syntax element is coded as ue(v), ue(v) being an unsigned integer using an exponential Golomb code.
[0380] 4. The method according to solution 1, wherein the format rule further specifies that the SEI message further includes a list of random access point (RAP) picture identifiers of IRAP pictures and DRAP pictures within the same coded layer video sequence (CLVS) as the DRAP picture.
[0381] 5. The method according to solution 4, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of the DRAP picture.
[0382] 6. The method according to solution 4, wherein each of the list of RAP picture identifiers is coded to be the same as a RAP picture identifier of the DRAP picture associated with the SEI message.
[0383] 7. The method according to solution 4, wherein the identifiers in the list have values corresponding to the i-th RAP picture, i being equal to or larger than 0, and wherein the values of the RAP picture identifiers are increasing in an increasing order of the values i.
[0384] 8. The method according to solution 7, wherein each identifier in the list is coded using a value of the i-th DRAP picture identifier and a ue(v) of 1) an increment between a value of the (i-1)-th DRAP picture or IRAP picture identifier, where i is larger than 0, or 2) an increment between 0, where i is equal to 0.
[0385] 9. The method according to solution 4, wherein each identifier in the list is coded to represent a picture order count (POC) value of the RAP picture.
[0386] 10. The method of solution 4, wherein each identifier in the list is coded to represent POC delta information relative to a picture order count (POC) value of an IRAP picture associated with the SEI message.
[0387] 11. The method of solution 4, wherein each identifier in the list is coded to represent POC delta information between a picture order count (POC) value of a current picture and 1) a POC value of an (i-1)th DRAP picture or IRAP picture, where i is greater than 0, or 2) a POC value of an IRAP picture associated with the SEI message.
[0388] 12. The method of solution 4, wherein the list includes identifiers corresponding to an i-th RAP picture and a j-th RAP picture, where i is less than j, and where the i-th RAP picture precedes the j-th RAP picture in decoding order.
[0389] 13. The method of any of solutions 1 to 12, wherein performing the conversion comprises generating a bitstream from the visual media data.
[0390] 14. The method of any of solutions 1 to 12, wherein performing the conversion comprises reconstructing the visual media data from the bitstream.
[0391] 15. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a conversion between visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies that a supplemental enhancement information (SEI) message referring to a dependent random access point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating a number of intra random access point (IRAP) pictures or dependent random access point (DRAP) pictures that are within a same coded layer video sequence (CLVS) as the DRAP picture.
[0392] 16. The apparatus according to solution 15, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of a DRAP picture, wherein the syntax element is coded as u(3), u(3) being an unsigned integer using 3 bits, or the syntax element is coded as ue(v), ue(v) being an unsigned integer using an exponential Golomb code, and wherein the format rule further specifies that the SEI message further includes a list of random access point (RAP) picture identifiers of IRAP pictures or DRAP pictures that are within the same coded layer video sequence (CLVS) as the DRAP picture, wherein the IRAP picture or the DRAP picture is allowed to be included in a valid entry of a reference picture list of the DRAP picture, and wherein each identifier in the list of RAP picture identifiers is coded as the same as a RAP picture identifier of a DRAP picture associated with the SEI message.
[0393] 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a conversion between a visual media data and a bitstream of the visual media data according to a format rule, wherein the format rule specifies that a supplemental enhancement information (SEI) message referring to a dependent random access point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating a number of intra random access point (IRAP) pictures or dependent random access point (DRAP) pictures that are within the same coded layer video sequence (CLVS) as the DRAP picture.
[0394] 18. The non-transitory computer-readable storage medium according to solution 17, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of a DRAP picture, wherein the syntax element is coded as u(3), u(3) being an unsigned integer using 3 bits, or the syntax element is coded as ue(v), ue(v) being an unsigned integer using an exponential Golomb code, and wherein the format rule further specifies that the SEI message further includes a list of random access point (RAP) picture identifiers of IRAP pictures or DRAP pictures that are within the same coded layer video sequence (CLVS) as the DRAP picture, wherein the IRAP picture or the DRAP picture is allowed to be included in a valid entry of a reference picture list of the DRAP picture, and wherein each identifier in the list of RAP picture identifiers is coded as the same as a RAP picture identifier of a DRAP picture associated with the SEI message.
[0395] 19. A non-transitory computer-readable storage medium storing a bitstream of visual media data generated by a method performed by a visual media data processing apparatus, wherein the method comprises determining that a supplemental enhancement information (SEI) message referring to a dependent random access point (DRAP) picture is included in the bitstream, and generating the bitstream based on the determination.
[0396] 20. The non-transitory computer-readable storage medium of solution 19, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of the DRAP picture, wherein a syntax element is coded as u(3), u(3) being an unsigned integer using 3 bits, or a syntax element is coded as ue(v), ue(v) being an unsigned integer using an exponential Golomb code, or a format rule further specifies that the SEI message further comprises a list of random access point (RAP) picture identifiers of IRAP pictures or DRAP pictures that are within the same coded layer video sequence (CLVS) as the DRAP picture, wherein the IRAP picture or the DRAP picture is allowed to be included in a valid entry of a reference picture list of the DRAP picture, and wherein each identifier in the list of RAP picture identifiers is coded as the same as a RAP picture identifier of the DRAP picture associated with the SEI message.
[0397] 21. A visual media data processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 14.
[0398] 22. A method of storing a bitstream of visual media data, comprising a method recited in any one of solutions 1 to 14, further comprising storing the bitstream to a non-transitory computer-readable storage medium.
[0399] 23. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of solutions 1 to 14.
[0400] 24. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0401] 25. A visual media data processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement a method recited in any one or more of solutions 1 to 14.
[0402] 26. A computer-readable medium having a bitstream thereon conforming to a format rule recited in any one of solutions 1 to 14.
[0403] In the solutions described herein, the visual media data corresponds to video or images. In the solutions described herein, an encoder can conform to a format rule by generating a coded representation according to the format rule. In the solutions described herein, a decoder can use a format rule to parse syntax elements in a coded representation to generate decoded video, given knowledge of the presence and absence of syntax elements according to the format rule.
[0404] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block can for example correspond to collocated or scattered bits within the bitstream as defined by the syntax. For example, a macroblock can be encoded according to transformed and coded error residual values and also using bits in headers and other fields in the bitstream. Furthermore, during conversion, a decoder can parse the bitstream based on the determination, given knowledge of the presence or absence of some fields, as described in the above solutions. Similarly, an encoder can determine to include or not include certain syntax fields and generate a coded representation accordingly by including or excluding syntax fields from the coded representation.
[0405] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of one or more of them, or a combination of one or more of them and other computer program products. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated for the purpose of encoding information for transmission to suitable receiver apparatus.
[0406] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be run on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0407] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and / or devices that are
[0408] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0409] While this patent document contains many details, these should not be construed as limiting the scope of any subject matter or of any embodiment, but as merely describing features that are specific to certain embodiments of the specific technology described in this patent document. Certain features described in the context of separate embodiments in this patent document can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any appropriate subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0410] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring any particular order among the operations or that all of the operations be performed, to achieve desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0411] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method of processing visual media data, comprising: performing a conversion between a visual media data and a bitstream of the visual media data comprising multiple layers according to a format rule, wherein the format rule specifies that a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream, and wherein the first type of SEI message indicates a first type of dependent random access point (DRAP) picture, and the second type of SEI message indicates a second type of DRAP picture, wherein the first type of DRAP picture is a picture that depends on an IRAP picture and is associated with the first type of SEI message, the first type of SEI message being a first DRAP indication SEI message, wherein the second type of DRAP picture is a picture that is allowed to depend on an IRAP picture or another DRAP picture and is associated with the second type of SEI message, the second type of SEI message being a second DRAP indication SEI message; wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier variable (RapPicld), and the value of the RAP picture identifier variable (RapPicld) of an IRAP picture is inferred to be equal to 0; wherein the format rule further specifies that the second type of SEI message includes a list of random access point (RAP) picture identifiers of IRAP pictures or second type of DRAP pictures within the same coded layer video sequence (CLVS) as the second type of DRAP picture, and wherein each of the list of RAP picture identifiers is identically coded with the RAP picture identifier of the second type of DRAP picture associated with the second type of SEI message.
2. The method of claim 1, wherein, the format rule further specifies that the first type of SEI message or the second type of SEI message includes information about a picture order count (POC) value of the first type of DRAP picture or the second type of DRAP picture.
3. The method of claim 1, wherein, the format rule further specifies that the values of the RAP picture identifiers of any two IRAP pictures or DRAP pictures within a coded layer video sequence (CLVS) are different from each other.
4. The method of claim 1, wherein, the format rule further specifies that the value of the RAP picture identifier of an IRAP picture or a DRAP picture within a coded layer video sequence (CLVS) is incremented in accordance with the increasing decoding order of the IRAP picture or the DRAP picture.
5. The method of claim 1, wherein, the format rule further specifies that the value of the RAP picture identifier of the DRAP picture is greater by one than the value of a previous IRAP picture or DRAP picture in decoding order within a coded layer video sequence (CLVS).
6. The method of any one of claims 1-5, wherein, the performing of the conversion comprises generating the bitstream from the visual media data.
7. The method of any one of claims 1-5, wherein, the performing of the conversion comprises reconstructing the visual media data from the bitstream.
8. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: performing a conversion between visual media data and a bitstream of the visual media data comprising multiple layers according to a format rule, wherein the format rule specifies that a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream, and wherein the first type of SEI message indicates a first type of dependent random access point (DRAP) picture, and the second type of SEI message indicates a second type of DRAP picture, wherein the first type of DRAP picture is a picture dependent on an IRAP picture and associated with the first type of SEI message, the first type of SEI message being a first DRAP indication SEI message, wherein the second type of DRAP picture is a picture allowed to be dependent on an IRAP picture or another DRAP picture and associated with the second type of SEI message, the second type of SEI message being a second DRAP indication SEI message; wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier variable (RapPicId), and the value of the RAP picture identifier variable (RapPicId) of an IRAP picture is inferred to be equal to 0; wherein the format rule further specifies that the second type of SEI message includes a list of random access point (RAP) picture identifiers of IRAP pictures or second type of DRAP pictures within the same coded layer video sequence (CLVS) as the second type of DRAP picture, and wherein each of the list of RAP picture identifiers is identically coded with the RAP picture identifier of the second type of DRAP picture associated with the second type of SEI message.
9. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between visual media data and a bitstream of the visual media data comprising multiple layers according to a format rule, wherein, the format rule specifies that a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream, and wherein the first type of SEI message indicates a first type of dependent random access point (DRAP) picture, and the second type of SEI message indicates a second type of DRAP picture, wherein the first type of DRAP picture is a picture dependent on an IRAP picture and associated with the first type of SEI message, the first type of SEI message being a first DRAP indication SEI message, wherein the second type of DRAP picture is a picture allowed to be dependent on an IRAP picture or another DRAP picture and associated with the second type of SEI message, the second type of SEI message being a second DRAP indication SEI message; wherein the format rule further specifies that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier variable (RapPicld), and the value of the RAP picture identifier variable (RapPicld) of an IRAP picture is inferred to be equal to 0; wherein the format rule further specifies that the second type of SEI message includes a list of random access point (RAP) picture identifiers of IRAP pictures or second type of DRAP pictures within the same coded layer video sequence (CLVS) as the second type of DRAP picture, and wherein each of the list of RAP picture identifiers is identically coded with the RAP picture identifier of the second type of DRAP picture associated with the second type of SEI message.
10. A non-transitory computer-readable storage medium storing a bitstream comprising multiple layers of visual media data, the bitstream generated by a method performed by a visual media data processing apparatus, wherein, The method comprises: determining that a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream; and generating the bitstream based on the determination, wherein the first type of SEI message indicates a first type of dependent random access point (DRAP) picture, and the second type of SEI message indicates a second type of DRAP picture, wherein the first type of DRAP picture is a picture that depends on an IRAP picture and is associated with the first type of SEI message, the first type of SEI message being a first DRAP indication SEI message, wherein the second type of DRAP picture is a picture that is allowed to depend on an IRAP picture or another DRAP picture and is associated with the second type of SEI message, the second type of SEI message being a second DRAP indication SEI message; wherein each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier variable (RapPicld), and the value of the RAP picture identifier variable (RapPicld) of an IRAP picture is inferred to be equal to 0; wherein the second type of SEI message includes a list of random access point (RAP) picture identifiers of IRAP pictures or second type of DRAP pictures within the same coded layer video sequence (CLVS) as the second type of DRAP picture, and wherein each of the list of RAP picture identifiers is identically coded with the RAP picture identifier of the second type of DRAP picture associated with the second type of SEI message.
11. A method for storing a bitstream including multiple layers of a video, comprising: determining that a second type of supplemental enhancement information (SEI) message different from a first type of SEI message is included in the bitstream; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the first type of SEI message indicates a first type of dependent random access point (DRAP) picture, and the second type of SEI message indicates a second type of DRAP picture, wherein the first type of DRAP picture is a picture that depends on an IRAP picture and is associated with the first type of SEI message, the first type of SEI message being a first DRAP indication SEI message, wherein the second type of DRAP picture is a picture that is allowed to depend on an IRAP picture or another DRAP picture and is associated with the second type of SEI message, the second type of SEI message being a second DRAP indication SEI message; wherein each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier variable (RapPicld), and the value of the RAP picture identifier variable (RapPicld) of an IRAP picture is inferred to be equal to 0; wherein the second type of SEI message includes a list of random access point (RAP) picture identifiers of IRAP pictures or second type of DRAP pictures within the same coded layer video sequence (CLVS) as the second type of DRAP picture, and wherein each of the list of RAP picture identifiers is identically coded with the RAP picture identifier of the second type of DRAP picture associated with the second type of SEI message.
Citation Information
Patent Citations
Dependent random access point pictures
WO2015192990A1