Supplemental enhancement information for multi-layer video streams

By modifying the semantics of the DRAP instruction SEI message and introducing a type 2 DRAP instruction, the problem of cross-random access point reference in DASH and ISOBMFF applications for multi-layer bitstreams is solved, enabling efficient video streaming and decoding.

CN114339244BActive Publication Date: 2026-01-06FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111142836.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-29
Filing Date
2021-09-28
Publication Date
2026-01-06
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

In existing technologies, multi-layer video stream bitstreams are difficult to effectively support cross-random access point references in DASH and ISOBMFF applications. In particular, the semantics of DRAP instruction SEI messages are only applicable to single-layer bitstreams, which means that file and DASH media representation editors need to parse and derive a lot of information to achieve CRR or EDR streaming operations.

Method used

By modifying the semantics of the DRAP indication SEI message to make it applicable to multi-layer bitstreams, and introducing a new SEI message type 2 DRAP indication, the dependencies and reference order of images are clearly defined, allowing the decoder to correctly decode DRAP images and their subsequent images without decoding associated IRAP images.

Benefits of technology

It enables efficient decoding of DRAP images and their subsequent images in multi-layer bitstreams, simplifies streaming operations in DASH and ISOBMFF applications, and improves encoding efficiency and decoding flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114339244B_ABST
    Figure CN114339244B_ABST
Patent Text Reader

Abstract

Systems, methods, and apparatuses of supplemental enhancement information, encoding, decoding, or transcoding visual media data of a multi-layer video stream are described. An example method of processing visual media data includes performing a conversion between visual media data and a bitstream of the visual media data including multiple layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in decoding order and output order without having to decode other pictures in the layer except for an intra random access point (IRAP) picture associated with the DRAP picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is made to promptly claim priority and interest in U.S. Provisional Patent Application No. 63 / 084,953, filed September 29, 2020, in accordance with applicable patent law and / or the rules of the Paris Convention. For all purposes under the law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to digital video encoding and decoding technologies, including video encoding, transcoding, or decoding. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process video or image representations according to file formats.

[0006] In one example aspect, a method for processing visual media data is disclosed. The method includes: performing a conversion between the visual media data and a multi-layered bitstream of the visual media data according to format rules; wherein the format rules specify that Supplemental Enhancement Information (SEI) messages are included in the bitstream to indicate that a decoder is permitted to decode 1) a Dependent Random Access Point (DRAP) picture in the layer associated with the SEI message and / or 2) pictures included in that layer and following the DRAP picture in both decoding and output order, without having to decode other pictures in that layer besides the Intra-Frame Random Access Point (IRAP) picture associated with the DRAP picture.

[0007] In another example, a different method for processing visual media data is disclosed. This method includes performing a conversion between visual media data and a bitstream of visual media data according to format rules, wherein the format rules specify whether and how a second type of SEI message, different from a first type of Supplemental Enhancement Information (SEI) message, is included in the bitstream, and wherein the first type of SEI message and the second type of SEI message respectively indicate a first type of Dependent Random Access Point (DRAP) image and a second type of DRAP image.

[0008] In another example, a different method for processing visual media data is disclosed. This method includes performing a conversion between visual media data and a bitstream of visual media data according to format rules, wherein the format rules specify that a Supplemental Enhancement Information (SEI) message referring to a Dependent Random Access Point (DRAP) picture is included in the bitstream, and wherein the format rules further specify that the SEI message includes a syntax element indicating the number of Intra-Random Access Point (IRAP) pictures or DRAP pictures within the same codec layer video sequence (CLVS) as the DRAP picture.

[0009] In yet another example, a video processing apparatus is disclosed. This video processing apparatus includes a processor configured to implement the methods described above.

[0010] In yet another example, a method for storing visual media data into a file comprising one or more bitstreams is disclosed. This method corresponds to the method described above and further includes storing the one or more bitstreams into a non-transitory computer-readable recording medium.

[0011] In yet another example, a computer-readable medium for storing a bitstream is disclosed. The bitstream is generated according to the method described above.

[0012] In yet another example, a video processing apparatus for storing bitstreams is disclosed, wherein the video processing apparatus is configured to implement the above-described method.

[0013] In yet another example, a computer-readable medium is disclosed on which a bitstream conforms to a file format generated according to the method described above.

[0014] These and other features are described throughout this document. Attached Figure Description

[0015] Figure 1 This is a block diagram of an example video processing system.

[0016] Figure 2 This is a block diagram of a video processing device.

[0017] Figure 3 This is a flowchart of an example method for video processing.

[0018] Figure 4 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0019] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0020] Figure 6This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0021] Figures 7 to 9 This is a flowchart of an example method for processing visual media data based on some embodiments of the disclosed technology. Detailed Implementation

[0022] For ease of understanding, chapter headings are used in this document, and the applicability of the techniques and embodiments disclosed in each chapter is not limited to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Thus, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editable changes to text are indicated by strikethrough indicating deleted text and highlighting (including bold and italics) indicating added text, relative to the current draft of the VVC specification.

[0023] 1. Preliminary Discussion

[0024] This document relates to video codec technologies. Specifically, it relates to support for cross-random access point (RAP) references in video codecs based on Supplemental Enhancement Information (SEI) messages. These ideas can be applied individually or in various combinations to any standard or non-standard video codec, such as the recently finalized Universal Video Codec (VVC).

[0025] 2. Abbreviation

[0026] ACT Adaptive Color Transformation

[0027] ALF Adaptive Loop Filter

[0028] AMVR Adaptive Motion Vector Resolution

[0029] APS Adaptive Parameter Set

[0030] AU Access Unit

[0031] AUD Access Unit Separator

[0032] AVC Advanced Video Codec (Rec.ITU-T H.264|ISO / IEC 14496-10)

[0033] B Two-way prediction

[0034] BCW bidirectional prediction with CU-level weights

[0035] BDOF bidirectional optical flow

[0036] BDPCM (Block-based Incremental Pulse Codec Modulation)

[0037] BP buffer period

[0038] CABAC Context-Based Adaptive Binary Arithmetic Encoding and Decoding

[0039] CB codec block

[0040] CBR Constant Bit Rate

[0041] CCALF Cross-Component Adaptive Loop Filter

[0042] CLVS codec layer video sequence

[0043] CLVSS codec layer video sequence start

[0044] CPB image buffer

[0045] CRA (Clean Random Access)

[0046] CRC Cyclic Redundancy Check

[0047] CTB codec tree block

[0048] CTU (Codec Tree Unit)

[0049] CU encoding / decoding unit

[0050] CVS codec video sequence

[0051] CVSS codec video sequence start

[0052] DPB Decoding Image Buffer

[0053] DCI decoding capability information

[0054] DRAP relies on random access points

[0055] DU decoding unit

[0056] DUI Decoding Unit Information

[0057] EG Index Columbus

[0058] EGk k-order exponent Columbus

[0059] EOB bitstream end

[0060] EOS sequence ends

[0061] FD fill data

[0062] FIFO (First In First Out)

[0063] FL fixed length

[0064] GBR Green, Blue and Red

[0065] GCI General Constraint Information

[0066] GDR gradually decoded and refreshed

[0067] GPM Geometric Segmentation Mode

[0068] HEVC High-Efficiency Video Codec (Rec.ITU-T H.265|ISO / IEC 23008-2)

[0069] HRD Assumption Reference Decoder

[0070] HSS Assumption Stream Scheduler

[0071] I within the frame

[0072] IBC Intra-block Copy

[0073] IDR Instant Decoding and Refresh

[0074] ILRP interlayer reference image

[0075] IRAP Intra-Frame Random Access Point

[0076] LFNST Low-Frequency Inseparable Transform

[0077] LPS least likely symbol

[0078] LSB (Least Significant Bit)

[0079] LTRP Long-Term Reference Image

[0080] LMCS Luminance Map with Chroma Scaling

[0081] MIP (Matrix-Based Intra-Frame Prediction)

[0082] MPS Maximum Possible Symbol

[0083] MSB Most significant bit

[0084] MTS Multiple Transformation Selection

[0085] MVP Motion Vector Prediction

[0086] NAL Network Abstraction Layer

[0087] OLS Output Layer Set

[0088] OP operation point

[0089] OPI Operation Point Information

[0090] P prediction

[0091] PH Image Header

[0092] POC Image Sequential Counting

[0093] PPS Image Parameter Set

[0094] PROF refines predictions using optical flow.

[0095] PT Image Time Sequence

[0096] PU Image Unit

[0097] QP Quantization Parameters

[0098] RADL Random Access Decodable Bootstrap (Image)

[0099] RAP Random Access Point

[0100] RASL Random Access Skip to Bootstrap (Image)

[0101] RBSP raw byte sequence payload

[0102] RGB red, green and blue

[0103] RPL Reference Image List

[0104] SAO Sample Adaptive Offset

[0105] SAR sample aspect ratio

[0106] SEI Supplemental Enhancement Information

[0107] SH strip head

[0108] SLI sub-image level information

[0109] SODB data bit string

[0110] SPS Sequence Parameter Set

[0111] STRP Short-Term Reference Image

[0112] STSA Stepwise Temporal Sublayer Access

[0113] TR (truncated rice)

[0114] TU Transformer

[0115] VBR Variable Bit Rate

[0116] VCL (Video Codec Layer)

[0117] VPS Video Parameter Set

[0118] VSEI General Supplemental Enhancement Information (Rec.ITU-T H.274|ISO / IEC 23002-7)

[0119] VUI Video Availability Information

[0120] VVC Universal Video Codec (Rec.ITU-T H.266|ISO / IEC 23090-3)

[0121] 3. Introduction to Video Encoding and Decoding

[0122] 3.1. Video codec standards

[0123] Video codec standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Group (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). When the Universal Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Group (JVET). VVC is a new codec standard finalized by JVET at its 19th meeting, which concluded on July 1, 2020. Its goal is to reduce the bit rate by 50% compared to HEVC.

[0124] The Universal Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and the associated Universal Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple codec video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360° immersive media.

[0125] 3.2. Image Sequence Counting in HEVC and VVC (POC)

[0126] In HEVC and VVC, in many parts of the decoding process, including DPB management, one part of which is reference image management, POC is basically used as an image ID to identify the image.

[0127] For the newly introduced PH, in VVC, the least significant bit (LSB) information of the POC is signaled in the PH, unlike HEVC, where it is signaled in the SH. This LSB is used to derive the POC value and has the same value for all stripes of the picture. VVC also allows signaling the period value of the most significant bit (MSB) of the POC in the PH to enable the derivation of the POC value without tracking the POC MSB, which relies on the POC information of the earlier decoded picture. For example, this allows mixing IRAP and non-IRAP within an AU in a multi-layer bitstream. An additional difference between the POC signaling of HEVC and VVC is that in HEVC, there is no signaling notification of the POC LSB for IDR pictures. This presented some disadvantages during the later development of multi-layer extensions of HEVC to enable mixing IDR and non-IDR pictures within an AU. Therefore, in VVC, the POC LSB information is signaled for each picture, including IDR pictures. The signaling notification of POC LSB information for IDR images also makes it easier to support merging IDR and non-IDR images from different bitstreams into a single image, which would otherwise require some complex design to handle the POC LSB in the merged image.

[0128] 3.3. Random Access in HEVC and VVC and its Support

[0129] Random access refers to accessing and decoding the bitstream starting from the image that is not the first image in the bitstream according to the decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, searching in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points. These random access points are typically intra-frame encoded images, but can also be inter-frame encoded images (e.g., in the case of progressive decoding refresh).

[0130] HEVC includes signaling notifications for Intra-Random Access Point (IRAP) pictures in the NAL unit header through NAL unit types. Three types of IRAP pictures are supported: Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any pictures preceding the current group of pictures (GOP), traditionally referred to as closed GOP random access points. CRA pictures are less restrictive by allowing some pictures to reference pictures preceding the current GOP; in the case of random access, all pictures are discarded. CRA pictures are traditionally referred to as open GOP random access points. BLA pictures are typically derived from splicing of two bitstreams or portions thereof at the CRA picture, for example, during stream switching. To enable the system to better utilize IRAP pictures, a total of six different NAL units are defined to signal attributes of IRAP pictures. These attributes can be used to better match the stream access point types defined in the ISO Base Media File Format (ISOBMFF), which is used to support HTTP's Dynamic Adaptive Streaming (DASH).

[0131] VVC supports three types of IRAP pictures, two types of IDR pictures (RADL pictures with or without an association with one type), and one type of CRA picture. These are essentially the same as in HEVC. VVC does not include the BLA picture type from HEVC for two main reasons: i) the basic functionality of a BLA picture can be achieved by adding a sequence end NAL unit to a CRA picture, the presence of which indicates that a new CVS begins in a single-layer bitstream; ii) during the development of VVC, it was desirable to specify fewer NAL unit types than in HEVC, as indicated by the use of five bits instead of six bits in the NAL unit type field in the NAL unit header.

[0132] Another key difference between VVC and HEVC in terms of random access support is that GDR support in VVC is implemented in a more canonical manner. In GDR, bitstream decoding can begin with an inter-frame encoded picture, and although not the entire picture region can be correctly decoded at the beginning, it will be correct after several pictures. AVC and HEVC also support GDR, using the Recovery Point SEI message for signaling notification of GDR random access points and recovery points. In VVC, a new NAL unit type is specified to indicate GDR pictures, and recovery points are signaled in the picture header syntax structure. This allows CVS and bitstreams to begin with GDR pictures. This means that an entire bitstream can contain only inter-frame encoded pictures, without any single intra-frame encoded pictures. The main benefit of specifying GDR support in this way is providing consistent behavior for GDR. GDR enables encoders to smooth the bitrate of a bitstream by distributing intra-frame encoded stripes or blocks across multiple images, rather than intra-frame encoding and decoding the entire image, thus significantly reducing end-to-end latency. This is considered more important today than ever as wireless displays, online gaming, and drone-based applications become increasingly popular.

[0133] Another feature related to GDR in VVC is virtual boundary signaling notification. On the image between the GDR image and its recovery points, the boundary between refreshed areas (i.e., correctly decoded areas) and unrefreshed areas can be signaled as a virtual boundary. When this signaling is notified, in-loop filtering across the boundary is not applied, thus preventing decoding mismatches at or near the boundary. This can be useful when the application determines which areas to display during the GDR process.

[0134] IRAP images and GDR images can be collectively referred to as random access point (RAP) images.

[0135] 3.4.VUI and SEI messages

[0136] The VUI is a syntax structure sent as part of the SPS (and possibly in the HEVC VPS). The information carried by the VUI does not affect the canonical decoding process, but it may be important for the correct rendering of the encoded and decoded video.

[0137] SEI assists in processes related to decoding, display, or other purposes. Like VUI, SEI does not affect the specification's decoding process. SEI is carried within SEI messages. Decoder support for SEI messages is optional. However, SEI messages do affect bitstream consistency (e.g., if the syntax of SEI messages in the bitstream does not conform to the specification, the bitstream is inconsistent), and some SEI messages are required in the HRD specification.

[0138] The VUI syntax structures and most SEI messages used with VVC are not specified in the VVC specification, but rather in the VSEI specification. The VVC specification defines the SEI information required for HRD conformance testing. VVC v1 defines five SEI messages related to HRD conformance testing, and VSEI v1 specifies 20 additional SEI messages. The SEI messages carried in the VSEI specification do not directly affect the behavior of the conforming decoder and are already defined, therefore they can be used in a codec-agnostic manner, allowing VSEI to be used with other video codec standards besides VVC in the future. The VSEI specification does not refer to specific VVC syntax element names, but rather to variables whose values ​​are set in the VVC specification.

[0139] Compared to HEVC, VVC's VUI syntax structure focuses only on information related to the correct rendering of the image, without including any timing information or bitstream limit indications. In VVC, the VUI is signaled in the SPS (Signaling Partition Script), which includes a length field preceding the VUI syntax structure, signaling the length of the VUI payload in bytes. This allows the decoder to easily skip information, and more importantly, it allows for convenient future VUI syntax extensions by adding new syntax elements directly to the end of the VUI syntax structure in a manner similar to SEI message syntax extensions.

[0140] The VUI syntax structure contains the following information:

[0141] • The content is either line-spaced or line-by-line;

[0142] • Does the content include frame-packed stereoscopic video or projected omnidirectional video?

[0143] • Aspect ratio of sample points;

[0144] • Is the content suitable for overscan display?

[0145] • Color description, including primary colors, matrix, and transmission characteristics, is particularly important for signaling communication of Ultra High Definition (UHD) and High Definition (HD) color spaces as well as High Dynamic Range (HDR);

[0146] • Chromaticity position relative to luminance (for progressive content clarification signaling notification compared to HEVC).

[0147] When the SPS does not contain any VUI, this information is considered unspecified, and if the content of the bitstream is intended to be displayed on the screen, the information must be transmitted by external means or specified by the application.

[0148] Table 1 lists all SEI messages specified for VVC v1, along with the specification containing their syntax and semantics. Of the 20 SEI messages specified in the VSEI specification, many are inherited from HEVC (e.g., padding payload and user data SEI messages). Some SEI messages are crucial for the proper processing or presentation of encoded or decoded video content. For example, this includes SEI messages for primary display color volume, content light level information, or alternative transport characteristics particularly relevant to HDR content. Other examples include equirectangular projection, spherical rotation, region packing, or omnidirectional viewport SEI messages, which are related to signaling notification and processing of 360° video content.

[0149] Table 1: SEI Message List in VVC v1

[0150]

[0151]

[0152] The new SEI messages specified for VVC v1 include frame-field information SEI messages, sample aspect ratio information SEI messages, and sub-picture level information SEI messages.

[0153] The Frame-Field Information (SEI) message contains information indicating how the associated picture should be displayed (e.g., field parity or frame repetition period), the source scan type of the associated picture, and whether the associated picture is a copy of a previous picture. In previous video codec standards, this information, along with the timing information of the associated picture, could be used for signaling notification in the Picture Timing SEI message. However, it was observed that Frame-Field Information and timing information are two distinct pieces of information, and they are not necessarily signaled together. A typical example is that timing information is signaled at the system level, but Frame-Field Information is signaled at the bitstream level. Therefore, it was decided to remove Frame-Field Information from the Picture Timing SEI message and instead signal it in a dedicated SEI message. This change also allows for modification of the syntax of the Frame-Field Information to convey additional and clearer instructions to the display, such as pairing fields together or setting more values ​​for frame repetition.

[0154] The Sample Aspect Ratio (SEI) message can notify different images within the same sequence of different sample aspect ratios, while the corresponding information contained in the VUI is applied to the entire sequence. When using a reference image resampling function with a scaling factor, this may result in different images within the same sequence having different sample aspect ratios.

[0155] The Sub-Picture Level Information (SEI) message provides level information for a sequence of sub-pictures.

[0156] 3.5. Cross-RAP Reference

[0157] JVET-M0360, JVET-N0119, JVET-O0149, and JVET-P0114 propose a video coding and decoding method based on cross-RAP reference (CRR), also known as external decode refresh (EDR).

[0158] The basic idea behind this video encoding / decoding method is as follows: Instead of encoding and decoding random access points (RIPs) as intra-frame encoded IRAP pictures (except for the first picture in the bitstream), they are encoded and decoded using inter-frame prediction to overcome the unavailability of earlier pictures if RIPs were encoded and decoded as IRAP pictures. The trick is to provide a limited number of earlier pictures, typically representing different scenes of the video content, through a separate video bitstream, which can be called an external device. These earlier pictures are called external pictures. Therefore, each external picture can be used as an inter-frame prediction reference by the pictures across RIPs. The encoding / decoding efficiency gain comes from encoding and decoding RIPs as inter-frame prediction pictures and having more usable reference pictures for pictures that follow the EDR pictures in the decoding order.

[0159] As described below, bitstreams encoded and decoded using this video encoding and decoding method can be used in applications based on ISOBMFF and DASH.

[0160] DASH Content Preparation Operation

[0161] 1) The video content is encoded into one or more representations, each with a specific spatial resolution, temporal resolution, and quality.

[0162] 2) Each specific representation of the video content can be represented by a main stream or an external stream. A main stream contains a codec image, which may or may not contain an EDR image. When the main stream includes at least one EDR image, the external stream also exists and contains the external image. When the main stream does not contain an EDR image, the external stream does not exist.

[0163] 3) Each mainstream is carried in the Main Stream Representation (MSR). Each EDR image in the MSR is the first image of a segment.

[0164] 4) Each external stream (if it exists) is carried in an External Stream Representation (ESR).

[0165] 5) For each segment in the MSR that begins with an EDR picture, there exists a segment in the corresponding ESR with the same segment start time derived from the MPD, carrying the external picture required to decode the EDR picture, as well as the subsequent pictures in the bitstream carried by the MSR in the order of decoding.

[0166] 6) MSRs of the same video content are included in an Adaptation Set (AS). ESRs of the same video content are included in an AS.

[0167] DASH streaming operation

[0168] 1) The client obtains the MPD of the DASH media representation, parses the MPD, selects the MSR, and determines the start representation time of the content to be consumed.

[0169] 2) The client requests segments of MSR, starting with the segment containing the image whose representation time is equal to (or sufficiently close to) the starting representation time.

[0170] a. If the first image in the starting segment is an EDR image, then the corresponding segment in the associated ESR (with the same segment start time derived from MPD) is also requested, preferably before requesting the MSR segment. Otherwise, no segments in the associated ESR are requested.

[0171] 3) When switching to a different MSR, the client requests to switch to the MSR segment starting from the first segment, and the segment start time of the first segment is greater than the start time of the last segment requested to switch from the MSR.

[0172] a. If the first image in the starting segment of the MSR is an EDR image, the corresponding segment in the associated ESR will also be requested, preferably before the MSR segment is requested. Otherwise, no segments in the associated ESR will be requested.

[0173] 4) When operating on the same MSR consecutively (after decoding the starting segment after a search or stream switching operation), it is not necessary to request any segment of the associated ESR, including when requesting any segment that starts with an EDR image.

[0174] 3.6. DRAP Instruction SEI Message

[0175] The VSEI specification includes the DRAP instruction SEI message, as follows:

[0176] dependent_rap_indication(payloadSize){ descriptor }

[0177] The image associated with the Dependency Random Access Point (DRAP) Indicator SEI message is called a DRAP image.

[0178] The presence of the DRAP Indication SEI message indicates that the constraints on picture order and picture references specified in this clause apply. These constraints enable the decoder to correctly decode the DRAP picture and subsequent pictures in the decoding and output order, without needing to decode any other pictures besides the associated IRAP picture of the DRAP picture.

[0179] The presence of the DRAP-indicating SEI message indicates the following constraints, all of which apply:

[0180] –DRAP images are trailing pictures;

[0181] –The temporal sublayer identifier for the DRAP image is equal to 0;

[0182] – Apart from the associated IRAP image, a DRAP image does not include any images in the active entry of its reference image list;

[0183] - Any image following the DRAP image in decoding and output order is excluded from the valid entries in its reference image list from any image preceding the DRAP image in decoding or output order, except for the associated IRAP image of the DRAP image.

[0184] 4. The technical problem solved by the disclosed technical solution

[0185] The functionality of the DRAP Indication SEI message can be considered a subset of the CRR method. For simplicity, the picture associated with the DRAP Indication SEI message is referred to as a Type 1 DRAP picture.

[0186] From an encoding perspective, although the CRR method proposed in JVET-P0114 or earlier JVET contributions was not adopted by VVC, the encoder can still encode the video bitstream in such a way that some pictures depend only on the associated IRAP picture (e.g., a type 1 DRAP picture indicated by the DRAP SEI message) used as a reference for inter-frame prediction, and some other pictures (e.g., referred to as type 2 DRAP pictures) depend only on some pictures in a set of pictures consisting of the associated IRAP picture and some other (type 1 or type 2) DRAP pictures.

[0187] However, given a VVC bitstream, it is unknown whether such a Type 2 DRAP image exists within the bitstream. Furthermore, even if the existence of such a Type 2 DRAP image in the bitstream is known, in order to assemble a media file based on the ISOBMFF and DASH media representations of this VVC bitstream for CRR or EDR streaming operations, the file and DASH media representation editors would need to parse and derive a significant amount of information, including POC values ​​and valid entries in the reference image list, to determine whether a particular image is a Type 2 DRAP image, and if so, which earlier IRAP or DRAP image is needed for random access from that particular image, so that the appropriate set of images can be included in separate, time-synchronized file tracks and DASH representations.

[0188] Another issue is that the semantics of DRAP-indicated SEI messages only apply to single-layer bitstreams.

[0189] 5. Solution List

[0190] To address the above issues, and others, the following summarized methods are disclosed. These items should be considered as examples for explaining general concepts and should not be interpreted narrowly. Furthermore, these items can be applied individually or in any combination.

[0191] 1) In one example, the semantics of the DRAP Indicator SEI message are changed so that the SEI message can be applied to multi-layer bitstreams. That is, the semantics enable the decoder to correctly decode the DRAP picture (i.e., the picture associated with the DRAP Indicator SEI message) and the pictures in the same layer that follow the DRAP in the decoding and output order, without needing to decode any other pictures in the same layer other than the associated IRAP picture of the DRAP picture.

[0192] a. For example, it is required that, apart from the associated IRAP image of the DRAP image, the DRAP image does not include any images in the same layer in the valid entries of its reference image list.

[0193] b. In one example, it is required that any image in the same layer that follows the DRAP image in the decoding and output order is not included in the valid entries of its reference image list as any image in the same layer that precedes the DRAP image in the decoding or output order, except for the associated IRAP image of the DRAP image.

[0194] 2) In one example, the RAP image ID for a DRAP image is signaled in the DRAP Instruction SEI message to specify the identifier of the RAP image. The RAP image can be an IRAP image or a DRAP image.

[0195] a. In one example, a presence flag indicating whether a RAP image ID exists in a DRAP indication is signaled, and when the flag is equal to a specific value, such as 1, the RAP image ID is signaled in the DRAP indication SEI message, and when the flag is equal to another value, such as 0, the RAP image ID is not signaled in the DRAP indication SEI message.

[0196] 3) In one example, the DRAP picture associated with the DRAP indication SEI message can refer to the associated IRAP picture or the previous picture in the decoding order, which is the GDR picture for inter-frame prediction reference with ph_recovery_poc_cnt equal to 0.

[0197] 4) In one example, a new SEI message is named, for example, a Type 2 DRAP Indication SEI message, and each picture associated with this new SEI message is called a special type picture, such as a Type 2 DRAP picture.

[0198] 5) In one example, Type 1 DRAP pictures (associated with DRAP indication SEI messages) and Type 2 DRAP pictures (associated with Type 2 DRAP indication SEI messages) are collectively referred to as DRAP pictures.

[0199] 6) In one example, the Type 2 DRAP instruction SEI message includes a RAP picture ID, such as RapPicId, to specify the identifier of the RAP picture, which can be an IRAP picture or a DRAP picture, and a syntax element (e.g., t2drap_num_ref_rap_pics_minus1) that indicates the number of IRAP or DRAP pictures in the same CLVS as the Type 2 DRAP picture and which can be included in the list of reference pictures of the Type 2 DRAP picture.

[0200] a. In one example, the syntax elements indicating that number (e.g., t2drap_num_ref_rap_pics_minus1) are encoded as u(3) using 3 bits.

[0201] b. Optionally, indicate that the number of syntax elements (e.g., t2drap_num_ref_rap_pics_minus1) are encoded as ue(v).

[0202] 7) In one example, for the RAP image ID of a DRAP image, in a DRAP Indication SEI message or a Type 2 DRAP Indication SEI message, one or more of the following methods apply:

[0203] a. In one example, the syntax element of the signaling notification RAP image ID is encoded as u(16) using 16-bit encoding.

[0204] i. Alternatively, the syntax elements of the signaling notification RAP image ID are encoded and decoded using ue(v).

[0205] b. In one example, instead of signaling the RAP image ID in the DRAP Indication SEI message, the POC value of the DRAP image is signaled using se(v) or i(32).

[0206] i. Optionally, for example, use ue(v) or u(16) signaling to notify the POC increment relative to the associated IRAP picture's POC value.

[0207] 8) In one example, each IRAP or DRAP image that is designated as an IRAP or DRAP is associated with a RAP image IDRapPicId.

[0208] a. In one example, it is specified that the value of RapPicId for an IRAP image is inferred to be equal to 0.

[0209] b. In one example, it is specified that the RapPicId values ​​of any two IRAP or DRAP images within CLVS should be different.

[0210] c. In addition, in one example, the RapPicId value of the IRAP and DRAP images within CLVS will increase as the decoding order of the IRAP or DRAP images increases.

[0211] d. Additionally, in one example, within the same CLVS, the RapPicId of a DRAP image should be 1 greater than the RapPicId of the IRAP or DRAP image that comes first in the decoding order. 9) In one example, the Type 2 DRAP Indicator SEI message also includes a list of RAP image IDs.

[0212] A list for each of the valid entries in the reference image list of Type 2 DRAP images that is in the same CLVS as the Type 2 DRAP image and can be included in the Type 2 DRAP image.

[0213] a. In one example, each of the RAP image IDs in the list is encoded as having the same RAP image ID as the DRAP image associated with the Type 2 DRAP Indication SEI message.

[0214] b. Optionally, the values ​​of the RAP image ID list are required to be incremented in ascending order of the values ​​of the list index i, and the increment between the RapPicId value of the i-th DRAP pic and the RapPicId value of the (i-1)-th DRAP or IRAP pic (when i is greater than 0) or 2) 0 (when i is equal to 0) is encoded using ue(v).

[0215] c. Optionally, each of the RAP image IDs in the list is encoded to represent the POC value of the RAP image, for example, encoded as se(v) or i(32).

[0216] d. Optionally, each of the RAP image IDs in the list is encoded to represent the POC increment relative to the POC value of the associated IRAP image, for example, using ue(v), u(16) for signaling notification.

[0217] e. Optionally, each of the RAP image IDs in the list is encoded to represent the POC increment of the current image relative to 1) the POC value of the (i-1)th DRAP or IRAP image (when i is greater than 0) or 2) the POC value of the IRAP image (when i is equal to 0), for example using ue(v) or u(16).

[0218] f. Optionally, it is also required that for any two values ​​of list index values ​​i and j, for a list of RAP image IDs, when i is less than j, the i-th IRAP or DRAP image should precede the j-th IRAP or DRAP image in the decoding order.

[0219] 6. Example

[0220] The following are some example embodiments of the inventions summarized in Section 5 above, which can be applied to the VSEI specification. The changed text is based on the latest VSEI text in JVET-S2007-v7. Most relevant parts that have been added or modified are highlighted in bold and italics, and some deleted parts are marked with double brackets (e.g., [[a]] indicates the deletion of the letter "a"). There may be some other editable changes, which are not highlighted.

[0221] 6.1. First Embodiment

[0222] This embodiment is a modification of the existing DRAP instruction SEI message.

[0223] 6.1.1. Dependency on Random Access Point Indicator (SEI) Message Syntax

[0224]

[0225] 6.1.2. Reliance on Random Access Point Indicators (SEI) Message Semantics

[0226] The image associated with the Dependency Random Access Point (DRAP) Indicator SEI message is called a Type 1 DRAP image.

[0227] Type 1 DRAP images and Type 2 DRAP images (associated with Type 2 DRAP Indication SEI messages) are collectively referred to as DRAP images.

[0228] The presence of the DRAP Indication SEI message indicates that the constraints on picture order and picture references specified in this sub-clause apply. These constraints enable the decoder to correctly decode Type 1 DRAP pictures and pictures in the same layer that follow the Type 1 DRAP picture in the decoding and output order, without needing to decode any other pictures in the same layer besides the associated IRAP picture of the Type 1 DRAP picture.

[0229] The constraints indicated by the presence of the DRAP-indicating SEI message are as follows, and these constraints all apply:

[0230] – Type 1 DRAP images are rear images.

[0231] – The temporal sublayer identifier for a Type 1 DRAP image is equal to 0.

[0232] – Apart from the associated IRAP images of a Type 1 DRAP image, a Type 1 DRAP image is not included in any valid entries in its reference image list within the same layer.

[0233] - Any image in the same layer that follows the Type 1 DRAP image in decoding and output order is not included in the valid entries of its reference image list as any image in the same layer that precedes the Type 1 DRAP image in decoding or output order, except for the associated IRAP image of the Type 1 DRAP image.

[0234] drap_rap_id_in_clvs specifies the RAP image ID of type 1 Drap images, represented as RapPicId.

[0235] Each IRAP or DRAP image, whether IRAP or DRAP, is associated with a RapPicId. The RapPicId value of an IRAP image is inferred to be equal to 0. Any two IRAP or DRAP images within CLVS should have different RapPicId values.

[0236] 6.2. Second Embodiment

[0237] This embodiment is for the new Type 2 DRAP Indication SEI message.

[0238] 6.2.1. Type 2 DRAP Indicator SEI Message Syntax

[0239]

[0240]

[0241] 6.2.2. Type 2 DRAP Indicator SEI Message Semantics

[0242] The image associated with a Type 2 DRAP Indication SEI message is called a Type 2 DRAP image.

[0243] Type 1 DRAP images (associated with DRAP indication SEI messages) and Type 2 DRAP images are collectively referred to as DRAP images.

[0244] The presence of a Type 2 DRAP indication SEI message indicates that the constraints on picture order and picture references specified in this subsection apply. These constraints enable the decoder to correctly decode Type 2 DRAP pictures and pictures in the same layer that follow the Type 2 DRAP picture in decoding and output order, without needing to decode any other pictures in the same layer, in addition to the picture list referenceablePictures, which consists of a list of IRAP or DRAP pictures in the same CLVS, identified by the t2drap_ref_rap_id[i] syntax element, in decoding order.

[0245] The constraints indicated by the presence of the Type 2 DRAP indication SEI message are as follows, and all of these constraints apply:

[0246] – Type 2 DRAP images are rear-image images;

[0247] – The temporal sublayer identifier for a Type 2 DRAP image is equal to 0;

[0248] – Type 2 DRAP images do not include any images in the same layer, except for referenceable Pictures, in the valid entries of their reference picture list;

[0249] - Any picture in the same layer that follows the Type 2 DRAP picture in the decoding and output order, in the valid entries of its reference picture list, excluding any picture in the same layer that precedes the Type 2 DRAP picture in the decoding or output order, except for referenceablePictures;

[0250] -Any picture in the list of referenceablePictures is not included in the valid entries of its list of referenceablePictures except for any picture that is at the same level and is not an earlier picture in the list of referenceablePictures.

[0251] Note – Therefore, the first picture in referenceablePictures, even if it is a DRAP picture rather than an IRAP picture, does not include any pictures of the same layer in the valid entries of its reference picture list.

[0252] t2drap_rap_id_in_clvs specifies the RAP image identifier for 2drap images, represented as RapPicId.

[0253] Each IRAP or DRAP image, as either an IRAP or DRAP image, is associated with a RapPicId. The RapPicId value of an IRAP image is inferred to be equal to 0. Any two IRAP or DRAP images within CLVS should have different RapPicId values.

[0254] In this version of the bitstream conforming to this specification, t2drap_reserved_zero_13bits should be equal to 0. Other values ​​for t2drap_reserved_zero_13bits are reserved for future use by ITU-T|ISO / IEC. The decoder should ignore the value of t2drap_reserved_zero_13bits.

[0255] The increment of 1 in t2drap_num_ref_rap_pics_minus1 indicates the number of IRAP or DRAP images that are in the same CLVS as the Type 2 DRAP images and can be included in the valid entries of the reference image list for Type 2 DRAP images.

[0256] t2drap_ref_rap_id[i] indicates the RapPicId of the i-th IRAP or DRAP image that is in the same CLVS as the type 2 DRAP image and can be included in the list of reference images for type 2 DRAP images.

[0257] Figure 1This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0258] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via a communication connection, as represented by component 1906. The bitstream (or codec) representation of the video received at input 1902, whether stored or communicated, can be used by component 1908 to generate pixel values ​​or transmit as displayable video to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.

[0259] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0260] Figure 2This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors(multiple) 3602 may be configured to implement one or more methods described herein. The memories(multiple) 604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 606 may be used to implement some of the techniques described herein in a hardware circuit system. In some embodiments, the video processing hardware 3606 may be at least partially included in the processor 3602 (e.g., a graphics coprocessor).

[0261] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0262] like Figure 4 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, and this source device 110 may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110, and the target device 120 may be referred to as a video decoding device.

[0263] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0264] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via I / O interface 116 through network 130a. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0265] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0266] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interface with an external display device.

[0267] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or additional standards.

[0268] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 4 The video encoder 114 in the system 100 shown.

[0269] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0270] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0271] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0272] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 5 The examples are shown separately.

[0273] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0274] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).

[0275] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0276] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0277] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0278] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in lists 0 and 1 containing the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0279] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.

[0280] In some examples, the motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0281] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0282] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0283] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling Notification.

[0284] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0285] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0286] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0287] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0288] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0289] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in buffer 213.

[0290] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.

[0291] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0292] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 4 The video decoder 114 in the system 100 shown.

[0293] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0294] exist Figure 6 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 5 The encoding process described is the opposite of the decoding process.

[0295] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and based on the entropy-coded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.

[0296] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0297] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0298] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0299] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0300] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates the decoded video for presentation on the display device.

[0301] The following is a list of preferred solutions for some embodiments.

[0302] The first set of solutions is provided below, which illustrates example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0303] 1. A video processing method (e.g., Figure 3 The method 700 described herein includes: performing (702) a conversion between a video comprising multiple layers and a codec representation of the video, wherein the codec representation is organized according to a format rule; wherein the format rule specifies that supplementary enhancement information (SEI) is included in the codec representation, wherein the SEI information carries information sufficient to enable the decoder to decode dependent random access point (DRAP) pictures and / or decode pictures in a layer in the order of decoding and output, without needing to decode other pictures in that layer, except for intra-frame random access pictures (IRAP) of DRAP pictures.

[0304] 2. According to the method of Solution 1, DRAP images are excluded from the list of reference images for any image in that layer, except for IRAP images.

[0305] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0306] 3. A video processing method, comprising: performing a conversion between a multi-layered video and a codec representation of the video, wherein the codec representation is organized according to a format rule; wherein the format rule specifies that a Supplemental Enhancement Information (SEI) message is included in a codec representation of a Random Access Point (RAP) image, wherein the SEI message includes an identifier of the RAP image.

[0307] 4. According to the method of Solution 3, where RAP is an intra-frame random access image.

[0308] 5. Based on the approach of Solution 3, where RAP is based on Random Access Image (DRAP).

[0309] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).

[0310] 6. According to the method of Solution 5, DRAP images are allowed to refer to the associated intra-frame random access images or previous images in the decoding order, wherein the previous images are progressively decoded refresh images.

[0311] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 4-6).

[0312] 7. A video processing method comprising: performing a conversion between a multi-layered video and a video codec representation, wherein the codec representation is organized according to format rules; wherein the format rules specify whether and how Type 2 Supplemental Enhancement Information (SEI) messages referring to Random Access Attached Images (DRAP) are included in the codec representation.

[0313] 8. According to the method of Solution 7, the format rules specify that Type 2SEI messages and each picture associated with the message are treated as special type pictures.

[0314] 9. According to the method of Solution 7, wherein the format rules specify that the Type 2 SEI message includes an identifier for a random access picture (RAP) referred to as a Type 2 RAP picture, and a syntax element indicating the number of pictures in the same codec video layer as the random access picture, such that the picture is included in the list of valid reference pictures for the Type 2 RAP picture.

[0315] 10. The method according to any one of solutions 1-9, wherein the transformation includes generating a codec representation from the video.

[0316] 11. The method according to any one of solutions 1-9, wherein the conversion includes decoding the encoding / decoding representation to generate video.

[0317] 12. A video decoding apparatus, including a processor configured to implement the method described in one or more of solutions 1 to 11.

[0318] 13. A video encoding apparatus, including a processor configured to implement the method described in one or more of solutions 1 to 11.

[0319] 14. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of solutions 1 to 11.

[0320] 15. A computer-readable medium storing a codec representation generated according to any one of solutions 1 to 11.

[0321] 16. The methods, apparatus or systems described in this document.

[0322] The second set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., items 1, 1.a, 1.b, 2, 2.a, 3).

[0323] 1. A method for processing visual media data (e.g., such as...) Figure 7 The method shown (710) includes: performing (712) a conversion between visual media data and a bitstream of visual media data comprising multiple layers, according to a format rule; wherein the format rule specifies that a Supplemental Enhancement Information (SEI) message is included in the bitstream to indicate that the decoder is allowed to decode 1) a Dependent Random Access Point (DRAP) picture in the layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in the decoding and output order, without having to decode other pictures in the layer other than the Intra-Frame Random Access Point (IRAP) picture associated with the DRAP picture.

[0324] 2. According to the method of Solution 1, DRAP images are excluded from the valid entries of the reference image list of DRAP images in this layer, in addition to IRAP images.

[0325] 3. According to the method of Solution 1, the first image included in the layer and following the DRAP image in the decoding and output order is excluded from the valid entries of the reference image list of the first image, except for the IRAP image. The second image included in the layer and preceding the DRAP image in the decoding and output order is excluded.

[0326] 4. According to the method described in Solution 1, the format rules further specify that the SEI message includes an identifier for the Random Access Point (RAP) image.

[0327] 5. According to the method of Solution 4, the RAP image is either an IRAP image or a DRAP image.

[0328] 6. According to the method of Solution 4, the format rules also stipulate that a presence flag indicating the presence of a RAP image in the SEI message is included in the bitstream.

[0329] 7. According to the method of Solution 6, the presence of a flag with a value equal to the first value indicates that the identifier of the RAP image exists in the SEI message.

[0330] 8. According to the method of Solution 6, the presence of a flag having a value equal to the second value indicates that the identifier of the RAP image is omitted from the SEI message.

[0331] 9. According to the method of Solution 1, wherein a DRAP image is allowed to refer to an IRAP image or a previous image in the decoding order, wherein the previous image is a progressive decode refresh (GDR) image whose recovery point is equal to 0 in the output order.

[0332] 10. The method according to any one of solutions 1 to 9, wherein the bitstream is a general video codec bitstream.

[0333] 11. The method according to any one of solutions 1 to 10, wherein performing the conversion includes generating a bitstream from visual media data.

[0334] 12. The method according to any one of solutions 1 to 10, wherein performing the transformation includes reconstructing visual media data from the bitstream.

[0335] 13. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between the visual media data and a multi-layered bitstream of the visual media data according to a format rule; wherein the format rule specifies that Supplemental Enhancement Information (SEI) messages are included in the bitstream to indicate that a decoder is permitted to decode 1) a Dependent Random Access Point (DRAP) picture in the layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in the decoding and output order, without having to decode other pictures in the layer other than the Intra-Frame Random Access Point (IRAP) picture associated with the DRAP picture.

[0336] 14. The apparatus according to solution 13, wherein, in addition to IRAP images, DRAP images are excluded from valid entries in the reference image list of DRAP images in that layer.

[0337] 15. The apparatus according to solution 13, wherein a first image included in the layer and following the DRAP image in the decoding and output order is excluded from the valid entries of the reference image list of the first image, except for the IRAP image, from the second image included in the layer and preceding the DRAP image in the decoding and output order.

[0338] 16. The apparatus according to solution 13, wherein the bitstream is a universal video codec bitstream.

[0339] 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform conversion between visual media data and a multi-layered bitstream of visual media data according to a format rule; wherein the format rule specifies that Supplemental Enhancement Information (SEI) messages are included in the bitstream to instruct a decoder to decode 1) a Dependent Random Access Point (DRAP) picture in the layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in the decoding and output order, without having to decode other pictures in the layer besides the Intra-Frame Random Access Point (IRAP) picture associated with the DRAP picture.

[0340] 18. The non-transitory computer-readable storage medium according to Solution 17, wherein the bitstream is a general video codec bitstream.

[0341] 19. A non-transitory computer-readable storage medium for storing a bitstream of visual media data generated by a method performed by a visual media data processing apparatus, wherein the method comprises: determining that a Supplemental Enhancement Information (SEI) message is included in the bitstream, allowing a decoder to decode 1) a Dependent Random Access Point (DRAP) picture in a layer associated with the SEI message and / or 2) a picture included in the layer and following the DRAP picture in both decoding and output order, without having to decode other pictures in the layer besides the Intra-Frame Random Access Point (IRAP) picture associated with the DRAP picture; and generating the bitstream based on the determination.

[0342] 20. The non-transitory computer-readable storage medium according to Solution 19, wherein the bitstream is a universal video codec bitstream.

[0343] 21. A visual media data processing apparatus, comprising a processor configured to implement the method described in any one or more of solutions 1 to 12.

[0344] 22. A method for storing a bitstream of visual media data, comprising the method of any one of solutions 1 to 12, further comprising storing the bitstream to a non-transitory computer-readable storage medium.

[0345] 23. A computer-readable medium storing program code that, when executed, causes a processor to perform any one or more of the methods described in solutions 1 to 12.

[0346] 24. A computer-readable medium for storing a bitstream generated according to any of the methods described above.

[0347] 25. A visual media data processing apparatus for storing bit streams, wherein the visual media data processing apparatus is configured to implement the method described in any one or more of solutions 1 to 12.

[0348] 26. A computer-readable medium having a bit stream thereon conforming to the format rules according to any one of solutions 1 to 12.

[0349] The third set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., items 4 to 8).

[0350] 1. A method for processing visual media data (e.g., such as...) Figure 8 The method 800 shown includes: performing a conversion between visual media data and a bitstream of visual media data according to a format rule, wherein the format rule specifies whether and how a second type of SEI message, different from a first type of Supplemental Enhancement Information (SEI) message, is included in the bitstream, and wherein the first type of SEI message and the second type of SEI message respectively indicate a first type of Dependent Random Access Point (DRAP) image and a second type of DRAP image.

[0351] 2. The method according to Solution 1, wherein the format rule further specifies that the second type of SEI message includes a Random Access Point (RAP) image identifier.

[0352] 3. According to the method of Solution 1, for either the first type of DRAP image or the second type of DRAP image, the Random Access Point (RAP) image identifier is included in the bitstream.

[0353] 4. According to the method of Solution 3, the RAP image identifier is encoded as u(16), u(16) is an unsigned integer using 16 bits, or the RAP image identifier is encoded as ue(v), ue(v) is an unsigned integer using exponential Golomb code.

[0354] 5. The method according to Solution 1, wherein the format rule further specifies that the first type of SEI message or the second type of SEI message includes information about the Picture Order Count (POC) value of the first type of DRAP picture or the second type of DRAP picture.

[0355] 6. According to the method of Solution 1, the format rules also specify that each IRAP picture or DRAP picture is associated with a Random Access Point (RAP) picture identifier.

[0356] 7. According to the method of Solution 6, the format rules also stipulate that the value of the RAP image identifier of the IRAP image is inferred to be equal to 0.

[0357] 8. According to the method of Solution 6, the format rules also stipulate that the values ​​of the RAP picture identifiers of any two IRAP or DRAP pictures within the codec layer video sequence (CLVS) are different from each other.

[0358] 9. According to the method of Solution 6, the format rules also stipulate that the value of the RAP picture identifier of the IRAP or DRAP picture in the codec layer video sequence (CLVS) increases in the order of decoding the IRAP or DRAP pictures.

[0359] 10. According to the method of Solution 6, the format rules further specify that the value of the RAP picture identifier of the DRAP picture is one greater than the value of the previous IRAP picture or DRAP picture in the decoding order within the codec layer video sequence (CLVS).

[0360] 11. The method according to any one of solutions 1 to 10, wherein performing the conversion includes generating a bitstream from visual media data.

[0361] 12. The method according to any one of solutions 1 to 10, wherein performing the transformation includes reconstructing visual media data from the bitstream.

[0362] 13. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between visual media data and a bitstream of visual media data according to a format rule, wherein the format rule specifies whether and how a second type of SEI message, different from a first type of Supplemental Enhancement Information (SEI) message, is included in the bitstream, and wherein the first type of SEI message and the second type of SEI message respectively indicate a first type of Dependent Random Access Point (DRAP) image and a second type of DRAP image.

[0363] 14. The apparatus according to solution 13, wherein the format rule further specifies that the second type of SEI message includes a random access point (RAP) picture identifier.

[0364] 15. The apparatus according to solution 13, wherein for a first type of DRAP picture or a second type of DRAP picture, a random access point (RAP) picture identifier is included in the bitstream, the RAP picture identifier is encoded as u(16), u(16) is an unsigned integer using 16 bits, or the RAP picture identifier is encoded as ue(v), ue(v) is an unsigned integer using exponential Golomb code.

[0365] 16. The apparatus according to solution 13, wherein the format rule further specifies that each IRAP image or DRAP image is associated with a random access point (RAP) image identifier, and the format rule further specifies that the value of the RAP image identifier of the IRAP image is inferred to be equal to 0.

[0366] 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between visual media data and a bitstream of visual media data according to a format rule, wherein the format rule specifies whether and how a second type of SEI message, different from a first type of Supplemental Enhancement Information (SEI) message, is included in the bitstream, and wherein the first type of SEI message and the second type of SEI message respectively indicate a first type of Dependent Random Access Point (DRAP) image and a second type of DRAP image.

[0367] 18. A non-transitory computer-readable storage medium according to Solution 17, wherein the format rules further specify that the second type of SEI message includes a random access point (RAP) picture identifier, and wherein, for a first type of DRAP picture or a second type of DRAP picture, the random access point (RAP) picture identifier is included in the bitstream, the RAP picture identifier is encoded as u(16), which is an unsigned integer using 16 bits, or is encoded as ue(v), which is an unsigned integer using exponential Golomb codes, and wherein the format rules further specify that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier, and the format rules further specify that the value of the RAP picture identifier of the IRAP picture is inferred to be equal to 0.

[0368] 19. A non-transitory computer-readable storage medium for storing a bitstream of visual media data generated by a method performed by a visual media data processing apparatus, wherein the method includes determining whether and how a second type of SEI message, different from a first type of Supplemental Enhancement Information (SEI) message, is included in the bitstream; and generating the bitstream based on the determination.

[0369] 20. A non-transitory computer-readable storage medium according to Solution 19, wherein the format rules further specify that the second type of SEI message includes a random access point (RAP) picture identifier, and wherein, for a first type of DRAP picture or a second type of DRAP picture, the random access point (RAP) picture identifier is included in the bitstream, the RAP picture identifier is encoded as u(16), which is an unsigned integer using 16 bits, or is encoded as ue(v), which is an unsigned integer using exponential Golomb codes, and wherein the format rules further specify that each IRAP picture or DRAP picture is associated with a random access point (RAP) picture identifier, and the format rules further specify that the value of the RAP picture identifier of the IRAP picture is inferred to be equal to 0.

[0370] 21. A visual media data processing apparatus, comprising a processor configured to implement the method described in any one or more of solutions 1 to 12.

[0371] 22. A method for storing a bitstream of visual media data, comprising the method of any one of solutions 1 to 12, further comprising storing the bitstream to a non-transitory computer-readable storage medium.

[0372] 23. A computer-readable medium storing program code that, when executed, causes a processor to perform any one or more of the methods described in solutions 1 to 12.

[0373] 24. A computer-readable medium for storing a bitstream generated according to any of the methods described above.

[0374] 25. A visual media data processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions 1 to 12.

[0375] 26. A computer-readable medium having a bit stream thereon conforming to the format rules according to any one of solutions 1 to 12.

[0376] The fourth set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., items 6 and 9).

[0377] 1. A method for processing visual media data (e.g., such as...) Figure 9The method 900 shown includes: performing a conversion between visual media data and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a Supplemental Enhancement Information (SEI) message referring to a Dependent Random Access Point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating the number of Intra-IRAP (IRAP) pictures or Dependent Random Access Point (DRAP) pictures within the same codec layer video sequence (CLVS) as the DRAP picture.

[0378] 2. According to the method of Solution 1, IRAP images or DRAP images are allowed to be included in valid entries of the reference image list of DRAP images.

[0379] 3. According to the method of Solution 1, the syntax element is encoded as u(3), u(3) is an unsigned integer using 3 bits, or the syntax element is encoded as ue(v), ue(v) is an unsigned integer using exponential Golomb code.

[0380] 4. According to the method of Solution 1, the format rules further specify that the SEI message also includes a list of IRAP images and random access point (RAP) image identifiers of the DRAP images that are in the same codec layer video sequence (CLVS) as the DRAP images.

[0381] 5. According to the method of Solution 4, IRAP images or DRAP images are allowed to be included in valid entries of the reference image list of DRAP images.

[0382] 6. According to the method of Solution 4, each of the RAP image identifiers in the list is encoded as having the same RAP image identifier as the DRAP image associated with the SEI message.

[0383] 7. According to the method of Solution 4, wherein the identifiers in the list have values ​​corresponding to the i-th RAP image, i is equal to or greater than 0, and wherein the values ​​of the RAP image identifiers are incremented in ascending order of the value i.

[0384] 8. According to the method of Solution 7, each identifier in the list is encoded and decoded using the increment ue(v) between the value of the i-th DRAP image identifier and the value of the (i-1)-th DRAP image or IRAP image identifier, where i is greater than 0, or the increment ue(v) between i and 0, where i is equal to 0.

[0385] 9. According to the method of Solution 4, each identifier in the list is encoded and decoded to represent the Picture Order Count (POC) value of the RAP image.

[0386] 10. According to the method of Solution 4, each identifier in the list is encoded and decoded to represent POC increment information relative to the picture order count (POC) value of the IRAP picture associated with the SEI message.

[0387] 11. According to the method of Solution 4, each identifier in the list is encoded and decoded to represent the Picture Order Count (POC) value of the current picture and 1) the POC increment information between the POC value of the (i-1)th DRAP picture or IRAP picture, where i is greater than 0, or 2) the POC increment information between the POC values ​​of the IRAP picture associated with the SEI message.

[0388] 12. According to the method of Solution 4, wherein the list includes identifiers corresponding to the i-th RAP image and the j-th RAP image, where i is less than j, and wherein the i-th RAP image precedes the j-th RAP image in the decoding order.

[0389] 13. The method according to any one of solutions 1 to 12, wherein performing the conversion includes generating a bitstream from visual media data.

[0390] 14. The method according to any one of solutions 1 to 12, wherein performing the transformation includes reconstructing visual media data from the bitstream.

[0391] 15. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between visual media data and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a Supplemental Enhancement Information (SEI) message referring to a Dependent Random Access Point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating the number of Intra-Random Access Point (IRAP) pictures or Dependent Random Access Point (DRAP) pictures within the same codec layer video sequence (CLVS) as the DRAP picture.

[0392] 16. The apparatus according to solution 15, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of DRAP pictures, wherein the syntax element is encoded as u(3), u(3) is an unsigned integer using 3 bits, or the syntax element is encoded as ue(v), ue(v) is an unsigned integer using exponential Golomb codes, wherein the format rules further specify that the SEI message also includes a list of random access point (RAP) picture identifiers of IRAP pictures or DRAP pictures within the same codec layer video sequence (CLVS) as the DRAP picture, wherein the IRAP picture or the DRAP picture is allowed to be included in a valid entry of the reference picture list of DRAP pictures, and wherein each identifier in the list of RAP picture identifiers is encoded as being identical to the RAP picture identifier of the DRAP picture associated with the SEI message.

[0393] 17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between visual media data and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a Supplemental Enhancement Information (SEI) message referring to a Dependent Random Access Point (DRAP) picture is included in the bitstream, and wherein the format rule further specifies that the SEI message includes a syntax element indicating the number of Intra-IRAP (IRAP) pictures or Dependent Random Access Point (DRAP) pictures within the same codec layer video sequence (CLVS) as the DRAP picture.

[0394] 18. A non-transitory computer-readable storage medium according to Solution 17, wherein an IRAP picture or a DRAP picture is allowed to be included in a valid entry of a reference picture list of DRAP pictures, wherein the syntax element is encoded as u(3), u(3) is an unsigned integer using 3 bits, or the syntax element is encoded as ue(v), ue(v) is an unsigned integer using exponential Golomb codes, wherein the format rules further specify that the SEI message also includes a list of random access point (RAP) picture identifiers of IRAP pictures or DRAP pictures within the same codec layer video sequence (CLVS) as the DRAP picture, wherein the IRAP picture or the DRAP picture is allowed to be included in a valid entry of the reference picture list of DRAP pictures, and wherein each identifier in the list of RAP picture identifiers is encoded as being identical to the RAP picture identifier of the DRAP picture associated with the SEI message.

[0395] 19. A non-transitory computer-readable storage medium storing a bitstream of visual media data generated by a method performed by a visual media data processing apparatus, wherein the method includes: determining that a supplementary enhancement information (SEI) message referring to a dependent random access point (DRAP) picture is included in the bitstream; and generating the bitstream based on the determination.

[0396] 20. A non-transitory computer-readable storage medium according to Solution 19, wherein an IRAP picture or DRAP picture is allowed to be included in a valid entry of a reference picture list of DRAP pictures, wherein the syntax element is encoded as u(3), u(3) is an unsigned integer using 3 bits, or the syntax element is encoded as ue(v), ue(v) is an unsigned integer using exponential Golomb codes, wherein the format rules further specify that the SEI message also includes a list of random access point (RAP) picture identifiers of IRAP pictures or DRAP pictures within the same codec layer video sequence (CLVS) as the DRAP picture, wherein the IRAP picture or DRAP picture is allowed to be included in a valid entry of the reference picture list of DRAP pictures, and wherein each identifier in the list of RAP picture identifiers is encoded as being identical to the RAP picture identifier of the DRAP picture associated with the SEI message.

[0397] 21. A visual media data processing apparatus, comprising a processor configured to implement the method described in any one or more of solutions 1 to 14.

[0398] 22. A method for storing a bitstream of visual media data, comprising the method of any one of solutions 1 to 14, further comprising storing the bitstream to a non-transitory computer-readable storage medium.

[0399] 23. A computer-readable medium storing program code that, when executed, causes a processor to perform any one or more of the methods described in solutions 1 to 14.

[0400] 24. A computer-readable medium for storing a bitstream generated according to any of the methods described above.

[0401] 25. A visual media data processing apparatus for storing bitstreams, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions 1 to 14.

[0402] 26. A computer-readable medium having a bitstream thereon conforming to the format rules according to any one of solutions 1 to 14.

[0403] In the solution described in this paper, visual media data corresponds to video or images. In the solution described in this paper, the encoder can conform to the format rules by generating a codec representation based on those rules. In the solution described in this paper, the decoder can parse the syntax elements in the codec representation using the format rules, knowing whether or not syntax elements exist, to generate the decoded video.

[0404] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to its corresponding bitstream representation, and vice versa. The bitstream representation of the current video block can, for example, correspond to bits juxtaposed or scattered in different places within the bitstream, as defined by the syntax. For example, a macroblock can be encoded based on the error residual values ​​after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0405] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for use by a data processing apparatus to operate or control the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting machine-readable propagation signals, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an operating environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.

[0406] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code sections). Computer programs can be deployed to run on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected through a communications network.

[0407] The processes and logic described in this document can be executed by one or more programmable processors running one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0408] Processors suitable for running computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0409] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0410] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0411] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A method of processing visual media data, comprising: performing a conversion between visual media data and a bitstream of the visual media data comprising a plurality of layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and 2) a picture included in a same layer as the DRAP picture and following the DRAP picture in a decoding order and an output order without having to decode other pictures in the same layer except for an intra random access point (IRAP) picture associated with the DRAP picture, wherein the format rule further specifies that the SEI message comprises an identifier of a random access point (RAP) picture, wherein the format rule further specifies that a presence flag indicating a presence of the identifier of the RAP picture in the SEI message is included in the bitstream.

2. The method of claim 1, wherein, the DRAP picture excludes, from valid entries of a reference picture list of the DRAP picture, pictures in the same layer except for the IRAP picture.

3. The method of claim 1, wherein, a first picture included in the same layer and following the DRAP picture in the decoding order and the output order excludes, from valid entries of a reference picture list of the first picture, a second picture included in the same layer and preceding the DRAP picture in the decoding order and the output order except for the IRAP picture.

4. The method of claim 1, wherein, the RAP picture is the IRAP picture or the DRAP picture.

5. The method of claim 1, wherein, the presence flag having a value equal to a first value indicates that the identifier of the RAP picture is present in the SEI message.

6. The method of claim 1, wherein, the presence flag having a value equal to a second value indicates that the identifier of the RAP picture is omitted from the SEI message.

7. The method of any one of claims 1-3, wherein, the DRAP picture is allowed to refer to the IRAP picture or a previous picture in decoding order that is a gradual decoding refresh (GDR) picture whose decoding picture recovery point is equal to 0 in output order.

8. The method of any one of claims 1-3, wherein, the bitstream is a multi-functional video coding bitstream.

9. The method of any one of claims 1-3, wherein, the performing of the conversion comprises generating the bitstream from the visual media data.

10. The method of any one of claims 1-3, wherein, the performing of the conversion comprises reconstructing the visual media data from the bitstream.

11. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: perform a conversion between visual media data and a bitstream of the visual media data comprising a plurality of layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and 2) a picture included in a same layer as the DRAP picture and following the DRAP picture in a decoding order and an output order without having to decode other pictures in the same layer except for an intra random access point (IRAP) picture associated with the DRAP picture, wherein the format rule further specifies that the SEI message includes an identifier of a random access point (RAP) picture, wherein the format rule further specifies that a presence flag indicating a presence of the identifier of the RAP picture in the SEI message is included in the bitstream.

12. The apparatus of claim 11, wherein, The DRAP picture excludes, from valid entries of a reference picture list of the DRAP picture, pictures in the same layer other than the IRAP picture.

13. The apparatus of claim 11, wherein, A first picture included in the same layer and following the DRAP picture in the decoding order and the output order excludes, from valid entries of a reference picture list of the first picture, a second picture included in the same layer and preceding the DRAP picture in the decoding order and the output order other than the IRAP picture.

14. The apparatus of any of claims 11-13, wherein, The bitstream is a multi-functional video coding bitstream.

15. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between visual media data and a bitstream of the visual media data including multiple layers according to a format rule; wherein the format rule specifies that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and 2) a picture included in a same layer as the DRAP picture and following the DRAP picture in a decoding order and an output order without having to decode other pictures in the same layer other than an intra random access point (IRAP) picture associated with the DRAP picture, wherein the format rule further specifies that the SEI message includes an identifier of a random access point (RAP) picture, wherein the format rule further specifies that a presence flag indicating a presence of the identifier of the RAP picture in the SEI message is included in the bitstream.

16. The non-transitory computer-readable storage medium of claim 15, wherein, The bitstream is a multi-functional video coding bitstream.

17. A non-transitory computer-readable recording medium storing a bitstream of visual media data, the bitstream generated by a method performed by a visual media data processing apparatus, wherein, The method comprises: determining that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and 2) a picture included in a same layer as the DRAP picture and following the DRAP picture in a decoding order and an output order without having to decode other pictures in the same layer other than an intra random access point (IRAP) picture associated with the DRAP picture; and generating the bitstream based on the determining, wherein the SEI message includes an identifier of a random access point (RAP) picture, wherein a presence flag indicating a presence of the identifier of the RAP picture in the SEI message is included in the bitstream.

18. The non-transitory computer-readable storage medium of claim 17, wherein, The bitstream is a multi-functional video coding bitstream.

19. A method for storing a video bitstream, comprising: determining that a supplemental enhancement information (SEI) message is included in the bitstream to indicate that a decoder is allowed to decode 1) a dependent random access point (DRAP) picture in a layer associated with the SEI message and 2) a picture included in a same layer as the DRAP picture to which the DRAP picture belongs and following the DRAP picture in decoding order and output order, without having to decode other pictures in the same layer except for an intra random access point (IRAP) picture associated with the DRAP picture; and generating the bitstream based on the determining; storing the bitstream in a non-transitory computer-readable recording medium, wherein the SEI message includes an identifier of a random access point (RAP) picture, wherein a presence flag indicating a presence of the identifier of the RAP picture in the SEI message is included in the bitstream.

Citation Information

Patent Citations

  • Dependent random access point pictures

    WO2015192990A1