Constraints on reference picture lists

By redefining and standardizing the order and output order of sub-image types, the problem of complex and inefficient processing of sub-image types in the prior art is solved, and a more efficient and accurate video encoding and decoding process is achieved.

CN115462070BActive Publication Date: 2025-05-16DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180029943.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-20
Filing Date
2021-04-19
Publication Date
2025-05-16
Estimated Expiration
2041-04-19

AI Technical Summary

Technical Problem

Existing video encoding and decoding techniques have confusion and insufficient constraints when dealing with multiple sub-picture types and decoding sequences, resulting in complex and inefficient encoding and decoding processes.

Method used

By redefining and regulating the order and output order of sub-image types, new format rules and parameter sets are introduced to ensure correct decoding and output of sub-image types in single- and multi-layer contexts.

Benefits of technology

Improves the efficiency and accuracy of the video encoding and decoding process, simplifies the design of encoder and decoder, and reduces complexity and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115462070B_ABST
    Figure CN115462070B_ABST
Patent Text Reader

Abstract

Methods and apparatus for processing video are described. The processing may include video encoding, decoding, or transcoding. An example video processing method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying that, in response to the sub-picture not being a preceding sub-picture of an intra random access point sub-picture, the sub-picture cannot be a random access type sub-picture, and wherein the preceding sub-picture precedes the intra random access point sub-picture in output order.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on International Patent Application No. PCT / US2021 / 027963 filed on April 19, 2021, which claims priority and benefits to International Patent Application No. 63 / 012,713 filed on April 20, 2020. All of the above patent applications are hereby incorporated by reference in their entirety. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of video using various rules of a grammar.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying that, in response to the sub-picture not being a preceding sub-picture of an intra random access point sub-picture, the sub-picture cannot be a random access type sub-picture, and wherein the preceding sub-picture precedes the intra random access point sub-picture in output order.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including a plurality of sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying that a first sub-picture precedes a second sub-picture in an output order in a recovery point picture in response to the following: the first sub-picture and the second picture have the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index, and the first sub-picture precedes the second sub-picture in a decoding order.

[0008] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture containing a current slice, wherein the bitstream conforms to a format rule, wherein the format rule specifies an order in which pictures are indicated in the bitstream, wherein the format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture that precedes the current picture according to a second order according to a first order, wherein the second picture includes an intra random access point sub-picture having a layer identifier of a network abstraction unit (NAL) unit and a same sub-picture index as the current sub-picture, and wherein the current sub-picture is a fully random access sub-picture.

[0009] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture containing a current slice, wherein the bitstream conforms to a format rule, wherein the format rule specifies an order in which pictures are indicated in the bitstream, wherein the format rule does not allow an active entry in a reference picture list of the current slice to include a first picture that precedes a second picture according to a first order, wherein the second picture includes an intra random access point sub-picture having a layer identifier of a network abstraction unit (NAL) unit and a same sub-picture index as the current sub-picture, and wherein the current sub-picture is after the intra random access point sub-picture according to the second order.

[0010] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture containing a current slice, wherein the bitstream conforms to a format rule, wherein the format rule specifies an order in which pictures are indicated in the bitstream, wherein the format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture in a first order or a second order, wherein the second picture includes an intra random access point sub-picture having zero or more associated preceding sub-pictures and having a layer identifier of a network abstraction unit (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra random access point sub-picture and the zero or more associated preceding sub-pictures in the first order and the second order.

[0011] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture containing a current slice, wherein the bitstream conforms to a format rule, the format rule specifying that, in response to the current sub-picture being a random access decodable preceding sub-picture, active entries of a reference picture list of the current slice are not allowed to include any one or more of the following: a first picture including a random access skipped preceding sub-picture having the same sub-picture index as the current sub-picture, and a second picture preceding, in decoding order, a third picture including an intra random access point sub-picture associated with the random access decodable preceding sub-picture.

[0012] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video. The codec representation conforms to a format rule that specifies that the one or more pictures including the one or more sub-pictures are included in the codec representation according to a network abstraction layer (NAL) unit, wherein the type NAL unit indicated in the codec representation includes a codec slice of a particular type of picture or a codec slice of a particular type of sub-picture.

[0013] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies that two adjacent sub-pictures having different network abstraction layer unit types are to have the same indication of being treated as sub-pictures of a picture flag.

[0014] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule defining an order of a first type of sub-picture and a second type of sub-picture, wherein the first sub-picture is a post-sub-picture or a pre-sub-picture or a random access skip pre-sub (RASL) sub-picture type, and the second sub-picture is a RASL type or a random access decodable pre-sub (RADL) type or an instantaneous decoding refresh (IDR) type or a gradual decoding refresh (GDR) type sub-picture.

[0015] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule defining conditions for allowing or not allowing a sub-picture of a first type to appear with a sub-picture of a second type.

[0016] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more video pictures including one or more sub-pictures and a codec representation of the video; wherein the codec representation includes one or more layers of the video pictures in an order according to a rule.

[0017] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.

[0018] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.

[0019] In yet another exemplary aspect, a computer readable medium having stored thereon code is disclosed. The code is in the form of processor executable code embodying one of the methods described herein.

[0020] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.

[0022] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0023] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0024] Figure 4 A picture partitioned into 15 slices, 24 slices and 24 sub-pictures is shown.

[0025] Figure 5 is a block diagram of an example video processing system.

[0026] Figure 6 is a block diagram of a video processing device.

[0027] Figure 7 is a flow chart of an example method of video processing.

[0028] Figure 8 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0029] Fig. 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0030] Fig.10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0031] Figures 11 to 16 is a flow chart of an example method of video processing. DETAILED DESCRIPTION

[0032] Section headings are used in this document for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, H.266 technical terms are used in some descriptions only for ease of understanding and are not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes to text relative to the current draft of the VVC specification are indicated by left and right double square brackets (e.g., [[ ]]), deleted text between the double square brackets indicates canceled text, and Indicates the text that was added.

[0033] 1. Introduction

[0034] This document relates to video codec technology. Specifically, it is about the definition of sub-picture types and the relationship between different types of sub-pictures in single-layer and multi-layer contexts in terms of decoding order, output order, and prediction relationships. The key is to clearly specify the meaning of mixed sub-picture types within a picture through a set of constraints on decoding order, output order, and prediction relationships. The idea can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codec, such as the Versatile Video Codec (VVC) under development.

[0035] 2. Abbreviations

[0036] APS Adaptive Parameter Set

[0037] AU Access Unit

[0038] AUD Access Unit Delimiter

[0039] AVC Advanced Video Codec

[0040] CLVS Codec Layer Video Sequence

[0041] CPB codec picture buffer

[0042] CRA Completely Random Access

[0043] CTU Codec Tree Unit

[0044] CVS codec video sequence

[0045] DCI decoding capability information

[0046] DPB decoded picture buffer

[0047] EOB End of bitstream

[0048] EOS sequence end

[0049] GDR Gradual Decode Refresh

[0050] HEVC High Efficiency Video Codec

[0051] HRD Hypothesized Reference Decoder

[0052] IDR Instant Decode Refresh

[0053] JEM Joint Exploration Model

[0054] MCTS Motion Constraints Episode Set

[0055] NAL Network Abstraction Layer

[0056] OLS output layer set

[0057] PH Image Header

[0058] PPS Picture Parameter Set

[0059] PTL grades, tiers and levels

[0060] PU Picture Unit

[0061] RADL Random Access Decodable Preamble (Image)

[0062] RAP Random Access Point

[0063] RASL Random Access Skip Preamble (picture)

[0064] RBSP Raw Byte Sequence Payload

[0065] RPL Reference Image List

[0066] SEI auxiliary enhancement information

[0067] SPS Sequence Parameter Set

[0068] STSA Stepwise Temporal Sublayer Access

[0069] SVC Scalable Video Codec

[0070] VCL video codec layer

[0071] VPS Video Parameter Set

[0072] VTM VVC test model

[0073] VUI Video Availability Information

[0074] VVC Multi-functional Video Codec

[0075] 3. Preliminary Discussion

[0076] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec structure, in which temporal prediction plus transform codec is used. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and put them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held simultaneously every quarter, and the goal of the new codec standard is to reduce the bit rate by 50% compared to HEVC. The new video codec standard was officially named Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts on VVC standardization, new codec technologies are adopted into the VVC standard at each JVET meeting. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the July 2020 meeting.

[0077] 3.1. Picture segmentation scheme in HEVC

[0078] HEVC includes four different picture partitioning schemes, namely normal slice, dependent slice, tile, and wavefront parallel processing (WPP), which can be applied for maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.

[0079] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although there may still be interdependencies due to loop filtering operations).

[0080] Normal slices are the only tool that can be used for parallelization, which is also available in H.264 / AVC in almost the same form. Parallelization based on normal slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive codec pictures, which is generally much more difficult than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, the use of normal slices may lead to a large codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. In addition, due to the intra-picture independence of normal slices and the fact that each normal slice is encapsulated in its own NAL unit, normal slices (in contrast to the other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting requirements on the layout of slices in a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.

[0081] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Basically, dependent slices provide segmentation of regular slices into multiple NAL units to provide reduced end-to-end latency by allowing a portion of a regular slice to be transmitted before the encoding of the entire regular slice is completed.

[0082] In WPP, a picture is partitioned into a single row of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding of CTB rows is delayed by two CTBs, thereby ensuring that data related to the CTBs above and to the right of the main CTB is available before the main CTB is decoded. Using this staggered start (which looks like a wavefront when represented graphically), parallelization is possible using as many processors / cores as the picture contains CTB rows. Because intra-picture prediction between adjacent tree block rows within a picture is allowed, the required inter-frame processor / inter-frame core communication to implement intra-picture prediction may be large. WPP segmentation does not result in the generation of additional NAL units compared to when it is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, conventional slices can be used with WPP, with certain encoding and decoding overhead.

[0083] Slices define the horizontal and vertical boundaries that divide the image into slice columns and rows. Slice columns extend from the top of the image to the bottom of the image. Similarly, slice rows extend from the left side of the image to the right side of the image. The number of slices in an image can be simply derived as the number of slice columns multiplied by the number of slice rows.

[0084] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of the CTB raster scan of the slice). Similar to conventional slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in separate NAL units (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header when the slice spans more than one slice, and loop filtering related to sharing of reconstruction samples and metadata. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first is signaled in the slice header.

[0085] For simplicity, restrictions on the application of four different picture partitioning schemes have been specified in HEVC. A given codec video sequence cannot simultaneously include slices and wavefronts from most of the profiles specified in HEVC. For each slice and slice, one or both of the following conditions must be met: 1) all codec tree blocks in a slice belong to the same slice; 2) all codec tree blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is used, if a slice starts inside a CTB row, it must end in the same CTB row.

[0086] The latest amendment to HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), "HEVC Additional Supplemental Enhancement Information (Draft 4)", dated October 24, 2017, publicly available here: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. By including this amendment, HEVC specifies three MCTS-related SEI messages, namely, the time-domain MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nested SEI message.

[0087] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vector is restricted to point to the full sample position within the MCTS and the fractional sample position that only needs the full sample position within the MCTS for interpolation, and the motion vector candidate derived from the block outside the MCTS for temporal motion vector prediction is not allowed. In this way, each MCTS can be decoded independently without the slices not included in the MCTS.

[0088] The MCTS extraction information set SEI message provides auxiliary information (specified as part of the semantics of the SEI message) that can be used in the MCTS sub-bitstream extraction to generate a conformant bitstream for the MCTS set. The information consists of multiple extraction information sets, each defining multiple MCTS sets and containing RBSP bytes to replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.

[0089] 3.2. Image Segmentation in VVC

[0090] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​a picture. The CTUs in a slice are scanned in raster scan order within the slice.

[0091] A slice consists of an integer number of consecutive complete CTU rows within a slice of an integer number of complete slices or a picture.

[0092] Two strip modes are supported, namely raster scan strip mode and rectangular strip mode. In raster scan strip mode, a strip contains a complete sequence of strips in a slice raster scan of a picture. In rectangular strip mode, a strip contains multiple complete slices that together form a rectangular area of ​​a picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of ​​a picture. The slices within a rectangular strip are scanned in a slice raster scan order within the rectangular area corresponding to the strip.

[0093] A sub-picture consists of one or more strips that together cover a rectangular area of ​​the picture.

[0094] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.

[0095] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0096] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0097] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, the 12 slices on the left hand side each cover one stripe of a 4×4 CTU, and the 6 slices on the right hand side each cover 2 vertically stacked stripes of a 2×2 CTU, resulting in a total of 24 stripes and 24 sub-pictures of different dimensions (each stripe is a sub-picture).

[0098] 3.3. Image resolution changes within a sequence

[0099] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC enables picture resolution changes within a sequence at locations where IRAP pictures are not encoded, which are always intra-coded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling of reference pictures used for inter prediction when the reference picture has a different resolution than the current picture being decoded.

[0100] The scaling ratio is restricted to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three resampling filter sets with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three resampling filter sets are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each resampling filter set has 16 phases for luminance and 32 phases for chrominance, which is the same as the case of motion compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process, in which the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.

[0101] Other aspects of the VVC design that support this feature that differ from HEVC include: i) picture resolutions and corresponding consistency windows are signaled in the PPS rather than in the SPS, where the maximum picture resolution is signaled. ii) For a single-layer bitstream, each picture store (variable slot in the DPB for storing one decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.

[0102] 3.4. Scalable Video Codec (SVC) in General and VVC

[0103] Scalable video codec (SVC, sometimes also referred to as scalability in video codecs) refers to video codecs that use a base layer (BL), sometimes referred to as a reference layer (RL), and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously coded layers. For example, the bottom layer can serve as a BL, while the top layer can serve as an EL. Intermediate layers can serve as an EL or RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest layer nor the highest layer) can be the EL of a layer below the intermediate layer (such as a base layer or any intervening enhancement layer), and at the same time serve as the RL of one or more enhancement layers above the intermediate layer. Similarly, in a multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to encode (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0104] In SVC, parameters used by an encoder or decoder are grouped into parameter sets based on the codec level (e.g., video level, sequence level, picture level, slice level, etc.) at which they can be utilized. For example, parameters that can be utilized by one or more codec video sequences of different layers in a bitstream can be included in a video parameter set (VPS), and parameters that can be utilized by one or more pictures in a codec video sequence can be included in a sequence parameter set (SPS). Similarly, parameters utilized by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters specific to a single slice can be included in a slice header. Similarly, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.

[0105] Due to the support for reference picture resampling (RPR) in VVC, support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) can be designed without any additional signal processing level codec tools, because the upsampling required for spatial scalability support can use only RPR upsampling filters. However, for scalability support, high-level syntax changes are required (compared to not supporting scalability). Scalability support is specified in VVC version 1. Unlike the scalability support in any earlier video codec standards (including extensions of AVC and HEVC), the design of VVC scalability has been as friendly to single-layer decoder design as possible. The decoding capability of multi-layer bitstreams is specified as if there is only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a manner independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not need too many changes to be able to decode multi-layer bitstreams. Compared with the design of multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, the IRAP AU is required to contain pictures of each layer present in the CVS.

[0106] 3.5. Random access and support in HEVC and VVC

[0107] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture in the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include closely spaced random access points, which are typically intra-codec pictures, but can also be inter-codec pictures (for example, in the case of gradual decoding refresh).

[0108] HEVC includes signaling of intra-frame random access point (IRAP) pictures in the NAL unit header via the NAL unit type. Three types of IRAP pictures are supported, namely instant decoder refresh (IDR), complete random access (CRA), and broken link access (BLA) pictures. IDR pictures constrain the inter-frame picture prediction structure to not reference any pictures before the current picture group (GOP), which is traditionally referred to as a closed GOP random access point. CRA pictures are less restricted by allowing specific pictures to reference pictures before the current GOP, where all pictures are discarded in the case of random access. CRA pictures are traditionally referred to as open GOP random access points. BLA pictures typically originate from the splicing of two bitstreams or parts thereof at a CRA picture, such as during stream switching. In order to better enable the system to use IRAP pictures, a total of six different NAL units are defined to signal the properties of IRAP pictures, which can be used to better match the stream access point type defined in the ISO Base Media File Format (ISOBMFF) for random access support in Dynamic Adaptive Streaming over HTTP (DASH).

[0109] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or without an associated RADL picture), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type in HEVC is not included in VVC for two main reasons: i) The basic functionality of a BLA picture can be implemented by a CRA picture plus a sequence NAL unit end, the presence of which indicates that subsequent pictures start a new CVS in a single-layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than in HEVC, as indicated by the use of 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.

[0110] Another key difference between VVC and HEVC in random access support is that GDR is supported in a more standardized way in VVC. In GDR, decoding of the bitstream can start from an inter-frame coded picture, although not the entire picture area can be correctly decoded at the beginning, but after multiple pictures, the entire picture area will be correct. GDR random access points and recovery points are signaled using the recovery point SEI message, and AVC and HEVC also support GDR. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. CVS and bitstreams are allowed to start with GDR pictures. This means that the entire bitstream is allowed to contain only inter-frame coded pictures, without a single intra-frame coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded strips or blocks across multiple pictures, as opposed to intra-coding the entire picture, allowing for significant end-to-end latency reduction, which is considered more important today than ever before as ultra-low latency applications such as wireless displays, online gaming, and drone-based applications become more popular.

[0111] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refresh area (i.e., correctly decoded area) and the unrefreshed area at the picture between the GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so there will be no decoding mismatch of some samples at or near the boundary. This can be useful when the application determines to display the correctly decoded area during the GDR process.

[0112] IRAP pictures and GDR pictures may be collectively referred to as random access point (RAP) pictures.

[0113] 3.6. Reference Picture Management and Reference Picture List (RPL)

[0114] Reference picture management is a core functionality required for any video codec using inter prediction. It manages the storage and removal of reference pictures in the decoded picture buffer (DPB) and places the reference pictures in their correct order in the RPL.

[0115] HEVC's reference picture management, including reference picture marking and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC), is different from AVC. Instead of the reference picture marking mechanism based on sliding window plus adaptive memory management control operation (MMCO) in AVC, HEVC specifies a reference picture management and marking mechanism based on the so-called reference picture set (RPS), and RPLC is therefore based on the RPS mechanism. The RPS consists of a reference picture set associated with a picture, which consists of all reference pictures before the associated picture in decoding order, which can be used for inter-frame prediction of the associated picture or any picture after the associated picture in decoding order. The reference picture set consists of five reference picture lists. The first three lists contain all reference pictures that can be used for inter-frame prediction of the current picture and can be used for inter-frame prediction of one or more pictures after the current picture in decoding order. The other two lists consist of all reference pictures that are not used for inter-frame prediction of the current picture but can be used for inter-frame prediction of one or more pictures after the current picture in decoding order. RPS provides "intra-frame codec" signaling of DPB status, rather than "inter-frame codec" signaling as in AVC, mainly to improve error tolerance. The RPLC process in HEVC is based on RPS, and the index of the RPS subset is notified by signaling for each reference index; this process is simpler than the RPLC process in AVC.

[0116] Reference picture management in VVC is more similar to HEVC than AVC, but slightly simpler and more robust. As in those standards, two RPLs are derived, List 0 and List 1, but they are not based on the reference picture set concept used in HEVC or the automatic sliding window process used in AVC; instead, they are signaled more directly. Reference pictures are listed for the RPL as active and inactive entries, and only active entries can be used as reference indexes in inter prediction of CTUs of the current picture. Inactive entries indicate other pictures to be saved in the DPB for reference by other pictures arriving later in the bitstream.

[0117] 3.7. Parameter Set

[0118] AVC, HEVC, and VVC specify parameter sets. Types of parameter sets include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced from HEVC and is included in HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.

[0119] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, information that does not change frequently does not need to be repeated for each sequence or picture, so redundant signaling of that information can be avoided. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission, but also improving fault tolerance.

[0120] VPS is introduced to carry sequence level header information common to all layers in a multi-layer bitstream.

[0121] APS is introduced to carry such picture-level or slice-level information, which requires quite a lot of bits to encode and decode, can be shared by multiple pictures, and can have quite a lot of different variations in a sequence.

[0122] 3.8. Related definitions in VVC

[0123] The relevant definitions in the latest VVC text (JVET-Q2001-vE / v15) are as follows.

[0124] Associated IRAP picture (of a specific picture): The previous IRAP picture in decoding order (when present) has the same nuh_layer_id value as the specific picture.

[0125] Completely random access (CRA) PU: The PU of the coded picture is a CRA picture.

[0126] Completely random access (CRA) picture: An IRAP picture with nal_unit_type equal to CRA_NUT per VCL NAL unit.

[0127] Coded Video Sequence (CVS): a sequence of AUs consisting of CVSS AUs, followed by zero or more AUs that are not CVSS AUs, in decoding order, including all subsequent AUs up to but not including any subsequent AU that is a CVSS AU.

[0128] Codec Video Sequence Start (CVSS) AU: An AU in which there is a PU for each layer in CVS and the coded picture in each PU is a CLVSS picture.

[0129] Gradual Decoding Refresh (GDR) AU: An AU in which every coded picture present in a PU is a GDR picture.

[0130] Gradual Decoding Refresh (GDR) PU: The PU of the coded picture is a GDR picture.

[0131] Gradual Decoding Refresh (GDR) picture: A picture where each VCL NAL unit has nal_unit_type equal to GDR_NUT.

[0132] Instantaneous Decoding Refresh (IDR) PU: A PU whose coded picture is an IDR picture.

[0133] Instantaneous Decoding Refresh (IDR) picture: An IRAP picture with nal_unit_type equal to IDR_W_RADL or IDR_N_LP per VCL NAL unit.

[0134] Intra Random Access Point (IRAP) AU: An AU in which there is a PU for each layer in the CVS and the codec picture in each PU is an IRAP picture.

[0135] Intra Random Access Point (IRAP) PU: A PU whose codec picture is an IRAP picture.

[0136] Intra random access point (IRAP) picture: A codec picture in which all VCL NAL units have the same value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive.

[0137] Preceding picture: A picture that is in the same layer as the associated IRAP picture and precedes the associated IRAP picture in output order.

[0138] Output order: The order in which decoded pictures are output from the DPB (for decoded pictures to be output from the DPB).

[0139] Random Access Decodable Front (RADL) PU: A PU whose coded picture is a RADL picture.

[0140] Random Access Decodable Leading (RADL) picture: A picture with nal_unit_type equal to RADL_NUT per VCL NAL unit.

[0141] Random Access Skip Leading (RASL) PU: The PU of the coded picture is a RASL picture.

[0142] Random Access Skip Leading (RASL) picture: A picture with nal_unit_type equal to RASL_NUT per VCL NAL unit.

[0143] Step-by-step temporal sub-layer access (STSA) PU: A PU whose codec picture is a STSA picture.

[0144] Step-by-Step Temporal Sub-Layer Access (STSA) picture: A picture with nal_unit_type equal to STSA_NUT per VCL NAL unit.

[0145] NOTE – STSA pictures do not use pictures with the same TemporalId as the STSA picture for inter prediction reference. Pictures that follow the STSA picture in decoding order with the same TemporalId as the STSA picture do not use pictures that precede the STSA picture in decoding order with the same TemporalId as the STSA picture for inter prediction reference. STSA pictures implement an upward switch from the immediately lower sublayer to the sublayer containing the STSA picture at the STSA picture. STSA pictures must have a TemporalId greater than 0.

[0146] Sub-picture: A rectangular area of ​​one or more strips within a picture.

[0147] Posting picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not a STSA picture.

[0148] NOTE – Posting pictures associated with an IRAP picture also follow the IRAP picture in decoding order. Pictures that follow the associated IRAP picture in output order and precede the associated IRAP picture in decoding order are not allowed.

[0149] 3.9. NAL unit header syntax and semantics in VVC

[0150] In the latest VVC text (in JVET-Q2001-vE / v15), the NAL unit header syntax and semantics are as follows.

[0151] 7.3.1.2 NAL unit header syntax

[0152]

[0153] 7.4.2.2 NAL unit header semantics

[0154] forbidden_zero_bit should be equal to 0.

[0155] nuh_reserved_zero_bit shall be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified by ITU-T|ISO / IEC in the future. A decoder shall ignore (ie, remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1.

[0156] nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which the non-VCL NAL unit applies. The value of nuh_layer_id shall be in the range of 0 to 55, inclusive. Other values ​​of nuh_layer_id are reserved for future use by ITU-T | ISO / IEC.

[0157] The value of nuh_layer_id shall be the same for all VCL NAL units of a codec picture.The value of nuh_layer_id of a codec picture or PU is the value of nuh_layer_id of the VCL NAL unit of the codec picture or PU.

[0158] The value of nuh_layer_id for AUD, PH, EOS, and FD NAL units is constrained as follows:

[0159] - If nal_unit_type is equal to AUD_NUT, nuh_layer_id shall be equal to vps_layer_id[0].

[0160] Otherwise, when nal_unit_type is equal to PH_NUT, EOS_NUT or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit.

[0161] NOTE 1 – The value of nuh_layer_id for DCI, VPS and EOB NAL units is not constrained.

[0162] The value of nal_unit_type should be the same for all pictures of a CVSS AU.

[0163] nal_unit_type specifies the NAL unit type, ie, the type of RBSP data structure contained in the NAL unit as specified in Table 5.

[0164] NAL units with nal_unit_type in the range UNSPEC_28..UNSPEC_31 (unspecified semantics), inclusive, shall not affect the decoding process as specified in this specification.

[0165] NOTE 2 – NAL unit types in the range of UNSPEC_28..UNSPEC_31 may be used as determined by the application. The decoding process for these values ​​of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, special care must be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define any management of these values. These nal_unit_type values ​​may be appropriate for use only in contexts where "conflicts" in usage (i.e., different definitions of the meaning of the content of NAL units for the same nal_unit_type value) are not important, or are not possible, or are managed (e.g., defined or managed in a controlling application or transport specification, or by an environment where the control bitstream is distributed).

[0166] For purposes other than determining the amount of data in a DU of a bitstream (as specified in Annex C), a decoder should ignore (remove from the bitstream and discard) the contents of all NAL units that use reserved values ​​of nal_unit_type.

[0167] NOTE 3 – This requirement allows for the future definition of compatible extensions to this specification.

[0168] Table 5 – NAL unit type codes and NAL unit type classifications

[0169]

[0170]

[0171] NOTE 4 - A completely random access (CRA) picture may have an associated RASL or RADL picture present in the bitstream.

[0172] NOTE 5 - An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP does not have an associated preceding picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream.

[0173] The value of nal_unit_type shall be the same for all VCL NAL units of a sub-picture. A sub-picture is said to have the same NAL unit type as the VCL NAL units of the sub-picture.

[0174] For any particular picture's VCL NAL unit, the following applies:

[0175] - If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and a picture or PU is said to have the same NAL unit type as the codec slice NAL units of the picture or PU.

[0176] Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture shall have at least two sub-pictures, and the VCL NAL units of the picture shall have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one sub-picture of the picture shall all have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of the other sub-pictures in the picture shall all have different specific values ​​of nal_unit_type equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.

[0177] For single-layer bitstreams, the following constraints apply:

[0178] - Every picture except the first picture in the bitstream in decoding order is considered to be associated with the previous IRAP picture in decoding order.

[0179] - When a picture is the preceding picture of an IRAP picture, it shall be a RADL or RASL picture.

[0180] - When a picture is a post-IRAP picture, it shall not be a RADL or RASL picture.

[0181] - There shall be no RASL pictures associated with an IDR picture in the bitstream.

[0182] - There shall be no RADL pictures associated with an IDR picture with nal_unit_type equal to IDR_N_LP in the bitstream.

[0183] NOTE 6 – Random access can be performed at the location of an IRAP PU (and correctly decode the IRAP picture and all subsequent non-RASL pictures in decoding order) by discarding all PUs preceding the IRAP PU, provided that each parameter set is available at the time it is referenced (either in the bitstream or by external means not specified in this specification).

[0184] - Any picture that precedes an IRAP picture in decoding order shall precede the IRAP picture in output order, and shall precede any RADL pictures associated with the IRAP picture in output order.

[0185] - Any RASL pictures associated with a CRA picture shall precede any RADL pictures associated with the CRA picture in output order.

[0186] - Any RASL pictures associated with a CRA picture shall follow, in output order, any IRAP pictures that precede the CRA picture in decoding order.

[0187] - If field_seq_flag is equal to 0, and the current picture is a leading picture associated with an IRAP picture, then it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, let picA and picB be the first and last leading pictures associated with the IRAP picture in decoding order, respectively, there shall be at most one non-leading picture before picA in decoding order, and there shall be no non-leading pictures between picA and picB in decoding order.

[0188] nuh_temporal_id_plus1 minus 1 specifies the temporal identifier of the NAL unit.

[0189] The value of nuh_temporal_id_plus1 shall not be equal to 0.

[0190] The variable TemporalId is derived as follows:

[0191] TemporalId=nuh_temporal_id_plus1-1 (36)

[0192] When nal_unit_type is in the range of IDR_W_RADL to RSV_IRAP_12 (including IDR_W_RADL and RSV_IRAP_12), TemporalId shall be equal to 0.

[0193] When nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, TemporalId shall not be equal to 0.

[0194] The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId of a codec picture, PU, ​​or AU is the value of TemporalId of the VCL NAL unit of the codec picture, PU, ​​or AU. The value of TemporalId of a sublayer representation is the maximum value of the TemporalId of all VCL NAL units in the sublayer representation.

[0195] The value of TemporalId for non-VCL NAL units is constrained as follows:

[0196] - If nal_unit_type is equal to DCI_NUT, VPS_NUT or SPS_NUT, TemporalId shall be equal to 0, and the TemporalId of the AU containing the NAL unit shall be equal to 0.

[0197] - Otherwise, if nal_unit_type is equal to PH_NUT, TemporalId shall be equal to the TemporalId of the PU containing the NAL unit.

[0198] - Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0.

[0199] - Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT or SUFFIX_SEI_NUT, TemporalId shall be equal to the TemporalId of the AU containing the NAL unit.

[0200] - Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU containing the NAL unit.

[0201] NOTE 7 – When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values ​​of all AUs to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing AU, since all PPS and APS may be included at the beginning of the bitstream (e.g., when they are transmitted out-of-band and the receiver places them at the beginning of the bitstream), with the first codec picture having TemporalId equal to 0.

[0202] 3.10. Mixed NAL unit types within a picture

[0203] 7.4.3.4 Picture parameter set semantics ...

[0204] mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture of the referenced PPS has more than one VCL NAL unit, the VCL NAL units do not have the same value of nal_unit_type, and the picture is not an IRAP picture. mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same value of nal_unit_type.

[0205] When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_types_in_pic_flag shall be equal to 0.

[0206] For each slice with a nal_unit_type value nalUnitTypeA in the range IDR_W_RADL to CRA_NUT (inclusive), in a picture picA that also contains one or more slices with another value of nal_unit_type (i.e., the value of mixed_nalu_types_in_pic_flag of picture picA is equal to 1), the following applies:

[0207] - The slice shall belong to a sub-picture subpicA for which the value of subpic_treated_as_pic_flag[i] is equal to 1.

[0208] - A slice shall not belong to a sub-picture picA that contains VCL NAL units whose nal_unit_type is not equal to nalUnitTypeA.

[0209] - If nalUnitTypeA is equal to CRA, then for all subsequent PUs that follow the current picture in CLVS in decoding order and in output order, neither RefPicList[0] nor RefPicList[1] of the slices in subpicA in these PUs shall include any pictures in the active entry that precede picA in decoding order.

[0210] - Otherwise (i.e., nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS that follow the current picture in decoding order, neither RefPicList[0] nor RefPicList[1] of the slices in subpicA in these PUs should include any pictures in the active entry that precede picA in decoding order.

[0211] NOTE 1 – mixed_nalu_types_in_pic_flag equal to 1 indicates that the picture referencing the PPS contains slices with different NAL unit types, e.g., a codec picture resulting from a sub-picture bitstream merge operation where the encoder must ensure matching bitstream structure and further alignment of parameters of the original bitstream. An example of such alignment is as follows: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_in_pic_flag is equal to 1, the picture referencing the PPS shall not have slices with nal_unit_type equal to IDR_W_RADL or IDR_N_LP.

[0212] 3.11. Picture Header Structure Syntax and Semantics in VVC

[0213] In the latest VVC text (in JVET-Q2001-vE / v15), the picture header structure syntax and semantics most relevant to the present invention in this context are as follows.

[0214] 7.3.2.7 Picture header structure syntax

[0215]

[0216]

[0217] 7.4.3.7 Image header structure semantics

[0218] The PH syntax structure contains information common to all slices of the coded picture associated with the PH syntax structure.

[0219] gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag equal to 0 specifies that the current picture may or may not be a GDR or IRAP picture.

[0220] gdr_pic_flag equal to 1 specifies that the picture associated with the PH is a GDR picture. gdr_pic_flag equal to 0 specifies that the picture associated with the PH is not a GDR picture. When not present, the value of gdr_pic_flag is inferred to be equal to 0. When gdr_enabled_flag is equal to 0, the value of gdr_pic_flag shall be equal to 0.

[0221] NOTE 1 – When gdr_or_irap_pic_flag is equal to 1 and gdr_pic_flag is equal to 0, the picture associated with the PH is an IRAP picture. ...

[0222] ph_pic_order_cnt_lsb specifies the picture order count of the current picture modulo MaxPicOrderCntLsb. The length of the ph_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits. The value of ph_pic_order_cnt_lsb shall be in the range of 0 to MaxPicOrderCntLsb-1, inclusive.

[0223] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB as specified in Annex C after decoding a CLVSS picture that is not the first picture in the bitstream.

[0224] recovery_poc_cnt specifies the recovery point of the decoded picture in the output order. If the current picture is a GDR picture associated with PH, and there is a picture picA that follows the current GDR picture in the decoding order in the CLVS with PicOrderCntVal equal to the value of PicOrderCntVal of the current GDR picture plus recovery_poc_cnt, then the picture picA is called the recovery point picture. Otherwise, the first picture in the output order with a PicOrderCntVal greater than the value of PicOrderCntVal of the current picture plus recovery_poc_cnt is called the recovery point picture. The recovery point picture should not be before the current GDR picture in the decoding order. The value of recovery_poc_cnt should be in the range of 0 to MaxPicOrderCntLsb-1 (including 0 and MaxPicOrderCntLsb-1).

[0225] When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:

[0226] RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81)

[0227] NOTE 2 - When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the RpPicOrderCntVal of the associated GDR picture, the current decoded picture and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (when present) that precedes the associated GDR picture in decoding order. ...

[0228] 3.12. Constraints on RPL in VVC

[0229] In the latest VVC text (in JVET-Q2001-vE / v15), the constraints on RPL in VVC are as follows (as part of the decoding process of clause 8.3.2 Reference picture list construction of VVC).

[0230] 8.3.2 Decoding process of reference picture list construction ...

[0231] For each i equal to 0 or 1, the first NumRefIdxActive[i] entries in RefPicList[i] are referred to as active entries in RefPicList[i], and the other entries in RefPicList[i] are referred to as inactive entries in RefPicList[i].

[0232] NOTE 2 – A particular picture may be referenced by an entry in RefPicList[0] and an entry in RefPicList[1]. A particular picture may also be referenced by more than one entry in RefPicList[0] or more than one entry in RefPicList[1].

[0233] NOTE 3 – The active entries in RefPicList[0] and the active entries in RefPicList[1] jointly reference all reference pictures that can be used for inter prediction of the current picture and one or more pictures after the current picture in decoding order. The inactive entries in RefPicList[0] and the inactive entries in RefPicList[1] jointly reference all reference pictures that are not used for inter prediction of the current picture but can be used for inter prediction of one or more pictures after the current picture in decoding order.

[0234] NOTE 4 – There may be one or more entries in RefPicList[0] or RefPicList[1] equal to “no reference picture” because the corresponding pictures are not present in the DPB. Each inactive entry in RefPicList[0] or RefPicList[0] equal to “no reference picture” should be ignored. An unintentional picture loss should be inferred for each active entry in RefPicList[0] or RefPicList[1] equal to “no reference picture”.

[0235] The requirement for bitstream conformance is that the following constraints apply:

[0236] - For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i].

[0237] - The picture referenced by each active entry in RefPicList[0] or RefPicList[1] shall be present in the DPB and shall have a TemporalId less than or equal to the TemporalId of the current picture.

[0238] - Each entry in RefPicList[0] or RefPicList[1] shall refer to a picture that is not the current picture and shall have non_reference_picture_flag equal to 0.

[0239] - A STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and a LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture.

[0240] - There shall be no difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry in RefPicList[0] or RefPicList[1] greater than or equal to 2 24 LTRP entry.

[0241] - Let setOfRefPics be the set of unique pictures referenced by all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize-1 (inclusive), where MaxDpbSize is as specified in clause A.4.2, and setOfRefPics shall be the same for all slices of the picture.

[0242] - When the current slice has nal_unit_type equal to STSA_NUT, there shall be no active entry in RefPicList[0] or RefPicList[1] with TemporalId equal to the current picture's TemporalId and nuh_layer_id equal to the current picture's nuh_layer_id.

[0243] - When the current picture is a picture after the STSA picture in decoding order with TemporalId equal to the current picture's TemporalId and nuh_layer_id equal to the current picture's nuh_layer_id, the picture before the STSA picture in decoding order with TemporalId equal to the current picture's TemporalId and nuh_layer_id equal to the current picture shall not be included as the active entry in RefPicList[0] or RefPicList[1].

[0244] - When the current picture is a CRA picture, there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede any preceding IRAP picture in decoding order (when present), in either output order or decoding order.

[0245] - When the current picture is a post-picture, there shall be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that were generated by the decoding process for generating unavailable reference pictures for the IRAP picture associated with the current picture.

[0246] - When the current picture is a subsequent picture that follows one or more preceding pictures (if any) associated with the same IRAP picture in both decoding order and output order, there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that were generated by the decoding process for generating unavailable reference pictures for the IRAP picture associated with the current picture.

[0247] - When the current picture is a recovery point picture or a picture following the recovery point picture in output order, there should be no entries in RefPicList[0] or RefPicList[1] that contain pictures generated by the decoding process of the unavailable reference pictures of the GDR picture used to generate the recovery point picture.

[0248] - When the current picture is a subsequent picture, there shall be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.

[0249] - When the current picture is a subsequent picture that follows one or more preceding pictures (if any) associated with the same IRAP picture in both decoding order and output order, there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.

[0250] - When the current picture is a RADL picture, there shall be no active entries in RefPicList[0] or RefPicList[1] that are any of the following:

[0251] o RASL images

[0252] oPictures generated by the decoding process used to generate unavailable reference pictures

[0253] o The picture that precedes the associated IRAP picture in decoding order

[0254] - The picture referenced by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be in the same AU as the current picture.

[0255] - The picture referenced by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB and shall have a nuh_layer_id less than the nuh_layer_id of the current picture.

[0256] - Each ILRP entry in either RefPicList[0] or RefPicList[1] of a stripe shall be an active entry. ...

[0257] 3.13.PictureOutputFlag settings

[0258] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting the value of the variable PictureOutputFlag is as follows (as part of clause 8.1.2 decoding process for codec pictures).

[0259] 8.1.2 Decoding process of coded images

[0260] The decoding process specified in this clause applies to each coded picture in BitstreamToDecode (called the current picture and represented by the variable CurrPic).

[0261] Depending on the value of chroma_format_idc, the number of sample arrays of the current picture is as follows:

[0262] - If chroma_format_idc is equal to 0, the current picture consists of 1 sample array S L composition.

[0263] - Otherwise (chroma_format_idc is not equal to 0), the current picture consists of 3 sample arrays S L , S Cb , S Cr composition.

[0264] The decoding process of the current picture takes as input the syntax elements and uppercase variables from clause 7. When explaining the semantics of each syntax element in each NAL unit and in the rest of clause 8, the term "bitstream" (or a portion thereof, such as a CVS of a bitstream) refers to BitstreamToDecode (or a portion thereof).

[0265] Depending on the value of separate_colour_plane_flag, the decoding process is structured as follows:

[0266] - If separate_colour_plane_flag is equal to 0, the decoding process is called once with the current picture as output.

[0267] - Otherwise (separate_colour_plane_flag is equal to 1), the decoding process is called three times. The input to the decoding process is all NAL units of the coded picture with the same value of colour_plane_id. The decoding process for NAL units with a particular value of colour_plane_id is specified as if only CVSs in monochrome color format with that particular value of colour_plane_id will be present in the bitstream. The output of each of the three decoding processes is assigned to one of the 3 sample arrays of the current picture, where NAL units with colour_plane_id equal to 0, 1 and 2 are assigned to S respectively. L , S Cb and S Cr .

[0268] NOTE – When separate_colour_plane_flag is equal to 1 and chroma_format_idc is equal to 3, the variable ChromaArrayType is derived to be equal to 0. During decoding, the value of this variable is evaluated, resulting in the same operation as for monochrome pictures (when chroma_format_idc is equal to 0).

[0269] For the current picture CurrPic, the decoding process is as follows:

[0270] 1. The decoding of NAL units is specified in clause 8.2.

[0271] 2. The process in clause 8.3 specifies the following decoding process using syntax elements in the slice header layer and above:

[0272] - Derive variables and functions related to the picture order count as specified in clause 8.3.1. This only needs to be called for the first slice of a picture.

[0273] - At the beginning of the decoding process of each slice of a non-IDR picture, the decoding process of the reference picture list construction specified in clause 8.3.2 is called to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).

[0274] - Invoke the decoding process of reference picture marking in clause 8.3.3, where a reference picture can be marked as "unused for reference" or "used for long-term reference". This only needs to be invoked for the first slice of a picture.

[0275] - When the current picture is a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 or a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, the decoding process for generating unusable reference pictures specified in subclause 8.3.4 is called, which only needs to be called for the first slice of the picture.

[0276] -PictureOutputFlag is set as follows:

[0277] -PictureOutputFlag is set equal to 0 if one of the following conditions is true:

[0278] - The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.

[0279] -gdr_enabled_flag is equal to 1, and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0280] - gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.

[0281] -sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:

[0282] - PicA has PictureOutputFlag equal to 1.

[0283] - PicA has a nuh_layer_id nuhLid greater than the nuh_layer_id of the current picture.

[0284] - PicA belongs to the output layer of OLS (ie, OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).

[0285] -sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.

[0286] - Otherwise, PictureOutputFlag is set equal to pic_output_flag.

[0287] 3. The procedures in clauses 8.4, 8.5, 8.6, 8.7, and 8.8 specify the decoding process using syntax elements in all syntax structure layers. A bitstream conformance requirement is that a codec slice of a picture shall contain slice data for each CTU of the picture such that the picture is divided into slices, and the slices are divided into CTUs that each form a partition of the picture.

[0288] 4. After all slices of the current picture are decoded, the current decoded picture is marked as "used for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "used for short-term reference".

[0289] 4. Technical problems solved by the disclosed technical solutions

[0290] The existing design in the latest VVC text (JVET-Q2001-vE / v15) has the following problems:

[0291] 1) Since different types of sub-pictures are allowed to be mixed within one picture, it is confusing to refer to the contents of a NAL unit with a VCLNAL unit type as a codec slice of a particular type of picture. For example, a NAL unit with nal_unit_type equal to CRA_NUT is a codec slice of a CRA picture only if all slices of the picture have nal_unit_type equal to CRA_NUT; when one slice of the picture has nal_unit_type not equal to CRA_NUT, the picture is not a CRA picture.

[0292] 2) Currently, if a sub-picture contains a VCL NAL unit with nal_unit_type in the range of IDR_W_RADL to CRA_NUT (inclusive), then for the sub-picture, the value of subpic_treated_as_pic_flag[] is required to be equal to 1, and for the picture, mixed_nalu_types_in_pic_flag is equal to 1. In other words, for an IRAP sub-picture that is mixed with a sub-picture of another type in a picture, the value of subpic_treated_as_pic_flag[] is required to be equal to 1. However, as more mixtures of VCL NAL unit types are supported, this requirement is not sufficient.

[0293] 3) Currently, only up to two different types of VCL NAL units (and two different types of sub-pictures) are allowed within a picture.

[0294] 4) In both single-layer and multi-layer contexts, there is a lack of constraints on the output order of post-sub-pictures relative to the associated IRAP or GDR sub-pictures.

[0295] 5) Currently, it is specified that when a picture is a preceding picture of an IRAP picture, it shall be a RADL or RASL picture. This constraint, together with the definition of preceding / RADL / RASL pictures, does not allow mixing of RADL and RASL NAL unit types within a picture resulting from a mix of two CRA pictures and their non-AU aligned associated RADL and RASL pictures.

[0296] 6) In both single-layer and multi-layer contexts, there is a lack of constraints on the sub-picture type of the preceding sub-picture (ie, the NAL unit type of the VCL NAL units in the sub-picture).

[0297] 7) In both single-layer and multi-layer contexts, there is a lack of constraints on whether RASL sub-pictures can exist and be associated with IDR sub-pictures.

[0298] 8) In both single-layer and multi-layer contexts, there is a lack of constraints on whether a RADL sub-picture can exist and be associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP.

[0299] 9) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between sub-pictures that precede an IRAP sub-picture in decoding order and RADL sub-pictures associated with an IRAP sub-picture.

[0300] 10) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between sub-pictures that precede a GDR sub-picture in decoding order and sub-pictures that are associated with a GDR sub-picture.

[0301] 11) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between RASL sub-pictures associated with CRA sub-pictures and RADL sub-pictures associated with CRA sub-pictures.

[0302] 12) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between RASL sub-pictures associated with a CRA sub-picture and IRAP sub-pictures that precede the CRA sub-picture in decoding order.

[0303] 13) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative decoding order between the associated non-previous pictures and the preceding pictures of an IRAP picture.

[0304] 14) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL activity entries for sub-pictures that follow the STSA sub-picture in decoding order.

[0305] 15) There is a lack of constraints on the RPL entries of CRA subgraphs in both single-layer and multi-layer contexts.

[0306] 16) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL activity entries that reference sub-pictures of pictures generated by the decoding process used to generate unavailable reference pictures.

[0307] 17) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL entries that reference sub-pictures of a picture generated by a decoding process used to generate unavailable reference pictures.

[0308] 18) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL activity entries for sub-pictures that are associated with an IRAP picture and follow the IRAP picture in output order.

[0309] 19) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL entries for sub-pictures that are associated with an IRAP picture and follow the IRAP picture in output order.

[0310] 20) There is a lack of constraints on the RPL active entries of RADL sub-pictures in both single-layer and multi-layer contexts.

[0311] 5. Examples of solutions and implementations

[0312] In order to solve the above problems and other problems, the following summarized methods are disclosed. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow way. In addition, these items can be applied alone or combined in any way.

[0313] 1) To solve problem 1, instead of specifying the content of a NAL unit with a VCL NAL unit type as "a codec slice of a specific type of picture", specify it as "a codec slice of a specific type of picture or sub-picture". For example, the content of a NAL unit with nal_unit_type equal to CRA_NUT is specified as "a codec slice of a CRA picture or sub-picture".

[0314] a. In addition, one or more of the following terms are defined: associated GDR sub-picture, associated IRAP sub-picture, CRA sub-picture, GDR sub-picture, IDR sub-picture, IRAP sub-picture, leading sub-picture, RADL sub-picture, RASL sub-picture, STSA sub-picture, trailing sub-picture.

[0315] 2) To solve problem 2, a constraint is added to require that any two adjacent sub-pictures with different NAL unit types should have subpic_treated_as_pic_flag[] equal to 1.

[0316] a. In one example, the constraint is specified as follows: For any two adjacent sub-pictures with sub-picture indices i and j in a picture, when subpic_treated_as_pic_flag[i] or subpic_treated_as_pic_flag[j] is equal to 0, the two sub-pictures shall have the same NAL unit type.

[0317] a. Alternatively, it is required that all sub-pictures in a picture shall have the same NAL unit type (i.e., all VCL NAL units in a picture shall have the same NAL unit type, i.e., the value of mixed_nalu_types_in_pic_flag shall be equal to 0) when any sub-picture with sub-picture index i has subpic_treated_as_pic_flag[i] equal to 0. This means that mixed_nalu_types_in_pic_flag can only be equal to 1 when the corresponding subpic_treated_as_pic_flag[] of all sub-pictures is equal to 1.

[0318] 3) To solve issue 3, when mixed_nalu_types_in_pic_flag is equal to 1, a picture may be allowed to contain more than two different types of VCL NAL units.

[0319] 4) To solve problem 4, it is specified that the post-picture should be after the associated IRAP or GDR sub-picture in the output order.

[0320] 5) To address issue 5, in order to allow a mix of RADL and RASL NAL unit types within a picture resulting from a mix of two CRA pictures and their non-AU aligned associated RADL and RASL pictures, the existing constraint specifying that the leading picture of an IRAP picture shall be a RADL or RASL picture is changed as follows: When the picture is the leading picture of an IRAP picture, the nal_unit_type value of all VCL NAL units in the picture shall be equal to RADL_NUT or RASL_NUT. In addition, during the decoding of a picture with mixed nal_unit_type values ​​of RADL_NUT and RASL_NUT, when the layer containing the picture is an output layer, the PictureOutputFlag of the picture is set equal to pic_output_flag.

[0321] Thus, RADL sub-pictures within such pictures are guaranteed by the constraint that all pictures being output require correctness for a conforming decoder, although guarantees of "correctness" of "intermediate" RASL sub-pictures within such pictures are also appropriate, but not actually required, when the associated CRA picture has NoOutputBeforeRecoveryFlag equal to 1. The unnecessary portion of this guarantee is inconsequential and does not increase the complexity of implementing a conforming encoder or decoder. In this case, it is useful to add a note clarifying that although such RASL sub-pictures associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 can be output by the decoding process, they are not intended for display and therefore should not be used for display.

[0322] 6) To solve problem 6, it is specified that when a sub-picture is a preceding sub-picture of an IRAP sub-picture, it should be a RADL or RASL sub-picture.

[0323] 7) To solve issue 7, it is specified that RASL sub-pictures associated with IDR sub-pictures should not exist in the bitstream.

[0324] 8) To solve problem 8, it is specified that RADL sub-pictures associated with IDR sub-pictures with nal_unit_type equal to IDR_N_LP should not be present in the bitstream.

[0325] 9) To solve problem 9, it is specified that any sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx that precedes the IRAP sub-picture with nuh_layer_id equal to a specific value layerId and sub-picture index equal to a specific value subpicIdx in the decoding order should precede the IRAP sub-picture and all its associated RADL sub-pictures in the output order.

[0326] 10) To solve problem 10, it is specified that any sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx that precedes the GDR sub-picture with nuh_layer_id equal to a specific value layerId and sub-picture index equal to a specific value subpicIdx in the decoding order should precede the GDR sub-picture and all its associated sub-pictures in the output order.

[0327] 11) To address issue 11, specify that any RASL sub-picture associated with a CRA sub-picture should precede any RADL sub-picture associated with the CRA sub-picture in output order.

[0328] 12) To address issue 12, specify that any RASL sub-picture associated with a CRA sub-picture should follow, in output order, any IRAP sub-picture that precedes the CRA sub-picture in decoding order.

[0329] 13) In order to solve problem 13, it is specified that if field_seq_flag is equal to 0, and the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a leading sub-picture associated with an IRAP sub-picture, then it should be before all non-leading sub-pictures associated with the same IRAP sub-picture in the decoding order; otherwise, let subpicA and subpicB be the first and last leading sub-pictures associated with the IRAP sub-picture in the decoding order, respectively, then there should be at most one non-leading sub-picture before subpicA in the decoding order, with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx, and there should be no non-leading picture between picA and picB in the decoding order, with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx.

[0330] 14) In order to solve problem 14, it is specified that when the current sub-picture whose TemporalId is equal to a specific value tId, nuh_layer_id is equal to a specific value layerId, and the sub-picture index is equal to a specific value subpicIdx is a sub-picture after the STSA sub-picture whose TemporalId is equal to tId, nuh_layer_id is equal to layerId, and the sub-picture index is equal to subpicIdx in the decoding order, the picture whose TemporalId is equal to tId and nuh_layer_id is equal to layerId that is before the picture containing the STSA sub-picture in the decoding order should not be included as an active entry in RefPicList[0] or RefPicList[1].

[0331] 15) To address issue 15, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a CRA sub-picture, there should be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede, in output order or decoding order, any picture (when present) containing a previous IRAP sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx in decoding order.

[0332] 16) To address issue 16, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is not a RASL sub-picture associated with a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR sub-picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a sub-picture of a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there should be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] generated by the decoding process for generating unavailable reference pictures.

[0333] 17) To solve problem 17, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is not a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a sub-picture that precedes the preceding sub-picture associated with the same CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 in decoding order, a preceding sub-picture associated with a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR sub-picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a sub-picture of a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there should be no pictures referenced by entries in RefPicList[0] or RefPicList[1] generated by the decoding process for generating unavailable reference pictures.

[0334] 18) To address issue 18, it is specified that when the current sub-picture is associated with an IRAP sub-picture and follows the IRAP sub-picture in output order, there should be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that precede the picture containing the associated IRAP sub-picture in output order or decoding order.

[0335] 19) To address issue 19, it is specified that when the current sub-picture is associated with an IRAP sub-picture, follows the IRAP sub-picture in output order, and follows the preceding sub-picture (if any) associated with the same IRAP sub-picture in both decoding order and output order, there should be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede the picture containing the associated IRAP sub-picture in output order or decoding order.

[0336] 20) To address issue 20, specify that when the current sub-picture is a RADL sub-picture, there should be no active entry in RefPicList[0] or RefPicList[1] that is any of the following:

[0337] a. Pictures containing RASL sub-pictures

[0338] b. The picture that precedes the picture containing the associated IRAP sub-picture in decoding order

[0339] 21) Specifies that when a sub-picture is not a preceding sub-picture of an IRAP sub-picture, it should not be a RADL or RASL sub-picture.

[0340] 22) Alternatively, to solve problem 10, it is specified that any sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx that precedes the sub-picture with nuh_layer_id equal to a specific value layerId and sub-picture index equal to a specific value subpicIdx in the decoding order in the recovery point picture should precede the sub-picture in the output order in the recovery point picture.

[0341] 23) Alternatively, to address issue 15, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a CRA sub-picture, there should be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede any previous picture in decoding order (when any) that contains an IRAP sub-picture with nuh_layer_id equal to layerId and a sub-picture index equal to subpicIdx in output order or decoding order.

[0342] 24) Alternatively, to address issue 18, it is specified that when the current sub-picture is after an IRAP sub-picture with the same value of nuh_layer_id and the same value of sub-picture index in both decoding and output order, there should be no picture referenced by an active entry in RefPicList[0] or RefPicList[1] that is before the picture containing the IRAP sub-picture in output order or decoding order.

[0343] 25) Alternatively, to address issue 19, it is specified that when the current sub-picture is after an IRAP sub-picture with the same value of nuh_layer_id and the same value of sub-picture index and the preceding sub-picture associated with the IRAP sub-picture (if any) in both decoding and output order, there should be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede the picture containing the IRAP sub-picture in output order or decoding order.

[0344] 26) Alternatively, to solve problem 20, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a RADL sub-picture, there should be no active entry in RefPicList[0] or RefPicList[1] that is any of the following:

[0345] a. A picture containing a RASL sub-picture with a sub-picture index equal to subpicIdx and nuh_layer_id equal to layerId

[0346] b. The picture that precedes the picture containing the associated IRAP sub-picture in decoding order

[0347] 6. Examples

[0348] The following are some example embodiments of some aspects of the invention summarized in Section 5 above, which may be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-Q2001-vE / v15. The most relevant parts that have been added or modified are Highlight, and some of the deleted parts are highlighted with left and right double brackets (e.g., [[]]), where the deleted text is between the double brackets. There are also some other changes of an editorial nature or that are not part of the present technical solution, so they are not highlighted.

[0349] 6.1. First embodiment

[0350] This embodiment is for items 1, 1a, 2, 2a, 4 and 6 to 20.

[0351] 3 Definition ...

[0352] Associated GDR pictures (of a specific picture with a specific value of nuh_layer_id layerId): the previous GDR picture in decoding order (when present) with nuh_layer_id equal to layerId, where there is no IRAP picture with nuh_layer_id equal to layerId between this picture and the specific picture in decoding order.

[0353]

[0354] Associated IRAP pictures (of a specific picture with a specific value of nuh_layer_id layerId): the previous IRAP picture in decoding order (when present) with nuh_layer_id equal to layerId, where there is no GDR picture with nuh_layer_id equal to layerId between this picture and the specific picture in decoding order.

[0355]

[0356] Completely random access (CRA) picture: An IRAP picture with nal_unit_type equal to CRA_NUT per VCL NAL unit.

[0357]

[0358] Gradual Decoding Refresh (GDR) AU: An AU in which there is a PU for each layer in the CVS and each coded picture stored in the PU is a GDR picture.

[0359] Gradual Decoding Refresh (GDR) picture: A picture where each VCL NAL unit has nal_unit_type equal to GDR_NUT.

[0360]

[0361] Instantaneous Decoding Refresh (IDR) picture: An IRAP picture with nal_unit_type equal to IDR_W_RADL or IDR_N_LP per VCL NAL unit.

[0362]

[0363] Intra random access point (IRAP) picture: A picture whose all VCL NAL units have the same value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive.

[0364]

[0365] Preceding picture: A picture that precedes the associated IRAP picture in output order.

[0366]

[0367] Output order: order, Output of decoded pictures from DPB

[0368] Random Access Decodable Leading (RADL) picture: A picture with nal_unit_type equal to RADL_NUT per VCL NAL unit.

[0369]

[0370] Random Access Skip Leading (RASL) picture: A picture with nal_unit_type equal to RASL_NUT per VCL NAL unit.

[0371]

[0372] Step-by-Step Temporal Sub-Layer Access (STSA) picture: A picture with nal_unit_type equal to STSA_NUT per VCL NAL unit.

[0373]

[0374] Trailing picture: A picture where each VCL NAL unit has nal_unit_type equal to TRAIL_NUT.

[0375] NOTE – A posting picture associated with an IRAP or GDR picture also follows the IRAP or GDR picture in decoding order. Pictures that follow the associated IRAP or GDR picture in output order and precede the associated IRAP or GDR picture in decoding order are not allowed.

[0376] ...

[0377] 7.4.2.2 NAL unit header semantics ...

[0378] nal_unit_type specifies the NAL unit type, ie, the type of RBSP data structure contained in the NAL unit as specified in Table 5.

[0379] NAL units with nal_unit_type in the range UNSPEC_28..UNSPEC_31 (unspecified semantics), inclusive, shall not affect the decoding process as specified in this specification.

[0380] NOTE 2 – NAL unit types in the range of UNSPEC_28..UNSPEC_31 may be used as determined by the application. The decoding process for these values ​​of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, special care must be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define any management of these values. These nal_unit_type values ​​may be appropriate for use only in contexts where "conflicts" in usage (i.e., different definitions of the meaning of the content of NAL units for the same nal_unit_type value) are not important, or are not possible, or are managed (e.g., defined or managed in a controlling application or transport specification, or by an environment where the control bitstream is distributed).

[0381] For purposes other than determining the amount of data in a DU of a bitstream (as specified in Annex C), a decoder should ignore (remove from the bitstream and discard) the contents of all NAL units that use reserved values ​​of nal_unit_type.

[0382] NOTE 3 – This requirement allows for the future definition of compatible extensions to this specification.

[0383] Table 5 – NAL unit type codes and NAL unit type classifications

[0384]

[0385]

[0386] NOTE 4 - A completely random access (CRA) picture may have an associated RASL or RADL picture present in the bitstream.

[0387] NOTE 5 - An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP does not have an associated preceding picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream.

[0388] The value of nal_unit_type shall be the same for all VCL NAL units of a sub-picture. A sub-picture is said to have the same NAL unit type as the VCL NAL units of the sub-picture.

[0389]

[0390] For any particular picture's VCL NAL unit, the following applies:

[0391] - If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and a picture or PU is said to have the same NAL unit type as the VCL NAL units of the picture or PU.

[0392] Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture shall have at least two sub-pictures, and the VCL NAL units of the picture shall have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one sub-picture of the picture shall all have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of the other sub-pictures in the picture shall all have different specific values ​​of nal_unit_type equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.

[0393] The requirement for bitstream conformance is that the following constraints apply:

[0394] - A post-image shall follow the associated IRAP or GDR image in output order.

[0395]

[0396] - When a picture is the preceding picture of an IRAP picture, it shall be a RADL or RASL picture.

[0397]

[0398] - There shall be no RASL pictures associated with an IDR picture in the bitstream.

[0399]

[0400] - There shall be no RADL pictures associated with an IDR picture with nal_unit_type equal to IDR_N_LP in the bitstream.

[0401] NOTE 6 – Random access can be performed at the location of an IRAP PU (and correctly decode the IRAP picture and all subsequent non-RASL pictures in decoding order) by discarding all PUs preceding the IRAP PU, provided that each parameter set is available at the time it is referenced (either in the bitstream or by external means not specified in this specification).

[0402]

[0403] - Any pictures with nuh_layer_id equal to layerId that precede the IRAP picture with nuh_layer_id equal to a particular value layerId in decoding order shall precede the IRAP picture and all its associated RADL pictures in output order.

[0404]

[0405] - Any pictures with nuh_layer_id equal to layerId that precede the GDR picture with nuh_layer_id equal to a particular value layerId in decoding order shall precede the GDR picture and all its associated pictures in output order.

[0406]

[0407] - Any RASL pictures associated with a CRA picture shall precede any RADL pictures associated with the CRA picture in output order.

[0408]

[0409] - Any RASL pictures associated with a CRA picture shall follow, in output order, any IRAP pictures that precede the CRA picture in decoding order.

[0410]

[0411] - If field_seq_flag is equal to 0, and the current picture with nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, then it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, let picA and picB be the first and last leading pictures associated with the IRAP picture in decoding order, respectively, then there shall be at most one non-leading picture with nuh_layer_id equal to layerId before picA in decoding order, and there shall be no non-leading pictures with nuh_layer_id equal to layerId between picA and picB in decoding order.

[0412] ...

[0413] 7.4.3.4 Picture parameter set semantics ...

[0414] mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture of the reference PPS has more than one VCL NAL unit, The VCL NAL units do not have the same value of nal_unit_type [[and the picture is not an IRAP picture]]. mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture referencing a PPS has one or more VCL NAL units, and that the VCL NAL units of each picture referencing a PPS have the same value of nal_unit_type.

[0415] When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_types_in_pic_flag shall be equal to 0.

[0416] [[For each slice with nal_unit_type value nalUnitTypeA in the range of IDR_W_RADL to CRA_NUT (inclusive), in a picture picA that also contains one or more slices with another value of nal_unit_type (i.e., the value of mixed_nalu_types_in_pic_flag of picture picA is equal to 1), the following applies:

[0417] - The slice shall belong to a sub-picture subpicA for which the value of subpic_treated_as_pic_flag[i] is equal to 1.

[0418] - A slice shall not belong to a sub-picture picA that contains VCL NAL units whose nal_unit_type is not equal to nalUnitTypeA.

[0419] - If nalUnitTypeA is equal to CRA, then for all subsequent PUs that follow the current picture in CLVS in decoding order and in output order, neither RefPicList[0] nor RefPicList[1] of the slices in subpicA in these PUs shall include any pictures in the active entry that precede picA in decoding order.

[0420] - Otherwise (i.e., nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS that follow the current picture in decoding order, neither RefPicList[0] nor RefPicList[1] for slices in subpicA in these PUs shall include any pictures in the active entry that precede picA in decoding order. ]]

[0421] NOTE 1 – mixed_nalu_types_in_pic_flag equal to 1 indicates that the picture referencing the PPS contains slices with different NAL unit types, e.g., a codec picture resulting from a sub-picture bitstream merge operation where the encoder must ensure matching bitstream structure and further alignment of parameters of the original bitstream. An example of such alignment is as follows: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_in_pic_flag is equal to 1, the picture referencing the PPS shall not have slices with nal_unit_type equal to IDR_W_RADL or IDR_N_LP. ...

[0422] 7.4.3.7 Image header structure semantics ...

[0423] recovery_poc_cnt specifies the recovery point of the decoded picture in output order.

[0424]

[0425] If the current picture is a GDR picture [[associated with PH]] and has equal If the CLVS of PicOrderCntVal of [[PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt]] contains a picture picA that follows the current GDR picture in decoding order, then the picture picA is called the recovery point picture. Otherwise, Having greater than The first picture in the output order whose PicOrderCntVal is [[the current picture's PicOrderCntVal plus the value of recovery_poc_cnt]] is called the recovery point picture. The recovery point picture should not precede the current GDR picture in the decoding order. The value of recovery_poc_cnt shall be in the range of 0 to MaxPicOrderCntLsb-1, inclusive.

[0426] [[When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:

[0427] RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81)]]

[0428] NOTE 2 – When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to that of the associated GDR picture, [[RpPicOrderCntVal]], the current decoded picture and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (when present) that precedes the associated GDR picture in decoding order. ...

[0429] 8.3.2 Decoding process of reference picture list construction ...

[0430] The requirement for bitstream conformance is that the following constraints apply:

[0431] - For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i].

[0432] - The picture referenced by each active entry in RefPicList[0] or RefPicList[1] shall be present in the DPB and shall have a TemporalId less than or equal to the TemporalId of the current picture.

[0433] - Each entry in RefPicList[0] or RefPicList[1] shall refer to a picture that is not the current picture and shall have non_reference_picture_flag equal to 0.

[0434] - A STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and a LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture.

[0435] - There shall be no difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry in RefPicList[0] or RefPicList[1] greater than or equal to 2 24 LTRP entry.

[0436] - Let setOfRefPics be the set of unique pictures referenced by all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize-1 (inclusive), where MaxDpbSize is as specified in clause A.4.2, and setOfRefPics shall be the same for all slices of the picture.

[0437] - When the current slice has nal_unit_type equal to STSA_NUT, there shall be no active entry in RefPicList[0] or RefPicList[1] with TemporalId equal to the current picture's TemporalId and nuh_layer_id equal to the current picture's nuh_layer_id.

[0438] - When the current picture is a picture after the STSA picture in decoding order with TemporalId equal to the current picture's TemporalId and nuh_layer_id equal to the current picture's nuh_layer_id, the picture before the STSA picture in decoding order with TemporalId equal to the current picture's TemporalId and nuh_layer_id equal to the current picture shall not be included as the active entry in RefPicList[0] or RefPicList[1].

[0439]

[0440] - When the current picture with nuh_layer_id equal to a particular value layerId is a CRA picture, there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede, in output order or decoding order, any preceding IRAP picture with nuh_layer_id equal to layerId in decoding order (when present).

[0441]

[0442] - When the current picture with nuh_layer_id equal to a specific value layerId is not a RASL picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there shall be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] generated by the decoding process for generating unavailable reference pictures.

[0443]

[0444] - When the current picture with nuh_layer_id equal to a specific value layerId is not a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a picture that precedes, in decoding order, the preceding picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a preceding picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] generated by the decoding process for generating unusable reference pictures.

[0445]

[0446] - When the current picture is associated with an IRAP picture and follows the IRAP picture in output order, there shall be no pictures referenced by active entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.

[0447]

[0448] - When the current picture is associated with an IRAP picture, follows the IRAP picture in output order, and follows the preceding picture (if any) associated with the same IRAP picture in both decoding order and output order, there shall be no pictures referenced by entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.

[0449]

[0450] - When the current picture is a RADL picture, there shall be no active entries in RefPicList[0] or RefPicList[1] that are any of the following:

[0451] o RASL images

[0452] o The picture that precedes the associated IRAP picture in decoding order

[0453] - The picture referenced by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be in the same AU as the current picture.

[0454] - The picture referenced by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB and shall have a nuh_layer_id less than the nuh_layer_id of the current picture.

[0455] - Each ILRP entry in either RefPicList[0] or RefPicList[1] of a stripe shall be an active entry.

[0456] Figure 5 1 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10 bit multi-component pixel values, or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0457] System 1900 may include a codec component 1904 that can implement various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to generate a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 can be stored or sent via a communication connection as represented by component 1906. The bitstream (or codec) representation of the storage or communication transmission of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or transmit to a displayable video of display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tool or operation is used at the encoder, and the corresponding decoding tool or operation of the inverse codec result will be performed by the decoder.

[0458] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smart phones, or other devices capable of performing digital data processing and / or video display.

[0459] Figure 6 3600. The apparatus 3600 may be used to implement one or more methods described herein. The apparatus 3600 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 may be configured to implement one or more methods described in this document. The memory(s) 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in a hardware circuit system.

[0460] Figure 8 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0461] like Figure 8 As shown, the video coding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, wherein the source device 110 may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110, wherein the target device 120 may be referred to as a video decoding device.

[0462] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0463] The video source 112 may include a source, such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bit stream. The bit stream may include a bit sequence that forms a codec representation of the video data. The bit stream may include a codec picture and related data. The codec picture is a codec representation of the picture. Related data may include a sequence parameter set, a picture parameter set, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be directly sent to the target device 120 via the network 130a via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130b for access by the target device 120.

[0464] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0465] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain coded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the coded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the target device 120, or may be outside the target device 120 configured to interface with an external display device.

[0466] The video encoder 114 and the video decoder 124 may operate in accordance with a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or additional standards.

[0467] Fig. 9 is a block diagram showing an example of a video encoder 200, which may be Figure 8 The video encoder 114 in the system 100 is shown.

[0468] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Fig. 9 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared between various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0469] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0470] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is a picture in which the current video block is located.

[0471] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the purpose of explanation, they are not shown in FIG. Fig. 9 are represented separately in the examples.

[0472] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0473] The mode selection unit 203 may select one of the coding modes (e.g., intra or inter) based on the error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction modes (CIIP), where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of the motion vector of the block (e.g., sub-pixel or integer pixel precision).

[0474] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block of the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.

[0475] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0476] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search the reference picture of list 0 or list 1 for the reference video block of the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1, the reference index including the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0477] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 containing the reference video block and a motion vector indicating a spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0478] In some examples, motion estimation unit 204 may output a complete set of motion information for use in a decoding process of a decoder.

[0479] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0480] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0481] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0482] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0483] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.

[0484] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0485] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0486] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[0487] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0488] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0489] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0490] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.

[0491] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of a block of a video, but may not necessarily modify the generated bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on the decision or determination, the conversion from a block of a video to a bitstream (or bitstream representation) of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream of the video to the block of the video will be performed using the video processing tool or mode enabled based on the decision or determination.

[0492] Fig.10 is a block diagram showing an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown.

[0493] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Fig.10 In the example of , video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared between various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0494] exist Fig.10 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those generally performed for the video encoder 200 ( Fig. 9 ) is the reverse of the encoding process described in .

[0495] The entropy decoding unit 301 may retrieve a coded bitstream. The coded bitstream may include entropy-coded video data (e.g., coded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and Merge modes.

[0496] The motion compensation unit 302 may generate a motion compensated block, and interpolation may be performed based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in a syntax element.

[0497] The motion compensation unit 302 may calculate interpolation of sub-integer pixels of the reference block using an interpolation filter as used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.

[0498] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the coded video sequence.

[0499] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0500] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307 to provide a reference block for subsequent motion compensation / intra prediction and also to generate a decoded video for presentation on a display device.

[0501] A list of some preferred solutions for the embodiments is provided next.

[0502] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 1).

[0503] 1. A video processing method (for example, in Figure 7), comprising: performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video (702), wherein the codec representation conforms to a format rule that specifies that the one or more pictures including the one or more sub-pictures are included in the codec representation according to a network abstraction layer (NAL) unit, wherein the type NAL unit indicated in the codec representation includes a codec slice of a specific type of picture or a codec slice of a specific type of sub-picture.

[0504] The following solution illustrates an example embodiment of the technique discussed in the previous section (eg, item 2).

[0505] 2. A video processing method, comprising: performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule specifying that two adjacent sub-pictures having different network abstraction layer unit types will have the same indication of sub-pictures being considered as picture flags.

[0506] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 4, 5, 6, 7, 9, 1, 11, 12).

[0507] 3. A video processing method, comprising: performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule defining an order of a first type of sub-picture and a second type of sub-picture, wherein the first sub-picture is a post-sub-picture or a pre-sub-picture or a random access skip pre-sub (RASL) sub-picture type, and the second sub-picture is a RASL type or a random access decodable pre-sub (RADL) type or an instantaneous decoding refresh (IDR) type or a gradual decoding refresh (GDR) type sub-picture.

[0508] 4. The method of solution 3, wherein the rule specifies that the post-sub-picture follows the associated intra-frame random access point or GDR sub-picture in output order.

[0509] 5. The method according to solution 3, wherein the rule specifies that when a picture is a preceding picture of an intra random access point picture, the nal_unit_type value of all network abstraction layer units in the picture is equal to RADL_NUT or RASL_NUT.

[0510] 6. The method of solution 3, wherein the rule specifies that a given sub-picture that is a preceding sub-picture of an IRAP sub-picture must also be a RADL or RASL sub-picture.

[0511] 7. The method of solution 3, wherein the rule specifies that a given sub-picture that is a RASL sub-picture is not allowed to be associated with an IDR sub-picture.

[0512] 8. The method of solution 3, wherein the rule specifies that a given sub-picture with the same layer id and sub-picture index as an IRAP sub-picture must precede the IRAP sub-picture and all its associated RADL sub-pictures in output order.

[0513] 9. The method of solution 3, wherein the rule specifies that a given sub-picture with the same layer id and sub-picture index as a GDR sub-picture must precede the GDR sub-picture and all its associated RADL sub-pictures in output order.

[0514] 10. The method of solution 3, wherein the rule specifies that a given sub-picture that is a RASL sub-picture associated with a CRA sub-picture precedes all RADL sub-pictures associated with the CRA sub-picture in output order.

[0515] 11. The method of solution 3, wherein the rule specifies that a given sub-picture that is a RASL sub-picture associated with a CRA sub-picture precedes all IRAP sub-pictures associated with the CRA sub-picture in output order.

[0516] 12. The method of solution 3, wherein the rule specifies that a given sub-picture is a preceding sub-picture with an IRAP sub-picture, and the given sub-picture precedes all non-preceding sub-pictures associated with the IRAP picture in decoding order.

[0517] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 8, 14, 15).

[0518] 13. A video processing method, comprising: performing conversion between a video including one or more pictures containing one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to format rules defining conditions for allowing or not allowing a first type of sub-picture to appear together with a second type of sub-picture.

[0519] 14. The method of solution 13, wherein the rule specifies that in the case of an IDR sub-picture with a network abstraction layer type IDR_N_LP, the codec representation is not allowed to have a RADP sub-picture.

[0520] 15. The method of solution 13, wherein the rule does not allow a picture to be included in a reference list of a picture including a step-by-step temporal sub-layer access (STSA) sub-picture such that the picture precedes the picture including the STSA sub-picture.

[0521] 16. The method of solution 13, wherein the rule does not allow a picture to be included in a reference list of a picture including an intra random access point (IRAP) sub-picture such that the picture precedes the picture including the IRAP sub-picture.

[0522] The following solution illustrates an example embodiment of the techniques discussed in the previous section (eg, items 21-26).

[0523] 17. A method of video processing, comprising: performing conversion between a video comprising one or more video pictures including one or more sub-pictures and a codec representation of the video; wherein the codec representation comprises one or more layers of the video pictures in an order according to a rule.

[0524] The following solution illustrates an example implementation of the technique discussed in the previous section (eg, item 21).

[0525] 18. The method of solution 17, wherein the rule specifies that a sub-picture that is not a pre-picture of type intra random access point cannot have a random access decodable pre-picture (RADL) or random access skip pre-picture (RASL) sub-picture type.

[0526] The following solution illustrates an example implementation of the technique discussed in the previous section (eg, item 22).

[0527] 19. A method according to any one of solutions 17-18, wherein the rule specifies that a first sub-picture that precedes a second sub-picture in decoding order in a recovery point picture must also precede the second sub-picture in output order, wherein the first sub-picture and the second sub-picture belong to the same layer and have the same sub-picture index.

[0528] The following solution illustrates an example implementation of the technique discussed in the previous section (eg, item 23).

[0529] 20. A method according to any one of solutions 17-19, wherein the rule specifies that the video picture referenced by the reference picture list is before the completely random access (CRA) sub-picture having an intra-frame random access point sub-picture with the same layer id and sub-picture index as the CRA sub-picture in output order or decoding order.

[0530] The following solution illustrates an example implementation of the technique discussed in the previous section (eg, item 24).

[0531] 21. A method according to any one of solutions 17-20, wherein the rule specifies that when the current sub-picture is after an intra random access point (IRAP) sub-picture with the same value of nuh_layer_id and the same value of sub-picture index in both decoding and output order, the picture referenced by the active entry in the reference picture list is not allowed to be before the picture containing the IRAP sub-picture in output order or decoding order.

[0532] The following solution illustrates an example implementation of the technique discussed in the previous section (eg, item 25).

[0533] 22. A method according to any one of solutions 17-21, wherein the rule specifies that when the current sub-picture is after an intra random access point IRAP sub-picture with the same value of nuh_layer_id and the same value of sub-picture index and a preceding sub-picture associated with the IRAP sub-picture in both decoding and output order, there should be no picture referenced by an entry in the reference picture list that is before the picture containing the IRAP sub-picture in output order or decoding order.

[0534] The following solution illustrates an example implementation of the technique discussed in the previous section (eg, item 26).

[0535] 23. A method according to any one of solutions 17-22, wherein the rule specifies that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a RADL sub-picture, there should be no active entry in the reference picture list that is any of the following: (a) a picture containing a RADL sub-picture with a sub-picture index equal to subpicIdx and nuh_layer_id equal to layerId, or (b) a picture that precedes the picture containing the associated IRAP sub-picture in decoding order.

[0536] 24. A method according to any one of solutions 1 to 23, wherein the conversion includes encoding the video into a codec representation.

[0537] 25. A method according to any one of solutions 1 to 23, wherein the conversion includes decoding the codec representation to generate pixel values ​​of the video.

[0538] 26. A video decoding device comprising a processor configured to implement the method according to one or more of solutions 1 to 25.

[0539] 27. A video encoding device comprising a processor configured to implement the method according to one or more of solutions 1 to 25.

[0540] 28. A computer program product storing a computer code, which, when executed by a processor, causes the processor to implement the method according to any one of solutions 1 to 25.

[0541] 29. A method, apparatus or system as described in this document.

[0542] In the solution described herein, an encoder may comply with the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder may parse the syntax elements in the codec representation using the format rules to generate decoded video, knowing the presence and absence of the syntax elements according to the format rules.

[0543] Fig.11 A flow chart of an example method 1100 of video processing is shown. Operation 1102 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule that specifies that, in response to a sub-picture not being a preceding sub-picture of an intra random access point sub-picture, the sub-picture cannot be a random access type sub-picture, and wherein the preceding sub-picture precedes the intra random access point sub-picture in output order.

[0544] In some embodiments of the method 1100, an intra random access point sub-picture is a sub-picture for which all video coding layer (VCL) network abstraction layer (NAL) units have the same value of NAL unit type in the range of IDR_W_RADL to CRA_NUT, the range including IDR_W_RADL and CRA_NUT. In some embodiments of the method 1100, a sub-picture of a random access type includes a random access decodable leading sub-picture. In some embodiments of the method 1100, a random access decodable leading sub-picture is a sub-picture for which each video coding layer (VCL) network abstraction layer (NAL) unit has a NAL unit type equal to RADL_NUT. In some embodiments of the method 1100, a sub-picture of a random access type includes a random access skip leading sub-picture. In some embodiments of the method 1100, a random access skip leading sub-picture is a sub-picture for which each video coding layer (VCL) network abstraction layer (NAL) unit has a NAL unit type equal to RASL_NUT.

[0545] Fig.12A flow chart of an example method 1200 for video processing is shown. Operation 1202 includes performing conversion between a video including one or more pictures including a plurality of sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a first sub-picture precedes a second sub-picture in an output order in a recovery point picture in response to the first sub-picture and the second picture having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index, and the first sub-picture precedes the second sub-picture in decoding order.

[0546] Fig.13 A flow diagram of an example method 1300 of video processing is shown. Operation 1302 includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture that includes a current slice, wherein the bitstream conforms to a format rule, wherein the format rule specifies an order in which pictures are indicated in the bitstream, wherein the format rule does not allow an entry in a reference picture list for the current slice to include a first picture that precedes a second picture that precedes the current picture according to a second order according to a first order, wherein the second picture includes an intra random access point sub-picture having a same layer identifier of a network abstraction unit (NAL) unit and a same sub-picture index as the current sub-picture, and wherein the current sub-picture is a fully random access sub-picture.

[0547] In some embodiments of the method 1300, the first order comprises a decoding order or an output order. In some embodiments of the method 1300, the second order comprises a decoding order. In some embodiments of the method 1300, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 1300, the reference picture list comprises a list 1 reference picture list. In some embodiments of the method 1300, the fully random access sub-picture is an intra random access point sub-picture per video coding layer (VCL) NAL unit having a NAL unit type equal to CRA_NUT.

[0548] Fig.14 A flow chart of an example method 1400 for video processing is shown. Operation 1402 includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture that includes a current slice, wherein the bitstream conforms to a format rule, wherein the format rule specifies an order in which pictures are indicated in the bitstream, wherein the format rule does not allow an active entry in a reference picture list for the current slice to include a first picture that precedes a second picture according to a first order, wherein the second picture includes an intra random access point sub-picture having a same layer identifier of a network abstraction unit (NAL) unit and a same sub-picture index as the current sub-picture, and wherein the current sub-picture is after the intra random access point sub-picture according to the second order.

[0549] In some embodiments of the method 1400, the first order comprises a decoding order or an output order. In some embodiments of the method 1400, the active entry corresponds to an entry that can be used as a reference index in inter prediction of the current slice. In some embodiments of the method 1400, the second order comprises a decoding order and an output order. In some embodiments of the method 1400, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 1400, the reference picture list comprises a list 1 reference picture list.

[0550] Fig.15 A flow chart of an example method 1500 for video processing is shown. Operation 1502 includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture including a current slice, wherein the bitstream conforms to a format rule, wherein the format rule specifies an order in which pictures are indicated in the bitstream, wherein the format rule does not allow an entry in a reference picture list for the current slice to include a first picture that precedes a second picture in a first order or a second order, wherein the second picture includes an intra random access point sub-picture that has zero or more associated preceding sub-pictures and has a same layer identifier of a network abstraction unit (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra random access point sub-picture and the zero or more associated preceding sub-pictures in a first order and a second order.

[0551] In some embodiments of the method 1500, the first order comprises a decoding order. In some embodiments of the method 1500, the second order comprises an output order. In some embodiments of the method 1500, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 1500, the reference picture list comprises a list 1 reference picture list.

[0552] Fig.16 A flow chart of an example method 1600 for video processing is shown. Operation 1602 includes performing conversion between a video including a current picture and a bitstream of the video, the current picture including a current sub-picture comprising a current slice, wherein the bitstream conforms to a format rule that specifies that, in response to the current sub-picture being a random access decodable preceding sub-picture, active entries of a reference picture list of the current slice are not allowed to include any one or more of: a first picture including a random access skipped preceding sub-picture having the same sub-picture index as the current sub-picture, and a second picture that precedes, in decoding order, a third picture including an intra random access point sub-picture associated with the random access decodable preceding sub-picture.

[0553] In some embodiments of the method 1600, the active entry corresponds to an entry that can be used as a reference index in inter prediction of the current slice. In some embodiments of the method 1600, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 1600, the reference picture list comprises a list 1 reference picture list.

[0554] In some embodiments of methods 1100-1600, performing the conversion includes encoding the video into a bitstream. In some embodiments of methods 1100-1600, performing the conversion includes generating a bitstream from the video, and the method also includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of methods 1100-1600, performing the conversion includes decoding the video from the bitstream.

[0555] In some embodiments, a video decoding device includes a processor configured to implement the operations described for methods 1100-1600. In some embodiments, a video encoding device includes a processor configured to implement the operations described for methods 1100-1600. In some embodiments, a computer program product storing computer instructions, which when executed by a processor causes the processor to implement the operations described for methods 1100-1600. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to the operations described for methods 1100-1600. In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause the processor to implement the operations described for methods 1100-1600. In some embodiments, a method of bitstream generation includes: generating a bitstream of a video according to the operations described for methods 1100-1600, and storing the bitstream on a computer-readable program medium. Some embodiments include a method, an apparatus, a bitstream generated according to the disclosed method, or a system as described in this document.

[0556] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation of the current video block may correspond, for example, to bits that are juxtaposed or interspersed in different places within the bitstream. For example, macroblocks may be encoded according to error residual values ​​of transforms and codecs and also using bits in headers and other fields in the bitstream. In addition, during conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include or not include certain syntax fields, and generate a codec representation accordingly by including syntax fields or excluding syntax fields from the codec representation.

[0557] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, which are used to be executed by a data processing device or control the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that affect machine-readable propagation signals, or a combination of one or more of them. The term "data processing device" includes all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagation signal is an artificially generated signal generated to encode information for transmission to a suitable receiver device, such as a machine-generated electrical signal, an optical signal, or an electromagnetic signal.

[0558] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0559] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits, and the apparatus can also be implemented as special purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application specific integrated circuits).

[0560] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operably coupled to receive data from or transfer data to or from the one or more mass storage devices. However, a computer does not require such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory can be supplemented by or incorporated into a dedicated logic circuit.

[0561] Although this patent document contains many details, these details should not be interpreted as limitations on any subject matter or the scope of possible protection, but rather as descriptions of features of specific embodiments specified for specific technologies. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations and even initially claimed as such, one or more features from the claimed combination may be excluded from the combination in some cases, and the claimed combination may be directed to a sub-combination or a variation of the sub-combination.

[0562] Similarly, although operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the operations shown be performed, to achieve the desired results. In addition, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0563] Only a few implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and shown in this patent document.

Claims

1. A video processing method, comprising: performing conversion between a video comprising one or more pictures including one or more sub-pictures and a bitstream of said video, wherein the bitstream complies with a format rule that specifies that when a sub-picture is not a preceding sub-picture of an intra-frame random access point sub-picture, the sub-picture cannot be a random access decodable preceding sub-picture or a random access skipped preceding sub-picture, and Wherein, the preceding sub-picture precedes the random access point sub-picture in the frame according to the output order; wherein the format rule specifies an order in which pictures are indicated in the bitstream, The format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture according to a first order or a second order, wherein the second picture comprises an intra random access point sub-picture, the intra random access point sub-picture having zero or more associated preceding sub-pictures and having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra-frame random access point sub-picture and the zero or more associated preceding sub-pictures in the first order and the second order; The first order is a decoding order, the second order is an output order, and the reference picture list includes a list 0 reference picture list or a list 1 reference picture list.

2. The method according to claim 1, wherein: The random access point sub-picture is a sub-picture whose all video codec layer (VCL) network abstraction layer (NAL) units have the same value of NAL unit type in the range of IDR_W_RADL to CRA_NUT, including IDR_W_RADL and CRA_NUT, wherein the random access decodable leading sub-picture is a sub-picture with a NAL unit type equal to RADL_NUT for each VCL NAL unit, and The random access skipped leading sub-picture is a sub-picture in which each VCL NAL has a NAL unit type equal to RASL_NUT.

3. The method according to claim 1, wherein: The format rule further specifies that, in response to the following situation, the first sub-picture precedes the second sub-picture in the recovery point picture in the output order: The first sub-picture and the second sub-picture have the same layer identifier and the same sub-picture index of a network abstraction layer (NAL) unit, and The first sub-picture precedes the second sub-picture in the decoding order.

4. The method according to claim 1, wherein: in, When the current sub-picture is a completely random access sub-picture, the format rule does not allow an entry in a reference picture list of the current slice to include a third picture preceding the current picture according to a third order, a fourth picture preceding the current picture according to a fourth order, wherein the fourth picture includes an intra random access point sub-picture having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and The fully random access sub-picture is an intra-frame random access point sub-picture in which each VCL NAL unit has a nal_unit_type equal to CRA_NUT.

5. The method according to claim 4, wherein: The third order includes the decoding order or the output order, wherein the fourth order includes the decoding order, and The reference picture list includes the list 0 reference picture list or the list 1 reference picture list.

6. The method according to claim 1, wherein: in, The format rule does not allow an active entry in a reference picture list of the current slice to include a fifth picture that precedes a sixth picture according to a fifth order, wherein the sixth picture includes an intra random access point sub-picture having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and The current sub-picture is located after the random access point sub-picture in the frame in the sixth order.

7. The method according to claim 6, wherein: the fifth order includes the decoding order or the output order, and the sixth order includes the decoding order and the output order, wherein the active entry corresponds to an entry that can be used as a reference index in inter prediction of the current slice, and The reference picture list includes the list 0 reference picture list or the list 1 reference picture list.

8. The method according to claim 1, wherein: The format rule further specifies that, in response to the current sub-picture being a random access decodable preceding sub-picture, active entries of the reference picture list of the current slice are prohibited from including any one or more of the following: a seventh picture comprising a random access skipped preceding sub-picture having the same sub-picture index as the current sub-picture, and An eighth picture precedes, in decoding order, a ninth picture including an intra random access point sub-picture associated with the random access decodable preceding sub-picture.

9. The method according to claim 8, wherein: The active entry corresponds to an entry that can be used as a reference index in inter prediction of the current slice, and The reference picture list includes the list 0 reference picture list or the list 1 reference picture list.

10. The method according to claim 1, wherein: Performing the conversion includes encoding the video into the bitstream.

11. The method according to claim 1, wherein: Performing the conversion includes decoding the video from the bitstream.

12. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: performing conversion between a video comprising one or more pictures including one or more sub-pictures and a bitstream of said video, in, The bitstream conforms to a format rule that specifies that when a sub-picture is not a leading sub-picture of an intra random access point sub-picture, the sub-picture cannot be a random access decodable leading sub-picture or a random access skipped leading sub-picture, and Wherein, the preceding sub-picture precedes the random access point sub-picture in the frame according to the output order; wherein the format rule specifies an order in which pictures are indicated in the bitstream, The format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture according to a first order or a second order, wherein the second picture comprises an intra random access point sub-picture, the intra random access point sub-picture having zero or more associated preceding sub-pictures and having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra-frame random access point sub-picture and the zero or more associated preceding sub-pictures in the first order and the second order; The first order is a decoding order, the second order is an output order, and the reference picture list includes a list 0 reference picture list or a list 1 reference picture list.

13. The device according to claim 12, wherein: The random access point sub-picture is a sub-picture whose all video codec layer (VCL) network abstraction layer (NAL) units have the same value of NAL unit type in the range of IDR_W_RADL to CRA_NUT, including IDR_W_RADL and CRA_NUT, wherein the random access decodable leading sub-picture is a sub-picture with a NAL unit type equal to RADL_NUT for each VCL NAL unit, and The random access skipped leading sub-picture is a sub-picture whose NAL unit type is equal to RASL_NUT in each view VCL NAL.

14. The device according to claim 12, wherein: The format rule further specifies that, in response to the following situation, the first sub-picture precedes the second sub-picture in the recovery point picture in the output order: The first sub-picture and the second sub-picture have the same layer identifier and the same sub-picture index of a network abstraction layer (NAL) unit, and The first sub-picture precedes the second sub-picture in the decoding order.

15. The device according to claim 12, wherein, When the current sub-picture is a completely random access sub-picture, the format rule does not allow an entry in a reference picture list of the current slice to include a third picture preceding the current picture according to a third order, a fourth picture preceding the current picture according to a fourth order, wherein the fourth picture includes an intra random access point sub-picture having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and Wherein, the completely random access sub-picture is an intra-frame random access point sub-picture in which each VCL NAL unit has nal_unit_type equal to CRA_NUT; wherein the third order includes the decoding order or the output order, wherein the fourth order comprises a decoding order, and The reference picture list includes the list 0 reference picture list or the list 1 reference picture list.

16. The device according to claim 12, wherein, The format rule does not allow an active entry in a reference picture list of the current slice to include a fifth picture that precedes a sixth picture according to a fifth order, wherein the sixth picture includes an intra random access point sub-picture having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture is after the random access point sub-picture in the frame according to the sixth order; wherein the fifth order includes the decoding order or the output order, and the sixth order includes the decoding order and the output order, wherein the active entry corresponds to an entry that can be used as a reference index in inter prediction of the current slice, and The reference picture list includes the list 0 reference picture list or the list 1 reference picture list.

17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: performing conversion between a video comprising one or more pictures including one or more sub-pictures and a bitstream of said video, in, The bitstream conforms to a format rule that specifies that when a sub-picture is not a leading sub-picture of an intra random access point sub-picture, the sub-picture cannot be a random access decodable leading sub-picture or a random access skipped leading sub-picture, and Wherein, the preceding sub-picture precedes the random access point sub-picture in the frame according to the output order; wherein the format rule specifies an order in which pictures are indicated in the bitstream, The format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture according to a first order or a second order, wherein the second picture comprises an intra random access point sub-picture, the intra random access point sub-picture having zero or more associated preceding sub-pictures and having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra-frame random access point sub-picture and the zero or more associated preceding sub-pictures in the first order and the second order; The first order is a decoding order, the second order is an output order, and the reference picture list includes a list 0 reference picture list or a list 1 reference picture list.

18. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing device, wherein the method comprises: generating said bitstream of said video comprising one or more pictures including one or more sub-pictures, wherein the bitstream complies with a format rule that specifies that when a sub-picture is not a preceding sub-picture of an intra-frame random access point sub-picture, the sub-picture cannot be a random access decodable preceding sub-picture or a random access skipped preceding sub-picture, and Wherein, the preceding sub-picture precedes the random access point sub-picture in the frame according to the output order; wherein the format rule specifies an order in which pictures are indicated in the bitstream, The format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture according to a first order or a second order, wherein the second picture comprises an intra random access point sub-picture, the intra random access point sub-picture having zero or more associated preceding sub-pictures and having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra-frame random access point sub-picture and the zero or more associated preceding sub-pictures in the first order and the second order; The first order is a decoding order, the second order is an output order, and the reference picture list includes a list 0 reference picture list or a list 1 reference picture list.

19. A method for storing a bitstream of a video, comprising: generating said bitstream of said video comprising one or more pictures including one or more sub-pictures, and storing the bitstream in a non-transitory computer-readable recording medium, in, The bitstream conforms to a format rule that specifies that when a sub-picture is not a leading sub-picture of an intra random access point sub-picture, the sub-picture cannot be a random access decodable leading sub-picture or a random access skipped leading sub-picture, and Wherein, the preceding sub-picture precedes the random access point sub-picture in the frame according to the output order; wherein the format rule specifies an order in which pictures are indicated in the bitstream, The format rule does not allow an entry in a reference picture list of the current slice to include a first picture that precedes a second picture according to a first order or a second order, wherein the second picture comprises an intra random access point sub-picture, the intra random access point sub-picture having zero or more associated preceding sub-pictures and having the same layer identifier of a network abstraction layer (NAL) unit and the same sub-picture index as the current sub-picture, and wherein the current sub-picture follows the intra-frame random access point sub-picture and the zero or more associated preceding sub-pictures in the first order and the second order; The first order is a decoding order, the second order is an output order, and the reference picture list includes a list 0 reference picture list or a list 1 reference picture list.

20. A video decoding apparatus comprising a processor configured to implement the method according to any one of claims 1-11.

21. A video encoding apparatus comprising a processor configured to implement the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • An apparatus, a method and a computer program for video coding and decoding

    CN104604223A

  • Signaling of regions of interest and gradual decoding refresh in video coding

    CN104823449A