Encoding and decoding of adjacent sub-pictures

By clarifying the sub-picture type and decoding order constraints, the problems of low video encoding and decoding efficiency and errors in the prior art are solved, and more efficient and accurate video encoding and decoding are achieved.

CN115380530BActive Publication Date: 2025-06-06DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180022907.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2021-03-18
Publication Date
2025-06-06
Estimated Expiration
2041-03-18

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have confusion and insufficient constraints when dealing with multiple sub-picture types and complex decoding sequence relationships, resulting in inefficient encoding and decoding and errors.

Method used

By defining clear sub-image types and decoding order constraints, such as allowing multiple different types of VCL NAL units in the picture, adjusting the output order of the pre-image and post-images, limiting the association relationship between IDR and RASL sub-images, etc.

Benefits of technology

It improves the efficiency and accuracy of video encoding and decoding, allows more complex video structures to be effectively coded, and reduces error rate and calculation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115380530B_ABST
    Figure CN115380530B_ABST
Patent Text Reader

Abstract

Methods and apparatus for processing video are described. The processing may include video encoding, decoding, or transcoding. An example video processing method includes performing conversion between a video including a picture including two adjacent sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that two adjacent sub-pictures having different types of network abstraction layer (NAL) units have a syntax element having a same first value, the syntax element indicating whether each of the two adjacent sub-pictures in a codec layer video sequence is considered a picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on International Patent Application No. PCT / US2021 / 022978 filed on March 18, 2021, which claims priority and benefits to U.S. Provisional Patent Application No. 62 / 992,724 filed on March 20, 2020. All of the above patent applications are hereby incorporated by reference in their entirety. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Art

[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth requirements for digital video usage are expected to continue to grow. Summary of the invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of video using various grammar rules.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies a syntax of a network abstraction layer (NAL) unit in the bitstream, and wherein the format rule specifies that a NAL unit of a video codec layer (VCL) NAL unit type includes content associated with a specific type of picture or a specific type of sub-picture.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture including a sub-picture and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that in response to the sub-picture being a leading sub-picture of an intra random access point sub-picture, the sub-picture is a random access type sub-picture.

[0008] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture including a sub-picture and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that in response to one or more random access skip leading sub-pictures being associated with an instantaneous decoding refresh sub-picture, the one or more random access skip leading sub-pictures are not present in the bitstream.

[0009] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture including a sub-picture and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that in response to one or more random access decodable leading sub-pictures being associated with an immediate decoding refresh sub-picture having a type of network abstraction layer (NAL) unit indicating that the immediate decoding refresh sub-picture is not associated with a preceding picture, the one or more random access decodable leading sub-pictures are not present in the bitstream.

[0010] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture including two adjacent sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that two adjacent sub-pictures having different types of network abstraction layer (NAL) units have a syntax element having a same first value, the syntax element indicating whether each of the two adjacent sub-pictures in a codec layer video sequence is considered a picture.

[0011] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture including two adjacent sub-pictures and a bitstream of the video, wherein the format rule specifies that the two adjacent sub-pictures include a first adjacent sub-picture having a first sub-picture index and a second adjacent sub-picture having a second sub-picture index, and wherein the format rule specifies that in response to a first syntax element associated with the first sub-picture index indicating that the first adjacent sub-picture is not considered as a picture or a second syntax element associated with the second sub-picture index indicating that the second adjacent sub-picture is not considered as a picture, the two adjacent sub-pictures have the same type of network abstraction layer (NAL) unit.

[0012] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a picture including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule providing that, in response to a syntax element indicating a reference picture parameter set (PPS), each picture of the video has a plurality of video codec layer (VCL) network abstraction layer (NAL) units without the same type of VCL NAL units, allowing the picture to include more than two different types of VCL NAL units.

[0013] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a post-subpicture associated with an intra random access point subpicture or a gradual decoding refresh subpicture in sequence follows the intra random access point subpicture or the gradual decoding refresh subpicture.

[0014] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule providing that in response to: (1) the sub-picture precedes the intra random access point sub-picture in the second order, (2) the sub-picture and the intra random access point sub-picture have the same first value for a layer to which network abstraction layer (NAL) units of the sub-picture and the intra random access point sub-picture belong, and (3) the sub-picture and the intra random access point sub-picture have the same second value of a sub-picture index, the sub-picture precedes the intra random access point sub-picture and one or more random access decodable pre-sub-pictures associated with the intra random access point sub-picture in the first order.

[0015] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a random access skipped leading sub-picture associated with a clean random access sub-picture precedes one or more random access decodable leading sub-pictures associated with the clean random access sub-picture in order.

[0016] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a random access skip preamble subpicture associated with a pure random access subpicture in a first order follows one or more intra random access point subpictures preceding the pure random access subpicture in a second order.

[0017] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule providing that in response to: (1) a syntax element indicating that a codec layer video sequence conveys pictures representing a frame, and (2) a current sub-picture is a leading sub-picture associated with an intra random access point sub-picture, and the current sub-picture precedes one or more non-leading sub-pictures associated with the intra random access point sub-picture in decoding order.

[0018] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule, the format rule specifying that, in response to the picture being a preceding picture of an intra random access point picture, one or more types of NAL units of all video codec layer (VCL) network abstraction layer (NAL) units in the picture include RADL_NUT or RASL_NUT.

[0019] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including a plurality of sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule providing that in response to: (1) at least one sub-picture precedes a progressively decoded refresh sub-picture in a second order, (2) the at least one sub-picture and the progressively decoded refresh sub-picture have the same first value for a layer to which a network abstraction layer (NAL) unit of the at least one sub-picture and the progressively decoded refresh sub-picture belongs, and (3) the at least one sub-picture and the progressively decoded refresh picture have the same second value of a sub-picture index, the at least one sub-picture precedes the progressively decoded refresh sub-picture and one or more sub-pictures associated with the progressively decoded refresh sub-picture in a first order.

[0020] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to: (a) a first picture having the same temporal identifier and the same layer identifier of a network abstraction layer (NAL) unit as the current sub-picture, and (b) the current sub-picture is after a stepwise temporal sub-layer access sub-picture in decoding order, and (c) the current sub-picture and the stepwise temporal sub-layer access sub-picture have the same temporal identifier, the same layer identifier, and the same sub-picture index, the format rule not allowing an active entry in a reference picture list of the current slice to include a first picture that is before a second picture in decoding order, the second picture including the stepwise temporal sub-layer access sub-picture.

[0021] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, and in response to the current sub-picture not being a sub-picture of a particular type, the format rule does not allow an active entry in a reference picture list of the current slice to include a first picture generated by a decoding process that generates an unavailable reference picture.

[0022] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the sub-picture including a current slice, wherein the bitstream complies with a format rule, and in response to the current sub-picture not being a sub-picture of a particular type, the format rule does not allow an entry in a reference picture list for the current slice to include a first picture generated by a decoding process that generates an unavailable reference picture.

[0023] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to (a) a first picture including a previous intra random access point sub-picture, the previous intra random access point sub-picture preceding the current sub-picture in a second order, (b) the previous intra random access point sub-picture has a layer identifier of a network abstraction layer (NAL) unit and a same sub-picture index as the current sub-picture, and (c) the current sub-picture is a pure random access sub-picture, the format rule not allowing an entry in a reference picture list of the current slice to include a first picture preceding the current picture in a first order or a second order.

[0024] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to (a) the current sub-picture is associated with an intra random access point sub-picture, and (b) the current sub-picture is after the intra random access point sub-picture in a first order, the format rule does not allow an active entry in a reference picture list of the current slice to include a first picture that is before the current picture in a first order or a second order.

[0025] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to: (a) the current sub-picture is after an intra random access point sub-picture in a first order, (b) the current sub-picture is after one or more preceding sub-pictures associated with an IRAP sub-picture in a first order and a second order, the format rule not allowing an entry in a reference picture list of the current slice to include a first picture that is before the current picture including the intra random access point sub-picture associated with the current sub-picture in the first order or the second order.

[0026] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream conforms to a format rule, the format rule providing that, in response to the current sub-picture being a random access decodable leading sub-picture, a reference picture list of the current slice does not include any one or more of the following active entries: a first picture including a random access skipped leading sub-picture and a second picture preceding, in decoding order, a third picture including an associated intra-frame random access point sub-picture.

[0027] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video. The codec representation complies with a format rule that specifies that the one or more pictures including the one or more sub-pictures are included in the codec representation according to a network abstraction layer (NAL) unit, wherein a type of NAL unit indicated in the codec representation includes a codec slice of a specific type of picture or a codec slice of a specific type of sub-picture.

[0028] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies that two adjacent sub-pictures having different network abstraction layer unit types will have the same indication of a sub-picture that is considered a picture flag.

[0029] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule, the format rule defining an order of a first type of sub-picture and a second type of sub-picture, wherein the first sub-picture is a post-sub-picture or a pre-sub-picture or a random access skip pre-sub (RASL) sub-picture type, and the second sub-picture is a RASL type or a random access decodable pre-sub (RADL) type or an instantaneous decoding refresh (IDR) type or a gradual decoding refresh (GDR) type sub-picture.

[0030] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that defines conditions for allowing or not allowing a sub-picture of a first type to exist with a sub-picture of a second type.

[0031] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.

[0032] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.

[0033] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor executable code.

[0034] These and other features will be described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan strips.

[0036] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0037] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0038] Figure 4 A picture partitioned into 15 slices, 24 slices and 24 sub-pictures is shown.

[0039] Figure 5 is a block diagram of an example video processing system.

[0040] Figure 6 is a block diagram of a video processing device.

[0041] Figure 7 is a flow chart of an example method of video processing.

[0042] Figure 8 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0043] Fig. 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0044] Fig.10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0045] Figures 11 to 31 is a flow chart of an example method of video processing. DETAILED DESCRIPTION

[0046] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed technology. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes to the text relative to the current draft of the VVC specification are shown by left and right square brackets (e.g., [[ ]]), deleted text in square brackets indicates obsolete text, and bold italics indicates added text.

[0047] 1. Initial Discussion

[0048] This document is about video codec technology. Specifically, it is about the definition of sub-picture types and their relationships in terms of decoding order, output order, and prediction relationships between sub-pictures of different types in both single-layer and multi-layer contexts. The key is to clearly specify the meaning of mixed sub-picture types within a picture through a set of constraints on decoding order, output order, and prediction relationships. These ideas can be applied, alone or in various combinations, to any video codec standard or non-standard video codec that supports multi-layer video codec, such as the Versatile Video Codec (VVC) under development.

[0049] 2. Abbreviations

[0050] APS (Adaptation Parameter Set) Adaptive Parameter Set

[0051] AU (Access Unit) Access Unit

[0052] AUD (Access Unit Delimiter) Access Unit Delimiter

[0053] AVC (Advanced Video Coding) Advanced Video Codec

[0054] CLVS (Coded Layer Video Sequence) Coded Layer Video Sequence

[0055] CPB (Coded Picture Buffer) Coded Picture Buffer

[0056] CRA (Clean Random Access) Pure Random Access

[0057] CTU (Coding Tree Unit)

[0058] CVS (Coded Video Sequence) Coded Video Sequence

[0059] DCI (Decoding Capability Information) Decoding Capability Information

[0060] DPB (Decoded Picture Buffer) Decoded Picture Buffer

[0061] EOB (End Of Bitstream) bitstream end

[0062] EOS (End Of Sequence) sequence end

[0063] GDR (Gradual Decoding Refresh) Gradual Decoding Refresh

[0064] HEVC (High Efficiency Video Coding) High Efficiency Video Coding

[0065] HRD (Hypothetical Reference Decoder) Hypothetical Reference Decoder

[0066] IDR (Instantaneous Decoding Refresh)

[0067] JEM (Joint Exploration Model)

[0068] MCTS (Motion-Constrained Tile Sets)

[0069] NAL (Network Abstraction Layer) Network Abstraction Layer

[0070] OLS (Output Layer Set) Output Layer Set

[0071] PH (Picture Header) Picture Header

[0072] PPS (Picture Parameter Set) Picture Parameter Set

[0073] PTL (Profile, Tier and Level) Profile, Tier and Level

[0074] PU (Picture Unit)

[0075] RADL (Random Access Decodable Leading (Picture)) Random Access Decodable Leading (Picture)

[0076] RAP (Random Access Point)

[0077] RASL (Random Access Skipped Leading (Picture)) Random Access Skipped Leading (Picture)

[0078] RBSP (Raw Byte Sequence Payload)

[0079] RPL (Reference Picture List) Reference Picture List

[0080] SEI (Supplemental Enhancement Information) Supplemental Enhancement Information

[0081] SPS (Sequence Parameter Set) Sequence Parameter Set

[0082] STSA (Step-wise Temporal Sublayer Access)

[0083] SVC (Scalable Video Coding) Scalable Video Codec

[0084] VCL (Video Coding Layer) Video Coding Layer

[0085] VPS (Video Parameter Set) Video Parameter Set

[0086] VTM (VVC Test Model) VVC Test Model

[0087] VUI (Video Usability Information) Video Usability Information

[0088] VVC (Versatile Video Coding) Multi-functional video codec

[0089] 3. Introduction to video encoding and decoding

[0090] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations worked together to develop H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec structure, where temporal prediction plus transform codec is utilized. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal for new codec standards is to reduce the bit rate by 50% compared to HEVC. The new video codec standard was officially named Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts to contribute to VVC standardization, new codec technologies are adopted into the VVC standard at each JVET meeting. Then, the working draft of VVC and the test model VTM are updated after each meeting. The VVC project now strives to be technically completed (FDIS) at the July 2020 meeting.

[0091] 3.1. Picture segmentation scheme in HEVC

[0092] HEVC includes four different picture partitioning schemes, namely normal slice, dependent slice, tile and Wavefront Parallel Processing (WPP), which can be applied to Maximum Transfer Unit (MTU) size matching, parallel processing and reduced end-to-end delay.

[0093] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although there may still be interdependencies due to loop filter operations).

[0094] Normal slices are the only tool that can be used for parallelization, and are also available in H.264 / AVC in almost the same form. Parallelization based on normal slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive codec pictures, which is usually much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, the use of normal slices incurs a large codec overhead due to the bit cost of the slice header and due to the lack of prediction across slice boundaries. In addition, due to the intra-picture independence of normal slices and the fact that each normal slice is encapsulated in its own NAL unit, normal slices (in contrast to the other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting requirements on the layout of slices within a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.

[0095] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Basically, dependent slices provide fragmentation of a regular slice into multiple NAL units to provide reduced end-to-end latency by allowing a portion of a regular slice to be sent out before encoding of the entire regular slice is complete.

[0096] In WPP, a picture is partitioned into a single row of codec tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding of CTB rows is delayed by two CTBs to ensure that data related to CTBs above and to the right of the main CTB can be obtained before the main CTB being decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores as pictures containing CTB rows can be parallelized. Because intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication required to enable intra-picture prediction can be substantial. WPP partitioning does not result in the generation of additional NAL units compared to when it is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP, with some degree of codec overhead.

[0097] Slices define the horizontal and vertical boundaries that divide the image into slice columns and slice rows. Slice columns extend from the top of the image to the bottom of the image. Similarly, slice rows extend from the left side of the image to the right side of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.

[0098] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of the CTB raster scan of the slice). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in separate NAL units (same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where the slice spans multiple slices, the inter-processor / inter-core communication required to decode intra-picture prediction between processing units of adjacent slices is limited to transmitting a shared slice header and loop filtering associated with sharing of reconstruction samples and metadata. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first is signaled in the slice header.

[0099] For simplicity, restrictions on the application of four different picture partitioning schemes have been specified in HEVC. A given codec video sequence cannot include both slices and wavefronts of most profiles specified in HEVC. For each slice and slice, one or both of the following conditions must be met: 1) all codec tree blocks in a slice belong to the same slice; 2) all codec tree blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is being used, if a slice starts within a CTB row, it must end in the same CTB row.

[0100] The latest revisions to HEVC are specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), “HEVC Additional Supplemental Enhancement Information (Draft 4)”, publicly released on October 24, 2017:

[0101] http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip By including this amendment, HEVC specifies three MCTS-related SEI messages, namely, the time-domain MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nesting SEI message.

[0102] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vector is restricted to point to the full sample position inside the MCTS and the fractional sample position of the full sample position inside the MCTS that is only needed for interpolation, and the use of motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction is not allowed. In this way, each MCTS can be independently decoded in the absence of slices not included in the MCTS.

[0103] The MCTS extraction information set SEI message provides supplementary information (defined as part of the semantics of the SEI message) that can be used in MCTS sub-bitstream extraction to generate conforming bitstreams for MCTS groups. The information consists of many extraction information sets, each defining many MCTS groups and containing RBSP bytes that replace VPS, SPS and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all slice address related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.

[0104] 3.2. VVC Image Segmentation

[0105] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​a picture. The CTUs in a slice are scanned in a raster scan order within a slice.

[0106] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of a picture.

[0107] Two strip modes are supported, namely raster scan strip mode and rectangular strip mode. In raster scan strip mode, a strip contains a complete sequence of slices in a slice raster scan of a picture. In rectangular strip mode, a strip contains a number of complete slices that together form a rectangular area of ​​a picture, or a number of consecutive complete CTU rows that together form a slice of a rectangular area of ​​a picture. The slices within a rectangular strip are scanned in a slice raster scan order within the rectangular area corresponding to the strip.

[0108] A sub-picture consists of one or more strips that together cover a rectangular area of ​​the picture.

[0109] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan strips.

[0110] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0111] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0112] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, 12 slices on the left (each covering a strip of 4x4 CTUs) and 6 slices on the right (each covering 2 vertically stacked strips of 2x2 CTUs), resulting in a total of 24 strips and 24 sub-pictures of different dimensions (each strip is a sub-picture).

[0113] 3.3. Changes in image resolution within a sequence

[0114] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC enables changing the resolution of pictures within a sequence at locations where IRAP pictures are not encoded, which are always intra-coded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling the reference pictures used for inter prediction when the reference picture has a different resolution than the current picture being decoded.

[0115] The scaling ratio is restricted to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luminance and 32 phases for chrominance, which is the same as the case of motion compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process with scaling ratios ranging from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.

[0116] Other aspects of the VVC design that support this feature that differs from HEVC include: i) Picture resolution and corresponding consistency window are signaled in the PPS instead of in the SPS, where the maximum picture resolution is signaled. ii) For a single-layer bitstream, each picture store (a slot in the DPB for the storage of one decoded picture) occupies the buffer size required to store decoded pictures with the maximum picture resolution.

[0117] 3.4. Scalable Video Codec (SVC) in General and in VVC

[0118] Scalable video codec (SVC, sometimes also referred to as scalability in video codecs) refers to a video codec in which a base layer (BL), sometimes referred to as a reference layer (RL), and one or more scalable enhancement layers (EL) are used. In SVC, the base layer can carry video data with a base level of quality. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal and / or signal-to-noise (SNR) levels. The enhancement layers can be defined relative to previously coded layers. For example, the bottom layer can be used as the BL, while the top layer can be used as the EL. The middle layer can act as the EL or the RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest layer nor the highest layer) can be the EL of a layer below the intermediate layer (e.g., the base layer or any intervening enhancement layer), and at the same time, the RL of one or more enhancement layers above the intermediate layer. Similarly, in a multi-view or 3D extension of the HEVC standard, there may be multiple views, and information from one view may be used to encode (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0119] In SVC, parameters used by an encoder or decoder are grouped into parameter sets in which they can be utilized based on the codec level (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be utilized by one or more codec video sequences of different layers in a bitstream can be included in a video parameter set (VPS), and parameters utilized by one or more pictures in a codec video sequence can be included in a sequence parameter set (SPS). Similarly, parameters utilized by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters specific to a single slice can be included in a slice header. Similarly, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.

[0120] Due to the support for reference picture resampling (RPR) in VVC, support for bitstreams containing multiple layers, for example, two layers with SD and HD resolutions in VVC, can be designed to not require any additional signaling level codec tools, because the upsampling required for spatial scalability support can use only RPR upsampling filters. However, a high level of syntax changes are required for scalability support (compared to not supporting scalability). Scalability support is specified in VVC version 1. Unlike scalability support in any earlier video codec standards (including extensions of AVC and HEVC), the design of VVC scalability is as friendly to single-layer decoder design as possible. The decoding capabilities of multi-layer bitstreams are specified as if there is only a single layer in the bitstream. For example, the decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not require much change to be able to decode multi-layer bitstreams. Compared with the multi-layer extension design of AVC and HEVC, the HLS aspect has been significantly simplified at the expense of some flexibility. For example, the IRAP AU needs to contain pictures of each layer present in the CVS.

[0121] 3.5. Random access and its support in HEVC and VVC

[0122] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. In order to support tuning and channel switching in broadcast / multicast and multi-party video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points, which are usually intra-frame coded pictures, but can also be inter-frame coded pictures (for example, in the case of gradual decoding refresh).

[0123] HEVC includes signaling of intra random access point (IRAP) pictures in the NAL unit header via the NAL unit type. Three types of IRAP pictures are supported, namely instantaneous decoder refresh (IDR), pure random access (CRA), and broken link access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any pictures before the current group-of-pictures (GOP), which is often referred to as a closed-GOP random access point. CRA pictures are less restrictive by allowing certain pictures to reference pictures before the current GOP, and in the case of random access, all pictures are discarded. CRA pictures are often referred to as open-GOP random access points. BLA pictures typically originate from the concatenation of two bitstreams or parts thereof in a CRA picture, such as during stream switching. To enable better system usage of IRAP pictures, a total of six different NAL units are defined to signal the properties of IRAP pictures, which can be used to better match the stream access point type defined in the ISO base media file format (ISOBMFF), which is used for random access support in dynamic adaptive streaming over HTTP (DASH).

[0124] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or without associated RADL pictures), and one type of CRA pictures. These are essentially the same as HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) The basic functionality of the BLA picture can be implemented by a CRA picture plus a sequence end NAL unit, the presence of which indicates that the subsequent picture starts a new CVS in a single-layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than HEVC, as indicated by using five bits instead of six bits for the NAL unit type field in the NAL unit header.

[0125] Another key difference between VVC and HEVC in random access support is that GDR is supported in a more standardized way in VVC. In GDR, decoding of the bitstream can start from an inter-coded picture, and although not the entire picture area can be correctly decoded at the beginning, after multiple pictures, the entire picture area will be correct. AVC and HEVC also support GDR, using the recovery point SEI message to signal the GDR random access point and recovery point. In VVC, a new NAL unit type is specified to indicate a GDR picture, and the recovery point is signaled in the picture header syntax structure. CVS and bitstreams are allowed to start with a GDR picture. This means that the entire bitstream is allowed to contain only inter-coded pictures, without a single intra-coded picture. The main benefit of specifying GDR support in this way is that consistent behavior is provided for GDR. GDR-enabled encoders smooth the bitrate of the bitstream by distributing intra-coded strips or blocks across multiple pictures, instead of intra-coding the entire picture, allowing for significantly reduced end-to-end latency, which is considered even more important today as ultra-low latency applications such as wireless displays, online gaming, drone-based applications, etc. become increasingly popular.

[0126] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refresh area (i.e., correctly decoded area) and the unrefreshed area at the picture between the GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, the loop filter across the boundary will not be applied, so there will be no decoding mismatch of some samples at or near the boundary. This can become useful when the application determines to display the correctly decoded area during GDR processing.

[0127] IRAP pictures and GDR pictures may be collectively referred to as random access point (RAP) pictures.

[0128] 3.6. Reference Picture Management and Reference Picture List (RPL)

[0129] Reference picture management is a core function required for any video codec using inter-frame prediction. It manages the storage and removal of reference pictures in the decoded picture buffer (DPB) and places the reference pictures in the RPL in their correct order.

[0130] The reference picture management of HEVC, including reference picture marking and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC), is different from that of AVC. Instead of the reference picture marking mechanism based on sliding window plus adaptive memory management control operation (MMCO) in AVC, HEVC specifies a reference picture management and marking mechanism based on the so-called reference picture set (RPS), and therefore RPLC is based on the RPS mechanism. The RPS consists of a set of reference pictures associated with a picture, including all reference pictures before the associated picture in decoding order, which can be used for inter-frame prediction of the associated picture or any picture after the associated picture in decoding order. The reference picture set includes five reference picture lists. The first three lists contain all reference pictures that can be used for inter-frame prediction of the current picture and can be used for inter-frame prediction of one or more pictures after the current picture in decoding order. The other two lists consist of all reference pictures that are not used in inter-frame prediction of the current picture but can be used for inter-frame prediction of one or more pictures after the current picture in decoding order. RPS provides "intra-frame codec" signaling notification of DPB status, rather than "inter-frame codec" signaling notification like in AVC, mainly to improve error tolerance. The RPLC process in HEVC is based on RPS, by signaling the index for each reference index to the RPS subgroup; this process is simpler than the RPLC process in AVC.

[0131] Reference picture management in VVC is more similar to HEVC than AVC, but simpler and more robust. As in those standards, two RPLs are derived, List 0 and List 1, but they are not based on the reference picture set concept used in HEVC or the automatic sliding window processing used in AVC; instead, they are signaled more directly. The reference pictures of the RPL are listed as active and inactive entries, and only active entries can be used as reference indexes for inter prediction of CTUs of the current picture. Inactive entries indicate other pictures to be hosted in the DPB for reference by other pictures arriving later in the bitstream.

[0132] 3.7. Parameter Set

[0133] AVC, HEVC, and VVC specify parameter sets. Types of parameter sets include SPS, PPS, APS, and VPS. AVC, HEVC, and VVC all support SPS and PPS. VPS was introduced starting with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.

[0134] SPS is designed to carry sequence-level header information, while PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, information that does not change frequently does not need to be repeated for each sequence or picture, so redundant signaling of that information can be avoided. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, which not only avoids the need for redundant transmission but also improves fault tolerance.

[0135] VPS is introduced to carry sequence level header information common to all layers in a multi-layer bitstream.

[0136] APS is introduced to carry such picture-level or slice-level information, which requires quite a lot of bits to encode and decode, can be shared by multiple pictures, and can have quite a lot of different variations in a sequence.

[0137] 3.8. Related definitions in VVC

[0138] The relevant definitions in the latest VVC text (JVET-Q2001-vE / v15) are as follows.

[0139] Associated IRAP picture (of a specific picture): The previous IRAP picture (when present) in decoding order has the same value of nuh_layer_id as that of the specific picture.

[0140] Pure Random Access (CRA) PU: The coded picture is a PU of a CRA picture.

[0141] Pure Random Access (CRA) pictures: Each VCL NAL unit includes an IRAP picture with nal_unit_type equal to CRA_NUT.

[0142] Coded Video Sequence (CVS): A sequence of AUs consisting of a CVSS AU, followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs, but excluding any subsequent AU that is a CVSS AU, in decoding order.

[0143] Coded Video Sequence Start (CVSS) AU: Each layer in CVS has PUs and the coded pictures in each PU are AUs of CLVSS pictures.

[0144] Gradual Decoding Refresh (GDR) AU: Each AU in which the codec picture in the current PU is a GDR picture.

[0145] Gradual Decoding Refresh (GDR) PU: The PU of the coded picture is a GDR picture.

[0146] Gradual Decoding Refresh (GDR) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to GDR_NUT.

[0147] Instantaneous Decoding Refresh (IDR) PU: A PU whose coded picture is an IDR picture.

[0148] Instantaneous Decoding Refresh (IDR) picture: Each VCL NAL unit includes an IRAP picture with nal_unit_type equal to IDR_W_RADL or IDR_N_LP.

[0149] Intra-frame random access point (IRAP) AU: Each layer in CVS has a PU and the coded picture in each PU is an AU of an IRAP picture.

[0150] Intra Random Access Point (IRAP) PU: A PU whose coded picture is an IRAP picture.

[0151] Intra Random Access Point (IRAP) picture: A codec picture in which all VCL NAL units contain the same value of nal_unit_type in the range IDR_W_RADL to CRA_NUT, inclusive.

[0152] Preceding picture: A picture that is in the same layer as the associated IRAP picture and precedes the associated IRAP picture in output order.

[0153] Output order: The order in which decoded pictures are output from the DPB (for decoded pictures to be output from the DPB).

[0154] Random Access Decodable Front (RADL) PU: A PU whose coded picture is a RADL picture.

[0155] Random Access Decodable Lead (RADL) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to RADL_NUT.

[0156] Random Access Skip Preamble (RASL) PU: The PU whose coded picture is a RASL picture.

[0157] Random Access Skip Leading (RASL) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to RASL_NUT.

[0158] Step-by-step temporal sublayer access (STSA) PU: The coded picture is a PU of a STSA picture.

[0159] Step-by-Step Temporal Sublayer Access (STSA) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to STSA_NUT.

[0160] Note – STSA pictures do not use pictures with the same TemporalId as the STSA picture for inter prediction reference. Pictures with the same TemporalId as the STSA picture that follow the STSA picture in decoding order do not use pictures with the same TemporalId as the STSA picture that precede the STSA picture in decoding order for inter prediction reference. STSA pictures enable switching from the immediately lower sublayer at the STSA picture to the sublayer containing the STSA picture. The TemporalId of an STSA picture must be greater than 0.

[0161] Sub-picture: A rectangular area of ​​one or more strips within a picture.

[0162] Post-picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not a STSA picture.

[0163] Note – Post-pictures associated with an IRAP picture also follow the IRAP picture in decoding order. Pictures that follow the associated IRAP picture in output order and precede the associated IRAP picture in decoding order are not allowed.

[0164] 3.9. NAL unit header syntax and semantics in VVC

[0165] In the latest VVC text (in JVET-Q2001-vE / v15), the NAL unit header syntax and semantics are as follows.

[0166] 7.3.1.2 NAL unit header syntax

[0167]

[0168]

[0169] 7.4.2.2 NAL unit header semantics

[0170] forbidden_zero_bit should be equal to 0.

[0171] nuh_reserved_zero_bit shall be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified by ITU-T|ISO / IEC in the future. The decoder shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1.

[0172] nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which the non-VCL NAL unit applies. The value of nuh_layer_id shall be in the range of 0 to 55 (inclusive). Other values ​​of nuh_layer_id are reserved for future use by ITU-T | ISO / IEC.

[0173] The value of nuh_layer_id shall be the same for all VCL NAL units of a codec picture. The value of nuh_layer_id of a codec picture or PU is the value of nuh_layer_id of the VCL NAL unit of the codec picture or PU.

[0174] The nuh_layer_id value for AUD, PH, EOS, and FD NAL units is subject to the following constraints:

[0175] – If nal_unit_type is equal to AUD_NUT, nuh_layer_id shall be equal to vps_layer_id[0].

[0176] – Otherwise, when nal_unit_type is equal to PH_NUT, EOS_NUT or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit.

[0177] NOTE 1 – The nuh_layer_id value of DCI, VPS, and EOB NAL units is not constrained.

[0178] The value of nal_unit_type shall be the same for all pictures of a CVSS AU.

[0179] nal_unit_type specifies the NAL unit type, ie the type of the RBSP data structure contained in the NAL unit as specified in Table 5.

[0180] NAL units with unspecified semantics and nal_unit_type in the range UNSPEC_28..UNSPEC_31 (inclusive) shall not affect the decoding process specified in this specification.

[0181] NOTE 2 - NAL unit types in the range UNSPEC_28..UNSPEC_31 MAY be used as determined by the application. The decoding treatment of these values ​​of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, special care must be exercised when designing encoders that generate NAL units with these nal_unit_type values, and when designing decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define any management of these values. These nal_unit_type values ​​MAY only be used in contexts where usage "conflicts" (i.e., different definitions of the meaning of the content of NAL units for the same nal_unit_type value) are not important, or are not possible, or are managed (e.g., as defined or managed in a controlling application or transport specification, or by the context of controlling bitstream distribution).

[0182] For purposes other than determining the amount of data in a DU of a bitstream (as specified in Annex C), a decoder shall ignore (remove from the bitstream and discard) the contents of all NAL units that use the reserved value of nal_unit_type.

[0183] NOTE 3 – This requirement allows for the future definition of compatible extensions to this specification.

[0184] Table 5 – NAL unit type codes and NAL unit type classifications

[0185]

[0186]

[0187] NOTE 4 - A pure random access (CRA) picture may have an associated RASL or RADL picture present in the bitstream.

[0188] NOTE 5 – An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP has no associated preceding picture in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL has no associated RASL picture present in the bitstream.

[0189] But there may be associated RADL pictures in the bitstream.

[0190] The nal_unit_type value of all VCL NAL units in a sub-picture shall be the same. A sub-picture is said to have the same NAL unit type as the VCL NAL units of the sub-picture.

[0191] For any particular picture's VCL NAL unit, the following applies:

[0192] – If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and the picture or PU is considered to have the same NAL unit type as the codec slice NAL units of the picture or PU.

[0193] – Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture shall have at least two sub-pictures, and the VCL NAL units of the picture shall have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one sub-picture of the picture shall all have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of the other sub-pictures in the picture shall all have different specific values ​​of nal_unit_type equal to TRAIL_NUT, RADL_NUT, RASL_NUT.

[0194] For single-layer bitstreams, the following constraints apply:

[0195] - Except for the first picture in the bitstream in decoding order, each picture is considered to be associated with the previous IRAP picture in decoding order.

[0196] – When a picture is the preceding picture of an IRAP picture, it shall be a RADL or RASL picture.

[0197] – When a picture is a post-IRAP picture, it shall not be a RADL or RASL picture.

[0198] – There shall be no RASL pictures associated with an IDR picture in the bitstream.

[0199] – A RADL picture associated with an IDR picture with nal_unit_type equal to IDR_N_LP shall not be present in the bitstream.

[0200] NOTE 6 – Random access can be performed at the location of an IRAP PU (and correctly decode the IRAP picture and all subsequent non-RASL pictures in decoding order) by discarding all PUs preceding the IRAP PU, provided that each parameter set is available (either in the bitstream or by external means not specified in this specification) when referenced.

[0201] – Any pictures that precede an IRAP picture in decoding order shall precede the IRAP picture in output order, and shall precede any RADL pictures associated with the IRAP picture in output order.

[0202] – Any RASL pictures associated with a CRA picture shall precede any RADL pictures associated with the CRA picture in output order.

[0203] – Any RASL pictures associated with a CRA picture in output order shall follow any IRAP pictures that precede the CRA picture in decoding order.

[0204] – If field_seq_flag is equal to 0, and the current picture is a pre-previous picture associated with an IRAP picture, then it shall precede, in decoding order, all non-previous pictures associated with the same IRAP picture. Otherwise, let picA and picB be the first and last pre-previous pictures associated with the IRAP picture in decoding order, respectively, there shall be at most one non-previous picture before picA in decoding order, and there shall be no non-previous pictures between picA and picB in decoding order.

[0205] nuh_temporal_id_plus1 minus 1 specifies the temporal identifier of the NAL unit.

[0206] The value of nuh_temporal_id_plus1 shall not be equal to 0.

[0207] The derivation of the variable TemporalId is as follows:

[0208] TemporalId=nuh_temporal_id_plus1-1 (36)

[0209] When nal_unit_type is in the range of IDR_W_RADL to RSV_IRAP_12 (inclusive), TemporalId shall be equal to 0.

[0210] When nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, TemporalId shall not be equal to 0.

[0211] The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId of a codec picture, PU, ​​or AU is the value of TemporalId of the VCL NAL unit of the codec picture, PU, ​​or AU. The value of TemporalId of a sublayer representation is the maximum value of the TemporalId of all VCL NAL units in the sublayer representation.

[0212] The value of TemporalId for non-VCL NAL units is subject to the following constraints:

[0213] – If nal_unit_type is equal to DCI_NUT, VPS_NUT or SPS_NUT, TemporalId shall be equal to 0, and the TemporalId of the AU containing the NAL unit shall be equal to 0.

[0214] – Otherwise, if nal_unit_type is equal to PH_NUT, TemporalId shall be equal to the TemporalId of the PU containing the NAL unit.

[0215] – Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0.

[0216] – Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT or SUFFIX_SEI_NUT, TemporalId shall be equal to the TemporalId of the AU containing the NAL unit.

[0217] – Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU containing the NAL unit.

[0218] NOTE 7 – When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values ​​of all AUs to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing AU, since all PPS and APS may be included at the beginning of the bitstream (e.g., when they are transmitted out-of-band and the receiver places them at the beginning of the bitstream), where the TemporalId of the first codec picture is equal to 0.

[0219] 3.10. Mixed NAL unit types in a picture

[0220] 7.4.3.4 Picture parameter set semantics ...

[0222] mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture of the referenced PPS has more than one VCL NAL unit, the VCL NAL units do not have the same nal_unit_type value, and the picture is not an IRAP picture. mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same nal_unit_type value.

[0223] When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_types_in_pic_flag shall be equal to 0.

[0224] For each slice with a nal_unit_type value nalUnitTypeA in the range IDR_W_RADL to CRA_NUT (inclusive), in picture picA that also contains one or more slices with another value of nal_unit_type (i.e., the value of mixed_nalu_types_in_pic_flag of picture picA is equal to 1), the following applies:

[0225] – The slice shall belong to the sub-picture subpicA for which the value of subpic_treated_as_pic_flag[i] is equal to 1.

[0226] – A slice shall not belong to a picA sub-picture containing VCL NAL units with nal_unit_type not equal to nalUnitTypeA.

[0227] – If nalUnitTypeA is equal to CRA, for all subsequent PUs following the current picture in CLVS in decoding order and in output order, neither RefPicList[0] nor RefPicList[1] of slices in subpicA in these PUs shall include any pictures in the active entries that precede picA in decoding order.

[0228] – Otherwise (i.e., nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS that follow the current picture in decoding order, neither the RefPicList[0] nor the RefPicList[1] of the slices in subpicA in these PUs shall include any pictures in the active entries that precede picA in decoding order.

[0229] NOTE 1 – mixed_nalu_types_in_pic_flag equal to 1 indicates that the picture referencing the PPS contains slices with different NAL unit types, for example, a codec picture originating from a sub-picture bitstream merge operation, where the encoder must ensure bitstream structure matching and further alignment of the original bitstream parameters. An example of such alignment is as follows: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_in_pic_flag is equal to 1, the picture referencing the PPS shall not have slices with nal_unit_type equal to IDR_W_RADL or IDR_N_LP. ...

[0231] 3.11. Picture Header Structure Syntax and Semantics in VVC

[0232] In the latest VVC text (in JVET-Q2001-vE / v15), the picture header structure syntax and semantics most relevant to the technical solution of this article are as follows.

[0233] 7.3.2.7 Picture header structure syntax

[0234]

[0235]

[0236] 7.4.3.7 Image header structure semantics

[0237] The PH syntax structure contains information common to all slices of the coded picture associated with the PH syntax structure.

[0238] gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag equal to 0 specifies that the current picture may or may not be a GDR or IRAP picture.

[0239] gdr_pic_flag equal to 1 specifies that the picture associated with the PH is a GDR picture. gdr_pic_flag equal to 0 specifies that the picture associated with the PH is not a GDR picture. When not present, the value of gdr_pic_flag is inferred to be equal to 0. When gdr_enabled_flag is equal to 0, the value of gdr_pic_flag shall be equal to 0.

[0240] NOTE 1 - When gdr_or_irap_pic_flag is equal to 1 and gdr_pic_flag is equal to 0, the picture associated with the PH is an IRAP picture. ...

[0242] ph_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the ph_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits. The value of ph_pic_order_cnt_lsb shall be in the range of 0 to MaxPicOrderCntLsb-1, inclusive.

[0243] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding a CLVSS picture that is not the first picture in the bitstream as specified in Annex C.

[0244] recovery_poc_cnt specifies the recovery point of the decoded picture in output order. If the current picture is a GDR picture associated with PH, and there is a picture picA that follows the current GDR picture in decoding order in the CLVS with PicOrderCntVal, and PicOrderCntVal is equal to the value of PicOrderCntVal of the current GDR picture plus recovery_poc_cnt, then the picture picA is called the recovery point picture. Otherwise, the first picture whose PicOrderCntVal is greater than the value of PicOrderCntVal of the current picture plus recovery_poc_cnt in output order is called the recovery point picture. The recovery point picture should not be before the current GDR picture in decoding order. The value of recovery_poc_cnt should be in the range of 0 to MaxPicOrderCntLsb-1 (including this number).

[0245] When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:

[0246] RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81)

[0247] NOTE 2 - When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the RpPicOrderCntVal of the associated GDR picture, the current and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (when present) that precedes the associated GDR picture in decoding order. ...

[0249] 3.12. Constraints on RPL in VVC

[0250] In the latest VVC text (in JVET-Q2001-vE / v15), VVC's constraints on RPL are as follows (as part of the decoding process of VVC clause 8.3.2 Reference picture list construction).

[0251] 8.3.2 Decoding Processing of Reference Picture List Construction ...

[0253] For each i equal to 0 or 1, the first NumRefIdxActive[i] entries in RefPicList[i] are referred to as active entries in RefPicList[i], and the other entries in RefPicList[i] are referred to as inactive entries in RefPicList[i].

[0254] Note 2 - A particular picture may be referred to by both an entry in RefPicList[0] and an entry in RefPicList[1]. It is also possible that a particular picture is referred to by more than one entry in RefPicList[0] or more than one entry in RefPicList[1].

[0255] Note 3 – The active entries in RefPicList[0] and the active entries in RefPicList[1] refer together to all reference pictures that can be used for inter prediction of the current picture and one or more pictures after the current picture in decoding order. The inactive entries in RefPicList[0] and the inactive entries in RefPicList[1] refer together to all reference pictures that are not used for inter prediction of the current picture but can be used for inter prediction of one or more pictures after the current picture in decoding order.

[0256] NOTE 4 – One or more entries in RefPicList[0] or RefPicList[1] may be equal to “No reference picture” because the corresponding pictures are not present in the DPB. Each inactive entry in RefPicList[0] or RefPicList[0] that is equal to “No reference picture” shall be ignored. Unexpected picture loss shall be inferred for each active entry in RefPicList[0] or RefPicList[1] that is equal to “No reference picture”.

[0257] A requirement for bitstream conformance is that the following constraints apply:

[0258] – For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i].

[0259] – The picture referred to by each active entry in RefPicList[0] or RefPicList[1] shall exist in the DPB and the TemporalId shall be less than or equal to the TemporalId of the current picture.

[0260] – Each entry in RefPicList[0] or RefPicList[1] refers to a picture that shall not be the current picture and non_reference_picture_flag shall be equal to 0.

[0261] – A STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and a LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture.

[0262] – There should not be a difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture involved in this entry in RefPicList[0] or RefPicList[1] that is greater than or equal to 2 24 LTRP entry.

[0263] – Let setOfRefPics be the unique set of pictures involving all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize-1 (inclusive), where MaxDpbSize is specified in clause A.4.2 and setOfRefPics shall be the same for all slices of a picture.

[0264] – When nal_unit_type of the current slice is equal to STSA_NUT, there shall be no active entry in RefPicList[0] or RefPicList[1] with TemporalId equal to the TemporalId of the current picture and nuh_layer_id equal to the nuh_layer_id of the current picture.

[0265] – When the current picture is a picture that follows a STSA picture in decoding order, whose TemporalId is equal to the TemporalId of the current picture and whose nuh_layer_id is equal to the nuh_layer_id of the current picture, then there shall be no picture that precedes the STSA picture in decoding order, whose TemporalId is equal to the TemporalId of the current picture and whose nuh_layer_id is equal to the nuh_layer_id of the current picture, included as an active entry in RefPicList[0] or RefPicList[1].

[0266] – When the current picture is a CRA picture, there shall be no picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes any previous IRAP picture (when present) in decoding order, either in output order or in decoding order.

[0267] – When the current picture is a post-picture, there shall be no pictures referred to by active entries in RefPicList[0] or RefPicList[1] generated by the decoding process of generating unavailable reference pictures for the IRAP picture associated with the current picture.

[0268] – When the current picture is a subsequent picture that follows one or more preceding pictures associated with the same IRAP picture (if any) both in decoding order and in output order, there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] generated by the decoding process of generating unavailable reference pictures for the IRAP picture associated with the current picture.

[0269] – When the current picture is a recovery point picture or a picture that follows the recovery point picture in output order, there shall be no entry in RefPicList[0] or RefPicList[1] containing a picture generated by decoding processing of an unavailable reference picture of the GDR picture that generated the recovery point picture.

[0270] – When the current picture is a subsequent picture, there shall be no pictures referred to by active entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or in decoding order.

[0271] – When the current picture is a subsequent picture that follows one or more preceding pictures associated with the same IRAP picture (if any) both in decoding order and in output order, there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or in decoding order.

[0272] – When the current picture is a RADL picture, it is not in RefPicList[0] or RefPicList[1].

[0273] There should be any one of the following active entries:

[0274] οRASL pictures

[0275] o Pictures generated by decoding processes that generate unavailable reference pictures

[0276] o The picture that precedes the associated IRAP picture in decoding order

[0277] – Each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture refers to a picture in the same AU as the current picture.

[0278] – The picture referred to by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall exist in the DPB and the nuh_layer_id shall be less than the nuh_layer_id of the current picture.

[0279] – Each ILRP entry in either RefPicList[0] or RefPicList[1] of a stripe shall be an active entry. ...

[0281] 3.13.PictureOutputFlag settings

[0282] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting the value of the variable PictureOutputFlag is as follows (as part of the clause 8.1.2 decoding process of a codec picture).

[0283] 8.1.2 Decoding of Coded Images

[0284] The decoding process specified in this clause applies to each coded picture in BitstreamToDecode, called the current picture and represented by the variable CurrPic.

[0285] Depending on the value of chroma_format_idc, the number of sample arrays for the current picture is as follows:

[0286] – If chroma_format_idc is equal to 0, the current picture consists of 1 sample array S L composition.

[0287] – Otherwise (chroma_format_idc is not equal to 0), the current picture consists of 3 sample arrays S L , S Cb , S Cr composition.

[0288] The decoding process of the current picture takes as input the syntax elements and uppercase variables from clause 7. When interpreting the semantics of each syntax element in each NAL unit and the rest of clause 8, the term "bitstream" (or a portion thereof, such as a CVS of a bitstream) refers to BitstreamToDecode (or a portion thereof).

[0289] Depending on the value of separate_colour_plane_flag, the decoding process is structured as follows:

[0290] – If separate_colour_plane_flag is equal to 0, the decoding process is called once and the current picture is taken as output.

[0291] – Otherwise (separate_colour_plane_flag is equal to 1), the decoding process is called three times. The input to the decoding process is all NAL units of the coded picture with the same colour_plane_id value. The decoding process for NAL units with a specific colour_plane_id value is specified as if only CVSs in monochrome color format with that specific colour_plane_id value are present in the bitstream. The output of each of the three decoding processes is assigned to one of the 3 sample arrays of the current picture, and the NAL units with colour_plane_id equal to 0, 1 and 2 are assigned to S L , S Cb and S Cr .

[0292] Note – When separate_colour_plane_flag is equal to 1 and chroma_format_idc is equal to 3, the derived variable ChromaArrayType is equal to 0. In the decoding process, the value of this variable is evaluated, resulting in the same operation as for monochrome pictures (when chroma_format_idc is equal to 0).

[0293] For the current picture CurrPic, the decoding process is as follows:

[0294] 1. The decoding of NAL units is specified in clause 8.2.

[0295] 2. The processing in clause 8.3 specifies the following decoding processing using syntax elements in the slice header layer and above:

[0296] – Derives variables and functions related to the picture order count as specified in clause 8.3.1. It only needs to be called for the first slice of a picture.

[0297] – At the beginning of the decoding process of each slice of a non-IDR picture, the decoding process of the reference picture list construction specified in clause 8.3.2 is called to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).

[0298] - Invoke the decoding process of reference picture marking in clause 8.3.3, where a reference picture can be marked as "unused for reference" or "used for long-term reference". This only needs to be invoked for the first slice of a picture.

[0299] – When the current picture is a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 or a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, the decoding process for generating unusable reference pictures as specified in subclause 8.3.4 is invoked, which only needs to be invoked for the first slice of the picture.

[0300] –PictureOutputFlag is set as follows:

[0301] – PictureOutputFlag is set equal to 0 if one of the following conditions is true:

[0302] – The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.

[0303] – gdr_enabled_flag is equal to 1, and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0304] – gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.

[0305] –sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:

[0306] –PictureOutputFlag of PicA is equal to 1.

[0307] – nuh_layer_id nuhLid of PicA is greater than nuh_layer_id of the current picture.

[0308] –PicA belongs to the output layer of OLS (that is, OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).

[0309] –sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.

[0310] – Otherwise, PictureOutputFlag is set equal to pic_output_flag.

[0311] 3. The processes in clauses 8.4, 8.5, 8.6, 8.7 and 8.8 specify decoding processes using syntax elements in all syntax structure layers. A bitstream conformance requirement is that a codec slice of a picture shall contain slice data for each CTU of the picture, such that the division of the picture into slices and the division of slices into CTUs each form a partition of the picture.

[0312] 4. After all slices of the current picture are decoded, the currently decoded picture is marked as "used for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "used for short-term reference".

[0313] 4. Technical problems solved by the disclosed technical solutions

[0314] The existing design in the latest VVC text (JVET-Q2001-vE / v15) has the following problems:

[0315] 1) Since different types of sub-pictures are allowed to be mixed in one picture, it is confusing to refer to the contents of a NAL unit of the VCL NAL unit type as a codec slice of a particular type of picture. For example, a NAL unit with nal_unit_type equal to CRA_NUT is a codec slice of a CRA picture only if the nal_unit_type of all slices of the picture is equal to CRA_NUT; when the nal_unit_type of one slice of the picture is not equal to CRA_NUT, the picture is not a CRA picture.

[0316] 2) Currently, if a sub-picture contains VCL NAL units with nal_unit_type in the range of IDR_W_RADL to CRA_NUT (inclusive), the value of subpic_treated_as_pic_flag[] of the sub-picture needs to be equal to 1, and mixed_nalu_types_in_pic_flag of the picture needs to be equal to 1. In other words, for IRAP sub-pictures mixed with sub-pictures of another type in the picture, the value of subpic_treated_as_pic_flag[] needs to be equal to 1. However, with the support of more mixed VCL NAL unit types, this requirement is not sufficient.

[0317] 3) Currently, at most two different types of VCL NAL units (and two different types of sub-pictures) are allowed in a picture.

[0318] 4) In both single-layer and multi-layer contexts, there is a lack of constraints on the output order of subsequent sub-pictures relative to the associated IRAP or GDR sub-picture.

[0319] 5) Currently, it is specified that when a picture is a preceding picture of an IRAP picture, it shall be a RADL or RASL picture. This constraint, together with the definition of preceding / RADL / RASL pictures, does not allow mixing of RADL and RASL NAL unit types within a picture resulting from a mixture of two CRA pictures and their non-AU aligned associated RADL and RASL pictures.

[0320] 6) In both single-layer and multi-layer contexts, there is a lack of constraints on the sub-picture type of the preceding sub-picture (ie, the NAL unit type of the VCL NAL units in the sub-picture).

[0321] 7) In both single-layer and multi-layer contexts, there is a lack of constraints on whether RASL sub-pictures can exist and be associated with IDR sub-pictures.

[0322] 8) In both single-layer and multi-layer contexts, there is a lack of constraints on whether a RADL sub-picture can exist and be associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP.

[0323] 9) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between sub-pictures that precede an IRAP sub-picture in decoding order and RADL sub-pictures associated with an IRAP sub-picture.

[0324] 10) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between sub-pictures that precede a GDR sub-picture in decoding order and sub-pictures that are associated with a GDR sub-picture.

[0325] 11) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between RASL sub-pictures associated with CRA sub-pictures and RADL sub-pictures associated with CRA sub-pictures.

[0326] 12) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between RASL sub-pictures associated with a CRA sub-picture and IRAP sub-pictures that precede the CRA sub-picture in decoding order.

[0327] 13) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative decoding order between the associated non-previous pictures and the preceding pictures of an IRAP picture.

[0328] 14) In both single-layer and multi-layer contexts, there is a lack of constraints on the RPL activity entries for sub-pictures that follow the STSA sub-picture in decoding order.

[0329] 15) In both single-layer and multi-layer contexts, there are no constraints on the RPL entries of CRA sub-pictures.

[0330] 16) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL activity entries that reference sub-pictures of pictures generated by the decoding process used to generate unavailable reference pictures.

[0331] 17) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL entries that reference sub-pictures of pictures generated by the decoding process used to generate unavailable reference pictures.

[0332] 18) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL activity entries for sub-pictures that are associated with an IRAP picture and follow the IRAP picture in output order.

[0333] 19) In both single-layer and multi-layer contexts, there are no constraints on RPL entries for sub-pictures that are associated with an IRAP picture and follow the IRAP picture in output order.

[0334] 20) In single-layer and multi-layer contexts, there is a lack of constraints on the RPL active entries of RADL sub-pictures.

[0335] 5. Examples of technical solutions and embodiments

[0336] To solve the above problems and other problems, the following summarized methods are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow manner. In addition, these items can be used alone or in combination in any way.

[0337] 1) To solve problem 1, instead of specifying the content of a NAL unit with a VCL NAL unit type as a "codec slice of a specific type of picture", it is specified as a "codec slice of a specific type of picture or sub-picture". For example, the content of a NAL unit with a nal_unit_type equal to CRA_NUT is specified as a "codec slice of a CRA picture or sub-picture".

[0338] a. In addition, one or more of the following terms are defined: associated GDR sub-picture, associated IRAP sub-picture, CRA sub-picture, GDR sub-picture, IDR sub-picture, IRAP sub-picture, leading sub-picture, RADL sub-picture, RASL sub-picture, STSA sub-picture, trailing sub-picture.

[0339] 2) To solve problem 2, add a constraint to require that the subpic_treated_as_pic_flag[] of any two adjacent sub-pictures of different NAL unit types should be equal to 1.

[0340] a. In one example, the constraint states that for any two adjacent sub-pictures with sub-picture indices i and j in a picture, when subpic_treated_as_pic_flag[i] or subpic_treated_as_pic_flag[j] is equal to 0, the two sub-pictures shall have the same NAL unit type.

[0341] a. Optionally, it is required that when subpic_treated_as_pic_flag[i] of any subpicture with subpicture index i is equal to 0, all subpictures in the picture shall have the same NAL unit type (i.e., all VCL NAL units in the picture shall have the same NAL unit type, i.e., the value of mixed_nalu_types_in_pic_flag shall be equal to 0). And this means that mixed_nalu_types_in_pic_flag can only be equal to 1 when all subpictures have their corresponding subpic_treated_as_pic_flag[] equal to 1.

[0342] 3) To solve issue 3, when mixed_nalu_types_in_pic_flag is equal to 1, a picture may be allowed to contain more than two different types of VCL NAL units.

[0343] 4) In order to solve problem 4, it is stipulated that the post-sub-picture should be after the associated IRAP or GDR sub-picture in the output order.

[0344] 5) To address issue 5, in order to allow mixing of RADL and RASL NAL unit types in a picture resulting from a mixture of two CRA pictures and their non-AU aligned associated RADL and RASL pictures, the existing constraint specifying that the leading picture of an IRAP picture shall be a RADL or RASL picture is changed as follows: When the picture is the leading picture of an IRAP picture, the nal_unit_type values ​​of all VCL NAL units in the picture shall be equal to RADL_NUT or RASL_NUT. Furthermore, in the decoding process of a picture using mixed nal_unit_type values ​​of RADL_NUT and RASL_NUT, when the layer containing the picture is an output layer, the PictureOutputFlag of the picture is set equal to pic_output_flag.

[0345] Thus, although guarantees of "correctness" for "intermediate" RASL sub-pictures in an associated CRA picture are appropriate when the NoOutputBeforeRecoveryFlag of such a picture is equal to 1, they are not actually required, but can be guaranteed by the constraint that all output pictures need to be correct for a conforming decoder. The unnecessary parts of the guarantee are insignificant and do not increase the complexity of implementing a conforming encoder or decoder. In this case, it would be useful to add a note (NOTE) to clarify that although such RASL sub-pictures associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 can be output by the decoding process, they are not intended to be used for display and therefore are not applied for display.

[0346] 6) In order to solve problem 6, it is stipulated that when a sub-picture is a preceding sub-picture of an IRAP sub-picture, it should be a RADL or RASL sub-picture.

[0347] 7) In order to solve problem 7, it is specified that there should be no RASL sub-picture associated with the IDR sub-picture in the bitstream.

[0348] 8) To solve problem 8, it is specified that there should be no RADL sub-picture associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP in the bitstream.

[0349] 9) In order to solve problem 9, it is stipulated that any sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx should precede the IRAP sub-picture and all its associated RADL sub-pictures in the output order, and any sub-picture should precede the IRAP sub-picture with nuh_layer_id equal to layerId and a sub-picture index equal to subpicIdx in the decoding order.

[0350] 10) In order to solve problem 10, it is stipulated that any sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx should be before the GDR sub-picture and all its associated sub-pictures in the output order, and any sub-picture should be before the GDR sub-picture with nuh_layer_id equal to layerId and a sub-picture index equal to subpicIdx in the decoding order.

[0351] 11) To address issue 11, it is specified that any RASL sub-picture associated with a CRA sub-picture should precede any RADL sub-picture associated with the CRA sub-picture in output order.

[0352] 12) To address issue 12, it is specified that any RASL sub-picture associated with a CRA sub-picture should follow, in output order, any IRAP sub-picture that precedes the CRA sub-picture in decoding order.

[0353] 13) In order to solve problem 13, it is stipulated that if field_seq_flag is equal to 0, and the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a leading sub-picture associated with an IRAP sub-picture, then it should precede all non-leading sub-pictures associated with the same IRAP sub-picture in the decoding order; otherwise, let subpicA and subpicB be the first and last leading sub-pictures associated with the IRAP sub-picture in the decoding order, respectively, and there should be at most one non-leading sub-picture with nuh_layer_id equal to layerId and subpicIdx sub-picture index before subpicA in the decoding order, and there should be no non-leading picture with nuh_layer_id equal to layerId and subpicIdx sub-picture index between picA and picB in the decoding order.

[0354] 14) In order to solve problem 14, it is stipulated that when the current sub-picture having a TemporalId equal to a specific value tId, a nuh_layer_id equal to a specific value layerId, and a sub-picture index equal to a specific value subpicIdx is a sub-picture after a STSA sub-picture having a TemporalId equal to tId, a nuh_layer_id equal to layerId, and a sub-picture index equal to subpicIdx in decoding order, then there should not be a picture having a TemporalId equal to tId and a nuh_layer_id equal to layerId before the picture containing the STSA sub-picture in decoding order, which is included as an active entry in RefPicList[0] or RefPicList[1].

[0355] 15) To address issue 15, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a CRA sub-picture, there should be no picture referenced by an entry in RefPicList[0] or RefPicList[1] that precedes any picture in output order or in decoding order that contains a previous IRAP sub-picture (when present) with nuh_layer_id equal to layerId and a sub-picture index equal to subpicIdx.

[0356] 16) In order to solve problem 16, it is stipulated that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is not a RASL sub-picture associated with a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR sub-picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a sub-picture of a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, then there should be no picture involved in the active entry in RefPicList[0] or RefPicList[1] generated by the decoding process of generating an unavailable reference picture.

[0357] 17) In order to solve problem 17, it is stipulated that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is not a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a sub-picture that precedes the preceding sub-picture associated with the same CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 in decoding order, a preceding sub-picture associated with a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR sub-picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a sub-picture of a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, then there should be no picture referred to by an entry in RefPicList[0] or RefPicList[1] generated by a decoding process that generates an unavailable reference picture.

[0358] 18) In order to solve problem 18, it is stipulated that when the current sub-picture is associated with an IRAP sub-picture and follows the IRAP sub-picture in the output order, there should be no pictures referred to by active entries in RefPicList[0] or RefPicList[1] that precede the picture containing the associated IRAP sub-picture in the output order or in the decoding order.

[0359] 19) To address issue 19, it is provided that when the current sub-picture is associated with an IRAP sub-picture, after the IRAP sub-picture in output order and after the preceding sub-picture associated with the same IRAP sub-picture (if any) in both decoding order and in output order, there should not be any picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the associated IRAP sub-picture in output order or in decoding order.

[0360] 20) To address issue 20, it is specified that when the current sub-picture is a RADL sub-picture, there should not be any of the following active entries in RefPicList[0] or RefPicList[1]:

[0361] a. Pictures containing RASL sub-pictures

[0362] b. The picture that precedes the picture containing the associated IRAP sub-picture in decoding order

[0363] 6. Examples

[0364] The following are some example embodiments of some of the technical solutions summarized in Section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-Q2001-vE / v15. Most of the relevant parts that have been added or modified are highlighted in bold italics, and some deleted parts are highlighted in a highlighted manner with left and right double brackets (e.g., [[ ]]), with the deleted text between the double brackets. There are also some other changes that are editorial or are not part of the present technical solution and are therefore not highlighted.

[0365] 6.1. First embodiment

[0366] This embodiment applies to items 1, 1a, 2, 2a, 4 and 6 to 20.

[0367] 3 Definition ...

[0369] Associated GDR pictures (of a specific picture with nuh_layer_id with a specific value of layerId): the previous GDR picture with nuh_layer_id equal to layerId in decoding order (when present), with no IRAP picture with nuh_layer_id equal to layerId between this picture and the specific picture in decoding order.

[0370]

[0371] Associated IRAP picture (of a specific picture with nuh_layer_id with a specific value of layerId): the previous IRAP picture in decoding order (when present) with nuh_layer_id equal to layerId, with no GDR picture with nuh_layer_id equal to layerId between this picture and the specific picture in decoding order.

[0372]

[0373] Pure Random Access (CRA) picture: An IRAP picture with nal_unit_type equal to CRA_NUT for each VCL NAL unit.

[0374]

[0375] Gradual Decoding Refresh (GDR) AU: Each layer in CVS has a PU and each coded picture in the current PU is an AU of a GDR picture.

[0376] Gradual Decoding Refresh (GDR) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to GDR_NUT.

[0377]

[0378] Instantaneous Decoding Refresh (IDR) picture: Each VCL NAL unit includes an IRAP picture with nal_unit_type equal to IDR_W_RADL or IDR_N_LP.

[0379]

[0380] Intra Random Access Point (IRAP) picture: A picture whose all VCL NAL units include the same value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive.

[0381]

[0382] Preceding picture: The picture that comes before the associated IRAP picture in the output order.

[0383] order, The decoded picture output from DPB

[0384] Random Access Decodable Lead (RADL) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to RADL_NUT.

[0385]

[0386] Random Access Skip Leading (RASL) picture: Each VCL NAL unit includes a picture with nal_unit_type equal to RASL_NUT.

[0387]

[0388] Step-by-Step Temporal Sublayer Access (STSA) pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to STSA_NUT.

[0389]

[0390] Trailing pictures: Each VCL NAL unit includes a picture with nal_unit_type equal to TRAIL_NUT.

[0391] Note – A post-picture associated with an IRAP or GDR picture also follows the IRAP or GDR picture in decoding order. Pictures that follow the associated IRAP or GDR picture in output order and precede the associated IRAP or GDR picture in decoding order are not allowed.

[0392]

[0393] ...

[0395] 7.4.2.2 NAL unit header semantics ...

[0397] nal_unit_type specifies the NAL unit type, ie the type of the RBSP data structure contained in the NAL unit as specified in Table 5.

[0398] NAL units with unspecified semantics whose nal_unit_type is in the range UNSPEC_28..UNSPEC_31 (inclusive) shall not affect the decoding process specified in this specification.

[0399] NOTE 2 – NAL unit types in the range UNSPEC_28..UNSPEC_31 MAY be used as determined by the application. The decoding treatment of these values ​​of nal_unit_type is not specified in this specification. Because different applications may use these NAL unit types for different purposes, special care must be exercised when designing encoders that produce NAL units with these nal_unit_type values, and when designing decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define any management of these values.

[0400] These nal_unit_type values ​​may only be used in contexts where usage "conflicts" (i.e., different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not important, or are not possible, or are managed (e.g., as defined or managed in a controlling application or transport specification, or by controlling the distribution of the bitstream).

[0401] For purposes other than determining the amount of data in a DU of a bitstream (as specified in Annex C), a decoder shall ignore (remove from the bitstream and discard) the contents of all NAL units that use the reserved value of nal_unit_type.

[0402] NOTE 3 – This requirement allows for the future definition of compatible extensions to this specification.

[0403] Table 5 – NAL unit type codes and NAL unit type classifications

[0404]

[0405]

[0406]

[0407] NOTE 4 - A pure random access (CRA) picture may have an associated RASL or RADL picture present in the bitstream.

[0408] NOTE 5 – An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP has no associated preceding picture in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL has no associated RASL picture present in the bitstream.

[0409] But there may be associated RADL pictures in the bitstream.

[0410] The nal_unit_type value of all VCL NAL units in a sub-picture shall be the same. A sub-picture is said to have the same NAL unit type as the VCL NAL units of the sub-picture.

[0411]

[0412] For any particular picture's VCL NAL unit, the following applies:

[0413] – If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and a picture or PU is considered to have the same NAL unit type as the VCL NAL units of the picture or PU.

[0414] – Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture shall have at least two sub-pictures, and the VCL NAL units of the picture shall have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one sub-picture of the picture shall all have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of the other sub-pictures in the picture shall all have different specific values ​​of nal_unit_type equal to TRAIL_NUT, RADL_NUT, RASL_NUT.

[0415] A requirement for bitstream conformance is that the following constraints apply:

[0416] – Poster pictures should come after the associated IRAP or GDR picture in output order.

[0417]

[0418] – When a picture is the preceding picture of an IRAP picture, it shall be a RADL or RASL picture.

[0419]

[0420] – There shall be no RASL pictures associated with an IDR picture in the bitstream.

[0421]

[0422] – A RADL picture associated with an IDR picture with nal_unit_type equal to IDR_N_LP shall not be present in the bitstream.

[0423] NOTE 6 – Random access can be performed at the location of an IRAP PU (and correctly decode the IRAP picture and all subsequent non-RASL pictures in decoding order) by discarding all PUs preceding the IRAP PU, provided that each parameter set is available (either in the bitstream or by external means not specified in this specification) when referenced.

[0424]

[0425] – Any pictures (with nuh_layer_id equal to a particular value of layerId) that precede the IRAP picture with nuh_layer_id equal to layerId in decoding order shall precede the IRAP picture and all its associated RADL pictures in output order.

[0426]

[0427] – Any picture (with nuh_layer_id equal to a particular value of layerId) that precedes the GDR picture with nuh_layer_id equal to layerId in decoding order shall precede the GDR picture and all its associated pictures in output order.

[0428]

[0429] – Any RASL pictures associated with a CRA picture shall precede any RADL pictures associated with the CRA picture in output order.

[0430]

[0431] – Any RASL pictures associated with a CRA picture in output order shall follow any IRAP pictures that precede the CRA picture in decoding order.

[0432]

[0433] – If field_seq_flag is equal to 0, and the current picture with nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, then it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, let picA and picB be the first and last leading pictures associated with the IRAP picture in decoding order, respectively, there shall be at most one non-leading picture with nuh_layer_id equal to layerId before picA in decoding order, and there shall be no non-leading pictures with nuh_layer_id equal to layerId between picA and picB in decoding order.

[0434]

[0435] ...

[0437] 7.4.3.4 Picture parameter set semantics ...

[0439] mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture of the referenced PPS has more than one VCL NAL unit and the VCL NAL units do not have the same nal_unit_type value [[and the picture is not an IRAP picture]]. mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same nal_unit_type value.

[0440] When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_types_in_pic_flag shall be equal to 0.

[0441] [[For each slice with nal_unit_type value nalUnitTypeA in the range IDR_W_RADL to CRA_NUT (inclusive), in picture picA that also contains one or more slices with another value of nal_unit_type (i.e., the value of mixed_nalu_types_in_pic_flag of picture picA is equal to 1), the following applies:

[0442] – The slice shall belong to the sub-picture subpicA for which the value of subpic_treated_as_pic_flag[i] is equal to 1.

[0443] – A slice shall not belong to a picA sub-picture containing VCL NAL units with nal_unit_type not equal to nalUnitTypeA.

[0444] – If nalUnitTypeA is equal to CRA, for all subsequent PUs following the current picture in CLVS in decoding order and in output order, neither RefPicList[0] nor RefPicList[1] of slices in subpicA in these PUs shall include any pictures in the active entries that precede picA in decoding order.

[0445] – Otherwise (i.e., nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS that follow the current picture in decoding order, neither RefPicList[0] nor RefPicList[1] for slices in subpicA in those PUs shall include any pictures in the active entry that precede picA in decoding order. ]]

[0446] NOTE 1 – mixed_nalu_types_in_pic_flag equal to 1 indicates that the picture referencing the PPS contains slices with different NAL unit types, for example, a codec picture originating from a sub-picture bitstream merge operation, where the encoder must ensure bitstream structure matching and further alignment of the original bitstream parameters. An example of such alignment is as follows: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_in_pic_flag is equal to 1, the picture referencing the PPS shall not have slices with nal_unit_type equal to IDR_W_RADL or IDR_N_LP. ...

[0448] 7.4.3.7 Image header structure semantics ...

[0450] recovery_poc_cnt specifies the recovery point of the decoded picture in output order.

[0451]

[0452] If the current picture is a GDR picture [[associated with PH]], and there is a picture picA in CLVS that follows the current GDR picture in decoding order, which has a value equal to If the value of [[the current GDR picture's PicOrderCntVal plus the value of recovery_poc_cnt]] is equal to PicOrderCntVal, the picture picA is called the recovery point picture. Otherwise, In output order greater than The first picture whose PicOrderCntVal is [[the current picture's PicOrderCntVal plus the value of recovery_poc_cnt]] is called the recovery point picture. The recovery point picture should not precede the current GDR picture in decoding order. The value of recovery_poc_cnt should be in the range of 0 to MaxPicOrderCntLsb-1 (inclusive).

[0453] [[When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:

[0454] RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81)]]

[0455] Note 2 – When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to that of the associated GDR picture [[RpPicOrderCntVal]], the current and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (when present) that precedes the associated GDR picture in decoding order. ...

[0457] 8.3.2 Decoding Processing of Reference Picture List Construction ...

[0459] A requirement for bitstream conformance is that the following constraints apply:

[0460] – For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i].

[0461] – The picture referred to by each active entry in RefPicList[0] or RefPicList[1] shall exist in the DPB and the TemporalId shall be less than or equal to the TemporalId of the current picture.

[0462] – Each entry in RefPicList[0] or RefPicList[1] refers to a picture that shall not be the current picture and non_reference_picture_flag shall be equal to 0.

[0463] – A STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and a LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture.

[0464] – There should not be a difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture involved in this entry in RefPicList[0] or RefPicList[1] that is greater than or equal to 2 24 LTRP entry.

[0465] – Let setOfRefPics be the unique set of pictures involving all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize-1 (inclusive), where MaxDpbSize is specified in clause A.4.2 and setOfRefPics shall be the same for all slices of a picture.

[0466] – When the current slice has nal_unit_type equal to STSA_NUT, there shall be no active entry in RefPicList[0] or RefPicList[1] with TemporalId equal to the TemporalId of the current picture and nuh_layer_id equal to the nuh_layer_id of the current picture.

[0467] – When the current picture is a picture that follows a STSA picture in decoding order, whose TemporalId is equal to the TemporalId of the current picture and whose nuh_layer_id is equal to the nuh_layer_id of the current picture, then there shall be no picture that precedes the STSA picture in decoding order, whose TemporalId is equal to the TemporalId of the current picture and whose nuh_layer_id is equal to the nuh_layer_id of the current picture, included as an active entry in RefPicList[0] or RefPicList[1].

[0468]

[0469]

[0470] – When the current picture with nuh_layer_id equal to a particular value layerId is a CRA picture, then there shall be no picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes any previous IRAP picture (when present) in decoding order with nuh_layer_id equal to layerId in output order or in decoding order.

[0471]

[0472] – When the current picture with nuh_layer_id equal to a specific value layerId is not a RASL picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, then there shall be no pictures referred to by the active entry in RefPicList[0] or RefPicList[1] generated by the decoding process that generates unusable reference pictures.

[0473]

[0474] – When the current picture with nuh_layer_id equal to a particular value layerId is not a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a picture preceding, in decoding order, the preceding picture associated with the same CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a preceding picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, then there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] generated by a decoding process that generates unusable reference pictures.

[0475]

[0476] – When the current picture is associated with an IRAP picture and follows that IRAP picture in output order, there shall be no pictures referred to by active entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or in decoding order.

[0477]

[0478] – When the current picture is associated with an IRAP picture, after the IRAP picture in output order and after the preceding picture associated with the same IRAP picture (if any) in both decoding order and in output order, there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or in decoding order.

[0479]

[0480]

[0481] – When the current picture is a RADL picture, there shall not be any of the following active entries in RefPicList[0] or RefPicList[1]:

[0482] οRASL pictures

[0483] o The picture that precedes the associated IRAP picture in decoding order

[0484]

[0485] – The pictures referred to by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be in the same AU as the current picture.

[0486] – The picture referred to by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB and the nuh_layer_id shall be less than the nuh_layer_id of the current picture.

[0487] – Each ILRP entry in either RefPicList[0] or RefPicList[1] of a stripe shall be an active entry.

[0488] Figure 5 1 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values), or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0489] System 1900 may include a codec component 1904, which may implement various codecs or encoding methods described in this document. Codec component 1904 may reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 may be stored or transmitted through the communication of the connection represented by component 1906. Component 1908 may use the bitstream (or codec) representation of the storage or communication of the video received at input 1902 to generate pixel values ​​or displayable video sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations of the inverse encoding results will be performed by the decoder.

[0490] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0491] Figure 6 3600 is a block diagram of a video processing device 3600. Device 3600 may be used to implement one or more methods described herein. Device 3600 may be implemented in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor(s) 3602 may be configured to implement one or more methods described in this document. Memory(s) 3604 may be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuits.

[0492] Figure 8 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0493] like Figure 8As shown, the video codec system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.

[0494] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0495] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bit stream. The bit stream may include a bit sequence that forms a codec representation of the video data. The bit stream may include a codec picture and related data. The codec picture is a codec representation of the picture. Related data may include a sequence parameter set, a picture parameter set, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.

[0496] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0497] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the target device 120, or may be outside the target device 120, which is configured to be connected to an external display device.

[0498] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.

[0499] Fig. 9 is a block diagram showing an example of a video encoder 200, which may be Figure 8 The video encoder 114 in the system 100 is shown in FIG.

[0500] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Fig. 9 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared between various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0501] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0502] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, wherein at least one reference picture is a picture in which the current video block is located.

[0503] Furthermore, some components (eg, motion estimation unit 204 and motion compensation unit 205) may be highly integrated but are not shown for purposes of explanation. Fig. 9 In the example, they are represented separately.

[0504] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0505] The mode selection unit 203 may select one of the coding modes (intra or inter) (e.g., based on the error result), and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).

[0506] In order to perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block of the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 213 and decoded samples.

[0507] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0508] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block of the current video block in a reference picture in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1, the reference picture including the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information of the current video block. Motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0509] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 that include the reference video block, and a motion vector indicating a spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0510] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder.

[0511] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0512] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0513] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0514] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0515] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0516] The residual generation unit 207 may generate residual data for the current video block by subtracting (eg, indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0517] In other examples, the current video block may not have residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform a subtraction operation.

[0518] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0519] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0520] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0521] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0522] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0523] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of a video block, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from a video block to a bitstream (or bitstream representation) of a video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from a bitstream of a video to a block of a video will be performed using a video processing tool or mode enabled based on the decision or determination.

[0524] Fig.10 is a block diagram showing an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown.

[0525] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Fig.10 In the example of , video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared between various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0526] exist Fig.10In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform operations generally similar to those for the video encoder 200 ( Fig. 9 ) is the decoding channel that is the opposite of the encoding process described.

[0527] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 may decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 302 may determine motion information, including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and merge modes.

[0528] Motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0529] Motion compensation unit 302 may calculate interpolation of sub-integer pixels of a reference block using interpolation filters as used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 according to received syntax information and use the interpolation filters to generate a prediction block.

[0530] The motion compensation unit 302 may use some syntax information to determine the sizes of blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0531] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0532] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates decoded video for presentation on a display device.

[0533] Next, a list of technical solutions preferred by some embodiments is provided.

[0534] The following technical scheme illustrates an example implementation of the techniques discussed in the previous sections (eg, Item 1).

[0535] 1. A video processing method (eg, Figure 7 ), comprising: performing (702) conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation format rule specifies that the one or more pictures including the one or more sub-pictures are included in a codec representation according to a network abstraction layer (NAL) unit, wherein a type of NAL unit indicated in the codec representation includes a codec slice of a specific type of picture or a codec slice of a specific type of sub-picture.

[0536] The following technical scheme illustrates an example implementation of the technology discussed in the previous section (eg, Item 2).

[0537] 2. A video processing method, comprising: performing conversion between a video including one or more pictures containing one or more sub-pictures and a codec representation of the video, wherein the codec representation complies with a format rule that specifies that two adjacent sub-pictures with different network abstraction layer unit types will have the same indication of a sub-picture that is considered a picture marker.

[0538] The following technical solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., items 4, 5, 6, 7, 9, 1, 11, 12).

[0539] 3. A video processing method, comprising: performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation complies with a format rule, the format rule defining an order of a first type of sub-picture and a second type of sub-picture, wherein the first sub-picture is a post-sub-picture or a pre-sub-picture or a random access skip pre-sub (RASL) sub-picture type, and the second sub-picture is a RASL type or a random access decodable pre-sub (RADL) type or an instantaneous decoding refresh (IDR) type or a gradual decoding refresh (GDR) type sub-picture.

[0540] 4. The method according to technical solution 3, wherein the rule stipulates that the post-sub-picture is after the random access point or GDR sub-picture in the associated frame in the output order.

[0541] 5. The method according to technical solution 3, wherein the rule stipulates that when a picture is a preceding picture of an intra-frame random access point picture, the nal_unit_type value of all network abstraction layer units in the picture is equal to RADL_NUT or RASL_NUT.

[0542] 6. The method according to technical solution 3, wherein the rule stipulates that a given sub-picture that is a preceding sub-picture of an IRAP sub-picture must also be a RADL or RASL sub-picture.

[0543] 7. The method according to Technical Solution 3, wherein the rule stipulates that a given sub-picture as a RASL sub-picture is not allowed to be associated with an IDR sub-picture.

[0544] 8. A method according to technical solution 3, wherein the rule stipulates that a given sub-picture with the same layer id and sub-picture index as the IRAP sub-picture must precede the IRAP sub-picture and all its associated RADL sub-pictures in output order.

[0545] 9. A method according to technical solution 3, wherein the rule stipulates that a given sub-picture with the same layer id and sub-picture index as the GDR sub-picture must be before the GDR sub-picture and all its associated RADL sub-pictures in output order.

[0546] 10. The method according to technical solution 3, wherein the rule stipulates that a given sub-picture that is a RASL sub-picture associated with a CRA sub-picture precedes all RADL sub-pictures associated with the CRA sub-picture in output order.

[0547] 11. A method according to technical solution 3, wherein the rule stipulates that a given sub-picture that is a RASL sub-picture associated with a CRA sub-picture precedes all IRAP sub-pictures associated with the CRA sub-picture in output order.

[0548] 12. A method according to technical solution 3, wherein the rule stipulates that if a given sub-picture is a preceding sub-picture with an IRAP sub-picture, then the given sub-picture is before all non-preceding sub-pictures related to the IRAP sub-picture in decoding order.

[0549] The following technical scheme illustrates an example embodiment of the techniques discussed in previous sections (eg, items 8, 14, 15).

[0550] 13. A video processing method, comprising: performing conversion between a video including one or more pictures containing one or more sub-pictures and a codec representation of the video, wherein the codec representation complies with a format rule, which defines conditions for allowing or not allowing a first type of sub-picture to exist together with a second type of sub-picture.

[0551] 14. The method according to technical solution 13, wherein the rule stipulates that if there is an IDR sub-picture of the network abstraction layer type IDR_N_LP, the codec representation with RADP sub-picture is not allowed.

[0552] 15. A method according to technical solution 13, wherein the rule does not allow a picture to be included in a reference list of a picture containing a step-by-step temporal sublayer access (STSA) sub-picture, so that the picture is before the picture containing the STSA sub-picture.

[0553] 16. A method according to technical solution 13, wherein the rule does not allow a picture to be included in a reference list of a picture of a picture containing an intra random access point (IRAP) sub-picture, so that the picture is before the picture containing the IRAP sub-picture.

[0554] 17. A method according to any one of technical solutions 1 to 16, wherein the conversion includes encoding the video into a codec representation.

[0555] 18. A method according to any one of technical solutions 1 to 16, wherein the conversion includes decoding a codec representation to generate pixel values ​​of a video.

[0556] 19. A video decoding device, comprising a processor configured to implement the method described in one or more of Technical Solutions 1 to 18.

[0557] 20. A video encoding device, comprising a processor configured to implement the method described in one or more of Technical Solutions 1 to 18.

[0558] 21. A computer program product having computer code stored thereon, which, when executed by a processor, enables the processor to implement the method described in any one of Technical Solutions 1 to 18.

[0559] 22. The methods, devices, or systems described in this document.

[0560] In the technical scheme described herein, an encoder may comply with the format rules by generating a codec representation according to the format rules. In the technical scheme described herein, a decoder may parse syntax elements in the codec representation using the format rules, understand the presence and absence of syntax elements according to the format rules, and generate decoded video.

[0561] Fig.11 1 is a flow diagram of an example method 1100 of video processing. Operation 1102 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies a syntax of a network abstraction layer (NAL) unit in the bitstream, and wherein the format rule specifies that a NAL unit of a video codec layer (VCL) NAL unit type includes content associated with a particular type of picture or a particular type of sub-picture.

[0562] In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with a pure random access picture or a pure random access sub-picture. In some embodiments of the method 1100, the pure random access sub-picture is an intra random access point sub-picture with a pure random access type per VCL NAL unit. In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with an associated gradual decoding refresh picture or an associated gradual decoding refresh sub-picture. In some embodiments of the method 1100, the associated gradual decoding refresh sub-picture is, in decoding order, a previous gradual decoding refresh sub-picture having an identifier of a layer to which the VCL NAL unit belongs or an identifier of a layer to which the non-VCL NAL unit applies equal to a first specific value and having a second specific value of a sub-picture index, and wherein, in decoding order, between the previous gradual decoding refresh sub-picture and the specific sub-picture having the first specific value of the identifier and the second specific value of the sub-picture index, there is no intra random access point sub-picture having the first specific value of the identifier identifier and the second specific value of the sub-picture index.

[0563] In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with an associated intra random access point picture or an associated intra random access point sub-picture. In some embodiments of the method 1100, the associated intra random access point sub-picture is, in decoding order, a previous intra random access point sub-picture having an identifier of a layer to which the VCL NAL unit belongs or an identifier of a layer to which the non-VCL NAL unit applies equal to a first specific value and having a second specific value of a sub-picture index, and wherein, in decoding order, between the previous intra random access point sub-picture and the specific sub-picture having the first specific value of the identifier and the second specific value of the sub-picture index, there is no progressive decoding refresh sub-picture having the first specific value of the identifier and the second specific value of the sub-picture index. In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with an immediate decoding refresh picture or an immediate decoding refresh sub-picture. In some embodiments of the method 1100, an immediate decoding refresh sub-picture is an intra random access point sub-picture in which each VCL NAL unit has an immediate decoding refresh type.

[0564] In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with a leading picture or a leading sub-picture. In some embodiments of the method 1100, the leading sub-picture is a sub-picture that precedes the random access point sub-picture in the associated frame in output order. In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with a random access decodable leading picture or a random access decodable leading sub-picture. In some embodiments of the method 1100, the random access decodable leading sub-picture is a sub-picture with a random access decodable leading type per VCL NAL unit. In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with a random access skip leading picture or a random access skip leading sub-picture. In some embodiments of the method 1100, the random access skip leading sub-picture is a sub-picture with a random access skip leading type per VCL NAL unit. In some embodiments of the method 1100 , the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with a step-by-step temporal sub-layer access picture or a step-by-step temporal sub-layer access sub-picture.

[0565] In some embodiments of the method 1100, a step-by-step temporal sub-layer access sub-picture is a sub-picture with a step-by-step temporal sub-layer access type per VCL NAL unit. In some embodiments of the method 1100, the content of the NAL unit of the VCL NAL unit type indicates that the codec slice is associated with a post-picture or a post-sub-picture. In some embodiments of the method 1100, a post-sub-picture is a sub-picture with a post-type per VCL NAL unit.

[0566] Fig.12 12 is a flow chart of an example method 1200 for video processing. Operation 1202 includes performing conversion between a video including a picture including a sub-picture and a bitstream of the video, wherein the bitstream complies with a format rule, and wherein the format rule specifies that in response to the sub-picture being a preceding sub-picture of an intra random access point sub-picture, the sub-picture is a random access type sub-picture.

[0567] In some embodiments of the method 1200, the random access type of the sub-picture is a random access decodable leading sub-picture. In some embodiments of the method 1200, the random access type of the sub-picture is a random access skip leading sub-picture.

[0568] Fig.13 13 is a flow chart of an example method 1300 for video processing. Operation 1302 includes performing conversion between a video including a picture including a sub-picture and a bitstream of the video, wherein the bitstream complies with a format rule, and wherein the format rule specifies that in response to one or more random access skip leading sub-pictures being associated with an instantaneous decoding refresh sub-picture, one or more random access skip leading sub-pictures are not present in the bitstream.

[0569] Fig.14 14 is a flow diagram of an example method 1400 for video processing. Operation 1402 includes performing conversion between a video including a picture including a sub-picture and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that, in response to one or more random access decodable leading sub-pictures being associated with an immediate decoding refresh sub-picture having a type of network abstraction layer (NAL) unit indicating that the immediate decoding refresh sub-picture is not associated with a preceding picture, the one or more random access decodable leading sub-pictures are not present in the bitstream.

[0570] Fig.1515 is a flow chart of an example method 1500 for video processing. Operation 1502 includes performing conversion between a video including a picture including two adjacent sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule that specifies that two adjacent sub-pictures having different types of network abstraction layer (NAL) units have a syntax element having a same first value, the syntax element indicating whether each of the two adjacent sub-pictures in a codec layer video sequence is considered a picture.

[0571] In some embodiments of the method 1500 , the format rule specifies that the syntax element of two adjacent sub-pictures indicates that each of the two adjacent sub-pictures in the codec layer video sequence is considered as a picture.

[0572] Fig.16 is a flow chart of an example method 1600 for video processing. Operation 1602 includes performing conversion between a video including a picture including two adjacent sub-pictures and a bitstream of the video, wherein the format rule specifies that the two adjacent sub-pictures include a first adjacent sub-picture having a first sub-picture index and a second adjacent sub-picture having a second sub-picture index, and wherein the format rule specifies that in response to a first syntax element associated with the first sub-picture index indicating that the first adjacent sub-picture is not considered as a picture or a second syntax element associated with the second sub-picture index indicating that the second adjacent sub-picture is not considered as a picture, the two adjacent sub-pictures have the same type of network abstraction layer (NAL) unit.

[0573] In some embodiments of the method 1600, the picture includes a plurality of sub-pictures including two adjacent sub-pictures, and wherein the format rule provides that, in response to a sub-picture from the plurality of sub-pictures having a syntax element indicating that the sub-picture is not considered as a picture, the plurality of sub-pictures have NAL units of the same type. In some embodiments of the method 1600, the picture includes a plurality of sub-pictures including two adjacent sub-pictures, and wherein the format rule provides that, in response to the plurality of sub-pictures having a corresponding syntax element indicating that each of the plurality of sub-pictures in a CLVS is considered as a picture, the syntax element indicates that each picture of a video that references a picture parameter set (PPS) has a plurality of VCL NAL units that do not have video codec layer (VCL) NAL units of the same type.

[0574] Fig.17is a flow chart of an example method 1700 for video processing. Operation 1702 includes performing conversion between a video including a picture including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that, in response to a syntax element indicating a reference picture parameter set (PPS), each picture of the video has a plurality of video codec layer (VCL) network abstraction layer (NAL) units without the same type of VCL NAL units, allowing the picture to include more than two different types of VCL NAL units.

[0575] Fig.18 18 is a flow chart of an example method 1800 for video processing. Operation 1802 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule that specifies that a post-subpicture associated with an intra random access point subpicture or a gradual decoding refresh subpicture in sequence follows the intra random access point subpicture or the gradual decoding refresh subpicture.

[0576] In some embodiments of method 1800, the order is an output order.

[0577] Fig.19 is a flow chart of an example method 1900 for video processing. Operation 1902 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule, the format rule providing that in response to: (1) the sub-picture precedes the intra random access point sub-picture in the second order, (2) the sub-picture and the intra random access point sub-picture have the same first value for a layer to which network abstraction layer (NAL) units of the sub-picture and the intra random access point sub-picture belong, and (3) the sub-picture and the intra random access point sub-picture have the same second value of a sub-picture index, the sub-picture precedes the intra random access point sub-picture and one or more random access decodable pre-sub-pictures associated with the intra random access point sub-picture in the first order.

[0578] In some embodiments of the method 1800, the first order is an output order. In some embodiments of the method 1800, the second order is a decoding order.

[0579] Fig. 20 2000 is a flow chart of an example method 2000 of video processing. Operation 2002 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a random access skipped leading sub-picture associated with a pure random access sub-picture precedes one or more random access decodable leading sub-pictures associated with the pure random access sub-picture in order.

[0580] In some embodiments of method 2000, the order is an output order.

[0581] Fig.21 is a flow chart of an example method 2100 for video processing. Operation 2102 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a random access skip preamble sub-picture associated with a pure random access sub-picture in a first order is after one or more intra random access point sub-pictures that precede the pure random access sub-picture in a second order.

[0582] In some embodiments of the method 2100, the first order is an output order. In some embodiments of the method 2100, the second order is a decoding order.

[0583] Fig. 22 is a flow chart of an example method 2200 for video processing. Operation 2202 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, the format rule specifying that in response to: (1) a syntax element indicating that a codec layer video sequence conveys pictures representing a frame, and (2) a current sub-picture is a leading sub-picture associated with an intra random access point sub-picture, and the current sub-picture precedes one or more non-leading sub-pictures associated with the intra random access point sub-picture in decoding order.

[0584] In some embodiments of method 2200, in response to (1) a syntax element indicating that a codec layer video sequence conveys a picture representing a field, and (2) a current sub-picture is not a leading sub-picture, a format rule provides that: there is at most one non-leading sub-picture before a first leading sub-picture associated with an intra-frame random access point sub-picture in decoding order, and there is no non-leading sub-picture between the first leading sub-picture associated with the intra-frame random access point sub-picture and the last leading sub-picture in decoding order, wherein the current sub-picture, the at most one non-leading sub-picture, and the non-leading picture have the same first value for the layer to which the network abstraction layer (NAL) units of the current sub-picture, the at most one non-leading sub-picture, and the non-leading picture belong, and wherein the current sub-picture, the at most one non-leading sub-picture, and the non-leading picture have the same second value of the sub-picture index.

[0585] Fig.2323 is a flow chart of an example method 2300 for video processing. Operation 2302 includes performing conversion between a video including one or more pictures including one or more sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule that specifies that, in response to the picture being a preceding picture of an intra random access point picture, one or more types of NAL units of all video codec layer (VCL) network abstraction layer (NAL) units in the picture include RADL_NUT or RASL_NUT.

[0586] In some embodiments of method 2300, the format rules provide that in response to: (1) one or more types of NAL units of all VCL NAL units in a picture include RADL_NUT and RASL_NUT, and (2) the layer including the picture is an output layer, a variable of the picture is set equal to the value of a picture output flag.

[0587] Fig.24 is a flow chart of an example method 2400 for video processing. Operation 2402 includes performing conversion between a video including one or more pictures including a plurality of sub-pictures and a bitstream of the video, wherein the bitstream complies with a format rule, the format rule providing that in response to: (1) at least one sub-picture precedes a progressively decoding refresh sub-picture in a second order, (2) the at least one sub-picture and the progressively decoding refresh sub-picture have the same first value for a layer to which a network abstraction layer (NAL) unit of the at least one sub-picture and the progressively decoding refresh sub-picture belongs, and (3) the at least one sub-picture and the progressively decoding refresh picture have the same second value of a sub-picture index, the at least one sub-picture precedes the progressively decoding refresh sub-picture and one or more sub-pictures associated with the progressively decoding refresh sub-picture in the first order.

[0588] In some methods of embodiment 2400, the first order is an output order. In some methods of embodiment 2400, the second order is a decoding order.

[0589] Fig.25is a flow chart of an example method 2500 for video processing. Operation 2502 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to: (a) a first picture having a same temporal identifier and a same layer identifier of a network abstraction layer (NAL) unit as the current sub-picture, and (b) the current sub-picture is after a stepwise temporal sub-layer access sub-picture in decoding order, and (c) the current sub-picture and the stepwise temporal sub-layer access sub-picture have the same temporal identifier, the same layer identifier, and the same sub-picture index, the format rule not allowing an active entry in a reference picture list of the current slice to include a first picture that is before a second picture in decoding order that includes the stepwise temporal sub-layer access sub-picture.

[0590] In some methods of embodiment 2500, the reference picture list comprises a list 0 reference picture list. In some methods of embodiment 2500, the reference picture list comprises a list 1 reference picture list.

[0591] Fig.26 is a flow chart of an example method 2600 of video processing. Operation 2602 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, and in response to the current sub-picture not being a sub-picture of a particular type, the format rule does not allow an active entry in a reference picture list for the current slice to include a first picture generated by a decoding process that generates an unavailable reference picture.

[0592] In some embodiments of the method 2600, the current sub-picture is not a random access skip preamble sub-picture associated with a pure random access sub-picture of a pure random access picture, the pure random access sub-picture having a flag value equal to 1 indicating no output before recovery. In some embodiments of the method 2600, the current sub-picture is not a gradually decoded refresh sub-picture of a gradually decoded refresh picture, the gradually decoded refresh sub-picture having a flag value equal to 1 indicating no output before recovery. In some embodiments of the method 2600, the current sub-picture is not a sub-picture of a recovery picture of a gradually decoded refresh picture, the sub-picture of the recovery picture of the gradually decoded refresh picture having a flag value equal to 1 indicating no output before recovery and having a layer identifier of a network abstraction layer (NAL) unit that is the same as the current sub-picture. In some embodiments of the method 2600, the reference picture list includes a list 0 reference picture list. In some embodiments of the method 2600, the reference picture list includes a list 1 reference picture list.

[0593] Fig. 27is a flow chart of an example method 2700 of video processing. Operation 2702 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, and in response to the current sub-picture not being a sub-picture of a particular type, the format rule does not allow an entry in a reference picture list for the current slice to include a first picture generated by a decoding process that generates an unavailable reference picture.

[0594] In some embodiments of the method 2700, the current picture is not a pure random access sub-picture of a pure random access picture, the pure random access sub-picture having a flag value equal to 1 indicating no output before recovery. In some embodiments of the method 2700, the current sub-picture is not a sub-picture that precedes one or more preceding sub-pictures associated with a pure random access sub-picture of a pure random access picture in decoding order, the sub-pictures preceding the one or more preceding sub-pictures having a flag value equal to 1 indicating no output before recovery. In some embodiments of the method 2700, the current sub-picture is not a preceding sub-picture associated with a pure random access sub-picture of a pure random access picture, the preceding sub-picture having a flag value equal to 1 indicating no output before recovery. In some embodiments of the method 2700, the current sub-picture is not a gradually decoded refresh sub-picture of a gradually decoded refresh picture, the gradually decoded refresh sub-picture having a flag value equal to 1 indicating no output before recovery.

[0595] In some embodiments of the method 2700, the current sub-picture is not a sub-picture of a recovery picture of a progressive decoding refresh picture, the sub-picture having a flag value equal to 1 indicating no output before recovery and having a layer identifier of a network abstraction layer (NAL) unit that is the same as the current sub-picture. In some embodiments of the method 2700, the reference picture list includes a list 0 reference picture list. In some embodiments of the method 2700, the reference picture list includes a list 1 reference picture list.

[0596] Fig.28 is a flow chart of an example method 2800 for video processing. Operation 2802 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to (a) a first picture including a previous intra random access point sub-picture, the previous intra random access point sub-picture preceding the current sub-picture in the second order, (b) the previous intra random access point sub-picture has a layer identifier of a network abstraction layer (NAL) unit and a same sub-picture index as the current sub-picture, and (c) the current sub-picture is a pure random access sub-picture, the format rule not allowing an entry in a reference picture list of the current slice to include a first picture preceding the current picture in the first order or the second order.

[0597] In some embodiments of the method 2800, the first order comprises an output order. In some embodiments of the method 2800, the second order comprises a decoding order. In some embodiments of the method 2800, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 2800, the reference picture list comprises a list 1 reference picture list.

[0598] Fig.29 is a flow chart of an example method 2900 for video processing. Operation 2902 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to (a) the current sub-picture being associated with an intra random access point sub-picture, and (b) the current sub-picture being after the intra random access point sub-picture in a first order, the format rule not allowing an active entry in a reference picture list of the current slice to include a first picture that is before the current picture in a first order or a second order.

[0599] In some embodiments of the method 2900, the first order comprises an output order. In some embodiments of the method 2900, the second order comprises a decoding order. In some embodiments of the method 2900, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 2900, the reference picture list comprises a list 1 reference picture list.

[0600] Fig.30 is a flow chart of an example method 3000 for video processing. Operation 3002 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, in response to: (a) the current sub-picture is after an intra random access point sub-picture in a first order, (b) the current sub-picture is after one or more preceding sub-pictures associated with an IRAP sub-picture in a first order and a second order, the format rule not allowing an entry in a reference picture list for the current slice to include a first picture that is before the current picture including the intra random access point sub-picture associated with the current sub-picture in the first order or the second order.

[0601] In some embodiments of the method 3000, the first order comprises an output order. In some embodiments of the method 3000, the second order comprises a decoding order. In some embodiments of the method 3000, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 3000, the reference picture list comprises a list 1 reference picture list.

[0602] Fig.31is a flow chart of an example method 3100 for video processing. Operation 3102 includes performing conversion between a video including a current picture including a current sub-picture and a bitstream of the video, the current sub-picture including a current slice, wherein the bitstream complies with a format rule, the format rule providing that, in response to the current sub-picture being a random access decodable leading sub-picture, a reference picture list for the current slice does not include an active entry for any one or more of: a first picture including a random access skipped leading sub-picture and a second picture preceding, in decoding order, a third picture including an associated intra-frame random access point sub-picture.

[0603] In some embodiments of the method 3100, the reference picture list comprises a list 0 reference picture list. In some embodiments of the method 3100, the reference picture list comprises a list 1 reference picture list.

[0604] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block may correspond to bits that are co-located or propagated at different positions in a bitstream defined by syntax. For example, a macroblock may be encoded based on an error residual value that has been converted and encoded and decoded, and bits in headers and other fields in the bitstream may also be used. In addition, during the conversion, the decoder may parse the bitstream based on the determination, knowing that some fields may or may not exist, as described in the above technical solution. Similarly, the encoder may determine whether to include or not include certain syntax fields, and generate a codec representation accordingly by including or not including the syntax fields from the codec representation.

[0605] The disclosed and other technical solutions, examples, embodiments, modules and functional operations described herein can be implemented in digital electronic circuits or computer software, firmware or hardware, including the structures disclosed herein and their structural equivalents, or a combination of one or more thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for a data processing device to execute or control its operation. The computer-readable medium can be a combination of a machine-readable storage device, a machine-readable storage substrate, a storage device, a material that affects a machine-readable propagation signal, or one or more combinations thereof. The term "data processing device" includes all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer or multiple processors or computers. In addition to hardware, the device may also include code that creates an execution environment for a computer program, for example, code that constitutes a processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. The propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.

[0606] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language (including compiled or interpreted languages) and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed for execution on one or more computers that are located at one site or distributed across multiple sites and interconnected by a communications network.

[0607] The processes and logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and the apparatus may also be implemented as, special purpose logic circuits, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0608] For example, processors suitable for executing computer programs include general and special purpose microprocessors, and any one or more of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as hard disks within a frame or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and memory may be supplemented by, or incorporated into, dedicated logic circuits.

[0609] Although this patent document contains many details, they should not be interpreted as limitations on any subject matter or the scope of the claims, but rather as descriptions of features of specific embodiments of specific technologies. Certain features described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment can also be implemented separately in multiple embodiments, or in any suitable subcombination. In addition, although the above-mentioned features may be described as working in certain combinations, or even initially required to be so, in some cases, one or more features in the claim combination can be removed from the combination, and the combination of claims can be directed to subcombinations or variations of subcombinations.

[0610] Likewise, although operations are described in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order or order shown, or that all illustrated operations be performed, in order to achieve the desired results. In addition, the separation of various system components in the embodiments of this patent document should not be understood as requiring such separation in all embodiments.

[0611] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing video data, include: performing conversion between the video and the bitstream according to a format rule to which the bitstream of the video including the picture containing two adjacent sub-pictures complies, wherein the format rule specifies that when a first syntax element has a first value for the picture, the two adjacent sub-pictures having different types of network abstraction layer NAL units have a second syntax element having the same second value, wherein the first syntax element having the first value indicates that each picture of the reference picture parameter set contains more than one video codec layer VCL NAL unit, and the more than one VCL NAL units of each picture of the reference picture parameter set do not have the same NAL unit type, and The second syntax element having the second value indicates that the corresponding sub-picture in the coding layer video sequence is regarded as a picture.

2. The method according to claim 1, in, The two adjacent sub-pictures include at least one of a P slice, a B slice, or an I slice.

3. The method according to claim 2, in, When the two adjacent sub-pictures have the second syntax element with the same third value, the two adjacent sub-pictures have the same type of NAL units.

4. The method according to claim 1, in, The first syntax element is included in the bitstream, and The first syntax element has a fourth value indicating that each picture referencing the picture parameter set contains one or more VCL NAL units, and the one or more VCL NAL units of each picture referencing the picture parameter set have the same NAL unit type.

5. The method according to claim 4, in, When the first syntax element has the first value, the value of the second syntax element of all sub-pictures in the picture and including at least one of a P slice, a B slice, or an I slice is equal to the second value.

6. The method according to claim 4, in, When the first syntax element has the fourth value, the picture is referred to as having the same NAL unit type as a codec slice NAL unit of the picture.

7. The method according to claim 1, in, The converting includes encoding the video into the bitstream.

8. The method according to claim 1, in, The converting includes decoding the video from the bitstream.

9. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, in, The instructions, when executed by the processor, cause the processor to: performing conversion between the video and the bitstream according to a format rule to which the bitstream of the video including the picture containing two adjacent sub-pictures complies, wherein the format rule specifies that when a first syntax element has a first value for the picture, the two adjacent sub-pictures having different types of network abstraction layer NAL units have a second syntax element having the same second value, wherein the first syntax element having the first value indicates that each picture of the reference picture parameter set contains more than one video codec layer VCL NAL unit, and the more than one VCL NAL units of each picture of the reference picture parameter set do not have the same NAL unit type, and The second syntax element having the second value indicates that the corresponding sub-picture in the coding layer video sequence is regarded as a picture.

10. The device according to claim 9, in, The two adjacent sub-pictures include at least one of a P slice, a B slice, or an I slice.

11. The device according to claim 10, in, When the two adjacent sub-pictures have the second syntax element with the same third value, the two adjacent sub-pictures have the same type of NAL units.

12. The device according to claim 9, in, The first syntax element is included in the bitstream, and The first syntax element has a fourth value indicating that each picture referencing the picture parameter set contains one or more VCL NAL units, and the one or more VCL NAL units of each picture referencing the picture parameter set have the same NAL unit type.

13. The device according to claim 12, in, When the first syntax element has the first value, the value of the second syntax element of all sub-pictures in the picture and including at least one of a P slice, a B slice, or an I slice is equal to the second value.

14. A non-transitory computer-readable storage medium storing instructions that cause a processor to: performing conversion between the video and the bitstream according to a format rule to which the bitstream of the video including the picture containing two adjacent sub-pictures complies, in, The format rule specifies that when a first syntax element has a first value for the picture, the two adjacent sub-pictures having network abstraction layer NAL units of different types have a second syntax element having the same second value, wherein the first syntax element having the first value indicates that each picture of the reference picture parameter set contains more than one video codec layer VCL NAL unit, and the more than one VCL NAL units of each picture of the reference picture parameter set do not have the same NAL unit type, and The second syntax element having the second value indicates that the corresponding sub-picture in the coding layer video sequence is regarded as a picture.

15. The non-transitory computer-readable storage medium of claim 14, in, The two adjacent sub-pictures include at least one of a P slice, a B slice, or an I slice.

16. The non-transitory computer-readable storage medium of claim 15, in, When the two adjacent sub-pictures have the second syntax element with the same third value, the two adjacent sub-pictures have the same type of NAL units.

17. The non-transitory computer-readable storage medium of claim 14, in, The first syntax element is included in the bitstream, and The first syntax element has a fourth value indicating that each picture referencing the picture parameter set contains one or more VCL NAL units, and the one or more VCL NAL units of each picture referencing the picture parameter set have the same NAL unit type.

18. The non-transitory computer-readable storage medium of claim 17, in, When the first syntax element has the first value, the value of the second syntax element of all sub-pictures in the picture and including at least one of a P slice, a B slice, or an I slice is equal to the second value.

19. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing device, in, The method comprises: generating a bitstream according to a format rule to which a bitstream of a video including a picture containing two adjacent sub-pictures complies, wherein the format rule specifies that when a first syntax element has a first value for the picture, the two adjacent sub-pictures having different types of network abstraction layer NAL units have a second syntax element having the same second value, wherein the first syntax element having the first value indicates that each picture of the reference picture parameter set contains more than one video codec layer VCL NAL unit, and the more than one VCL NAL units of each picture of the reference picture parameter set do not have the same NAL unit type, and The second syntax element having the second value indicates that the corresponding sub-picture in the coding layer video sequence is regarded as a picture.

20. The non-transitory computer-readable recording medium according to claim 19, in, The two adjacent sub-pictures include at least one of a P slice, a B slice, or an I slice.

21. A method for storing a bit stream of a video, include: generating a bitstream according to a format rule to which a bitstream of a video including a picture containing two adjacent sub-pictures complies; as well as storing the bitstream in a non-transitory computer-readable recording medium, wherein the format rule specifies that when a first syntax element has a first value for the picture, the two adjacent sub-pictures having different types of network abstraction layer NAL units have a second syntax element having the same second value, wherein the first syntax element having the first value indicates that each picture of the reference picture parameter set contains more than one video codec layer VCL NAL unit, and the more than one VCL NAL units of each picture of the reference picture parameter set do not have the same NAL unit type, and The second syntax element having the second value indicates that the corresponding sub-picture in the coding layer video sequence is regarded as a picture.

22. A video decoding device, comprising a processor configured to implement the method of claim 6 or 8.

23. A video encoding apparatus comprising a processor configured to implement the method of claim 6 or 7.

24. A non-transitory computer-readable storage medium storing a bit stream, wherein the bit stream is generated by a video processing device according to the method according to any one of claims 3 to 8.

25. A non-transitory computer-readable storage medium storing instructions that cause a processor to implement the method of any one of claims 6 to 8.

26. A method for generating a bit stream, include: Generating a video bitstream according to the method of any one of claims 2 to 8, and The bitstream is stored on a computer readable program medium.

Citation Information

Patent Citations

  • Picture alignments in multi-layer video coding

    US20140301437A1

  • Parameter set coding

    US20150195577A1