USE OF SUB-IMAGE INFORMATION IN VIDEO CODING

MX430946BActive Publication Date: 2026-02-25BYTEDANCE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2022011424
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2022-09-13
Publication Date
2026-02-25
Estimated Expiration
2041-03-18

AI Technical Summary

Technical Problem

Existing video coding standards, particularly VVC, lack clear definitions and constraints for managing mixed subimage types within an image, leading to confusion and inefficiencies in decoding processes, especially regarding output order and reference image management.

Method used

Introduce specific constraints and definitions for subimage types, such as requiring consistent NAL unit types within subimages and clarifying output orders, to ensure clear decoding processes and efficient handling of mixed subimage types.

Benefits of technology

Enhances decoding efficiency by providing clear guidelines for handling mixed subimage types, ensuring correct output order and reference image management, thereby improving the overall decoding process in video coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX430946B0
    Figure MX430946B0
Patent Text Reader

Abstract

Methods and devices for processing video are described. Processing may include encoding, decoding, or transcoding of video. An exemplary video processing method involves performing a conversion between a video comprising one or more images comprising one or more sub-images and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies a syntax of network abstraction layer (NAL) units in the bitstream, and wherein the format rule specifies that an NAL unit of a video encoding layer (VCL) NAL unit type includes content associated with a particular image type or a particular sub-image type.
Need to check novelty before this filing date? Find Prior Art

Description

USE OF SUB-IMAGE INFORMATION IN VIDEO ENCODING Cross reference to related requests Under patent law and / or rules applicable pursuant to the Paris Convention, this application is made to timely claim the priority and benefits of United States Provisional Patent Application number 62 / 992,724, filed on March 20 2020. For all purposes under the law, the entire description of the aforementioned applications are incorporated by reference as part of the description of this application. Field of Invention This patent document refers to image and video encoding and decoding. Background of the invention Digital video represents the largest use of bandwidth on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth for digital video use is expected to continue to grow. Brief description of the invention This document discloses techniques that can be used by video encoders and decoders to process the encoded representation of video using different syntax rules. In an exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule specifying a layer unit syntax. of network abstraction (NAL) in the bitstream, and where the format rule specifies that a NAL unit of a type of video coding layer (VCL) NAL unit acronym) includes content associated with a particular type of image or a particular type of sub-image. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising an image comprising a subimage and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that the subimage is a type of sub-image random access in response to the sub-image being an initial sub-image of an intra-random access point sub-image. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising an image comprising a subimage and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that one or more random access skipped initial subimages are absent from the bitstream in response to said one or more random access skipped initial subimages that are associated with an instantaneous decode update subimage. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising an image comprising a subimage and a bitstream of the video, wherein the bitstream conforms to a format rule specifying that one or more initial random access decodable subimages are absent from the bitstream in response to said one or more random access decodable initial subimages that are associated with an instantaneous decoding update subimage that has a type of network abstraction layer (NAL) unit indicating that the subimage Instant decode update is not associated with an initial image. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising an image comprising two neighboring subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that the two neighboring subimages with different types of Network abstraction layer (NAL) units have syntax elements with the same first value that indicates whether each of the two neighboring subimages in an encoded layer video sequence is treated as one image. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising an image comprising two neighboring subimages and a bitstream of the video, wherein the format rule specifies that the two neighboring subimages include a first neighboring subimage with a first subimage index and a second neighboring subimage with a second subimage index, and wherein the format rule specifies that the two neighboring subimages have the same type of network abstraction layer (NAL) units in response to a first syntax element associated with the first subimage index indicating that the first neighboring subimage is not treated as a subimage or a second syntax element associated with the second subimage index indicating that the second neighboring subimage is not treated as an image. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that an image is allowed to include more than two different types of video coding layer (VCL) network abstraction layer (NAL) units in response to a syntax element that indicates that each image of the video that references a set of image parameters (PPS, for its acronym in English) has a plurality of NAL VCL units that do not have the same type of NAL VCL unit. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a final subimage that is associated with an intra-random access point sub-image or a gradual decoding update sub-image follows the intra-random access point sub-image or the gradual decoding update sub-image in an order. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a subimage precedes in a first order to an intra-random hotspot subimage and one or more random access decodable initial subimages associated with the intrarandom hotspot subimage in response to: (1) that the subimage precedes the intrarandom hotspot subimage by a second order, (2) that the subimage and the intra-random access point subimage have the same first value for a layer to which a network abstraction layer (NAL) unit of the subimage and the access point subimage belong intrarandom (3) that the subimage and the intrarandom hotspot subimage have the same second value of a subimage index. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that an initial subimage omitted from random access subimage associated with a clean random access subimage precedes in an order one or more decodable initial random access subimages associated with the clean random access subimage. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that an initial subimage omitted from Random access associated with a clean random access subimage follows in a first order one or more intra-random access point subimages that precede the clean random access subimage in a second order. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a current subimage precedes in a decoding order to one or more non-initial subimages associated with an intra-random hotspot subimage in response to: (1) that a syntax element indicates that an encoded layer video sequence transmits images representing frames, and (2) ) that the current subimage is an initial subimage associated with the intra-random hotspot subimage. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that one or more types of network abstraction layer (NAL) unit for all Video Coding Layer (VCL) NALs in an image include RADL_NUT or RASL_NUT in response to the image being an initial image of an intra-random hotspot image. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising a plurality of subimages and a bitstream of the video, wherein the bitstream conforms to a format rule specifying that said at least one subimage precedes in a first order a gradual decoding update subimage and one or more subimages associated with the gradual decoding update subimage in response to: (1) that said at least one subimage precedes the gradual decoding update subimage in a second order, (2) that said at least one subimage and the gradual decoding update subimage have the same first value for a layer to which a network abstraction layer (NAL) unit of said at least one subimage belongs and the gradual decoding update subimage, and (3) that said at least one subimage and the gradual decoding update image have the same second value of a subimage index. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits an input active in a list of reference images of the current slice includes a first image that precedes in a decoding order a second image that includes a staged temporal sublayer access subimage in response to: (a) that the first image has a same temporal identifier and a same layer identifier of a network abstraction layer (NAL) unit as that of the current sub-image, and (b) that the current sub-image follows in encoding order the temporal sub-layer access sub-image staged, and (c) that the current subimage and the staged temporal sublayer access subimage have the same temporal identifier, the same layer identifier, and the same subimage index. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits an input active in a list of reference images of the current slice includes a first image that is generated by a decoding process to generate unavailable reference images in response to the current subimage not being of a particular type of subimage. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits an input in a list of reference images of the current slice include a first image that is generated by a decoding process to generate unavailable reference images in response to the current subimage not being of a particular type of subimage. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits an input in a list of reference images of the current slice include a first image that precedes in a first order or a second order the current image in response to: (a) that the first image includes a preceding intra-random hotspot subimage that precedes in the second order to the current subimage, (b) that the preceding intra-random access point subimage has the same layer identifier of a network abstraction layer (NAL) unit and the same subimage index as that of the subimage current, and (c) that the current subimage is a clean random access subimage. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits an input active in a list of reference images of the current slice includes a first image that precedes in a first order or a second order the current image in response to: (a) that the current subimage is associated with an intra-random hotspot subimage , (b) that the current subimage follows the intra-random hotspot subimage in the first order. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits an input in a list of reference images of the current slice include a first image that precedes in a first order or the second order the current image that includes an intra-random hotspot subimage associated with the current subimage in response to: (a) that the current sub-image follows the intra-random access point sub-image in the first order, (b) that the current sub-image follows one or more of the initial sub-images associated with the intra-random access point (IRAP) image sub-image in English) in the first order and the second order. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that in response Because the current subimage is a decodable random access initial subimage, a list of reference images of the current slice excludes an active entry for any or more of: a first image that includes an omitted random access initial subimage, and a second image preceding a third image that includes an associated intra-random hotspot subimage in a decoding order. In another exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video. The encoded representation conforms to a format rule specifying that said one or more images comprising one or more subimages are included in the encoded representation according to network abstraction layer (NAL) units, wherein a unit of type NAL is indicated by the encoded representation that includes an encoded slice of a particular type of image or an encoded slice of a particular type of a subimage. In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule that specifies that two neighboring subimages with different types of network abstraction layer unit will have the same indication of subimages that are treated as flag images. In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule that defines an order of a first type of subimage. and a second type of subimage, wherein the first subimage is a final subimage or an initial subimage or a random access skipped initial subimage (RASL) type and the second subimage is of type RASL or a type random access decodable initial subpicture (RADL) or an instantaneous decoding update (IDR) type or a gradual decoding update (GDR) type . In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule that defines a condition under which it is allowed or prohibits a first type of subimage from occurring with a second type of subimage. In yet another exemplary aspect, a video encoding apparatus is disclosed. The video encoder comprises a processor configured to implement the methods described above. In yet another exemplary aspect, a video decoding apparatus is disclosed. The video decoder comprises a processor configured to implement the methods described above. In yet another exemplary aspect, a computer readable medium having code stored therein is disclosed. The code incorporates one of the methods described herein in the form of processor-executable code. These, and other, features are described throughout this document. Brief description of the drawings Figure 1 shows an example of raster scan slice partitioning of an image, where the image is divided into 12 tiles and 3 raster scan slices. Figure 2 shows an example of rectangular slice partitioning of an image, where the image is divided into 24 tiles (6 columns of tiles and 4 rows of tiles) and 9 rectangular slices. Figure 3 shows an example of an image divided into tiles and rectangular slices, where the image is divided into 4 tiles (2 columns of tiles and 2 rows of tiles) and 4 rectangular slices. Figure 4 shows an image that is divided into 15 mosaics, 24 slices, and 24 subimages. Figure 5 is a block diagram of an exemplary video processing system. Figure 6 is a block diagram of a video processing apparatus. Figure 7 is a flow chart for an example method of video processing. Figure 8 is a block diagram illustrating a video encoder system according to some embodiments of the present disclosure. Figure 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure. Figure 10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure. Figures 11 to 31 are flow charts for example video processing methods. Detailed description of the invention Section headings are used herein for ease of understanding and do not limit the applicability of the techniques and modalities disclosed in each section to that section only. Additionally, H.266 terminology is used in some descriptions only for ease of understanding and not to limit the scope of the techniques disclosed. As such, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editing changes are shown in the text using open and closed double brackets (for example, [[ ]]) with deleted text between double brackets indicating canceled text and bold italic indicating added text, with regarding the current draft VVC specification. Introduction This document refers to video coding technologies. Specifically, it is about the definitions of subimage types and the relationships in terms of decoding order, output order, and prediction relationship between different types of subimages, in both single-layer and multi-layer contexts. The key is to clearly specify the meaning of mixed subimage types within an image through a set of constraints on decoding order, output order, and prediction relationship. The ideas can be applied individually, or in various combinations, to any standard video coding or non-standard video codec that supports multi-layer video coding, for example, Versatile Video Coding (VVC) that is developing. Abbreviations APS Adaptation Parameter Set AU Access Unit AUD Access Unit Delimiter 5 AVC Advanced Video Coding CLVS Encoded Layer Video Sequence CPB Encoded Image Buffer CRA Clean Random Access CTU Coding Tree Unit 10 CVS Video Sequence encoded DCI Decoding Capability Information DPB Decoded Image Buffer EOB End of Bitstream EOS End of Stream 15 GDR Incremental Decoding Update HEVC High Efficiency Video Coding HRD Hypothetical Reference Decoder IDR Instant Decoding Update JEM Scanning Model Joint 20 MCTS Motion Constrained Tile Sets NAL Network Abstraction Layer OLS Output Layer Set PH Image Header PPS Image Parameter Set 25 PTL Profile, Category and Level PU Image Unit Initial RADL Decodable random access RAP Random Access Point Random Access Initial RASL (Image) Skipped 30 RBSP Raw Byte Stream Payload RPL Reference Image List SEI Supplemental Enhancement Information SPS Stream Parameter Set STSA Temporal Sublayer Access by 35 stages SVC Scalable Video Coding VCL Video Coding Layer VPS Video Parameter Set VTM VVC Test Model VUI Video Usability Information VVC Versatile Video Coding Initial discussion Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 encoding standards. Advanced Video (AVC) and H.265 / HEVC. Since H.262, video coding standards are based on the hybrid video coding structure where temporal prediction plus transformation coding are used. In order to explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded in 2015 jointly between the Video Coding Expert Group (VCEG). for its acronym in English) and the Moving Picture Expert Group (MPEG, for its acronym in English). Since then, JVET has adopted many new methods and included them in reference software called the Joint Exploration Model (JEM). The concurrent JVET board is held once every three months and the new encoding standard targets a 50% bitrate reduction compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the April 2018 JVET meeting and the first version of VVC Test Mode (VTM) was released at that time. As there is an ongoing effort to contribute to the standardization of VVC, new coding techniques have been adopted to the VVC standard at each JVET board. The VVC working project and test model VTM are then updated after each meeting. The VVC project is now heading towards technical completion (Final draft International standard (FDIS)) at the July 2020 meeting. Image partitioning schemes in the HEVC HEVC includes four different image partitioning schemes, namely regular slices, dependent slices, mosaics, and wavefront parallel processing (WPP), which can be applied for unit size matching. maximum transfer rate (MTU), parallel processing and reduced end-to-end delay. Regular cuts are similar to those of H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and in-image prediction (intrasample prediction, motion information prediction, coding mode prediction) and entropy coding dependence across slice boundaries are disabled . Therefore, a regular slice can be reconstructed independently of other regular slices within the same image (although there may still be interdependencies due to loop filtering operations). Regular slicing is the only tool that can be used for parallelization that is also available, in almost identical form, in H.264 / AVC. Parallelization based on regular slices does not require much communication between processors or between cores (except for data exchange between processors or between cores for motion compensation when decoding a predictively encoded image, which is typically much heavier than the data exchange between processors or between cores due to prediction in the image). However, for the same reason, using regular slices can incur substantial coding overhead due to the bit cost of the slice header and due to the lack of prediction across slice boundaries. Additionally, regular slices (in contrast to the other tools mentioned below) also serve as the key mechanism for bitstream partitioning to match MTU size requirements, due to the image independence of the slices. regular and that each regular cut is encapsulated in its own NAL unit. In many cases, the goal of parallelization and the goal of MTU size matching place conflicting demands on the slice layout on an image. Carrying out this situation led to the development of the parallelization tools mentioned below. Dependent slices have short slice headers and allow bitstream partitioning across tree block boundaries without breaking any predictions in the image. Basically, dependent slices provide fragmentation of regular slices into multiple NAL units, to provide reduced end-to-end delay by allowing a portion of a regular slice to be sent before the encoding of the entire regular slice is completed. In WPP, the image is divided into individual rows of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use CTB data in other divisions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding of a CTB row is delayed by two CTBs, to ensure that data related to a CTB above and to the right of the Subject CTB are available before the subject CTB is decoded. By using this staggered start (which appears as a wavefront when plotted), parallelization is possible with up to as many processors / cores as the image contains rows of CTBs. Because in-image prediction is allowed between rows of neighboring tree blocks within an image, the inter-processor / inter-core communication required to enable in-image prediction can be substantial. WPP partitioning does not result in the production of additional NAL drives compared to when it is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP, with some coding overhead. Tiles define horizontal and vertical boundaries that divide an image into columns and rows of tiles. A tile column runs from the top of an image to the bottom of the image. Similarly, a row of tiles runs from the left of the image to the right of the image. The number of tiles in an image can be derived simply as the number of tile columns multiplied by the number of tile rows. The scan order of the CTBs is changed to be local within a tile (in the order of a CTB raster scan of a tile), before decoding the top left CTB of the next tile in the raster scan order mosaic of an image. Like the cuts IVIA / a / ZUZZ / U I 1 regular, mosaics break the prediction dependencies in the image as well as the entropy decoding dependencies. However, they do not need to be included in the individual NAL units (same as WPP in this regard); therefore, tiles cannot be used for MTU size matching. Each tile can be processed by one processor / core, and the interprocessor / intercore communication required for in-image prediction between processing units decoding neighboring tiles is limited to transmitting the shared slice header in cases where a cut spans more than one mosaic, and loop filtering related to the sharing of samples and reconstructed metadata. When more than one WPP tile or slice is included in a slice, the entry point byte offset for each WPP tile or slice other than the first in the slice is indicated in the slice header. For simplicity, restrictions on the application of the four different image partitioning schemes have been specified in the HEVC. A given encoded video sequence cannot include both mosaics and wavefronts for most of the profiles specified in the HEVC. For each slice and tile, one or both of the following conditions must be met: 1) all tree blocks encoded in a slice belong to the same tile; 2) all encoded tree blocks in a mosaic belong to the same slice. Finally, a wavefront slice contains exactly one CTB row, and when WPP is in use, if a slice starts within a CTB row, it must end in the same CTB row. A recent amendment to HEVC is specified in the Joint Collaborative Group for Video Coding (JCT-VC) output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan , A. Tourapis, Y.-K. Wang (eds.), HEVC Additional Supplemental Enhancement Information (Draft 4), October 24, 2017, publicly available at the following link: http: / / phenlx.lnt-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005v2.zip. With this amendment included, HEVC specifies three MCTS-related SEI messages, namely, MCTS temporary SEI message, MCTS extraction information set SEI message, and MCTS extraction information nesting SEI message. The Temporary MCTS SEI message indicates the existence of MCTS in the bitstream and flags the MCTS. For each MCTS, motion vectors are limited to pointing to full sample locations within the MCTS and to fractional sample locations requiring only full sample locations within the MCTS for interpolation, and the use of motion vector candidates is prohibited. for temporal motion vector prediction derived from blocks outside the MCTS. In this way, each MCTS can be decoded independently without the existence of tiles not included in the MCTS. The MCTS Extraction Information Sets SEI message provides supplementary information that can be used in the extraction of secondary MCTS bitstreams (specified as part of the SEI message semantics) to generate a conforming bitstream for a MCTS set. The information consists of a number of extract information sets, each of which defines a number of MCTS sets and contains RBSP bytes of the VPS, SPS and PPS of IVIA / a / ZUZZ / U I 1 replacement to be used during the MCTS secondary bitstream extraction process. When a secondary bitstream is extracted according to the MCTS secondary bitstream extraction process, the parameter sets (VPS, SPS and PPS) must be rewritten or replaced, the slice headers must be updated slightly because one or all syntax elements related to the slice address (including first_slice_segment_¡n_p¡c_flag and slice_segment_address) would normally need to have different values. Partitioning images in VVC In VVC, an image is divided into one or more rows of tiles and one or more columns of tiles. A mosaic is a CTU sequence that covers a rectangular region of an image. CTUs in a tile are scanned in raster scan order within that tile. A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of an image. Both slicing modes are supported, i.e. raster scan slicing mode and rectangular slicing mode. In raster scan slice mode, a slice contains a sequence of complete tiles in a mosaic raster scan of an image. In rectangular slice mode, a slice contains a series of complete tiles that collectively form a rectangular region of the image or a series of consecutive complete CTU rows of a tile that collectively form a rectangular region of the image. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to that slice. A subimage contains one or more slices that collectively cover a rectangular region of an image. Figure 1 shows an example of raster scan slice partitioning of an image, where the image is divided into 12 tiles and 3 raster scan slices. Figure 2 shows an example of rectangular slice partitioning of an image, where the image is divided into 24 tiles (6 columns of tiles and 4 rows of tiles) and 9 rectangular slices. Figure 3 shows an example of an image divided into tiles and rectangular slices, where the image is divided into 4 tiles (2 columns of tiles and 2 rows of tiles) and 4 rectangular slices. Figure 4 shows an example of subimage partitioning of an image, where an image is divided into 18 tiles, 12 on the left side, each covering a 4 by 4 CTU slice, and 6 tiles on the right side, each one covering 2 vertically stacked 2 by 2 CTU slices, resulting in total 24 slices and 24 subimages of varying dimensions (each slice is a subimage). Changing image resolution within a sequence In AVC and HEVC, the spatial resolution of images cannot change unless you start a new sequence using a new SPS, with an IRAP image. VVC allows the image resolution to change within a sequence at one position without encoding an IRAP image, which is always intra-encoded. This feature is sometimes referred to as a reference image resampling (RPR), since the feature needs to resample a reference image used for interprediction when that reference image has a different resolution than the current image that it is being encoded. The scaling ratio is constrained to be greater than or equal to 1 / 2 (2 times resolution reduction from the reference image to the current image), and less than or equal to 8 (8 times resolution increase). Three sets of resampling filters with different frequency cutoffs are specified to handle different scaling relationships between a reference image and the current image. The three sets of resampling filters are respectively applied for the scaling ratio ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same in the case of motion compensation interpolation filters. Actually, the normal MC interpolation process is a special case of resampling process with scale ratio ranging from 1 / 1.25 to 8. The horizontal and vertical scale ratios are derived based on the image width and height, and the specified left, right, top, and bottom scale offsets for the reference image and the current image. Other aspects of the VVC design to support this feature that are different from HEVC include: i) The image resolution and corresponding compliance window are noted in the PPS rather than the SPS, while the SPS indicates the maximum image resolution, i) For a single-layer bitstream, each image store (a slot in the DPB for storage of a decoded image) occupies a buffer size as required to store an encoded image that has the maximum image resolution. Scalable Video Coding (SVC) in General and VVC Scalable video coding (SVC, also sometimes referred to as simply scalable video coding) refers to video coding that uses a base layer (BL) sometimes referred to as a reference layer (RL), and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a base quality level. Such one or more enhancement layers may carry additional video data to support, for example, higher spatial, temporal and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to a previously encoded layer. For example, a lower layer can serve as a BL, while an upper layer can serve as an EL. Intermediate layers can serve as either EL or RL, or both. For example, an intermediate layer (e.g., a layer that is neither the bottommost layer nor the topmost layer) may be an EL for the layers below the intermediate layer, such as the base layer or any enhancement layer. that intervenes, and at the same time serve as an RL for one or more enhancement layers above the intermediate layer. Similarly, in the Multiview or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to encode (e.g., encode or decode) information from another view (e.g., estimation of motion, motion vector prediction and / or other redundancies). In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding level (for example, video level, sequence level, image level, slice level, etc.) in the that can be used. For example, parameters that can be used by one or more video sequences encoded from different layers in the bitstream can be included in a video parameter set (VPS), and parameters that are used by one or more images in an encoded video sequence can be included in a sequence parameter set (SPS). Similarly, parameters that are used by one or more slices in an image can be included in a picture parameter set (PPS), and other parameters that are specific to a single slice can be included in a slice header. Similarly, the indication of which set(s) of parameters a particular layer is using at a given time can be provided at different coding levels. With reference image resampling (RPR) support in VVC, you can design support for a bitstream containing multiple layers, for example two layers with standard definition (SD) and HD resolutions. High Definition (HD) in VVC without the need for any additional signal processing level encoding tools, as the upscaling required to support spatial scalability only uses the upscaling filter. RPR. However, high-level syntax changes (compared to not supporting scalability) are necessary for scalability support. Scalability support is specified in VVC version 1. Unlike scalability supports in any previous video coding standards, including extensions to AVC and HEVC, VVC's scalability design has been made friendly to decoder designs. Single layer as much as possible. The decoding capability for multi-layer bitstreams is specified as if there were only one layer in the bitstream. For example, decoding capability, such as DPB size, is specified in a manner that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for single-layer bitstreams doesn't need much change to be able to decode multi-layer bitstreams. Compared to AVC and HEVC multi-layer extension designs, high-level syntax (HLS) aspects have been significantly simplified at the sacrifice of some flexibilities. For example, an IRAP AU is required to contain an image for each of the layers present in the CVS. Random access and its supports in HEVC and VVC Random access refers to initiating access and decoding of a bitstream of an image that is not the first image of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, searching in local playback and streaming, as well as streaming adaptation, the bitstream needs to include frequent random access points, which are generally intra-coded images but They can also be intercoded images (for example, in the case of gradual decoding update). HEVC includes intra-random access point (IRAP) image signaling in the NAL unit header, via NAL unit types. Three types of IRAP images are supported, i.e. Instant Decode Update (IDR), Clean Random Access (ORA), and Broken Link Access (BLA) images. IDR images restrict the inter-image prediction structure to not reference any image before the current group of images (GOP), conventionally referred to as closed GOP random access points. ORA images are less restrictive by allowing certain images to reference images before the current GOP, all of which are discarded in the case of random access. ORA images are conventionally referred to as open GOP random access points. BLA images generally originate from splitting two bitstreams or part of them into an ORA image, for example during transmission switching. To enable better use of IRAP imaging systems, in total six different NAL units are defined to denote the properties of IRAP images, which can be used to better relate transmission access point types as defined in the ISO-based media file format (ISOBMFF), which are used for random access support in dynamic adaptive streaming over HTTP (DASH). VVC supports three types of IRAP images, two types of IDR images (one type with or the other type without associated RADL images), and one type of CRA image. These are basically the same as in HEVC. BLA image types in HEVC are not included in VVC, mainly due to two reasons: i) The basic functionality of BLA images can be realized by means of CRA images plus the end-of-sequence NAL unit, the presence of which indicates that the subsequent image initiates a new CVS in a single-layer bitstream, ii) There was a desire to specify fewer NAL unit types than in HEVC during the development of VVC, as indicated by the use of five instead of six bits for the NAL unit type field in the NAL unit header. Another key difference in random access support between VVC and HEVC is the support of GDR in a more prescriptive manner in VVC. In GDR, the decoding of a bitstream can start from an intercoded image and, although initially the entire image region cannot be correctly decoded but after a number of images, the entire image region would be correct. AVC in HEVC also supports GDR, using the recovery point SEI message to signal GDR random access points and recovery points. In VVC, a new NAL unit type is specified for GDR image indication and the recovery point is pointed out in the image header syntax structure. It is allowed to start a CVS and a bitstream with a GDR image. This means that a complete bitstream is allowed to contain only intercoded images without a single intracoded image. The primary benefit of specifying GDR support in this way is to provide compliant behavior for GDR. GDR allows encoders to smooth the bitrate of a bitstream by distributing intra-coded slices or blocks over multiple images as opposed to intra-coding entire images, thus enabling a significant reduction in end-to-end delay, which is considered more important today than sooner, as ultra-low lag applications such as wireless display, online gaming, and drone-based applications become more popular. Another feature related to GDR in VVC is virtual boundary signaling. The boundary between the updated region (that is, the correctly decoded region) and the non-updated region in an image between a GDR image and its recovery point can be marked as a virtual boundary, and when marked, the loop filtration across the entire boundary, therefore a decoding mismatch would not occur for some samples at or near the boundary. This can be useful when the application determines to display correctly decoded regions during the GDR process. IRAP images and GDR images can be collectively referred to as random access point (RAP) images. Management of reference images and reference image lists (RPL) Reference image management is a core functionality that is necessary for any video coding scheme that uses interprediction. It manages the storage and removal of reference images in and from a decoded image buffer (DPB) and places reference images in their appropriate order in the RPLs. HEVC reference image management, including marking and removing reference images from the decoded image buffer (DPB) as well as reference image list construction (RPLC), differs from that of AVC. Instead of the sliding window plus adaptive memory management control operation (MMCO) based reference image marking mechanism in AVC, HEVC specifies a reference image marking and management mechanism based on the so-called reference picture set (RPS), and RPLC is consequently based on the RPS mechanism. An RPS consists of a set of reference images associated with an image, consisting of all reference images that are prior to the associated image in decoding order, which can be used for interprediction for the associated image or any image that follows to the associated image in decoding order. The reference image set consists of five lists of reference images. The first three lists contain all reference images that can be used in interprediction of the current image and that can be used in interprediction of one or more of the images that follow the current image in decoding order. The other two lists consist of all reference images that are not used in interprediction of the current image but can be used in interprediction of one or more of the images following the current image in decoding order. RPS provides “intracoded” signaling of DPB status, rather than “intercoded” signaling as in AVC, primarily for improved error resiliency. The RPLC process in HEVC is based on the RPS, signaling an index in a subset of the RPS for each reference index; This process is simpler than the RPLC process in AVC. Reference image management in VVC is more similar to HEVC than AVC, but is somewhat simpler and more robust. As in those standards, two RPLs, list 0 and list 1, are derived, but they may not be based on the reference image set concept used in HEVC or the automatic sliding window process used in AVC; instead they are pointed out more directly. Reference images are listed for RPLs as any of active and active inputs, and only active inputs can be used as reference indices in interprediction of the CTUs of the reference image. Inactive entries indicate other images that will be held in the DPB for reference by other images that arrive later in the bitstream. Parameter sets AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in all AVC, HEVC and VVC. VPS was introduced from HEVC and was included in both HEVC and VVC. APS was not included in AVC or HEVC but was included in the latest VVC draft text. SPS was designed to carry sequence-level header information, and PPS was designed to carry image-level header information that changes infrequently. With SPS and PPS, information that changes infrequently does not need to be repeated for each image sequence, therefore redundant signaling of this information can be avoided. Additionally, the use of SPS and PPS allows out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmissions but also improving resilience to errors. VPS was introduced to carry stream-level header information that is common to all layers in multi-layer bitstreams. APS was introduced to carry such image-level or slice-level information that takes quite a few bits to encode, can be shared by multiple images, and in a sequence there can be many different variations. Related definitions in VVC The definitions related in the latest VVC text (in JVET-Q2001-vE / v15) are as follows. associated intra-random access point (IRAP) image (of a particular image): The previous IRAP image in decoding order (when present) that has the same nuhjayerjd value as the particular image. Clean Random Access (CRA) PU: A PU where the encoded image is a CRA image. clean random access (CRA) image: An IRAP image for which each NAL VCL unit has nal_unit_type equal to CRA_NUT. coded video sequence (CVS): A sequence of AUs consisting, in decoding order, of a coded video sequence (CVSS) start AU, followed by zero or more non-AUs CVSS, including all subsequent AUs but not including any subsequent AU that is a CVSS AU. Coded Video Sequence (CVSS) Start AU: An AU where there is one PU for each layer in the CVS and the encoded image in each PU is a Coded Layer Video Sequence (CLVSS) start image. English). Gradient Decoding Refresh (GDR) AU: An AU where the encoded image on each current PU is a GDR image. Gradient Decode Refresh (GDR) PU: A PU where the encoded image is a GDR image. rolling decode refresh (GDR) image: An image for which each NAL VCL unit has nal_unit_type equal to GDR_NUT. Instant Decode Refresh (IDR) PU: A PU where the encoded image is an IDR image. instant decode update (IDR) image: An IRAP image for which each NAL VCL unit has nal_unit_type equal to IDR W RADL or IDR N LP. Intra Random Access Point (IRAP) AU: An AU where there is one PU for each layer in the CVS and the encoded image in each PU is an IRAP image. Intra-Random Access Point (IRAP) PU: A PU where the encoded image is an IRAP image. intra-random access point (IRAP) image: An encoded image for which all NAL VCL units have the same nal_unit_type value in the range of IDR_W_RADL to CRA_NUT, inclusive. initial image: An image that is on the same layer as the associated IRAP image and precedes the associated IRAP image in output order. output order: The order in which decoded images are generated from the DPB (for decoded images to be output from the DPB). Random Access Decodable Initial PU (RADL): A PU where the encoded image is a RADL image. random access decodable initial image (RADL): An image for which each NAL VCL unit has nal_unit_type equal to RADL_NUT. Random Access Skipped Initial PU (RASL): A PU where the encoded image is a RASL image. random access skipped initial image (RASL): An image for which each NAL VCL unit has nal_unit_type equal to RASL_NUT. Staged Temporal Sublayer Access (STSA) PU: A PU where the encoded image is an STSA image. staged temporal sublayer access (STSA) image: An image for which each NAL VCL unit has nal_unit_type equal to STSA NUT. NOTE – An STSA image does not use images with the same Temporalld as the STSA image for the interprediction reference. Images following an STSA image in decoding order with the same Temporalld as the STSA image do not use images preceding the STSA image in decoding order with the same Temporalld as the STSA image for the interprediction reference. An STSA image allows upshifting, in the STSA image, to the sublayer containing the STSA image, from the sublayer immediately below. STSA images must have Temporalld greater than 0. ΜΛ / a / ZUZZ / U I Ί subimage: A rectangular region of one or more slices within an image. final image: A non-IRAP image that follows the associated IRAP image in output order and is not a STSA image. NOTE – The final images associated with an IRAP image also follow the IRAP image in decoding order. Images that follow the associated IRAP image in output order and precede the associated IRAP image in decoding order are not allowed. NAL unit header syntax and semantics in VVC In the latest VVC text (in JVET-Q2001-vE / v15), the NAL unit header syntax and semantics are as follows. 7.3.1.2 NAL Unit Header Syntax nal_unit_header() {Descriptor forbidden_zero_bit f(1) nuh_reserved_zero_bit u(1) nuhlayerid u(6) nal_unit_type u(5) nuh_temporaljd_plus1 u(3)} 7.4.2.2 Semantics of the NAL unit header forbidden_zero_bit shall be equal to 0. nuh_reserved_zero_bit shall be equal to 0. The value 1 of nuh_reserved_zero_bit may be specified in the future by ITU-T | ISO / IEC. Decoders will ignore (i.e. remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1. nuhjayerjd specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of a layer to which a non-VCL NAL unit applies. The value of nuh Jayer_id will be in the range 0 to 55, inclusive. Other values ​​for nuhjayerjd are reserved for future use by ITU-T | ISO / IEC. The value of nuhjayerjd will be the same for all NAL VCL units in an encoded image. The nuhjayerjd value of an encoded image or PU is the nuh_layer_id value of the NAL VCL units of the encoded image or PU. The value of nuhjayerjd for NAL access unit delimiter (AUD), PH, EOS, and fill data (FD) units is limited as follows: If nal_unitjype is equal to AUD_NUT, nuhjayerjd will be equal to vpsjayer idj 0 ]. Otherwise, when nal_unitjype is equal to PH_NUT, EOS_NUT, or FD_NUT, nuhjayerjd will be equal to nuhjayerjd of the associated NAL VCL unit. NOTE 1- The value of nuhjayerjd of DCI units, video parameter set (VPS) and NAL end of bitstream (EOB) is not restricted. The value of nal_unit_type will be the same for all images in a CVSS ALI. nal_unit_type specifies the type of the NAL unit, that is, the type of RBSP data structure contained in the NAL unit as specified in Table 5. NAL units having nal_unit_type in the range of UNSPEC_28..UNSPEC_31, inclusive, for which semantics are not specified, will not affect the decoding process specified in this specification. NOTE 2- NAL drive types in the range of UNSPEC28..UNSPEC31 may be used as determined by the application. No decoding process is specified in this specification for these nal_unit_type values. Since different applications may use these types of NAL units for different purposes, special care must be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the contents of NAL units with these values. nal_unit_type. This specification does not define any management for these values. These nal_unit_type values ​​may only be suitable for use in contexts where usage collisions (i.e. different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not important, or are not possible, or are managed, for example, defined or managed in the control application or transport specification, or by controlling the environment where the bitstreams are distributed. For purposes other than determining the amount of data in the DUs of the bitstream (as specified in Annex C), decoders must ignore (remove from the bitstream and discard) the contents of all NAL units that they use reserved nal_unit_type values. NOTE 3 – This requirement allows for the future definition of extensions compatible with this specification. Table 5 - NAL Unit Type Codes and NAL Unit Type Classes nal_unit_typ e Name of nal_unit_type Contents of the NAL unit and the RBSP syntax structure NAL unit type class 0 TRAILNUT Encoded cut of an image final slice layer rbsp() VCL 1 STSANUT Encoded cut of an image STSA slice layer rbsp() VCL 2 RADLNUT Coded cut of a RADL image slice layer rbsp() VCL 3 RASLNUT Coded cut of a RASL image slice layer rbsp() VCL 4..6 RSVVCL4.. RSV VCL 6 Reserved non-IRAP NAL VCL unit types VCL 7 8 IDR_W_RADL IDR N LP Coded cut of an image IDR slice layer rbsp() VCL 9 CRANUT Coded cut of an image CRA silce layer rbsp() VCL 10 GDR_NUT Coded cut of an image GDR slice layer rbsp() VCL 11 12 RSV_IRAP_11 RSV IRAP 12 Types of NAL VCL IRAP units reserved VCL 13 DCI_NUT decoding capability information decoding capability information rbsp() non-VCL 14 VPS_NUT video parameter set video parameter set rbsp() non-VCL 15 SPS_NUT Seq parameter set seq parameter set rbsp() non-VCL 16 PPS_NUT Picture parameter set pie parameter set rbsp() non-VCL 17 18 PREFIXAPSNUT SUFFIX APS NUT Adaptation parameter set adaptation parameter set rbsp() non-VCL 19 PHNUT Picture header picture header rbsp() non-VCL 20 AUDNUT AU of delimiter access unit delimiter rbsp() non-VCL 21 EOS_NUT End of sequence end of seq rbsp() non-VCL 22 EOBNUT End of bitstream end of bitstream rbsp() non-VCL 23 24 PREFIXSEINUT SUFFIX SEI NUT Information supplementary enhancement sei rbsp() non-VCL 25 FDNUT Filler data filler data rbsp() non-VCL 26 27 RSVNVCL26 RSV NVCL 27 Non-VCL NAL unit types reserved non-VCL 28..31 UNSPEC_28.. UNSPEC_31 NAL unit types non-VCL not specified non-VCL NOTE 4— A clean random access (ORA) image may have associated RASL or RADL images present in the bitstream. NOTE 5 – An instant decode update (IDR) image that has nal_unit_type equal to IDR_N_LP has no associated initial images present in the bitstream. An IDR image that has nal_unit_type equal to IDR_W_RADL does not have associated RASL images present in the bitstream, but may have associated RADL images in the bitstream. The value of nal_unit_type will be the same for all NAL VCL units in a subimage. A subimage is named as having the same NAL unit type as the NAL VCL units in the subimage. For NAL VCL units of any particular image, the following applies: If mixed_nalu_types_¡n_pic_flag is equal to 0, the value of nal_unit_type will be the same for all NAL VCL units in an image, and an image or PU is named as having the same NAL unit type as NAL units coded cutting of the image or PU. Otherwise (mixed_nalu_types_¡n_p¡c_flag is equal to 1), the image will have at least two subimages and the NAL VCL units in the image will have exactly two different nal_unit_type values ​​as follows: all NAL VCL units of at least a subimage of the image will have a particular value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA NUT, while all NAL VCL units of other subimages in the image will have a different particular value of nal_unit_type equal to TRAIL NUT, RADL_NUT, or RASL_NUT. For a single-layer bitstream, the following restrictions apply: Each image, other than the first image in the bitstream in decoding order, is considered associated with the previous IRAP image in decoding order. - When an image is an initial image of an IRAP image, it will be a RADL or RASL image. - When an image is a final image of an IRAP image, it will not be a RADL or RASL image. - There will be no RASL images present in the bitstream associated with an IDR image. - RADL images will not be presented in the bitstream associated with an IDR image that has nal_unit_type equal to IDR_N_LP. NOTE 6 – It is possible to perform random access on the position of an IRAP PU by discarding all PUs before the IRAP PU (and correctly decode the IRAP image and all subsequent non-RASL images in decoding order), provided each parameter set is available (either in the bitstream or by external means not specified in this specification) when referenced. Any image that precedes an IRAP image in decoding order will precede the IRAP image in output order and will precede any RADL image associated with the IRAP image in output order. Any RASL image associated with a CRA image must precede any RADL image associated with the CRA image in output order. Any RASL image associated with a CRA image must follow, in output order, any IRAP image that precedes the CRA image in decoding order. - If field_seq_flag is equal to 0 and the current image is an initial image associated with an IRAP image, it will precede, in decoding order, all non-initial images associated with the same IRAP image. Otherwise, if picA and picB are the first and last leading image, in decoding order, associated with an IRAP image, respectively, there will be at most one non-leading image preceding picA in decoding order, and there will be none non-main image between picA and picB in decoding order. nuh_temporal_id_plus1 minus 1 specifies a temporary identifier for the NAL unit. The value of nuh_temporal_id_plus1 will not be equal to 0. The Temporalld variable is derived as follows: Temporalld = nuhtemporaljdplusl - 1 (36) When nal_unit_type is in the range of IDRWRADL to RSV_IRAP_12, inclusive, Temporalld will be equal to 0. When nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[ GeneralLayerldx[ nuhjayerjd ] ] is equal to 1, Temporalld will not be equal to 0. The Temporalld value will be the same for all NAL VCL units in an AU. The Temporalld value of an encoded image, PU or AU is the Temporalld value of the NAL units VCL of the encoded image, PU or AU. The Temporalld value of a sublayer representation is the largest Temporalld value of all NAL VCL units in the sublayer representation. The Temporalld value for non-VCL NAL units is restricted as follows: - If nal_unit_type is equal to DCI_NUT, VPS_NUT or SPS_NUT, Temporalld will be equal to 0 and the Temporalld of the AU containing the NAL unit will be equal to 0. Otherwise, if nal_unit_type is equal to PH_NUT, Temporalld will be equal to the Temporalld of the PU containing the NAL unit. - Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, Temporalld will be equal to 0. - Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT or SUFFIX_SEI_NUT, Temporalld will be equal to the Temporalld of the AU containing the NAL unit. Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIXAPSNUT, Temporalld will be greater than or equal to the Temporalld of the PU containing the NAL unit. NOTE 7— When the NAL unit is a non-VCL NAL unit, the Temporalld value is equal to the minimum value of the Temporalld values ​​of all AUs to which the non-VCL NAL unit is applied. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT or SUFFIX_APS_NUT, Temporalld can be greater than or equal to the Temporalld of the containing AU, since all PPS and APS can be included at the beginning of the bitstream (for example, when transported out of band and the receiver places them at the beginning of the bit stream), where the first encoded image has Temporalld equal to 0. Mixed NAL drive types within an image 7.4.3.4 Semantics of the image parameter set mixed_nalu_types_in_pic_flag equal to 1 specifies that each image that refers to the PPS has more than one NAL VCL unit, NAL VCL units do not have the same nal_unit_type value, and the image is not an image intra-random access point (IRAP). mixed_nalu_types_in_pic_flag equal to 0 specifies that each image referring to the PPS has one or more NAL VCL units and the NAL VCL units of each image referring to the PPS have the same nal_unit_type value. When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of rect_slice_flag will be equal to 0. For each slice with a nal_unit_type value of nalUnitTypeA in the range of IDR_W_RADL to CRA_NUT, inclusive, in an image picA that also contains one or more slices with another value of nal_unit_type (i.e., the value of mixed_nalu_types_in_pic_flag for image picA is equal to 1 ), the following applies: The slice will belong to a subimage subpicA for which the value of the corresponding subpic_treated_as_pic_flag[ i ] is equal to 1. The slice will not belong to a subimage subpicA containing NAL VCL units with nal_unit_type not equal to nalUnitTypeA. If nalUnitTypeA is equal to CRA, for all subsequent PUs that follow the current image in the CLVS in decoding order and output order, neither RefPicList[0] nor RefPicList[1] of a cut in subpicA in those PUs will include any image before picA in decoding order on an active input. - Otherwise (i.e. nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS following the current image in decoding order, neither RefPicList[ 0 ] nor RefPicList[ 1 ] give a cut in subpicA in those PUs will include any image before picA in decoding order into an active input. NOTE 1- mixed_nalu_types_in_pic_flag equal to 1 indicates that images referring to the PPS contain slices with different types of NAL units, for example, encoded images originating from a subpicture bitstream merging operation for which the encoders have which ensure a matching bitstream structure and greater alignment of the parameters of the original bitstreams. An example of such alignments is the following: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_¡n_p¡c_flag is equal to 1, an image referring to the PPS cannot have slices with nal_unit_type equal to IDRWRADL or IDR_N_LP. Syntax and semantics of image header structure in VVC In the latest VVC text (in JVET-Q2001-vE / v15), the image header structure syntax and semantics that are most relevant to the technical solutions in this document are as follows. 7.3.2.7 Image Header Structure Syntax picture header structure() {Descriptor gdr or irap pie flag u(1) if( gdr or irap pie flag ) gdr pie flag u(1) ph pie order cnt Isb u(v) if( gdr or irap pie flag ) no output of prior feet flag u(1) if( gdr foot flag ) recovery poc cnt ue(v) ue(v)} 7.4.3.7 Semantics of image header structure The PH syntax structure contains information that is common to all slices of the encoded image associated with the PH syntax structure. gdrorirappicflag equal to 1 specifies that the current image is a GDR or IRAP image. gdr_or_irap_pic_flag equal to 0 specifies that the current image may or may not be a GDR or IRAP image. gdr_pic_flag equal to 1 specifies that the image associated with the PH is a GDR image. gdr_pic_flag equal to 0 specifies that the image associated with the PH is not a GDR image. When not present, the value of gdr_pic_flag is inferred to be equal to 0. When gdr_enabled_flag is equal to 0, the value of gdr_pic_flag will be equal to 0. NOTE 1 – When gdr_or_irap_pic_flag is equal to 1 and gdr_pic_flag is equal to 0, the image associated with the PH is an IRAP image. ph_pic_order_cnt_lsb specifies the MaxPicOrderCntLsb image order counting module for the current image. The length of the ph_pic_order_cnt_lsb syntax element is Iog2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of the ph_pic_order_cnt_lsb will be in the range 0 to MaxPicOrderCntLsb - 1, inclusive. no_output_of_prior_pics_flag affects the output of previously decoded images in the DPB after decoding a coded layer video sequence (CLVSS) start image that is not the first image in the bitstream as specified in Annex C. recovery_poc_cnt specifies the recovery point for decoded images in output order. If the current image is a GDR image that is associated with the PH, and there is an image picA that follows the current GDR image in decoding order in the CLVS that has PicOrderCntVal equal to PicOrderCntVal of the current GDR image plus the value of recovery_poc_cnt, the picA image is called the recovery point image. Otherwise, the first image in output order that has PicOrderCntVal greater than the PicOrderCntVal of the current image plus the recovery_poc_cnt value is called the recovery point image. The recovery point image should not precede the current GDR image in decoding order. The value of recovery_poc_cnt will be in the range 0 to MaxPicOrderCntLsb - 1, inclusive. When the current image is a GDR image, the RpPicOrderCntVal variable is derived as follows: RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt (81) NOTE 2: When gdr_enabled_flag is equal to 1 and PicOrderCntVal of the current image is greater than or equal to RpPicOrderCntVal of the associated GDR image, the current and subsequent decoded images in output order exactly match the corresponding images produced when starting the decoding process of the previous IRAP image, when present, preceding the associated GDR image in decoding order. Restrictions on RPLs in VVC In the latest VVC text (in JVET-Q2001 -vE / v15), the restrictions on RPLs in VVC are as follows (as part of VVC clause 8.3.2 Decoding process for construction of reference image lists) . 8.3.2 Decoding process for building reference image lists For each i equal to 0 or 1, the first entries of NumRefldxActive[ i] in RefPicList[ i] are named as the active entries in RefPicList[ i ], and the other entries in RefPicList[ i ] are named as the inactive entries in RefPicList[ i ]. NOTE 2 – It is possible for a particular image to be named by both an entry in RefPicList[ 0 ] and an entry in RefPicList[ 1 ]. It is also possible for a particular image to be named by more than one entry in RefPicList[ 0 ] by more than one entry in RefPicList[ 1 ]. NOTE 3 – Active entries in RefPicList[ 0 ] and active entries in RefPicList[ 1 ] are collectively referred to as all reference images that can be used for interprediction of the current image and one or more images following the current image in decoding order. Inactive entries in RefPicList[ 0 ] and inactive entries in RefPicList[ 1 ] are collectively referred to as all reference images that are not used for interprediction of the current image but can be used in interprediction for one or more images following the current image in decoding order. NOTE 4 – There may be one or more entries in RefPicList[ 0 ] or RefPicList[ 1 ] that are equal to “no reference image” because the corresponding images are not present in the DPB. Each inactive entry in RefPicList[ 0 ] or RefPicList[ 0 ] that is equal to “no reference image” should be ignored. An unintentional image loss must be inferred for each active entry in RefPicList[ 0 ] or RefPicList[ 1 ] as being equal to “no reference image”. It is a bitstream compliance requirement that the following restrictions apply: For each i equal to 0 or 1, num_ref_entries[ i ][ Rplsldx[ i ] ] will not be less than NumRefldxActive[ i ]. The image referenced by each active entry in RefPicList[ 0 ] or RefPicList[ 1 ] will be present in the DPB and will have Temporalld less than or equal to that of the current image. The image referenced by each entry in RefPicList[ 0 ] or RefPicList[ 1 ] will not be the current image and will have non_reference_picture_flag equal to 0. A short-term reference image (STRP) entry in RefPicList[ 0 ] orRefPicList[ 1 ] of a slice of an image and a long-term reference image (LTRP) entry English) in RefPicList[ 0 ] or RefPicList[ 1 ] of the same slice or a different slice of the same image will not refer to the same image. There may be no LTRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] for which the difference between the PicOrderCntVal of the current image and the PicOrderCntVal of the image referenced by the entry is greater than or equal to 224. Let setOfRefPics be the set of unique images referenced by all entries in RefPicList[ 0 ] that have the same nuhjayerjd as the current image and all entries in RefPicListj 1 ] that have the same nuhjayerjd as the current image . The number of images in setOfRefPics shall be less than or equal to MaxDpbSize - 1, inclusive, where MaxDpbSize is as specified in clause A.4.2, and setOfRefPics shall be the same for all slices of an image. When the current slice has nal_unit_type equal to STSA_NUT, there will be no active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that has Temporalld equal to that of the current image and nuh_layer_id equal to that of the current image. When the current image is an image that follows, in decoding order, a STSA image that has Temporalld equal to that of the current image and nuh_layer_id equal to that of the current image, there will be no image that precedes the STSA image in decoding order, has Temporalld equal to that of the current image, and has nuh layerjd equal to that of the current image that is included as an active entry in RefPicList[ 0 ] or RefPicList[ 1 ]. When the current image is a CRA image, there will be no image referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes, in output order or decoding order, any IRAP image in decoding order (when present). When the current image is a final image, there will be no image referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that was generated by the decoding process to generate reference images not available for the IRAP image associated with the current image. When the current image is a final image that follows, in both the decoding order and the output order, one or more initial images associated with the same IRAP image, if any, there will be no image to which reference by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that was generated by the decoding process to generate reference images not available for the IRAP image associated with the current image. When the current image is a recovery point image or an image that follows the recovery point image in output order RefPicList[ 0 ] or RefPicl_ist[ 1 ] containing an image that was generated by the decoding process to generate Reference images not available for the GDR image of the recovery point image. When the current image is a final image, there will be no image referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes the associated IRAP image in the output order or decoding order . When the current image is a final image that follows, in both the decoding order and the output order, one or more initial images associated with the same IRAP image, if any, there will be no image to which referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes the associated IRAP image in the output order or decoding order. When the current image is a RADL image, there will be no active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that is any of the following: An image of RASL An image that was generated by the decoding process to generate unavailable reference images An image preceding the associated IRAP image in decoding order The image referenced by each interlayer reference image (ILRP) entry in RefPicList[ 0 ] or RefPicList[ 1 ] of a slice of the current image will be in the same AU as the current image. The image referenced by each ILRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] of a slice of the current image will be present in the DPB and will have nuh_layer_id less than that of the current image. Each ILRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] of a slice will be an active entry. PictureOutputFIag Settings In the latest VVC text (in JVET-Q2001 -vE / v15), the specification for setting the value of the PictureOutputFIag variable is as follows (as part of clause 8.1.2 Decoding process for an encoded image). 8.1.2 Decoding process for an encoded image The decoding processes specified in this clause apply to each encoded image, named as the current image and denoted by the CurrPic variable, in BitstreamToDecode. Depending on the value of chroma_format_idc, the number of sample arrays of the current image is as follows: - If chroma_format_idc is equal to 0, the current image consists of 1 sample array Sl. - Otherwise (chroma_format_idc is not equal to 0), the current image consists of 3 sample arrays Sl, Scb, Ser. The decoding process for the current image, as inputs the uppercase syntax elements and variables of clause 7. When interpreting the semantics of each syntax element in each NAL unit, and in the remaining parts of clause 8, The term “the bitstream” (or part thereof, for example, a CVS bitstream) refers to BitstreamToDecode (or part thereof). Depending on the value of separate_colour_plane_flag, the decoding process is structured as follows: If separate_colour_plane_flag is equal to 0, the decoding process is invoked only once with the current image as the output. Otherwise (separate_colour_plane_flag is equal to 1), the decoding process is invoked three times. The inputs to the decoding processes are all NAL units of the encoded image with identical color_plane_id value. The process of decoding NAL units with a particular color_plane_id value is specified as if only one monochrome color formatted CVS with that particular color_plane_id value were present in the bitstream. The output of each of the three decoding processes is mapped to one of the 3 sample arrays of the current image, with the NAL units with color_plane_id equal to 0, 1, and 2 mapping to Sl, Scb, and Ser, respectively. NOTE-The ChromaArrayType variable is derived as equal to 0 when separate_colour_plane_flag is equal to 1 and chroma_format_idc is equal to 3. In the decoding process, the value of this variable is evaluated as resulting in operations identical to those for monochrome images ( when chroma_format_idc is equal to 0). The decoding process operates as follows for the current CurrPic image: Decoding of NAL units is specified in clause 8.2. The processes in clause 8.3 specify the following decoding processes using syntax elements at the slice header layer and above: - The variables and functions that relate to the image order count are derived as specified in clause 8.3.1. This needs to be invoked only for the first cut of an image. At the beginning of the decoding process for each slice of a non-IDR image, the reference image list construction decoding process specified in clause 8.3.2 is invoked for the derivation of reference image list 0 (RefPicList [ 0 ]) and the reference image list 1 (RefPicList[ 1 ]). The decoding process for reference image marking is invoked in clause 8.3.3, where reference images can be marked as “not used for reference” or “used for long-term reference”. This needs to be invoked only for the first cut of an image. When the current image is a CRA image with NoOutputBeforeRecoveryFlag equal to 1 or GDR image with NoOutputBeforeRecoveryFlag equal to 1, the decoding process is invoked to generate unavailable reference images specified in subclause 8.3.4, which needs to be invoked only for the first cut of an image. - PictureOutputFIag is set as follows: -If one of the following conditions is true, PictureOutputFIag is set equal to 0 - the current image is a RASL image and NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1. - gdr_enabled_flag is equal to 1 and the current image is a GDR image with NoOutputBeforeRecoveryFlag equal to 1. - gdr_enabled_flag is equal to 1, the current image is associated with a GDR image with NoOutputBeforeRecoveryFlag equal to 1, and PicOrderCntVal of the current image is less than RpPicOrderCntVal of the associated GDR image. - sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picA image that satisfies all of the following conditions: - PicA has PictureOutputFIag equal to 1. - PicA has nuh layerjd nuhLid greater than that of the current image. - PicA belongs to the output layer of the OLS (i.e. OutputLayerldlnOls[ TargetOlsIdx ][ 0 ] is equal to nuhLid). - sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[ TargetOlsIdx ][ GeneralLayerldx[ nuhjayerjd ] ] is equal to 0. -Otherwise, PictureOutputFIag is set equal to pic_output_flag. The processes in clauses 8.4, 8.5, 8.6, 8.7, and 8.8 specify decoding processes that use syntax elements at all syntax structure layers. It is a requirement of bitstream compliance that the encoded slices of the image contain slice data for each CTU of the image, such that each of the division of the image into slices, and the division of the slices into CTU form an image partition. After all slices of the current image have been decoded, the current decoded image is marked as “used for short-term reference,” and each ILRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] is marked as “used as a short-term reference. Technical problems resolved by the disclosed technical solutions The existing design in the latest VVC text (in JVET-Q2001-vE / v15) has the following problems: Because mixing different types of subimages within an image is permitted, it is confusing to call the contents of a NAL unit with a VCL NAL unit type a coded slice of a particular type of image. For example, a NAL unit with nal_unit_type equal to CRA_NUT is an encoded slice of a CRA image only when all slices in the image have nal_unit_type equal to CRA_NUT; when a slice of this image has nal_unit_type other than CRA NUT, then the image is not a CRA image. Currently, the value of subpic_treated_as_pic_flag[ ] is required to be equal to 1 for a subimage if the subimage contains a NAL VCL unit with nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive, and mixed_nalu_types_in_pic_flag equal to 1 for the image. In other words, the value of subpic_treated_as_pic_flag[ ] is required to be equal to 1 for an IRAP subimage mixed with another type of subimage in an image. However, with the support of more mixes of NAL VCL drive types, this requirement is not sufficient. Currently, only up to two different types of NAL VCL units (and two different types of subimages) are allowed within an image. There is a missing restriction on the output order of a final subimage relative to the associated IRAP or GDR subimage, in both single-layer and multi-layer contexts. Currently, it is specified that when an image is a seed image of an IRAP image, it will be a RADL or RASL image. This restriction, together with the initial / RADL / RASL image definitions, prohibits the mixing of RADL and RASL NAL unit types within an image resulting from mixing two CRA images and their associated non-aligned RADL and RASL images. AU. A constraint on the subimage type (that is, the NAL unit type of the NAL VCL units in a subimage) is missing for an initial subimage, in both single-layer and multi-layer contexts. There is a missing restriction on whether a RASL subimage can be present and associated with an IDR subimage, in both single-layer and multi-layer contexts. A restriction is missing on whether a RADL subimage can be present and associated with an IDR subimage that has nal_unit_type equal to IDR_N_LP, in both single-layer and multi-layer contexts. There is a missing restriction on the relative output order between a subimage preceding an IRAP subimage in decoding order and the RADL subimages associated with the IRAP subimage, in both single-layer and multi-layer contexts. There is a missing constraint on the relative output order between a subimage preceding a GDR subimage in decoding order and the subimages associated with the GDR subimage, in both single-layer and multi-layer contexts. There is a missing constraint on the relative output order between a RASL subimage associated with a CRA subimage and a RADL subimage associated with the CRA subimage, in both single-layer and multi-layer contexts. There is a missing restriction on the relative output order between a RASL subimage associated with a CRA subimage and an IRAP subimage preceding the CRA subimage in decoding order, in both single-layer and multi-layer contexts. There is a missing constraint on the relative decoding order between an IRAP image associated with non-initial images and initial images, in both single-layer and multi-layer contexts. A constraint is missing on active RPL entries for a subimage following an STSA subimage in decoding order, in both single-layer and multi-layer contexts. A constraint is missing on the RPL inputs for a CRA subimage in both single-layer and multi-layer contexts. There is a missing constraint on the active RPL entries for a subimage that refers to an image that was generated by the decoding process to generate unavailable reference images, in both single-layer and multi-layer contexts. A constraint is missing on the RPL entries for a subimage that refers to an image that was generated by the decoding process to generate unavailable reference images, in both single-layer and multi-layer contexts. A constraint is missing on the active RPL entries for a subimage associated with an IRAP image and following the IRAP image in output order, in both single-layer and multi-layer contexts. A constraint is missing on the RPL inputs for a subimage associated with an IRAP image and following the IRAP image in output order, in both single-layer and multi-layer contexts. A constraint is missing on the active RPL entries for a RADL subimage in both single-layer and multi-layer contexts. Examples of solutions and modalities To solve the above problems, and others, methods are disclosed as summarized below. The points should be considered as examples to explain general concepts and should not be interpreted in a restrictive manner. Furthermore, these points can be applied individually or combined in any way. To resolve issue 1, instead of specifying the content of a NAL unit with a VCL NAL unit type as “encoded slice of a particular image type,” specify “encoded slice of a particular image or subimage type.” ”. For example, the content of a NAL unit with nal_unit_type equal to CRA_NUT is specified as “encoded slice of a CRA image or subimage”. Additionally, one or more of the following terms are defined: associated GDR subimage, associated IRAP subimage, CRA subimage, GDR subimage, IDR subimage, IRAP subimage, initial subimage, RADL subimage, RASL subimage, subimage from STSA, final subimage. To solve problem 2, add a constraint to require that any pair of neighboring subpictures with different NAL unit landmarks have both subpic_treated_as_pic_flag[ ] equal to 1. In one example, the constraint is specified as follows: For any pair of neighboring subimages with subimage indices i and j in an image, when subpic_treated_as_pic_flag[ i ] or subpic_treated_as_pic_flag[ j ] is equal to 0, the two subimages will have the same unit type from NAL. Alternatively, it is required that, when any subimage with subimage index i has subpic_treated_as_pic_flag[ i ] equal to 0, all subimages in an image will have the same NAL unit type (i.e., all NAL VCL units in an image will have the same NAL unit type, i.e. the value of mixed_nalu_types_in_pic_flag will be equal to 0). And this means that mixed_nalu_typesjn_pic_flag can only be equal to 1 when all subpictures have their corresponding subpic_treated_as_pic_flag[ ] equal to 1. To solve problem 3, when mixed_nalu_types_in_pic_flag is equal to 1, an image can be allowed to contain more than two different types of NAL VCL units. To resolve issue 4, we specify that a final subimage will follow the associated IRAP or GDR subimage in output order. To resolve issue 5, to allow mixing of RADL and RASL NAL unit types within an image resulting from mixing the two CRA images and their associated non-AU-aligned RADL and RASL images, the existing constraint specifying that an initial image of an IRAP image will be a RADL or RASL image is changed to be as follows: When an image is an initial image of an IRAP image, the value of nal_unit_type for all NAL VCL units in the image will be equal to RADL_NLJT or RASL_NUT. Additionally, in the decoding process for an image with mixed nal_unitjype values ​​of RADL_NUT and RASL_NUT, the PictureOutputFIag of the image is set equal to pic_output_flag when the layer containing the image is an output layer. In this way, by constraining that all images being output need to be corrected for conformance decoders, RADL subimages within such images can be guaranteed, although the guarantee of the “accuracy” of the RASL subimages of “value medium” within such images when the associated CRA image has NoOutputBeforeRecoveryFlag equal to 1 is also in place but is not actually required. The unnecessary part of the guarantee does not matter and does not add complexity to implementing compliance encoders or decoders. In this case, it would be useful to add a NOTE clarifying that although such RASL subimages associated with a CRA image with NoOutputBeforeRecoveryFlag equal to 1 may be output by the decoding process, they are not intended to be used for display and therefore Therefore they should not be used for display. To resolve issue 6, we specify that when a subimage is an initial subimage of an IRAP subimage, it must be a RADL or RASL subimage. To resolve issue 7, it is specified that no RASL subimage must be present in the bitstream that is associated with an IDR subimage. To resolve issue 8, we specify that no RADL subimage must be present in the bitstream that is associated with an IDR subimage that has nal_unit_type equal to IDRNLP. To solve problem 9, we specify that any subimage, with nuhjayerjd equal to a particular layerld value and subimage index equal to a particular subpicldx value, that precedes, in decoding order, an IRAP subimage with nuhjayerjd equal to layerld and subimage index equal to subpicldx will precede, in output order, the IRAP subimage and all its associated RADL subimages. To solve problem 10, we specify that any subpicture, with nuhjayerjd equal to a particular layerld value and subpicture index equal to a particular subpicldx value, that precedes, in decoding order, a GDR subpicture with nuhjayerjd equal to layerld and subimage index equal to subpicldx will precede, in output order, the GDR subimage and all its associated subimages. To resolve issue 11, we specify that any RASL subimage associated with a CRA subimage will precede any RADL subimage associated with the CRA subimage in output order. To resolve problem 12, we specify that any RASL subimage associated with a CRA subimage will follow, in output order, any IRAP subimage that precedes the CRA subimage in decoding order. To solve problem 13, we specify that if field_seqjlag is equal to 0 and the current subimage, with nuhjayerjd equal to a particular subimage index layerld value equal to a particular subpicldx value, is an initial subimage associated with an IRAP subimage, shall precede, in decoding order, all non-initial subpictures that are associated with the same ILRP subpicture; Otherwise, let subpicA and subpicB be the first and last initial subimages, in decoding order, associated with an IRAP subimage, respectively, there will be at most one non-initial subimage with nuhjayerjd equal to layerld and subimage index equal to subpicldx preceding subpicA in decoding order, and there will be no non-initial image with nuhjayerjd equal to layerld the index of subimage id equal to subpicldx between picA and picB in decoding order. To solve problem 14, we specify that when the current subimage, with Temporalld equal to a particular tld value, nuhjayerjd equal to a particular layerld value, and subimage index equal to a particular subpicldx value, is a subimage that follows, in order of decoding, an STSA subimage with Temporalld equal to tld, nuhjayerjd equal to layerld, and subimage index equal to subpicldx, there will be no image with Temporalld equal to tld and nuhjayerjd equal to layerld that precedes the image containing the subimage of STSA in decoding order included as an active entry in RefPicList[ 0 ] or RefPicList[ 1 ]. To solve problem 15, we specify that when the current subimage, with nuhjayerjd equal to a particular layerld value and subimage indices equal to a particular subpicldx value, is a CRA subimage, there will be no image referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes, in output order or decoding order, any image containing an IRAP subimage with nuhjayerjd equal to layerld and subimage indices equal to subpicldx in decoding order ( when present). To resolve problem 16, we specify that when the current subimage, with nuhjayerjd equal to a particular layerld value and subimage index equal to a particular subpicldx value, is not a RASL subimage associated with a CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a GDR subimage of a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or a subimage of a recovery image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1 and nuhjayerjd equal to layerld, there will be no image a which is referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that was generated by the decoding process to generate unavailable reference images. To solve problem 17, we specify that when the current subimage, with nuhjayerjd equal to a particular layerld value and subimage index equal to a particular subpicldx value, is not a CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a subimage preceding, in decoding order, the initial subimages associated with the same CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, an initial subimage associated with a CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a GDR subimage of a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or a subimage of a recovery image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1 and nuhjayerjd equal to layerld, there will be no image referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] which was generated by the decoding process to generate unavailable reference images. To resolve issue 18, we specify that when the current subimage is associated with an IRAP subimage and follows the IRAP subimage in output order, there will be no image referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] preceding the image containing the associated IRAP subimage in output order or decoding order. To resolve problem 19, we specify that when the current subimage is associated with an IRAP subimage, it follows the IRAP subimage in output order, and follows, in both decoding order and output order, the initial subimages associated with the same IRAP subimage, if any, there will be no image referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes the image containing the associated IRAP subimage in order of output or decoding order. To resolve issue 20, we specify that when the current subimage is a RADL subimage, there will be no active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that is any of the following: An image containing a RASL subimage An image preceding the image containing the associated IRAP subimage in decoding order Examples of modality Below are some example modalities for some of the technical solution aspects summarized above in Section 5, which can be applied to the VVC specification. The modified texts are based on the latest VVC text in JVET-Q2001-vE / v15). The most relevant parts that have been added and modified are highlighted in bold italics, and some of the removed parts are highlighted in double open and closed brackets (for example, [[ ]]) with the deleted text between the double brackets. There are some other changes that are editorial in nature or are not part of this technical fix and therefore are not highlighted. First modality This modality is for points 1, 1a, 2, 2a, 4, and 6 to 20. Definitions associated GDR image (of a particular image with a particular value of nuhjayerjd layerld): The previous GDR image in decoding order with nuhjayerjd equal to layerld (when present) between which and the particular image in decoding order no There is no IRAP image with nuhjayerjd equal to layerld. associated GDR subimage (of a particular subimage with a particular value of nuh_layer_ld layerld and a particular value of subimage index subplcture Index subpicldx): The previous GDR subimage in decoding order with nuh_layer_id equal to layerld and subimage index equal to subpicldx (when present) between which and the particular subimage in decoding order there is no IRAP subimage with nuh_layer_id equal to layerld and subimage index equal to subpicldx. associated IRAP image (of a particular image with a particular value of nuh_layer_id layerld): The previous IRAP image in decoding order with nuhjayerjd equal to layerld (when present) between which and the particular image in decoding order there is no no GDR image with nuhjayerjd equal to layerld. associated IRAP subpicture (of a particular subpicture with a particular value of nuh_layer_id layerld and a particular value of subpicture index subpicture Index subpicldx): The previous IRAP subpicture in decoding order with nuh_layer_id equal to layerld and subpicture index equal to subpicldx (when present) between which and the particular subpicture in decoding order there is no GDR subpicture with nuh_layer_id equal to layerld and subpicture index equal to subpicldx. clean random access (CRA) image: An IRAP image for which each NAL VCL unit has nal_unitjype equal to CRA_NUT. Clean Random Access (CRA) Subimage: An IRAP subimage for which each NAL VCL unit has nal_unit_type equal to CRA_NUT. Gradient Decoding Refresh (GDR) AU: An AU where there is one PU for each layer in the CVS and the encoded image in each PU is a GDR image. rolling decode refresh (GDR) image: An image for which each NAL VCL unit has nal_unitjype equal to GDRNUT. gradual decoding update (GDR) subimage: A subimage for which each NAL VCL unit has nal_unit_type equal to GDR_NUT. instant decode update (IDR) image: An IRAP image for which each NAL VCL unit has nal_unitjype equal to IDR_W_RADL or IDR_N_LP. instant decode update (IDR) subimage: An IRAP subimage for which each NAL VCL unit has nal_unit_type equal to IDR_W_RADL or IDR_N_LP. intra-random access point (IRAP) image: An image for which all NAL VCL units have the same nal_unitjype value in the range of IDR_W_RADL to CRA NUT, inclusive. intra-random access point (IRAP) subimage: A subimage for which all NAL VCL units have the same nal_unit_type value in the range of IDR_W_RADL to CRA_NUT, inclusive. initial image: An image that precedes the associated IRAP image in output order. initial subimage: An image that precedes the associated IRAP subimage in output order. output order: The order of images or subimages within a CL VS indicated by increasing picture order count (POC) values, and for decoded images that are DPB output, this is the order in which which the decoded images are output from the DPB. random access decodable initial image (RADL): An image for which each NAL VCL unit has nal_unit_type equal to RADL_NUT. random access decodable initial subimage (RADL): An encoded image for which each NAL VCL unit has nal_unit_type equal to RADL_NUT. initial random access image skipped (RASL): An image for which each NAL unit VCLWene nal_unit_type equal to RASL_NUT. initial random access subimage skipped (RASL): An encoded image for which each NAL VCL unit has nal_unit_type equal to RASL_NUT. staged temporal sublayer access (STSA) image: An image for which each NAL VCL unit has nal_unit_type equal to STSA_NUT. staged temporal sublayer access (STSA) subimage: An encoded subimage for which each NAL VCL unit has nal_unit_type equal to STSA_NUT. final image: An image for which each VCL NAL unit has nal_unit_type equal to TRAIL_NUT. NOTE – Final images associated with an IRAP or GDR image also follow the IRAP or GDR image in decoding order. Images that follow the associated IRAP or GDR image in output order and precede the associated IRAP or GDR image in decoding order are not permitted. final subimage: A subimage for which each NAL VCL unit has nal_unit_type equal to TRAIL_NUT. NOTE – Final subpictures associated with an IRAP or GDR subpicture also follow the IRAP or GDR subpicture in decoding order. Subpictures that follow the associated IRAP or GDR subpicture in output order and precede the associated IRAP or GDR subpicture in decoding order are not permitted. 7.4.2.2 Semantics of the NAL unit header nal_unit_type specifies the type of the NAL unit, that is, the type of RBSP data structure contained in the NAL unit as specified in Table 5. NAL units having nal_unit_type in the range of UNSPEC_28..UNSPEC_31, inclusive, for which semantics are not specified, will not affect the decoding process specified in this specification. NOTE 2- NAL unit types in the range of UNSPEC_28..UNSPEC_31 may be used as determined by the application. No decoding process is specified in this specification for these nal_unit_type values. Since different applications may use these types of NAL units for different purposes, special care must be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the contents of NAL units with these values. nal_unit_type. This specification does not define any management for these values. These nal_unit_type values ​​may only be suitable for use in contexts where usage collisions (i.e. different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not important, or are not possible, or are managed, for example, defined or managed in the control application or transport specification, or by controlling the environment where the bitstreams are distributed. For purposes other than determining the amount of data in the DUs of the bitstream (as specified in Annex C), decoders must ignore (remove from the bitstream and discard) the contents of all NAL units that they use reserved nal_unit_type values. NOTE 3 – This requirement allows for the future definition of extensions compatible with this specification. Table 5- NAL unit type codes and NAL unit type classes nal_unit_typ e Name of nal_unit_type Contents of the NAL unit and the RBSP syntax structure NAL unit type class 0 TRAILNUT Encoded cut of a final image or subimage slice layer rbsp() VCL 1 STSA_NUT Encoded cut of a STSA image or subimage slice layer rbsp () VCL 2 RADLNUT Coded cut of a RADL image or sub-image slice layer rbsp() VCL 3 RASLNUT Coded cut of a RASL image or sub-image slice layer rbsp() VCL 4..6 RSVVCL4.. RSV VCL 6 Unit types NAL VCL no IRAP reserved VCL 7 8 IDRWRADL IDR N LP Coded cut of an IDR image or sub-image slice layer rbsp() VCL 9 CRANUT Coded cut of a CRA image or sub-image silce layer rbsp() VCL 10 GDRNUT Coded cut of a GDR image or sub-image slice layer rbsp() VCL 11 12 RSV_IRAP_11 RSV IRAP 12 NAL unit types VCL IRAP reserved VCL 13 DCINUT Decoding capability information decoding capability Information rbsp() non-VCL 14 VPS_NUT Video parameter set video parameter set rbsp() non -VCL 15 SPS_NUT Sequence parameter set seq parameter set rbsp() non-VCL 16 PPS_NUT Footer image parameter set .parameter. set. rbsp() non-VCL 17 18 PREFIX_APS_NUT SUFFIX APS NUT adaptation parameter set adaptation parameter set rbsp() non-VCL 19 PHNUT Picture header picture header rbsp() non-VCL 20 AUD_NUT AU delimiter access unit delimiter rbsp() non-VCL 21 EOS_NUT End of sequence end of seq rbsp() non-VCL 22 EOBNUT End of bitstream end of bitstream rbsp() non-VCL 23 24 PREFIXSEINUT SUFFIX SEI NUT Supplementary enhancement information sei rbsp() non-VCL 25 FDNUT Padding data filler data rbsp() non-VCL 26 27 RSVNVCL26 RSV NVCL 27 Non-VCL NAL unit types reserved non-VCL 28..31 UNSPEC_28.. UNSPEC 31 Non-VCL NAL unit types not specified non-VCL NOTE 4 – A clean random access (ORA) image may have associated RASL or RADL images present in the bitstream. NOTE 5 – An instant decode update (IDR) image that has nal_unit_type equal to IDR_N_LP has no associated initial images present in the bitstream. An IDR image that has nal_unit_type equal to IDR_W_RADL does not have associated RASL images present in the bitstream, but may have associated RADL images in the bitstream. The value of nal_unit_type will be the same for all NAL VCL units in a subimage. A subimage is designated as having the same NAL unit type as the NAL VCL units in the subimage. For any pair of neighboring subpictures with subpicture indices i and j in an image, when subpic_treated_as_pic_flag[ i] or subpic_treated_as_pic_flag[ / ] is equal to 0, the two subpictures will have the same NAL unit type. For NAL VCL units of any particular image, the following applies: - If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type will be the same for all NAL VCL units in an image, and an image or a PU is named as having the same NAL unit type as the NAL VCL units of the image or PU. - Otherwise (mixed_nalu_types_in_p¡c_flag is equal to 1), the image will have at least two subimages and the NAL VCL units in the image will have exactly two different nal_unit_type values ​​as follows: all NAL VCL units of at least one subimage of the image will have a particular nal unit type value equal to STSANUT, RADL NUT, RASLNUT, IDR W RADL, IDR_N_LP, or CRA NUT, while all NAL VCL units of other subimages in the image will have a different particular value of nal_unit_type equal to TRAIL NUT, RADL_NUT, or RASL_NUT. It is a bitstream compliance requirement that the following restrictions apply: A final image will follow the associated IRAP or GDR image in output order. A final subimage will follow the associated IRAP or GDR subimage in output order. When an image is a seed image of an IRAP image, it will be a RADL or RASL image. - When a subimage is an initial image of an IRAP subimage, it will be a RADL or RASL subimage. - There will be no RASL images present in the bitstream associated with an IDR image. There will be no RASL subimages present in the bitstream associated with an IDR subimage. - RADL images will not be presented in the bitstream associated with an IDR image that has nal_unit_type equal to IDR_N_LP. NOTE 6 – It is possible to perform random access on the position of an IRAP PU by discarding all PUs before the IRAP PU (and correctly decode the IRAP image and all subsequent non-RASL images in decoding order), provided each parameter set is available (either in the bitstream or by external means not specified in this specification) when referenced. No RADL subimages will be presented in the bitstream associated with an IDR subimage that has nal_unit_type equal to IDR_N_LP. Any image, with nuhjayerjd equal to a particular layerld value, that precedes, in decoding order, an IRAP image with nuhjayerjd equal to layerld will precede, in output order, the IRAP image and all its associated RADL images. - Any subimage, with nuh_layer_id equal to a particular layerld value and subimage index equal to a particular subpicldx value, that precedes, in decoding order, an IRAP subimage with nuh_layer_id equal to layerld and subimage index equal to subpicldx shall precede, in order of output, to the IRAP subimage and all its associated RADL subimages. Any image, with nuhjayerjd equal to a particular layerld value, that precedes, in decoding order, a GDR image with nuhjayerjd equal to layerld will precede, in output order, the GDR image and all its associated images. Any subimage, with nuh_layer_id equal to a particular layerld value and subimage index equal to a particular subpicldx value, that precedes, in decoding order, a GDR subimage with nuh_layer_id equal to layerld and subimage index equal to subpicldx will precede, in output order, to the GDR subimage and all its associated subimages. Any RASL image associated with a CRA image must precede any RADL Image associated with the CRA image in output order. Any RASL subimage associated with a CRA subimage shall precede any RADL subimage associated with the CRA subimage in output order. Any RASL image associated with a CRA image must follow, in output order, any IRAP image that precedes the CRA image in decoding order. - Any RASL subimage associated with a CRA subimage must follow, in output order, any IRAP subimage that precedes the CRA subimage in decoding order. - If field_seq_flag is equal to 0 and the current image, with nuhjayerjd equal to a particular layerld value, is an initial image associated with an IRAP image, it will precede, in decoding order, all non-initial images that are associated with the same image of IRAP. Otherwise, let picA and picB be the first and last initial images, in decoding order, associated with an IRAP image, respectively, there must be at least one non-initial image with nuhjayerjd equal to layerld preceding picA in order decoding order, and there will be no non-initial image with nuhjayerjd equal to layerld between picA and picB in decoding order. If field_seq_flag is equal to 0 and the current subimage, with nuh_layer_id equal to a particular layerld value and subimage index equal to a particular subpicldx value, is an initial subimage associated with an IRAP subimage, it will precede, in decoding order, all non-initial subimages that are associated with the same IRAP subimage. Otherwise, let subpicA and subpicB be the first and last initial subimages, in decoding order, associated with an IRAP subimage, respectively, there must be at least one non-initial subimage with nuh_layer_id equal to layerld subimage index equal to subpicldx preceding subpicA in decoding order, and there will be no non-initial image with nuh_layer_id subimage index equal to layerld subimage index equal to subpicldx between picA and picB in decoding order. 7.4.3.4 Semantics of the image parameter set mixed_nalu_typesjn_pic_flag equal to 1 specifies that each image that refers to the PPS has more than one NAL VCL unit and the NAL VCL units do not have the same nal_unitjype value, and the image is not an image of intra-random access point (IRAP). mixed_nalujypesjn_picjlag equal to 0 specifies that each image that refers to the PPS has one or more NAL VCL units and the NAL VCL units of each image that refers to the PPS have the same nal_unitjype value. When no_mixed_nalujypesjn_p¡c_constraintjlag is equal to 1, the value of mixed_nalujypesjn_picjlag will be equal to 0. [[For each slice with a nal_unitjype nalUnitTypeA value in the range of IDR_W_RADL to CRA_NUT, inclusive, in a picA image that also contains one or more slices with another nal_unitjype value (i.e., the value of mixed_nalujypesjn_picjlag for the picA image is equal to 1), the following applies: The slice will belong to a subimage subpicA for which the value of the corresponding subpicjreated_as_picjlag[ i ] is equal to 1. The slice will not belong to a subpicA subimage containing NAL VCL units with nal_unitjype not equal to nalUnitTypeA. If nalUnitTypeA is equal to CRA, for all subsequent PUs that follow the current image in the CLVS in decoding order and output order, neither RefPicList[0] nor RefPicList[1] of a cut in subpicA in those PUs will include any image before picA in decoding order on an active input. - Otherwise (i.e. nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS following the current image in decoding order, neither RefPicList[ 0 ] nor RefPicList[ 1 ] give a cut in subpicA in those PUs will include any image before picA in decoding order in an active input.]] NOTE 1- mixed_nalu_typesjn_pic_flag equal to 1 indicates that images referring to the PPS contain slices with different types of NAL units, for example, encoded images originating from a subpicture bitstream merging operation for which the encoders have which ensure a matching bitstream structure and greater alignment of the parameters of the original bitstreams. An example of such alignments is the following: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_in_pic_flag is equal to 1, an image referring to the PPS cannot have slices with nal_unit_type equal to IDR_W_RADL or IDR_N_LP. 7.4.3 .7 Image Header Structure Semantics recovery_poc_cnt specifies the recovery point for decoded images in output order. When the current image is a GDR image, the recoveryPointPocVal variable is derived as follows: recoveryPointPocVal = PícOrderCntVal + recovery_poc_cnt (81) If the current image is an image of the GDR that is associated with the PH, and there is an image picA that follows the current image of the GDR in decoding order in the CLVS that has PícOrderCntVal equal to recoveryPointPocValde the current image of the GDR plus the value of recovery_poc_cnt, the picA image is called the recovery point image. Otherwise, the first image in the output order that has PicOrderCntVal greater than recoveryPointPocVal [[the PicOrderCntVal of the current image plus the value of recovery_poc_cnt]] in the CLVS is referred to as the recovery point image. The recovery point image should not precede the current GDR image in decoding order. Images that are associated with the current GDR image and have PicOrderCntVal less than recoveryPointPocVal are referred to as the recovery images of the GDR image. The value of recovery_poc_cnt will be in the range 0 to MaxPicOrderCntLsb - 1, inclusive. [[When the current image is a GDR image, the RpPicOrderCntVal variable is derived as follows: RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt (81)]] NOTE 2 – When gdr_enabled_flag is equal to 1 and PicOrderCntVal of the current image is greater than or equal to recoveryPointPocVal [[RpPicOrderCntVal]] of the associated GDR image, the current and subsequent images decoded in output order are exact matches with the corresponding images produced when starting the decoding process from the previous IRAP image, when present, preceding the associated GDR image in decoding order. 8.3.2 Decoding process for building reference image lists It is a bitstream compliance requirement that the following restrictions apply: For each i equal to 0 or 1, num_ref_entries[ i ][ Rplsldx[ i ] ] shall not be less than NumRefldxActive[ i ]. The image referenced by each active entry in RefPicList[ 0 ] or RefPicList[ 1 ] will be present in the DPB and will have Temporalld less than or equal to that of the current image. The image referenced by each entry in RefPicList[ 0 ] or RefPicList[ 1 ] will not be the current image and will have non_reference_picture_flag equal to 0. An STRP entry in RefPicList[ 0 ] orRefPicList[ 1 ] of a slice of an image and an LTRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] of the same slice or a different slice of the same image will not refer to the same image. There may be no LTRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] for which the difference between the PicOrderCntVal of the current image and the PicOrderCntVal of the image referenced by the entry is greater than or equal to 224. Let setOfRefPics be the set of unique images referenced by all entries in RefPicList[ 0 ] that have the same nuhjayerjd as the current image and all entries in RefPicList[ 1 ] that have the same nuhjayerjd as the image current. The number of images in setOfRefPics shall be less than or equal to MaxDpbSize - 1, inclusive, where MaxDpbSize is as specified in clause A.4.2, and setOfRefPics shall be the same for all slices of an image. When the current slice has nal_unitjype equal to STSA_NUT, there will be no active entry in RefPicListj 0 ] or RefPicList[ 1 ] that has Temporalld equal to that of the current image and nuhjayerjd equal to that of the current image. When the current image is an image that follows, in decoding order, a STSA image that has Temporalld equal to that of the current image and nuhjayerjd equal to that of the current image, there will be no image that precedes the STSA image in decoding order, has Temporalld equal to that of the current image, and has nuhjayerjd equal to that of the current image that is included as an active entry in RefPicList[ 0 ] or RefPicListf 1 ]. When the current subimage, with Temporalld equal to a particular tld value, nuh_layer_id equal to a particular layerld value, and subimage index equal to a particular subpicldx value, is a subimage that follows, in decoding order, an STSA subimage with Temporalld equal to tld, nuh_layer_id equal to layerld, and subimage index equal to subpicldx, there will be no image with Temporalld Equal to tld and nuh_layer_ld equal to layerld that precedes the Image containing the STSA subimage in decoding order included as an input active in RefPícList[ 0] or RefPícList[ 1 ]. When the current image, with nuhjayerjd equal to a particular layerld value, is a CRA image, there will be no image referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes it, in output order or decode order, to any IRAP image with nuhjayerjd equal to layerld in decode order (when present). When the current subimage, with nuh_layer_id equal to a particular layerld value and subimage indices equal to a particular subpicldx value, is a CRA subimage, there will be no image referenced by an entry in RefPicList[ 0] or RefPicList [ 1 ] that precedes, in output order or decoding order, any image containing an IRAP subimage with nuh_layer_id equal to layerld and subimage indices equal to subpicldx in decoding order (when present). When the current image, with nuhjayerjd equal to a particular layerld value, is not a RASL image associated with a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or an image recovery image of GDR with NoOutputBeforeRecoveryFlag equal to 1 and nuhjayerjd equal to layerld, there will be no image referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that was generated by the decoding process to generate reference images not available. When the current subimage, with nuh_layer_id equal to a particular layerld value and subimage index equal to a particular subpicldx value, is not a RASL subimage associated with a CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a GDR of a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or a sub-image of a recovery image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerld, there will be no image referenced by an active entry in RefPicList[ 0]or RefPicList[ 1 ] that was generated by the decoding process to generate unavailable reference images. When the current image, with nuhjayerjd equal to a particular layerld value, is not a CRA image with NoOutputBeforeRecoveryFlag equal to 1, an image that precedes, in decoding order, the initial images associated with the same CRA image with NoOutputBeforeRecoveryFlag equal to 1, an initial image associated with a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a GDR image with NoOutputBeforeRecoveryFlag equal to 1, or a recovery image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1 and nuhjayerjd equal to layerld, there will be no no image referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that was generated by the decoding process to generate unavailable reference images. When the current subimage, with nuh_layer_id equal to a particular layerld value and subimage index equal to a particular subpicldx value, is not a CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a preceding subimage, in decoding order , to initial subimages associated with the same CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, an initial subimage associated with a CRA subimage of a CRA image with NoOutputBeforeRecoveryFlag equal to 1, a GDR subimage of an image of GDR with NoOutputBeforeRecoveryFlag equal to 1, or a subimage of a recovery image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerld, there will be no image referenced by an entry in RefPicList[ Note RefPicList [ 1 ] which was generated by the decoding process to generate unavailable reference images. When the current image is associated with an IRAP image and follows the IRAP subimage in output order, there will be no image referenced by an active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes the associated IRAP image in output order or decoding order. When the current subimage is associated with an IRAP subimage and follows the IRAP subimage in output order, there will be no image referenced by an active entry in RefPicList[ Note RefPicListJ 1 ] that precedes the image that contains the associated IRAP subimage in output order or decoding order. When the current image is associated with an IRAP image, it follows the IRAP image in output order, and follows, in both decoding order and output order, the initial images associated with the same IRAP image, if If there is, there will be no image referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] that precedes the associated IRAP image in output order or decoding order. When the current subpicture is associated with an IRAP subpicture, it follows the IRAP subpicture in output order, and follows, in both decoding order and output order, the initial subpictures associated with the same IRAP subpicture, if If there is, there will be no image referenced by an entry in RefPicListJ Note RefPicListJ 1 ] that precedes the image containing the associated IRAP subimage in output order or decoding order. When the current image is a RADL image, there will be no active entry in RefPicList[ 0 ] or RefPicList[ 1 ] that is any of the following: An image of RASL An image preceding the associated IRAP image in decoding order When the current subimage is a RADL subimage, there will be no active entry in RefPicListJ Eye RefPicListJ 1 ] that is any of the following: An image containing a RASL subimage An image preceding the image containing the associated IRAP subimage in decoding order The image referenced by each ILRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] of a slice of the current image will be in the same AU as the current image. The image referenced by each ILRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] of a slice of the current image will be present in the DPB and will have nuhjayerjd less than that of the current image. Each ILRP entry in RefPicList[ 0 ] or RefPicList[ 1 ] of a slice will be an active entry. Figure 5 is a block diagram showing an exemplary video processing system 1900 in which different techniques disclosed herein can be implemented. Different implementations may include some or all of the components of system 1900. System 1900 may include input 1902 for receiving video content. Video content may be received in an unprocessed or uncompressed format, for example, 8-bit or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interface include wired interfaces such as Ethernet, passive optical network (PON), etc. and wireless interfaces such as Wi-Fi or cellular interfaces. The system 1900 may include an encryption component 1904 that may implement the different encryption methods described herein. The encoding component 1904 may reduce the average bit rate of the video from the input 1902 to the output of the encoding component 1904 to produce an encoded representation of the video. Therefore, encoding techniques are sometimes referred to as video compression techniques or video transcoding. The output of the encoding component 1904 may be stored, or transmitted via connected communication, as represented by the component 1906. The stored or communicated bitstream (or encoded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values ​​or display video that is sent to a display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “encoding” operations or tools, it will be appreciated that the encoding tools or operations are used in an encoder and the corresponding decoding tools or operations that invert the results of the Encoding will be carried out by means of a decoder. Examples of a peripheral bus interface or a display interface may include universal serial bus (USB) or high definition multimedia interface (HDMI) or display port, etc. Storage interfaces include Serial Advanced Technology Accessory Interface (SATA), Peripheral Component Interconnect (PCI), Integrated Electronic Device Interface (IDE) and the like. The techniques described herein can be incorporated into different electronic devices such as mobile phones, laptops, smartphones or other devices that are capable of carrying out digital data processing and / or video display. Figure 6 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more of the methods described herein. The device 3600 can be incorporated into a smartphone, electronic tablet, computer, IoT receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 may be configured to implement one or more of the methods described herein. . Memory (memories) 3604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 can be used to implement, in hardware circuits, some of the techniques described herein. Figure 8 is a block diagram illustrating an example video coding system 100 that may use the techniques of this disclosure. As shown in Figure 8, the video encoder system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 which may be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116. The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider and / or a computer graphics system for generating video data, or a combination from these sources. The video data may comprise one or more images. The video encoder 114 encodes the video data from the video source 112 to generate a bit stream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include encoded images and associated data. The encoded image is an encoded representation of an image. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. The l / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120. The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The l / O interface 126 may include a receiver and / or a modem. The l / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the target device 120, or may be external to the target device 120 that is configured to interact with an external display device. The video encoder 114 and video decoder 124 may operate in accordance with a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and others. current and / or additional standards. Figure 9 is a block diagram illustrating an example of video encoder 200, which may be video encoder 114 in the system 100 illustrated in Figure 8. The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In the example of Figure 9, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure. The functional components of the video encoder 200 may include a partition unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and a of intraprediction 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213 and a entropic coding 214. In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference image is an image where the current video block is located. In addition, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are depicted in the example of Figure 9 separately for purposes of explanation. The partition unit 201 may partition an image into one or more video blocks. The video encoder 200 and video decoder 300 can support various video block sizes. The mode selection unit 203 may select one of the coding modes, intra or inter, for example, based on the error results, and provide the resulting intra- or inter-coded block to a residual generation unit 207 to generate data from residual blocks and to a reconstruction unit 212 to reconstruct the encoded block for use as a reference image. In some example, the mode selection unit 203 may select a Combined Inter-Intra Prediction (CIIP) mode in which the prediction is based on an inter-prediction signal and an intra-prediction signal. The mode selection unit 203 may also select a resolution for a motion vector (e.g., sub-pixel or integer pixel precision) for the block in the case of interprediction. To perform interprediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames of the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded image samples of the buffer 213 other than the image associated with the current video block. The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for a current video block, for example, depending on whether the current video block is in an I-slice, a P-slice, or a b. In some examples, the motion estimation unit 204 may perform one-way prediction for the current video block, and the motion estimation unit 204 may search for reference images from list 0 or list 1 for a reference video block. for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference image in list 0 or list 1 containing the reference video block and a motion vector indicating a spatial offset between the video block. current and the reference video block. The motion estimation unit 204 may generate the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current block based on the reference video block indicated by the motion information of the current video block. In other examples, the motion estimation unit 204 may perform bidirectional prediction for the current video block, the motion estimation unit 204 may search the reference images in list 0 for a reference video block for the current video and you can also search the reference images in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indices indicating the reference images in list 0 and list 1 containing the reference video blocks and motion vectors indicating spatial displacements between the reference video blocks. and the current video block. The motion estimation unit 204 may generate the reference indices and motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate the predicted video block of the current video block based on the reference video blocks indicated by the motion information of the current video block. In some examples, the motion estimation unit 204 may generate a complete set of motion information for decoding processing of a decoder. In some examples, the motion estimation unit 204 may not generate a complete set of motion information for the current video. Rather, the motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block. In one example, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as the another video block. In another example, the motion estimation unit 204 may identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). Motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the difference of the motion vector to determine the motion vector of the current video block. As discussed above, the video encoder 200 may predictively indicate the motion vector. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and fusion mode signaling. The intra-prediction unit 206 may perform intra-prediction on the current video block. When the intra-prediction unit 206 performs intra-prediction on the current video block, the intra-prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same image. The prediction data for the current video block may include a predicted video block and various syntax elements. The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block. In other examples, there may be no residual data for the current video block, for example, in a skip mode, and the residual generation unit 207 may not perform the subtraction operation. The transformation processing unit 208 may generate one or more transformation coefficient video blocks for the current video block by applying one or more transformations to a residual video block associated with the current video block. After the transformation processing unit 208 generates a transformation coefficient video block associated with the current video block, the quantization unit 209 may quantize the transformation coefficient video block associated with the current video block with based on one or more quantization parameter (QP) values ​​associated with the current video block. The inverse quantization unit 210 and the inverse transformation unit 211 may apply inverse quantization and inverse transformations to the transformation coefficient video block, respectively, to reconstruct a residual video block of the transformation coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213. After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block. The entropic coding unit 214 may receive data from other functional components of the video encoder 200. When the entropic coding unit 214 receives the data, the entropic coding unit 214 may perform one or more entropic coding operations to generate data encoded by entropy and generate a bitstream that includes the entropy-encoded data. Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In an example, when the video processing tool or mode is enabled, the encoder will use or implement the tool or mode in processing a video block, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, a conversion of the video block to the bitstream (or bitstream representation) of the video will use the video processing tool or mode when enabled based on the decision or determination. In another example, when the video processing tool or mode is enabled, the decoder will process the bitstream with the knowledge that the bitstream has been modified based on the video processing tool or mode. That is, a conversion of the video bitstream to the video block will be performed using the video processing tool or mode that was enabled based on the decision or determination. Figure 10 is a block diagram illustrating an example of video decoder 300, which may be video decoder 114 in the system 100 illustrated in Figure 8. The video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In the example of Figure 10, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure. In the example of Figure 10, the video decoder 300 includes an entropic decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305 and a reconstruction unit 306 and a buffer 307. The video decoder 300 may, in some examples, perform a decoding pass generally reciprocal to the encoding pass described with respect to the video encoder 200 (FIG. 9). The entropic decoding unit 301 may recover an encoded bit stream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, and motion vector precision. motion, reference image list indexes and other motion information. The motion compensation unit 302 may, for example, determine this information when performing the blending mode and AMVP. The motion compensation unit 302 may produce motion compensated blocks, possibly performing interpolation based on interpolation filters. Identifiers for interpolation filters to be used with subpixel precision can be included in the syntax elements. The motion compensation unit 302 may use interpolation filters as used by the video encoder 200 during encoding of the video block to calculate interpolated values ​​for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 according to the received syntax information and use the interpolation filters to produce predictive blocks. The motion compensation unit 302 may use some of the syntax information to determine the sizes of the blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of an image is divided. of the encoded video sequence, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each intercoded block, and other information for decoding the encoded video sequence. The intra-prediction unit 303 may use intra-prediction modes, e.g., received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inversely quantizes, that is, dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropic decoding unit 301. The inverse transformation unit 305 applies an inverse transformation. The reconstruction unit 306 may add the residual blocks with the corresponding prediction blocks generated by the motion compensation unit 302 or the intra-prediction unit 303 to form decoded blocks. If desired, a deblocking filter can also be applied to filter the decoded blocks to remove blocking artifacts. The decoded video blocks are stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra-prediction and also produces decoded video for presentation on a display device. Below is a list of solutions preferred by some modalities. The following solutions show exemplary modalities of the techniques addressed in the previous section (for example, point 1). 1. A video processing method (e.g., method 700 shown in Figure 7), comprising: performing (702) a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule specifying that said one or more images comprising one or more subimages are included in the encoded representation according to network abstraction layer (NAL) units, wherein a NAL type unit is indicated in the coded representation that includes a coded slice of a particular type of image or a coded slice of a particular type of a sub-image. The following solutions show exemplary modalities of the techniques discussed in the previous section (for example, point 2). 2. A video processing method, comprising: performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule that specifies that two neighboring subimages with different network abstraction layer unit types will have the same subimage indication that is treated as an image flag. The following solutions show exemplary modalities of techniques discussed in the previous section (for example, points 4, 5, 6, 7, 9, 1, 11, 12). 3. A video processing method, comprising: performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule that defines an order of a first subimage type and a second subimage type, where the first subimage is a final subimage or an initial subimage or a random access skipped initial subimage (RASL) type and the second subimage is of type RASL or a random access decodable initial subpicture (RADL) type or an instantaneous decoding update (IDR) type or a gradual decoding update (GDR) type. 4. The method according to solution 3, wherein the rule specifies that the final subimage follows an associated intra-random hotspot or GDR subimage in an output order. 5. The method according to solution 3, where the rule specifies that when an image is an initial image of an intra-random hotspot image, the value of nal_unit_type for all network abstraction layer units in the image is equal to RADL_NUT or RASL NUT. 6. The method according to solution 3, wherein the rule specifies that a given subimage that is an initial subimage of an IRAP subimage must also be a RADL or RASL subimage. 7. The method according to solution 3, wherein the rule specifies that a given subimage that is a RASL subimage is prohibited from being associated with an IDR subimage. 8. The method according to solution 3, wherein the rule specifies that a given subimage having the same layer id and subimage index as an IRAP subimage must precede, in an output order, the subimage of IRAP and all associated RADL subimages thereof. 9. The method according to solution 3, wherein the rule specifies that a given subimage having the same layer id and subimage index as a GDR subimage must precede, in an output order, the subimage of GDR and all associated RADL subimages thereof. 10. The method according to solution 3, wherein the rule specifies that a given subimage that is a RASL subimage associated with a CRA subimage precedes in an output order all RADL subimages associated with the CRA subimage . 11. The method according to solution 3, wherein the rule specifies that a given subimage that is a RASL subimage associated with a CRA subimage precedes in an output order all IRAP subimages associated with the CRA subimage . 12. The method according to solution 3, where the rule specifies that a given subimage is an initial subimage with an IRAP subimage, then the given subimage precedes, in a decoding order, all non-initial subimages associated with the image of IRAP. The following solutions show exemplary modalities of the techniques discussed in the previous section (e.g. points 8, 14, 15). 13. A video processing method, comprising: performing a conversion between a video comprising one or more images comprising one or more subimages and an encoded representation of the video, wherein the encoded representation conforms to a format rule that defines a condition under which a first type of subimage is allowed or prohibited from occurring with a second type of subimage. 14. The method according to solution 13, wherein the rule specifies that, in case there is an IDR subimage of network abstraction layer type IDR_N_LP, then the encoded representation is prohibited from having a RADP subimage. 15. The method according to solution 13, wherein the rule prohibits including an image in a reference list of an image comprising a staged temporal sublayer access (STSA) subimage such that the image precedes a image that comprises the STSA subimage. 16. The method according to solution 13, wherein the rule prohibits including an image in a reference list of an image comprising an intra-random access point (IRAP) subimage such that the image precedes an image that understand the IRAP subimage. 17. The method according to any of solutions 1 to 16, wherein the conversion comprises encoding the video into the encoded representation. 18. The method according to any of solutions 1 to 16, wherein the conversion comprises decoding the encoded representation to generate pixel values ​​of the video. 19. A video decoding apparatus comprising a processor configured to implement a method mentioned in one or more of solutions 1 to 18. 20. A video encoding apparatus comprising a processor configured to implement a method mentioned in one or more of solutions 1 to 18. 21. A computer program product that has computer code stored therein, the code, when executed by a processor, causes the processor to implement a method mentioned in any of solutions 1 to 18. 22. A method, apparatus or system described herein. In the solutions described herein, an encoder may conform to the format rule by producing an encoded representation according to the format rule. In the solutions described herein, a decoder may use the format rule to parse syntax elements in the encoded representation with knowledge of the presence and absence of syntax elements according to the format rule to produce decoded video. Figure 11 is a flow chart for an example video processing method 1100. Operation 1102 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies a syntax of units of network abstraction layer (NAL) in the bitstream, and wherein the format rule specifies that a NAL unit of a video coding layer (VCL) NAL unit type includes content associated with a type particular image or a particular type of subimage. In some embodiments of method 1100, the content of the NAL unit of the NAL unit type VCL indicates that an encoded slice is associated with a clean random access image or a clean random access subimage. In some embodiments of method 1100, the clean random access subimage is an intra-random access point subimage for which each NAL VCL unit has a clean random access type. In some embodiments of method 1100, the content of the NAL unit of the NAL unit type VCL indicates that an encoded slice is associated with an associated rolling decoding update image or an associated rolling decoding update subimage. In some embodiments of method 1100, the associated gradual decoding update subimage is a previous gradual decoding update subimage in a decoding order with an identifier of a layer to which a NAL VCL unit or an identifier of a layer belongs. to which applies a non-VCL NAL unit equal to a first particular value and with a second particular value of sub-picture index, and wherein between the previous gradual decoding update sub-picture and a particular sub-picture with the first particular value of identifier and the second particular subimage index value in the decoding order, there is no intra-random hotspot subimage with the first particular identifier value and the second particular subimage index value. In some embodiments of method 1100, the content of the NAL unit of the NAL unit type VCL indicates that an encoded slice is associated with an associated intra-random hotspot image or an associated intra-random hotspot subimage. In some embodiments of method 1100, the associated intra-random access point sub-image is a previous intra-random access point sub-image in a decoding order with an identifier of a layer to which a NAL VCL unit or an identifier of a layer belongs. to which applies a non-VCL NAL unit equal to a first particular value and with a second particular value of sub-image index, and where between the previous intra-random access point sub-image and a particular sub-image with the first particular value of identifier and the second particular sub-picture index value in the decoding order, there is no gradual decoding update sub-picture with the first particular identifier value and the second particular sub-picture index value. In some embodiments of method 1100, the contents of the NAL unit of the NAL unit type VCL indicate that an encoded slice is associated with an instant decoding update image or an instant decoding update subimage. In some embodiments of method 1100, the instantaneous decoding update subimage is an intra-random access point subimage for which each NAL VCL unit has an instantaneous decoding update type. In some embodiments of method 1100, the contents of the NAL unit of NAL unit type VCL indicate that an encoded slice is associated with an initial image or an initial subimage. In some embodiments of method 1100, the initial subimage is a subimage that precedes the associated intra-random hotspot subimage in output order. In some embodiments of method 1100, the content of the NAL unit of NAL unit type VCL indicates that an encoded slice is associated with a random access decodable initial image or a random access decodable initial subimage. In some embodiments of method 1100, the initial random access decodable subimage is a subimage for which each NAL VCL unit has an initial random access decodable type. In some embodiments of method 1100, the content of the NAL unit of the NAL unit type VCL indicates that an encoded slice is associated with a random access skipped initial image or a random access skipped initial subimage. In some embodiments of method 1100, the random access skipped initial subimage is a subimage for which each NAL VCL unit has a random access skipped initial type. In some embodiments of method 1100, the content of the NAL unit of the NAL unit type VCL indicates that an encoded slice is associated with a staged temporal sublayer access image or a staged temporal sublayer access subimage. In some embodiments of method 1100, the staged temporal sublayer access subimage is a subimage for which each NAL VCL unit has a staged temporal sublayer access type. In some embodiments of method 1100, the content of the NAL unit of NAL unit type VCL indicates that an encoded slice is associated with a final image or a final subimage. In some embodiments of method 1100, the final subimage is a subimage for which each NAL VCL unit has a final type. Figure 12 is a flow chart for an example video processing method 1200. Operation 1202 includes performing a conversion between a video comprising an image comprising a subimage and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that the subimage is a type of subimage random access in response to the subimage being an initial subimage of an intra-random hotspot subimage. In some embodiments of method 1200, the subimage random access type is an initial random access decodable subimage. In some embodiments of method 1200, the subimage random access type is an initial random access skipped subimage. Figure 13 is a flow chart for an example video processing method 1300. Operation 1302 includes performing a conversion between a video comprising an image comprising a subimage and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that a or more random access skipped initial subimages are absent from the bitstream in response to said one or more random access skipped initial subimages that are associated with an instantaneous decode update subimage. Figure 14 is a flow chart for an example video processing method 1400. Operation 1402 includes performing a conversion between a video comprising an image comprising a subimage and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that one or more access decodable initial subimages random access are absent from the bitstream in response to said one or more initial random access decodable subimages that are associated with an instantaneous decoding update subimage that has a type of network abstraction layer (NAL) unit that indicates that the instant decode update subimage is not associated with an initial image. Figure 15 is a flow chart for an example video processing method 1500. Operation 1502 includes performing a conversion between a video comprising an image comprising two neighboring subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that the two neighboring subimages with different types of network abstraction layer (NAL) units have syntax elements with the same first value that indicates whether each of the two neighboring subimages in an encoded layer video sequence is treated as one image. In some embodiments of method 1500, the format rule specifies that the syntax elements of the two neighboring subimages indicate that each of the two neighboring subimages in the encoded layer video sequence is treated as one image. Figure 16 is a flow chart for an example video processing method 1600. Operation 1602 includes performing a conversion between a video comprising an image comprising two adjacent subimages and a bitstream of the video, wherein the format rule specifies that the two adjacent subimages include a first adjacent subimage with a first subimage index. and a second neighboring subimage with a second subimage index, and wherein the format rule specifies that the two neighboring subimages have the same type of network abstraction layer (NAL) units in response to a first syntax element associated with the first subimage index indicating that the first neighboring subimage is not treated as a subimage or a second syntax element associated with the second subimage index indicating that the second neighboring subimage is not treated as an image. In some embodiments of the method 1600, the image comprises a plurality of subimages that include the two neighboring subimages, and wherein the format rule specifies that the plurality of subimages have the same type of NAL units in response to a subimage of the plurality. of subimages that have a syntax element that indicates that the subimage is not treated like the image. In some embodiments of the method 1600, the image comprises a plurality of subimages that include the two neighboring subimages, and wherein the format rule specifies that a syntax element indicates that each image of the video that references a set of image parameters (PPS) has a plurality of video coding layer (VCL) NAL units that do not have a same type of NAL VCL unit in response to the plurality of subimages that have corresponding syntax elements indicating that each of the plurality of subimages in the CLVS are treated as the image. Figure 17 is a flow chart for an example video processing method 1700. Operation 1702 includes performing a conversion between a video comprising images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that an image is allowed to include more of two different types of video coding layer (VCL) network abstraction layer (NAL) units in response to a syntax element that indicates that each image of the video that references a set of image parameters (PPS ) has a plurality of NAL VCL units that do not have the same NAL VCL unit type. Figure 18 is a flow chart for an example video processing method 1800. Operation 1802 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a final subimage that is associated with an intra-random access point sub-image or a gradual decoding update sub-image follows the intra-random access point sub-image or the gradual decoding update sub-image in an order. In some embodiments of method 1800, the order is an output order. Figure 19 is a flow chart for an example video processing method 1900. Operation 1902 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a subimage precedes in a first order to an intra-random hotspot subimage and one or more random access decodable initial subimages associated with the intrarandom hotspot subimage in response to: (1) that the subimage precedes the intrarandom hotspot subimage in a second order, (2) that the subimage and the intra-random hotspot subimage have the same first value for a layer to which a network abstraction layer (NAL) unit of the subimage and the hotspot subimage belong. intra-random access (3) that the sub-image and the intra-random access point sub-image have the same second value of a sub-image index. In some embodiments of method 1900, the first order is an output order. In some embodiments of method 1900, the second order is a decoding order. Figure 20 is a flow chart for an example video processing method 2000. Operation 2002 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that an initial subimage omitted random access subimage associated with a clean random access subimage precedes in order one or more initial decodable random access subimages associated with the clean random access subimage. In some embodiments of method 2000, the order is an output order. Figure 21 is a flow chart for an example video processing method 2100. Operation 2102 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that an initial subimage omitted random access point associated with a clean random access subimage follows in a first order one or more intra-random access point subimages that precede the clean random access subimage in a second order. In some embodiments of method 2100, the first order is an output order. In some embodiments of method 2100, the second order is a decoding order. Figure 22 is a flow chart for an example video processing method 2200. Operation 2202 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that a current subimage precedes in a decoding order to one or more non-initial subimages associated with an intra-random hotspot subimage in response to: (1) that a syntax element indicates that an encoded layer video sequence transmits images representing frames, and ( 2) that the current subimage is an initial subimage associated with the intra-random hotspot subimage. In some embodiments of method 2200, in response to (1) the syntax element indicating that the encoded layer video stream transmits images representing fields, and (2) the current subimage is not the initial subimage, the format rule specifies: a presence of at most one non-initial sub-image preceding a first initial sub-image associated with the intra-random hotspot sub-image in the decoding order, and an absence of a non-initial image between the first initial sub-image and a last sub-image initial associated with the intra-random hotspot subimage in the decoding order, where the current subimage, at most one non-initial subimage, and the non-initial image have the same first value for a layer to which a unit of network abstraction layer (NAL) of the current subimage, at most one non-initial subimage, and the non-initial image, and where the current subimage, at most one non-initial subimage, and the non-initial image have the same second value of a subimage index. Figure 23 is a flow chart for an example video processing method 2300. Operation 2302 includes performing a conversion between a video comprising one or more images comprising one or more subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that one or more types network abstraction layer (NAL) unit number for all video coding layer (VCL) NAL units in an image include RADL_NUT or RASL_NUT in response to the image being a seed image of an intra-random hotspot image. In some embodiments of method 2300, the format rule specifies that an image variable is set equal to a value of an image output flag in response to: (1) said one or more NAL unit types for all NAL VCL units in the image include RADL_NUT and RASL_NUT, and (2) a layer that includes the image which is an output layer. Figure 24 is a flow chart for an example video processing method 2400. Operation 2402 includes performing a conversion between a video comprising one or more images comprising a plurality of subimages and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that said at least one subpicture precedes in a first order a gradual decoding update subpicture and one or more subpictures associated with the gradual decoding update subpicture in response to: (1) said at least one subpicture preceding the gradual decoding update subpicture in a second order, (2) that said at least one subimage and the gradual decoding update subimage have the same first value for a layer to which a network abstraction layer (NAL) unit of said at least one belongs subpicture and the gradual decoding update subpicture, and (3) that said at least one subpicture and the gradual decoding update image have the same second value of a subpicture index. In some embodiments of method 2400, the first order is an output order. In some embodiments of method 2400, the second order is a decoding order. Figure 25 is a flow chart for an example video processing method 2500. Operation 2502 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits a active entry in a list of reference images of the current slice includes a first image that precedes in a decoding order a second image that includes a stepwise temporal sublayer access subimage in response to: (a) that the first image has a same temporal identifier and a same layer identifier of a network abstraction layer (NAL) unit as that of the current sub-image, and (b) that the current sub-image follows in encoding order the sub-layer access sub-image staged temporal subimage, and (c) that the current subimage and the staged temporal sublayer access subimage have the same temporal identifier, the same layer identifier, and the same subimage index. In some embodiments of method 2500, the list of reference images includes a list of List 0 reference images. In some embodiments of method 2500, the list of reference images includes a list of List 1 reference images. Figure 26 is a flow chart for an example video processing method 2600. Operation 2602 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits a active entry in a list of reference images of the current slice includes a first image that is generated by a decoding process to generate unavailable reference images in response to the current subimage not being of a particular type of subimage. In some embodiments of method 2600, the current subimage is not an initial skipped random access subimage associated with a clean random access subimage of a clean random access image with a value of a flag indicating that no output before recovery is equal to 1. In some embodiments of method 2600, the current subpicture is not a gradual decoding update subpicture of a gradual decoding update image with a flag value indicating that no output before recovery is equal to 1. In some embodiments of method 2600, the current subimage is not a subimage of a recovery image of a gradual decoding update image with a value of a flag indicating that no output before recovery is equal to 1 and that it has a same layer identifier of a network abstraction layer (NAL) unit as that of the current subimage. In some embodiments of method 2600, the list of reference images includes a list of List 0 reference images. In some embodiments of method 2600, the list of reference images includes a list of List 1 reference images. Figure 27 is a flow chart for an example video processing method 2700. Operation 2702 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits a entry in a list of reference images of the current slice includes a first image that is generated by a decoding process to generate unavailable reference images in response to the current subimage not being of a particular type of subimage. In some embodiments of the method 2700, the current image is not a clean random access subimage of a clean random access image with a value of a flag indicating that no output before recovery is equal to 1. In some embodiments of the method 2700, the current subimage is not a subimage that precedes, in a decoding order, one or more initial subimages associated with the clean random access subimage of the clean random access image with a value of a flag indicating that no output before recovery is equal to 1. In some embodiments of method 2700, the current subimage is not an initial subimage associated with the clean random access subimage of the clean random access image with a value of a flag indicating that no output before recovery is equal to 1. In some embodiments of method 2700, the current subimage is not a gradual decoding update subimage of a gradual decoding update image with a flag value indicating that no output before recovery is equal to 1. In some embodiments of method 2700, the current subimage is not a subimage of a recovery image of a gradual decoding update image with a value of a flag indicating that no output before recovery is equal to 1 and that it has a same layer identifier of a network abstraction layer (NAL) unit as that of the current subimage. In some embodiments of method 2700, the list of reference images includes a list of List 0 reference images. In some embodiments of method 2700, the list of reference images includes a list of List 1 reference images. Figure 28 is a flow chart for an example video processing method 2800. Operation 2802 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits a entry in a list of reference images of the current slice includes a first image that precedes in a first order or a second order the current image in response to: (a) that the first image includes a preceding intra-random hotspot subimage that precedes in the second order the current subimage, (b) that the preceding intra-random access point subimage has the same layer identifier of a network abstraction layer (NAL) unit and the same subimage index as that of the current subimage, and (c) that the current subimage is a clean random access subimage. In some embodiments of method 2800, the first order includes an output order. In some embodiments of method 2800, the second order includes a decoding order. In some embodiments of method 2800, the list of reference images includes a list of List 0 reference images. In some embodiments of method 2800, the list of reference images includes a list of List 1 reference images. Figure 29 is a flow chart for an example video processing method 2900. Operation 2902 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits a active entry in a list of reference images of the current slice includes a first image that precedes in a first order or a second order the current image in response to: (a) that the current subimage is associated with a hotspot subimage intra-random, (b) that the current sub-image follows the intra-random hotspot sub-image in the first order. In some embodiments of method 2900, the first order includes an output order. In some embodiments of method 2900, the second order includes a decoding order. In some embodiments of method 2900, the list of reference images includes a list of List 0 reference images. In some embodiments of method 2900, the list of reference images includes a list of List 1 reference images. Figure 30 is a flow chart for an example video processing method 3000. Operation 3002 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that prohibits a entry in a list of reference images of the current slice includes a first image that precedes in a first order or the second order the current image that includes an intra-random hotspot subimage associated with the current subimage in response to: (a) that the current subimage follows the intra-random hotspot subimage in the first order, (b) that the current subimage follows one or more of the initial subimages associated with the IRAP subimage in the first order and the second order. In some embodiments of method 3000, the first order includes an output order. In some embodiments of method 3000, the second order includes a decoding order. In some embodiments of method 3000, the list of reference images includes a list of List 0 reference images. In some embodiments of method 3000, the list of reference images includes a list of List 1 reference images. Figure 31 is a flow chart for an example video processing method 3100. Operation 3102 includes performing a conversion between a video comprising a current image comprising a current sub-image comprising a current slice and a bitstream of the video, wherein the bitstream conforms to a format rule that specifies that in In response to the current subimage being a decodable random access initial subimage, a list of reference images of the current slice excludes an active entry for any or more of: a first image that includes a random access omitted initial subimage, and a second image preceding a third image that includes an associated intra-random hotspot subimage in a decoding order. In some embodiments of method 3100, the list of reference images includes a list of List 0 reference images. In some embodiments of method 3100, the list of reference images includes a list of List 1 reference images. As used herein, the term “video processing” may refer to video encoding, video decoding, video compression, or video decompression. For example, video compression algorithms can be applied during the conversion of the pixel representation of a video to a corresponding bitstream representation or vice versa. The bitstream representation of a current video block may, for example, correspond to bits that are co-located or scattered at different locations within the bitstream, as defined in the syntax. For example, a macroblock can be encoded in terms of transformed and encoded error residual values ​​and also use bits in headers and other fields of the bitstream. Additionally, during conversion, a decoder can analyze a bitstream with the knowledge that some fields may be present, or absent, depending on the determination, as described in the solutions above. Similarly, an encoder may determine that certain syntax fields are or are not included and generates the encoded representation accordingly by including or excluding the syntax fields from the encoded representation. The disclosed or different solutions, examples, modalities, modules and functional operations described herein may be implemented in digital electronic circuits, or in computer software, firmware or hardware, including the disclosed structures. herein and their structural equivalents, or in combinations of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, that is, one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. . The term “data processing apparatus” encompasses all apparatus, devices and machines for processing data, including, but not limited to, a programmatic processor, a computer, or multiple processors or computers. The apparatus may include, in addition to the hardware, code for creating an execution environment for the computer program in question, for example, code that constitutes a processor firmware, a protocol stack, a database management system , an operating system or a combination of one or more of them. A propagated signal is a signal that is artificially generated, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiving apparatus. A computer program (also known as a program, software, sequence, software application, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be developed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that contains other programs or data (for example, one or more stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (for example (for example, files that store one or more modules, applets, or chunks of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected through a communication network. The processes and logical flows described in this document may be carried out by means of one or more programmable processors that execute one or more computer programs to carry out the functions by operating input data and generating output. The logic processes and flows may also be carried out by, and the apparatus may also be implemented as, special purpose logic circuits, for example, a field-programmable logic gate array (FPGA) or a application-specific integrated circuit (ASIC). Processors suitable for the execution of a computer program include, but are not limited to, general and special purpose microprocessors, and one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from read-only memory or random access memory or amdas. The essential elements of a computer are a processor to carry out instructions and one or more memory devices to store instructions and data. Generally, a computer will also include, or be coupled

Claims

1. A method for video processing comprising: performing a conversion between a video comprising one or more images comprising one or more sub-images and a bitstream of the video, wherein the bitstream conforms to a format rule specifying a syntax of network abstraction layer (NAL) units in the bitstream, and wherein the format rule specifies that an NAL unit of a video coding layer (VCL) NAL unit type includes content associated with a particular image type or a particular sub-image type.

2. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with a clean random access image or a clean random access sub-image.

3. The method according to claim 2, wherein the clean random access subimage is an intra-random access point subimage for which each NAL VCL unit has a clean random access type.

4. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with an associated graded decoding update image or an associated graded decoding update sub-image.

5. The method according to claim 4, wherein the associated gradual decoding update subimage is a previous gradual decoding update subimage in a decoding order with an identifier of a layer to which a VCL NAL unit belongs or an identifier of a layer to which a non-VCL NAL unit applies equal to a particular first value and with a particular second value of subimage index, and wherein between the previous gradual decoding update subimage and a particular subimage with the particular first identifier value and the particular second value of subimage index in the decoding order, there is no intra-random hotspot subimage with the particular first identifier value and the particular second value of subimage index.

6. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with an associated intra-random access point image or an associated intra-random access point sub-image.

7. The method according to claim 6, wherein the associated intra-random hotspot subimage is a previous intra-random hotspot subimage in a decoding order having an identifier of a layer to which a VCL NAL unit belongs or an identifier of a layer to which a non-VCL NAL unit applies equal to a particular first value and having a particular second value of subimage index, and wherein between the previous intra-random hotspot subimage and a particular subimage having the particular first identifier value and the particular second value of subimage index in the decoding order, there is no stepped-up decoding update subimage having the particular first identifier value and the particular second value of subimage index.

8. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with an instantaneous decoding update image or an instantaneous decoding update sub-image.

9. The method according to claim 8, wherein the instant decoding update subimage is an intra-random hotspot subimage for which each NAL VCL unit has an instant decoding update type.

10. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with an initial image or an initial sub-image.

11. The method according to claim 10, wherein the initial subimage is a subimage that precedes the associated intra-random hotspot subimage in output order.

12. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with a random access decodable initial image or a random access decodable initial sub-image.

13. The method according to claim 12, wherein the random-access decodable initial subimage is a subimage for which each NAL VCL unit has a random-access decodable initial type.

14. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with a random access skipped initial image or a random access skipped initial sub-image.

15. The method according to claim 14, wherein the initial omitted random access subimage is a subimage for which each NAL VCL unit has an initial omitted random access type.

16. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with a staged temporal sublayer access image or a staged temporal sublayer access subimage.

17. The method according to claim 16, wherein the staged temporal sublayer access subimage is a subimage for which each NAL VCL unit has a staged temporal sublayer access type.

18. The method according to claim 1, wherein the content of the NAL unit of the VCL type NAL unit indicates that a coded slice is associated with a final image or a final sub-image.

19. The method according to claim 18, wherein the final subimage is a subimage for which each NAL VCL unit has a final type.

20. A video processing method comprising: performing a conversion between a video comprising an image comprising a sub-image and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that the sub-image is a sub-image random access type in response to the sub-image being an initial sub-image of an intra-random access point sub-image.

21. The method according to claim 20, wherein the subimage random access type is a random access decodable initial subimage.

22. The method according to claim 20, wherein the subimage random access type is an omitted random access initial subimage.

23. A method for video processing comprising: performing a conversion between a video comprising an image comprising a sub-image and a bitstream of the video, wherein the bitstream conforms to a format rule, and wherein the format rule specifies that one or more random-access skipped initial sub-images are absent from the bitstream in response to said one or more random-access skipped initial sub-images being associated with an instant decoding update sub-image.

24. A method for video processing comprising: performing a conversion between a video comprising an image comprising a sub-image and a bitstream of the video, wherein the bitstream conforms to a format rule specifying that one or more random-access decodable initial sub-images are absent from the bitstream in response to such one or more random-access decodable initial sub-images being associated with an instant decoding update sub-image having a network abstraction layer (NAL) unit type indicating that the instant decoding update sub-image is not associated with an initial image.

25. The method according to any of claims 1 to 24, wherein performing the conversion comprises encoding the video in the bitstream representation.

26. The method according to any one of claims 1 to 24, wherein performing the conversion comprises generating the video bitstream, and the method further comprises storing the bitstream on a non-transient, computer-readable recording medium.

27. The method according to any of claims 1 to 24, wherein performing the conversion comprises decoding the video from the bitstream.

28. A video decoding apparatus comprising a processor configured to implement a method mentioned in one or more of claims 1 to 27.

29. A video encoding apparatus comprising a processor configured to implement a method mentioned in one or more of claims 1 to 27.

30. A computer program product having computer instructions stored therein, the instructions, when executed by means of a processor, cause the processor to implement a method mentioned in any one of claims 1 to 27.

31. A computer-readable non-transient storage medium that stores a bit stream generated according to the method of any one of claims 1 to 27.

32. A computer-readable, non-transient storage medium that stores instructions causing a processor to implement a method according to any of claims 1 to 27.

33. A bitstream generation method, comprising: generating a bitstream of a video according to a method according to any of claims 1 to 27, and storing the bitstream in a computer-readable program medium.

34. A method, apparatus, bit stream generated in accordance with a disclosed method or system described herein.