Virtual boundaries in video coding
By introducing virtual boundary signaling and more standardized random access point management, the difficult problems of decoding order and prediction relationship management in multi-layer video encoding and decoding in the existing technology are solved, the encoding and decoding efficiency and flexibility are improved, and the requirements of efficient parallel processing and low-latency video applications are met.
Patent Information
- Application Number
- CN202180031476.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-27
- Filing Date
- 2021-04-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-04-23
AI Technical Summary
Existing video codec technologies have difficulty effectively managing the decoding order, output order, and prediction relationships when processing multiple layers and different types of sub-pictures, resulting in low codec efficiency. In particular, there is a contradiction in efficient parallel processing and MTU size matching.
By introducing virtual boundary signaling and more standardized random access point management, it allows format rules to be used in video encoders and decoders to handle the codec representation of the video, define the order and prediction relationship of sub-picture types, support multi-layer video encoding and decoding, parallel processing and MTU size matching.
It improves the efficiency and flexibility of video encoding and decoding, reduces end-to-end latency, supports more efficient parallel processing and better random access point management, and adapts to the needs of modern video applications.
Smart Images

Figure CN115486082B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority under applicable patent laws and / or rules under the Paris Convention to and the benefit of U.S. Provisional Patent Application No. 63 / 016,122, filed on April 27, 2020. The entire disclosure of the above application is incorporated by reference as a part of the disclosure of this application for all purposes prescribed by law. Technical Field
[0003] The patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. Bandwidth demand for digital video usage is expected to continue to grow as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for processing a codec representation of a video using control information useful for decoding the codec representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video. The codec representation conforms to a format rule that specifies that the one or more pictures including the one or more sub-pictures are included in the codec representation in terms of network abstraction layer (NAL) units, wherein a type of NAL unit indicated in the codec representation includes a codec slice of a particular type of picture or a codec slice of a particular type of sub-picture.
[0007] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies that two adjacent sub-pictures having different network abstraction layer unit types are to have the same indication of being treated as sub-pictures of a picture flag.
[0008] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule, the format rule defining an order of sub-pictures of a first type and sub-pictures of a second type, wherein the first sub-picture is a trailing sub-picture, a leading sub-picture, or a random access skip leading (RASL) sub-picture type, and the second sub-picture is a RASL type, a random access decodable leading (RADL) type, an instantaneous decoding refresh (IDR) type, or a gradual decoding refresh (GDR) type sub-picture.
[0009] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule defining conditions under which a sub-picture of a first type is allowed or not allowed to appear with a sub-picture of a second type.
[0010] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video of multiple layers including one or more pictures and a bitstream of the video according to a format rule, wherein the format rule specifies that a reference picture referred to by each inter-layer reference picture entry in a reference picture list of a slice of a current picture of a current layer satisfies a constraint, wherein the constraint is at least one of: (a) the reference picture is an intra random access (IRAP) picture, or (b) the reference picture has a temporal identifier less than or equal to a specific value according to a maximum allowed value of the video layer that can be referenced by a slice from the current layer, wherein the maximum allowed value is indicated in a syntax element.
[0011] In another example aspect, another video processing method is disclosed. The method includes performing conversion between a video including a video picture and a bitstream according to a format rule, wherein the format rule specifies using one or more of a first syntax structure indicating a position of a vertical virtual boundary and a second syntax structure indicating a position of a horizontal virtual boundary to indicate a virtual boundary of a video region of the video picture, wherein the format rule specifies that a numerical value indicated by a first syntax element in the first syntax structure and a second syntax element in the second syntax structure is one less than the position of the vertical virtual boundary and the horizontal virtual boundary counted from a top left position of the video picture.
[0012] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.
[0013] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.
[0014] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code is in the form of processor-executable code for implementing one of the methods described herein.
[0015] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster scan strips.
[0017] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0018] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0019] Figure 4 A picture is shown partitioned into 15 slices, 24 slices, and 24 sub-pictures.
[0020] Figure 5 is a block diagram of an example video processing system.
[0021] Figure 6 It is a block diagram of a video processing device.
[0022] Figure 7 is a flow chart of an example method of video processing.
[0023] Figure 8 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0024] Figure 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0025] Figure 10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0026] Figure 11 is a flow chart of an example method of video processing based on some implementations of the disclosed technology.
[0027] Figure 12 is a flow chart of an example method of video processing based on some implementations of the disclosed technology. DETAILED DESCRIPTION
[0028] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, deletions to the current draft of the VVC specification are indicated by strikethrough, additions to the text are indicated by highlighting (including bold italics), and editorial changes are shown in the text.
[0029] 1. Introduction
[0030] This document relates to video coding techniques. Specifically, it is about the definition of sub-picture types and their relationship in terms of decoding order, output order, and prediction relationships between sub-pictures of different types, in both single-layer and multi-layer contexts. The key is to clearly specify the meaning of mixed sub-picture types within a picture through a set of constraints on decoding order, output order, and prediction relationships. These concepts can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Codec (VVC) under development.
[0031] 2. Abbreviation
[0032] APS Adaptive Parameter Set
[0033] AU Access Unit
[0034] AUD Access Unit Delimiter
[0035] AVC Advanced Video Codec
[0036] CLVS codec layer video sequence
[0037] CPB Coded Picture Buffer
[0038] CRA Clean Random Access
[0039] CTU Codec Tree Unit
[0040] CVS codec video sequence
[0041] DCI decoding capability information
[0042] DPB decoded picture buffer
[0043] EOB End of bitstream
[0044] EOS sequence end
[0045] GDR Gradual Decode Refresh
[0046] HEVC High-Efficiency Video Codec
[0047] HRD Hypothetical Reference Decoder
[0048] IDR Instant Decode Refresh
[0049] JEM Joint Exploration Model
[0050] MCTS Motion Constraint Patch
[0051] NAL Network Abstraction Layer
[0052] OLS output layer set
[0053] PH Image Header
[0054] PPS Picture Parameter Set
[0055] PTL profile, tier, and level
[0056] PU picture unit
[0057] RADL Random Access Decodable Preamble (Image)
[0058] RAP Random Access Point
[0059] RASL Random Access Skip Preamble (image)
[0060] RBSP Raw Byte Sequence Payload
[0061] RPL Reference Image List
[0062] SEI Supplemental Enhancement Information
[0063] SPS sequence parameter set
[0064] STSA Stepwise Temporal Sublayer Access
[0065] SVC Scalable Video Codec
[0066] VCL video codec layer
[0067] VPS Video Parameter Set
[0068] VTM VVC test model
[0069] VUI Video Availability Information
[0070] VVC multifunctional video codec
[0071] 3. Preliminary Discussion
[0072] Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are also held quarterly, and the goal for the new codec standard is to reduce bitrates by 50% compared to HEVC. The new video coding standard was officially named the Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to the ongoing efforts to standardize VVC, new codec technologies are adopted into the VVC standard at each JVET meeting. The VVC working draft and test model (VTM) are updated after each meeting. The VVC project is currently aiming for technical completion (FDIS) at the July 2020 meeting.
[0073] 3.1. Image Segmentation Scheme in HEVC
[0074] HEVC includes four different picture partitioning schemes, namely, regular slice, dependent slice, tile, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.
[0075] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although there may still be interdependencies due to loop filtering operations).
[0076] Regular slices are the only tool that can be used for parallelization, and they are also available in H.264 / AVC in a nearly identical form. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is generally much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, the use of regular slices incurs significant codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. In addition, due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit, regular slices (compared to the other tools mentioned below) also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting demands on the layout of slices within a picture. Recognition of this situation led to the development of the parallelization tools mentioned below.
[0077] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without destroying any intra-picture prediction. Essentially, dependent slices provide fragmentation of regular slices into multiple NAL units to provide reduced end-to-end latency by allowing part of a regular slice to be sent before coding of the entire regular slice is complete.
[0078] In WPP, a picture is partitioned into a single row of codec treeblocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible by decoding CTB rows in parallel, where the start of decoding a CTB row is delayed by two CTBs, ensuring that data associated with the CTB above and to the right of the subject CTB is available before the subject CTB is decoded. This staggered start (which, when represented graphically, looks like a wavefront) allows parallelization to utilize as many processors / cores as the picture contains CTB rows. Because intra-picture prediction is allowed between adjacent treeblock rows within a picture, the inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. WPP partitioning does not generate additional NAL units compared to when it is not used, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, but with some codec overhead.
[0079] Slices define the horizontal and vertical boundaries that divide an image into slice columns and rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0080] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of the slice's CTB raster scan). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in a single NAL unit (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header when a slice spans more than one slice, and loop filtering related to sharing of reconstruction samples and metadata. When more than one slice or WPP segment is included in a slice, the entry point byte offset of each slice or WPP segment in the slice except the first slice or WPP segment is signaled in the slice header.
[0081] For simplicity, HEVC specifies restrictions on the application of four different picture partitioning schemes. A given codec video sequence cannot contain both slices and wavefronts from most profiles specified in the HEVC standard. For each slice and slice, one or both of the following conditions must be met: 1) all coding tree blocks in a slice belong to the same slice; 2) all coding tree blocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when using WPP, if a slice starts on a CTB row, it must end on the same CTB row.
[0082] The latest amendments to HEVC are specified in the JCT-VC output document JCTVC-AC1005 "HEVC Additional Supplemental Enhancement Information (Draft 4)", published by J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, and Y.-K. Wang (eds.), October 24, 2017, and publicly available here: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this amendment, HEVC specifies three MCTS-related SEI messages, namely, the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.
[0083] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that only require full sample positions within the MCTS for interpolation, and motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction are not allowed. In this way, each MCTS can be decoded independently without the presence of slices not included in the MCTS.
[0084] The MCTS extraction information set SEI message provides supplementary information (specified as part of the SEI message semantics) that can be used in MCTS sub-bitstream extraction to generate a bitstream that conforms to the MCTS set. The information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes that replace the VPS, SPS, and PPS to be used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0085] 3.2. Image Segmentation in VVC
[0086] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of a picture. The CTUs in a slice are scanned in raster scan order within the slice.
[0087] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows in a slice of a picture.
[0088] Two striping modes are supported: raster scan striping mode and rectangular striping mode. In raster scan striping mode, a stripe contains a complete sequence of stripes in a stripe raster scan of a picture. In rectangular striping mode, a stripe contains multiple complete slices that together form a rectangular area of the picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of the picture. The slices within a rectangular stripe are scanned in slice raster scan order within the rectangular area corresponding to the stripe.
[0089] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.
[0090] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is divided into 12 slices and 3 raster scan strips.
[0091] Figure 2 An example of rectangular stripe partitioning of a picture is shown, where the picture is divided into 24 stripes (6 stripe columns and 4 stripe rows) and 9 rectangular stripes.
[0092] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0093] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, the 12 on the left hand side each covering one stripe of a 4×4 CTU, and the 6 slices on the right hand side each covering 2 vertically stacked strips of a 2×2 CTU, resulting in a total of 24 stripes and 24 sub-pictures of different dimensions (each stripe is a sub-picture).
[0094] 3.3. Changes in image accuracy within a sequence
[0095] In AVC and HEVC, the spatial precision of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC allows changing picture precision within a sequence without encoding an IRAP picture, which is always intra-coded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling the reference pictures used for inter prediction when the reference pictures have a different precision than the current picture being decoded.
[0096] The scaling ratio is restricted to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three resampling filter sets with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three resampling filter sets are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as the case of motion compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process, where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.
[0097] Other aspects of the VVC design that support this feature that differ from HEVC include: i) picture precision and the corresponding consistency window are signaled in the PPS rather than the SPS, where the maximum picture precision is signaled. ii) For a single-layer bitstream, each picture store (a slot in the DPB used to store a decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture precision.
[0098] Scalable Video Codec (SVC) in General and VVC
[0099] Scalable Video Codec (SVC, sometimes also referred to as scalability in video codecs) refers to video codecs that use a base layer (BL), sometimes called a reference layer (RL), and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously coded layers. For example, the bottom layer can serve as the BL, while the top layer can serve as the EL. Intermediate layers can serve as the EL, the RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can serve as the EL for layers below the intermediate layer (such as the base layer or any intervening enhancement layers) and simultaneously serve as the RL for one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to encode (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).
[0100] In SVC, parameters used by an encoder or decoder are grouped into parameter sets based on the codec level (e.g., video level, sequence level, picture level, slice level, etc.) at which they may be utilized. For example, parameters that may be utilized by one or more codec video sequences of different layers in a bitstream may be included in a video parameter set (VPS), and parameters that may be used by one or more pictures in a codec video sequence may be included in a sequence parameter set (SPS). Similarly, parameters used by one or more slices in a picture may be included in a picture parameter set (PPS), and other parameters specific to a single slice may be included in a slice header. Similarly, indications of which parameter set(s) a particular layer is using at a given time may be provided at various codec levels.
[0101] Due to VVC's support for reference picture resampling (RPR), support for bitstreams containing multiple layers (e.g., two layers with SD and HD precision in VVC) can be designed without the need for any additional signal processing-level codec tools, since the upsampling required for spatial scalability support can use only the RPR upsampling filter. However, for scalability support, a high level of syntax changes are required (compared to not supporting scalability). Scalability support is specified in VVC version 1. Unlike scalability support in any earlier video codec standards (including extensions of AVC and HEVC), VVC scalability is designed to be as friendly to single-layer decoder design as possible. The decoding capabilities of multi-layer bitstreams are specified as if there is only one layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream can decode a multi-layer bitstream without much change. Compared to the multi-layer extension design of AVC and HEVC, HLS is significantly simplified at the expense of some flexibility. For example, the IRAP AU needs to contain pictures from each layer present in the CVS.
[0102] 3.5. Random Access and Its Support in HEVC and VVC
[0103] Random access refers to accessing and decoding the bitstream starting from a picture that is not the first picture in the decoding order. In order to support tuning and channel switching in broadcast / multicast and multi-party video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points, which are usually intra-frame codec pictures, but can also be inter-frame codec pictures (for example, in the case of gradual decoding refresh).
[0104] HEVC includes signaling of intra random access points (IRAP) pictures in the NAL unit header by NAL unit types. Three types of IRAP pictures are supported, namely instantaneous decoder refresh (IDR), clean random access (CRA), and broken link access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any picture before the current group of pictures (GOP), often referred to as a closed-GOP random access point. CRA pictures are less constrained by allowing certain pictures to reference pictures before the current GOP, in the case of a random access where all pictures are discarded. CRA pictures are often referred to as an open-GOP random access point. BLA pictures typically result from splicing of two bitstreams or a portion thereof, e.g., during stream switching. To enable the system to better use IRAP pictures, a total of six different NAL units are defined to signal properties of IRAP pictures, which can be used to better match the stream access point types defined in the ISO base media file format (ISOBMFF), which are used for random access support in dynamic adaptive streaming over HTTP (DASH).
[0105] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or another type without associated RADL pictures), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) the basic functionality of BLA pictures can be achieved by a CRA picture plus a sequence NAL unit end, whose presence indicates that the following pictures start a new CVS in a single-layer bitstream. ii) during the development of VVC, it was desired to specify fewer NAL unit types than HEVC, as shown by the use of 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.
[0106] Another key difference in random access support between VVC and HEVC is that GDR is supported in a more standardized way in VVC. In GDR, decoding of the bitstream can start from an inter-frame coded picture, and although not the entire picture area can be correctly decoded at the beginning, after multiple pictures, the entire picture area will be correct. AVC and HEVC also support GDR, using the recovery point SEI message to signal the GDR random access point and recovery point. In VVC, a new NAL unit type is specified to indicate GDR pictures, and the recovery point is signaled in the picture header syntax structure. CVS and bitstreams are allowed to start with GDR pictures. This means that the entire bitstream is allowed to contain only inter-frame coded pictures, without a single intra-frame coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded slices or blocks across multiple pictures, rather than intra-coding the entire picture, thereby significantly reducing end-to-end latency, which is considered more important today than ever before as ultra-low latency applications such as wireless displays, online gaming, and drone-based applications become increasingly popular.
[0107] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refresh area (i.e., correctly decoded area) and the unrefreshed area at the picture between the GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so there will be no decoding mismatch of some samples at or near the boundary. This is very useful when the application decides to display the correctly decoded area during the GDR process.
[0108] IRAP pictures and GDR pictures may be collectively referred to as random access point (RAP) pictures.
[0109] 3.6. Reference Picture Management and Reference Picture List (RPL)
[0110] Reference picture management is a core function required by any video codec that uses inter-frame prediction. It manages the storage and removal of reference pictures in the decoded picture buffer (DPB) and places the reference pictures in the RPL in the correct order.
[0111] HEVC's reference picture management, including reference picture marking and removal from the decoded picture buffer (DPB) and reference picture list construction (RPLC), differs from AVC. Instead of the reference picture marking mechanism based on a sliding window plus adaptive memory management control operation (MMCO) in AVC, HEVC specifies a reference picture management and marking mechanism based on the so-called Reference Picture Set (RPS), and RPLC is therefore based on the RPS mechanism. The RPS consists of a reference picture set associated with a picture (consisting of all reference pictures that precede the associated picture in decoding order), which can be used for inter-frame prediction of the associated picture or any picture that follows the associated picture in decoding order. The reference picture set consists of five reference picture lists. The first three lists contain all reference pictures that can be used for inter-frame prediction of the current picture and for inter-frame prediction of one or more pictures that follow the current picture in decoding order. The other two lists consist of all reference pictures that are not used for inter-frame prediction of the current picture, but can be used for inter-frame prediction of one or more pictures that follow the current picture in decoding order. RPS provides "intra-frame codec" signaling of the DPB state, rather than "inter-frame codec" signaling as in AVC, mainly to improve error resilience. HEVC's RPLC process is based on RPS, by signaling the index to the RPS subset for each reference index; this process is simpler than the RPLC process in AVC.
[0112] VVC's reference picture management is more similar to HEVC than AVC, but simpler and more robust. As in those standards, two RPLs are derived, List 0 and List 1, but they are not based on the reference picture set concept used in HEVC or the automatic sliding window process used in AVC; instead, they are signaled more directly. Reference pictures used for RPLs are listed as active and inactive entries, and only active entries can be used as reference indices for inter prediction of CTUs of the current picture. Inactive entries indicate other pictures to be saved in the DPB for reference by other pictures arriving later in the bitstream.
[0113] Parameter Set
[0114] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced starting with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0115] The SPS is designed to carry sequence-level header information, and the PPS is designed to carry infrequently changing picture-level header information. With SPS and PPS, infrequently changing information does not need to be repeated for every sequence or picture, thus avoiding redundant signaling of this information. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, eliminating the need for redundant transmission and improving error resilience.
[0116] VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0117] APS is introduced to carry such picture-level or slice-level information, which requires quite a lot of bits to encode and decode, can be shared by multiple pictures, and can have quite a lot of different variations in the sequence.
[0118] 3.8. Related definitions in VVC
[0119] The relevant definitions in the latest VVC text (JVET-Q2001-vE / v15) are as follows.
[0120] Associated IRAP picture (of a specific picture): the previous IRAP picture in decoding order (when present) that has the same nuh_layer_id value as the specific picture.
[0121] Completely random access (CRA) PU: A PU in which the codec picture is a CRA picture.
[0122] Completely random access (CRA) picture: an IRAP picture in which the nal_unit_type of every VCL NAL unit is equal to CRA_NUT.
[0123] Coded Video Sequence (CVS): A sequence of AUs that, in decoding order, consists of a CVSS AU, followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs but excluding any subsequent AU that is a CVSS AU.
[0124] Codec Video Sequence Start (CVSS) AU: an AU in which each layer in the CVS has a PU and the codec picture in each PU is a CLVSS picture.
[0125] Gradual Decoding Refresh (GDR) AU: an AU in which the codec picture in each current PU is a GDR picture.
[0126] Gradual Decoding Refresh (GDR) PU: A PU in which the codec picture is a GDR picture.
[0127] Gradual Decoding Refresh (GDR) picture: A picture in which the nal_unit_type of each VCL NAL unit is equal to GDR_NUT.
[0128] Instantaneous Decoding Refresh (IDR) PU: A PU in which the codec picture is an IDR picture.
[0129] Instantaneous Decoding Refresh (IDR) picture: An IRAP picture with nal_unit_type equal to IDR_W_RADL or IDR_N_LP for each VCL NAL unit.
[0130] Intra Random Access Point (IRAP) AU: an AU in which each layer in the CVS has a PU and the codec picture in each PU is an IRAP picture.
[0131] Intra Random Access Point (IRAP) PU: A PU in which the codec picture is an IRAP picture.
[0132] Intra Random Access Point (IRAP) picture: A codec picture whose all VCL NAL units have the same value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive.
[0133] Leading picture: A picture that is in the same layer as the associated IRAP picture and precedes the associated IRAP picture in output order.
[0134] Output order: The order in which decoded pictures are output from the DPB (for decoded pictures to be output from the DPB).
[0135] Random Access Decodable Leading (RADL) PU: A PU in which the codec picture is a RADL picture.
[0136] Random Access Decodable Leading (RADL) picture: A picture with nal_unit_type equal to RADL_NUT for each VCL NAL unit.
[0137] Random Access Skip Leading (RASL) PU: A PU in which the codec picture is a RASL picture.
[0138] Random Access Skip Leading (RASL) picture: A picture with nal_unit_type equal to RASL_NUT for each VCL NAL unit.
[0139] Step-by-step temporal sub-layer access (STSA) PU: A PU in which the codec picture is a STSA picture.
[0140] Step-by-step temporal sub-layer access (STSA) picture: A picture with nal_unit_type equal to STSA_NUT for each VCL NAL unit.
[0141] Note: STSA pictures do not use pictures with the same temporal ID (TemporalId) as the STSA picture for inter-frame prediction reference. Pictures with the same TemporalId as the STSA picture that follow the STSA picture in decoding order do not use pictures with the same TemporalId as the STSA picture that precede the STSA picture in decoding order for inter-frame prediction reference. STSA pictures enable upward switching from the immediately lower sublayer to the sublayer containing the STSA picture at the STSA picture. The TemporalId of the STSA picture must be greater than 0.
[0142] Sub-image: A rectangular area of one or more strips within an image.
[0143] Trailing picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not an STSA picture.
[0144] Note: The trailing picture associated with an IRAP picture also follows the IRAP picture in decoding order. Pictures that follow the associated IRAP picture in output order and precede the associated IRAP picture in decoding order are not allowed.
[0145] 3.9. NAL Unit Header Syntax and Semantics in VVC
[0146] In the latest VVC text (in JVET-Q2001-vE / v15), the NAL unit header syntax and semantics are as follows.
[0147] 7.3.1.2 NAL unit header syntax
[0148]
[0149] 7.4.2.2 NAL unit header semantics
[0150] forbidden_zero_bit should be equal to 0.
[0151] nuh_reserved_zero_bit shall be equal to 0. The value of nuh_reserved_zero_bit 1 may be specified by ITUT|ISO / IEC in the future. The decoder shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1.
[0152] nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which the non-VCL NAL unit applies. The value of nuh_layer_id shall be in the range of 0 to 55 (inclusive). Other values of nuh_layer_id are reserved for future use by ITU-T | ISO / IEC.
[0153] The value of nuh_layer_id shall be the same for all VCL NAL units of a codec picture. The value of nuh_layer_id of a codec picture or PU shall be the value of nuh_layer_id of the VCL NAL unit of the codec picture or PU.
[0154] The nuh_layer_id values for AUD, PH, EOS, and FD NAL units are constrained as follows:
[0155] If nal_unit_type is equal to AUD_NUT, nuh_layer_id shall be equal to vps_layer_id[0].
[0156] Otherwise, when nal_unit_type is equal to PH_NUT, EOS_NUT, or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit.
[0157] NOTE 1 The value of nuh_layer_id of DCI, VPS and EOB NAL units is not constrained.
[0158] The value of nal_unit_type should be the same for all pictures in a CVSS AU.
[0159] nal_unit_type specifies the NAL unit type, ie, the type of RBSP data structure contained in the NAL unit as specified in Table 5.
[0160] NAL units with a nal_unit_type in the range UNSPEC_28..UNSPEC_31, inclusive, have unspecified semantics and should not affect the decoding process specified in this specification.
[0161] NOTE 2 NAL unit types in the range UNSPEC_28..UNSPEC_31 may be used, as determined by the application. The decoding process for these values of nal_unit_type is not specified in this specification. Because different applications may use these NAL unit types for different purposes, care must be taken when designing encoders that generate NAL units with these nal_unit_type values, and when designing decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define any management of these values. These nal_unit_type values may only be used in contexts where usage "collisions" (i.e., different definitions of the meaning of the content of NAL units with the same nal_unit_type value) are not important, are not possible, or are managed (e.g., as defined or managed in a controlling application or transport specification, or by an environment that controls bitstream distribution).
[0162] For purposes other than determining the amount of data in a DU of a bitstream (as specified in Annex C), a decoder shall ignore (remove from the bitstream and discard) the contents of all NAL units that use reserved values of nal_unit_type.
[0163] NOTE 3 This requirement allows for the definition of compatible extensions to this specification in the future.
[0164] Table 5 NAL unit type codes and NAL unit type categories
[0165]
[0166]
[0167] NOTE 4 A completely random access (CRA) picture may have an associated RASL or RADL picture present in the bitstream.
[0168] NOTE 5 An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP does not have an associated leading picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream.
[0169] The value of nal_unit_type shall be the same for all VCL NAL units in a sub-picture. A sub-picture is considered to have the same NAL unit type as the VCL NAL units of the sub-picture.
[0170] For any particular picture's VCL NAL unit, the following applies:
[0171] If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and a picture or PU is considered to have the same NAL unit type as the VCL NAL units of the picture or PU.
[0172] Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture shall have at least two sub-pictures, and the VCL NAL units of the picture shall have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one sub-picture of the picture shall all have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of the other sub-pictures in the picture shall all have different specific values of nal_unit_type equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.
[0173] For single-layer bitstreams, the following constraints apply:
[0174] Except for the first picture in the bitstream in decoding order, every picture is considered to be associated with the previous IRAP picture in decoding order.
[0175] When a picture is the leading picture of an IRAP picture, it shall be a RADL or RASL picture.
[0176] When a picture is the trailing picture of an IRAP picture, it shall not be a RADL or RASL picture.
[0177] There shall be no RASL pictures associated with an IDR picture in the bitstream.
[0178] There shall be no RADL pictures in the bitstream that are associated with an IDR picture with nal_unit_type equal to IDR_N_LP.
[0179] NOTE 6 It is possible to perform random access at the location of an IRAP PU (and correctly decode the IRAP picture and all subsequent non-RASL pictures in decoding order) by discarding all PUs preceding the IRAP PU, provided that every parameter set is available (in the bitstream or by external means not specified in this specification) when referenced.
[0180] Any picture that precedes an IRAP picture in decoding order shall precede the IRAP picture in output order, and shall precede any RADL pictures associated with the IRAP picture in output order.
[0181] Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order.
[0182] Any RASL picture associated with a CRA picture shall follow any IRAP picture that precedes the CRA picture in decoding order in output order.
[0183] If field_seq_flag is equal to 0 and the current picture is a leading picture associated with an IRAP picture, it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, let picA and picB be the first and last leading pictures associated with an IRAP picture in decoding order, there shall be at most one non-leading picture that precedes picA in decoding order and no non-leading picture between picA and picB in decoding order.
[0184] nuh_temporal_id_plus1 minus 1 specifies the temporal identifier of the NAL unit.
[0185] The value of nuh_temporal_id_plus1 shall not be equal to 0.
[0186] The variable Temporalld is derived as follows:
[0187] Temporalld = nuh_temporal_id_plus1 - 1 (36)
[0188] When nal_unit_type is in the range of IDR W RADL to RSV IRAP 12, inclusive, Temporalld shall be equal to 0.
[0189] When nal_unit_type is equal to STSA NUT and vps_independent_layer_flag[ GeneralLayerldx[ nuh_layer_id ] ] is equal to 1, Temporalld shall not be equal to 0.
[0190] The value of Temporalld shall be the same for all VCL NAL units of an AU. The value of Temporalld for a coded picture, PU, or AU is the value of Temporalld for the VCL NAL units of the coded picture, PU, or AU. The value of Temporalld for a sublayer representation is the maximum value of Temporalld for all VCL NAL units in the sublayer representation.
[0191] The TemporalId values of non-VCL NAL units are constrained as follows:
[0192] If nal_unit_type is equal to DCI_NUT, VPS_NUT, or SPS_NUT, TemporalId shall be equal to 0, and the TemporalId of the AU containing the NAL unit shall be equal to 0.
[0193] Otherwise, if nal_unit_type is equal to PH_NUT, TemporalId shall be equal to the TemporalId of the PU containing the NAL unit.
[0194] Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0.
[0195] Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, TemporalId shall be equal to the TemporalId of the AU containing the NAL unit.
[0196] Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU containing the NAL unit.
[0197] NOTE 7 When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values of all AUs to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId can be greater than or equal to the TemporalId of the containing AU, because all PPS and APS can be included at the beginning of the bitstream (for example, when they are transmitted out-of-band and the receiver places them at the beginning of the bitstream), where the first codec picture has a TemporalId equal to 0.
[0198] 3.10. Mixed NAL Unit Types within a Picture
[0199] 7.4.3.4 Picture parameter set semantics ...
[0201] mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture of the referenced PPS has more than one VCL NAL unit, the VCL NAL units do not have the same value of nal_unit_type, and the picture is not an IRAP picture. mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same value of nal_unit_type.
[0202] When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_types_in_pic_flag shall be equal to 0.
[0203] For each slice in a picture picA having a nal_unit_type value nalUnitTypeA in the range IDR_W_RADL to CRA_NUT (inclusive), the picture picA also contains one or more slices with another nal_unit_type value, i.e., the value of mixed_nalu_types_in_pic_flag of picture picA is equal to 1, the following applies:
[0204] The slice shall belong to the sub-picture subpicA whose corresponding subpic_treated_as_pic_flag[i] value is equal to 1.
[0205] A slice shall not belong to a sub-picture of picA containing a VCL NAL unit with nal_unit_type not equal to nalUnitTypeA.
[0206] If nalUnitTypeA is equal to CRA, then for all subsequent PUs that follow the current picture in CLVS in decoding order and output order, the RefPicList[0] and RefPicList[1] of the slices in subpicA in these PUs shall not include any pictures that precede picA in decoding order in the valid entries.
[0207] Otherwise (i.e., nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS that follow the current picture in decoding order, neither the RefPicList[0] nor the RefPicList[1] of the slices in subpicA in these PUs shall include any pictures in the valid entries that precede picA in decoding order.
[0208] NOTE 1 mixed_nalu_types_in_pic_flag equal to 1 indicates that the picture referencing the PPS contains slices with different NAL unit types, for example, a codec picture resulting from a sub-picture bitstream merge operation, for which the encoder must ensure bitstream structure matching and further alignment of parameters of the original bitstreams. An example of such alignment is as follows: When the value of sps_idr_rpl_present_flag is equal to 0 and mixed_nalu_types_in_pic_flag is equal to 1, the picture referencing the PPS shall not have slices with nal_unit_type equal to IDR_W_RADL or IDR_N_LP. ...
[0210] 3.11. VVC Picture Header Structure Syntax and Semantics
[0211] In the latest VVC specification (in JVET-Q2001-vE / v15), the picture header structure syntax and semantics most relevant to the present invention are as follows.
[0212] 7.3.2.7 Picture header structure syntax
[0213]
[0214]
[0215] 7.4.3.7 Image header structure semantics
[0216] The PH syntax structure contains information common to all slices of the codec picture associated with the PH syntax structure.
[0217] gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag equal to 0 specifies that the current picture may or may not be a GDR or IRAP picture.
[0218] gdr_pic_flag equal to 1 specifies that the picture associated with the PH is a GDR picture. gdr_pic_flag equal to 0 specifies that the picture associated with the PH is not a GDR picture. When not present, the value of gdr_pic_flag is inferred to be equal to 0. When gdr_enabled_flag is equal to 0, the value of gdr_pic_flag shall be equal to 0.
[0219] NOTE 1 When gdr_or_irap_pic_flag is equal to 1 and gdr_pic_flag is equal to 0, the picture associated with the PH is an IRAP picture. ...
[0221] ph_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the ph_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of ph_pic_order_cnt_lsb shall be in the range of 0 to MaxPicOrderCntLsb-1, inclusive.
[0222] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding a CLVSS picture that is not the first picture in the bitstream as specified in Annex C.
[0223] recovery_poc_cnt specifies the recovery point of a decoded picture in output order. If the current picture is a GDR picture associated with PH, and there is a picture picA that follows the current GDR picture in decoding order in CLVS and whose PicOrderCntVal is equal to the PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt, then picture picA is called the recovery point picture. Otherwise, the first picture in output order whose PicOrderCntVal is greater than the PicOrderCntVal of the current picture plus the value of recovery_poc_cnt is called the recovery point picture. The recovery point picture should not precede the current GDR picture in decoding order. The value of recovery_poc_cnt should be in the range of 0 to MaxPicOrderCntLsb-1, inclusive.
[0224] When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:
[0225] RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt (81)
[0226] NOTE 2 When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the RpPicOrderCntVal of the associated GDR picture, the current and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (when present) that precedes the associated GDR picture in decoding order. ...
[0228] 3.12. Restrictions on RPL in VVC
[0229] In the latest VVC text (in JVET-Q2001-vE / vl5), the restrictions on RPL in VVC are as follows (as part of the decoding process of the construction of reference picture lists in Section 8.3.2 of VVC). 8.3.2 Decoding process of construction of reference picture lists ...
[0231] For each i equal to 0 or 1, the first NumRefIdxActive[i] entries in RefPicList[i] are considered to be active entries in RefPicList[i], and other entries in RefPicList[i] are considered to be inactive entries in RefPicList[i].
[0232] NOTE 2 A particular picture can be referred to by both an entry in RefPicList[0] and an entry in RefPicList[1]. It is also possible that a particular picture is referred to by more than one entry in RefPicList[0] or more than one entry in RefPicList[1].
[0233] NOTE 3 The active entries in RefPicList[0] and the active entries in RefPicList[1] together refer to all reference pictures that can be used for inter prediction of the current picture and one or more pictures following the current picture in decoding order. The inactive entries in RefPicList[0] and the inactive entries in RefPicList[1] together refer to all reference pictures that are not used for inter prediction of the current picture but can be used for inter prediction of one or more pictures following the current picture in decoding order.
[0234] NOTE 4 There can be one or more entries in RefPicList[0] or RefPicList[1] equal to “no reference picture” because the corresponding picture does not exist in the DPB. Each inactive entry in RefPicList[0] or RefPicList[0] equal to “no reference picture” shall be ignored. An unexpected picture loss shall be inferred for each active entry in RefPicList[0] or RefPicList[1] equal to “no reference picture”.
[0235] The requirement of bitstream conformance is that the following constraints apply:
[0236] For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i].
[0237] The picture referred to by each valid entry in RefPicList[0] or RefPicList[1] shall exist in the DPB and shall have a TemporalId less than or equal to the current picture.
[0238] Each entry in RefPicList[0] or RefPicList[1] shall refer to a picture that shall not be the current picture and shall have non_reference_picture_flag equal to 0.
[0239] A STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and an LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture.
[0240] There shall be no LTRP entry in RefPicList[0] or RefPicList[1] for which the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referred to by the entry is greater than or equal to 2 24 .
[0241] Let setOfRefPics be the set of unique pictures referred to by all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize - 1 (including itself), where MaxDpbSize is specified in clause A.4.2, and setOfRefPics shall be the same for all slices of the picture.
[0242] When nal_unit_type of the current slice is equal to STSA_NUT, there shall be no valid entry in RefPicList[0] or RefPicList[1] with TemporalId equal to TemporalId of the current picture and nuh_layer_id equal to nuh_layer_id of the current picture.
[0243] When the current picture is a picture that follows the STSA picture in decoding order and has a TemporalId equal to the TemporalId of the current picture and a nuh_layer_id equal to the nuh_layer_id of the current picture, there should be no pictures that precede the STSA picture in decoding order and have a TemporalId equal to the TemporalId of the current picture and have a nuh_layer_id equal to the nuh_layer_id of the current picture included as valid entries in RefPicList[0] or RefPicList[1].
[0244] When the current picture is a CRA picture, there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] that precede any previous IRAP picture (when present) in output order or decoding order.
[0245] When the current picture is the last picture, there should be no RefPicList[0] or RefPicList[1]
[0246] The picture referred to by the valid entry in is generated by the decoding process of the unavailable reference picture used to generate the IRAP picture associated with the current picture.
[0247] When the current picture is the trailing picture that follows one or more leading pictures (if any) associated with the same IRAP picture in both decoding order and output order, there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] that were generated by the decoding process for generating an unavailable reference picture for the IRAP picture associated with the current picture.
[0248] When the current picture is a recovery point picture or a picture that follows the recovery point picture in output order, there shall be no entries in RefPicList[0] or RefPicList[1] that contain pictures generated by the decoding process of unavailable reference pictures of the GDR picture used to generate the recovery point picture.
[0249] When the current picture is the trailing picture, there shall be no pictures referred to by valid entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.
[0250] When the current picture is the trailing picture of one or more leading pictures (if any) associated with the same IRAP picture in decoding order and output order, there shall be no pictures referred to by entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.
[0251] When the current picture is a RADL picture, there shall not be any valid entries in RefPicList[0] or RefPicList[1]:
[0252] οRASL pictures
[0253] o Pictures generated by the decoding process used to generate unavailable reference pictures
[0254] o pictures that precede the associated IRAP picture in decoding order
[0255] Each ILRP entry in the RefPicList[0] or RefPicList[1] of the slice of the current picture shall refer to a picture in the same AU as the current picture.
[0256] The picture referred to by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB and shall have a nuh_layer_id that is less than the nuh_layer_id of the current picture.
[0257] Each ILRP entry in the stripe's RefPicList[0] or RefPicList[1] shall be a valid entry. ...
[0259] 3.13.PictureOutputFlag setting
[0260] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting the value of the variable PictureOutputFlag is as follows (as part of clause 8.1.2 decoding process for codec pictures).
[0261] 8.1.2 Decoding of Coded Images
[0262] The decoding process specified in this clause applies to each coded picture in BitstreamToDecode, called the current picture and represented by the variable CurrPic.
[0263] Depending on the value of chroma_format_idc, the number of sample arrays in the current picture is as follows:
[0264] If chroma_format_idc is equal to 0, the current picture consists of 1 sample array S L .
[0265] Otherwise (chroma_format_idc is not equal to 0), the current picture consists of 3 sample arrays S L , S Cb , S Cr .
[0266] The decoding process of the current picture takes as input the syntax elements from clause 7 and the capital variables. When interpreting the semantics of each syntax element in each NAL unit and in the remainder of clause 8, the term "bitstream" (or a part thereof, e.g. the CVS of the bitstream) refers to BitstreamToDecode (or a part thereof).
[0267] Depending on the value of separate_colour_plane_flag, the decoding process is constructed as follows:
[0268] If separate_colour_plane_flag is equal to 0, the decoding process is invoked once, outputting the current picture.
[0269] Otherwise (separate_colour_plane_flag is equal to 1), the decoding process is invoked three times. The input to the decoding process is all NAL units of the coded picture with the same value of colour_plane_id. The decoding process of the NAL units with a particular colour_plane_id value is specified as a CVS in the bitstream that only contains monochrome colour format with that particular colour_plane_id value. The output of each of the three decoding processes is assigned to one of the 3 sample arrays of the current picture, the NAL units with colour_plane_id equal to 0, 1 and 2 are assigned to S L , S Cb and S Cr , respectively.
[0270] NOTE When separate_colour_plane_flag is equal to 1 and chroma_format_idc is equal to 3, the variable ChromaArrayType is equal to 0. In the decoding process, the value of this variable is evaluated, resulting in the same operation as for monochrome pictures (when chroma_format_idc is equal to 0).
[0271] For the current picture CurrPic, the decoding process operates as follows:
[0272] 1. The decoding of NAL units is specified in clause 8.2.
[0273] 2. The process in clause 8.3 uses syntax elements in the slice header layer and above to specify the following decoding process:
[0274] The variables and functions related to the picture sequence count are derived as specified in clause 8.3.1. This only needs to be called for the first slice of a picture.
[0275] At the beginning of the decoding process of each slice of a non-IDR picture, the decoding process of the reference picture list construction specified in clause 8.3.2 is called to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).
[0276] The decoding process for reference picture marking in clause 8.3.2 is invoked, where a reference picture can be marked as “unused for reference” or “used for long-term reference.” This only needs to be invoked for the first slice of a picture.
[0277] When the current picture is a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 or a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, the decoding process for generating unusable reference pictures as specified in subclause 8.3.4 is invoked, which only needs to be invoked for the first slice of the picture.
[0278] PictureOutputFlag is set as follows:
[0279] PictureOutputFlag is set equal to 0 if one of the following conditions is true:
[0280] The current picture is a RASL picture, and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.
[0281] gdr_enabled_flag is equal to 1, and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0282] gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.
[0283] sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that meets all of the following conditions:
[0284] PicA's PictureOutputFlag is equal to 1.
[0285] PicA's nuh_layer_id nuhLid is greater than the current picture's nuh_layer_id.
[0286] PicA belongs to the output layer of OLS (ie, OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0287] sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[target ols idx][general layeridx[nuh_layer_id]] is equal to 0.
[0288] Otherwise, PictureOutputFlag is set equal to pic_output_flag.
[0289] 3. The procedures in clauses 8.4, 8.5, 8.6, 8.7, and 8.8 use syntax elements from all syntax structure layers to specify the decoding process. A bitstream conformance requirement is that the codec slice of a picture shall contain slice data for each CTU of the picture such that the picture is divided into slices, and the slices are divided into CTUs, each forming a partition of the picture.
[0290] 4. After all slices of the current picture are decoded, the currently decoded picture is marked as "used for short-term reference" and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "used for short-term reference".
[0291] 4. Technical problems solved by the disclosed technical solutions
[0292] The existing design in the latest VVC text (in JVET-Q2001-vE / v15) has the following problems:
[0293] 1) Since different types of sub-pictures within a picture are allowed to be mixed, it is confusing to refer to the contents of a NAL unit with a VCL NAL unit type as a codec slice of a particular type of picture. For example, a NAL unit with nal_unit_type equal to CRA_NUT is a codec slice of a CRA picture only if the nal_unit_type of all slices of the picture is equal to CRA_NUT; when the nal_unit_type of a slice of the picture is not equal to CRA_NUT, the picture is not a CRA picture.
[0294] 2) Currently, if a sub-picture contains VCL NAL units with nal_unit_type in the range IDR_W_RADL to CRA_NUT (inclusive), and for the picture, mixed_nalu_types_in_pic_flag is equal to 1, then for the sub-picture, the value of subpic_treated_as_pic_flag[] needs to be equal to 1. In other words, for an IRAP sub-picture that is mixed with a sub-picture of another type in the picture, the value of subpic_treated_as_pic_flag[] needs to be equal to 1. However, with support for more mixed VCL NAL unit types, this requirement is not sufficient.
[0295] 3) Currently, only two different types of VCL NAL units (and two different types of sub-pictures) are allowed within a picture.
[0296] 4) In both single-layer and multi-layer contexts, there is a lack of constraints on the output order of the trailing sub-pictures relative to the associated IRAP or GDR sub-pictures.
[0297] 5) Currently, it is specified that a picture should be a RADL or RASL picture when it is the leading picture of an IRAP picture. This constraint, together with the definition of leading / RADL / RASL pictures, does not allow mixing of RADL and RASL NAL unit types within a picture resulting from a mixture of two CRA pictures and their non-AU-aligned associated RADL and RASL pictures.
[0298] 6) In both single-layer and multi-layer contexts, for the leading sub-picture, there is a lack of constraints on the sub-picture type (i.e., the NAL unit type of the VCL NAL unit in the sub-picture).
[0299] 7) In both single-layer and multi-layer contexts, there is a lack of constraints on whether RASL sub-pictures can exist and be associated with IDR sub-pictures.
[0300] 8) In single-layer and multi-layer contexts, there is a lack of constraints on whether a RADL sub-picture can exist and be associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP.
[0301] 9) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between sub-pictures that precede an IRAP sub-picture in decoding order and RADL sub-pictures associated with the IRAP sub-picture.
[0302] 10) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between sub-pictures that precede a GDR sub-picture in decoding order and sub-pictures that are associated with a GDR sub-picture.
[0303] 11) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between RASL sub-pictures associated with CRA sub-pictures and RADL sub-pictures associated with CRA sub-pictures.
[0304] 12) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative output order between RASL sub-pictures associated with a CRA sub-picture and IRAP sub-pictures that precede the CRA sub-picture in decoding order.
[0305] 13) In both single-layer and multi-layer contexts, there is a lack of constraints on the relative decoding order between the associated non-leading pictures and the leading pictures of an IRAP picture.
[0306] 14) In single-layer and multi-layer contexts, there is a lack of constraints on RPL validity entries for sub-pictures that follow the STSA sub-picture in decoding order.
[0307] 15) In both single-layer and multi-layer contexts, there are no constraints on the RPL entries of CRA subgraphs.
[0308] 16) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL valid entries referring to sub-pictures of pictures generated by the decoding process used to generate unavailable reference pictures.
[0309] 17) In both single-layer and multi-layer contexts, there is a lack of constraints on RPL entries referring to sub-pictures of a picture generated by the decoding process used to generate unavailable reference pictures.
[0310] 18) In single-layer and multi-layer contexts, there is a lack of constraints on RPL validity entries for sub-pictures that are associated with an IRAP picture and follow the IRAP picture in output order.
[0311] 19) In both single-layer and multi-layer contexts, there are no constraints on RPL entries for sub-pictures that are associated with an IRAP picture and follow the IRAP picture in output order.
[0312] 20) In single-layer and multi-layer contexts, there is a lack of constraints on the RPL valid entries for RADL sub-pictures.
[0313] 5. Examples of solutions and implementations
[0314] To address the above-mentioned and other problems, the methods outlined below are disclosed. These items should be considered as examples to explain the general concepts and should not be interpreted in a narrow sense. Furthermore, these items can be used alone or in any combination.
[0315] 1) To address issue 1, instead of specifying the content of a NAL unit with a VCL NAL unit type as a "codec slice of a specific type of picture," it is specified as a "codec slice of a specific type of picture or sub-picture." For example, the content of a NAL unit with nal_unit_type equal to CRA_NUT is specified as a "codec slice of a CRA picture or sub-picture."
[0316] a. In addition, define one or more of the following terms: associated GDR sub-picture, associated IRAP sub-picture, CRA sub-picture, GDR sub-picture, IDR sub-picture, IRAP sub-picture, leading sub-picture, RADL sub-picture, RASL sub-picture, STSA sub-picture, and trailing sub-picture.
[0317] 2) To solve problem 2, a constraint is added to require that any two adjacent sub-pictures with different NAL unit types should have subpic_treated_as_pic_flag[] equal to 1.
[0318] a. In one example, the constraints are specified as follows: for any two adjacent sub-pictures with sub-picture indices i and j in a picture, when subpic_treated_as_pic_flag[i] or subpic_treated_as_pic_flag[j] is equal to 0, the two sub-pictures should have the same NAL unit type.
[0319] a.Alternatively, it is required that when any subpicture with subpicture index i has subpic_treated_as_pic_flag[ i ] equal to 0, all subpictures in the picture should have the same NAL unit type (i.e., all VCL NAL units in the picture should have the same NAL unit type, i.e., the value of mixed_nalu_types_in_pic_flag should be equal to 0). This means that when all subpictures have their corresponding subpic_treated_as_pic_flag[ ] equal to 1, mixed_nalu_types_in_pic_flag can only be equal to 1.
[0320] 3) To solve problem 3, when mixed_nalu_types_in_pic_flag is equal to 1, it can be allowed that a picture contains more than two different types of VCL NAL units.
[0321] 4) To solve problem 4, it is specified that trailing subpictures shall follow the associated IRAP or GDR subpicture in output order.
[0322] 5) To solve problem 5, to allow intra-picture mixing of RADL and RASL NAL unit types resulting from a mix of two CRA pictures and their non-AU-aligned associated RADL and RASL pictures, the existing constraint that a RADL or RASL picture shall be the leading picture of an IRAP picture is changed as follows: When a picture is the leading picture of an IRAP picture, the nal_unit_type values of all VCL NAL units in the picture shall be equal to RADL NUT or RASL NUT. In addition, in the decoding process of a picture with mixed nal_unit_type values of RADL NUT and RASL NUT, the PictureOutputFlag of the picture is set equal to pic_output_flag when the layer containing the picture is an output layer.
[0323] Thus, by requiring all output pictures to be corrected to conform to the decoder's constraints, RADL sub-pictures in such pictures can be guaranteed, although guarantees of "correctness" of "mid-valued" RASL sub-pictures in such pictures are also appropriate when the associated CRA picture has NoOutputBeforeRecoveryFlag equal to 1, but are not actually required. The unnecessary parts of the guarantee are insignificant and do not increase the complexity of implementing a conforming encoder or decoder. In this case, a note is added to clarify that although such RASL sub-pictures associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 can be output by the decoding process, they are not intended for display and should not be used for display.
[0324] 6) To solve problem 6, it is stipulated that when a sub-picture is a leading sub-picture of an IRAP sub-picture, it should be a RADL or RASL sub-picture.
[0325] 7) To solve problem 7, it is specified that there should not be RASL sub-pictures associated with IDR sub-pictures in the bitstream.
[0326] 8) To solve problem 8, it is specified that there should not be any RADL sub-picture associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP in the bitstream.
[0327] 9) To address issue 9, it is specified that, in decoding order, any sub-picture with nuh_layer_id equal to a specific value layerId and sub-picture index equal to a specific value subpicIdx that precedes an IRAP sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx shall precede the IRAP sub-picture and all its associated RADL sub-pictures in output order.
[0328] 10) To address issue 10, it is specified that, in decoding order, any sub-picture with nuh_layer_id equal to a specific value layerId and sub-picture index equal to a specific value subpicIdx that precedes a GDR sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx shall precede the GDR sub-picture and all its associated sub-pictures in output order.
[0329] 11) To address issue 11, it is specified that any RASL sub-picture associated with a CRA sub-picture should precede any RADL sub-picture associated with the CRA sub-picture in output order.
[0330] 12) To address issue 12, it is specified that any RASL sub-picture associated with a CRA sub-picture should follow any IRAP sub-picture in output order, and any IRAP sub-picture should precede the CRA sub-picture in decoding order.
[0331] 13) To address issue 13, it is specified that if field_seq_flag is equal to 0, and the current sub-picture (where nuh_layer_id is equal to a specific value layerId and the sub-picture index is equal to a specific value subpicIdx) is a leading sub-picture associated with an IRAP sub-picture, then it should precede all non-leading sub-pictures associated with the same IRAP sub-picture in decoding order; otherwise, let subpicA and subpicB be the first and last leading sub-pictures associated with the IRAP sub-picture in decoding order, respectively, there should be at most one non-leading sub-picture with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx that precedes subpicA in decoding order, and there should not be any non-leading pictures with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx between picA and picB in decoding order.
[0332] 14) In order to solve problem 14, it is stipulated that when the current sub-picture whose TemporalId is equal to a specific value tId, nuh_layer_id is equal to a specific value layerId, and the sub-picture index is equal to a specific value subpicIdx is a sub-picture after the STSA sub-picture whose TemporalId is equal to tId, nuh_layer_id is equal to layerId, and the sub-picture index is equal to subpicIdx in decoding order, there should not be a picture whose TemporalId is equal to tId and nuh_layer_id is equal to layerId before the picture containing the STSA sub-picture in decoding order, which should be included as a valid entry in RefPicList[0] or RefPicList[1].
[0333] 15) To address issue 15, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is a CRA sub-picture, there should not be any picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes, in output order or decoding order, any picture that contains a preceding IRAP sub-picture (when present) with nuh_layer_id equal to layerId and sub-picture index equal to subpicIdx in decoding order.
[0334] 16) To address issue 16, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is not a RASL sub-picture associated with a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR sub-picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a sub-picture of a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there should not be a picture generated by the decoding process to generate an unavailable reference picture referred to by a valid entry in RefPicList[0] or RefPicList[1].
[0335] 17) To address issue 17, it is specified that when the current sub-picture with nuh_layer_id equal to a specific value layerId and a sub-picture index equal to a specific value subpicIdx is not a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a sub-picture that precedes in decoding order the leading sub-picture associated with the same CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a leading sub-picture associated with a CRA sub-picture of a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR sub-picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a sub-picture of a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, the picture referred to by the entry in RefPicList[0] or RefPicList[1] that was generated by the decoding process to generate an unavailable reference picture should not exist.
[0336] 18) To address issue 18, it is specified that when the current sub-picture is associated with an IRAP sub-picture and follows the IRAP sub-picture in output order, there should not be a picture referred to by a valid entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the associated IRAP sub-picture in output order or decoding order.
[0337] 19) To address issue 19, it is specified that when a current sub-picture is associated with an IRAP sub-picture, follows the IRAP sub-picture in output order, and follows the leading sub-picture (if any) associated with the same IRAP sub-picture in both decoding order and output order, there shall not be a picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the associated IRAP sub-picture in output order or decoding order.
[0338] 20) To address issue 20, it is specified that when the current sub-picture is a RADL sub-picture, there shall not be any valid entries in RefPicList[0] or RefPicList[1] for any of the following:
[0339] a. Pictures containing RASL sub-pictures
[0340] b. The picture that precedes the picture containing the associated IRAP sub-picture in decoding order
[0341] 6. Examples
[0342] The following are some example embodiments of some aspects of the invention outlined in Section 5 above, which may be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-Q2001-vE / v15. Most of the relevant parts that have been added or modified are The text of the present invention is highlighted, and some deleted parts are marked with double brackets (e.g., [[a]] indicates the deletion of the character "a"). Some other changes are editorial in nature or not part of the present invention and are therefore not highlighted.
[0343] 6.1. First embodiment
[0344] This embodiment applies to items 1, 1a, 2, 2a, 4 and 6 to 20.
[0345] 3 Definition ...
[0347] Associated GDR pictures (of a specific picture with a specific value of nuh_layer_id layerId): in decoding order, there is no IRAP picture with nuh_layer_id equal to layerId between the previous GDR picture with nuh_layer_id equal to layerId (when present) and the specific picture in decoding order.
[0348]
[0349] Associated IRAP pictures (of a specific picture with a specific value of nuh_layer_id layerId): There are no GDR pictures with nuh_layer_id equal to layerId between the previous IRAP picture (when present) and the specific picture in decoding order with nuh_layer_id equal to layerId.
[0350]
[0351] Completely random access (CRA) picture: An IRAP picture with nal_unit_type equal to CRA_NUT for each VCL NAL unit.
[0352] Completely random access (CRA) sub-picture:
[0353] Gradual Decoding Refresh (GDR) AU: An AU in which each layer in the CVS has a PU and each codec picture in the current PU is a GDR picture.
[0354] Gradual Decoding Refresh (GDR) picture: A picture in which the nal_unit_type of each VCL NAL unit is equal to GDR_NUT.
[0355] Gradual Decoding Refresh (GDR) sub-picture:
[0356] Instantaneous Decoding Refresh (IDR) picture: An IRAP picture with nal_unit_type equal to IDR_W_RADL or IDR_N_LP for each VCL NAL unit.
[0357]
[0358] Intra Random Access Point (IRAP) picture: A picture whose all VCL NAL units have the same nal_unit_type value in the range of IDR_W_RADL to CRA_NUT, inclusive.
[0359]
[0360] Leading picture: A picture that precedes the associated IRAP picture in output order.
[0361]
[0362] Output order:
[0363] Random Access Decodable Leading (RADL) picture: A picture with nal_unit_type equal to RADL_NUT for each VCL NAL unit.
[0364]
[0365] Random Access Skip Leading (RASL) picture: A picture with nal_unit_type equal to RASL_NUT for each VCL NAL unit.
[0366]
[0367] Step-by-step temporal sub-layer access (STSA) picture: A picture with nal_unit_type equal to STSA_NUT for each VCL NAL unit.
[0368]
[0369] Trailer picture: A picture whose nal_unit_type of each VCL NAL unit is equal to TRAIL_NUT.
[0370] NOTE: The trailing picture associated with an IRAP or GDR picture also follows the IRAP or GDR picture in decoding order. Pictures that follow the associated IRAP or GDR picture in output order and precede the associated IRAP or GDR picture in decoding order are not allowed.
[0371]
[0372] NOTE: The trailing picture associated with an IRAP or GDR sub-picture also follows the IRAP or GDR sub-picture in decoding order. Sub-pictures that follow the associated IRAP or GDR sub-picture in output order and precede the associated IRAP or GDR sub-picture in decoding order are not allowed. ...
[0374] 7.4.2.2 NAL unit header semantics ...
[0376] nal_unit_type specifies the NAL unit type, ie, the type of RBSP data structure contained in the NAL unit as specified in Table 5.
[0377] NAL units with a nal_unit_type in the range UNSPEC_28..UNSPEC_31, inclusive, have unspecified semantics and should not affect the decoding process specified in this specification.
[0378] NOTE 2 NAL unit types in the range of UNSPEC_28..UNSPEC_31 can be used as determined by the application. No decoding process is specified in this Specification for these values of nal_unit_type. Since different applications can use these NAL unit types for different purposes, special care must be taken when designing encoders that generate NAL units with these nal_unit_type values, and when designing decoders that interpret the content of NAL units with these nal_unit_type values. This Specification does not define any management of these values. These nal_unit_type values can only be appropriate in contexts where "conflict" (i.e., different definitions of the meaning of the content of NAL units of the same nal_unit_type value) is either unimportant, or impossible, or managed (e.g., defined or managed in a control application or transport specification, or managed by the environment that controls the distribution of the bitstream).
[0379] For purposes other than determining the amount of data in DUs of the bitstream (as specified in Annex C), decoders shall ignore (remove from the bitstream and discard) the content of all NAL units that use reserved values of nal_unit_type.
[0380] NOTE 3 This requirement allows for the future definition of compatible extensions of this Specification.
[0381] Table 5 - NAL unit type codes and NAL unit type categories
[0382]
[0383]
[0384] NOTE 4 A clean random access (CRA) picture can have an associated RASL or RADL picture present in the bitstream.
[0385] NOTE 5 An instantaneous decoding refresh (IDR) picture with nal_unit_type equal to IDR N LP has no associated leading picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR W RADL has no associated RASL picture present in the bitstream, but can have an associated RADL picture in the bitstream.
[0386] The value of nal_unit_type shall be the same for all VCL NAL units in a subpicture. A subpicture is considered to have the same NAL unit type as the VCL NAL units of the subpicture.
[0387]
[0388] For any particular picture's VCL NAL unit, the following applies:
[0389] If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and a picture or PU is considered to have the same NAL unit type as the VCL NAL units of the picture or PU.
[0390] Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), the picture shall have at least two sub-pictures, and the VCL NAL units of the picture shall have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one sub-picture of the picture shall all have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP, or CRA_NUT, while the VCL NAL units of the other sub-pictures in the picture shall all have a different specific value of nal_unit_type equal to TRAIL_NUT, RADL_NUT, or RASL_NUT.
[0391] The requirements for bitstream conformance are that the following constraints apply:
[0392] The trailer picture shall follow the associated IRAP or GDR picture in output order.
[0393]
[0394] When a picture is the leading picture of an IRAP picture, it shall be a RADL or RASL picture.
[0395]
[0396] There shall be no RASL pictures associated with an IDR picture in the bitstream.
[0397] There shall not be any RASL sub-pictures associated with an IDR sub-picture in the bitstream.
[0398] There shall be no RADL pictures in the bitstream that are associated with an IDR picture with nal_unit_type equal to IDR_N_LP.
[0399] NOTE 6 Random access can be performed at the location of the IRAP PU (and the IRAP picture and all subsequent non-RASL pictures are correctly decoded in decoding order) by discarding all PUs before the IRAP PU, provided that each parameter set (in the bitstream or by external means not specified in this Specification) is available when referenced.
[0400]
[0401] Any picture with nuh layer id equal to a particular value layerld that precedes the IRAP picture with nuh layer id equal to layerld in decoding order shall precede the IRAP picture and all its associated RADL pictures in output order.
[0402]
[0403] Any picture with nuh layer id equal to a particular value layerld that precedes the IRAP picture with nuh layer id equal to layerld in decoding order shall precede the IRAP picture and all its associated RADL pictures in output order.
[0404]
[0405] Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order.
[0406]
[0407] Any RASL picture associated with a CRA picture shall precede any IRAP picture that precedes the CRA picture in decoding order in output order.
[0408]
[0409] If field_seq_flag is equal to 0 and the current picture with nuh layer id equal to a particular value layerld is a leading picture associated with an IRAP picture, it shall precede all non-leading pictures associated with the same IRAP picture in decoding order. Otherwise, let picA and picB be the first and last leading pictures associated with an IRAP picture in decoding order, there shall be at most one non-leading picture with nuh layer id equal to layerld that precedes picA in decoding order, and there shall be no non-leading picture with nuh layer id equal to layerld between picA and picB in decoding order.
[0410]
[0411] 7.4.3.4 Picture parameter set semantics ...
[0413] mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture of the reference PPS has more than one VCL NAL unit, The VCL NAL units do not have the same nal_unit_type value [[and the picture is not an IRAP picture]]. mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture of the referenced PPS has one or more VCL NAL units, and the VCL NAL units of each picture of the referenced PPS have the same nal_unit_type value.
[0414] When no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_types_in_pic_flag shall be equal to 0.
[0415] [[For each slice in a picture picA having a nal_unit_type value nalUnitTypeA in the range IDR_W_RADL to CRA_NUT, inclusive (i.e., the value of mixed_nalu_types_in_pic_flag for picture picA is equal to 1), that picture picA also contains one or more slices having another nal_unit_type value, the following applies:
[0416] The slice shall belong to the sub-picture subpicA whose corresponding subpic_treated_as_pic_flag[i] value is equal to 1.
[0417] A slice shall not belong to a sub-picture of picA containing a VCL NAL unit with nal_unit_type not equal to nalUnitTypeA.
[0418] If nalUnitTypeA is equal to CRA, then for all subsequent PUs that follow the current picture in CLVS in decoding order and output order, the RefPicList[0] and RefPicList[1] of the slices in subpicA in these PUs shall not include any pictures that precede picA in decoding order among the valid entries.
[0419] Otherwise (i.e., nalUnitTypeA is equal to IDR_W_RADL or IDR_N_LP), for all PUs in the CLVS that follow the current picture in decoding order, neither the RefPicList[0] nor the RefPicList[1] for the slices in subpicA in those PUs shall include any pictures that precede picA in decoding order among the valid entries. ....
[0421] 7.4.3.7 Image Header Structure Semantics ...
[0423] recovery_poc_cnt specifies the recovery point of the decoded picture in output order.
[0424]
[0425]
[0426] If the current picture is a GDR picture [[associated with PH]] and there is a picture after the current GDR picture in decoding order in CLVS with PicOrderCntVal equal to If the value of [[PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt]] is the picture picA, then the picture picA is called the recovery point picture. Otherwise, PicOrderCntVal is greater than The first picture in output order that is [[the current picture's PicOrderCntVal plus the value of recovery_poc_cnt]] is called the recovery point picture. The recovery point picture should not precede the current GDR picture in decoding order. The value of recovery_poc_cnt shall be in the range of 0 to MaxPicOrderCntLsb-1, inclusive.
[0427] [[When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:
[0428] RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt(81)]]
[0429] NOTE 2 When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to that of the associated GDR picture, [[RpPicOrderCntVal]], the current and subsequent decoded pictures in output order exactly match the corresponding pictures produced by starting the decoding process from the previous IRAP picture (when present) that precedes the associated GDR picture in decoding order. ...
[0431] 8.3.2 Decoding Process of Reference Picture List Construction ...
[0433] The requirements for bitstream conformance are that the following constraints apply:
[0434] For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] should not be less than NumRefIdxActive[i].
[0435] The picture referred to by each valid entry in RefPicList[0] or RefPicList[1] shall exist in the DPB and shall have a TemporalId less than or equal to the TemporalId of the current picture.
[0436] The picture referred to by each entry in RefPicList[0] or RefPicList[1] shall not be the current picture, and non_reference_picture_flag shall be equal to 0. A STRP entry in RefPicList[0] or RefPicList[1] of a slice of a picture and an LTRP entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture.
[0437] There shall be no LTRP entry in RefPicList[0] or RefPicList[1] for which the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referred to by the entry is greater than or equal to 2 24 .
[0438] Let setOfRefPics be the set of unique pictures referred to by all entries in RefPicList[0] with the same nuh_layer_id as the current picture and all entries in RefPicList[1] with the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize - 1 (including itself), where MaxDpbSize is specified in clause A.4.2, and setOfRefPics shall be the same for all slices of the picture.
[0439] When nal_unit_type of the current slice is equal to STSA_NUT, there shall be no valid entry in RefPicList[0] or RefPicList[1] with TemporalId equal to TemporalId of the current picture and nuh_layer_id equal to nuh_layer_id of the current picture.
[0440] When the current picture is a picture that follows the STSA picture in decoding order and has a TemporalId equal to the TemporalId of the current picture and a nuh_layer_id equal to the nuh_layer_id of the current picture, there should be no pictures that precede the STSA picture in decoding order and have a TemporalId equal to the TemporalId of the current picture and have a nuh_layer_id equal to the nuh_layer_id of the current picture included as valid entries in RefPicList[0] or RefPicList[1].
[0441]
[0442] When the current picture with numh_layer_id equal to the specific value layerId is a CRA picture, there shall not be any pictures referred to by entries in RefPicList[0] or RefPicList[1] that precede any previous IRAP picture (when present) with numh_layer_id equal to the specific value layerId in output order or decoding order.
[0443]
[0444] When the current picture with nuh_layer_id equal to a specific value layerId is not a RASL picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId, there should not be any pictures generated by the decoding process used to generate unusable reference pictures referred to by valid entries in RefPicList[0] or RefPicList[1].
[0445]
[0446] The picture generated by the decoding process used to generate the unusable reference picture referred to by the entry in RefPicList[0] or RefPicList[1] shall not be present when the current sub-picture with nuh_layer_id equal to the particular value layerId is not a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a leading picture associated with the same CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a leading picture associated with a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, or a recovery picture of a GDR picture with NoOutputBeforeRecoveryFlag equal to 1 and nuh_layer_id equal to layerId.
[0447]
[0448] When the current picture is associated with an IRAP picture and follows the IRAP picture in output order, there shall be no pictures referred to by valid entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.
[0449]
[0450] When the current picture is associated with an IRAP picture, follows the IRAP picture in output order, and follows the leading picture (if any) associated with the same IRAP picture in both decoding order and output order, there shall not be any pictures referred to by entries in RefPicList[0] or RefPicList[1] that precede the associated IRAP picture in output order or decoding order.
[0451]
[0452] When the current picture is a RADL picture, there shall not be any valid entries in RefPicList[0] or RefPicList[1]:
[0453] οRASL pictures
[0454] o pictures that precede the associated IRAP picture in decoding order
[0455]
[0456] o
[0457] o
[0458] Each ILRP entry in the RefPicList[0] or RefPicList[1] of the slice of the current picture shall refer to a picture in the same AU as the current picture.
[0459] The picture referred to by each ILRP entry in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB and shall have a nuh_layer_id that is less than the nuh_layer_id of the current picture.
[0460] Each ILRP entry in the stripe's RefPicList[0] or RefPicList[1] shall be a valid entry.
[0461] Figure 5 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, for example, 8 or 10 bit multi-component pixel values, or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical networks (PONs), etc.), and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0462] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. As represented by component 1906, the output of codec component 1904 can be stored or sent via a connected communication. Component 1908 can use the bitstream (or codec) representation of the video received at input 1902, which is stored or communicated, to generate pixel values or displayable video that is sent to display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations, which are opposite to the encoding results, will be performed by the decoder.
[0463] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document may be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0464] Figure 6 36 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be implemented in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described in this document. Memory(s) 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0465] Figure 8 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.
[0466] like Figure 8 As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0467] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0468] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0469] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0470] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, with the destination device 120 being configured to interface with an external display device.
[0471] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or future standards.
[0472] Figure 9 is a block diagram illustrating an example of a video encoder 200, which may be Figure 8 The video encoder 114 in the system 100 is shown.
[0473] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 9 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0474] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0475] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.
[0476] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the purpose of explanation, they are not shown in FIG. Figure 9 are represented separately in the example.
[0477] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0478] The mode selection unit 203 may, for example, select one of the coding modes (intra or inter) based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select the precision of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0479] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from the buffer 213.
[0480] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0481] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0482] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 containing the reference video block and the motion vector indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0483] In some examples, motion estimation unit 204 may output the entire motion information set for use in the decoding process of the decoder.
[0484] In some examples, motion estimation unit 204 may not output the entire set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for the neighboring video block.
[0485] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0486] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0487] As described above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0488] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0489] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0490] In other examples, there may be no residual data for the current video block, eg, in skip mode, and the residual generation unit 207 may not perform a subtraction operation.
[0491] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0492] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0493] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0494] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0495] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0496] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown.
[0497] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 10 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0498] exist Figure 10 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those generally performed for the video encoder 200 ( Figure 9 ) is a decoding process that is the inverse of the encoding process described.
[0499] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and merge mode.
[0500] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0501] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.
[0502] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is coded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the coded video sequence.
[0503] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0504] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0505] A list of examples of preferred embodiments is provided below.
[0506] The first set of items illustrates example embodiments of the techniques discussed in the previous section (eg, item 1).
[0507] 1. A video processing method (e.g., Figure 7 ), comprising: performing (702) a conversion between a video comprising one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies that the one or more pictures including the one or more sub-pictures are included in a codec representation based on network abstraction layer (NAL) units, wherein a type of NAL unit indicated in the codec representation includes a codec slice of a particular type of picture or a codec slice of a particular type of sub-picture.
[0508] The following items illustrate example embodiments of the techniques discussed in the previous section (eg, item 2).
[0509] 2. A video processing method comprising: performing a conversion between a video comprising one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that specifies that two adjacent sub-pictures having different network abstraction layer unit types will have the same indication of sub-pictures that are considered to be picture flags.
[0510] The following items illustrate example embodiments of the techniques discussed in the previous section (eg, items 4, 5, 6, 7, 9, 1, 11, 12).
[0511] 3. A video processing method, comprising: performing conversion between a video comprising one or more pictures including one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule, the format rule defining an order of sub-pictures of a first type and sub-pictures of a second type, wherein the first sub-picture is a trailing sub-picture or a leading sub-picture or a random access skip leading (RASL) sub-picture type, and the second sub-picture is a RASL type or a random access decodable leading (RADL) type or an instantaneous decoding refresh (IDR) type or a gradual decoding refresh (GDR) type sub-picture.
[0512] 4. A method according to clause 3, wherein the rule specifies that the tail sub-picture follows the associated intra random access point or GDR sub-picture in output order.
[0513] 5. A method as described in clause 3, wherein the rule specifies that when the picture is a leading picture of an intra random access point picture, the nal_unit_type value of all network abstraction layer units in the picture is equal to RADL_NUT or RASL_NUT.
[0514] 6. A method as described in clause 3, wherein the rule specifies that a given sub-picture that is a leading sub-picture of an IRAP sub-picture must also be a RADL or RASL sub-picture.
[0515] 7. A method as described in clause 3, wherein the rule specifies that a given sub-picture that is a RASL sub-picture is not allowed to be associated with an IDR sub-picture.
[0516] 8. A method as described in clause 3, wherein the rule specifies that a given sub-picture with the same layer id and sub-picture index as an IRAP sub-picture must precede the IRAP sub-picture and all its associated RADL sub-pictures in output order.
[0517] 9. A method according to clause 3, wherein the rule specifies that a given sub-picture with the same layer id and sub-picture index as a GDR sub-picture must precede the GDR sub-picture and all its associated RADL sub-pictures in output order.
[0518] 10. A method as recited in clause 3, wherein the rule specifies that a given sub-picture that is a RASL sub-picture associated with a CRA sub-picture precedes all RADL sub-pictures associated with the CRA sub-picture in output order.
[0519] 11. A method as described in clause 3, wherein the rule specifies that a given sub-picture that is a RASL sub-picture associated with a CRA sub-picture precedes all IRAP sub-pictures associated with the CRA sub-picture in output order.
[0520] 12. A method as described in clause 3, wherein the rule specifies that if a given sub-picture is a leading sub-picture associated with an IRAP sub-picture, then the given sub-picture precedes all non-leading sub-pictures associated with the IRAP sub-picture in decoding order.
[0521] The following items illustrate example embodiments of the techniques discussed in the previous section (eg, items 8, 14, 15).
[0522] 13. A video processing method comprising: performing a conversion between a video comprising one or more pictures containing one or more sub-pictures and a codec representation of the video, wherein the codec representation conforms to a format rule that defines conditions under which a first type of sub-picture is allowed or not allowed to appear together with a second type of sub-picture.
[0523] 14. A method according to clause 13, wherein the rule specifies that in the presence of an IDR sub-picture of the network abstraction layer type IDR_N_LP, then the codec representation is not allowed to have a RADP sub-picture.
[0524] 15. A method according to clause 13, wherein the rule does not allow a picture to be included in a reference list of a picture containing a step-by-step temporal sub-layer access (STSA) sub-picture such that the picture precedes the picture containing the STSA sub-picture.
[0525] 16. A method as described in clause 13, wherein the rule does not allow a picture to be included in a reference list of a picture containing an intra random access point (IRAP) sub-picture such that the picture precedes the picture containing the IRAP sub-picture.
[0526] 17. A method according to any of clauses 1 to 16, wherein the converting comprises encoding the video into a codec representation.
[0527] 18. A method according to any of clauses 1 to 16, wherein the converting comprises decoding the codec representation to generate pixel values of the video.
[0528] 19. A video decoding apparatus comprising a processor configured to implement the method of one or more of clauses 1 to 18.
[0529] 20. A video encoding apparatus comprising a processor configured to implement the method of one or more of clauses 1 to 18.
[0530] 21. A computer program product having computer code stored thereon which, when executed by a processor, causes the processor to carry out the method of any one of clauses 1 to 18.
[0531] 22. The method, apparatus, or system described in this document.
[0532] The second set of items illustrates example embodiments of the techniques discussed in the previous section (eg, items 8-11).
[0533] 1. A method for video processing (e.g., Figure 11 The method 1100 shown comprises performing 1102 conversion between a video of multiple layers including one or more pictures and a bitstream of the video according to a format rule, and wherein the format rule specifies that the reference picture referred to by each inter-layer reference picture entry in the reference picture list of the slice of the current picture of the current layer satisfies a constraint, wherein the constraint is at least one of: (a) the reference picture is an intra random access (IRAP) picture, or (b) the reference picture has a temporal identifier less than or equal to a specific value, the specific value being a maximum allowed value of the video layer that can be referenced by the slice of the current layer, wherein the maximum allowed value is indicated in a syntax element.
[0534] 2. A method according to clause 1, wherein the one or more indices of the syntax element are layer indices.
[0535] 3. A method according to clause 1 or 2, wherein the syntax element is a two-dimensional syntax element.
[0536] 4. A method according to any of clauses 1 to 3, wherein the temporal identifier is TemporalId and the syntax element is max_tid_il_ref_pics_plus1[i][j], where i and j are integers.
[0537] 5. A method according to any of clauses 1 to 3, wherein the specific value corresponds to the value of Max(0, max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively, GeneralLayerIdx[] specifies the layer index of the layer with the specific layer identifier corresponding to nuh_layer_id, and refpicLayerId is the layer identifier of the picture.
[0538] 6. A method according to clause 5, wherein the format rule further specifies that the picture referred to by each inter-layer reference picture entry is present in the decoded picture buffer and has a layer identifier that is less than the layer identifier of the current picture.
[0539] 7. A method according to any of clauses 1 to 3, wherein the specific value corresponds to the value of max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx], where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively, GeneralLayerIdx[] specifies the layer index of the layer with the specific layer identifier corresponding to nuh_layer_id, and refpicLayerId is the layer identifier of the picture.
[0540] 8. A method according to clause 7, wherein the rule further specifies that the picture referred to by each inter-layer reference picture entry is present in the decoded picture buffer and has a layer identifier that is less than the layer identifier of the current picture.
[0541] 9. A method according to any one of clauses 1 to 8, wherein the converting comprises encoding the video into a bitstream.
[0542] 10. A method according to any of clauses 1 to 8, wherein the converting comprises decoding the video from a bitstream.
[0543] 11. The method of any one of clauses 1 to 8, wherein the converting comprises generating a bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0544] 12. A video processing apparatus comprising a processor configured to implement the method of any one or more of clauses 1 to 11.
[0545] 13. A method of storing a video bitstream, comprising the method of any one of clauses 1 to 11, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0546] 14. A computer-readable medium storing program code which, when executed, causes a processor to implement the method of any one or more of clauses 1 to 11.
[0547] 15. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0548] 16. A video processing device for storing a bitstream, wherein the video processing device is configured to implement the method of any one or more of clauses 1 to 11.
[0549] The third group of items illustrates example embodiments of the techniques discussed in the previous section (eg, item 12).
[0550] 1. A method for video processing (e.g., Figure 12 The method 1200 shown in FIG. 1 comprises performing 1202 conversion between a video including a video picture and a bitstream according to a format rule, wherein the format rule specifies using one or more of a first syntax structure indicating a position of a vertical virtual boundary and a second syntax structure indicating a position of a horizontal virtual boundary to indicate a virtual boundary of a video area of the video picture, wherein the format rule specifies that a numerical value indicated by a first syntax element in the first syntax structure and a second syntax element in the second syntax structure is 1 less than the position of the vertical virtual boundary and the horizontal virtual boundary counted from the upper left position of the video picture.
[0551] 2. A method as recited in clause 1, wherein the first syntax element comprises sps_virtual_boundary_pos_x_minus1[i] and the second syntax element comprises sps_virtual_boundary_pos_y_minus1[i], where i is an integer.
[0552] 3. A method as described in clause 1, wherein the first syntax element comprises ph_virtual_boundary_pos_x_minus1[i] and the second syntax element comprises ph_virtual_boundary_pos_y_minus1[i], where i is an integer.
[0553] 4. A method according to any of clauses 1 to 3, wherein at least one of the first syntax element and the second syntax element is encoded or decoded using a descriptor, the descriptor being ue(v) or u(v).
[0554] 5. A method according to any of clauses 1-3, wherein the first syntax element comprises sps_virtual_boundary_pos_x_minus1[i] plus 1, sps_virtual_boundary_pos_x_minus1[i] plus 1 specifies the position of the i-th vertical virtual boundary in units of luma samples divided by 8, and has a value in the range of 0 to Ceil(pic_width_max_in_luma_samples ÷ 8) - 2, where pic_width_max_in_luma_samples indicates the maximum width of the picture in units of luma samples.
[0555] 6. A method according to any of clauses 1-3, wherein the second syntax element comprises sps_virtual_boundary_pos_y_minus1[i] plus 1, sps_virtual_boundary_pos_y_minus1[i] plus 1 specifies the position of the i-th horizontal virtual boundary in units of luma samples divided by 8, and has a value in the range of 0 to Ceil(pic_height_max_in_luma_samples ÷ 8) - 2, where pic_height_max_in_luma_samples indicates the maximum height of the picture in units of luma samples.
[0556] 7. A method according to any of clauses 1-3, wherein the first syntax element comprises ph_virtual_boundary_pos_x_minus1[i] plus 1, ph_virtual_boundary_pos_x_minus1[i] plus 1 specifies the position of the i-th vertical virtual boundary in units of luma samples divided by 8, and has a value in the range of 0 to Ceil(pic_width_in_luma_samples ÷ 8) - 2, where pic_width_in_luma_samples indicates the width of the picture in units of luma samples.
[0557] 8. A method according to clause 7, wherein, for i ranging from 0 to NumVerVirtualBoundaries-1, the list VirtualBoundariesPosX[i] specifying the positions of the vertical virtual boundaries in units of luma samples is derived using the following condition: VirtualBoundariesPosX[i] = (sps_virtual_boundaries_present_flag?(sps_virtual_boundary_pos_x_minus1[i]+1):(ph_virtual_boundary_pos_x_minus1[i]+1))*8.
[0558] 9. The method of clause 8, wherein the distance between any two vertical virtual boundaries is greater than or equal to CtbSizeY luma samples, where CtbSizeY indicates the size of the codec tree unit.
[0559] 10. A method according to any of clauses 1-3, wherein the second syntax element comprises ph_virtual_boundary_pos_y_minus1[i] plus 1 specifying the position of the i-th horizontal virtual boundary in units of luma samples divided by 8, and has a value in the range of 0 to Ceil(pic_height_in_luma_samples ÷ 8)-2, where pic_height_in_luma_samples indicates the height of the picture in units of luma samples.
[0560] 11. A method according to clause 10, wherein, for i ranging from 0 to NumVerVirtualBoundaries-1, the list VirtualBoundariesPosX[i] specifying the positions of the vertical virtual boundaries in units of luma samples is derived using the following condition: VirtualBoundariesPosY[i] = (sps_virtual_boundaries_present_flag?(sps_virtual_boundary_pos_y_minus1[i]+1):(ph_virtual_boundary_pos_y_minus1[i]+1))*8.
[0561] 12. The method of clause 11, wherein the distance between any two horizontal virtual boundaries is greater than or equal to CtbSizeY luma samples, where CtbSizeY indicates the size of the codec tree unit.
[0562] 13. A method according to any of clauses 1 to 12, wherein the vertical virtual boundary and the horizontal virtual boundary correspond to boundaries between refreshed and non-refreshed areas within the video picture.
[0563] 14. A method according to any of clauses 1 to 13, wherein the vertical virtual boundaries and the horizontal virtual boundaries correspond to boundaries at which no loop filtering is applied during conversion.
[0564] 15. A method according to any of clauses 1 to 14, wherein the converting comprises encoding the video into a bitstream.
[0565] 16. A method according to any of clauses 1 to 14, wherein the converting comprises decoding the video from a bitstream.
[0566] 17. The method of clauses 1 to 14, wherein the converting comprises generating a bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0567] 18. A video processing apparatus comprising a processor configured to implement the method of any one or more of clauses 1 to 17.
[0568] 19. A method of storing a video bitstream, comprising the method of any one of clauses 1 to 17, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0569] 20. A computer-readable medium storing program code which, when executed, causes a processor to implement the method of any one or more of clauses 1 to 17.
[0570] 21. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0571] 22. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method of any one or more of clauses 1 to 17.
[0572] In the provisions described herein, an encoder may conform to a format rule by generating a codec representation according to the format rule. In the provisions described herein, a decoder may use the format rule to parse syntax elements of the codec representation using knowledge of the presence and absence of syntax elements according to the format rule to produce decoded video.
[0573] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa, a video compression algorithm may be applied. As defined by the syntax, the bitstream representation of the current video block may (for example) correspond to bits that are co-located or scattered at different locations within the bitstream. For example, a macroblock may be encoded based on error residual values from the transform and encoding, and also using bits from the header and other fields in the bitstream. Furthermore, during conversion, the decoder may parse the bitstream knowing that some fields may or may not be present based on this determination, as described in the solution above. Similarly, the encoder may determine whether to include or not include certain syntax fields, and generate the codec representation accordingly by including or excluding the syntax fields from the codec representation.
[0574] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of materials that implement a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0575] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one location or distributed across multiple locations and interconnected by a communications network.
[0576] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0577] By way of example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, to receive data from or transfer data to the mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0578] While this patent document contains many details, these should not be construed as limiting the scope of any subject matter or of any appended claims in which exclusive rights are claimed. Rather, these details are included for the purpose of describing particular embodiments of particular technology in a way that enables others working in the art to not only make and use them, but also to understand the particular features of the technology so they can be worked around. Certain features described in the context of separate embodiments in this patent document can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0579] Similarly, while operations are described in a particular order in the drawings, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0580] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Perform conversion between video including video pictures and bitstream according to format rules, wherein the format rule specifies that a virtual boundary of a video area of a video picture is indicated using one or more of a first syntax structure indicating a position of a vertical virtual boundary and a second syntax structure indicating a position of a horizontal virtual boundary, The format rule stipulates that the values indicated by the first syntax element in the first syntax structure and the second syntax element in the second syntax structure are 1 less than the positions of the vertical virtual boundary and the horizontal virtual boundary counted from the upper left position of the video picture.
2. The method according to claim 1, wherein The first syntax element includes sps_virtual_boundary_pos_x_minus1[i], and the second syntax element includes sps_virtual_boundary_pos_y_minus1[i], where i is an integer.
3. The method according to claim 1, wherein The first syntax element includes ph_virtual_boundary_pos_x_minus1[i], and the second syntax element includes ph_virtual_boundary_pos_y_minus1[i], where i is an integer.
4. The method according to any one of claims 1 to 3, wherein At least one of the first syntax element and the second syntax element is encoded and decoded using a descriptor, the descriptor being ue(v) or u(v).
5. The method according to any one of claims 1 to 3, wherein The first syntax element includes sps_virtual_boundary_pos_x_minus1[i] plus 1, which specifies the position of the i-th vertical virtual boundary in units of luma samples divided by 8 and has a value in the range of 0 to Ceil(pic_width_max_in_luma_samples÷8)-2, where pic_width_max_in_luma_samples indicates the maximum width of the picture in units of luma samples.
6. The method according to any one of claims 1 to 3, wherein The second syntax element includes sps_virtual_boundary_pos_y_minus1[i] plus 1, which specifies the position of the i-th horizontal virtual boundary in units of luma samples divided by 8 and has a value in the range of 0 to Ceil(pic_height_max_in_luma_samples÷8)-2, where pic_height_max_in_luma_samples indicates the maximum height of the picture in units of luma samples.
7. The method according to any one of claims 1 to 3, wherein The first syntax element includes ph_virtual_boundary_pos_x_minus1[i] plus 1, which specifies the position of the i-th vertical virtual boundary in units of luma samples divided by 8 and has a value in the range of 0 to Ceil(pic_width_in_luma_samples÷8)-2, where pic_width_in_luma_samples indicates the width of the picture in units of luma samples.
8. The method according to claim 7, wherein: For i ranging from 0 to NumVerVirtualBoundaries-1, the list VirtualBoundariesPosX[i] specifying the positions of vertical virtual boundaries in units of luma samples is derived using the following condition: VirtualBoundariesPosX[i]=(sps_virtual_boundaries_present_flag?(sps_virtual_boundary_pos_x_minus1[i]+1):(ph_virtual_boundary_pos_x_minus1[i]+1))*8.
9. The method according to claim 8, wherein The distance between any two vertical virtual boundaries is greater than or equal to CtbSizeY luma samples, where CtbSizeY indicates the size of the codec tree unit.
10. The method according to any one of claims 1 to 3, wherein: The second syntax element includes ph_virtual_boundary_pos_y_minus1[i] plus 1 to specify the position of the i-th horizontal virtual boundary in units of luma samples divided by 8, and has a value in the range of 0 to Ceil(pic_height_in_luma_samples÷8)-2, where pic_height_in_luma_samples indicates the height of the picture in units of luma samples.
11. The method according to claim 10, wherein: For i ranging from 0 to NumVerVirtualBoundaries-1, the list VirtualBoundariesPosY[i] specifying the positions of horizontal virtual boundaries in units of luma samples is derived using the following condition: VirtualBoundariesPosY[i]=(sps_virtual_boundaries_present_flag?(sps_virtual_boundary_pos_y_minus1[i]+1):(ph_virtual_boundary_pos_y_minus1[i]+1))*8.
12. The method according to claim 11, wherein The distance between any two horizontal virtual boundaries is greater than or equal to CtbSizeY luma samples, where CtbSizeY indicates the size of the codec tree unit.
13. The method according to any one of claims 1 to 3, 8, 9, 11 and 12, wherein The vertical virtual boundary and the horizontal virtual boundary correspond to boundaries between a refreshed area and a non-refreshed area within the video picture.
14. The method according to any one of claims 1 to 3, 8, 9, 11 and 12, wherein The vertical virtual boundary and the horizontal virtual boundary correspond to boundaries where loop filtering is not applied during the conversion.
15. The method according to any one of claims 1 to 3, 8, 9, 11 and 12, wherein The converting includes encoding the video into the bitstream.
16. The method according to any one of claims 1 to 3, 8, 9, 11 and 12, wherein The converting includes decoding the video from the bitstream.
17. The method according to any one of claims 1 to 3, 8, 9, 11 and 12, wherein The converting includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. 18 . A video processing device, comprising a processor, wherein the processor is configured to implement the method according to claim 1 .
19. A computer-readable medium storing program code, which, when executed, causes a processor to implement the method of any one of claims 1 to 17.
Citation Information
Patent Citations
Loop filtering method, loop filtering device, electronic equipment and readable medium
CN109600611A
Image filter device, image decoding device, and image encoding device
WO2019131400A1