Signaling notification of the existence of interlayer reference pictures
By modifying the signaling design of sub-pictures, slices and strips in VVC, the number of sub-pictures, unclear signaling notifications of sub-picture IDs and bitstream consistency are solved, and the effect of supporting the consistency of more than 256 sub-pictures and bitstreams is achieved.
Patent Information
- Application Number
- CN202180008798.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-09
- Filing Date
- 2021-01-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-01-08
AI Technical Summary
There are many problems in the signaling design of sub-pictures, slices and stripes in VVC, including too small limit on the number of sub-pictures, unclear signaling notification of sub-picture ID, insufficient bitstream consistency requirements, etc., which leads to the inability to meet the maximum number of sub-pictures and bitstream consistency problems in some applications.
By modifying the encoding method of sps_num_subpics_minus1, it supports more than 256 sub-pictures; adjusting the signaling notification conditions of sub-picture IDs to ensure explicit signaling notification of sub-picture IDs in SPS or PPS; adding syntax elements to clarify the sub-picture layout and striping mode to ensure bitstream consistency.
It supports encoding of more than 256 sub-pictures in VVC, clarify the signaling notification method of sub-picture ID, ensure the consistency of the bitstream, and meet the wider application needs.
Smart Images

Figure CN114946174B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority and benefit of U.S. Provisional Patent Application No. US 62 / 959,108, filed on January 9, 2020, and International Patent Application No. PCT / US2021 / 012835, filed on January 8, 2021. All of the above - mentioned patent applications are hereby incorporated by reference in their entirety. Technical Field
[0003] This application document relates to image and video encoding and decoding. Background Art
[0004] Digital video occupies the largest bandwidth usage on the Internet and other digital communication networks. With the increase in the number of connected user devices capable of receiving and displaying video, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses a video encoder and a decoder for video encoding and decoding respectively, and includes constraints, limitations, and signaling for sub - pictures, slices, and strips.
[0006] In one example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video, where the bitstream includes one or more access units according to format rules, and where the format rules specify the order in which a first message and a second message for an operation point OP appear in the access unit AU such that the first message precedes the second message in the decoding order.
[0007] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video, where the bitstream includes one or more access units according to format rules, and where the format rules specify the order in which a plurality of messages for an operation point OP appear in the access unit such that a first message of the plurality of messages precedes a second message of the plurality of messages in the decoding order.
[0008] In another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including a picture and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify whether to signal at the start of a picture header associated with the picture an indication of a first flag, and where the first flag indicates whether the picture is an Intra Random Access Point (IRAP) picture or an Asymptotic Decoding Refresh (GDR) picture.
[0009] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules do not allow encoding or decoding of the pictures in the one or more pictures to include coded slice network abstraction layer (NAL) units having an asymptotic decoding refresh type, and is related to a flag indicating that the picture includes NAL units of a hybrid type.
[0010] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules allow encoding or decoding of the pictures in the one or more pictures to include coded slice network abstraction layer (NAL) units having an asymptotic decoding refresh type, and is related to a flag indicating that the picture does not include NAL units of a hybrid type.
[0011] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to format rules, the format rules specifying whether to signal a first syntax element in a picture parameter set (PPS) related to the picture, wherein the picture includes one or more slices having a slice type, and wherein the first syntax element indicates signaling of the slice type in a picture header if the first syntax element is equal to 0, otherwise the first syntax element indicates signaling of the slice type in a slice header.
[0012] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a picture of a video and a bitstream of the video according to rules, wherein the conversion includes a loop filtering process, and wherein the rules specify that the total number of vertical virtual boundaries and the total number of horizontal virtual boundaries related to the loop filtering operation are signaled at a picture level or a sequence level.
[0013] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules conditionally allow encoding or decoding of pictures within one layer by using reference pictures from other layers based on a first syntax element, the first syntax element indicating whether the reference pictures from the other layer exist in the bitstream, and wherein the first syntax element is conditionally signaled in the bitstream based on a second syntax element, the second syntax element indicating whether an identifier of a parameter set related to the picture is not equal to 0.
[0014] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a picture of a video and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify a first syntax element that causes (a) synchronization processing of context variables before decoding a coding tree unit (CTU) in the picture and (b) storage processing of the context variables after decoding the CTU, and where the first syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
[0015] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a picture of a video and a bitstream of the video, where the bitstream conforms to format rules, where the format rules specify a syntax element that indicates whether there is signaling for an entry point offset for a slice or slice-specific coding tree unit (CTU) row in a slice header of the picture, and where the syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
[0016] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules specify that a first syntax element is less than a first preset threshold, and the first syntax element indicates that the number of parameters for an output layer set (OLS) hypothetical reference decoder (HRD) in a video parameter set (VPS) associated with the video is less than the first preset threshold.
[0017] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules specify that a syntax element is less than a preset threshold, and the syntax element indicates the number of profile / tier / level (PTL) syntax structures in a video parameter set (VPS) associated with the video.
[0018] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules specify a first syntax element that indicates that the number of decoded picture buffer parameter syntax structures in a video parameter set (VPS) must be less than or equal to a second syntax element, and the second syntax element indicates the number of layers specified by the VPS.
[0019] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules allow a decoder to obtain a network abstraction layer (NAL) unit of a terminal by signaling in the bitstream or by providing it through external means.
[0020] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video, wherein the bitstream conforms to format rules, and wherein, due to a syntax element being equal to 0, the format rules restrict each layer in the bitstream to include only one sub-picture, indicating that each layer is configured to use inter-layer prediction.
[0021] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between a video and a bitstream of the video according to rules, wherein the rules specify implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the rules specify that in the extraction process, when removing video coding layer (VCL) network abstraction layer (NAL) units, padding data units and padding supplementary enhancement information (SEI) messages in SEI VCL units associated with the VCL NAL units are also deleted.
[0022] In yet another example aspect, a video processing method is disclosed. The method includes: performing a conversion between video units of a video and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify that the bitstream includes a first syntax element that indicates whether the video units are coded in a lossy mode or a lossless mode, and wherein a second syntax element is signaled to indicate that escape samples in a palette mode applied to the video units are selectively included based on the value of the first syntax element.
[0023] In yet another example aspect, a video encoding device is disclosed. The video encoding device includes a processor, wherein the processor is configured to perform the above method.
[0024] In yet another example aspect, a video decoding device is disclosed. The video decoding device includes a processor, wherein the processor is configured to perform the above method.
[0025] In yet another example aspect, a computer-readable medium storing code is disclosed. The encoding and decoding implement the above method in the form of processor-executable code.
[0026] These and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 An example of partitioning a picture using luminance coding tree units (CTUs) is shown.
[0028] Figure 2Shows another example of segmenting a picture using luminance CTUs.
[0029] Figure 3 Shows an example of picture segmentation.
[0030] Figure 4 Shows another example of picture segmentation.
[0031] Figure 5 Is a block diagram of an example video processing system that can implement the disclosed technology.
[0032] Figure 6 Is a block diagram of an example hardware platform for video processing.
[0033] Figure 7 Is a block diagram illustrating a video codec system according to some embodiments of the present disclosure.
[0034] Figure 8 Is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0035] Figure 9 Is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0036] Figures 10 - 26 Shows a flowchart of an example method for video processing. Detailed implementation
[0037] In this document, section headings are used for ease of understanding, and the applicability of the technologies and embodiments disclosed in each section is not limited to that section. Additionally, in some descriptions, H.266 terms are used only for ease of understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described here are also applicable to other video codec protocols and designs.
[0038] 1. Introduction
[0039] This document relates to video codec technology. Specifically, regarding signaling for sub-pictures, slices, and strips. These concepts can be applied, either alone or in various combinations, to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the upcoming Versatile Video Coding (VVC).
[0040] 2. Abbreviations
[0041] APS Adaptive Parameter Set
[0042] AU Access Unit
[0043] AUD Access Unit Delimiter
[0044] AVC Advanced Video Coding
[0045] CLVS Coding and Decoding Layer Video Sequence
[0046] CPB Coding and Decoding Picture Buffer
[0047] CRA Clear Random Access
[0048] CTU Coding and Decoding Tree Unit
[0049] CVS Coding and Decoding Video Sequence
[0050] DPB Decoding Picture Buffer
[0051] DPS Decoding Parameter Set
[0052] EOB End of Bitstream
[0053] EOS End of Sequence
[0054] GDR Gradual Decoding Refresh
[0055] HEVC High Efficiency Video Coding
[0056] HRD Hypothetical Reference Decoder
[0057] IDR Instantaneous Decoding Refresh
[0058] JEM Joint Exploration Model
[0059] MCTS Motion Constrained Tile Set
[0060] NAL Network Abstraction Layer
[0061] OLS Output Layer Set
[0062] PH Picture Header
[0063] PPS Picture Parameter Set
[0064] PTL Profile, Tier and Level
[0065] PU Picture Unit
[0066] RBSP Raw Byte Sequence Payload
[0067] SEI Supplemental Enhancement Information
[0068] SPS Sequence Parameter Set
[0069] SVC Scalable Video Coding
[0070] VCL Video Coding and Decoding Layer
[0071] VPS Video Parameter Set
[0072] VTM VVC Test Model
[0073] VUI Video Usability Information
[0074] VVC Versatile Video Coding
[0075] 3. Preliminary Discussion
[0076] Video coding standards have mainly evolved by developing well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 video, and the two organizations jointly developed H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, which employs temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly simultaneously. The goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, when the first version of the VVC Test Model (VTM) was released. With continuous efforts in VVC standardization, new coding technologies have been adopted into the VVC standard at each JVET meeting. The VVC working draft and the test model VTM are updated after each meeting. The VVC project now aims for Feature Complete (FDIS) at the meeting in July 2020.
[0077] 3.1. Picture Partitioning Schemes in HEVC
[0078] HEVC includes four different picture partitioning schemes, namely regular slices, non-independent slices, tiles, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reducing end-to-end latency.
[0079] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, a regular slice can be reconstructed independently of other regular slices within the same picture (although there may still be dependencies due to loop filter operations).
[0080] Regular slices are the only tool available for parallelization, and this tool is also available in H.264 / AVC in almost the same form. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for the inter-processor or inter-core data sharing for motion compensation when decoding predicted coded pictures, which is usually much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, using regular slices may incur a large amount of coding and decoding overhead due to the bit cost of slice headers and the prediction loss across slice boundaries. In addition, due to the intra-picture independence of regular slices and each regular slice being encapsulated in its own NAL, regular slices (compared with other tools mentioned below) can also serve as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the requirements for slice layout in a picture from the goals of parallelization and MTU size matching are contradictory. The implementation of this situation has led to the development of the parallelization tools mentioned below.
[0081] Non-independent slices have short slice headers and allow partitioning of the bitstream at tree block boundaries without breaking any intra-picture prediction. Basically, non-independent slices divide a regular slice into multiple NAL units, reducing the end-to-end latency by allowing a part of the regular slice to be sent before the encoding of the entire regular slice is completed.
[0082] In WPP, a picture is segmented into single rows of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other segments. Parallel processing can be carried out through parallel decoding of CTB rows, where the decoding of a CTB row starts with a delay of two CTBs to ensure that data related to CTBs above and to the right of the main CTB can be obtained before the main CTB being decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores can be parallelized as there are CTB rows in the picture. Since intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication required for implementing intra-picture prediction may be substantial. Compared with not applying WPP segmentation, WPP segmentation does not result in the generation of additional NAL units, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular slices can be used together with WPP, but with a certain amount of coding and decoding overhead.
[0083] Slice definitions define the horizontal and vertical boundaries that divide a picture into slice columns and slice rows. Slice columns extend from the top to the bottom of the picture. Similarly, slice rows extend from the left to the right of the picture. The number of slices in a picture can be simply obtained by multiplying the number of slice columns by the number of slice rows.
[0084] Before decoding the top-left CTB of the next slice in the raster scan order of a picture, the scan order of CTBs is changed to the local scan order within the slice (in the order of the CTB raster scan of the slice). Similar to a regular strip, a slice breaks the prediction dependency and entropy decoding dependency within a picture. However, they do not need to be contained in separate NAL units (the same as WPP in this regard); thus, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where a strip spans multiple slices, the inter-processor / inter-core communication required for intra-picture prediction between the processing units decoding adjacent strips is limited to transmitting the shared strip header and loop filtering related to reconstructed samples and metadata sharing. When a strip contains more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment except the first one in the strip is signaled in the strip header.
[0085] For simplicity, restrictions on the application of four different picture partitioning schemes are specified in HEVC. For most profiles specified in HEVC, a given coded video sequence cannot contain both slices and wavefronts simultaneously. For each strip and slice, one or both of the following conditions must be satisfied: 1) all coded tree blocks in the strip belong to the same slice; 2) all coded tree blocks in a slice belong to the same strip. Finally, a wavefront segment exactly contains one CTB row, and when using WPP, if a strip starts within a CTB row, the strip must end in the same CTB row.
[0086] The most recent revision of HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K Wang (editors). "HEVC Additional Supplemental Enhancement Information (Draft4)", publicly available on October 24, 2017 at http: / / phenix.intevry.fr / jct / doc_end_user / documents- / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Included within this revision, HEVC specifies three SEI messages related to MCT, namely the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nested SEI message.
[0087] The time-domain MCTSs SEI message indicates the presence of MCTSs in the bitstream and signals the MCTSs. For each MCTS, the motion vectors are restricted to point to full-sample positions within the MCTS and fractional-sample positions that only require full-sample positions within the MCTS for interpolation, and motion vector candidates predicted from time-domain motion vectors for blocks outside the MCTS are not allowed. In this way, each MCTS can be decoded independently without the presence of slices not included in the MCTS.
[0088] The MCTSs extraction information set SEI message provides supplementary information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a bitstream conforming to the MCTS set. This information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains the RBSP bytes for replacing VPSs, SPSs, and PPSs to be used in the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPSs, SPSs, and PPSs) need to be rewritten or replaced, and the slice headers need to be slightly updated because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0089] 3.2. Segmentation of Pictures in VVC
[0090] In VVC, a picture is segmented into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that covers a rectangular region of the picture. The CTUs within a slice are scanned in raster scan order within that slice.
[0091] A strip consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of the picture.
[0092] Two strip modes are supported, namely the raster scan strip mode and the rectangular strip mode. In the raster scan strip mode, a strip contains the complete strip sequence in the strip raster scan of the picture. In the rectangular strip mode, a strip contains multiple complete slices that together form a rectangular region of the picture or multiple consecutive complete CTU rows of a single slice that together form a rectangular region of the picture. The strips within a rectangular strip are scanned in strip raster scan order within the rectangular region corresponding to that strip.
[0093] A sub-picture contains one or more strips that together cover a rectangular region of the picture.
[0094] Figure 1 An example of the raster scan strip segmentation of a picture is shown, where the picture is segmented into 12 slices and 3 raster scan strips.
[0095] Figure 2 Shows an example of rectangular strip segmentation of a picture, where the picture is segmented into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0096] Figure 3 Shows an example of a picture segmented into slices and rectangular strips, where the picture is segmented into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0097] Figure 4 Shows an example of sub - picture segmentation of a picture, where the picture is segmented into 18 slices, 12 slices on the left, each covering a strip with 4x4 CTUs, and 6 slices on the right, each covering 2 vertically stacked strips with 2x2 CTUs, resulting in a total of 24 strips and 24 sub - pictures of different dimensions (each strip is a sub - picture).
[0098] 3.3. Signaling of Sub - pictures, Slices, and Strips in VVC
[0099] In the latest VVC draft text, the information of sub - pictures is signaled in the SPS. The information of sub - pictures includes the sub - picture layout (i.e., the number of sub - pictures per picture and the position and size of each sub - picture) and other sequence - level sub - picture information. The order of sub - pictures signaled in the SPS defines the sub - picture index. The list of sub - picture IDs that each sub - picture has can be explicitly signaled in the SPS or PPS, for example.
[0100] Slices in VVC are conceptually the same as those in HEVC, i.e., each picture is segmented into slice columns and slice rows, but has different syntax for signaling slices in the PPS.
[0101] In VVC, the strip mode is also signaled in the PPS. When the strip mode is the rectangular strip mode, the strip layout of each picture (i.e., the number of strips per picture and the position and size of each strip) is signaled in the PPS. The order of rectangular strips within a picture signaled in the PPS defines the picture - level strip index. The sub - picture - level strip index is defined as the order of the strips within a sub - picture in the ascending order of its picture - level strip index. The position and size of the rectangular strips are sent / derived based on the sub - picture position and size signaled in the SPS (when each sub - picture contains only one strip), or based on the slice position and size signaled in the PPS (when a sub - picture may contain multiple strips). When the strip mode is the raster - scan strip mode, similar to in HEVC, the strip layout within a picture is signaled with different details in the strip itself.
[0102] The SPS, PPS, and strip headers and semantics in the latest VVC draft text most relevant to the present invention are as follows.
[0103] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0104]
[0105]
[0106] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...
[0107] A subpics_present_flag equal to 1 indicates that subpicture parameters are present in the SPS RBSP syntax. A subpics_present_flag equal to 0 indicates that subpicture parameters are not present in the SPS RBSP syntax.
[0108] Note 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the subpictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag in the RBSP of the SPS to be equal to 1.
[0109] sps_num_subpics_minus1 plus 1 indicates the number of subpictures. sps_num_subpics_minus1 shall be in the range of 0 to 254. When not present, the value of sps_num_subpics_minus1 is inferred to be equal to 0.
[0110] subpic_ctu_top_left_x[i] indicates the horizontal position of the top-left CTU of the i-th subpicture in a unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.
[0111] subpic_ctu_top_left_y[i] indicates the vertical position of the top-left CTU of the i-th subpicture in a unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.
[0112] subpic_width_minus1[i] plus 1 refers to the width of the i-th sub-picture in the unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.
[0113] subpic_height_minus1[i] plus 1 refers to the height of the i-th sub-picture in the unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it is not present, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.
[0114] subpic_treated_as_pic_flag[i] being equal to 1 indicates that the i-th sub-picture of each coded picture in CLVS is treated as a picture during the decoding process except for the loop filter operation. subpic_treatment_as_pic_flag[i] being equal to 0 indicates that the i-th sub-picture of each coded picture in CLVS is not treated as a picture during the decoding process except for the loop filter operation. When it is not present, the value of subpic_treatment_as_pic_flag[i] is inferred to be equal to 0.
[0115] loop_filter_across_subpic_enabled_flag[i] being equal to 1 means that the loop filter operation can be performed across the boundaries of the i-th sub-picture in each coded picture in CLVS. loop_filter_across_subpic_enabled_flag[i] being equal to 0 means that the loop filter operation is not performed across the boundaries of the i-th sub-picture in each coded picture in CLVS. When it is not present, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0116] The following constraints apply to the requirements for bitstream consistency:
[0117] For any two sub - pictures subpicA and subpicB, when the sub - picture index of subpicA is less than the sub - picture index of subpicB, any coded or decoded slice NAL unit of subPicA shall precede any coded or decoded slice NAL unit of subPicB in decoding order.
[0118] The shape of the sub - picture shall be such that when each sub - picture is decoded, its entire left boundary and its entire upper boundary are formed by the picture boundary or by the boundaries of previously decoded sub - pictures.
[0119] The sps_subpic_id_present_flag being equal to 1 indicates that there is a sub - picture ID mapping in the SPS. The sps_subpic_id_present_flag being equal to 0 means that the sub - picture ID mapping does not exist in the SPS.
[0120] The sps_subpic_id_signalling_present_flag being equal to 1 means that the sub - picture ID mapping is signalled in the SPS. The sps_subpic_id_signalling_present_flag being equal to 0 means that the sub - picture ID mapping is not signalled in the SPS. When it does not exist, the value of the sps_subpic_id_signalling_present_flag is inferred to be equal to 0.
[0121] sps_subpic_id_len_minus1 plus 1 indicates the number of bits used to represent the syntax element sps_subpic_id[i]. The value of sps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive of the end values.
[0122] sps_subpic_id[i] indicates the sub - picture ID of the i - th sub - picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1 + 1 bits. When it does not exist and when sps_subpic_id_present_flag is equal to 0, the value of sps_subpic_id[i] is inferred to be equal to i, for each i in the range from 0 to sps_num_subpics_minus1, inclusive of the end values. ...
[0123] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0124]
[0125]
[0126]
[0127] 7.4.3.4 Picture Parameter Set RBSP semantics ...
[0128] The pps_subpic_id_signalling_present_flag being equal to 1 indicates that the sub-picture ID is signalled in the PPS. The pps_subpic_id_signalling_present_flag being equal to 0 indicates that the sub-picture ID is not signalled in the PPS. When sps_subpic_id_present_flag is 0 or sps_subpic_id_signalling_present_flag is equal to 1, the pps_subpic_id_signalling_present_flag shall be equal to 0.
[0129] pps_num_subpics_minus1 plus 1 indicates the number of sub-pictures in the coded pictures that refer to the PPS.
[0130] The value of pps_num_subpic_minus1 shall be equal to sps_num_subpics_minus1, which is a requirement for bitstream conformance.
[0131] pps_subpic_id_len_minus1 plus 1 indicates the number of bits used to represent the syntax element pps_subpic_id[i]. The value of pps_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive of the end values.
[0132] The value of pps_subpic_id_len_minus1 for the coded pictures referenced in the CLVS shall be the same for all PPSs, which is a requirement for bitstream conformance.
[0133] pps_subpic_id[i] indicates the sub-picture ID of the i-th sub-picture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1 + 1 bits.
[0134] The no_pic_partition_flag being equal to 1 indicates that picture partitioning is not applied to each picture that refers to the PPS. The no_pic_partition_flag being equal to 0 indicates that each picture that refers to the PPS can be partitioned into multiple slices or strips.
[0135] For all PPSs referenced by the coded pictures within CLVS, the value of no_pic_partition_flag shall be the same, which is a requirement for bitstream consistency.
[0136] When the value of sps_num_subpics_minus1 + 1 is greater than 1, the value of no_pic_partition_flag shall not be equal to 1, which is a requirement for bitstream consistency.
[0137] pps_log2_ctu_size_minus5 plus 5 refers to the size of the luma coded tree blocks of each CTU. pps_log2_ctu_size_minus5 shall be equal to sps_log2_ctu_size_minus5.
[0138] num_exp_tile_columns_minus1 plus 1 refers to the number of explicitly provided tile column widths. The value of num_exp_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be 0.
[0139] num_exp_tile_rows_minus1 plus 1 refers to the number of explicitly provided tile row heights. The value of num_exp_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive. When no_pic_partition_flag is equal to 1, the value of num_exp_tile_columns_minus1 is inferred to be 0.
[0140] tile_column_width_minus1[i] plus 1 refers to the width of the i-th tile column of the units in CTBs, where i is in the range of 0 to num_exp_tile_columns_minus1 - 1, inclusive. tile_column_width_minus1[num_exp_tile_columns_minus1] is used to derive the widths of tile columns with indices greater than or equal to num_exp_tile_columns_minus1, as described in Clause 6.5.1. When it does not exist, the value of tile_column_width_minus1[0] is inferred to be PicWidthInCtbsY - 1.
[0141] tile_row_height_minus1[i] + 1 refers to the height of the i-th tile row of the units in the CTBs, where i ranges from 0 to num_exp_tile_rows_minus1 - 1, inclusive. tile_row_height_minus1[num_exp_tile_rows_minus1] is used to derive the height of tile rows with indices greater than or equal to num_exp_tile_rows_minus1, as described in Clause 6.5.1. When it is absent, the value of tile_row_height_minus1[0] is inferred to be equal to PicHeightInCtbsY - 1.
[0142] rect_slice_flag being equal to 0 means that the slices within each slice group are arranged in raster scan order, and the slice group information is not signaled in the PPS. rect_slice_flag being equal to 1 means that the slices within each slice group cover a rectangular region of the picture, and the slice group information is signaled in the PPS. When it is absent, it is inferred that rect_slice_flag is equal to 1. When subpics_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0143] single_slice_per_subpic_flag being equal to 1 means that each sub - picture consists of one and only one rectangular slice group. single_slice_per_subpic_flag being equal to 0 means that each sub - picture can contain one or more rectangular slice groups. When subpics_present_flag is equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, it is inferred that num_slices_in_pic_minus1 is equal to sps_num_subpics_minus1.
[0144] num_slices_in_pic_minus1 + 1 refers to the number of rectangular slice groups in each picture with reference to the PPS. The value of num_slices_in_pic_minus1 shall be in the range of 0 to MaxSlicesPerPicture - 1, inclusive, where MaxSlicesPerPicture is specified in Annex A. When no_pic_partition_flag is equal to 1, it is inferred that the value of num_slices_in_pic_minus1 is equal to 0.
[0145] When tile_idx_delta_present_flag equals 0, it means that the tile_idx_delta value does not exist in the PPS, and all rectangular stripes in the picture referring to the PPS are specified in raster scan order according to the process defined in Clause 6.5.1. When tile_idx_delta_present_flag equals 1, it means that the tile_idx_delta value may exist in the PPS, and all rectangular stripes in the picture referring to the PPS are specified in the order indicated by the tile_idx_delta value.
[0146] slice_width_in_tiles_minus1[i] plus 1 refers to the width of the i-th rectangular stripe in the unit of the slice column. The value of slice_width_in_tiles_minus1[i] shall be in the range of 0 to NumTileColumns - 1, inclusive of the end values. When it does not exist, the value of slice_width_in_tiles_minus1[i] shall be inferred according to the provisions in Section 6.5.1.
[0147] slice_height_in_tiles_minus1[i] plus 1 refers to the height of the i-th rectangular stripe in the unit of the slice row. The value of slice_height_in_tiles_minus1[i] shall be in the range of 0 to NumTileRows - 1, inclusive of the end values. When it does not exist, the value of slice_height_in_tiles_minus1[i] shall be inferred according to the provisions in Section 6.5.1.
[0148] num_slices_in_tile_minus1[i] plus 1 refers to the number of stripes in the current slice, applicable to the case where the i-th stripe contains a subset of CTU rows from a single slice. The value of num_slices_in_tile_minus1[i] shall be in the range of 0 to RowHeight[tileY] - 1, inclusive of the end values, where tileY is the slice row index containing the i-th stripe. When it does not exist, the value of num_slices_in_tile_minus1[i] is inferred to be equal to 0.
[0149] slice_height_in_ctu_minus1[i] plus 1 refers to the height of the i-th rectangular stripe in the unit of the CTU row, applicable to the case where the i-th stripe contains a subset of CTU rows from a single slice. The value of slice_height_in_ctu_minus1[i] shall be in the range of 0 to RowHeight[tileY] - 1, inclusive of the end values, where tileY is the slice row index containing the i-th stripe.
[0150] tile_idx_delta[i] refers to the tile index difference between the i-th rectangular strip and the (i + 1)-th rectangular strip. The value of tile_idx_delta[i] shall be in the range of –NumTilesInPic + 1 to NumTilesInPic - 1, inclusive of the end values. When it does not exist, the value of tile_idx_delta[i] is inferred to be equal to 0. In all other cases, the value of tile_idx_delta[i] shall not be equal to 0.
[0151] loop_filter_across_tiles_enabled_flag being equal to 1 means that loop filtering operations can be performed across tile boundaries in the picture of the reference PPS. loop_filter_across_tiles_enabled_flag being equal to 0 means that loop filtering operations are not performed across tile boundaries in the picture of the reference PPS. Loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When it does not exist, the value of loop_filter_across_tiles_enabled_flag is inferred to be equal to 1.
[0152] loop_filter_across_slices_enabled_flag being equal to 1 means that loop filtering operations can be performed across slice boundaries in the picture of the reference PPS. loop_filter_across_slice_enabled_flag being equal to 0 means that loop filtering operations are not performed across slice boundaries in the picture of the reference PPS. Loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When it does not exist, the value of loop_filter_across_slices_enabled_flag is inferred to be equal to 0.
[0153] 7.3.7.1 General Slice Header Syntax
[0154]
[0155] 7.4.8.1 General Slice Header Semantics ...
[0156] The slice_subpic_id refers to the sub-picture identifier of the sub-picture containing the slice. If the slice_subpic_id exists, the value of the variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to the slice_subpic_id. Otherwise (the slice_subpic_id does not exist), the variable SubPicIdx is derived to be equal to 0. The length of the slice_subpic_id, in bits, is derived as follows:
[0157] — If the sps_subpic_id_signalling_present_flag is equal to 1, the length of the slice_subpic_id is equal to sps_subpic_id_len_minus1 + 1.
[0158] — Otherwise, if the ph_subpic_id_signalling_present_flag is equal to 1, the length of the slice_subpic_id is equal to ph_subpic_id_len_minus1 + 1.
[0159] — Otherwise, if the pps_subpic_id_signalling_present_flag is equal to 1, the length of the slice_subpic_id is equal to pps_subpic_id_len_minus1 + 1.
[0160] — Otherwise, the length of the slice_subpic_id is equal to Ceil(Log2(sps_num_subpics_minus1 + 1)).
[0161] The slice_address refers to the slice address of the slice. When it does not exist, the value of the slice_address is inferred to be equal to 0.
[0162] If the rect_slice_flag is equal to 0, the following applies:
[0163] — The slice address is the raster scan slice index.
[0164] — The length of the slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0165] — The value of the slice_address should be in the range of 0 to NumTilesInPic - 1, inclusive of the end values.
[0166] Otherwise (the rect_slice_flag is equal to 1), the following applies:
[0167] — The slice address is the slice index of the slice within the SubPicIdx-th sub-picture.
[0168] — The length of slice_address is Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits.
[0169] — The value of slice_address shall be in the range of 0 to NumSlicesInSubpic[SubPicIdx] - 1, inclusive.
[0170] The following constraints apply to the requirements for bitstream compliance:
[0171] — If rect_slice_flag is equal to 0 or subpics_present_flag is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture.
[0172] — Otherwise, the pair of slice_subpic_id and slice_address values shall not be equal to the pair of slice_subpic_id and slice_address values of any other coded slice NAL unit of the same coded picture.
[0173] — When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.
[0174] — The shape of the picture slices shall be such that when each CTU is decoded, its entire left boundary and entire upper boundary shall be composed of the picture boundary or the boundary of a previously decoded CTU.
[0175] num_tiles_in_slice_minus1 plus 1 (if present) indicates the number of slices in the slice. The value of num_tiles_in_slice_minus1 shall be in the range of 0 to NumTilesInPic - 1, inclusive.
[0176] The variable NumCtuInCurrSlice, which indicates the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i], which indicates the picture raster scan address of the i-th CTB within the slice, where i ranges from 0 to NumCtuInCurrSlice - 1, inclusive, are derived as follows:
[0177]
[0178] The variables SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:
[0179]
[0180]
[0181] 3.4 Embodiment of JVET-Q0075
[0182] Escape samples are used to handle exceptional cases in the palette mode.
[0183] The binarization of escape samples is EG3 in the current VTM. However, for uniformly distributed signaling notifications, fixed-length binarization may be superior to EG3 in terms of both distortion and bitrate measurement.
[0184] JVET-Q0075 proposes to use fixed-length binarization for escape samples and also modify the quantization and dequantization processes accordingly.
[0185] In the proposed method, the maximum bit depth of escape samples depends on the quantization parameter and is derived as follows.
[0186] max(1, bitDepth – (max(QpPrimeTsMin, Qp) – 4) / 6)
[0187] Here, bitDepth is the internal bit depth, Qp is the current quantization parameter (QP), QpPrimeTsMin is the minimum QP of the transform skip block, and max is an operation to obtain the larger value between two inputs.
[0188] In addition, only a shift operation is required in the dequantization process of escape samples. Let escapeVal be the decoded escape value and recon be the reconstructed value of the escape sample. The value is derived as follows.
[0189] shift = min(bitDepth – 1, (max(QpPrimeTsMin, Qp) – 4) / 6)
[0190] recon = (escapeVal << shift)
[0191] It is ensured that the distortion of the reconstructed value is always less than or equal to the distortion of the current design, e.g., ((escapeVal * levelScale[qP % 6]) << (qP / 6) + 32) >> 6.
[0192] At the encoder, quantization is implemented as follows:
[0193] escapeVal = (p + (1 << (shift – 1))) >> shift
[0194] escapeVal = clip3(0, (1 << bd) – 1, escapeVal)
[0195] Compared with the current design that uses EG3 and anti-quantization tables, one addition, one multiplication, and two shift operations for quantization, the method proposed in this application is much simpler and only requires one shift operation.
[0196] 3.5 Examples of the implementation of JVET-Q0294
[0197] To achieve effective compression in hybrid lossy and lossless coding / decoding, JVET-Q0294 proposes signaling a flag at each coding tree unit (CTU) to indicate whether the CTU is coded / decoded in lossless or lossy mode. If the CTU is losslessly coded / decoded, an additional CTU-level flag is signaled to specify the residual coding method used for that CTU, either conventional residual coding or transform skip residual coding.
[0198] 4. Examples of technical problems solved by the solutions in this article
[0199] The existing designs for signaling sub-pictures, slices, and stripes in VVC have the following problems:
[0200] 1) The coding of sps_num_subpics_minus1 is u(8), which limits the number of sub-pictures per picture to no more than 256. However, in some applications, the maximum number of sub-pictures per picture may need to be greater than 256.
[0201] 2) It is allowed that subpics_present_flag is equal to 0 and sps_subpic_id_present_flag is equal to 1. However, this is meaningless because subpics_present_flag being equal to 0 means that CLVS has no information about sub-pictures at all.
[0202] 3) For each of the sub-pictures, the list of sub-picture IDs can be signaled in the picture header (PH). However, when signaling the list of sub-picture IDs in the PH and when extracting a subset of sub-pictures from the bitstream, all the PHs will need to be changed. This is not desirable.
[0203] 4) Currently, when the sub-picture ID is indicated to be signaled explicitly, by sps_subpic_id_present_flag equal to 1 (or the name of the syntax element is changed to subpic_ids_explicitly_signalled_flag), the sub-picture ID may not be signaled anywhere. This is problematic because when the sub-picture ID is indicated to be signaled explicitly, the sub-picture ID needs to be signaled explicitly in the SPS or PPS.
[0204] 5) When the sub-picture ID is not signaled explicitly, as long as subpics_present_flag is equal to 1, including when sps_num_subpics_minus1 is equal to 0, the slice_subpic_id syntax element in the slice header still needs to be signaled. However, the length of slice_subpic_id is currently specified as Ceil(Log2(sps_num_subpics_minus1 + 1)) bits, which is 0 bits when sps_num_subpics_minus1 is equal to 0. This is problematic because the length of any existing syntax element cannot be 0 bits.
[0205] 6) The sub-picture layout, including the number, size, and position of the sub-pictures, remains the same throughout the CLVS. Even if the sub-picture ID is not signaled explicitly in the SPS or PPS, the sub-picture ID length still needs to be signaled for the sub-picture ID syntax element in the slice header.
[0206] 7) Whenever rect_slice_flag is equal to 1, the syntax element slice_address is signaled in the slice header and specifies the slice index within the sub-picture containing the slice, including when the number of slices within the sub-picture (i.e., NumSlicesInSubpic[SubPicIdx]) is equal to 1. However, currently, when rect_slice_flag is equal to 1, the length of slice_address is specified as Ceil(Log2(NumSlicesInSubpic[SubPicIdx])) bits, which is 0 bits when NumSlicesInSubpic[SubPicIdx] is equal to 1. This is problematic because the length of any existing syntax element cannot be 0 bits.
[0207] 8) There is redundancy between the syntax element no_pic_partition_flag and pps_num_subpics_minus1, although the latest VVC text has the following constraint: when sps_num_subpics_minus1 is greater than 0, the value of no_pic_partition_flag shall be equal to 1.
[0208] 9) Within CLVS, the subpicture ID value for a specific subpicture position or index may vary from picture to picture. When this occurs, in principle, subpictures cannot use inter-layer prediction by referring to reference pictures in the same layer. However, currently, there is a lack of a constraint in the current VVC specification to prohibit this practice.
[0209] 10) In the current VVC design, reference pictures can be pictures in different layers to support various applications such as scalable video coding and multi-view video coding. If subpictures exist in different layers, it is necessary to study whether inter-layer prediction is allowed or not.
[0210] 5. Example Technologies and Embodiments
[0211] To solve the above problems and other problems, methods outlined below are disclosed. The present invention should be regarded as an example for explaining general concepts and should not be interpreted narrowly. In addition, these inventions can be applied individually or in any combined manner.
[0212] The following abbreviations have the same meanings as in JVET-P1001-vE.
[0213] BP (Buffering Period), BP SEI (Supplemental Enhancement Information),
[0214] PT (Picture Timing), PT SEI,
[0215] AU (Access Unit),
[0216] OP (Operating Point),
[0217] DUI (Decode Unit Information), DUI SEI,
[0218] NAL (Network Layer),
[0219] NUT (NAL Unit Type),
[0220] GDR (Gradual Decoding Refresh),
[0221] SLI (Subpicture Level Information), SLI SEI.
[0222] 1) To solve the first problem, the encoding and decoding of sps_num_subpics_minus1 is changed from u(8) to ue(v), so that each picture can have more than 256 sub-pictures.
[0223] a. In addition, the value of sps_num_subpics_minus1 is restricted to the range from 0 to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*Ceil(pic_height_max_in_-luma_samples÷CtbSizeY)-1.
[0224] b. In addition, the number of sub-pictures of each picture is further restricted in the definition of the level.
[0225] 2) To solve the second problem, the condition of the signalling syntax element sps_subpic_id_present_flag is set to "if(subpics_present_flag)", that is, when subpics_present_flag is equal to 0, the sps_subpic_id_present_flag syntax element is not signalled, and when it does not exist, it is inferred that the value of sps_subpic_id_present_flag is equal to 0.
[0226] a. Alternatively, when subpics_present_flag is equal to 0, the syntax element sps_subpic_id_present_flag is still signalled, but when subpics_present_flag is equal to 0, the value needs to be equal to 0.
[0227] b. In addition, the names of the syntax elements subpics_present_flag and sps_subpic_id_present_flag are changed to subpic_info_present_flag and subpic_ids_explicitly_signalled_flag respectively.
[0228] 3) To solve the third problem, the signalling of the sub-picture ID in the PH syntax is removed. Therefore, for i in the range from 0 to sps_num_subpics_minus1 (including the end values), the list SubpicIdList[i] is derived as follows:
[0229]
[0230] 4) To solve the fourth problem, when a sub - picture is signaled by explicit signaling, signal the sub - picture ID in the SPS or PPS.
[0231] a. This is achieved by adding the following constraint: If subpic_ids_explicitly_signalled_flag is 0 or subpic_ids_in_sps_flag equals 1, then subpic_ids_in_pps_flag should equal 0. Otherwise (subpic_ids_explicitly_signalled_flag is 1 or subpic_ids_in_sps_flag equals 0), subpic_ids_in_pps_flag should equal 1.
[0232] 5) To solve the fifth and sixth problems, regardless of the value of the SPS flag sps_subpic_id_present_flag (or renamed subpic_ids_explicitly_signalled_flag), signal the length of the sub - picture ID in the SPS. Although when the sub - picture ID is also signaled explicitly in the PPS, the length can also be signaled in the PPS to avoid parsing the PPS's dependence on the SPS. In this case, the length also specifies the length of the sub - picture ID in the slice header, even if the sub - picture ID is not signaled explicitly in the SPS or PPS. Therefore, when it exists, the length of slice_subpic_id is also specified by the length of the sub - picture ID signaled in the SPS.
[0233] 6) Alternatively, to solve the fifth and sixth problems, add a flag in the SPS syntax whose value is 1 to specify the presence of the sub - picture ID length in the SPS syntax. The presence of this flag is independent of the value of the flag indicating whether the sub - picture ID is signaled explicitly in the SPS or PPS. When subpic_ids_explicitly_signalled_flag equals 0, the value of this flag can equal 1 or 0, but when subpic_ids_explicitly_signalled_flag equals 1, the value of this flag must equal 1. When this flag equals 0, i.e., the sub - picture length does not exist, the length of slice_subpic_id is specified as Max(Ceil(Log2(sps_num_subpics_minus1 + 1)), 1) bits (instead of Ceil(Log2(sps_num_subpics_minus1 + 1)) bits in the latest VVC draft text).
[0234] a. Alternatively, this flag exists only when subpic_ids_explicitly_signalled_flag is equal to 0, and when subpic_ids_explicitly_signalled_flag is equal to 1, it is inferred that the value of this flag is equal to 1.
[0235] 7) To solve the seventh problem, when rect_slice_flag is equal to 1, the length of slice_address is specified as Max(Ceil(Log2(NumSlicesInSubpic[SubPicIdx])), 1) bits.
[0236] a. Alternatively, further, when rect_slice_flag is equal to 0, the length of slice_address is specified as Max(Ceil(Log2(NumTilesInPic)), 1) bits instead of Ceil(Log2(NumTilesInPic)) bits.
[0237] 8) To solve the eighth problem, the condition for signalling no_pic_partition_flag is set to "if(subpic_ids_in_pps_flag && pps_num_subpics-_minus1 > 0)", and the following inference is added: when it does not exist, it is inferred that the value of no_pic_partition_flag is equal to 1.
[0238] a. Alternatively, move the subpicture ID syntax (all four syntax elements) after the slice and strip syntax in the PPS, for example, immediately before the syntax element entropy_coding_sync_enabled_flag, and then set the condition for signalling pps_num_subpics_minus1 to "if(no_pic_partition_flag)".
[0239] 9) To solve the ninth problem, the following constraint is specified: for each specific subpicture index (or equivalently, subpicture position), when the subpicture ID value changes in picA compared to the previous picture in the decoding order at the same layer, for picA, unless picA is the first picture of CLVS, the subpicture at picA should only contain coded slice NAL units with nal_unit_type equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.
[0240] a. Alternatively, the above constraints only apply to sub-picture indices where the value of subpic_treatment_as_pic_flag[i] is equal to 1.
[0241] b. Alternatively, for the above 9 and 9a, change "IDR_W_RADL, IDR_N_LP or CRA_NUT" to "IDR_W_RADL, IDR_N_LP, CRA_NUT, RSV_IRAP_11 or RSV_IRAP_12".
[0242] c. Alternatively, the sub-pictures at picA may contain other types of coded and decoded slice NAL units. However, these coded and decoded slice NAL units only use one or more of inter-layer prediction, intra-layer block copy (IBC) prediction, and palette mode prediction.
[0243] d. Alternatively, the first video unit (such as a slice, picture, block, etc.) in the sub-picture of picA may refer to the second video unit in the previous picture. It is restricted that the second video unit and the first video unit can be in sub-pictures with the same sub-picture index, although their sub-picture IDs may be different. The sub-picture index is the unique number assigned to the sub-picture and cannot be changed in CLVS.
[0244] 10) For a specific sub-picture index (or equivalently, sub-picture position), it may be signaled in the bitstream which sub-pictures (identified by the layer ID value together with the sub-picture index or sub-picture ID value) are used as reference pictures.
[0245] 11) For the multi-layer case, when certain conditions are met (e.g., possibly depending on the number of sub-pictures, the position of the sub-pictures), inter-layer prediction (ILR) from sub-pictures in different layers is allowed, while when certain conditions are not met, ILR is disabled.
[0246] a. In one example, even when two sub-pictures in two layers have the same sub-picture index value but different sub-picture ID values, inter-layer prediction can still be allowed when certain conditions are met.
[0247] i. In one example, certain conditions are "if the two layers are associated with different view order indices / view order ID values".
[0248] b. If two sub-pictures have the same sub-picture index, the first sub-picture in the first layer and the second sub-picture in the second layer can be constrained to be in juxtaposed positions and / or reasonable widths / heights.
[0249] c. If the first sub - picture can refer to the second reference sub - picture, the first sub - picture in the first layer and the second sub - picture in the second layer can be restricted to be in a juxtaposed position and / or have a reasonable width / height.
[0250] 12) An indication of whether the current sub - picture can use inter - layer prediction (ILP) from sample values and / or other values (e.g., motion information and / or codec mode information) associated with a region or sub - picture of a reference layer is signaled in the bitstream, e.g., signaled in the VPS / DPS / SPS / PPS / APS / sequence header / picture header.
[0251] a. In one example, the reference region or sub - picture of the reference layer contains at least one collocated sample of the samples within the current sub - picture.
[0252] b. In one example, the reference region or sub - picture of the reference layer is outside the collocated region of the current sub - picture.
[0253] c. In one example, the indication is signaled in one or more SEI messages.
[0254] d. In one example, regardless of whether the reference layer has multiple sub - pictures, and when multiple sub - pictures exist in one or more reference layers, regardless of whether the splitting of the picture into sub - pictures is such that each sub - picture in the current picture is a corresponding sub - picture in the reference picture that covers the collocated region, and regardless of whether the corresponding / collocated sub - pictures have the same sub - picture ID value as the current sub - picture, the indication is signaled.
[0255] 13) When BP SEI messages and PT SEI messages applied to a specific OP appear in an AU, the BP SEI message shall be before the PT SEI message in the decoding order.
[0256] 14) When BP SEI messages and DUI SEI messages applied to a specific OP appear in an AU, the BP SEI message shall be before the DUI SEI message in the decoding order.
[0257] 15) When PT SEI messages and DUI SEI messages applied to a specific OP appear in an AU, the PT SEI message shall be before the DUI SEI message in the decoding order.
[0258] 16) An indication of whether a picture is an IRAP / GDR picture flag is signaled at the start of the picture header, and the no_output_of_prior_pics_flag can be signaled according to this indication.
[0259] Exemplary syntax design is as follows:
[0260]
[0261] When irap_or_gdr_pic_flag is equal to 1, it is specified that the picture associated with PH is an IRAP or GDR picture. When irap_or_gdr_pic_flag is equal to 0, it is specified that the picture associated with PH is neither an IRAP picture nor a GDR picture.
[0262] 17) It is not allowed that a picture with mixed_nalu_types_in_pic_flag equal to 0 does not contain coded strip NAL units with nal_unit_type equal to GDR_NUT.
[0263] 18) Signaling the syntax element (i.e., mixed_slice_types_in_pic_flag) in PPS. If mixed_slice_types_in_pic_flag is equal to 0, the slice type (B, P, or I) is coded in PH. Otherwise, the slice type is coded in SHs. The syntax values related to the unused slice types are further skipped in the picture header. Conditionally signaling the syntax element mixed_slice_types_in_pic_flag is as follows:
[0264] if (!rect_slice_flag || num_slices_in_pic_minus1 > 0) mixed_slice_types_in_pic_flag u(1)
[0265] 19) In SPS or PPS, signal at most N1 (e.g., 3) vertical virtual boundaries and at most N2 (e.g., 3) horizontal virtual boundaries. In PH, signal at most N3 (e.g., 1) additional vertical boundaries and at most N4 (e.g., 1) additional horizontal virtual boundaries, and it is restricted that the total number of vertical virtual boundaries should be less than or equal to N1, and the total number of horizontal virtual boundaries should be less than or equal to N2.
[0266] 20) The syntax element inter_layer_ref_pics_present_flag in SPS can be conditionally signaled as follows:
[0267] if (sps_video_parameter_set_id != 0) inter_layer_ref_pics_present_flag u(1)
[0268] 21) Signal the syntax elements entropy_coding_sync_enabled_flag and entry_point_offsets_present_flag in SPS instead of in PPS.
[0269] 22) It is required that the values of vps_num_ptls_minus1 and num_ols_hrd_params_minus1 must be less than the value T. For example, T can be equal to TotalNumOlss specified in JVET-2001-vE.
[0270] a. The difference between T and vps_num_ptls_minus1 or (vps_num_ptls_minus1 + 1) can be signaled.
[0271] b. The difference between T and hrd_params_minus1 or (hrd_params_minus1 + 1) can be signaled.
[0272] c. The above differences can be signaled through unary coding or exponential-Golomb coding.
[0273] 23) It is required that the value of vps_num_dpb_params_minus1 must be less than or equal to vps_max_layers_minus1.
[0274] a. The difference between T and vps_num_ptls_minus1 or (vps_num_ptls_minus1 + 1) can be signaled.
[0275] b. The difference can be signaled through unary coding or exponential-Golomb coding.
[0276] 24) The EOB NAL unit is allowed to be provided to the decoder by being included in the bitstream or in an external manner.
[0277] 25) The EOS NAL unit is allowed to be provided to the decoder by being included in the bitstream or in an external manner.
[0278] 26) For each layer with layer index i, when vps_independent_layer_flag[i] is equal to 0, each picture in that layer should contain only one sub-picture.
[0279] 27) During the sub-bitstream extraction process, it is specified that whenever a VCL NAL unit is removed, the padding data unit associated with the VCL NAL unit is also removed, and all padding SEI messages in the SEI NAL unit associated with the VCL NAL unit are removed.
[0280] 28) It is restricted that when an SEI NAL unit contains a padding SEI message, the SEI NAL unit does not contain any other SEI messages that are not padding SEI messages.
[0281] a. Alternatively, when a SEI NAL unit contains a filler SEI message, the SEI NAL unit shall not contain any other SEI messages.
[0282] 29) Restrict that a VCL NAL unit shall have at most one associated filler NAL unit.
[0283] 30) Restrict that a VCL NAL unit shall have at most one associated filler NAL unit.
[0284] 31) It is recommended to signal a flag to indicate whether a video unit is coded in lossless or lossy mode, where the video unit can be a CU or a CTU, and the signaling of escaped samples in palette mode may depend on the flag.
[0285] a. In one example, the binarization method of escaped samples in palette mode may depend on the flag.
[0286] b. In one example, the determination of the coding context in the arithmetic coding of escaped samples in palette mode may depend on the flag.
[0287] 6. Embodiments
[0288] The following are some example embodiments of all the inventive parts other than the 8 items summarized above in Section 5, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-P2001-v14. The most relevant added or modified parts are Underlined, bold, and italic text shown, and the most relevant deleted parts are highlighted in bold double brackets, e.g., indicating that "a" has been deleted. There are also some other changes that are editorial in nature and thus not highlighted.
[0289] 6.1. First Embodiment
[0290] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0291]
[0292] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...
[0293]
[0294]
[0295] Note 2 – When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream of the sub-bitstream extraction process, it may be necessary to set the value in the SPS to be equal to 1.
[0296] When it does not exist, the value of sps_num_subpics_minus1 is inferred to be equal to 0.
[0297] Refers to the horizontal position of the top-left CTU of the i-th sub-picture in a unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it does not exist, the value of subpic_ctu_top_left_x[i] is inferred to be equal to 0.
[0298] Refers to the vertical position of the top-left CTU of the i-th sub-picture in a unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it does not exist, the value of subpic_ctu_top_left_y[i] is inferred to be equal to 0.
[0299] Plus 1 refers to the width of the i-th sub-picture in a unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits. When it does not exist, the value of subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1.
[0300] Plus 1 refers to the height of the i-th sub-picture in a unit of CtbSizeY. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. When it does not exist, the value of subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1.
[0301] A value of 1 indicates that the i-th sub-picture of each coded picture in CLVS is considered as a picture during the decoding process excluding loop filter operations. A value of 0 for subpic_treatment_as_pic_flag[i] means that the i-th sub-picture of each coded picture in CLVS is not considered as a picture during the decoding process excluding loop filter operations. When it does not exist, it is inferred that the value of subpic_treatment_as_pic_flag[i] is equal to 0.
[0302] A value of 1 means that loop filter operations can be performed across the boundaries of the i-th sub-picture in each coded picture in CLVS. A value of 0 for loop_filter_across_subpic_enabled_flag[i] indicates that loop filter operations are not performed across the boundaries of the i-th sub-picture in each coded picture in CLVS. When it does not exist, the value of loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.
[0303] The following constraints apply to the requirements for bitstream consistency:
[0304] — For any two sub-pictures subpicA and subpicB, when the sub-picture index of subpicA is less than that of subpicB, any coded strip NAL unit of subPicA shall be prior in decoding order to any coded strip NAL unit of subPicB.
[0305] — The shape of the sub-picture shall be such that when each sub-picture is decoded, its entire left boundary and entire upper boundary are composed of the picture boundary or the boundary of a previously decoded sub-picture.
[0306]
[0307] Refers to the sub-picture ID of the i-th sub-picture. The length of the sps_subpic_id[i] syntax element is sps_subpic_id_len_minus1 + 1 bits. ...
[0308] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0309]
[0310]
[0311]
[0312] 7.4.3.4 Picture Parameter Set RBSP Semantics ...
[0313]
[0314] shall be equal to sps_num_subpics_minus1.
[0315] shall be equal to sps_subpic_id_len_minus1.
[0316] refers to the subpicture ID of the i-th subpicture. The length of the pps_subpic_id[i] syntax element is pps_subpic_id_len_minus1 + 1 bits.
[0317]
[0318]
[0319] The requirement for bitstream consistency is that for any i and j within the range of 0 to sps_num_subpics_minus1 (including the end values), when i is less than j, SubpicIdList[i] shall be less than SubpicIdList[j]. ...
[0320] Equal to 0 means that the slices within each strip are arranged in raster scan order and the strip information is not signaled in the PPS. rect_slice_flag equal to 1 means that the slices within each strip cover a rectangular area of the picture and the strip information is signaled in the PPS. When it does not exist, it is inferred that rect_slice_flag is equal to 1. When is equal to 1, the value of rect_slice_flag shall be equal to 1.
[0321] Equal to 1 means that each subpicture consists of one and only one rectangular strip. single_slice_per_subpic_flag equal to 0 means that each subpicture can contain one or more rectangular strips. When is equal to 0, single_slice_per_subpic_flag shall be equal to 0. When single_slice_per_subpic_flag is equal to 1, it is inferred that num_slices_in_pic_minus1 is equal to sps_num_subpics_minus1. ...
[0322] 7.3.7.1 General strip header syntax
[0323]
[0324]
[0325] 7.4.8.1 General strip header semantics ...
[0326] Refers to the subpicture ID of the subpicture containing the strip.
[0327] When it does not exist, it is inferred that the value of slice_subpic_id is equal to 0.
[0328] The variable SubPicIdx is derived such that SubpicIdList[SubPicIdx] is equal to the value of slice_subpic_id.
[0329] Refers to the strip address of the strip. When it does not exist, it is inferred that the value of slice_address is equal to 0.
[0330] If rect_slice_flag is equal to 0, the following applies:
[0331] — The strip address is the raster scan slice index.
[0332] — The length of slice_address is Ceil(Log2(NumTilesInPic)) bits.
[0333] — The value of slice_address should be in the range from 0 to NumTilesInPic - 1, inclusive.
[0334] Otherwise (rect_slice_flag is equal to 1), the following applies:
[0335] — The strip address is the subpicture level strip index of the strip.
[0336] — The length of slice_address is bits.
[0337] — The value of slice_address should be in the range from 0 to NumSlicesInSubpic[SubPicIdx] - 1, inclusive.
[0338] The following constraints apply to the requirements for bitstream conformance:
[0339] — If rect_slice_flag is equal to 0 or is equal to 0, the value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture.
[0340] — Otherwise, the pair of slice_subpic_id and slice_address values shall not be equal to the pair of slice_subpic_id and slice_address values of any other coded slice NAL unit of the same coded picture.
[0341] — When rect_slice_flag is equal to 0, the slices of the picture shall be arranged in ascending order of their slice_address values.
[0342] — The shape of the picture slices shall be such that when each CTU is decoded, its entire left boundary and its entire upper boundary shall be composed of a picture boundary or by the boundaries of previously decoded CTU(s). ...
[0343] Figure 5 FIG. shows a block diagram of an example video processing system 500 that can implement various techniques of the present disclosure. Various implementations may include some or all of the components of system 500. System 500 may include an input 502 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8- or 10-bit multi-component pixel values), or it may be received in a compressed or coded format. Input 502 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0344] System 500 may include a codec component 504 that may implement various codec or encoding methods described in the present disclosure. The codec component 504 may reduce the average bit rate of the video from the input 502 to the output of the codec component 504 to produce a codec representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 504 may be stored or transmitted via a communication connection as represented by component 506. The stored or communicated bitstream (or codec) representation of the video received at the input 502 may be used by component 508, which is used to generate pixel values or a viewable video to be sent to the display interface 510. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Additionally, while certain video processing operations are referred to as "codec" operations or tools, it should be understood that the encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding results will be performed by the decoder.
[0345] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), DisplayPort, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, IDE interface, etc. The techniques described in the present disclosure may be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0346] Figure 6 is a block diagram of a video processing apparatus 600. The apparatus 600 may be used to implement one or more methods described in the present disclosure. The apparatus 600 may be located in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 600 may include one or more processors 602, one or more memories 604, and video processing hardware 606. The processor 602 may be configured to implement one or more methods described in the present disclosure. The memory 604 may be used to store data and code for implementing the methods and techniques described in the present disclosure. The video processing hardware 606 may be used in hardware circuits to implement some of the techniques described in the present disclosure. In some embodiments, the hardware 606 may be partially or entirely in the processor 602, such as a graphics processor.
[0347] Figure 7 is a block diagram of an example video codec system 100 that may utilize the techniques of the present disclosure.
[0348] As Figure 7As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may generate encoded video data, and the source device 110 may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the destination device 120 may be referred to as a video decoding device.
[0349] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0350] The video source 112 may include, for example, a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0351] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0352] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, and the destination device 120 is configured to interface with an external display device.
[0353] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or other standards.
[0354] Figure 8 is a block diagram illustrating an example of a video encoder 200, and the video encoder 200 may be Figure 7 the video encoder 114 in the system 100 illustrated in
[0355] Video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In the example as Figure 8 shown, video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0356] The functional components of video encoder 200 may include a splitting unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0357] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0358] In addition, some components, such as motion estimation unit 204 and motion compensation unit 205, for example, may be highly integrated, but are shown separately in the Figure 8 example for purposes of description.
[0359] Splitting unit 201 may split a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.
[0360] Mode selection unit 203 may select one of the coding / decoding modes (intra or inter) based on, for example, error results, and provide the resulting intra or inter coded / decoded block to residual generation unit 207 to generate residual block data and reconstruction unit 212 to reconstruct the coded / decoded block to be used as a reference picture. In some examples, mode selection unit 203 may select a combination of intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0361] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0362] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice or a B-slice.
[0363] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search for reference pictures in reference video block search list 0 or 1 for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture that contains the reference video block in list 0 or list 1 and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0364] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block, the motion estimation unit 204 may search for reference pictures in reference video block search list 0 for the current video block, and may also search for reference pictures in reference video block search list 1 for the current video block. The motion estimation unit 204 may then generate reference indices indicating the reference pictures in list 0 and list 1 that contain the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference indices and the motion vector for the current video block as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0365] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoder's decoding process.
[0366] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information for the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of an adjacent video block.
[0367] In one example, the motion estimation unit 204 may indicate in a syntax structure associated with the current video block that the current video block has a value of the same motion information as another video block to the video decoder 300.
[0368] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference represents the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0369] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0370] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0371] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0372] In other examples, for the current video block, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0373] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0374] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0375] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block stored in the buffer 213.
[0376] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0377] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0378] Figure 9 is a block diagram illustrating an example of a video decoder 300 that can be the Figure 7 video decoder 114 in the system 100 shown.
[0379] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 8 the example shown, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0380] In an example such as Figure 9 shown, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding channel that is generally opposite to the encoding channel described with respect to the video encoder 200 ( Figure 8 ).
[0381] The entropy decoding unit 301 can obtain the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy encoded video data, and based on the entropy decoded video data, the motion compensation unit 302 can determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge mode.
[0382] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. An identifier for the interpolation filter for sub-pixel accuracy may be included in the syntax element.
[0383] The motion compensation unit 302 may use the interpolation filter used by the video encoder 20 during the encoding of a video block to calculate the interpolation values of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information and use the interpolation filter to generate a prediction block.
[0384] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode the frames and / or slices of an encoded video sequence, the partitioning information describes how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicates how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0385] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301, i.e., dequantizes. The inverse transform unit 303 applies an inverse transform.
[0386] The reconstruction unit 306 may add a residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block in order to remove blocking artifacts. Then the decoded video block is stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.
[0387] Figures 10 - 26 An example method that may implement the above technical solution is shown, for example, Figures 5 - 9 The embodiment shown.
[0388] Figure 10 A flowchart of an example method 1000 of video processing is shown. Method 1000 includes, at operation 1010, performing a conversion between a video and a bitstream of the video, the bitstream including one or more access units according to format rules, and the format rules specify the order in which a first message and a second message for an operation point OP appear in the access unit AU such that the first message precedes the second message in the decoding order.
[0389] Figure 11A flowchart of an example method 1100 for video processing is shown. Method 1000 includes, at operation 1110, performing a conversion between a video and a bitstream of the video, the bitstream including one or more access units according to format rules, and the format rules specifying an order in which a plurality of messages appear in the access units for an operation point OP such that a first message of the plurality of messages precedes a second message of the plurality of messages in the decoding order.
[0390] Figure 12 A flowchart of an example method 1200 for video processing is shown. Method 1200 includes, at operation 1210, performing a conversion between a video including pictures and a bitstream of the video, the bitstream conforming to format rules that specify whether to signal an indication of a first flag at the start of a picture header associated with a picture, the first flag indicating whether the picture is an Intra Random Access Point (IRAP) picture or a Gradual Decoding Refresh (GDR) picture.
[0391] Figure 13 A flowchart of an example method 1300 for video processing is shown. Method 1300 includes, at operation 1310, performing a conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to format rules that do not allow encoding or decoding of pictures in the one or more pictures to include coded strip Network Abstraction Layer (NAL) units of a gradual decoding refresh type and that are associated with a flag indicating that the pictures include NAL units of a mixed type.
[0392] Figure 14 A flowchart of an example method 1400 for video processing is shown. Method 1400 includes, at operation 1410, performing a conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to format rules that allow encoding or decoding of pictures in the one or more pictures to include coded strip Network Abstraction Layer (NAL) units of a gradual decoding refresh type and that are associated with a flag indicating that the pictures do not include NAL units of a mixed type.
[0393] Figure 15 A flowchart of an example method 1500 for video processing is shown. Method 1500 includes, at operation 1510, performing a conversion between pictures of a video and a bitstream of the video, the bitstream conforming to format rules that specify whether to signal a first syntax element in a Picture Parameter Set (PPS) associated with a picture, the picture including one or more strips of a strip type, and since the first syntax element equals 0, the first syntax element indicates signaling of the strip type in a picture header, otherwise the first syntax element indicates signaling of the strip type in a strip header.
[0394] Figure 16A flowchart showing an example method 1600 for video processing is presented. Method 1600 includes, at operation 1610, performing a conversion between a picture of a video and a bitstream of the video according to rules, the conversion including a loop filtering process, and the rules specifying that the total number of vertical virtual boundaries and the total number of horizontal virtual boundaries related to the loop filtering operation are signaled at the picture level or the sequence level.
[0395] Figure 17 A flowchart showing an example method 1700 for video processing is presented. Method 1700 includes, at operation 1710, performing a conversion between a video including one or more pictures and a bitstream of the video, the bitstream conforming to format rules, the format rules conditionally allowing a picture within a layer to be encoded or decoded based on a first syntax element by using reference pictures from other layers, the first syntax element indicating whether reference pictures from other layers are present in the bitstream, and conditionally signaling the first syntax element in the bitstream based on a second syntax element, the second syntax element indicating whether an identifier of a parameter set associated with the picture is not equal to 0.
[0396] Figure 18 A flowchart showing an example method 1800 for video processing is presented. Method 1800 includes, at operation 1810, performing a conversion between a picture of a video and a bitstream of the video, the bitstream conforming to format rules, the format rules specifying a first syntax element that causes (a) context variables to be synchronized before decoding a coding tree unit (CTU) in a picture and (b) context variables to be stored after decoding the CTU, and the first syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
[0397] Figure 19 A flowchart showing an example method 1900 for video processing is presented. Method 1900 includes, at operation 1910, performing a conversion between a picture of a video and a bitstream of the video, the bitstream conforming to format rules, the format rules specifying a syntax element that indicates whether there is signaling of an entry point offset for a slice or slice-specified CTU row in a slice header of the picture, and the syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
[0398] Figure 20 A flowchart showing an example method 2000 for video processing is presented. Method 2000 includes, at operation 2010, performing a conversion between a video and a bitstream of the video according to rules, the rules specifying that a first syntax element is less than a first preset threshold, the first syntax element indicating the number of parameters for an output layer set (OLS) hypothetical reference decoder (HRD) in a video parameter set (VPS) associated with the video.
[0399] Figure 21A flowchart of an example method 2100 for video processing is shown. Method 2100 includes, at operation 2110, performing a conversion between a video and a bitstream of the video according to a rule that specifies that a syntax element is less than a preset threshold, the syntax element indicating the number of profile / tier / level (PTL) syntax structures in a video parameter set (VPS) associated with the video.
[0400] Figure 22 A flowchart of an example method 2200 for video processing is shown. Method 2200 includes, at operation 2210, performing a conversion between a video and a bitstream of the video according to a rule that specifies that a first syntax element, which indicates the number of decoded picture buffer parameter syntax structures in a video parameter set (VPS), must be less than or equal to a second syntax element, which indicates the number of layers specified by the VPS.
[0401] Figure 23 A flowchart of an example method 2300 for video processing is shown. Method 2300 includes, at operation 2310, performing a conversion between a video and a bitstream of the video according to a rule that allows a decoder to obtain a network abstraction layer (NAL) unit of a terminal either by signaling in the bitstream or by being provided by an external means.
[0402] Figure 24 A flowchart of an example method 2400 for video processing is shown. Method 2400 includes, at operation 2410, performing a conversion between a video and a bitstream of the video, the bitstream conforming to a format rule, and due to a syntax element being equal to 0, the format rule restricting that each layer in the bitstream includes only one sub-picture, indicating that each layer is configured to use inter-layer prediction.
[0403] Figure 25 A flowchart of an example method 2500 for video processing is shown. Method 2500 includes, at operation 2510, performing a conversion between a video and a bitstream of the video according to a rule that specifies implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and the rule specifies that during the extraction process, when removing a video coding layer (VCL) network abstraction layer (NAL) unit, padding data units and padding supplementary enhancement information (SEI) messages in SEI VCL units associated with the VCL NAL unit are also deleted.
[0404] Figure 26A flowchart showing an example method 2600 for video processing is presented. Method 2600 includes, at operation 2610, performing a conversion between a video unit of a video and a bitstream of the video, the bitstream conforming to format rules that specify that the bitstream includes a first syntax element that indicates whether the video unit is encoded / decoded in a lossy mode or a lossless mode, and signaling a second syntax element that indicates escape samples selectively included in a palette mode applied to the video unit based on the value of the first syntax element.
[0405] Next, a list of preferred solutions for some embodiments is provided.
[0406] A1. A video processing method includes: performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more access units according to format rules, and wherein the format rules specify the order in which a first message and a second message for an operation point OP appear in an access unit AU such that the first message precedes the second message in the decoding order.
[0407] A2. The method according to solution A1, wherein the first message includes a buffering period BP supplementary enhancement information SEI message, and the second message includes a picture timing (PI) SEI message.
[0408] A3. The method according to solution A1, wherein the first message includes a buffering period BP supplementary enhancement information SEI message, and the second message includes a decoding unit (DUI) SEI message.
[0409] A4. The method according to solution A1, wherein the first message includes a picture timing PT supplementary enhancement information SEI message, and the second message includes a decoding unit (DUI) SEI message.
[0410] A5. A video processing method includes: performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more access units according to format rules, and wherein the format rules specify the order in which a plurality of messages for an operation point OP appear in an access unit such that a first message of the plurality of messages precedes a second message of the plurality of messages in the decoding order.
[0411] A6. The method according to solution A5, wherein the plurality of messages includes a buffering period (BP) SEI message, a decoding unit information (DUI) SEI message, a picture timing (PT) SEI message, and a sub-picture level information (SLI) SEI message.
[0412] A7. The method according to solution A6, wherein the decoding order is the SLI SEI message, the BP SEI message, the PT SEI message, and the DUI message.
[0413] Next, another list of preferred solutions for some embodiments is provided.
[0414] B1. A video processing method, comprising: performing a conversion between a video including pictures and a bitstream of the video, wherein the bitstream conforms to format rules, and wherein the format rules specify whether to indicate a first flag in a start signaling of a picture header associated with the picture, and wherein the first flag indicates whether the picture is an Intra Random Access Point (IRAP) picture or a Gradual Decoding Refresh (GDR) picture.
[0415] B2. The method according to solution B1, wherein the IRAP picture is a picture that can be correctly decoded for the picture and all subsequent pictures in the output order when decoding of the bitstream starts from the picture.
[0416] B3. The method according to solution B1 or B2, wherein the IRAP picture only includes I slices.
[0417] B4. The method according to solution B1, wherein the GDR picture is a picture that can be correctly decoded for the relevant recovery point pictures and all subsequent pictures in the decoding order and the output order when decoding of the bitstream starts from the picture.
[0418] B5. The method according to solution B1, wherein the format rules further specify whether to signal a second flag based on an indication in the bitstream.
[0419] B6. The method according to solution B5, wherein the first flag is irap_or_gdr_pic_flag.
[0420] B7. The method according to solution B6, wherein the second flag is no_output_of_prior_pics_flag.
[0421] B8. The method according to any one of solutions B1 - B7, wherein the first flag being equal to 1 specifies whether the picture is an IRAP picture or a GDR picture.
[0422] B9. The method according to any one of solutions B1 - B7, wherein the first flag being equal to 0 specifies that the picture is neither an IRAP picture nor a GDR picture.
[0423] B10. The method according to any one of solutions B1 - B9, wherein each Network Access Layer (NAL) of the GDR picture has a nal_unit_type syntax element equal to GDR_NUT.
[0424] Next, another list of preferred solutions for some embodiments is provided.
[0425] C1. A video processing method, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules do not allow encoding or decoding of pictures in the one or more pictures to include codec slice network abstraction layer (NAL) units with an asymptotically decoded refresh type and is related to a flag indicating that the pictures include NAL units of a mixed type.
[0426] C2. A video processing method, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules allow encoding or decoding of pictures in the one or more pictures to include codec slice network abstraction layer (NAL) units with an asymptotically decoded refresh type and is related to a flag indicating that the pictures do not include NAL units of a mixed type.
[0427] C3. The method according to solution C1 or C2, wherein the flag is mixed_nalu_types_in_pic_flag.
[0428] C4. The method according to solution C3, wherein the flag is in the picture parameter set (PPS).
[0429] C5. The method according to any one of solutions C1 - C4, wherein the picture is an asymptotically decoded refresh (GDR) picture, and wherein each slice or NAL unit in the picture has a nal_unit_type equal to GDR_NUT.
[0430] C6. A video processing method, comprising: performing a conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to format rules, and the format rules specify whether to signal a first syntax element in a picture parameter set (PPS) related to the picture, wherein the picture includes one or more slices with a slice type, and wherein, since the first syntax element is equal to 0, the first syntax element indicates that the slice type is signaled in the picture header, otherwise the first syntax element indicates that the slice type is signaled in the slice header.
[0431] C7. The method according to solution C6, wherein the first syntax element is mixed_slice_types_in_pic_flag.
[0432] C8. The method according to solution C6 or C7, wherein the first syntax element is signaled because a second syntax element is equal to 0 and / or a third syntax element is greater than 0.
[0433] C9. The method according to solution C8, wherein the second syntax element specifies at least one characteristic of one or more slices in one or more slices.
[0434] C10. A method according to solution C9, wherein a second syntax element being equal to 0 specifies that one or more slices are in raster scan order, and strip information is not included in the PPS.
[0435] C11. A method according to solution C9, wherein a second syntax element being equal to 0 specifies that one or more slices cover a rectangular region of a picture, and strip information is signaled in the PPS.
[0436] C12. A method according to any one of solutions C9 - C11, wherein the second syntax element is rect_slice_flag.
[0437] C13. A method according to solution C8, wherein a third syntax element specifies the number of rectangular strips in a picture of a reference PPS.
[0438] C14. A method according to solution C13, wherein the third syntax element is num_slices_in_pic_minus1.
[0439] C15. A video processing method, comprising: performing a conversion between a picture of a video and a bitstream of the video according to rules, wherein the conversion includes a loop filter process, and wherein the rules specify that the total number of vertical virtual boundaries and the total number of horizontal virtual boundaries related to loop filter operations are signaled at the picture level or the sequence level.
[0440] C16. A method according to solution C15, wherein the total number of vertical virtual boundaries includes a first number N1 of vertical virtual boundaries up to those signaled in a picture parameter set PPS or a sequence parameter set SPS and a second number N3 of additional vertical virtual boundaries up to those signaled in a picture header PH, and wherein the total number of horizontal virtual boundaries includes a third number N2 of horizontal virtual boundaries up to those signaled in the PPS or the SPS and a fourth number N4 of additional horizontal virtual boundaries up to those signaled in the PH.
[0441] C17. A method according to solution C16, wherein N1 + N3 ≤ N1 and N2 + N4 ≤ N2.
[0442] C18. A method according to solution C16 or C17, wherein N1 = 3, N2 = 3, N3 = 1, and N4 = 1.
[0443] Next, another list of preferred solutions of some embodiments is provided.
[0444] D1. A video processing method, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules conditionally allow encoding and decoding of pictures within a layer by using reference pictures from other layers based on a first syntax element, the first syntax element indicating whether reference pictures from other layers exist in the bitstream, and wherein the first syntax element is conditionally signaled in the bitstream based on a second syntax element, the second syntax element indicating whether an identifier of a parameter set associated with the picture is not equal to 0.
[0445] D2. The method according to solution D1, wherein the first syntax element is inter_layer_ref_pics_present_flag.
[0446] D3. The method according to solution D1 or D2, wherein the first syntax element specifies whether inter-layer prediction is enabled and whether inter-layer reference pictures can be used.
[0447] D4. The method according to solution D1, wherein the first syntax element is sps_inter_layer_prediction_enabled_flag.
[0448] D5. The method according to any one of solutions D1-D4, wherein the first syntax element is in the sequence parameter set SPS.
[0449] D6. The method according to any one of solutions D1-D4, wherein the parameter set is the video parameter set VPS.
[0450] D7. The method according to solution D6, wherein the second syntax element is sps_video_parameter_set_id.
[0451] D8. The method according to solution D6 or D7, wherein the parameter set provides an identifier for the VPS for reference by other syntax elements.
[0452] D9. The method according to solution D7, wherein the second syntax element is in the sequence parameter set SPS.
[0453] Next, another list of preferred solutions for some embodiments is provided.
[0454] E1. A video processing method, comprising: performing a conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify a first syntax element, the first syntax element enabling (a) synchronization processing of context variables before decoding a coding tree unit (CTU) in a picture and (b) storage processing of context variables after decoding the CTU, wherein the first syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
[0455] E2. The method according to solution E1, wherein the first syntax element is not included in a picture parameter set (PPS) associated with the picture.
[0456] E3. The method according to solution E1 or E2, wherein the CTU comprises a first coding tree block (CTB) in a row of coding tree blocks in each slice of the picture.
[0457] E4. The method according to any one of solutions E1 - E3, wherein the first syntax element is entropy_coding_sync_enabled_flag.
[0458] E5. The method according to any one of solutions E1 - E3, wherein the first syntax element is sps_entropy_coding_sync_enabled_flag.
[0459] E6. The method according to any one of solutions E1 - E5, wherein the format rules specify a second syntax element, the second syntax element indicating whether there is signaling for an entry point offset for a slice or a slice - specified CTU row in a strip header of the picture, and wherein the second syntax element is signaled in the SPS associated with the picture.
[0460] E7. A video processing method, comprising: performing a conversion between a picture of a video and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules specify a syntax element, the syntax element indicating whether there is signaling for an entry point offset for a slice or a slice - specified coding tree unit (CTU) row in a strip header of the picture, and wherein the syntax element is signaled in a sequence parameter set (SPS) associated with the picture.
[0461] E8. The method according to solution E7, wherein the format rules allow signaling of the entry point offset based on a syntax element having a value equal to 1.
[0462] E9. The method according to solution E7 or E8, wherein the picture parameter set (PPS) associated with the picture does not include the syntax element.
[0463] Method according to any one of Solutions E7 - E9, wherein the syntax element is entry_point_offsets_present_flag.
[0464] Method according to any one of Solutions E7 - E9, wherein the syntax element is sps_entry_point_offsets_present_flag.
[0465] Next, another list of preferred solutions of some embodiments is provided.
[0466] F1. A video processing method, comprising: performing a conversion between a video and a bitstream of the video according to a rule, wherein the rule stipulates that a first syntax element is less than a first preset threshold, and the first syntax element indicates the number of parameters for an output layer set OLS assuming a reference decoder HRD in a video parameter set VPS associated with the video.
[0467] F2. The method according to Solution 1, wherein the first syntax element is vps_num_ols_timing_hrd_params_minus1.
[0468] F3. The method according to Solution F1 or F2, wherein the first preset threshold is the number of multi - layer output layer sets (denoted as NumMultiLayerOlss) minus 1.
[0469] F4. The method according to any one of Solutions F1 - F3, wherein when the first syntax element is not signaled in the bitstream, the first syntax element is inferred to be 0.
[0470] F5. The method according to any one of Solutions F1 - F4, wherein the rule stipulates a second syntax element, and the second syntax element indicates that the number of profile / tier / level PTL syntax structures in a video parameter set VPS associated with the video is less than a second preset threshold.
[0471] F6. A video processing method, comprising: performing a conversion between a video and a bitstream of the video according to a rule, wherein the rule stipulates a syntax element, and the syntax element indicates that the number of profile / tier / level PTL syntax structures in a video parameter set VPS associated with the video is less than a preset threshold.
[0472] F7. The method according to Solution F6, wherein the syntax element is vps_num_ptls_minus1.
[0473] F8. The method according to Solution F6 or F7, wherein the preset threshold is the total number of output layer sets (denoted as TotalNumOlss).
[0474] F9. A method according to any one of solutions F1 - F8, wherein the difference between a preset threshold and a syntax element is signaled in the bitstream.
[0475] F10. A method according to solution F9, wherein the difference is signaled using unary coding and decoding.
[0476] F11. A method according to solution F9, wherein the difference is signaled using exponential - Golomb (EG) coding and decoding.
[0477] F12. A video processing method, comprising: performing a conversion between a video and a bitstream of the video according to a rule, wherein the rule specifies a first syntax element, the first syntax element indicating that the number of decoded picture buffer parameter syntax structures in a video parameter set (VPS) must be less than or equal to a second syntax element, and the second syntax element indicating the number of layers specified by the VPS.
[0478] F13. A method according to solution F12, wherein the difference between the first syntax element and a preset threshold is signaled in the bitstream.
[0479] F14. A method according to solution F12, wherein the difference between the second syntax element and a preset threshold is signaled in the bitstream.
[0480] F15. A method according to solution F13 or F14, wherein the difference is signaled using unary coding and decoding.
[0481] F16. A method according to solution F13 or F14, wherein the difference is signaled using exponential - Golomb (EG) coding and decoding.
[0482] F17. A video processing method, comprising: performing a conversion between a video and a bitstream of the video according to a rule, wherein the rule allows the decoder to obtain a terminal network abstraction layer (NAL) unit by signaling in the bitstream or by an external means.
[0483] F18. A method according to solution F17, wherein the terminal NAL unit is the end - of - bitstream (EOB) of the bitstream NAL unit.
[0484] F19. A method according to solution F9, wherein the terminal NAL unit is the end - of - sequence (EOS) of the sequence NAL unit.
[0485] F20. A method according to any one of solutions F17 - F19, wherein the external means includes a parameter set.
[0486] F21. A video processing method includes: performing a conversion between a video and a bitstream of the video, where the bitstream conforms to format rules, and where the format rules limit each layer in the bitstream to include only one sub-picture due to a syntax element being equal to 0, indicating that each layer is configured to use inter-layer prediction.
[0487] F22. The method according to solution F21, where the syntax element is vps_independent_layer_flag.
[0488] Next, another list of preferred solutions for some embodiments is provided.
[0489] G1. A video processing method includes: performing a conversion between a video and a bitstream of the video according to rules, where the rules stipulate implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, where the sub-bitstream extraction process is configured to extract a sub-bitstream with a target highest temporal identifier from the bitstream, and where the rules stipulate that during the extraction process, when removing video coding layer (VCL) network abstraction layer (NAL) units, padding data units and padding supplementary enhancement information (SEI) messages in the SEI VCL units associated with the VCL NAL units are also deleted.
[0490] G2. The method according to solution G1, where the VCL NAL units are removed based on the identifier of the layer to which the VCL NAL belongs.
[0491] G3. The method according to solution G2, where the identifier is nuh_layer_id.
[0492] G4. The method according to solution G1, where the bitstream conforms to format rules, and the format rules stipulate that SEI NAL units including SEI messages with padding payloads do not include SEI messages with payloads different from the padding payloads.
[0493] G5. The method according to solution G1, where the bitstream conforms to format rules, and the format rules stipulate that SEI NAL units including SEI messages with padding payloads are configured not to include another SEI message with any other payload.
[0494] G6. The method according to solution G1, where a VCL NAL unit has at most one associated padding NAL unit.
[0495] G7. A video processing method includes performing a conversion between a video unit of a video and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule specifies that the bitstream includes a first syntax element that indicates whether the video unit is encoded / decoded in a lossy mode or a lossless mode, and wherein a second syntax element is signaled to indicate escape samples selectively included in a palette mode applied to the video unit based on the value of the first syntax element.
[0496] G8. The method according to solution G7, wherein a binarization method of the escape samples is based on the first syntax element.
[0497] G9. The method according to solution G7, wherein context encoding / decoding in arithmetic encoding / decoding of the escape samples is determined based on a flag.
[0498] G10. The method according to any one of solutions G7 - G9, wherein the video unit is a coding unit (CU) or a coding tree unit (CTU).
[0499] Next, another list of preferred solutions for some embodiments is provided.
[0500] P1. A video processing method includes performing a conversion between a picture of a video and a coded representation of the video, wherein the number of sub - pictures in the picture is included as a field in the coded representation, and the bit - width of the field depends on the value of the number of sub - pictures.
[0501] P2. The method according to solution P1, wherein the field represents the number of sub - pictures using a codeword.
[0502] P3. The method according to solution P2, wherein the codeword includes a Golomb codeword.
[0503] P4. The method according to any one of solutions P1 to P3, wherein the value of the number of sub - pictures is limited to an integer number less than or equal to the number of coding tree blocks suitable within the picture.
[0504] P5. The method according to any one of solutions P1 to P4, wherein the field depends on the coding level associated with the coded representation.
[0505] P6. A video processing method includes performing a conversion between a video region of a video and a coded representation of the video, wherein the coded representation conforms to a format rule, and wherein the format rule specifies that a syntax element indicating a sub - picture identifier is omitted since the video region does not contain any sub - pictures.
[0506] P7. The method according to solution P6, wherein the coded representation includes a field having a value of 0, and the field having a value of 0 indicates that the video region does not include any sub - pictures.
[0507] P8. A video processing method includes performing a conversion between a video region of a video and an encoded / decoded representation of the video, wherein the encoded representation conforms to format rules, and wherein the format rules specify omitting an identifier of a sub-picture in the video region at the video region header level of the encoded / decoded representation.
[0508] P9. The method according to solution P8, wherein the encoded / decoded representation numerically identifies sub-pictures according to the order in which the sub-pictures are listed in the video region header.
[0509] P10. A video processing method includes performing a conversion between a video region of a video and an encoded / decoded representation of the video, wherein the encoded representation conforms to format rules, and wherein the format rules specify including an identifier of a sub-picture in the video region and / or a length of the identifier of the sub-picture at the sequence parameter set level or the picture parameter set level.
[0510] P11. The method according to solution P10, wherein the length is included at the picture parameter set level.
[0511] P12. A video processing method includes performing a conversion between a video region of a video and an encoded / decoded representation of the video, wherein the encoded representation conforms to format rules, and wherein the format rules specify including a field in the encoded / decoded representation at the video sequence level to indicate whether a sub-picture identifier length field is included in the encoded / decoded representation at the video sequence level.
[0512] P13. The method according to solution P12, wherein the format rules specify setting the field to 1 in the case that another field in the encoded / decoded representation indicates that a length identifier of the video region is included in the encoded / decoded representation.
[0513] P14. A video processing method includes performing a conversion between a video region of a video and an encoded / decoded representation of the video, wherein the encoded representation conforms to format rules, and wherein the format rules specify including an indication in the encoded / decoded representation to indicate whether the video region can be used as a reference picture.
[0514] P15. The method according to solution P14, wherein the indication includes a layer ID and an index or ID value associated with the video region.
[0515] P16. A video processing method includes performing a conversion between a video region of a video and an encoded / decoded representation of the video, wherein the encoded representation conforms to format rules, and wherein the format rules specify including an indication in the encoded / decoded representation to indicate whether the video region can use inter-layer prediction (ILP) from a plurality of sample values associated with a video region of a reference layer.
[0516] P17. The method according to solution P16, wherein the indication is included at sequence level, picture level or video level.
[0517] P18. The method according to solution P16, wherein the video region of the reference layer includes at least one collocated sample of the samples within the video region.
[0518] P19. The method according to solution P16, wherein the indication is included in one or more supplementary enhancement information (SEI) messages.
[0519] P20. A video processing method, comprising configuring a codec representation in which a first message is superior to a second message in decoding order, for example, when it is determined that a first message and a second message for an operation point in an access unit of an applied video exist in a codec representation of the video; and based on this configuration, performing a conversion between a video region of the video and the codec representation.
[0520] P21. The method according to solution P20, wherein the first message includes a buffering period (BP) enhancement supplementary information SEI message, and the second message includes a picture timing (PT) SEI message.
[0521] P22. The method according to solution P20, wherein the first message includes a buffering period (BP) supplementary enhancement information SEI message, and the second message includes a decoding unit (DUI) SEI message.
[0522] P23. The method according to solution P20, wherein the first message includes a picture timing (PT) supplementary enhancement information SEI message, and the second message includes a decoding unit (DUI) SEI message.
[0523] P24. The method of any of the above solutions, wherein the video region includes sub - pictures of the video.
[0524] P25. The method of any of the above solutions, wherein the conversion includes parsing and decoding the codec representation to generate the video.
[0525] P26. The method of any of the above solutions, wherein the conversion includes encoding the video to generate the codec representation.
[0526] P27. A video decoding device, comprising a processor configured to execute the method of any one or more of solutions P1 to P26.
[0527] P28. A video encoding device, comprising a processor configured to execute the method of any one or more of solutions P1 to P26.
[0528] P29. A computer program product having computer code stored thereon, which when executed causes a processor to perform the method of any one of Solutions P1 to P26.
[0529] Next, another list of preferred solutions of some embodiments is provided.
[0530] O1. The method of any of the above solutions, wherein the conversion includes decoding a video from a bitstream.
[0531] O2. The method of any of the above solutions, wherein the conversion includes encoding a video into a bitstream.
[0532] O3. The method of any of the above solutions, wherein the conversion includes generating a bitstream from a video, and wherein the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.
[0533] O4. A method of storing a bitstream representing a video in a computer-readable recording medium, including: generating a bitstream from a video according to the method of any of the above solutions; and writing the bitstream into a computer-readable recording medium.
[0534] O5. A video processing apparatus including a processor, wherein the processor is configured to perform the method of any of the above solutions.
[0535] O6. A computer-readable medium having instructions stored thereon, which when executed cause a processor to perform the method of any of the above solutions.
[0536] O7. A computer-readable medium storing a bitstream generated by the method of any of the above solutions.
[0537] O8. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to perform the method of any of the above solutions.
[0538] O9. A bitstream generated using the method of any of the above solutions, which is stored on a computer-readable medium.
[0539] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this application document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The content disclosed in this specification and other embodiments can be implemented as one or more computer program products, that is, modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for a data processing apparatus to execute or control the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a substance composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing unit" or "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multi-processors or groups of computers. In addition to hardware, the apparatus may also include code for creating an execution environment for the computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, for example, an electrical, optical, or electromagnetic signal generated by a machine, which is generated to encode information for transmission to a suitable receiver device.
[0540] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (for example, one or more scripts in a markup language document), in a single file dedicated to the program, or in multiple coordinated files (for example, files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one or more computers, which are located at one site or distributed across multiple sites and interconnected by a communication network.
[0541] The processing and logical flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processing and logical flows can also be executed by special-purpose logic circuits, and the apparatus can also be implemented as special-purpose logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits).
[0542] For example, a processor suitable for executing a computer program includes general and special-purpose microprocessors, as well as any one or more of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CDROM and DVDROM disks. The processor and the memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0543] Although this patent document contains many details, it should not be construed as limiting any invention or the scope of any claims, but rather as a description of features of particular embodiments of a particular invention. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various functions described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Additionally, although the above features may be described as acting in certain combinations, and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0544] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood to mean that such operations must be performed in the particular order shown or in sequential order to obtain a desired result, or that all illustrated operations must be performed. Additionally, the separation of various system components described in the embodiments of this patent document should not be understood to be required in all embodiments.
[0545] Only some implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: performing a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to format rules, wherein the format rules conditionally allow, based on a first syntax element indicating whether inter-layer prediction is enabled and whether inter-layer reference pictures are available for use, encoding and decoding one or more pictures in a layer using reference pictures from other layers, and wherein the first syntax element conditionally exists in the bitstream based on a second syntax element indicating whether an identifier of a video parameter set (VPS) associated with the one or more pictures is not equal to zero, wherein the format rules stipulate that when a picture in the one or more pictures is a gradual decoding refresh (GDR) picture, all slices of the picture have a NAL unit type equal to the GDR network abstraction layer (NAL) unit type, and wherein the format rules stipulate implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the format rules stipulate that during the extraction process, when removing video coding layer (VCL) NAL units, padding data units associated with the VCL NAL units are also removed, and padding SEI messages in supplementary enhancement information (SEI) NAL units associated with the VCL NAL units are removed.
2. The method according to claim 1, wherein, the first syntax element is in a sequence parameter set SPS, and the first syntax element is sps_inter_layer_prediction_enabled_flag.
3. The method according to claim 1, wherein, the second syntax element is in a sequence parameter set SPS, and the second syntax element is sps_video_parameter_set_id.
4. The method according to claim 1, wherein, the format rules define a third syntax element for initiating (a) a synchronization process of context variables before decoding each coding tree unit (CTU) in each of the one or more pictures, and (b) a storage process of the context variables after decoding the CTU, wherein the third syntax element exists in a sequence parameter set SPS associated with the one or more pictures, and the third syntax element is not included in a picture parameter set PPS associated with the one or more pictures.
5. The method according to claim 4, wherein, the CTU includes a first coding tree block of one row of coding tree blocks in each slice of each picture.
6. The method according to claim 4, wherein, the third syntax element is sps_entropy_coding_sync_enabled_flag.
7. The method according to claim 1, wherein, The format rule further defines a fourth syntax element that is used to indicate whether signaling of an entry point offset for slices or slice-specified coding tree units (CTUs) rows is allowed to be present in the slice header of one of the one or more pictures. Wherein, the fourth syntax element is present in the sequence parameter set (SPS) associated with the one or more pictures, and the fourth syntax element is not included in the picture parameter set (PPS) associated with the one or more pictures.
8. The method according to claim 7, wherein, the format rule allows signaling of an entry point offset based on a value of the fourth syntax element equal to 1.
9. The method according to claim 7, wherein, the fourth syntax element is sps_entry_point_offsets_present_flag.
10. The method according to any one of claims 1-9, wherein the conversion includes decoding the video from the bitstream.
11. The method according to any one of claims 1-9, wherein, the conversion includes encoding the video into the bitstream.
12. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: perform a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to a format rule, wherein the format rule conditionally allows decoding and encoding of one or more pictures in one layer using reference pictures from other layers based on a first syntax element indicating whether inter-layer prediction is enabled and whether inter-layer reference pictures are available, and wherein the first syntax element is conditionally present in the bitstream based on a second syntax element that indicates whether an identifier of a video parameter set (VPS) associated with the one or more pictures is not equal to zero, wherein the format rule specifies that when a picture in the one or more pictures is a progressive decoding refresh (GDR) picture, all slices of the picture have a NAL unit type equal to the GDR network abstraction layer (NAL) unit type, and wherein the format rule specifies performing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the format rule specifies that during the extraction process, when removing video coding layer (VCL) NAL units, padding data units associated with the VCL NAL units are also removed, and padding SEI messages in supplementary enhancement information (SEI) NAL units associated with the VCL NAL units are removed.
13. The apparatus according to claim 12, wherein, the first syntax element is in the sequence parameter set (SPS), and the first syntax element is sps_inter_layer_prediction_enabled_flag; The second syntax element is in a sequence parameter set (SPS), and the second syntax element is sps_video_parameter_set_id.
14. The apparatus according to claim 12, wherein the formatting rule defines a third syntax element for initiating (a) a synchronization process of context variables before decoding each coding tree unit (CTU) in one or more pictures, and (b) a storage process of the context variables after decoding the CTU; wherein the formatting rule further defines a fourth syntax element for indicating whether signaling of an entry point offset for a slice or a slice-specified CTU row is allowed to exist in a slice header of one of the one or more pictures; wherein the third syntax element and the fourth syntax element are in the sequence parameter set (SPS) associated with the one or more pictures, and the third syntax element and the fourth syntax element are not included in the picture parameter set (PPS) associated with the one or more pictures.
15. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including one or more pictures and a bitstream of the video, wherein the bitstream conforms to a formatting rule, wherein the formatting rule conditionally allows decoding of one or more pictures in one layer using reference pictures from other layers based on a first syntax element indicating whether inter-layer prediction is enabled and whether inter-layer reference pictures are available, and wherein the first syntax element conditionally exists in the bitstream based on a second syntax element indicating whether an identifier of a video parameter set (VPS) associated with the one or more pictures is not equal to zero, wherein the formatting rule specifies that when a picture in the one or more pictures is a progressive decoding refresh (GDR) picture, all slices of the picture have a NAL unit type equal to the GDR network abstraction layer (NAL) unit type, and wherein the formatting rule specifies implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the formatting rule specifies that, during the extraction process, when removing video coding layer (VCL) NAL units, padding data units associated with the VCL NAL units are also removed, and padding SEI messages in supplementary enhancement information (SEI) NAL units associated with the VCL NAL units are removed.
16. The non-transitory computer-readable storage medium according to claim 15, wherein the first syntax element is in the sequence parameter set (SPS), and the first syntax element is sps_inter_layer_prediction_enabled_flag; The second syntax element is in a sequence parameter set (SPS), and the second syntax element is sps_video_parameter_set_id.
17. The non-transitory computer-readable storage medium according to claim 15, wherein the format rule defines a third syntax element for initiating (a) a synchronization process of context variables before decoding each coding tree unit (CTU) in each of the one or more pictures, and (b) a storage process of the context variables after decoding the CTU; wherein the format rule further defines a fourth syntax element for indicating whether signaling of an entry point offset for a slice or a slice-specified CTU row is allowed to exist in a slice header of a picture among the one or more pictures; wherein the third syntax element and the fourth syntax element are in a sequence parameter set (SPS) associated with the one or more pictures, and the third syntax element and the fourth syntax element are not included in a picture parameter set (PPS) associated with the one or more pictures.
18. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing apparatus, wherein the method includes: generating a bitstream of a video including one or more pictures, wherein the bitstream conforms to a format rule, wherein the format rule conditionally allows decoding of one or more pictures in one layer using reference pictures from other layers based on a first syntax element indicating whether inter-layer prediction is enabled and whether inter-layer reference pictures are available, and wherein the first syntax element conditionally exists in the bitstream based on a second syntax element indicating whether an identifier of a video parameter set (VPS) associated with the one or more pictures is not equal to zero, wherein the format rule specifies that when a picture among the one or more pictures is a progressive decoding refresh (GDR) picture, all slices of the picture have a NAL unit type equal to the GDR network abstraction layer (NAL) unit type, and wherein the format rule specifies implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the format rule specifies that, during the extraction process, when removing video coding layer (VCL) NAL units, padding data units associated with the VCL NAL units are also removed, and padding SEI messages in supplementary enhancement information (SEI) NAL units associated with the VCL NAL units are removed.
19. The non-transitory computer-readable recording medium according to claim 18, wherein the first syntax element is in a sequence parameter set (SPS), and the first syntax element is sps_inter_layer_prediction_enabled_flag; The second syntax element is in a sequence parameter set (SPS), and the second syntax element is sps_video_parameter_set_id.
20. The non-transitory computer-readable recording medium according to claim 18, wherein, the format rule defines a third syntax element for initiating (a) a synchronization process of context variables before decoding each coding tree unit (CTU) in one or more pictures, and (b) a storage process of the context variables after decoding the CTU, wherein the format rule further defines a fourth syntax element for indicating whether signaling for an entry point offset for a slice or a slice-specified CTU row is allowed to exist in a strip header of one of the one or more pictures, wherein the third syntax element and the fourth syntax element are in a sequence parameter set (SPS) associated with the one or more pictures, and the third syntax element and the fourth syntax element are not included in a picture parameter set (PPS) associated with the one or more pictures.
21. A method for storing a bitstream of video, comprising: generating a bitstream of video including one or more pictures, and storing the bitstream in a non-transitory computer-readable recording medium, wherein, the bitstream conforms to a format rule, wherein the format rule conditionally allows decoding one or more pictures in one layer using a reference picture from another layer based on a first syntax element indicating whether inter-layer prediction is enabled and whether an inter-layer reference picture is available, and wherein the first syntax element conditionally exists in the bitstream based on a second syntax element indicating whether an identifier of a video parameter set (VPS) associated with the one or more pictures is not equal to zero, wherein the format rule specifies that when a picture in the one or more pictures is a progressive decoding refresh (GDR) picture, all strips of the picture have a NAL unit type equal to the GDR network abstraction layer (NAL) unit type, and wherein the format rule specifies implementing a sub-bitstream extraction process to generate a sub-bitstream for decoding, wherein the sub-bitstream extraction process is configured to extract a sub-bitstream having a target highest temporal identifier from the bitstream, and wherein the format rule specifies that, during the extraction process, when removing video coding layer (VCL) NAL units, padding data units associated with the VCL NAL units are also removed, and padding SEI messages in supplementary enhancement information (SEI) NAL units associated with the VCL NAL units are removed.
Citation Information
Patent Citations
Video encoding and decoding
US20150156501A1
Parameter set coding
US20150195577A1