Mixed nal unit picture constraints in video coding
Patent Information
- Application Number
- KR1020257004453
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-08
- Filing Date
- 2020-07-07
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2040-07-07
Smart Images

Figure 112025015701814-PAT00006_ABST
Abstract
Description
Technology Field
[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 871,524 (Constraints for Mixed NAL Unit Types within One Picture in Video Coding), filed by Ye-Kui Wang et al. on July 8, 2019, which is incorporated herein by reference.
[0002] The present disclosure generally relates to video coding, and specifically to sub-picture coding of a picture in video coding. Background Technology
[0003] Even for relatively short videos, the amount of video data required to depict them can be substantial, which can cause difficulties when data is streamed or transmitted over communication networks with limited bandwidth. Therefore, video data is typically compressed before being transmitted over modern communication networks. Since memory resources may be limited, the size of the video can also be an issue when it is stored on a storage device. Video compression devices often use software and / or hardware at the source to code video data before transmission or storage, which can reduce the amount of data required to represent a digital video image. Subsequently, the compressed data is received at the destination by a video decompression device that decodes the video data. As network resources remain limited and the demand for higher video quality continues to grow, there is a need for improved compression and decompression technologies that enhance compression ratios with little to no sacrifice in image quality.
[0004] In an embodiment, the present disclosure comprises a method implemented in a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream comprising a current picture comprising a plurality of video coding layer (VCL) NAL units that do not have the same network abstraction layer (NAL) unit type; by a processor of the decoder, obtaining an active entry of a reference picture list for a slice located in a subpicA of a subsequent picture following the current picture in a decoding order — the active entry does not include a reference to any reference picture preceding the current picture in a decoding order where the subpicA of the current picture is associated with an intra-random access point (IRAP) NAL unit type —; and by a processor decoding the subsequent picture based on the active entry of the reference picture list; and by a processor delivering the subsequent picture for display as part of a decoded video sequence.
[0005] Video coding systems may also encode video by using IRAP and non-IRAP pictures. An IRAP picture is a picture coded according to intra-prediction that acts as a random access point for a video sequence. An IRAP picture can be decoded without first decoding any other picture. Therefore, a decoder can start decoding a video sequence from any IRAP picture. In contrast, a decoder generally cannot start decoding a video sequence from a non-IRAP picture. An IRAP picture can also refresh the DPB. This is because an IRAP picture can act as the starting point of a coded video sequence (CVS), and the picture in the CVS does not refer to the picture in the previous CVS. As such, an IRAP picture can also break / stop the inter-prediction chain and stop inter-prediction-related coding errors because such errors cannot propagate through the IRAP picture.
[0006] In some cases, video coding systems can be used to code virtual reality (VR) video. VR video may contain a sphere of video content displayed as if the user were at the center of the sphere. Only a portion of the sphere, referred to as the viewport, is displayed to the user. The rest of the picture is discarded without being rendered. Typically, the entire picture is transmitted so that different viewports can be dynamically selected and displayed in response to the user's head movements. With this approach, the video file size can become very large. To improve coding efficiency, some systems split the picture into sub-pictures. The video can be encoded in two or more resolutions. Each resolution is encoded into a different set of sub-bitstreams corresponding to the sub-picture. When a user streams VR video, the coding system can merge the sub-bitstreams into a bitstream for transmission based on the current viewport the user is using. Specifically, the current viewport is acquired from a high-resolution sub-bitstream, and unviewed viewports are acquired from low-resolution bitstream(s). In this way, the highest quality video is displayed to the user, and the lower quality video is discarded. If the user selects a new viewport, a lower-resolution video is displayed to the user. The decoder can request the new viewport to receive a higher-resolution video. The encoder can modify the merging process accordingly. Upon reaching an IRAP picture, the decoder can begin decoding the high-resolution video sequence in the new viewport. This approach significantly increases video compression without negatively impacting the user's viewing experience.
[0007] One concern regarding the aforementioned approach is that the time required to change resolution is based on the time taken to reach an IRAP picture. This is because the decoder cannot initiate decoding of different video sequences from non-IRAP pictures, as previously mentioned. One approach to reducing this latency is to include more IRAP pictures. However, this increases file size. To balance coding efficiency and functionality, different viewports / subpictures may include IRAP pictures at different frequencies. For example, a viewport more likely to be viewed may have more IRAP pictures than other viewports. This approach presents another problem. Specifically, a picture following an IRAP picture is restricted from referencing a picture preceding it. However, these constraints are applied at the picture level. A picture containing a mixed NAL unit that includes both IRAP and non-IRAP subpictures may not be considered an IRAP picture at the picture level. Therefore, these picture-level constraints may not apply. This can lead to a portion of a picture following an IRAP subpicture improperly referencing a picture preceding the IRAP picture. In this case, the IRAP subpicture will not function properly as an access point because the reference picture / subpicture may not be available, which will prevent the subpicture following the IRAP subpicture from being decoded. Additionally, the IRAP subpicture must not prevent such references that would render the purpose of having a NAL unit mixed with non-IRAP subpictures (e.g., an intercoded sequence of different lengths depending on the subpicture position) invalid.
[0008] This example includes a mechanism for mitigating coding errors when a picture contains both IRAP NAL units and non-IRAP NAL units. Specifically, a subpicture of the current picture may contain an IRAP NAL unit. When this occurs, slices of the picture following the current picture, which are also contained in the subpicture, are restricted from referencing a reference picture prior to the current picture. This ensures that the IRAP NAL unit stops inter-prediction propagation at the subpicture level. Thus, the decoder can begin decoding from the IRAP subpicture. Slices associated with the subpicture of the later picture can always be decoded because these slices do not refer to data prior to the IRAP subpicture (which is not decoded). These constraints do not apply to non-IRAP NAL units. Therefore, inter-prediction is not broken for subpictures containing non-IRAP data. As such, the disclosed mechanism allows for the implementation of additional features. For example, the disclosed mechanism supports dynamic resolution changes at the subpicture level when using subpicture bitstreams. Therefore, the disclosed mechanism allows for the transmission of lower-resolution sub-picture bitstreams when streaming VR video without significantly degrading the user experience. Consequently, the disclosed mechanism increases coding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0009] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, where the IRAP NAL unit type is a clean Random Access (CRA) NAL unit type.
[0010] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, wherein the IRAP NAL unit type is an Instantaneous Decoder Refresh (IDR) NAL unit type.
[0011] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, and further includes a step in which the processor determines that all slices of the current picture located in subpicA are associated with the same NAL unit type.
[0012] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, and further includes the step of determining by the processor that a first NAL unit type value for the VCL NAL unit of the current picture is different from a second NAL unit type value for the VCL NAL unit of the current picture based on a flag.
[0013] Optionally, in any one of the preceding aspects, another initiation of this aspect is provided, wherein the bitstream includes a picture parameter set (PPS) and a flag is obtained from the PPS.
[0014] Optionally, in any of the preceding aspects, another disclosure of this aspect is provided, where the flag is mixed_nalu_types_in_pic_flag, and mixed_nalu_types_in_pic_flag is equal to 1 if each picture referencing the PPS has one or more VCL NAL units and the VCL NAL units do not have the same value of NAL unit type (nal_unit_type).
[0015] In an embodiment, the present disclosure comprises a method implemented in an encoder, the method comprising: a step by which a processor of the encoder determines that a current picture comprises a plurality of VCL NAL units that do not have the same NAL unit type; a step by which a processor determines that a subpicA of the current picture is associated with an IRAP NAL unit type; a step by which a processor generates an active entry of a reference picture list for a slice located in a subpicA of a subsequent picture following the current picture in a decoding order — the active entry does not include a reference to any reference picture preceding the current picture in a decoding order when the subpicA of the current picture is associated with an IRAP NAL unit type —; a step by which a processor encodes the subsequent picture into a bitstream based on the reference picture list; and a step by which a memory coupled to the processor stores the bitstream for communication to a decoder.
[0016] Video coding systems can also encode video by using IRAP and non-IRAP pictures. An IRAP picture is a picture coded according to intra-prediction that acts as a random access point for a video sequence. An IRAP picture can be decoded without first decoding any other picture. Therefore, a decoder can start decoding a video sequence from any IRAP picture. In contrast, a decoder generally cannot start decoding a video sequence from a non-IRAP picture. IRAP pictures can also refresh the DPB. This is because an IRAP picture can act as the starting point of a coded video sequence (CVS), and the picture in the CVS does not refer to the picture in the previous CVS. As such, an IRAP picture can also break the inter-prediction chain and stop inter-prediction-related coding errors because such errors cannot propagate through the IRAP picture.
[0017] In some cases, video coding systems can be used to code virtual reality (VR) video. VR video may contain a sphere of video content that is displayed as if the user were at the center of the sphere. Only a portion of the sphere, referred to as the viewport, is displayed to the user. The rest of the picture is discarded without being rendered. Typically, the entire picture is transmitted so that different viewports can be dynamically selected and displayed in response to the user's head movements. With this approach, the video file size can become very large. To improve coding efficiency, some systems split the picture into sub-pictures. The video can be encoded in two or more resolutions. Each resolution is encoded into a different set of sub-bitstreams corresponding to the sub-picture. When a user streams VR video, the coding system can merge the sub-bitstreams into a bitstream for transmission based on the current viewport the user is using. Specifically, the current viewport is acquired from a high-resolution sub-bitstream, and unviewed viewports are acquired from low-resolution bitstream(s). In this way, the highest quality video is displayed to the user, and the lower quality video is discarded. If the user selects a new viewport, a lower-resolution video is displayed to the user. The decoder can request the new viewport to receive a higher-resolution video. The encoder can modify the merging process accordingly. Upon reaching an IRAP picture, the decoder can begin decoding the high-resolution video sequence in the new viewport. This approach significantly increases video compression without negatively impacting the user's viewing experience.
[0018] One concern regarding the aforementioned approach is that the time required to change resolution is based on the time taken to reach an IRAP picture. This is because the decoder cannot initiate decoding of different video sequences from non-IRAP pictures, as previously mentioned. One approach to reducing this latency is to include more IRAP pictures. However, this increases file size. To balance coding efficiency and functionality, different viewports / subpictures may include IRAP pictures at different frequencies. For example, a viewport more likely to be viewed may have more IRAP pictures than other viewports. This approach presents another problem. Specifically, a picture following an IRAP picture is restricted from referencing a picture preceding it. However, these constraints are applied at the picture level. A picture containing a mixed NAL unit that includes both IRAP and non-IRAP subpictures may not be considered an IRAP picture at the picture level. Therefore, these picture-level constraints may not apply. This can lead to a portion of a picture following an IRAP subpicture improperly referencing a picture preceding the IRAP picture. In this case, the IRAP subpicture will not function properly as an access point because the reference picture / subpicture may not be available, which will prevent the subpicture following the IRAP subpicture from being decoded. Additionally, the IRAP subpicture must not prevent such references that would render the purpose of having a NAL unit mixed with non-IRAP subpictures (e.g., an intercoded sequence of different lengths depending on the subpicture position) invalid.
[0019] This example includes a mechanism for mitigating coding errors when a picture contains both IRAP NAL units and non-IRAP NAL units. Specifically, a subpicture of the current picture may contain an IRAP NAL unit. When this occurs, slices of the picture following the current picture, which are also contained in the subpicture, are restricted from referencing a reference picture prior to the current picture. This ensures that the IRAP NAL unit stops inter-prediction propagation at the subpicture level. Thus, the decoder can begin decoding from the IRAP subpicture. Slices associated with the subpicture of the later picture can always be decoded because these slices do not refer to data prior to the IRAP subpicture (which is not decoded). These constraints do not apply to non-IRAP NAL units. Therefore, inter-prediction is not broken for subpictures containing non-IRAP data. As such, the disclosed mechanism allows for the implementation of additional features. For example, the disclosed mechanism supports dynamic resolution changes at the subpicture level when using subpicture bitstreams. Therefore, the disclosed mechanism allows for the transmission of lower-resolution sub-picture bitstreams when streaming VR video without significantly degrading the user experience. Consequently, the disclosed mechanism increases coding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0020] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, where the IRAP NAL unit type is the CRA NAL unit type.
[0021] Optionally, in any of the preceding aspects, another disclosure of this aspect is provided, where the IRAP NAL unit type is the IDR NAL unit type.
[0022] Optionally, in any one of the preceding aspects, another initiation of this aspect is provided, and further includes the step of encoding the current picture into a bitstream by the processor, ensuring that all slices of the current picture located in subpicA are associated with the same NAL unit type.
[0023] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, and further includes the step of encoding a flag into a bitstream by the processor that indicates that a first NAL unit type value for a VCL NAL unit of the current picture is different from a second NAL unit type value for a VCL NAL unit of the current picture.
[0024] Optionally, in any of the previous aspects, another initiation of this aspect is provided, and the flag is encoded as the bitstream's PPS.
[0025] Optionally, in any of the preceding aspects, another initiation of this aspect is provided, and the flag is mixed_nalu_types_in_pic_flag, and mixed_nalu_types_in_pic_flag is set to 1 if each picture referencing the PPS has one or more VCL NAL units and the VCL NAL units do not have the same value of nal_unit_type.
[0026] In an embodiment, the present disclosure comprises a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter comprise a video coding device configured to perform any one of the prior aspects.
[0027] In an embodiment, the present disclosure comprises a computer program product for use by a video coding device, and the computer program product comprises a non-transient computer-readable medium comprising computer-executable instructions stored in a non-transient computer-readable medium to enable the video coding device to perform any one of the preceding aspects when executed by a processor.
[0028] In an embodiment, the present disclosure comprises a decoder including: receiving means for receiving a bitstream comprising a current picture comprising a plurality of VCL NAL units that do not have the same NAL unit type; acquiring means for acquiring an active entry of a reference picture list for a slice located in a subpicA of a subsequent picture following the current picture in a decoding order; determining means for determining that, if the subpicA of the current picture is associated with an IRAP NAL unit type, the active entry does not contain a reference to any reference picture preceding the current picture in a decoding order; decoding means for decoding the subsequent picture based on the active entry of the reference picture list; and a transfer means for transferring the subsequent picture to be displayed as part of a decoded video sequence.
[0029] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, wherein the decoder is additionally configured to perform the method of any one of the preceding aspects.
[0030] In an embodiment, the present disclosure comprises: a determination means—the determination means is for determining that the current picture comprises a plurality of VCL NAL units that do not have the same NAL unit type and determining that the subpicA of the current picture is associated with an IRAP NAL unit type—; a generation means for generating an active entry of a reference picture list for a slice located in the subpicA of a subsequent picture following the current picture in a decoding order—the active entry does not include a reference to any reference picture preceding the current picture in a decoding order when the subpicA of the current picture is associated with an IRAP NAL unit type—; an encoding means for encoding the subsequent picture into a bitstream based on the reference picture list; and a storage means for storing a bitstream for communication to a decoder.
[0031] Optionally, in any one of the preceding aspects, another disclosure of this aspect is provided, wherein the encoder is additionally configured to perform the method of any one of the preceding aspects.
[0032] For clarity, any one of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create a new embodiment within the scope of the present disclosure.
[0033] These and other features will be more clearly understood from the following detailed description taken in relation to the attached drawings and claims. Brief explanation of the drawing
[0034] For a more complete understanding of the present disclosure, the following brief description taken in connection with the accompanying drawings and detailed description is referenced, wherein the same reference numerals indicate the same parts. Figure 1 is a flowchart of an exemplary method for coding a video signal. Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system for video coding. Figure 3 is a schematic diagram illustrating an exemplary video encoder. Figure 4 is a schematic diagram illustrating an exemplary video decoder. FIG. 5 is a schematic diagram illustrating multiple sub-picture video streams split from a virtual reality (VR) picture video stream. FIG. 6 is a schematic diagram illustrating the constraints used when the current picture includes a mixed Network Abstraction Layer (NAL) unit type. FIG. 7 is a schematic diagram illustrating an exemplary reference picture list structure including a reference picture list. FIG. 8 is a schematic diagram illustrating an exemplary bitstream including a picture having a mixed NAL unit type. Figure 9 is a schematic diagram of an exemplary video coding device. FIG. 10 is a flowchart of an exemplary method for encoding a video sequence containing a picture having a mixed NAL unit type into a bitstream. FIG. 11 is a flowchart of an exemplary method for decoding a video sequence containing a picture having a mixed NAL unit type from a bitstream. FIG. 12 is a schematic diagram of an exemplary system for coding a video sequence containing a picture having a mixed NAL unit type into a bitstream. Specific details for implementing the invention
[0035] While exemplary implementations of one or more embodiments are provided below, it should be understood that the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The present disclosure should not be limited to the exemplary implementations, drawings, and techniques depicted below, including the exemplary designs and implementations illustrated and described herein, and may be modified within the scope of the appended claims, along with the full range of equivalents.
[0036] The following terms are defined as follows unless otherwise specified herein. Specifically, the following definitions are intended to provide additional clarity to the present disclosure. However, terms may be described differently depending on the context. Accordingly, the following definitions should be considered supplementary and should not be construed as limiting other definitions of these terms provided herein.
[0037] A bitstream is a sequence of bits containing video data that is compressed for transmission between an encoder and a decoder. An encoder is a device configured to use an encoding process to compress video data into a bitstream. A decoder is a device configured to use a decoding process to reconstruct video data from a bitstream for display. A picture is an array of lumina samples and / or chroma samples that generate a frame or its field. For convenience of description, a picture being encoded or decoded may be referred to as the current picture, and a picture following the current picture may be referred to as the subsequent picture. A subpicture is a rectangular area of one or more slices within a picture sequence. Since a square is a type of rectangle, a subpicture may contain a square area. A slice is an integer representing a complete tile within a picture tile contained exclusively in a single network abstraction layer (NAL) unit, or an integer representing a consecutive complete coding tree unit (CTU) row. A NAL unit is a syntax structure containing data bytes and instructions for the data types contained therein. NAL units include video coding layer (VCL) NAL units containing video data and non-VCL NAL units containing supporting syntax data. The NAL unit type is the type of data structure contained in the NAL unit. The intra-random access point (IRAP) NAL unit type is a data structure containing data from an IRAP picture or sub-picture. An IRAP picture / sub-picture is a picture / sub-picture coded according to intra-prediction, which indicates that the decoder can start video sequence decoding at the corresponding picture / sub-picture without referring to the picture preceding the IRAP picture / sub-picture.The Clean Random Access (CRA) NAL unit type is a data structure containing data from a CRA picture or subpicture. A CRA picture / subpicture is an IRAP picture / subpicture that does not refresh the decoded picture buffer (DPB). The Instantaneous Decoding Refresh (IDR) NAL unit type is a data structure containing data from an IDR picture or subpicture. An IDR picture / subpicture is an IRAP picture / subpicture that refreshes the DPB. A reference picture is a picture containing reference samples that can be used when coding another picture as a reference according to inter-prediction. A reference picture list is a list of reference pictures used for inter-prediction and / or inter-layer prediction. Some video coding systems refer to two picture lists, which can be denoted as Reference Picture List 1 and Reference Picture List 0. The reference picture list structure is an addressable syntax structure containing multiple reference picture lists. An active entry is an entry in the reference picture list that refers to a reference picture available to the current picture when inter-predicting is performed. A flag is a data structure containing a sequence of bits that can be set to indicate corresponding data. A Picture Parameter Set (PPS) is a set of parameters containing picture-level data associated with one or more pictures. The decoding order is the order in which syntax elements are processed by the decoding process. A decoded video sequence is a picture sequence reconstructed by a decoder for display to a user.
[0038] The following acronyms, Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Instantaneous Decoding Refresh (IDR), Intra-Random Access Point (IRAP), Least Significant Bit (LSB), Most Significant Bit (MSB), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), and Working Draft (WD) are used here.
[0039] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques may include performing spatial (e.g., within a picture) prediction and / or temporal (e.g., between pictures) prediction to reduce or eliminate data redundancy in video sequences. In the case of block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be divided into video blocks, which may be referred to as tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU), and / or coding nodes. Video blocks of an intra-coded (I) slice of a picture are coded using spatial prediction for reference samples in neighboring blocks of the same picture. Video blocks of an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded using spatial prediction for reference samples in neighboring blocks of the same picture or temporal prediction for reference samples in other reference pictures. A picture may be referred to as a frame and / or image, and a reference picture may be referred to as a reference frame and / or reference image. Spatial or temporal prediction generates prediction blocks representing image blocks. Residual data represents the pixel difference between the original image block and the prediction block. Thus, intercoded blocks are encoded according to motion vectors pointing to blocks of reference samples forming the prediction block, and residual data indicating the difference between the coded block and the prediction block. Intracoded blocks are encoded according to the intracoding mode and residual data. For further compression, residual data can be transformed from the pixel domain to the transform domain. This generates residual transform coefficients that can be quantized. The quantized transform coefficients can initially be arranged in a two-dimensional array.Quantized transform coefficients can be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding can be applied to achieve greater compression. These video compression techniques are described in detail below.
[0040] To ensure that encoded video can be accurately decoded, the video is encoded and decoded according to the corresponding video coding standards. Video coding standards include ITU (International Telecommunication Union) Standardization Sector (ITU-T) H.261, ISO / IEC (International Organization for Standardization / International Electrotechnical Commission) MPEG (Motion Picture Experts Group)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, AVC (Advanced Video Coding) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and HEVC (High Efficiency Video Coding) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extensions such as SVC (Scalable Video Coding), MVC (Multiview Video Coding), MVC+D (Multiview Video Coding plus Depth), and three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as SHVC (Scalable HEVC), MV-HEVC (Multiview HEVC), and 3D HEVC. The Joint Video Experts Team (JVET) of the ITU-T and ISO / IEC began developing a video coding standard referred to as VVC (Versatile Video Coding). VVC is included in the Working Draft (WD) containing JVET-L1001-v10.
[0041] A video coding system can encode video by using IRAP pictures and non-IRAP pictures. An IRAP picture is a picture coded according to intra-prediction, which acts as a random access point for a video sequence. In intra-prediction, blocks of a picture are coded by referencing other blocks within the same picture. This contrasts with non-IRAP pictures, which use inter-prediction. In inter-prediction, blocks of the current picture are coded by referencing other blocks in a reference picture different from the current picture. Because IRAP pictures are coded without referencing other pictures, they can be decoded without first decoding any other picture. Therefore, a decoder can start decoding a video sequence from any IRAP picture. In contrast, non-IRAP pictures are coded by referencing other pictures, and thus decoders generally cannot start decoding a video sequence from non-IRAP pictures. IRAP pictures can also refresh DPBs. This is because an IRAP picture can serve as the starting point for a CVS, and the picture in the CVS does not refer to the picture of the previous CVS. As such, IRAP pictures can also stop inter-prediction-related coding errors because such errors cannot propagate through the IRAP picture. However, IRAP pictures are much larger than non-IRAP pictures in terms of data size. As such, video sequences typically contain many non-IRAP pictures along with a smaller number of interspersed IRAP pictures to balance coding efficiency and functionality. For example, a 60-frame CVS may contain one IRAP picture and 59 non-IRAP pictures.
[0042] In some cases, video coding systems can be used to code virtual reality (VR) video, which may also be referred to as 360-degree video. VR video may contain a sphere of video content displayed as if the user were at the center of the sphere. Only a portion of the sphere, referred to as a viewport, is displayed to the user. For example, a user may use a Head Mounted Display (HMD) that selects and displays a viewport of the sphere based on the user's head movements. This provides the sensation of being physically present in the virtual space, as depicted in the video. To achieve this result, each picture in the video sequence contains the entire sphere of video data at the corresponding point in time. However, only a small portion of the picture (e.g., a single viewport) is displayed to the user. The rest of the picture is discarded without being rendered. The entire picture is typically transmitted so that different viewports can be dynamically selected and displayed in response to the user's head movements. This approach can result in very large video file sizes.
[0043] To improve coding efficiency, some systems divide a picture into sub-pictures. A sub-picture is a defined spatial region of a picture. Each sub-picture contains the viewport of the corresponding picture. Video can be encoded in two or more resolutions. Each resolution is encoded into a different sub-bitstream. When a user streams VR video, the coding system can merge the sub-bitstreams into a bitstream for transmission based on the current viewport being used by the user. Specifically, the current viewport is acquired from a high-resolution sub-bitstream, and viewports not being viewed are acquired from low-resolution bitstream(s). In this way, the highest quality video is displayed to the user, and the lower quality video is discarded. If the user selects a new viewport, the lower-resolution video is displayed to the user. The decoder may request that the new viewport receive the higher-resolution video. The encoder may change the merging process accordingly. When an IRAP picture is received, the decoder may begin decoding the high-resolution video sequence in the new viewport. This approach significantly increases video compression without negatively impacting the user's viewing experience.
[0044] One concern regarding the aforementioned approach is that the time required to change resolution is based on the time it takes to reach an IRAP picture. This is because the decoder cannot initiate decoding of different video sequences from non-IRAP pictures, as previously mentioned. One approach to reducing this latency is to include more IRAP pictures. However, this increases file size. To balance coding efficiency and functionality, different viewports / sub-pictures may contain IRAP pictures at different frequencies. For example, viewports more likely to be viewed may have more IRAP pictures than other viewports. For instance, in a basketball scenario, viewports related to the basket and / or center court may contain IRAP pictures at a higher frequency than viewports looking at the stands or ceiling, as these viewports are less likely to be viewed by the user.
[0045] This approach gives rise to other problems. Specifically, a picture following an IRAP picture is restricted from referencing a picture preceding the IRAP picture. However, these constraints are applied at the picture level. A picture containing a mixed NAL unit that includes both IRAP and non-IRAP subpictures may not be considered an IRAP picture at the picture level. Consequently, these picture-level constraints may not apply. This can lead to a portion of a picture following an IRAP subpicture that improperly references a picture preceding the IRAP picture. In this case, the IRAP subpicture may not function properly as an access point because the referenced picture / subpicture may be unavailable, which would prevent the subpicture following the IRAP subpicture from being decoded. Furthermore, an IRAP subpicture should not prevent such references that could nullify the purpose of having a mixed NAL unit (e.g., an intercoded sequence of different lengths depending on the subpicture position) with a non-IRAP subpicture.
[0046] A mechanism for mitigating coding errors is disclosed herein when a picture contains both IRAP NAL units and non-IRAP NAL units. Specifically, a subpicture of the current picture may contain an IRAP NAL unit. When this occurs, slices in the picture following the current picture, which are also contained in the subpicture, are restricted from referencing a reference picture prior to the current picture. This causes the IRAP NAL unit to stop inter-predicting at the subpicture level (e.g., stopping the inter-predicting reference chain). Thus, the decoder can begin decoding in the IRAP subpicture. Slices associated with the subpicture of the later picture can always be decoded because these slices do not refer to data prior to the IRAP subpicture (which is not decoded). These constraints do not apply to non-IRAP NAL units. Therefore, inter-predicting is not stopped for subpictures containing non-IRAP data. As such, the disclosed mechanism allows for the implementation of additional functionality. For example, the disclosed mechanism supports dynamic resolution changes at the sub-picture level when using sub-picture bitstreams. Therefore, the disclosed mechanism allows lower-resolution sub-picture bitstreams to be transmitted when streaming VR video without significantly degrading the user experience. Consequently, the disclosed mechanism increases coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in encoders and decoders.
[0047] FIG. 1 is a flowchart of an exemplary operation method (100) for encoding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal using various mechanisms to reduce the video file size. The smaller the file size, the more the associated bandwidth overhead can be reduced while transmitting the compressed video file to the user. Then, a decoder decodes the compressed video file to reconstruct the original video signal to be displayed to the end user. The decoding process generally reflects the encoding process so that the decoder can consistently reconstruct the video signal.
[0048] In step 101, the video signal is input into the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device, such as a video camera, and may be encoded to support live streaming of the video. The video file may contain both an audio component and a video component. The video component contains a series of image frames that give a visual impression of motion when viewed sequentially. The frames contain pixels represented in terms of light, referred to herein as the luminance component (or luminance sample), and colors referred to as the chroma component (color sample). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0049] In step 103, the video is partitioned into blocks. Partitioning involves subdividing the pixels of each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be partitioned into Coding Tree Units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU contains both luminance and chroma samples. A coding tree can be used to partition the CTU into blocks and then recursively subdivide the blocks until a configuration supporting further encoding is achieved. For example, the luminance component of a frame can be subdivided until individual blocks contain relatively uniform lighting values. Additionally, the chroma component of a frame can be subdivided until individual blocks contain relatively uniform color values. Thus, the partitioning mechanism depends on the content of the video frame.
[0050] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, a block describing an object in a reference frame does not need to be described repeatedly in adjacent frames. Specifically, objects such as tables may remain in a constant position across multiple frames. Thus, a table is described once, and adjacent frames may refer back to the reference frame. Pattern matching mechanisms may be used to match objects across multiple frames. Additionally, moving objects may be represented across multiple frames due to, for example, object movement or camera movement. As a specific example, a video may show a car moving across the screen over multiple frames. Motion vectors may be used to describe this movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of an object in a reference frame. In this way, inter-prediction can encode an image block of the current frame as a set of motion vectors indicating an offset from the corresponding block of the reference frame.
[0051] Intra prediction encodes blocks of a common frame. Intra prediction leverages the fact that luminance and chroma components tend to cluster within a frame. For example, green patches within a part of a tree tend to be located next to similar green patches. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. The directional mode indicates that the current block is similar to / identical to neighboring block samples in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the edges of the row. In effect, the planar mode indicates smooth light / color transitions across rows / columns by using a relatively constant gradient when changing values. The DC mode is used for boundary smoothing and indicates that the block is similar to / identical to the average value associated with all neighboring block samples associated with the angular direction of the directional prediction mode. Therefore, intra prediction blocks can represent image blocks with various relational prediction mode values instead of actual values. Additionally, intra prediction blocks can represent image blocks with motion vector values instead of actual values. In both cases, the prediction block may not accurately represent the image block depending on the circumstances. All differences are stored in the residual block. Transformations can be applied to the residual block to further compress the file.
[0052] In step 107, various filtering techniques may be applied. In HEVC, filters are applied according to an in-loop filtering method. The aforementioned block-based prediction may result in the generation of block images in the decoder. Additionally, the block-based prediction method may reconstruct the encoded blocks after encoding them to be used later as reference blocks. The in-loop filtering method repeatedly applies noise suppression filters, de-blocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters can mitigate these blocking artifacts, allowing the encoded file to be accurately reconstructed. Furthermore, these filters can mitigate artifacts in the reconstructed reference blocks, reducing the likelihood that artifacts will generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.
[0053] Once the video signal is split, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream contains the data described above as well as desired signaling data to support appropriate video signal reconstruction in the decoder. For example, such data may include split data, prediction data, residual blocks, and various flags that provide coding commands to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. The generation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of discussion and is not intended to restrict the video coding process to a specific order.
[0054] The decoder receives the bitstream in step 111 and starts the decoding process. Specifically, the decoder uses an entropy decoding method to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the partitioning for the frame. The partitioning must match the result of the block partitioning in step 103. Now, the entropy encoding / decoding used in step 111 is described. The encoder makes many choices during the compression process, such as selecting a block partitioning method from several possible choices based on the spatial location of values in the input image(s). Many bins can be used to signal the correct choice. As used here, a bin is a binary value treated as a variable (e.g., a bit value that may vary depending on the context). Entropy coding allows the encoder to discard options that are clearly unfeasible in certain cases, leaving a set of acceptable options. Then, a codeword is assigned to each acceptable option. The length of the codeword is based on the quantity of acceptable options (e.g., one bin for two options, two bins for three to four options, etc.). Then, the encoder encodes the codeword for the selected option. This method reduces the size of the codeword because the codeword is as large as desired to uniquely indicate the selection from a small subset of acceptable options, as opposed to uniquely indicating the selection from a potentially large set of all possible options. Then, the decoder decodes the selection by determining the set of acceptable options in a manner similar to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0055] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate a residual block. Then, the decoder uses the residual block and the corresponding prediction block to reconstruct the image block according to the partition. The prediction block may include both an intra-prediction block and an inter-prediction block, as generated by the encoder in step 105. Subsequently, the reconstructed image block is positioned as a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax in step 113 may also be signaled as a bitstream via entropy coding as previously described.
[0056] In step 115, filtering is performed on frames of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal may be output to a display in step 117 so that the end user can view it.
[0057] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system (200) for video coding. Specifically, the codec system (200) provides a function that supports the implementation of the operation method (100). The codec system (200) is generalized to describe components used in both the encoder and the decoder. The codec system (200) receives and splits a video signal as discussed in relation to steps 101 and 103 in the operation method (100), which generates a split video signal (201). Then, the codec system (200) compresses the split video signal (201) into a coded bitstream when operating as an encoder as described in relation to steps 105, 107, and 109 in the method (100). When operating as a decoder, the codec system (200) generates an output video signal from the bitstream as discussed in steps 111, 113, 115, and 117 in the method of operation (100). The codec system (200) includes a general coder control component (211), a transform scaling and quantization component (213), an intra-picture estimation component (215), an intra-picture prediction component (217), a motion compensation component (219), a motion estimation component (221), a scaling and inverse transform component (229), a filter control analysis component (227), an in-loop filter component (225), a decoded picture buffer component (223), and a header formatting and context-adaptive binary arithmetic coding (CABAC) component (231). These components are combined as illustrated. In FIG. 2, the black line indicates the movement of data to be encoded / decoded, and the dotted line indicates the movement of control data that controls the operation of other components. All components of the codec system (200) may exist in the encoder. The decoder may include a subset of the components of the codec system (200).For example, the decoder may include an in-picture prediction component (217), a motion compensation component (219), a scaling and inverse transformation component (229), an in-loop filter component (225), and a decoded picture buffer component (223). These components will now be described.
[0058] The segmented video signal (201) is a captured video sequence that has been segmented into pixel blocks by a coding tree. The coding tree uses various segmentation modes to subdivide the pixel blocks into smaller pixel blocks. These blocks may be referred to as nodes in the coding tree. Larger parent nodes are segmented into smaller child nodes. The number of times a node is segmented is referred to as the depth of the node / coding tree. The segmented blocks may, in some cases, be included in a coding unit (CU). For example, a CU may be a sub-part of a CTU that includes a luminance block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block, along with corresponding syntax instructions for the CU. The segmentation modes may include binary trees (BT), triple trees (TT), and quad trees (QT) used to segment the nodes into two, three, or four child nodes of various shapes, respectively, depending on the segmentation mode used. The segmented video signal (201) is passed to a general coder control component (211), a transform scaling and quantization component (213), an intra-picture estimation component (215), a filter control analysis component (227), and a motion estimation component (221) for compression.
[0059] The general coder control component (211) is configured to make decisions regarding the coding of images of a video sequence into a bitstream according to application constraints. For example, the general coder control component (211) manages the optimization of bitrate / bitstream size versus reconstruction quality. These decisions may be made based on storage space / bandwidth availability and image resolution requests. The general coder control component (211) also manages buffer utilization in terms of transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general coder control component (211) manages splitting, prediction, and filtering by other components. For example, the general coder control component (211) can dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the general coder control component (211) controls other components of the codec system (200) to balance video signal reconstruction quality and bitrate issues. A general coder control component (211) generates control data that controls the operation of other components. The control data is also passed to a header formatting and CABAC component (231) to be encoded into a bitstream to signal parameters for decoding in the decoder.
[0060] The segmented video signal (201) is also transmitted to a motion estimation component (221) and a motion compensation component (219) for inter-prediction. A frame or slice of the segmented video signal (201) may be divided into multiple video blocks. The motion estimation component (221) and the motion compensation component (219) perform inter-prediction coding of the received video blocks for one or more blocks in one or more reference frames to provide temporal prediction. The codec system (200) may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0061] The motion estimation component (221) and the motion compensation component (219) may be highly integrated, but are exemplified separately for conceptual purposes. Motion estimation performed by the motion estimation component (221) is a process of generating motion vectors that estimate motion for a video block. For example, the motion vectors may indicate the displacement of a coded object relative to a prediction block. A prediction block is a block identified as closely matching the block to be coded in terms of pixel difference. A prediction block may also be referred to as a reference block. These pixel differences may be determined by the sum of absolute differences (SAD), the sum of square differences (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be subdivided into CTBs and then subdivided into CBs to be included in a CU. The CU may be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component (221) generates motion vectors, PUs, and TUs by using rate distortion analysis as part of a rate distortion optimization process. For example, the motion estimation component (221) may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and may select the reference blocks, motion vectors, etc. that have the best rate distortion characteristics. The best rate distortion characteristics balance both video reconstruction quality (e.g., amount of data loss due to compression) and coding efficiency (e.g., final encoding size).
[0062] In some examples, the codec system (200) can calculate values for sub-integer pixel locations of a reference picture stored in the decoded picture buffer component (223). For example, the video codec system (200) can interpolate values for 1 / 4 pixel locations, 1 / 8 pixel locations, or other fractional pixel locations of the reference pixel. Thus, the motion estimation component (221) can perform motion search for whole pixel locations and fractional pixel locations and output a motion vector with fractional pixel precision. The motion estimation component (221) calculates a motion vector for the PU of the video block in the intercoded slice by comparing the position of the PU with the position of the prediction block of the reference picture. The motion estimation component (221) outputs the calculated motion vector as motion data to the header formatting and CABAC component (231) for encoding and motion for the motion compensation component (219).
[0063] Motion compensation performed by the motion compensation component (219) may include fetching or generating a prediction block based on a motion vector determined by the motion estimation component (221). Again, the motion estimation component (221) and the motion compensation component (219) may be functionally integrated in some examples. Upon receiving a motion vector for the PU of the current video block, the motion compensation component (219) may find the prediction block indicated by the motion vector. Subsequently, the residual video block is formed by subtracting the pixel value of the prediction block from the pixel value of the current video block being coded to form a pixel difference value. Generally, the motion estimation component (221) performs motion estimation for the luminance component, and the motion compensation component (219) uses a motion vector calculated based on the luminance component for both the chroma component and the luminance component. The prediction block and the residual block are passed to the transform scaling and quantization component (213).
[0064] The segmented video signal (201) is also transmitted to the in-picture estimation component (215) and the in-picture prediction component (217). Like the motion estimation component (221) and the motion compensation component (219), the in-picture estimation component (215) and the in-picture prediction component (217) can be highly integrated, but are described separately for conceptual purposes. As described above, the in-picture estimation component (215) and the in-picture prediction component (217) inter-predict the current block for the current frame as an alternative to the inter-prediction performed by the motion estimation component (221) and the motion compensation component (219) between frames. In particular, the in-picture estimation component (215) determines the intra-prediction mode to use for encoding the current block. In some examples, the in-picture estimation component (215) selects an appropriate intra-prediction mode for encoding the current block from a number of tested intra-prediction modes. After that, the selected intra prediction mode is passed to the header formatting and CABAC component (231) for encoding.
[0065] For example, the in-picture estimation component (215) calculates a rate distortion value using rate distortion analysis for various tested intra-prediction modes and selects the intra-prediction mode having the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to generate the encoded block, as well as the bitrate (e.g., number of bits) used to generate the encoded block. The in-picture estimation component (215) calculates a ratio from the distortion and rate for various encoded blocks to determine whether the intra-prediction mode represents the best rate distortion value for the block. Additionally, the in-picture estimation component (215) may be configured to code the depth blocks of the depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0066] The in-picture prediction component (217), when implemented in the encoder, can generate a residual block from the prediction block based on the selected intra prediction mode determined by the in-picture estimation component (215), or when implemented in the decoder, can read the residual block from the bitstream. The residual block contains the difference in value between the prediction block and the original block, represented as a matrix. The residual block is then passed to the transform scaling and quantization component (213). The in-picture estimation component (215) and the in-picture prediction component (217) can operate in both the lumina and chroma components.
[0067] The transform scaling and quantization component (213) is configured to further compress the residual block. The transform scaling and quantization component (213) applies a transform, such as the discrete cosine transform (DT), discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform factor values. Wavelet transform, integer transform, subband transform, or other types of transform may also be used. The transform can convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component (213) is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different grain sizes, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component (213) is also configured to quantize the transform factors to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the transform scaling and quantization component (213) can perform a scan of a matrix containing quantized transform coefficients. The quantized transform coefficients are passed to the header formatting and CABAC component (231) and encoded into a bitstream.
[0068] The scaling and inverse transform component (229) applies the inverse operation of the transform scaling and quantization component (213) to support motion estimation. The scaling and inverse transform component (229) applies inverse scaling, transform, and / or quantization to reconstruct a residual block in the pixel domain for later use as a reference block that may be a prediction block for another current block, for example. The motion estimation component (221) and / or motion compensation component (219) may compute the reference block by adding the residual block back to the corresponding prediction block for use in motion estimation of a subsequent block / frame. A filter is applied to the reconstructed reference block to mitigate artifacts generated during scaling, quantization, and transform. Otherwise, these artifacts may cause inaccurate predictions (and generate additional artifacts) when a subsequent block is predicted.
[0069] The filter control analysis component (227) and the in-loop filter component (225) apply filters to residual blocks and / or reconstructed image blocks. For example, residual blocks transformed from the scaling and inverse transformation component (229) may be combined with corresponding prediction blocks from the in-picture prediction component (217) and / or motion compensation component (219) to reconstruct the original image block. Then, the filters may be applied to the reconstructed image block. In some examples, the filters may be applied to residual blocks instead. Like other components in FIG. 2, the filter control analysis component (227) and the in-loop filter component (225) may be highly integrated and implemented together, but are shown separately for conceptual purposes. Filters applied to the reconstructed reference block are applied to specific spatial regions and include several parameters that adjust how these filters are applied. The filter control analysis component (227) analyzes the reconstructed reference block to determine where such filters should be applied and sets the corresponding parameters. This data is passed to the header formatting and CABAC component (231) as filter control data for encoding. The in-loop filter component (225) applies these filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., reconstructed pixel blocks) or the frequency domain, depending on the example.
[0070] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component (223) for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component (223) stores and transmits the reconstructed and filtered blocks toward the display as part of the output video signal. The decoded picture buffer component (223) may be any memory device capable of storing the prediction blocks, residual blocks, and / or reconstructed image blocks.
[0071] The header formatting and CABAC component (231) receives data from various components of the codec system (200) and encodes this data into a coded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component (231) generates various headers to encode control data, such as general control data and filter control data. Additionally, prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data, are all encoded into a bitstream. The final bitstream contains all information required by the decoder to reconstruct the original segmented video signal (201). This information may also include an intra prediction mode index table (also referred to as a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra prediction mode, indications of segmentation information, etc. This data may be encoded using entropy coding. For example, information may be encoded using Context Adaptive Variable Length Coding (CAVLC), CABAC, Syntax-Based Context-Adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioning Entropy (PIPE) Coding, or other entropy coding techniques. After entropy coding, the encoded bitstream may be transmitted to another device (e.g., a video decoder) or stored for future transmission or retrieval.
[0072] FIG. 3 is a block diagram illustrating an exemplary video encoder (300). The video encoder (300) may be used to implement the encoding function of the codec system (200) and / or to implement steps 101, 103, 105, 107 and / or 109 of the method of operation (100). The encoder (300) divides the input video signal to generate a divided video signal (301), which is substantially similar to the divided video signal (201). The divided video signal (301) is then compressed by a component of the encoder (300) and encoded into a bitstream.
[0073] Specifically, the segmented video signal (301) is passed to the in-picture prediction component (317) for intra-prediction. The in-picture prediction component (317) may be substantially similar to the in-picture estimation component (215) and the in-picture prediction component (217). The segmented video signal (301) is also passed to the motion compensation component (321) for inter-prediction based on the reference block of the decoded picture buffer component (323). The motion compensation component (321) may be substantially similar to the motion estimation component (221) and the motion compensation component (219). The prediction block and residual block from the in-picture prediction component (317) and the motion compensation component (321) are passed to the transformation and quantization component (313) for the transformation and quantization of the residual block. The transformation and quantization component (313) may be substantially similar to the transformation scaling and quantization component (213). The converted and quantized residual blocks and the corresponding prediction blocks (along with associated control data) are passed to the entropy coding component (331) for coding into a bitstream. The entropy coding component (331) may be substantially similar to the header formatting and CABAC component (231).
[0074] The converted and quantized residual block and / or corresponding prediction block is also passed to the conversion and quantization component (313) for reconstruction into a reference block for use by the motion compensation component (321). The inverse conversion and quantization component (329) may be substantially similar to the scaling and inverse conversion component (229). The in-loop filter of the in-loop filter component (325) is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component (325) may be substantially similar to the filter control analysis component (227) and the in-loop filter component (225). The in-loop filter component (325) may include a number of filters as discussed for the in-loop filter component (225). The filtered block is then stored in the decoded picture buffer component (323) for use as a reference block by the motion compensation component (321). The decoded picture buffer component (323) may be substantially similar to the decoded picture buffer component (223).
[0075] FIG. 4 is a block diagram illustrating an exemplary video decoder (400). The video decoder (400) may be used to implement the decoding function of the codec system (200) and / or to implement steps 111, 113, 115 and / or 117 of the method of operation (100). The decoder (400) receives a bitstream, for example, from the encoder (300) and generates an output video signal reconstructed based on the bitstream for display to an end user.
[0076] The bitstream is received by the entropy decoding component (433). The entropy decoding component (433) is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component (433) may use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are passed to the inverse transform and quantization component (429) for reconstruction into the residual block. The inverse transform and quantization component (429) may be similar to the inverse transform and quantization component (329).
[0077] The reconstructed residual block and / or prediction block is passed to the in-picture prediction component (417) to be reconstructed into an image block based on an intra-prediction operation. The in-picture prediction component (417) may be similar to the in-picture estimation component (215) and the in-picture prediction component (217). Specifically, the in-picture prediction component (417) uses a prediction mode to find a reference block in the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and the corresponding inter-prediction data are passed to the decoded picture buffer component (423) via the in-loop filter component (425), which may be substantially similar to the decoded picture buffer component (223) and the in-loop filter component (225), respectively. The in-loop filter component (425) filters the reconstructed image blocks, residual blocks, and / or prediction blocks, and this information is stored in the decoded picture buffer component (423). The reconstructed image blocks from the decoded picture buffer component (423) are passed to the motion compensation component (421) for inter-prediction. The motion compensation component (421) may be substantially similar to the motion estimation component (221) and / or the motion compensation component (219). Specifically, the motion compensation component (421) uses motion vectors from the reference blocks to generate prediction blocks and to apply residual blocks to the result to reconstruct the image blocks. The resulting reconstructed blocks may also be passed to the decoded picture buffer component (423) via the in-loop filter component (425). The decoded picture buffer component (423) continues to store additional reconstructed image blocks that can be reconstructed into frames through the segmentation information. These frames may also be arranged in a sequence. The sequence is output to the display as a reconstructed output video signal.
[0078] FIG. 5 is a schematic diagram illustrating a plurality of sub-picture video streams (501, 502, 503) divided from a VR picture video stream (500). For example, the sub-picture video streams (501-503) and / or the VR picture video stream (500) may be encoded by an encoder such as a codec system (200) and / or an encoder (300) according to method (100). Additionally, the sub-picture video streams (501-503) and / or the VR picture video stream (500) may be decoded by a decoder such as a codec system (200) and / or a decoder (400).
[0079] A VR picture video stream (500) includes multiple pictures that are displayed over time. Specifically, VR operates by coding a sphere of video content that can be displayed as if the user were in the center of the sphere. Each picture contains the entire sphere. Meanwhile, only a portion of the picture, known as a viewport, is displayed to the user. For example, the user may use a Head Mounted Display (HMD) that selects and displays a viewport of the sphere based on the user's head movements. This provides the sensation of being physically present in the virtual space as depicted by the video. To achieve this result, each picture of the video sequence contains the entire sphere of video data at the corresponding point in time. However, only a small portion of the picture (e.g., a single viewport) is displayed to the user. The rest of the picture is not rendered and is discarded. Generally, the entire picture is transmitted so that different viewports can be dynamically selected and displayed in response to the user's head movements.
[0080] In the illustrated example, the pictures of the VR picture video stream (500) can be subdivided into sub-pictures based on the available viewports. Thus, each picture and its corresponding sub-picture includes a picture order count as part of the representation. Sub-picture video streams (501-503) are generated when the sub-division is applied consistently over time. This consistent sub-division generates sub-picture video streams (501-503) in which each stream contains a set of sub-pictures of predetermined size, shape, and spatial location for the corresponding picture of the VR video video stream (500). Additionally, the sub-pictures of the sub-picture video streams (501-503) change their picture order counts during the representation time. In this way, the sub-pictures of the sub-picture video streams (501-503) can be arranged based on the picture order over the representation time. Next, sub-pictures from the sub-picture video streams (501-503) at each picture order count value can be merged in a spatial domain based on predefined spatial positions to reconstruct the VR picture video stream (500) for display. Specifically, the sub-picture video streams (501-503) can each be encoded into separate sub-bitstreams. When these sub-bitstreams are merged together, they generate a bitstream containing the entire set of pictures for the display time. The resulting bitstream can be sent to a decoder for decoding and display based on the viewport currently selected by the user.
[0081] One of the problems with VR video is that all sub-picture video streams (501-503) can be transmitted to the user in high quality (e.g., high resolution). This allows the decoder to dynamically select the user's current viewport and display the sub-picture in real time from the corresponding sub-picture video streams (501-503). However, while the user can see only a single viewport from, for example, the sub-picture video stream (501), the sub-picture video streams (502-503) are discarded. As such, transmitting the sub-picture video streams (502-503) in high quality can waste a significant amount of bandwidth. To improve coding efficiency, VR video can be encoded into multiple video streams (500), each of which is encoded with different quality / resolution. In this way, the decoder can transmit a request for the current sub-picture video stream (501). In response, the encoder (or intermediate slicer or other content server) may select a higher quality sub-picture video stream (501) from a higher quality video stream (500) and select a lower quality sub-picture video stream (502-503) from a lower quality video stream (500). Then, the encoder may merge these sub-bitstreams into a complete encoded bitstream for transmission to the decoder. In this way, the decoder receives a series of pictures with higher quality in the current viewport and lower quality in other viewports. Additionally, the highest quality sub-picture is generally displayed to the user (no head movement), and the lower quality sub-picture is generally discarded, allowing a balance between functionality and coding efficiency.
[0082] When the user switches the view from the sub-picture video stream (501) to the sub-picture video stream (502), the decoder requests that the new current sub-picture video stream (502) be transmitted at a higher quality. Then, the encoder may change the merging mechanism accordingly. The decoder can start new CVS decoding only on the IRAP picture. This is because the IRAP picture is coded according to intra-prediction, which does not refer to other pictures. Therefore, the IRAP picture can be decoded even if the picture preceding the IRAP picture is unavailable. The non-IRAP picture is coded according to inter-prediction. As such, the non-IRAP picture cannot be decoded without first decoding the corresponding set of reference pictures based on the reference picture list. Therefore, the decoder generally cannot start decoding the video sequence on the non-IRAP picture. Due to these constraints, the sub-picture video stream (502) is displayed at a lower quality until the IRAP picture / sub-picture is reached. Then, the IRAP picture is decoded at a higher quality to initiate decoding of the higher quality version of the sub-picture video stream (502). This approach significantly increases video compression without negatively impacting the user's viewing experience.
[0083] One concern regarding the aforementioned approach is that the time required to change the resolution is based on the time until an IRAP picture is reached in the video stream. This is because the decoder cannot begin decoding different versions of the sub-picture video stream (502) from a non-IRAP picture. One approach to reducing this latency is to include more IRAP pictures. However, this increases the file size. To balance functionality and coding efficiency, different viewport / sub-picture video streams (501-503) may include IRAP pictures at different frequencies. For example, a viewport / sub-picture video stream (501-503) that is more likely to be viewed may have more IRAP pictures than a different viewport / sub-picture video stream (501-503). For example, in a basketball situation, a viewport / sub-picture video stream (501-503) related to the basket and / or center court may contain IRAP pictures more frequently than a viewport / sub-picture video stream (501-503) viewing the stand or ceiling because the viewport / sub-picture video stream (501-503) viewing the stand or ceiling is less likely to be viewed by the user.
[0084] This approach introduces additional problems. Specifically, sub-pictures from sub-picture video streams (501-503) sharing a POC are part of a single picture. As mentioned above, slices from a picture are included in NAL units based on the picture type. In some video coding systems, all NAL units associated with a single picture are restricted to containing the same NAL unit type. If different sub-picture video streams (501-503) have IRAP pictures at different frequencies, some of the pictures contain both IRAP sub-pictures and non-IRAP sub-pictures. This violates the constraint that each single picture must use only the same type of NAL unit.
[0085] This issue can be resolved by removing the constraint that all NAL units for a picture slice use the same NAL unit type. For example, a picture is contained within an access unit. By removing this constraint, the access unit can contain both IRAP NAL unit types and non-IRAP NAL unit types. Additionally, a flag may be encoded to indicate cases where the picture / access unit contains a mix of IRAP and non-IRAP NAL unit types. In some examples, the flag is the mixed NAL unit types in the picture flag (mixed_nalu_types_in_pic_flag). Furthermore, a constraint may be applied requiring that a single mixed picture / access unit can contain only one type of IRAP NAL unit and one type of non-IRAP NAL unit. This prevents unintended mixing of NAL unit types. If such mixing were permitted, the decoder would need to be designed to manage it. This would unnecessarily increase the required hardware complexity without providing additional benefits to the coding process. For example, the mixed picture may include one type of IRAP NAL unit selected from IDR_W_RADL, IDR_N_LP, or CRA_NUT. Additionally, the mixed picture may include one type of non-IRAP NAL unit selected from TRAIL_NUT, RADL_NUT, and RASL_NUT.
[0086] FIG. 6 is a schematic diagram illustrating a constraint (600) used when including a mixed NAL unit type of current picture. The constraint (600) can be applied when coding a VR picture video stream (500). As such, the constraint (600) can be used by an encoder, such as a codec system (200) and / or an encoder (300), that applies the method (100). Additionally, the constraint (600) can be used by a decoder, such as a codec system (200) and / or a decoder (400), that applies the method (100). The constraint (600) is a requirement imposed on the video data, imposed on supporting parameters, and / or imposed on the process related to video data coding and / or decoding.
[0087] Specifically, FIG. 6 illustrates a series of pictures divided into sub-pictures. The sub-pictures are denoted as sub-picture A (subpicA) (601) and sub-picture B (subpicB) (603). The sub-pictures, namely subpicA (601) and subpicB (603), are rectangular and / or square areas of one or more slices within the picture sequence. SubpicA (601) and subpicB (603) may be included in a sub-picture video stream such as any of the sub-picture video streams (501-503) of FIG. 5. For clarity of description, only two sub-pictures are illustrated, but any number of pictures may be used.
[0088] A picture is an array of luminance samples and / or chroma samples that generate a frame or its field. A picture is encoded and / or decoded in a decoding order (650). The decoding order is the order in which syntax elements are processed by the decoding process. As the decoding order progresses, the decoding process proceeds through the picture. For clarity of description, a picture currently being encoded / decoded at a specific point in time is referred to as the current picture (620). A picture that has already been encoded / decoded is the preceding picture (610). A picture that has not yet been decoded is the subsequent picture (630 and / or 640). As illustrated in FIG. 6, the picture is consistently divided into subpicA (601) and subpicB (603). Accordingly, the preceding picture (610), the current picture (620), and the subsequent pictures (630, 640) are each divided into subpicA (601) and subpicB (603) and / or include.
[0089] As described above, IRAP sub-pictures may be applied at different frequencies in some examples. In the illustrated example, the current picture (620) includes multiple VCL NAL units that do not have the same NAL unit type. Specifically, subpicA (601) of the current picture (620) includes an IRAP NAL unit (621), whereas subpicB (603) of the current picture (620) includes a non-IRAP NAL unit (623). The IRAP NAL unit (621) is a data structure containing data from an IRAP picture or sub-picture. The IRAP picture / sub-picture is a picture / sub-picture coded according to intra-prediction, which indicates that the decoder can start video sequence decoding from the corresponding picture / sub-picture without referring to the picture preceding the IRAP picture / sub-picture. The IRAP NAL unit (621) may include a CRA NAL unit and / or an IDR NAL unit. A CRA picture / subpicture is an IRAP picture / subpicture that does not refresh the DPB, and an IDR picture / subpicture is an IRAP picture / subpicture that refreshes the DPB. A non-IRAP NAL unit (623) is any VCL NAL unit that does not contain an IRAP picture / subpicture. For example, a non-IRAP NAL unit (623) may contain a leading picture, such as a Random Access Skipped Leading (RASL) picture or a Random Access Decodable Leading (RADL) picture, or a trailing picture. The leading picture occurs before the IRAP picture in the presentation order and after the IRAP picture in the decoding order (650). The trailing picture occurs after the IRAP picture in both the presentation order and the decoding order (650). The non-IRAP NAL unit (623) can be coded according to the inter prediction.In this way, the current picture (620) is coded according to both the inter prediction in subpicB (603) and the intra prediction in subpicA (601).
[0090] The preceding picture (610) includes a non-IRAP NAL unit (611) and a non-IRAP NAL unit (613), the subsequent picture (630) includes a non-IRAP NAL unit (631) and a non-IRAP NAL unit (633), and the subsequent picture (640) includes a non-IRAP NAL (641) and a non-IRAP NAL unit (643). The non-IRAP NAL units (611, 613, 631, 633, 641, 643) may be similar to the non-IRAP NAL unit (623) (e.g., may be coded according to inter-prediction), but may contain different video data.
[0091] Because the current picture (620) contains VCL NAL units that do not have the same NAL unit type, the current picture (620) is a mixed NAL unit picture. The presence of a mixed NAL unit picture can be signaled by a flag in the parameter set of the bitstream. The constraint (600) is applied when a mixed NAL unit picture, such as the current picture (620), occurs.
[0092] During encoding, the encoder may determine that the subpicA (601) of the current picture (620) is associated with an IRAP NAL unit type and thus contains an IRAP NAL unit (621). This ensures that the subpicA (601) of the current picture (620) acts as a boundary and prevents inter-predictions in the subpicA (601) from referencing through the current picture (620). Therefore, a non-IRAP NAL unit (631) cannot refer to any other NAL unit preceding the non-IRAP NAL unit (611) or the IRAP NAL unit (621). The encoder may encode the current picture (620) by ensuring that all slices of the current picture located in the subpicA (601) are associated with the same NAL unit type. For example, if subpicA (601) contains at least one CRA slice (or IDR slice), all slices of subpicA (601) must also be CRA (or IDR). Then, slices of the current picture (620) (containing IRAP NAL unit (621) and non-IRAP NAL unit (623)) are encoded based on the NAL unit type (e.g., intra-predict and inter-predict, respectively). Subsequent pictures (630, 640) following the current picture (620) in the decoding order (650) are also encoded. To ensure that the IRAP NAL unit (621) prevents inter-predict propagation, the non-IRAP NAL unit (631, 641) is prevented from referencing (632) the preceding picture (610). Reference (632) is controlled by a reference picture list, where the active entry of the reference picture list indicates a reference picture available for the picture currently being coded.Thus, the constraint (600) ensures that the active entry associated with the slice located in the subpicA (601) in the subsequent pictures (630 and 640) (e.g., non-IRAP NAL units (631 and 641)) does not reference / reference any reference picture (632) preceding the current picture in the decoding order (650) when the subpicA (601) of the current picture (620) is associated with an IRAP NAL unit (621) having an IRAP NAL unit type. The non-IRAP NAL unit (631) is still allowed to reference the current picture (620) and / or the subsequent picture (640). Additionally, the non-IRAP NAL unit (641) is still allowed to reference the current picture (620) and / or the subsequent picture (630). This stops inter-predictive propagation for non-IRAP NAL units (631, 641) because the non-IRAP NAL units (631, 641) follow the IRAP NAL unit (621) in subpicA (601). Then, subsequent pictures (630 and / or 640) can be encoded based on the NAL unit type and according to the constraint (600).
[0093] These constraints (600) do not apply because the non-IRAP NAL units (633, 643) are located in subpicB (603) and therefore do not follow the IRAP NAL units in subpicB (603). Thus, the non-IRAP NAL units (633, 643) can refer to the preceding picture (610). Thus, the active entry of the reference picture list associated with the non-IRAP NAL units (633, 643) can refer to the preceding picture (610), the current picture (620), and / or the subsequent picture (640 or 630), respectively. In this way, inter-prediction for subpicA (601) is interrupted by the IRAP NAL unit (621), but the IRAP NAL unit (621) does not interrupt the propagation of inter-prediction for subpicB (603). Accordingly, by using the constraint (600), the inter-prediction chain may be stopped for the sub-picture based on the sub-picture (or not). As mentioned above, only two sub-pictures are shown. However, the constraint (600) may be applied to any number of sub-pictures to stop the propagation of the inter-prediction reference chain for the sub-picture based on the sub-picture.
[0094] FIG. 7 is a schematic diagram illustrating an exemplary reference picture list (RPL) structure (700) containing a reference picture list. The RPL structure (700) can be used to store instructions for reference pictures used for unidirectional inter-predicting and / or bidirectional inter-predicting. Thus, the RPL structure (700) can be used by a codec system (200), an encoder (300), and / or a decoder (400) when performing method (100). Additionally, the RPL structure (700) can be used when coding a VR picture video stream (500), wherein the RPL structure (700) can be coded according to constraints (600).
[0095] The RPL structure (700) is an addressable syntax structure containing multiple reference picture lists, such as RPL 0 (711) and RPL 1 (712). The RPL structure (700) may be coded and / or derived for use when coding a corresponding slice. The RPL structure (700) may be stored in the slice header of an SPS and / or bitstream, depending on the example. Reference picture lists, such as RPL 0 (711) and RPL 1 (712), are lists indicating reference pictures used for inter-prediction. RPL 0 (711) and RPL 1 (712) may each contain multiple entries (715). An RPL structure entry (715) is an addressable location within the RPL structure (700) indicating a reference picture associated with a reference picture list, such as RPL 0 (711) and / or RPL 1 (712). Each entry (715) may contain a Picture Order Count (POC) value (or other pointer value) that references the picture used for inter prediction. Specifically, a reference to the picture used by unidirectional inter prediction is stored in RPL 0 (711), and a reference to the picture used by bidirectional inter prediction is stored in both RPL 0 (711) and RPL 1 (712). For example, bidirectional inter prediction may use one reference picture indicated by RPL 0 (711) and one reference picture indicated by RPL 1 (712).
[0096] In a specific example, the RPL structure (700) may be represented as ref_pic_list_struct(listIdx, rplsIdx), where the list index (listIdx) (721) identifies the reference picture list RPL 0 (711) and / or RPL 1 (712) and the reference picture list structure index (rplsIdx) (725) identifies the entry (715) of the reference picture list. Thus, ref_pic_list_struct is a syntax structure that returns the entry (715) based on listIdx (721) and rplsIdx (725). The encoder may encode a portion of the RPL structure (700) for each un-intracoded slice of the video sequence. Then, the decoder may resolve the corresponding portion of the RPL structure (700) before decoding each un-intracoded slice of the coded video sequence. When coding the current picture according to the inter prediction, the entry (715) indicating the available reference picture is referred to as the active entry. The entry (715) that cannot be used for the current picture is referred to as the inactive entry.
[0097] FIG. 8 is a schematic diagram illustrating an exemplary bitstream (800) comprising a picture having a mixed NAL unit type. For example, the bitstream (800) may be generated by a codec system (200) and / or an encoder (300) for decoding by a codec system (200) and / or a decoder (400) according to method (100). Additionally, the bitstream (800) may include a VR picture video stream (500) merged from a plurality of sub-picture video streams (501-503) at a plurality of video resolutions. Additionally, the bitstream (800) may include an RPL structure (700) coded according to a constraint (600).
[0098] The bitstream (800) includes a sequence parameter set (SPS) (810), a plurality of picture parameter sets (PPS) (811), a plurality of slice headers (815), and image data (820). The SPS (810) includes sequence data common to all pictures of the video sequence included in the bitstream (800). This data may include picture scaling, bit depth, coding tool parameters, bit rate limits, etc. The PPS (811) includes parameters applied to the entire picture. Thus, each picture of the video sequence may refer to the PPS (811). Although each picture refers to the PPS (811), a single PPS (811) may include data for multiple pictures in some examples. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS (811) may include data for such similar pictures. PPS (811) may indicate coding tools available for the slices, quantization parameters, offsets, etc. of the corresponding picture. The slice header (815) contains parameters specific to each slice of the picture. Thus, there may be one slice header (815) per slice of the video sequence. The slice header (815) may include slice type information, picture order count (POC), reference picture list, prediction weight, tile entry point, deblocking parameters, etc. The slice header (815) may also be referred to as a tile group header in some contexts.
[0099] Image data (820) includes video data encoded according to inter-prediction and / or intra-prediction, and corresponding transformed and quantized residual data. For example, a video sequence includes multiple pictures (821) coded with image data (820). Since a picture (821) is a single frame of a video sequence, it is generally displayed as a single unit when displaying a video sequence. However, to implement certain technologies such as virtual reality, sub-pictures (823) may be displayed. Each picture (821) refers to a PPS (811). A picture (821) may be divided into sub-pictures (823), tiles, and / or slices. A sub-picture (823) is a spatial region of a picture (821) that is consistently applied across the coded video sequence. Thus, sub-pictures (823) may be displayed by an HMD in a VR context. Additionally, a sub-picture (823) having a specific POC can be obtained from a sub-picture video stream (501-503) at a corresponding resolution. The sub-picture (823) may refer to the SPS (810). In some systems, a slice (825) is referred to as a group of tiles containing tiles. A slice (825) and / or a group of tiles refers to a slice header (815). A slice (825) may be defined as an integer of complete tiles within a tile of a picture (821) contained exclusively in a single NAL unit, or as an integer of a continuous row of complete Coding Tree Units (CTUs). Thus, the slice (825) is further subdivided into CTUs and / or Coding Tree Blocks (CTBs). The CTUs / CTBs are further subdivided into Coding Blocks based on the Coding Tree. Then, the Coding Blocks can be encoded / decoded according to a prediction mechanism.
[0100] A parameter set and / or slice (825) is coded as a NAL unit. A NAL unit may be defined as a syntax structure containing a byte and a subsequent instruction of a data type containing the corresponding data in the form of an RBSP with anti-emulation bytes interspersed as needed. More specifically, a NAL unit is a storage unit containing a parameter set or slice (825) of a picture (821) and a corresponding slice header (815). Specifically, a VCL NAL unit (840) is a NAL unit containing a slice (825) of a picture (821) and a corresponding slice header (815). Additionally, a non-VCL NAL unit (830) contains a parameter set such as an SPS (810) and a PPS (811). Several types of NAL units may be used. For example, SPS (810) and PPS (811) can each be contained in SPS NAL unit type (SPS_NUT) (831) and PPS NAL unit type (PPS_NUT) (832), which are both non-VCL NAL unit (830). In this way, the decoder can read SPS_NUT (831) from the bitstream (800) to obtain the SPS (810) coded by the encoder. Likewise, the decoder can read PPS_NUT (832) from the bitstream (800) to obtain the PPS (811) coded by the encoder.
[0101] A slice (825) of an IRAP picture / subpicture may be included in an IRAP NAL unit (845). A slice (825) of a non-IRAP picture / subpicture, such as a leading picture and a trailing picture, may be included in a non-IRAP NAL unit (849). For example, a slice (825) may be included in a single VCL NAL unit (840). Then, a type identifier may be assigned to the VCL NAL unit (840) based on the type of the picture (821) and / or subpicture (823) containing the slice (825). For example, a slice (825) taken from a subpicture (823) that is a CRA subpicture is included in CRA_NUT (843). The bitstream (800) includes several types of IRAP NAL units (845) and thus includes picture / subpicture types including an IDR without a leading picture, an IDR with a Random Access Decoding Capable Leading (RADL) picture, and a CRA picture. The bitstream (800) also includes several types of non-IRAP NAL units (849) and thus includes picture / subpicture types including a Random Access Skipped Leading (RASL) picture, a RADL picture, and a trailing picture.
[0102] A leading picture is a picture that is coded after an IRAP picture in the decoding order and before a picture in the representation order. An IRAP NAL unit (845) is any NAL unit containing a slice (825) taken from an IRAP picture or a sub-picture. A non-IRAP NAL unit (849) is any NAL unit containing a slice (825) taken from any picture that is not an IRAP picture or a sub-picture (e.g., a leading picture or a trailing picture). Both the IRAP NAL unit (845) and the non-IRAP NAL unit (849) are VCL NAL units (840) because they contain slice data. In an exemplary embodiment, an IRAP NAL unit (845) may contain a slice (825) from an IDR associated with an IDR picture or a RADL picture that does not have a leading picture in the IDR_N_LP NAL unit (841) or the IDR_w_RADL NAL unit (842), respectively. Additionally, the IRAP NAL unit (845) may include a slice (825) from a CRA picture of the CRA_NUT (843). In an exemplary embodiment, the non-IRAP NAL unit (849) may include a slice (825) from a RASL picture, a RADL picture, or a trailing picture in the RASL_NUT (846), RADL_NUT (847), or TRAIL_NUT (848), respectively. In an exemplary embodiment, a complete list of possible NAL units is shown below, sorted by NAL unit type.
[0103] nal_unit_type Name of nal_unit_type Content of the NAL unit RBSP syntax structure NAL unit type class 0 TRAIL_NUT Coded slice of the trailing picture slice_layer_rbsp() VCL 1 STSA_NUT Coded slice of STSA picture slice_layer_rbsp() VCL 2 RADL_NUT Coded slice of a RADL picture slice_layer_rbsp() VCL 3 RASL_NUT Coded slice of a RASL picture slice_layer_rbsp() VCL 4..6 RSV_VCL_4..RSV_VCL_6 Reserved non-IRAP VCL NAL unit type VCL 78 IDR_W_RADLIDR_N_LP Coded slice of the IDR picture slice_layer_rbsp() VCL 9 CRA_NUT Coded slice of CRA picture silce_layer_rbsp() VCL 10 GDR_NUT Coded slice of a GDR picture slice_layer_rbsp() VCL 1112 RSV_IRAP_11RSV_IRAP_12 Reserved IRAP VCL NAL unit type VCL 13 DCI_NUT Decoding capability information decoding_capability_information_rbsp( ) non-VCL 14 VPS_NUT Video parameter set video_parameter_set_rbsp( ) non-VCL 15 SPS_NUT Sequence parameter set seq_parameter_set_rbsp( ) non-VCL 16 PPS_NUT Picture parameter set pic_parameter_set_rbsp( ) non-VCL 1718 PREFIX_APS_NUTSUFFIX_APS_NUT Adaptation parameter set adaptation_parameter_set_rbsp( ) non-VCL 19 PH_NUT Picture Header picture_header_rbsp( ) non-VCL 20 AUD_NUT AU delimiter character access_unit_delimiter_rbsp( ) non-VCL 21 EOS_NUT End of sequence end_of_seq_rbsp() non-VCL 22 EOB_NUT End of bitstream end_of_bitstream_rbsp() non-VCL 2324 PREFIX_SEI_NUTSUFFIX_SEI_NUT Supplemental Improvement Information sei_rbsp( ) non-VCL 25 FD_NUT Filter data filler_data_rbsp( ) non-VCL 2627 RSV_NVCL_26RSV_NVCL_27 Reserved non-IRAP VCL NAL unit type non-VCL 28..31 UNSPEC_28..UNSPEC_31 Unspecified non-VCL NAL unit type non-VCL
[0104] As described above, a VR video stream may include sub-pictures (823) that have IRAP pictures at different frequencies. This allows fewer IRAP pictures to be used for spatial regions that the user is unlikely to view, and more IRAP pictures to be used for spatial regions that the user is likely to view frequently. In this way, spatial regions that the user is likely to switch back to regularly can be quickly adjusted to higher resolution. If this approach generates a picture (821) that includes both IRAP NAL units (845) and non-IRAP NAL units (849), the picture (821) is referred to as a mixed picture. This condition can be signaled by the mixed NAL unit type (mixed_nalu_types_in_pic_flag) (827) of the picture flag. The mixed_nalu_types_in_pic_flag (827) can be set in the PPS (811). Additionally, mixed_nalu_types_in_pic_flag (827) may be set to 1 if it specifies that each picture (821) referencing the PPS (811) has one or more VCL NAL units (840) and that the VCL NAL units (840) do not have the same value of NAL unit type (nal_unit_type). Additionally, mixed_nalu_types_in_pic_flag (827) may be set to 0 if each picture (821) referencing the PPS (811) has one or more VCL NAL units (840) and that the VCL NAL units (840) of each picture (821) referencing the PPS (811) all have the same value of nal_unit_type.Additionally, when mixed_nalu_types_in_pic_flag (827) is set, various constraints may be used such that one or more VCL NAL units (840) of the sub-pictures (823) of the picture (821) all have a first specific value of the NAL unit type and other VCL NAL units (840) of the picture (821) all have different second specific values of the NAL unit type. For example, the constraints may require that the mixed picture (821) contain a single type of IRAP NAL unit (845) and a single type of non-IRAP NAL unit (849). For example, the picture (821) may contain one or more IDR_N_LP NAL units (841), one or more IDR_w_RADL NAL units (842), or one or more CRA_NUT (843), but may not contain any combination of these IRAP NAL units (845). Additionally, the picture (821) may include one or more RASL_NUTs (846), one or more RADL_NUTs (847), or one or more TRAIL_8NUTs, but may not include any combination of these IRAP NAL units (845). Additionally, the sub-picture (823) may be limited to only one type of VCL NAL unit (840). Additionally, the constraint (600) may be applied to the sub-picture (823) based on the type of VCL NAL unit (840) used in the bitstream (800).
[0105] Prior information is now explained in more detail below. In video codec specifications, picture types may be identified to define the decoding process. This may include the derivation of picture identification (e.g., POC), indication of the reference picture status of (DPB), picture output from DPB, etc.
[0106] In AVC and HEVC, picture types can be identified from NAL unit types containing coded pictures. Picture types in AVC include IDR pictures and non-IDR pictures. Picture types in HEVC include trailing pictures, temporal sub-layer access pictures (TSA), step-wise temporal sub-layer access pictures (STSA), random access decodingable reading pictures (RADL), random access skipped reading pictures (RASL), broken-link access pictures (BLA), IDR, and CRA. In HEVC, each of these picture types can be further distinguished as a sub-layer referenced picture or a sub-layer non-referenced picture. BLA pictures may also include BLAs with reading pictures, BLAs with RADL pictures, and BLAs without reading pictures. IDR pictures may include IDRs with RADL pictures and IDRs without preceding pictures.
[0107] In HEVC, IDR, BLA, and CRA pictures are IRAP pictures. VVC uses IDR and CRA pictures as IRAP pictures. IRAP pictures provide the following functionality / advantages. The presence of an IRAP picture indicates that the decoding process can start at that picture. This functionality supports random access features that allow the decoding process to start at any location in the bitstream as long as the IRAP picture is at that location. The location may not be the beginning of the bitstream. The presence of an IRAP picture can also refresh the decoding process so that coded pictures following the IRAP picture, excluding RASL pictures, are coded without reference to the picture preceding the IRAP picture. Therefore, IRAP pictures prevent errors that occurred before the IRAP picture from propagating to the picture following the IRAP picture in the decoding order.
[0108] IRAP pictures provide the aforementioned functionality but result in a penalty to compression efficiency. The presence of IRAP pictures also causes a spike in bit rate. This penalty to compression efficiency has two causes. First, since IRAP pictures are intra-predicted pictures, they are represented by more bits than inter-predicted pictures. Second, the presence of IRAP pictures can break temporal prediction by refreshing the decoding process when a reference picture is removed from the DPB. This can lead to less efficient coding of pictures following IRAP pictures because there are fewer reference pictures available for inter-prediction.
[0109] HEVC IDR pictures can be derived and signaled differently from other picture types. Some differences are as follows: When signaling and deriving the POC value of an IDR picture, the Most Significant Bit (MSB) of the POC may be set to 0 instead of being derived from the previous key picture. Additionally, the slice header of an IDR picture may not contain information to assist in reference picture management. For other picture types, such as CRA and trailing, the Reference Picture Set (RPS) or Reference Picture List may be used in the reference picture marking process. This process is used to determine the status of reference pictures in a DPB that are used for reference or not. In the case of IDR pictures, this information may not be signaled because the presence of the IDR instructs the decoding process to mark all reference pictures in the DPB as not being used for reference.
[0110] In addition to picture types, picture identification is also used for various purposes. These include picture identification of reference pictures in inter-prediction, picture identification for output in DPB, picture identification for motion vector scaling, and picture identification for weighted prediction. In AVC and HEVC, pictures can be identified by POC. In AVC and HEVC, pictures in DPB can be marked as used for short-term reference, used for long-term reference, or not used for reference. If a picture is marked as not used for reference, it can no longer be used for prediction. If a picture is no longer needed for output, it can be removed from DPB. AVC uses short-term and long-term reference pictures. A reference picture can be marked as not used for reference if the picture is no longer needed for prediction reference. The conversion between short-term, long-term, and not used for reference is controlled by the decoded reference picture marking process. The implicit sliding window process and the explicit memory management control operation (MMCO) process can be used as mechanisms for marking decoded reference pictures. The sliding window process marks a short-reference picture as not being used for reference when the number of reference frames is equal to the maximum number specified in the SPS. Short-reference pictures are stored in a first-in, first-out manner so that the short-reference picture most recently decoded is retained in the DPB. The explicit MMCO process may include multiple MMCO commands. An MMCO command can mark one or more short or long-reference pictures as not being used for reference, or mark all pictures as not being used for reference. An MMCO command can also mark the current reference picture or an existing short-reference picture as long-reference and assign a long-reference picture index to a long-reference picture index.In AVC, the reference picture display operation and the picture output and removal process in DPB are performed after the picture is decoded.
[0111] HEVC uses RPS for reference picture management. For each slice, the RPS can contain a complete set of reference pictures used by the current picture or any subsequent picture. Therefore, the RPS signals the complete set of all pictures that must be maintained in the DPB for use by the current or subsequent picture. This differs from the AVC method, where only changes relative to the DPB are signaled. To maintain the accurate state of the reference picture in the DPB, the RPS may not retain information from the previous picture during decoding. In HEVC, the order of picture decoding and DPB operations leverages the advantages of the RPS and improves error resilience. In AVC, picture display and buffer operations may be applied after the current picture is decoded. In HEVC, the RPS is first decoded from the slice header of the current picture. Then, picture display and buffer operations are applied before decoding the current picture.
[0112] VVC can directly signal and derive Reference Picture List 0 and Reference Picture List 1. Reference Picture Lists are not based on RPS, sliding window, or MMCO processes as in HEVC and AVC. Reference picture indication is performed directly based on Reference Picture Lists 0 and 1 by utilizing both active and inactive entries of the Reference Picture Lists. Only active entries can be used as reference indices in the CTU's inter-prediction for the current picture. Information for deriving the two Reference Picture Lists is signaled by the syntax elements and syntax structures of the SPS, PPS, and slice headers. Predefined RPL structures are signaled in the SPS to be used as references in the slice header. Two Reference Picture Lists are generated for all types of slices, including bidirectional inter-prediction (B), unidirectional inter-prediction (P), and intra-prediction (I) slices. The two Reference Picture Lists are configured without using a Reference Picture List Initialization process or a Reference Picture List Modification process. The Long-term Reference Picture (LTRP) is identified by the POC LSB. Delta POC MSB cycles can be signaled for the LTRP as desired on a picture-by-picture basis.
[0113] HEVC can utilize normal slices, dependent slices, tiles, and Wavefront Parallel Processing (WPP) as partitioning methods. These partitioning methods can be applied to Maximum Transfer Unit (MTU) size matching, parallel processing, and end-to-end latency reduction. Each normal slice can be encapsulated into a separate NAL unit. Entropy coding dependencies and in-picture predictions, including in-sample prediction, motion information prediction, and coding mode prediction, can be disabled across slice boundaries. Therefore, a normal slice can be reconstructed independently of other normal slices within the same picture. However, slices may still have some interdependencies due to loop filtering operations.
[0114] Regular slice-based parallelization may not require significant inter-processor or inter-core communication. One exception is when inter-processor and / or inter-core data sharing can be critical for motion compensation when decoding predictive-coded pictures. This process may involve more processing resources than inter-processor or inter-core data sharing due to in-picture prediction. However, for the same reason, using regular slices can result in significant coding overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. Additionally, regular slices are also used as a mechanism for bitstream splitting to meet MTU size requirements due to the in-picture independence of regular slices and the fact that each regular slice is encapsulated in a separate NAL unit. In many cases, the goals of parallelization and MTU size matching conflict with the requirements for the picture slice layout.
[0115] Dependent slices have short slice headers and allow the splitting of bitstreams at tree-block boundaries without breaking in-picture prediction. Dependent slices fragment a regular slice into multiple NAL units. This reduces end-to-end latency by allowing parts of the regular slice to be transmitted before the encoding of the entire regular slice is complete.
[0116] In WPP, the picture is partitioned into a single row of CTBs. Entropy decoding and prediction can utilize data from CTBs in different partitions. Parallel processing is possible through the parallel decoding of CTB rows. The start of decoding for a CTB row is delayed by one or two CTBs, depending on the example, to ensure that data related to the CTBs above and to the right of the subject CTB is available before the subject CTB is decoded. This staggered start creates the appearance of wavefronts. This process supports parallelization up to the number of processors / cores corresponding to the number of CTB rows contained in the picture. Since in-picture prediction between neighboring treeblock rows within the picture is allowed, inter-processor / inter-core communication enabling in-picture prediction can be significant. WPP partitioning does not generate additional NAL units. Therefore, WPP may not be used for MTU size matching. However, if MTU size matching is required, regular slices can be used in conjunction with WPP, which has a specific coding overhead.
[0117] A tile defines horizontal and vertical boundaries that divide a picture into tile columns and rows. The scan order of the CTB can be local within the tile, in the order of the tile's CTB raster scan. Therefore, a tile can be fully decoded in the picture's tile raster scan order before the top-left CTB of the next tile is decoded. Similar to regular slices, tiles break in-picture prediction dependencies as well as entropy decoding dependencies. However, tiles may not be contained within individual NAL units. Consequently, tiles may not be used for MTU size matching. Each tile can be processed by a single processor / core. Inter-processor / inter-core communication used for in-picture prediction between processing units decoding neighboring tiles may be limited to loop filtering involving the sharing of reconstructed samples and metadata, and carrying a shared slice header if the slice contains more than one tile. If one or more tiles or WPP segments are included in a slice, the entry point byte offset for each tile or WPP segment other than the first tile or WPP segment of the slice may be signaled in the slice header.
[0118] For simplicity, HEVC uses specific restrictions on the application of four different picture segmentation schemes. A coded video sequence may not contain both tiles and wavefronts for most profiles specified in HEVC. Additionally, for each slice and / or tile, one or all of the following conditions must be met: All coded tree blocks of a slice are contained within the same tile. Also, all coded tree blocks of a tile are contained within the same slice. Additionally, a wavefront segment contains exactly one CTB row. If WPP is in use, a slice starting within a CTB row must end within the same CTB row.
[0119] HEVC may include a motion constrained tile set (MCTS) supplemental enhancement information (SEI) message, an MCTS extract information set SEI message, and an MCTS extract information nesting SEI message. The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to pointing to full sample locations within the MCTS and fragment sample locations using only full sample locations within the MCTS for interpolation. The use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is not permitted. In this way, each MCTS can be decoded independently without the presence of tiles not included in the MCTS. The MCTS extract information set SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction to generate a matching bitstream for the MCTS set. The information defines the number of MCTS sets and includes the number of extraction information sets, each containing RBSP bytes of the replacement video parameter set (VPS), SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting sub-bitstreams according to the MCTS sub-bitstream extraction process, parameter sets such as VPS, SPS, and PPS may be rewritten or replaced. The slice header may also be updated because one or more of the slice address-related syntax elements may contain different values after sub-bitstream extraction.
[0120] VVC can subdivide a picture as described below. A picture can be subdivided into tile groups and tiles. A tile can be a sequence of CTUs covering a rectangular area of a picture. A tile group, also called a slice, can contain multiple tiles of a picture. Slice / tile groups can be configured according to raster scan mode and rectangular mode. In raster scan mode, a tile group / slice contains a sequence of tiles in raster scan order for the picture boundaries. In rectangular mode, a tile group / slice contains multiple tiles that collectively form a rectangular area of a picture. Tiles within a rectangular tile group are included in raster scan order for the tile group / slice.
[0121] 360-degree video applications (e.g., VR) can display only a portion of the sphere of content, and consequently, only a subset of the entire picture. Viewport-dependent 360-degree delivery can be used to reduce bitrates when delivering VR video via DASH. Viewport-dependent coding can split the entire sphere / projected picture (e.g., using cubemap projection) into multiple MCTSs. Then, two or more bitstreams can be encoded at different spatial resolutions or qualities. The MCTS(s) from the high-resolution / quality bitstream are sent to a decoder for the viewport to be displayed (e.g., the front viewport). The remaining viewports use MCTS from the low-resolution / quality bitstream. These MCTS are packed in a specific way and then sent to a receiver for decoding. Typically, the viewport viewed by the user is represented by the high-resolution / quality MCTS, providing a positive viewing experience. When the user rotates to view a different viewport (e.g., the left or right viewport), the displayed content comes from the low-resolution / quality viewport. This continues for a short period until the system retrieves the high-resolution / quality MCTS for that viewport.
[0122] A delay occurs between the time the user rotates and the time the higher resolution / quality representation of the viewport is displayed. This delay depends on how quickly the system can fetch the higher resolution / quality MCTS for that viewport. This depends on the IRAP duration, which is the interval between two IRAP occurrences. This is because the MCTS for the new viewport can be decoded starting from the IRAP picture. If the IRAP duration is coded every second, the following applies. The best-case scenario for delay is equivalent to the network round-trip delay, where the user rotates to view the new viewport just before the system begins fetching a new segment / IRAP duration. In this scenario, the system can immediately request the higher resolution / quality MCTS for the new viewport. Therefore, the only delay is the network round-trip delay, which includes the fetching request delay plus the transmission time of the requested MCTS. This assumes that the minimum buffering delay can be set to 0 or other negligible value. For example, the network round-trip delay could be approximately 200 milliseconds. The worst-case scenario for latency involves adding network round-trip delay to the IRAP duration when the user rotates to view a new viewport immediately after the system has already requested the next segment. To mitigate this worst-case scenario, the bitstream can be encoded with more frequent IRAP pictures to allow for shorter IRAP durations. This reduces overall latency. However, increasing the number of IRAP pictures increases bandwidth, thereby lowering compression efficiency.
[0123] Mixed NAL unit types can be used within a picture by adding a PPS flag that specifies whether all VCL NAL units in the picture have the same NAL unit type. Additionally, constraints may be added requiring that for any specific picture, all VCL NAL units have the same NAL unit type, or that some VCL NAL units have a specific IRAP NAL unit type while others have a specific non-IRAP VCL NAL. An example of the mechanism description is as follows.
[0124] The IRAP picture is a coded picture in which mixed_nalu_types_in_pic_flag is equal to 0 and each VCL NAL unit has a NalUnitType in the range of IDR_W_RADL to CRA_NUT (inclusive).
[0125] An exemplary PPS RBSP syntax is as follows.
[0126] pic_parameter_set_rbsp( ) { explainer pps_pic_parameter_set_id ue(v) pps_seq_parameter_set_id ue(v) mixed_nalu_types_in_pic_flag u(1) ...
[0127] The following are exemplary NAL unit header semantics. An IDR picture with the same NalUnitType as IDR_N_LP does not have an associated leading picture in the bitstream. An IDR picture with the same NalUnitType as IDR_W_RADL does not have an associated RASL picture in the bitstream, but may have an associated RADL picture in the bitstream. For VCL NAL units of a specific picture, the following applies: If mixed_nalu_types_in_pic_flag is equal to 0, all VCL NAL units must have the same value for nal_unit_type. Otherwise, some VCL NAL units must have specific IRAP NAL unit type values (e.g., values of nal_unit_type in the range from IDR_W_RADL to CRA_NUT (inclusive), while all other VCL NAL units must have specific non-IRAP VCL NAL unit types (e.g., values of nal_unit_type in the range from TRAIL_NUT to RSV_VCL_15 (inclusive), equivalent to GRA_NUT). Exemplary PPS RBSP semantics are as follows. pps_seq_parameter_set_id specifies the sps_seq_parameter_set_id value for the active SPS. The value of pps_seq_parameter_set_id must be between 0 and 15 (inclusive). mixed_nalu_types_in_pic_flag can be set to 1 to specify that each picture referencing the PPS has multiple VCL NAL units and that these NAL units do not have the same value of nal_unit_type. mixed_nalu_types_in_pic_flag can be set to 0 to specify that the VCL NAL units of each picture referencing the PPS have the same value of nal_unit_type.If sps_idr_rpl_present_flag is equal to 0, the value of mixed_nalu_types_in_pic_flag must be equal to 0.
[0128] A picture can be divided into sub-pictures. The presence of sub-pictures can be indicated in the SPS along with other sequence-level information about the sub-pictures. Whether the boundaries of a sub-picture are treated as picture boundaries during the decoding process (excluding in-loop filtering operations) can be controlled by the bitstream. Whether in-loop filtering across sub-picture boundaries is disabled can be controlled by the bitstream for each sub-picture. The deblocking filter (DBF), SAO, and adaptive loop filter (ALF) processes are updated to control in-loop filtering operations across sub-picture boundaries. The sub-picture width, height, horizontal offset, and vertical offset can be signaled to the luminance sample unit of the SPS. Sub-picture boundaries can be limited to slice boundaries. Treating sub-pictures as pictures during the decoding process (excluding in-loop filtering operations) is specified by updating the coding_tree_unit() syntax, the derivation process for advanced temporal luminance motion vector prediction, the luminance sample bilinear interpolation process, the luminance sample 8-tab interpolation filtering process, and the chroma sample interpolation process. Sub-picture identifiers (IDs) are explicitly specified in the SPS and included in the tile group header, allowing sub-picture sequences to be extracted without needing to modify the VCL NAL unit. The Output sub-picture set (OSPS) can specify prescriptive extraction and match points for sub-pictures and their sets.
[0129] The preceding system has a specific problem. A bitstream may contain a picture that has both IRAP and non-IRAP slices. If such a picture follows in the decoding order and a slice of the picture covering the same picture area as the picture's IRAP slice references a picture that is earlier than the picture in the decoding order for inter-prediction, an error will occur.
[0130] The present disclosure includes improved techniques for supporting subpicture or MCTS-based random access in video coding. More specifically, the present disclosure discloses a method for imposing certain constraints on IRAP slices within a picture having both IRAP and non-IRAP slices. This is a description of the VVC technique. However, this technique may also be applied to other video / media codec specifications.
[0131] For example, to reference a picture faster than the mixed NAL unit picture in the decoding order for inter-prediction, a constraint is added to ensure that (1) the mixed NAL unit picture follows the decoding order and (2) the picture slice covering the same picture area as the picture's IRAP slice does not reference it. An exemplary implementation is as follows.
[0132] An IRAP picture can be defined as a coded picture where mixed_nalu_types_in_pic_flag is equal to 0 and each VCL NAL unit has a NalUnitType in the range of IDR_W_RADL to CRA_NUT (inclusive).
[0133] An exemplary PPS RBSP syntax is as follows.
[0134] pic_parameter_set_rbsp( ) { explainer pps_pic_parameter_set_id ue(v) pps_seq_parameter_set_id ue(v) mixed_nalu_types_in_pic_flag u(1) ...
[0135] The following are exemplary NAL unit header semantics. An IDR picture with the same NalUnitType as IDR_N_LP does not have an associated leading picture in the bitstream. An IDR picture with the same NalUnitType as IDR_W_RADL does not have an associated RASL picture in the bitstream, but may have an associated RADL picture in the bitstream. For VCL NAL units of a specific screen, the following applies: If mixed_nalu_types_in_pic_flag is equal to 0, all VCL NAL units must have the same value for nal_unit_type. Otherwise, some VCL NAL units must have specific IRAP NAL unit type values (e.g., values of nal_unit_type in the range from IDR_W_RADL to CRA_NUT (inclusive), while all other VCL NAL units must have specific non-IRAP VCL NAL unit types (e.g., values of nal_unit_type in the range from TRAIL_NUT to RSV_VCL_15 (inclusive), or equivalent to GRA_NUT). Exemplary PPS RBSP semantics are as follows. pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the active SPS. The value of pps_seq_parameter_set_id must be in the range from 0 to 15 (inclusive). mixed_nalu_types_in_pic_flag can be set to 1 to specify that each picture referencing the PPS has multiple VCL NAL units and that these NAL units do not have the same value of nal_unit_type. mixed_nalu_types_in_pic_flag can be set to 0 to specify that the VCL NAL units of each picture referencing the PPS have the same value of nal_unit_type.If sps_idr_rpl_present_flag is equal to 0, the value of mixed_nalu_types_in_pic_flag must be equal to 0. Additionally, for each IRAP slice of picture picA that has at least one non-IRAP slice, the following applies: The IRAP slice must belong to the sub-picture subpicA, and the boundaries of the sub-picture are treated as picture boundaries during the decoding process (excluding in-loop filtering operations). For example, the value of sub_pic_treated_as_pic_flag[i] for subpicA must be equal to 1. The IRAP slice must not belong to a sub-picture of the same picture that contains one or more non-IRAP slices. For all subsequent layer access units (AU) in the decoding order, neither RefPicList[0] nor RefPicList[1] of the slice in subpicA must contain any picture that precedes picA in the decoding order of the active entry.
[0136] FIG. 9 is a schematic diagram of an exemplary video coding device (900). The video coding device (900) is suitable for implementing the disclosed examples / executions as described herein. The video coding device (900) includes a downstream port (920), an upstream port (950), and / or a transceiver unit (Tx / Rx) (910) including a transmitter and / or receiver for communicating data upstream and / or downstream through a network. The video coding device (900) also includes a processor (930) including a logic unit and / or a central processing unit (CPU) for processing data, and a memory (932) for storing data. The video coding device (900) may also include an electric, optical-to-electrical (OE) component, an electrical-to-optical (EO) component, and / or a wireless communication component connected to an upstream port (950) and / or a downstream port (920) for communication of data through an electric, optical, or wireless communication network. The video coding device (900) may also include an input and / or output (I / O) device (960) for communicating data with a user. The I / O device (960) may include an output device such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device (960) may also include an input device such as a keyboard, a mouse, a trackball, etc., and / or a corresponding interface for interacting with these output devices.
[0137] The processor (930) is implemented by hardware and software. The processor (930) may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor (930) communicates with downstream ports (920), Tx / Rx (910), upstream ports (950), and memory (932). The processor (930) includes a coding module (914). The coding module (914) implements the disclosed embodiments described herein, such as methods (100, 1000, 1100), mechanisms (600), and / or applications (700) that can use images divided according to a bitstream (500) and / or a flexible video tiling scheme (800). The coding module (914) may also implement any other method / mechanism described herein. Additionally, the coding module (914) may implement a codec system (200), an encoder (300), and / or a decoder (400). For example, the coding module (914) may code VR video containing a current picture having both an IRAP NAL unit and a non-IRAP NAL unit. For example, the IRAP NAL unit may be included in a sub-picture. If this occurs, the coding module (914) may also limit slices in the picture following the current picture included in the sub-picture. Such slices may be prevented from referencing a reference picture preceding the current picture. Thus, the coding module (914) enables the video coding device (900) to provide additional functionality and / or coding efficiency when coding video data. In this way, the coding module (914) not only improves the functionality of the video coding device (900) but also solves problems specific to video coding technology.Additionally, the coding module (914) converts the video coding device (900) into a different state. Alternatively, the coding module (914) may be implemented as a command stored in memory (932) and executed by the processor (930) (e.g., as a computer program product stored on a non-transient medium).
[0138] Memory (932) includes one or more types of memory such as disk, tape drive, solid state drive, read-only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), static random-access memory (SRAM), etc. Memory (932) may be used as an overflow data storage device to store a program when such a program is selected for execution, and to store instructions and data read during program execution.
[0139] FIG. 10 is a flowchart of an exemplary method (1000) for encoding a video sequence comprising a picture having a NAL unit type mixed with a bitstream (800) such as a VR picture video stream (500) having an RPL structure (700). The method (1000) may encode such a bitstream according to constraints (600). The method (1000) may be used by an encoder such as a codec system (200), an encoder (300), and / or a video coding device (900) when performing the method (100).
[0140] The method (1000) may begin when an encoder receives a video sequence containing multiple pictures, such as a VR picture, and decides to encode the video sequence into a bitstream, for example, based on user input. The bitstream may contain VR video data. The VR video data may contain pictures representing a sphere of content at a corresponding moment of the video sequence. The pictures may be divided into sets of sub-pictures. For example, each sub-picture may contain video data corresponding to the viewport of the VR video. Additionally, the various sub-pictures may contain IRAP NAL units and non-IRAP NAL units at varying frequencies. In step 1001, the encoder determines that the current picture contains multiple VCL NAL units that do not have the same NAL unit type. For example, the VCL NAL units may include IRAP NAL units having an IRAP NAL unit type and non-IRAP NAL units having a non-IRAP NAL unit type. For example, IRAP NAL unit types may include IDR NAL unit types or CRA NAL unit types. Additionally, non-IRAP NAL unit types may include trailing NAL unit types, RASL NAL unit types and / or RADL NAL unit types.
[0141] In step 1003, the encoder encodes the flag into the bitstream. The flag indicates that the first NAL unit type value for the VCL NAL unit of the current picture is different from the second NAL unit type value for the VCL NAL unit of the current picture. In the example, the flag may be encoded into the bitstream's PPS. As a specific example, the flag may be mixed_nalu_types_in_pic_flag (827). mixed_nalu_types_in_pic_flag may be set to 1 if it specifies that each picture referencing the PPS has two or more VCL NAL units and that the VCL NAL units do not have the same value of nal_unit_type.
[0142] In step 1005, the encoder determines that subpicA of the current picture is associated with an IRAP NAL unit type. In step 1007, the encoder encodes the current picture into a bitstream. For example, the encoder ensures that all slices of the current picture located in subpicA are associated with the same NAL unit type (e.g., IRAP NAL unit type). The encoder can encode subpicA based on intra-prediction. The encoder can also determine that subpicB contains slices of a non-IRAP NAL unit type. Therefore, the encoder can encode the current subpicB based on inter-prediction.
[0143] In step 1009, the encoder may prepare to encode a subsequent picture that follows the current picture in the decoding order. For example, the encoder may generate an active entry in the reference picture list for a slice located at subpicA in the subsequent picture. An active entry for a specified subsequent picture indicates a picture that can be used as a reference picture when performing an inter-predictive encoding process for the specified subsequent picture. Specifically, an active entry for subpicA in the subsequent picture is restricted so that this active entry does not refer to any reference picture that precedes the current picture in the decoding order. This constraint is used when the subpicA of the current picture is associated with an IRAP NAL unit type. This constraint ensures that a slice of a subpicture following an IRAP subpicture does not refer to a picture preceding the IRAP subpicture, otherwise it may cause a coding error if the IRAP subpicture is used as a random access point. A slice of a subsequent picture that does not follow an IRAP subpicture (e.g., located in subpicB and followed by a non-IRAP NAL unit) may continue to refer to the picture preceding the current picture.
[0144] In step 1011, the encoder encodes the subsequent picture into a bitstream based on the reference picture list. For example, the encoder may encode the subsequent picture based on inter-prediction and / or intra-prediction according to the NAL type associated with the slice within the corresponding sub-picture. The inter-prediction process uses the reference picture list. The reference picture list may include reference picture list 0 and reference picture list 1. Additionally, the reference picture list may be coded into a reference picture list structure. The reference picture list structure may also be encoded into a bitstream. The encoder may also store the bitstream for communication to the decoder.
[0145] FIG. 11 is a flowchart of an exemplary method (1100) for decoding a video sequence containing a picture having a mixed NAL unit type from a bitstream such as a bitstream (800) containing a VR picture video stream (500) having an RPL structure (700). The method (1100) can decode such a bitstream according to constraints (600). The method (1100) can be used by a decoder such as a codec system (200), a decoder (400), and / or a video coding device (900) when performing the method (100).
[0146] Method (1100) may begin when a decoder begins to receive a bitstream of coded data representing a video sequence, for example, as a result of method (1000). The bitstream may include a VR video sequence containing multiple pictures, such as a VR picture. The bitstream may include VR video data. The VR video data may include pictures representing a sphere of content at a corresponding moment of the video sequence. A picture may be divided into a set of sub-pictures. For example, each sub-picture may include video data corresponding to the viewport of the VR video. Additionally, the various sub-pictures may include IRAP NAL units and non-IRAP NAL units at varying frequencies. In step 1101, the decoder receives the bitstream. The bitstream includes a current picture containing multiple VCL NAL units that do not have the same NAL unit type. For example, the VCL NAL units may include IRAP NAL units having an IRAP NAL unit type and non-IRAP NAL units having a non-IRAP NAL unit type. For example, IRAP NAL unit types may include IDR NAL unit types or CRA NAL unit types. Additionally, non-IRAP NAL unit types may include trailing NAL unit types, RASL NAL unit types and / or RADL NAL unit types.
[0147] In step 1103, the decoder determines, based on the flag, that the first NAL unit type value for the VCL NAL unit of the current picture is different from the second NAL unit type value for the VCL NAL unit of the current picture. In the example, the bitstream may include a PPS associated with the current picture. The flag may be obtained from the PPS. As a specific example, the flag may be mixed_nalu_types_in_pic_flag (827). mixed_nalu_types_in_pic_flag may be set to 1 if it specifies that each picture referencing the PPS has two or more VCL NAL units and that the VCL NAL units do not have the same value of nal_unit_type.
[0148] In step 1105, the decoder may determine that all slices of the current picture located in subpicA are associated with the same NAL unit type. The decoder may also decode subpicA and / or the current picture based on the NAL unit type of the slices. For example, subpicA of the current picture may contain an IRAP NAL unit. In this case, subpicA may be decoded according to intra prediction. The decoder may also determine that subpicB contains slices of a non-IRAP NAL unit type. Therefore, the decoder may decode the current subpicB according to inter prediction.
[0149] In step 1107, the decoder may obtain an active entry of a reference picture list for a slice located in subpicA in a subsequent picture that follows the current picture in the decoding order. An active entry for a specified subsequent picture indicates a picture that can be used as a reference picture when performing an inter-predictive decoding process for the specified subsequent picture.
[0150] In step 1109, the decoder may determine that if the subpicA of the current picture is associated with an intra-random access point (IRAP) NAL unit type, the active entry does not refer to any reference picture preceding the current picture in the decoding order. This constraint ensures that a slice of a subpicture following an IRAP subpicture does not refer to a picture preceding the IRAP subpicture, otherwise a coding error may occur when the IRAP subpicture is used as a random access point. A slice of a subsequent picture that does not follow an IRAP subpicture (e.g., located in subpicB and following a non-IRAP NAL unit) may continue to refer to a picture preceding the current picture.
[0151] In step 1111, the decoder can decode a subsequent picture based on the active entry of the reference picture list. For example, the decoder can decode the subsequent picture based on inter-prediction and / or intra-prediction according to the NAL type associated with the slice of the corresponding sub-picture. The inter-prediction process uses the reference picture list. The reference picture list may include reference picture list 0 and reference picture list 1. Additionally, the reference picture list may be obtained from a reference picture list structure coded into a bitstream. The decoder may pass the current picture, the subsequent picture, and / or its sub-picture (e.g., subpicA or subpicB) for display as part of the decoded video sequence.
[0152] FIG. 12 is a schematic diagram of an exemplary system (1200) for coding a video sequence comprising a picture having a NAL unit type mixed with a bitstream (800) containing a VR picture video stream (500) coded according to constraints (600) having an RPL structure (700). The system (1200) may be implemented by an encoder and a decoder such as a codec system (200), an encoder (300), a decoder (400), and / or a video coding device (900). Additionally, the system (1200) may be used when implementing the method (100, 1000, and / or 1100).
[0153] The system (1200) includes a video encoder (1202). The video encoder (1202) includes a determination module (1201) for determining that the current picture contains multiple VCL NAL units that do not have the same NAL unit type. The determination module (1201) is also for determining that a subpicA in the current picture is associated with an IRAP NAL unit type. The video encoder (1202) further includes a generation module (1203) for generating an active entry of a reference picture list for a slice located in the subpicA of a subsequent picture following the current picture in a decoding order, wherein the active entry does not refer to any reference picture preceding the current picture when the subpicA of the current picture is associated with an IRAP NAL unit type. The video encoder (1202) further includes an encoding module (1205) for encoding the subsequent picture into a bitstream based on the reference picture list. The video encoder (1202) further includes a storage module (1207) for storing a bitstream for communication toward a decoder. The video encoder (1202) further includes a transmission module (1209) for transmitting a bitstream toward a video decoder (1210). The video encoder (1202) may be further configured to perform any step of the method (1000).
[0154] The system (1200) also includes a video decoder (1210). The video decoder (1210) includes a receiving module (1211) for receiving a bitstream containing a current picture containing multiple VCL NAL units that do not have the same NAL unit type. The video decoder (1210) further includes an acquisition module (1213) for acquiring an active entry of a reference picture list for a slice located in subpicA in a subsequent picture following the current picture in the decoding order. The video decoder (1210) further includes a determination module (1215) for determining that the active entry does not refer to any reference picture preceding the current picture in the decoding order when the subpicA in the current picture is associated with an IRAP NAL unit type. The video decoder (1210) further includes a decoding module (1217) for decoding the subsequent picture based on the active entry of the reference picture list. The video decoder (1210) further includes a delivery module (1219) for delivering a subsequent picture to be displayed as part of the decoded video sequence. The video decoder (1210) may be further configured to perform any step of the method (1100).
[0155] The first component is directly coupled to the second component when there is no intermediate component, excluding lines, traces, or other media between the first component and the second component. When there is an intermediate component other than lines, traces, or other media between the first component and the second component, the first component is indirectly coupled to the second component. The term "coupled" and its variations include both direct and indirect coupling. The use of the term "approximately" implies a range including ±10% of the subsequent number, unless otherwise specified.
[0156] Furthermore, it should be understood that the steps of the exemplary method described herein must not necessarily be performed in the order described, and that the order of the steps of such method is understood to be merely exemplary. Likewise, additional steps may be included in such method, and specific steps may be omitted or combined in a method consistent with various embodiments of the present disclosure.
[0157] Although various embodiments have been provided in this disclosure, it will be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of this disclosure. These examples should be considered illustrative rather than restrictive, and their intent is not limited to the details provided herein. For example, various elements or components may be combined or integrated into other systems, or certain features may be omitted or not implemented.
[0158] Additionally, the technologies, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. Other examples of modifications, substitutions, and variations may be identified by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
Claim 1 A method for decoding a bitstream, comprising the step of receiving the bitstream — wherein the bitstream comprises coded data of a current picture and a subsequent picture following the current picture in a decoding order, and the current picture and the subsequent picture both comprise a first sub-picture and a second sub-picture, and the bitstream further comprises a picture parameter set (PPS), wherein the PPS comprises a mixed_nalu_types_in_pic_flag, and the mixed_nalu_types_in_pic_flag identical to 1 specifies that each picture referencing the PPS has two or more video coding layer (VCL) network abstraction layer (NAL) units, and the two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value —; the step of obtaining an active entry of a reference picture list for the first sub-picture of the subsequent picture — wherein the active entry is the current A method for decoding a bitstream comprising the step of: when a first sub-picture of a picture is associated with an intra-random access point (IRAP) NAL unit type, not including a reference to any reference picture preceding the current picture in the decoding order, and the first sub-picture of the subsequent picture and the first sub-picture of the current picture having the same sub-picture index value; and decoding the first sub-picture of the subsequent picture based on the reference picture list active entry. Claim 2 A method for decoding a bitstream, wherein, in claim 1, the IRAP NAL unit type is a Clean Random Access (CRA) NAL unit type. Claim 3 A method for decoding a bitstream according to claim 1, wherein the IRAP NAL unit type is an Instantaneous Decoder Refresh (IDR) NAL unit type. Claim 4 A method for decoding a bitstream according to claim 1, wherein all slices of the current picture located in the first sub-picture are associated with the same NAL unit type. Claim 5 A method for decoding a bitstream according to claim 1, wherein when the mixed_nalu_types_in_pic_flag is equal to 1, the VCL NAL units of one or more sub-pictures of the current picture all have specific values of the IRAP NAL unit type, and all other VCL NAL units of the current picture have specific values of the non-IRAP NAL unit type. Claim 6 A method for decoding a bitstream according to claim 1, wherein when the mixed_nalu_types_in_pic_flag is equal to 0, all VCL NAL units of the current picture have the same NAL unit type. Claim 7 A method for decoding a bitstream, wherein the second sub-picture of the current picture is associated with a non-IRAP NAL unit type. Claim 8 A method for encoding a bitstream, comprising the step of encoding a current picture and a subsequent picture following the current picture into a bitstream — wherein the current picture and the subsequent picture both include a first subpicture and a second subpicture —; and the step of encoding a picture parameter set (PPS) into the bitstream — wherein the PPS includes a mixed_nalu_types_in_pic_flag, and the mixed_nalu_types_in_pic_flag, which is identical to 1, specifies that each picture referencing the PPS has two or more Video Coding Layer (VCL) Network Abstraction Layer (NAL) units and that the two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value, and the active entry of the reference picture list for the first subpicture of the subsequent picture is any reference preceding the current picture in the decoding order when the first subpicture of the current picture is associated with an Intra Random Access Point (IRAP) NAL unit type A method for encoding a bitstream including: not including a reference to a picture, wherein the first sub-picture of the subsequent picture and the first sub-picture of the current picture have the same sub-picture index value. Claim 9 A method for encoding a bitstream, wherein the IRAP NAL unit type is a Clean Random Access (CRA) NAL unit type, in paragraph 8. Claim 10 A method for encoding a bitstream, wherein, in paragraph 8, the IRAP NAL unit type is an instantaneous decoder refresh (IDR) NAL unit type. Claim 11 A method for encoding a bitstream in claim 8, wherein all slices of the current picture located in the first sub-picture are associated with the same NAL unit type. Claim 12 A method for encoding a bitstream according to claim 8, wherein when the mixed_nalu_types_in_pic_flag is equal to 1, the VCL NAL units of one or more sub-pictures of the current picture all have specific values of the IRAP NAL unit type, and all other VCL NAL units of the current picture have specific values of the non-IRAP NAL unit type. Claim 13 A method for encoding a bitstream according to claim 8, wherein when the mixed_nalu_types_in_pic_flag is equal to 0, all VCL NAL units of the current picture have the same NAL unit type. Claim 14 In paragraph 8, a method for encoding a bitstream, wherein the second sub-picture of the current picture is associated with a non-IRAP NAL unit type. Claim 15 A video coding device comprising a processing circuit configured to perform the method of any one of claims 1 to 14. Claim 16 A non-transient computer-readable medium comprising a computer program for use by a video coding device, wherein the computer program comprises computer-executable instructions stored in the non-transient computer-readable medium to cause the video coding device to perform the method of any one of claims 1 to 14 when executed by a processor. Claim 17 As a decoder, receiving means for receiving a bitstream containing coded data of a current picture and a subsequent picture following the current picture in a decoding order ― wherein the current picture and the subsequent picture both include a first sub-picture and a second sub-picture, and the bitstream further includes a picture parameter set (PPS), wherein the PPS includes a mixed_nalu_types_in_pic_flag, and the mixed_nalu_types_in_pic_flag identical to 1 specifies that each picture referencing the PPS has two or more Video Coding Layer (VCL) Network Abstraction Layer (NAL) units and that the two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value ―; acquiring means for acquiring an active entry of a reference picture list for the first sub-picture of the subsequent picture ― wherein the first sub-picture of the current picture is associated with an Intra Random Access Point (IRAP) NAL unit type, the active A decoder comprising: an entry that does not include a reference to any reference picture preceding the current picture in the decoding order, and a first sub-picture of the subsequent picture and a first sub-picture of the current picture have the same sub-picture index value; and a decoding means for decoding the first sub-picture of the subsequent picture based on the reference picture list active entry. Claim 18 In paragraph 17, the decoder is further configured to perform the method of any one of paragraphs 2 through 7. Claim 19 As an encoder, the encoding means for encoding a current picture and a subsequent picture following the current picture into a bitstream — said current picture and said subsequent picture both include a first subpicture and a second subpicture — said encoding means further for encoding a picture parameter set (PPS) into the bitstream, said PPS includes a mixed_nalu_types_in_pic_flag, said mixed_nalu_types_in_pic_flag identical to 1 specifies that each picture referencing said PPS has two or more Video Coding Layer (VCL) Network Abstraction Layer (NAL) units, said two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value, and the active entry of the reference picture list for the first subpicture of said subsequent picture is, in the decoding order, said current An encoder that does not include a reference to any reference picture preceding the picture, wherein the first sub-picture of the subsequent picture and the first sub-picture of the current picture have the same sub-picture index value. Claim 20 In paragraph 19, the encoder is further configured to perform the method of any one of paragraphs 9 through 14. Claim 21 A decoder comprising: one or more processors as a decoder; and a non-transient computer-readable storage medium connected to said processors and storing a program for execution by said processors, wherein said program is configured to cause said decoder to perform a method of any one of claims 1 through 7 when executed by said processors. Claim 22 An encoder comprising: one or more processors as an encoder; and a non-transient computer-readable storage medium connected to said processors and storing a program for execution by said processors, wherein said program is configured to cause said encoder to perform the method of any one of claims 8 through 14 when executed by said processors. Claim 23 A device for storing a bitstream comprises a receiver and a storage medium, wherein the receiver is configured to receive the bitstream and the storage medium is configured to store the bitstream, wherein the bitstream comprises coded data of a current picture and a subsequent picture following the current picture in a decoding order, wherein the current picture and the subsequent picture both comprise a first sub-picture and a second sub-picture, wherein the bitstream further comprises a picture parameter set (PPS), wherein the PPS comprises a mixed_nalu_types_in_pic_flag, wherein the mixed_nalu_types_in_pic_flag, which is identical to 1, specifies that each picture referencing the PPS has two or more Video Coding Layer (VCL) Network Abstraction Layer (NAL) units and that the two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value, and the active entry of the reference picture list for the first sub-picture of the subsequent picture is the first sub of the current picture A device for storing a bitstream, wherein, when a picture is associated with an intra-random access point (IRAP) NAL unit type, it does not include a reference to any reference picture preceding the current picture in the decoding order, and the first sub-picture of the subsequent picture and the first sub-picture of the current picture have the same sub-picture index value. Claim 24 A method for storing a bitstream, comprising the step of acquiring a bitstream; and storing the bitstream in one or more storage media; wherein the bitstream includes a current picture and subsequent pictured data following the current picture in a decoding order, and both the current picture and the subsequent picture include a first sub-picture and a second sub-picture, and the bitstream further includes a picture parameter set (PPS), and the PPS includes a mixed_nalu_types_in_pic_flag, and the mixed_nalu_types_in_pic_flag identical to 1 specifies that each picture referencing the PPS has two or more Video Coding Layer (VCL) Network Abstraction Layer (NAL) units and that the two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value, and the active entry of the reference picture list for the first sub-picture of the subsequent picture is, if the first sub-picture of the current picture is associated with an Intra Random Access Point (IRAP) NAL unit type, the current in the decoding order A method for storing a bitstream, wherein the first sub-picture of the subsequent picture and the first sub-picture of the current picture have the same sub-picture index value, without including a reference to any reference picture preceding the picture. Claim 25 As a bitstream, the coded data of a current picture and a subsequent picture following the current picture in the decoding order ― both the current picture and the subsequent picture include a first sub-picture and a second sub-picture, and the bitstream further includes a picture parameter set (PPS), the PPS includes a mixed_nalu_types_in_pic_flag, and the mixed_nalu_types_in_pic_flag, which is identical to 1, specifies that each picture referencing the PPS has two or more Video Coding Layer (VCL) Network Abstraction Layer (NAL) units, and the two or more VCL NAL units do not have the same NAL unit type (nal_unit_type) value, and the active entry of the reference picture list for the first sub-picture of the subsequent picture does not include a reference to any reference picture preceding the current picture in the decoding order when the first sub-picture of the current picture is associated with an Intra Random Access Point (IRAP) NAL unit type. A bitstream including, without, the first sub-picture of the subsequent picture and the first sub-picture of the current picture have the same sub-picture index value.
Citation Information
Patent Citations
Concept for picture / video data streams allowing efficient reducibility or efficient random access
US20190014337A1