Techniques for Random Access Point Display and Picture Output in a Symbolized Video Stream
By signaling flags for IRAP and GDR pictures and constructing reference picture lists, the method addresses the challenge of managing random access and output processing in multi-layer video streams, enhancing decoder stability and playback efficiency.
Patent Information
- Application Number
- JP2024130815
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-14
- Filing Date
- 2024-08-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-05-18
AI Technical Summary
Existing video encoding and decoding technologies face challenges in managing random access points and output processing in encoded video streams with multiple layers, particularly in handling unavailable reference pictures and ensuring decoder stability during trick-mode playback.
The method involves signaling flags in the video bitstream to indicate intra random access point (IRAP) and hierarchical decoding refresh (GDR) pictures, constructing reference picture lists, and verifying bitstream compliance to ensure correct decoding and output processing of pictures, even in the absence of available reference pictures.
This approach enhances decoder stability and enables efficient random access and output processing in video streams with multiple layers, ensuring seamless playback and reducing the risk of decoder crashes during trick-mode operations.
Smart Images

Figure 0007701526000002 
Figure 0007701526000003 
Figure 0007701526000004
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 037,903, filed on June 11, 2020; U.S. Provisional Patent Application No. 63 / 036,335, filed on June 8, 2020; U.S. Provisional Patent Application No. 63 / 035,274, filed on June 5, 2020; U.S. Provisional Patent Application No. 63 / 027,826, filed on May 20, 2020; and U.S. Patent Application No. 17 / 320,764, filed on May 14, 2021, the disclosures of which are hereby incorporated by reference in their entireties.
[0002] Embodiments of the present disclosure relate to video encoding and decoding, and more particularly, to random access pictures and their output processing in an encoded video stream having multiple layers.
Background Art
[0003] Video encoding and decoding using inter - picture prediction with motion compensation has been used previously. Uncompressed digital video can include a series of pictures, and each picture can have spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate, for example, 60 pictures per second or 60 Hz (informally also known as the frame rate). Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. To use such video for one hour, a storage area of more than 600 GB is required.
[0004] One of the purposes of video encoding and decoding can be the reduction of redundancy in the input video signal by compression. Compression can help reduce the aforementioned bandwidth or storage requirements, sometimes by more than two orders of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique where an exact replica of the original signal can be restored from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal can be said to be small enough that the reconstructed signal is useful for the intended application. In the case of video, irreversible compression is widely adopted. The amount of distortion tolerated varies by application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television broadcast applications. The achievable compression ratio may reflect that higher compression ratios are obtained with higher allowable / tolerable distortion.
[0005] Video encoders and decoders can utilize several broad categories of techniques, such as motion compensation, transformation, quantization, entropy encoding, some of which are introduced below.
[0006] Previously, video encoders and decoders tended to operate at a given picture size that was defined for and remained constant for a coded video sequence (CVS), a Group of Pictures (GOP), or a similar multi-picture time frame. For example, in MPEG-2, the system design was used to change the horizontal resolution (and thus the picture size) according to factors such as the activity of the scene, but only for I pictures and thus usually for GOPs. Resampling of reference pictures for using different resolutions within a CVS is used, for example, in ITU-T Rec. H.263 Annex P. However, here the picture size does not change, only the reference pictures are resampled, and only a part of the picture canvas may be used (in the case of downsampling), or only a part of the scene may be captured (in the case of upsampling). Further, H.263 Annex Q enables resampling individual macroblocks by a factor of 2 up or down (in each dimension). Again, the picture size remains the same. Since the size of the macroblocks is fixed in H.263, there is no need to signal it.
[0007] The change of the picture size of the prediction picture has become more mainstream in the latest video coding. For example, VP9 enables resampling of the reference picture and changing the resolution of the entire picture. Similarly, according to a proposal made for VVC (e.g., Hendry et al., “On adaptive resolution change (ARC) for VVC”, Joint Video Team document JVET-M 0135-v1, January 9 - 19, 2019, which is incorporated herein in its entirety), resampling of the entire reference picture to different (higher or lower) resolutions becomes possible. In such documents, different candidate resolutions are proposed, which are encoded within the sequence parameter set and referenced by the per-picture syntax elements within the picture parameter set.
[0008] Bross et al., “Versatile Video Coding (Draft 9)”, Joint Video Experts Team document JVET-R 2001-vA, April 2020, which is incorporated herein in its entirety.
Summary of the Invention
Means for Solving the Problems
[0009] Within the encoded video stream. It is widely used to indicate random access point information in high-level syntax structures such as network abstraction layer (NAL) unit headers, parameter sets, picture headers, or slice headers. Based on the random access information, the decoded start picture associated with the random access picture is managed. In the present disclosure, some related syntax elements and constraints are described to clarify the decoded picture management related to random access processing.
[0010] When a video bitstream is randomly accessed by trick - mode play, an intra random access point (IRAP) picture can enable random access to an intermediate point of the bitstream and the successful decoding of the video bitstream at the random access point. One possible way is to gradually refresh the scene with a certain amount of restoration time. In VVC and other video codecs, gradual decoding refresh (GDR) pictures and access units (AUs) are defined to specify the syntax and semantics of the random access operation by gradual decoding refresh. In this disclosure, its syntax, semantics, and constraints are described to correctly specify the signaling and decoding process of GDR.
[0011] When one or more reference picture lists are constructed for inter - prediction in a P or B slice, one or more pictures may not be available for random access or due to unintentional picture loss. To avoid decoder crashes or unintentional behavior, it is desirable to generate unavailable pictures with initial setting values of pixels and parameters. After generating the unavailable pictures, it may be necessary to confirm the verification of all reference pictures in the reference picture list.
[0012] Embodiments of the present disclosure relate to random access pictures and output processing thereof in an encoded video stream having multiple layers. Embodiments of the present disclosure relate to random access pictures and leading picture output indications thereof in an encoded video stream having multiple layers. Embodiments of the present disclosure relate to signaling random access pictures using hierarchical decoding refresh and recovery points in an encoded video stream having multiple layers. Embodiments of the present disclosure relate to reference picture list construction and unavailable picture generation in an encoded video stream having multiple layers. Embodiments of the present disclosure include techniques for signaling adaptive picture sizes in a video bitstream.
[0013] One or more embodiments of the present disclosure include a method executed by at least one processor. The method includes receiving an encoded video stream including access units including pictures, signaling, at an access unit delimiter of the encoded video stream, a first flag indicating whether any one of an intra random access point (IRAP) picture and a hierarchical decoding refresh (GDR) picture is included in the access unit, signaling, within a picture header of the encoded video stream, a second flag indicating whether the picture is an IRAP picture, and decoding the picture as a current picture based on the signaling of the first flag and the second flag, wherein the values of the first flag and the second flag are aligned.
[0014] According to one embodiment, the method further includes signaling, within a picture header of the encoded video stream, a third flag indicating whether the picture is a GDR picture, wherein the values of the first flag and the third flag are aligned.
[0015] According to one embodiment, the third flag is signaled based on a second flag indicating that the picture is not an IRAP picture.
[0016] According to one embodiment, the first flag has a value indicating that the picture is one of an IRAP picture and a GDR picture, the second flag has a value indicating that the picture is an IRAP picture, and the method further includes a step of signaling, in a slice header of a slice of a picture of an encoded video stream, a third flag indicating whether any picture before the IRAP picture is output.
[0017] According to one embodiment, the method further includes a step of determining a network abstraction layer (NAL) unit type of a slice, wherein the third flag is signaled based on the determined NAL unit type.
[0018] According to one embodiment, the third flag is signaled based on a NAL unit type determined to be equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.
[0019] According to one embodiment, the method further includes a step of signaling, in a picture header of an encoded video stream, a fourth flag indicating whether the picture is a GDR picture, and the values of the first flag and the fourth flag are aligned.
[0020] According to one embodiment, the third flag is signaled based on a second flag indicating that the picture is not an IRAP picture.
[0021] According to one embodiment, the decoding step includes a step of constructing a reference picture list, a step of generating unavailable reference pictures in the reference picture list, and a step of checking bitstream compliance for the reference pictures in the reference picture list, where the number of entries indicated as being in the reference picture list is greater than or equal to the number of active entries indicated as being in the reference picture list, each picture referenced by an active entry in the reference picture list exists in a decoded picture buffer (DPB), has a time identifier value less than or equal to the time identifier value of the current picture, and the picture header flag indicates that each picture referenced by an entry in the reference picture list may be a reference picture rather than the current picture, and the following constraint is applied.
[0022] According to one embodiment, the step of checking bitstream compliance is performed based on a determination that the current picture is an independent decoder refresh (IDR) picture, a clean random access (CRA) picture, or a gradual decoding refresh (GDR) picture.
[0023] According to one or more embodiments, a system is provided. The system includes at least one processor configured to receive an encoded video stream including access units that include pictures, and a memory storing computer code. The computer code is configured to cause the at least one processor to signal, at an access unit delimiter of the encoded video stream, a first flag indicating whether the access unit includes either an Intra Random Access Point (IRAP) picture or a Gradual Decoding Refresh (GDR) picture; cause the at least one processor to signal, within a picture header of the encoded video stream, a second flag indicating whether the picture is an IRAP picture; and cause the at least one processor to decode a picture as a current picture based on signaling of the first flag and the second flag, the memory being configured such that a value of the first flag is aligned with a value of the second flag.
[0024] According to one embodiment, the computer code further includes a third signaling code configured to cause the at least one processor to signal, within a picture header of the encoded video stream, a third flag indicating whether the picture is a GDR picture, the value of the first flag being aligned with the value of the third flag.
[0025] According to one embodiment, the third flag is signaled based on a second flag indicating that the picture is not an IRAP picture.
[0026] According to one embodiment, the first flag has a value indicating that the picture is one of either an IRAP picture or a GDR picture, the second flag has a value indicating that the picture is an IRAP picture, and the computer code further includes a third signaling code configured to cause at least one processor to signal a third flag indicating whether any picture before the IRAP picture in the slice header of a slice of the encoded video stream has been output.
[0027] According to one embodiment, the computer code further includes determining a code configured to cause at least one processor to determine the network abstraction layer (NAL) unit type of the slice, and the third flag is signaled based on the determined NAL unit type.
[0028] According to one embodiment, the third flag is signaled based on the NAL unit type determined to be equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.
[0029] According to one embodiment, the computer code further includes a fourth signaling codec configured to cause at least one processor to signal a fourth flag indicating whether the picture in the picture header of the encoded video stream is a GDR picture, and the value of the first flag and the value of the fourth flag are aligned.
[0030] According to one embodiment, the third flag is signaled based on the second flag indicating that the picture is not an IRAP picture.
[0031] According to one embodiment, the decoding code includes a construction code configured to cause at least one processor to construct a reference picture list, a generation code configured to cause at least one processor to generate a reference picture that is not available for use in the reference picture list, and a verification code configured to cause at least one processor to verify the bitstream compliance of the reference pictures in the reference picture list, with the following constraints being applied: the number of entries indicated as being in the reference picture list is greater than or equal to the number of active entries indicated as being in the reference picture list; each picture referenced by an active entry in the reference picture list exists in a decoded picture buffer (DPB) and has a time identifier value less than or equal to the time identifier value of the current picture; each picture referenced by an entry in the reference picture list is not the current picture, and it is indicated that it may be a reference picture by a picture header flag.
[0032] According to one or more embodiments, a non-transitory computer-readable medium storing computer instructions is provided. When the computer instructions are executed by at least one processor that receives an encoded video stream including access units including pictures, the at least one processor is caused to perform the steps of signaling, at an access unit delimiter of the encoded video stream, a first flag indicating whether the access unit includes either an intra random access point (IRAP) picture or a hierarchical decoding refresh (GDR) picture; signaling, in a picture header of the encoded video stream, a second flag indicating whether the picture is an IRAP picture; and decoding the picture as the current picture based on the signaling of the first flag and the second flag, wherein the value of the first flag and the value of the second flag are aligned.
[0033] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0034]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15A
Figure 15B
Figure 16
Figure 17
Figure 18
Figure 19A
Figure 19B
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
DETAILED DESCRIPTION OF THE INVENTION
[0035] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In the case of unidirectional data transmission, a first terminal (110) can encode video data at a local location for transmission to another terminal (120) via the network (150). A second terminal (120) can receive the encoded video data of another terminal from the network (150), decode the encoded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0036] FIG. 1 shows, for example, a second pair of terminals (130, 140) provided to support bidirectional transmission of encoded video that may occur during a video conference. In the case of bidirectional data transmission, each terminal (130, 140) can encode video data captured at a local location for transmission to another terminal via the network (150). Each terminal (130, 140) can also receive the encoded video data transmitted by another terminal, decode the encoded data, and display the restored video data on a local display device.
[0037] In FIG. 1, the terminals (110-140) may be shown as a server, a personal computer, and a smart phone, and / or any other kind of terminal. For example, the terminals (110-140) may be a laptop computer, a tablet computer, a media player, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transmit encoded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (150) may not be important to the operation of the present disclosure, unless otherwise described herein below.
[0038] FIG. 2 shows the arrangement of video encoders and decoders in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media such as CDs, DVDs, memory sticks, and the like.
[0039] As shown in FIG. 2, the streaming system (200) may include a capture subsystem (213) that can include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). The uncompressed video sample stream (202) can provide a high data volume as compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bitstream (204) may include a lower data volume as compared to the sample stream and may be stored in a streaming server (205) for future use. One or more streaming clients (206) may access the streaming server (205) to obtain a video bitstream (209) that may be a replica of the encoded video bitstream (204).
[0040] In an embodiment, the streaming server (205) may also function as a Media-Aware Network Element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to match one or more of the streaming clients (206) with potentially different bitstreams. In an embodiment, the MANE may be provided separately from the streaming server (205) within the streaming system (200).
[0041] The streaming client (206) can include a video decoder (210) and a display (212). The video decoder (210) can decode, for example, a video bitstream (209) that is an input copy of an encoded video bitstream (204) and generate an output video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) can be encoded according to a specific video encoding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. In one example, a video encoding standard under development is informally known as Versatile Video Coding (VVC). Embodiments of the present disclosure can be used in the context of VVC.
[0042] FIG. 3 shows an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure.
[0043] The video decoder (210) can include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (). In at least one embodiment, the video decoder (210) can include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may also be partially or fully embodied in software executed on one or more CPUs having associated memory.
[0044] In this and other embodiments, a receiver (310) may receive one or more encoded video sequences to be decoded by a decoder (210), one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (312) that may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data along with other data, such as encoded audio data and / or auxiliary data streams, that may be transferred to respective using entities (not shown). The receiver (310) can separate the encoded video sequences from the other data. A buffer memory (315) may be coupled between the receiver (310) and an entropy decoder / parser (320) (hereinafter, "parser") to counter network jitter. The buffer (315) may not be used or may be small when the receiver (310) is receiving data from a sufficient bandwidth and controllable store-and-forward device or from an asynchronous network. The buffer (315) may be required for use in a best-effort packet network, such as the Internet, and may be relatively large and adaptable in size.
[0045] The video decoder (210) may include an analyzer (320) for reconstructing symbols (321) from an entropy - encoded video sequence. The categories of these symbols include, for example, information used to manage the operation of the decoder (210) and potentially information for controlling a rendering device such as a display (212) that may be coupled to the decoder as shown in FIG. 2. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The analyzer (320) can analyze / entropy - decode the received encoded video sequence. The encoding of the encoded video sequence can follow a video encoding technique or a video encoding standard and can follow principles well - known to those skilled in the art, including variable - length encoding, Huffman encoding, arithmetic encoding with or without context - dependence, etc. The analyzer (320) can extract a set of at least one sub - group parameter of at least one sub - group of pixels in the video decoder based on at least one parameter corresponding to a group. The sub - groups can include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, Coding Units (CUs), a block, Transform Units (TUs), Prediction Units (PUs), etc. The analyzer (320) can also extract from the encoded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.
[0046] The analyzer (320) can perform an entropy - decoding / analysis operation on the video sequence received from the buffer (315) to create symbols (321).
[0047] The reconstruction of symbol (321) may involve multiple different units depending on the type of the encoded video picture or a part thereof (such as inter-picture and intra-picture, inter-block and intra-block), and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by parser (320). Such a flow of subgroup control information between parser (320) and the following multiple units is not shown for clarity.
[0048] In addition to the function blocks already described, decoder 210 can conceptually be subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.
[0049] One unit can be a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) can receive from parser (320) quantization transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbol(s) (321). The scaler / inverse transform unit (351) can output a block with sample values that can be input to aggregator (355).
[0050] In some cases, the output samples of the scaler / inverse transform (351) may be related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use the surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory (358) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add the prediction information generated by the intra-prediction unit (352) to the output sample information from the scaler / inverse transform unit (351) on a sample-by-sample basis.
[0051] In other cases, the output samples of the scaler / inverse transform unit (351) may be related to inter-coded, potentially motion-compensated blocks. In such cases, the motion-compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols (321) associated with the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) to generate the output sample information (in this case, called residual samples or a residual signal). The address in the reference picture memory (357) from which the motion-compensation prediction unit (353) fetches the prediction samples can be controlled by a motion vector. The motion vector may be available to the motion-compensation prediction unit (353) in the form of, for example, a symbol (321) having X, Y, and reference picture components. Motion compensation may also include interpolation of the sample values fetched from the reference picture memory (357) when an exact sub-sample motion vector is used, a motion vector prediction mechanism, etc.
[0052] The output samples of the aggregator (355) can be subject to various loop filtering techniques in the loop filter unit (356). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and that can be used in the loop filter unit (356) as symbols (321) from the parser (320), but can also respond to meta information obtained during the decoding of an encoded picture or a previous (in decoding order) portion of an encoded video sequence, or to previously reconstructed and loop-filtered sample values.
[0053] The output of the loop filter unit (356) can be output to a rendering device such as a display (212) and can be a sample stream that can be stored in the reference picture memory (357) for use in future inter-picture prediction.
[0054] Once fully reconstructed, a particular encoded picture can be used as a reference picture for future prediction. When an encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture can become part of the reference picture memory (357), and a fresh current picture memory can be reallocated before starting the reconstruction of the next encoded picture.
[0055] The video decoder (210) can perform a decoding operation according to a predetermined video compression technique that can be documented in a standard such as ITU-T Rec.H.265. The encoded video sequence may conform to the syntax of the video compression technique or standard, in the sense that it conforms to the syntax specified by the video compression technique or standard being used, and to the syntax of the video compression technique document or standard, particularly the profile document therein. Also, in order to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstructed sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. The limits set by the level can, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specifications and the metadata for HRD buffer management signaled in the encoded video sequence.
[0056] In one embodiment, the receiver (310) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0057] FIG. 4 shows an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to one embodiment of the present disclosure.
[0058] The video encoder (203) can include, for example, an encoder such as a source encoder (430), an encoding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy encoder (445), a controller (450), and a channel (460).
[0059] The encoder (203) can receive video samples from a video source (201) (which is not part of the encoder) that can capture the video images encoded by the encoder (203).
[0060] The video source (201) can provide the source video sequence encoded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), can be in any color space (e.g., BT.601 Y CrCB, RGB,...), and can have a suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (201) can be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that give the appearance of motion when viewed in sequence. Each picture itself can be organized as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. One skilled in the art can easily understand the relationship between pixels and samples. Hereinafter, the description will focus on samples.
[0061] According to one embodiment, the encoder (203) can encode pictures of a source video sequence in real time or under any other time constraint as required by the application and compress them into an encoded video sequence (443). Enforcing an appropriate encoding speed is one function of the controller (450). The controller (450) may also control other functional units as described later and may be functionally coupled to these units. The coupling is not shown for clarity. Parameters set by the controller (450) may include rate control related parameters (such as picture skip, quantization, lambda value of the rate distortion optimization method, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) since it may relate to a video encoder (203) optimized for a specific system design.
[0062] Some video encoders operate in a way that is easily recognizable by those skilled in the art as an "encoding loop". As an overly simplified explanation, the encoding loop may consist of the encoding part of the source encoder (430) (which is responsible for creating symbols based on the input picture and reference pictures to be encoded), and the (local) decoder (433) embedded in the encoder (203) reconstructs the symbols to create sample data, as does the (remote) decoder (when the compression between the symbols and the encoded video bitstream is lossless for a particular video compression technique). The reconstructed sample stream is input into the reference picture memory (434). Since the decoding of the symbol stream results in a bit-exact result regardless of the location of the decoder (local or remote), the reference picture memory contents are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the decoder "sees" as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is known to those skilled in the art.
[0063] The operation of the "local" decoder (433) may be the same as the operation of the "remote" decoder (210), which has already been described in detail above in connection with FIG. 3. However, since symbols are available and the encoding / decoding of symbols to the encoded video sequence by the entropy encoder (445) and the parser (320) can be lossless, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) may not be fully implemented in the local decoder (433).
[0064] An observation that can be made at this point is that any decoder technology, except for syntax analysis / entropy decoding that exists within the decoder, may need to exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operation. The description of encoder technology can be omitted since they can be the reverse of the decoder technology described comprehensively. Only in certain areas is a more detailed description required and is provided below.
[0065] As part of the operation, the source encoder (430) may perform motion compensation prediction encoding that predictively encodes an input frame by referring to one or more previously encoded frames from a video sequence designated as a "reference frame". In this way, the encoding engine (432) encodes the difference between a pixel block of the input frame and a pixel block of the reference frame that can be selected as a prediction reference to the input frame.
[0066] The local video decoder (433) can decode the encoded video data of a frame that can be specified as a reference frame based on the symbols created by the source encoder (430). The operation of the encoding engine (432) may advantageously be an irreversible process. When the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) can replicate the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) can locally store a replica of the reconstructed reference frame having common content as the reconstructed reference frame obtained by the remote video decoder (without transmission errors).
[0067] The predictor (435) can perform the prediction search of the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference picture memory (434) for specific metadata that functions as an appropriate prediction reference for the new picture, such as sample data (as a candidate reference pixel block) or motion vectors and block shapes of the reference picture. The predictor (435) can operate on a sample block-by-pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search result obtained by the predictor (435).
[0068] The controller (450) can manage the encoding operation of the video encoder (430), including, for example, setting the parameters and subgroup parameters used for encoding the video data.
[0069] The outputs of all the foregoing functional units can undergo entropy encoding in an entropy encoder (445). The entropy encoder converts the symbols generated by the various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0070] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) and be prepared for transmission via a communication channel (460) which can be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0071] The controller (450) can manage the operation of the encoder (203). During encoding, the controller (450) can assign a specific encoded picture type to each encoded picture, which can affect the encoding technique applicable to each picture. For example, a picture is often assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture).
[0072] An intra picture (I picture) can be encoded and decoded without using other frames in the sequence as a source of prediction. Some video codecs allow for different types of intra pictures, such as independent decoder refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.
[0073] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0074] A bi-directionally predicted picture (B picture) can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0075] A source picture is typically subdivided spatially into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. The blocks can be encoded predictively by referring to other (already encoded) blocks as determined by the encoding assignment applied to each picture of the block. For example, blocks of an I picture may be encoded non-predictively, or they may be encoded predictively by referring to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded non-predictively via spatial prediction or via temporal prediction by referring to one previously encoded reference picture. Blocks of a B picture can be encoded non-predictively via spatial prediction or via temporal prediction by referring to one or two previously encoded reference pictures.
[0076] The video encoder (203) may perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec.H.265. In that operation, the video encoder (203) may perform various compression operations including a predictive encoding operation that exploits the temporal and spatial redundancy of the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard being used.
[0077] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video encoder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or slices, supplementary enhancement information (SEI) messages, visual user utility information (VUI) parameter set fragments, and the like.
[0078] Before describing specific aspects of embodiments of the present disclosure in more detail, some terms referred to in the remainder of this specification are introduced below.
[0079] Hereinafter, "sub-picture" refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that may optionally be meaningfully grouped and independently encoded at a modified resolution. One or more sub-pictures can form a picture. One or more encoded sub-pictures can form an encoded picture. One or more sub-pictures can be assembled into a picture, and one or more sub-pictures can be extracted from a picture. In certain environments, one or more encoded sub-pictures can be assembled in a compressed region without transcoding to the encoded picture at the sample level. And in the same or certain other cases, one or more encoded sub-pictures can be extracted from the encoded picture in the compressed region.
[0080] "Adaptive Resolution Change" (ARC) hereinafter refers to a mechanism that enables, for example, by reference picture resampling, the change of the resolution of pictures or sub-pictures within an encoded video sequence. Hereinafter, "ARC parameters" refer to the control information required to perform adaptive resolution change, which may include, for example, filter parameters, magnification, output and / or reference picture resolution, various control flags, etc.
[0081] The above description focuses on the encoding and decoding of a single semantically independent encoded video picture. Before explaining the implications of encoding / decoding of multiple sub-pictures with independent ARC parameters and the additional complexity thereby implied, embodiments for signaling ARC parameters will be described.
[0082] Referring to FIGS. 6A - 6C, several novel exemplary embodiments for signaling ARC parameters are shown. As described in each embodiment, they have certain advantages from the perspectives of encoding efficiency, complexity, and architecture. A video encoding standard or technology can implement one or more of these embodiments and can also include embodiments known from comparative techniques for signaling ARC parameters. Embodiments of comparative techniques include the examples shown in FIGS. 5A - 5B. The novel embodiments are not mutually exclusive and may be included in a standard or technology that also includes embodiments of comparative techniques such that any one can be used based on the needs of the application, related standard technologies, or the choice of the encoder.
[0083] The class of ARC parameters can include (1) up / downsample factors separated or combined in the X and Y dimensions, or (2) up / downsample factors indicating a constant speed zoom in / out of a given number of pictures with the addition of the time dimension. Either of the above two can include the encoding or decoding of one or more syntax elements that can refer to a table containing elements. Such syntax elements may be short in length in an embodiment.
[0084] "Resolution" can refer to the resolution in the X dimension or Y dimension in units of samples, blocks, macroblocks, CUs, or any other suitable granularity, for the input picture, output picture, reference picture, encoded picture, combined or individual. When there are two or more resolutions (for example, for the input picture, for the reference picture, etc.), in certain cases, one set of values can be inferred from another set of values. The resolution can be gated, for example, by the use of flags. More detailed examples of resolution are provided further below.
[0085] "Warping" coordinates may be of appropriate granularity as described above, similar to those used in H.263 Annex P. H.263 Annex P defines one efficient way to encode such warping coordinates, but other potentially more efficient ways are also conceivable. For example, the variable-length reversible "Huffman"-style encoding of the warping coordinates of Annex P can be replaced with a binary encoding of appropriate length, and the length of the binary codeword can be derived, for example, from the maximum picture size, optionally multiplied by a specific factor and offset by a specific value to allow for "warping" outside the boundaries of the maximum picture size.
[0086] Regarding upsampling filter parameters or downsampling filter parameters, in the simplest case, there may only be a single filter for upsampling and / or downsampling. However, in certain cases, it may be advantageous to allow for more flexibility in filter design that can be achieved by signaling the filter parameters. Such parameters may be selected via an index within a list of possible filter designs, the filter may be fully specified (for example, via a list of filter coefficients using appropriate entropy coding techniques), and / or the filter may be implicitly selected via the up / downsampling ratio signaled according to any of the mechanisms described above.
[0087] Hereinafter, a case of encoding a finite set of up / down sampling coefficients (the same coefficients used in both the X dimension and the Y dimension) represented by symbolic words will be described as an example. The symbolic words can advantageously be variable length encoded, for example, by using an Ext-Golomb code common to specific syntax elements in video encoding specifications such as H.264 and H.265. One suitable mapping of values to the up / down sampling coefficients can follow, for example, Table 1 below.
[0088] [Table 1]
[0089] Depending on the needs of the application and the capabilities of the upscaling and downscaling mechanisms available in the video compression technology or standard, many similar mappings can be devised. The table can be extended to more values. The values can also be represented, for example, by an entropy coding mechanism other than the Ext-Golomb code (e.g., using binary coding) that can have certain advantages when the resampling coefficients are of interest outside the video processing engine (first and foremost the decoder and encoder) itself. Note that in the most common case where resolution change is not required, a short (e.g., only a single bit as shown in the second row of Table 1) Ext-Golomb code can be selected that has an advantage in coding efficiency over using a binary code in the most common case.
[0090] The number of entries in the table, as well as their semantics, can be fully or partially configurable. For example, the basic outline of the table may be transmitted in a "high" parameter set such as a sequence or decoder parameter set. Alternatively or additionally, one or more such tables may be defined in the video encoding technology or standard and may be selected, for example, via a decoder or sequence parameter set.
[0091] Regarding how the upsampling / downsampling coefficients (ARC information) encoded as described above can be included in video encoding technologies or standard syntax, an explanation is provided below. Similar considerations may apply to one or several codewords that control the up / downsampling filter. Below, an explanation is also provided regarding cases where a relatively large amount of data is required for the filter or other data structures.
[0092] Referring to FIG. 5A, H.263 Annex P includes ARC information (502) in the form of four warping coordinates within the picture header (501), specifically within the H.263 PLUSPTYPE (503) header extension. Such a design can be sensible when (a) there is an available picture header and (b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be very high, and since the picture header can be of a temporary nature, the scaling coefficients may not be related to the picture boundaries.
[0093] Referring to FIG. 5B, JVCET-M 135-v1 includes ARC reference information (505) (index) located within the picture parameter set (504) that indexes a table (506) including the target resolution located within the sequence parameter set (507). The arrangement of possible solutions within the table (506) within the sequence parameter set (507) can be justified by using the SPS (507) as an interoperability negotiation point during functional interchange. The resolution can vary within the limits set by the values within the table (506) for each picture by referring to the appropriate picture parameter set (504).
[0094] Referring to FIGS. 6A-6C, the following embodiments of the present disclosure can communicate ARC information within a video bitstream, for example, to a decoder of the present disclosure. Each of these embodiments has certain advantages over the above-described comparative techniques. The embodiments may coexist within the same video coding technology or standard.
[0095] In the embodiment referring to FIG. 6A, ARC information (509), such as a resampling (zoom) factor, may be present in a header (508), such as a slice header, GOB header, tile header, or tile group header. As an example, FIG. 6A shows the header (508) as a tile group header. Such a configuration may be appropriate when the ARC information is small, for example, as shown in Table 1, a single variable length ue(v) or a few-bit fixed length codeword. Having the ARC information directly within the tile group header has the additional advantage that the ARC information may be applicable to a sub-picture represented by a tile group corresponding to the tile group header, rather than the entire picture. In addition, even when the video compression technology or standard uses only full-picture adaptive resolution changes (as opposed to, for example, tile group-based adaptive resolution changes), putting the ARC information in the tile group header (for example, in an H.263-style picture header) has certain advantages from the perspective of error resilience. In the above description, the case where the ARC information (509) is present in the tile group header has been described, but it goes without saying that the above description is equally applicable when the ARC information (509) is present in, for example, a slice header, GOB header, or tile header.
[0096] In the same or another embodiment with reference to FIG. 6B, the ARC information (512) itself may be present in a suitable parameter set (511), such as, for example, a picture parameter set, a header parameter set, a tile parameter set, an adaptive parameter set, etc. As an example, FIG. 6B shows the parameter set (511) as an adaptive parameter set (APS). The scope of the parameter set may advantageously be below the picture. For example, the scope of the parameter set may be a tile group. The use of the ARC information (512) may be implicit by the activation of the relevant parameter set. For example, if the video coding technology or standard contemplates only picture-based ARC, the picture parameter set or equivalent may be suitable as the relevant parameter set.
[0097] In the same or another embodiment with reference to FIG. 6C, the ARC reference information (513) may be present in a tile group header (514) or a similar data structure. The ARC reference information (513) can refer to a subset of the ARC information (515) available in a parameter set (516) having a scope that exceeds a single picture. For example, the parameter set (516) may be a sequence parameter set (SPS) or a decoder parameter set (DPS).
[0098] The implicit activation of an additional level of PPS from the tile group header, PPS, or SPS used in JVET-M 0135-v1 may not be necessary, as the picture parameter set can be used for function negotiation or announcement, similar to the sequence parameter set. However, if the ARC information should be applicable to sub-pictures, e.g., also represented by a tile group, a parameter set (e.g., an adaptive parameter set or a header parameter set) with an activation range limited to the tile group may be a better choice. Also, if the ARC information is of a size that cannot be ignored, e.g., contains filter control information such as a large number of filter coefficients, the parameter may be a better choice than directly using the header from the perspective of coding efficiency. Because these settings may be reusable by future pictures or sub-pictures by referring to the same parameter set.
[0099] When using a sequence parameter set or another higher parameter set with a scope spanning multiple pictures, certain considerations may apply.
[0100] (1) In some cases, the parameter set (516) for storing the ARC information (515) in a table can be the sequence parameter set, but in other cases, it can be advantageous to be the decoder parameter set. The decoder parameter set can have an activation range for a plurality of CVSs, i.e., the encoded video stream, i.e., all the encoded video bits from session start to session end. Such a range may be more appropriate as the possible ARC factor may be a decoder function likely implemented in hardware, and the hardware function tends not to change with the CVS (at least in some entertainment systems, the group of pictures has a length of 1 / 2 or less). Nevertheless, some embodiments can include the ARC information table in the sequence parameter set described herein, especially in relation to the following point (2).
[0101] (2) The ARC reference information (513) may preferably be placed directly in the header (514) (e.g., picture / slice style / GOB / tile group header; hereinafter, tile group header) rather than within the picture parameter set, such as JVCET-M 0135-v1. The reason is that when the encoder wants to change a single value within the picture parameter set, such as the ARC reference information, the encoder may have to create a new PPS and refer to that new PPS. Although only the ARC reference information changes, if other information, such as quantization matrix information within the PPS, remains, such information may be of a significant size and may need to be resent to complete the new PPS. Since the ARC reference information may be a single codeword, such as an index to the ARC information table, which is the only value that changes, it is cumbersome and wasteful to resent all the quantization matrix information, for example. Therefore, placing the ARC reference information directly in the header (e.g., header (514)), as proposed in JVET-M 0135-v1, can avoid indirection through the PPS and can be quite good from the perspective of coding efficiency. Also, putting the ARC reference information in the PPS has the further drawback that since the scope of picture parameter set activation is the picture, the ARC information referred to by the ARC reference information needs to be applied to the entire picture rather than a sub-picture.
[0102] In the same or another embodiment, the signaling of the ARC parameters can follow the detailed example as outlined in FIGS. 7A - 7B. FIGS. 7A - 7B show syntax diagrams. The notation of such syntax diagrams generally follows C-style programming. The thick lines indicate the syntax elements present in the bitstream, and the non-thick lines often indicate control flow or variable settings.
[0103] As an example of the syntax structure of a header applicable to a (presumably rectangular) portion of a picture, a tile group header (600) may conditionally include a variable-length Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold). The presence of this syntax element in the tile group header (600) can be gated by the use of an adaptive resolution (603). Here, the value of the adaptive resolution flag is not shown in bold, which means that the flag exists in the bitstream at the point where it occurs within the syntax diagram. Whether the adaptive resolution is used for this picture or a part thereof can be signaled in any high-level syntax structure, either internal or external to the bitstream. In the example shown in FIGS. 7A-7B, the adaptive resolution is signaled in the sequence parameter set (610) as outlined below.
[0104] FIG. 7B shows an excerpt of the sequence parameter set (610). The first syntax element shown is the adaptive_pic_resolution_change_flag (611). When true, such a flag can indicate the use of an adaptive resolution, which may require specific control information. In this example, such control information is conditionally present based on the value of the flag based on the if() statement (612) within the sequence parameter set (610) and the tile group header (600).
[0105] When adaptive resolution is being used, in this example, the encoding is the output resolution (613) in sample units. The output resolution (613) in this exemplary embodiment refers to both syntax elements output_pic_width_in_luma_samples and output_pic_height_in_luma_samples that can together define the resolution of the output picture. In other places in the video encoding technology or standard, specific restrictions for any value can be defined. For example, the level definition can limit the total number of output samples that can be the product of the values of the above two syntax elements. Also, a specific video encoding technology or standard, or an external technology or standard such as a system standard for example, can limit the numbering range (for example, one or both dimensions must be divisible by a power of 2), or the aspect ratio (for example, the width and height must be in a relationship such as 4:3 or 16:9). Such restrictions can be introduced to facilitate hardware implementation or for other reasons.
[0106] For certain applications, it may be desirable for the encoder to instruct the decoder to use a specific reference picture size rather than implicitly assuming a size that is the output picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture dimensions (615) (again, the numbers refer to both width and height in the exemplary embodiment).
[0107] FIG. 7B further shows a table of possible decoded picture widths and heights. Such a table can be represented, for example, by a table indication (616) (e.g., the syntax element num_dec_pic_size_in_luma_samples_minus1). The “minus1” of the syntax element can refer to the interpretation of the value of that syntax element. For example, if the encoded value of the syntax element is 0, there is one table entry. If the encoded value is 5, there are six table entries. For each “line” in the table, the decoded picture width and height are then included in the syntax as table entries (617).
[0108] The presented table entries (617) can be indexed using the syntax element dec_pic_size_idx (602) within the tile group header (600), thereby allowing different decoded sizes, actually zoom factors, for each tile group.
[0109] Some video encoding techniques or standards, such as VP9, support spatial scalability by performing a certain form of reference picture resampling (which can be signaled quite differently from the embodiments of the present disclosure) in combination with temporal scalability. In particular, certain reference pictures can be upsampled to a higher resolution using ARC-style techniques and can form the basis of a spatial enhancement layer. Such upsampled pictures can be refined using normal prediction mechanisms at high resolution to add details.
[0110] Embodiments of the present disclosure can be used in such an environment. In some cases, in the same or different embodiments, values within the Network Abstraction Layer (NAL) unit header, such as the time ID field, can be used to indicate not only the temporal layer but also the spatial layer. Doing so has certain advantages in a particular system design. For example, existing Selected Forwarding Units (SFUs) created and optimized for temporal layer selective forwarding based on the NAL unit header time ID value can be used without modification for a scalable environment. To enable this, embodiments of the present disclosure can include a mapping between the encoded picture size and the temporal layer indicated by the time ID field within the NAL unit header.
[0111] In some video coding techniques, an Access Unit (AU) can refer to encoded pictures, slices, tiles, NAL units, etc., that are incorporated into each picture / slice / tile / NAL unit bitstream at a given time instance. Such temporal examples can be presentation times.
[0112] In High Efficiency Video Coding (HEVC) and certain other video coding techniques, a picture order count (POC) value can be used to indicate a reference picture selected from among a plurality of reference pictures stored in a decoded picture buffer (DPB). When an access unit (AU) includes one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU can have the same POC value, from which it can be derived that they are created from content of the same composition time. In other words, in a scenario where two pictures / slices / tiles carry the same given POC value, it can be determined that the two pictures / slices / tiles belong to the same AU and have the same composition time. Conversely, two pictures / titles / slices having different POC values can indicate that those pictures / slices / tiles belong to different AUs and have different composition times.
[0113] In one embodiment of the present disclosure, the rigid relationship described above can be relaxed in that an access unit can include pictures, slices, or tiles having different POC values. By allowing different POC values within an AU, it becomes possible to use the POC value to identify potentially independently decodable pictures / slices / tiles having the same presentation time. Thus, embodiments of the present disclosure can enable support for multiple scalable layers without changing reference picture selection signaling (e.g., reference picture set signaling or reference picture list signaling), as described in more detail below.
[0114] In one embodiment, it is still desirable to be able to identify the AU to which a picture / slice / tile belongs, based only on the POC value, for other pictures / slices / tiles having different POC values. This can be achieved in the embodiments described below.
[0115] In the same or other embodiments, the Access Unit Count (AUC) may be signaled in a high-level syntax structure such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter. The value of the AUC can be used to identify which NAL unit, picture, slice, or tile belongs to a given AU. The value of the AUC may correspond to a separate composition time instance. The AUC value may be equal to a multiple of the POC value. The AUC value can be calculated by dividing the POC value by an integer value. In some cases, the division operation may impose a certain burden on the decoder implementation. In such cases, a small limitation on the numbering space of the AUC value can enable replacement of the division operation by a shift operation performed by embodiments of the present disclosure. For example, the AUC value may be equal to the Most Significant Bit (MSB) value of the POC value range.
[0116] In the same embodiment, the value of the POC cycle per AU (e.g., the syntax element poc_cycle_au) may be signaled in a high-level syntax structure such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter. The poc_cycle_au syntax element can indicate how many different consecutive POC values can be associated with the same AU. For example, if the value of poc_cycle_au is equal to 4, pictures, slices, or tiles with POC values equal to 0 - 3 are associated with an AU having an AUC value equal to 0, and pictures, slices, or tiles with POC values equal to 4 - 7 are associated with an AU having an AUC value equal to 1. Thus, the value of the AUC can be inferred by embodiments of the present disclosure by dividing the POC value by the value of poc_cycle_au.
[0117] In the same or another embodiment, the value of poc_cycle_au may be derived from information identifying the number of spatial or SNR layers in the encoded video sequence, for example, located in the video parameter set (VPS). Such a relationship will be briefly described below. The above-described derivation can save several bits in the VPS and thus improve the encoding efficiency. However, it may be advantageous to explicitly encode poc_cycle_au in a hierarchical high-level syntax structure appropriate below the video parameter set to minimize poc_cycle_au for a given small portion of the bitstream such as a picture. This optimization can save more bits than can be saved through the above-described derivation process because the POC value (and / or the value of the syntax element indirectly referring to the POC) can be encoded in a low-level syntax structure.
[0118] In the same or another embodiment, FIG. 9A shows an example of a syntax table for signaling the syntax element vps_poc_cycle_au(632) in the VPS(630) or SPS that indicates poc_cycle_au used for all pictures / slices within the encoded video sequence, and FIG. 9B shows an example of a syntax table for signaling the syntax element slice_poc_cycle_au(642) that indicates the poc_cycle_au of the current slice within the slice header(640). When the POC value increases uniformly for each AU, set vps_contant_poc_cycle_per_au(634) in the VPS(630) to 1 and signal vps_poc_cycle_au(632) in the VPS(630). In this case, slice_poc_cycle_au(642) is not explicitly signaled, and the value of AUC per AU is calculated by dividing the POC value by vps_poc_cycle_au(632). When the POC value does not increase uniformly for each AU, vps_contant_poc_cycle_per_au(634) in the VPS(630) is set to 0. In this case, vps_access_unit_cnt is not signaled, and slice_access_unit_cnt is signaled in the slice header for each slice or picture. Each slice or picture may have a different value of slice_access_unit_cnt. The value of AUC per AU is calculated by dividing the POC value by slice_poc_cycle_au(642).
[0119] FIG. 10 shows a block diagram for explaining a related work flow of an embodiment. For example, a decoder (or an encoder) analyzes the VPS / SPS to identify whether the POC cycle per AU is constant (652). Subsequently, the decoder (or the encoder) makes a determination based on whether the POC cycle per AU is constant within the encoded video sequence (654). That is, if the POC cycle per AU is constant, the decoder (or the encoder) calculates the access unit count value from the sequence level poc_cycle_au value and the POC value (656). Alternatively, if the POC cycle per AU is not constant, the decoder (or the encoder) calculates the access unit count value from the picture level poc_cycle_au value and the POC value (658). In either case, the decoder (or the encoder) can then repeat the process, for example, by analyzing the VPS / SPS and identifying whether the POC cycle per AU is constant (662).
[0120] In the same or other embodiments, even if the POC values of pictures, slices, or tiles are different, the pictures, slices, or tiles corresponding to AUs having the same AUC value can be associated with the same decoding or output time instance. Therefore, all or a subset of the pictures, slices, or tiles associated with the same AU can be decoded in parallel and output simultaneously without the dependency between analysis / decoding across the pictures, slices, or tiles within the same AU.
[0121] In the same or other embodiments, even if the POC values of pictures, slices, or tiles are different, the pictures, slices, or tiles corresponding to AUs having the same AUC value can be associated with the same composition / display time instance. When the composition time is included in a container format, even if the pictures correspond to different AUs, if the pictures have the same composition time, those pictures can be displayed at the same time instance.
[0122] In the same or other embodiments, each picture, slice, or tile can have the same time identifier (e.g., syntax element temporal_id) within the same AU. All or a subset of the pictures, slices, or tiles corresponding to a time instance can be associated with the same time sub-layer. In the same or other embodiments, each picture, slice, or tile can have the same or different spatial layer ids (e.g., syntax element layer_id) within the same AU. All or a subset of the pictures, slices, or tiles corresponding to a time instance can be associated with the same or different spatial layers.
[0123] Figure 8 shows an example of a video sequence structure (680) having a combination of temporal_id, layer_id, and POC values with adaptive resolution change and AUC values. In this example, a picture, slice, or tile within the first AU with AUC = 0 can have temporal_id = 0 and layer_id = 0 or 1, while a picture, slice, or tile within the second AU with AUC = 1 can have temporal_id = 1 and layer_id = 0 or 1 respectively. Regardless of the values of temporal_id and layer_id, the value of POC increases by 1 for each picture. In this example, the value of poc_cycle_au can be equal to 2. In one embodiment, the value of poc_cycle_au may be set equal to the number of (spatial scalability) layers. In this example, the value of POC increases by only 2 and the value of AUC increases by only 1. As an example, Figure 8 shows an I slice (681) having POC 0, TID 0, and LID 0 and a B slice (682) having POC 1, TID 0, and LID 1 within the first AU (AUC = 0). Within the second AU (AUC = 1), Figure 8 shows a B slice (683) having POC 2, TID 1, and LID 0 and a B slice (684) having POC 3, TID 1, and LID 1. Within the third AU (AUC = 3), Figure 8 shows a B slice (685) having POC 4, TID 0, and LID 0 and a B slice (686) having POC 5, TID 0, and LID 1.
[0124] In the above-described embodiments, all or a subset of the inter-picture or inter-layer prediction structures and reference picture indications can be supported by using the existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC. In the RPS or RPL, the selected reference picture is indicated by signaling the value of the picture order count (POC) or the delta value of the POC between the current picture and the selected reference picture. In embodiments of the present disclosure, the RPS and RPL can be used to indicate the inter-picture or inter-layer prediction structure without changing the signaling, but there are the following limitations. When the value of the temporal_id of the reference picture is greater than the value of the temporal_id of the current picture, the current picture may not use the reference picture for motion compensation or other prediction. When the value of the layer_id of the reference picture is greater than the value of the layer_id of the current picture, the current picture may not use the reference picture for motion compensation or other prediction.
[0125] In the same and other embodiments, the scaling of motion vectors based on the POC difference for temporal motion vector prediction can be disabled across multiple pictures within an access unit. Thus, each picture may have a different POC value within the access unit, but since reference pictures having different POCs within the same AU can be considered as reference pictures having the same temporal instance, the motion vectors are not scaled and are not used for temporal motion vector prediction within the access unit. Thus, in an embodiment, the motion vector scaling function can return 1 when the reference picture belongs to the AU associated with the current picture.
[0126] In the same and other embodiments, if the spatial resolution of the reference picture is different from that of the current picture, the scaling of the motion vectors based on the POC difference for temporal motion vector prediction can be optionally disabled over multiple pictures. When the scaling of the motion vectors is possible, the motion vectors can be scaled based on both the POC difference and the spatial resolution ratio between the current picture and the reference picture.
[0127] In the same or another embodiment, the motion vectors may be scaled based on the AUC difference instead of the POC difference for temporal motion vector prediction, especially when poc_cycle_au has a non-uniform value (when vps_contant_poc_cycle_per_au == 0). Otherwise (when vps_contant_poc_cycle_per_au == 1), the scaling of the motion vectors based on the AUC difference can be the same as the scaling of the motion vectors based on the POC difference.
[0128] In the same or another embodiment, when the motion vectors are scaled based on the AUC difference, the reference motion vectors within the same AU (having the same AUC value) as the current picture are not scaled based on the AUC difference and are used for motion vector prediction with or without scaling based on the spatial resolution ratio between the current picture and the reference picture.
[0129] In the same and other embodiments, the AUC value is used to identify the boundaries of the AU and is used for virtual reference decoder (HRD) operations that require input and output timings with AU granularity. In most cases, the decoded picture with the top layer within the AU can be output for display. The AUC value and the layer_id value can be used to identify the output picture.
[0130] In one embodiment, a picture can include one or more sub - pictures. Each sub - picture can cover a local or entire area of the picture. The area supported by a sub - picture may or may not overlap with the area supported by another sub - picture. The area composed of one or more sub - pictures may or may not cover the entire area of the picture. When the picture consists of sub - pictures, the area supported by the sub - pictures can be the same as the area supported by the picture.
[0131] In the same embodiment, the sub - pictures may be encoded by an encoding method similar to the encoding method used for the encoded picture. The sub - pictures can be encoded independently or can be encoded depending on another sub - picture or the encoded picture. The sub - pictures may or may not have an analytical dependency from another sub - picture or the encoded picture.
[0132] In the same embodiment, the encoded sub - pictures may be included in one or more layers. The encoded sub - pictures within a layer can have different spatial resolutions. The original sub - pictures can be spatially resampled (upsampled or downsampled), encoded with different spatial resolution parameters, and included in the bit - stream corresponding to the layer.
[0133] In the same or another embodiment, while a sub - picture having (W, H) can be included in the encoded bit - stream corresponding to layer 0, a sub - picture upsampled (or downsampled) from the original sub - picture having a spatial resolution of (W*S w,k ,H*S h,k ) can be encoded and included in the encoded bit - stream corresponding to layer k, where S w,k , S h,k indicates the resampling ratios in the horizontal and vertical directions. S w,k , S h,kIf the value of w,k S h,k is greater than 1, resampling is equal to upsampling. On the other hand, if the value of
[0134] S is less than 1, resampling is equal to downsampling. i,n In the same or another embodiment, the encoded sub - pictures within a layer may have a visual quality different from the visual quality of the encoded sub - pictures within another layer, whether in the same sub - picture or a different sub - picture. For example, sub - picture i within layer n is encoded with quantization parameter Q j,m and sub - picture j within layer m is encoded with quantization parameter Q
[0135] In the same or another embodiment, the encoded sub - pictures within a layer may be independently decodable without a syntax analysis or decoding dependency from the encoded sub - pictures within another layer of the same local region. A sub - picture layer that can be independently decodable without referring to another sub - picture layer of the same local region is an independent sub - picture layer. The encoded sub - pictures within an independent sub - picture layer may or may not have a decoding or syntax analysis dependency from previously encoded sub - pictures within the same sub - picture layer, but the encoded sub - pictures may not have any dependency from the encoded pictures within another sub - picture layer.
[0136] In the same or another embodiment, the encoded sub - pictures within a layer may be dependently decodable with any syntax analysis or decoding dependency from the encoded sub - pictures within another layer of the same local region. A sub - picture layer that can be dependently decodable by referring to another sub - picture layer of the same local region is a dependent sub - picture layer. The encoded sub - pictures within a dependent sub - picture can refer to the encoded sub - pictures belonging to the same sub - picture, the previously encoded sub - pictures within the same sub - picture layer, or both reference sub - pictures.
[0137] In the same or another embodiment, the encoded sub-picture includes one or more independent sub-picture layers and one or more dependent sub-picture layers. However, there may be at least one independent sub-picture layer for the encoded sub-picture. The independent sub-picture layer may have a value of a layer identifier (e.g., syntax element layer_id) that may exist in the NAL unit header or another high-level syntax structure equal to 0. The sub-picture layer with layer_id equal to 0 may be a basic sub-picture layer.
[0138] In the same or another embodiment, a picture can include one or more foreground sub-pictures and one background sub-picture. The area supported by the background sub-picture may be equal to the area of the picture. The area supported by the foreground sub-picture may overlap with the area supported by the background sub-picture. The background sub-picture may be a basic sub-picture layer, and the foreground sub-picture may be a non-base (extended) sub-picture layer. One or more non-base sub-picture layers can refer to the same base layer for decoding. Each non-base sub-picture layer with layer_id equal to a can refer to a non-base sub-picture layer with layer_id equal to b, where a is greater than b.
[0139] In the same or another embodiment, a picture can include one or more foreground sub-pictures regardless of the presence of a background sub-picture. Each sub-picture can have its own base sub-picture layer and one or more non-base (extended) layers. Each basic sub-picture layer can be referred to by one or more non-basic sub-picture layers. Each non-base sub-picture layer with layer_id equal to a can refer to a non-base sub-picture layer with layer_id equal to b, where a is greater than b.
[0140] In the same or another embodiment, a picture can include one or more foreground sub-pictures, regardless of the presence or absence of a background sub-picture. Each encoded sub-picture within a (base or non-base) sub-picture layer can be referenced by one or more non-base layer sub-pictures belonging to the same sub-picture and one or more non-base layer sub-pictures not belonging to the same sub-picture.
[0141] In the same or another embodiment, a picture can include one or more foreground sub-pictures, regardless of the presence or absence of a background sub-picture. The sub-pictures within layer a can be further divided into a plurality of sub-pictures within the same layer. One or more encoded sub-pictures within layer b can reference the divided sub-pictures within layer a.
[0142] In the same or another embodiment, an encoded video sequence (CVS) can be a group of encoded pictures. The CVS can include one or more coded sub-picture sequences (CSPS), and the CSPS can be a group of encoded sub-pictures covering the same local region of a picture. The CSPS can have a temporal resolution the same as or different from that of the encoded video sequence.
[0143] In the same or another embodiment, the CSPS can be encoded and can be included in one or more layers. The CSPS can include or consist of one or more CSPS layers. Decoding one or more CSPS layers corresponding to the CSPS can reconstruct a sequence of sub-pictures corresponding to the same local region.
[0144] In the same or another embodiment, the number of CSPS layers corresponding to a CSPS can be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0145] In the same or another embodiment, the CSPS layer may have a different temporal resolution (e.g., frame rate) than another CSPS layer. The original (uncompressed) sub-picture sequence may be temporally resampled (upsampled or downsampled), encoded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.
[0146] In the same or another embodiment, a sub-picture sequence having a frame rate F may be encoded and included in the encoded bitstream corresponding to layer 0, and the temporally upsampled (or downsampled) sub-picture sequence from the original sub-picture sequence having F*S t,k may be encoded and included in the encoded bitstream corresponding to layer k, where S t,k represents the temporal sampling ratio of layer k. When the value of S t,k is greater than 1, the temporal resampling process is equivalent to frame rate up-conversion. On the other hand, when the value of S t,k is less than 1, the temporal resampling process is equivalent to frame rate down-conversion.
[0147] In the same or another embodiment, when a sub-picture having CSPS layer a is referenced by a sub-picture having CSPS layer b for motion compensation or any inter-layer prediction, if the spatial resolution of CSPS layer a is different from the spatial resolution of CSPS layer b, the decoded pixels in CSPS layer a are resampled and used for the reference. The resampling process may require up-sampling filtering or down-sampling filtering.
[0148] FIG. 11 shows an exemplary video stream including a background video CSPS with layer_id equal to 0 and a plurality of foreground CSPS layers. The encoded sub-pictures may comprise one or more extended CSPS layers (704), while the background regions that do not belong to any of the foreground CSPS layers may comprise a base layer (702). The base layer (702) can include both background and foreground regions, while the extended CSPS layer (704) includes foreground regions. The extended CSPS layer (704) may have better visual quality than the base layer (702) in the same region. The extended CSPS layer (704) can refer to the reconstructed pixels and the motion vectors of the base layer (702) corresponding to the same region.
[0149] In the same or another embodiment, the video bitstream corresponding to the base layer (702) is included in a track, and the CSPS layers (704) corresponding to each sub-picture are included in separate tracks within the video file.
[0150] In the same or another embodiment, the video bitstream corresponding to the base layer (702) is included in a track, and the CSPS layers (704) having the same layer_id are included in separate tracks. In this example, the track corresponding to layer k includes only the CSPS layer (704) corresponding to layer k.
[0151] In the same or another embodiment, each CSPS layer (704) of each sub-picture is stored in a separate track. Each track may or may not have a syntax parsing or decoding dependency from one or more other tracks.
[0152] In the same or another embodiment, each track can include the bitstream corresponding to the CSPS layers (704) from layer i to layer j of all or a subset of the sub-pictures, where 0 < i <= j <= k and k is the topmost layer of the CSPS.
[0153] In the same or another embodiment, the picture includes or consists of one or more associated media data, such as a depth map, an alpha map, 3D geometry data, an occupancy map, etc. Such associated time-limited media data can be divided into one or more data sub-streams, each corresponding to one sub-picture.
[0154] In the same or another embodiment, FIG. 12 shows an example of a videoconference based on a multi-layer sub-picture method. The video stream includes one base layer video bitstream corresponding to the background picture and one or more enhancement layer video bitstreams corresponding to the foreground sub-pictures. Each enhancement layer video bitstream can correspond to a CSPS layer. On the display, a picture corresponding to the base layer (712) is initially displayed. The base layer (712) can include pictures of one or more users within the picture (PIP). When a specific user is selected under the control of the client, the enhanced CSPS layer (714) corresponding to the selected user is encoded and displayed with enhanced quality or spatial resolution.
[0155] FIG. 13 shows a diagram for the operation of an embodiment. In the embodiment, the decoder can decode a video bitstream including multiple layers, such as a certain base layer and one or more enhancement CSPS layers (722). Subsequently, the decoder can identify the background area and one or more foreground sub-pictures (724) and make a decision on whether a specific sub-picture area is selected (726). For example, if a specific sub-picture area corresponding to the user's PIP is selected (yes), the decoder can decode and display the enhanced sub-picture corresponding to the selected user (728). For example, the decoder can decode and display an image corresponding to the enhanced CSPS layer (714). If a specific sub-picture area is not selected (no), the decoder can decode and display the background area (730). For example, the decoder may decode and display an image corresponding to the base layer (712).
[0156] In the same or another embodiment, a network intermediate box (such as a router) can select a subset of the layers to send to the user according to its bandwidth. Picture / sub-picture composition can be used for bandwidth adaptation. For example, when the user has no bandwidth, the router selects layer strips or some sub-pictures due to their importance or based on the settings used. In one embodiment, such processing may be performed dynamically to adapt to the bandwidth.
[0157] FIG. 14 shows an exemplary use case of 360 video. When a spherical 360 picture (742) is projected onto a planar picture, the projected spherical 360 picture (742) can be divided into a plurality of sub-pictures (745) as a base layer (744). The enhancement layer (746) of a specific sub-picture among the sub-pictures (745) can be encoded and sent to the client. The decoder can decode both the base layer (744) including all the sub-pictures (745) and the enhancement layer (746) of one selected sub-picture among the sub-pictures (745). When the current viewport is the same as the selected one among the sub-pictures (745), the displayed picture can have higher quality using the decoded sub-picture (745) with the enhancement layer (746). Otherwise, the decoded picture with the base layer (744) can be displayed with lower quality.
[0158] In the same or another embodiment, any layout information for display may exist in the file as auxiliary information (such as SEI messages or metadata). One or more decoded sub-pictures can be rearranged and displayed according to the signaled layout information. The layout information may be signaled by a streaming server or a broadcasting station, may be regenerated by a network entity or a cloud server, or may be determined by the user's customized settings.
[0159] In one embodiment, when the input picture is divided into one or more (rectangular) sub-regions, each sub-region can be encoded as an independent layer. Each independent layer corresponding to a local region can have a unique layer_id value. For each independent layer, sub-picture size and position information can be signaled. For example, the picture size (width, height) and the offset information (x_offset, y_offset) of the upper left corner can be signaled. FIG. 15A shows an example of the layout of the divided sub-pictures (752), FIG. 15B shows an example of the corresponding sub-picture size and position information for one of the sub-pictures (752), and FIG. 16 shows the corresponding picture prediction structure. The layout information including the sub-picture size(s) and sub-picture position(s) can be signaled in a high-level syntax structure such as a parameter set(s), a slice or tile group header, or an SEI message.
[0160] In the same embodiment, each sub-picture corresponding to an independent layer can have its unique POC value within the AU. When indicating the reference pictures among the pictures stored in the DPB using the syntax elements of the RPS or RPL structure, the POC value of each sub-picture corresponding to the layer may be used.
[0161] In the same or another embodiment, in order to indicate the (inter-layer) prediction structure, the layer_id may not be used, and the POC (delta) value may be used.
[0162] In the same embodiment, a sub-picture having a POC value equal to N corresponding to a layer (or local region) may or may not be used as a reference picture for a sub-picture having a POC value equal to K+N corresponding to the same layer (or the same local region) for motion compensation prediction. In most cases, the value of the number K may be equal to the maximum number of (independent) layers, which may be the same as the number of sub-regions.
[0163] In the same or another embodiment, FIGS. 17-18 show the extended cases of FIGS. 15A-15B and FIG. 16. When the input picture is divided into a plurality of (e.g., 4) sub-regions, each local region can be encoded with one or more layers. In this case, the number of independent layers may be equal to the number of sub-regions, or one or more layers may correspond to the sub-regions. Thus, each sub-region can be encoded using one or more independent layers and zero or more dependent layers.
[0164] In the same embodiment, referring to FIG. 17, the input picture can be divided into four sub-regions including an upper left sub-region (762), an upper right sub-region (763), a lower left sub-region (764), and a lower right sub-region (765). The upper right sub-region (763) can be encoded as two layers, layer 1 and layer 4, and the lower right sub-region (765) can be encoded as two layers, layer 3 and layer 5. In this case, layer 4 can refer to layer 1 for motion compensation prediction, and layer 5 can refer to layer 3 for motion compensation.
[0165] In the same or another embodiment, in-loop filtering (such as deblocking filtering, adaptive in-loop filtering, reshaper, bilateral filtering, or any deep learning-based filtering, etc.) across layer boundaries can be (optionally) disabled.
[0166] In the same or another embodiment, motion compensation prediction or intra-block replication across layer boundaries can be (optionally) disabled.
[0167] In the same or another embodiment, boundary padding for motion compensation prediction or in-loop filtering at the sub-picture boundary may be optionally processed. A flag indicating whether the boundary padding is processed can be signaled in a high-level syntax structure such as a parameter set (VPS, SPS, PPS, or APS), a slice or tile group header, or an SEI message.
[0168] In the same or another embodiment, the layout information of the sub-region (or sub-picture) may be signaled by the VPS or SPS. FIG. 19A shows an example of the syntax elements of the VPS (770), and FIG. 19B shows an example of the syntax elements of the SPS (780). In this example, in the VPS (770), the vps_sub_picture_dividing_flag (772) is signaled. The flag can indicate whether the input picture is divided into a plurality of sub-regions. When the value of the vps_sub_picture_dividing_flag (772) is equal to 0, the input picture in the encoded video sequence corresponding to the current VPS may not be divided into a plurality of sub-regions. In this case, the input picture size may be equal to the encoded picture size (pic_width_in_luma_samples (786), pic_height_in_luma_samples (788)) signaled by the SPS (680). When the value of the vps_sub_picture_dividing_flag (772) is equal to 1, the input picture may be divided into a plurality of sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples (774) and vps_full_pic_height_in_luma_samples (776) are signaled by the VPS (770). The values of vps_full_pic_width_in_luma_samples (774) and vps_full_pic_height_in_luma_samples (776) may be equal to the width and height of the input picture, respectively.
[0169] In the same embodiment, the values of vps_full_pic_width_in_luma_samples (774) and vps_full_pic_height_in_luma_samples (776) may not be used for decoding, and may be used for synthesis and display.
[0170] In the same embodiment, when the value of vps_sub_picture_dividing_flag(772) is equal to 1, the syntax elements pic_offset_x(782) and pic_offset_y(784) may be signaled in the SPS(780) corresponding to a specific layer(s). In this case, the coded picture size (pic_width_in_luma_samples(786), pic_height_in_luma_samples(788)) signaled in the SPS(780) may be equal to the width and height of the sub-region corresponding to the specific layer. Also, the position (pic_offset_x(782), pic_offset_y(784)) of the upper left corner of the sub-region may be signaled in the SPS(780).
[0171] In the same embodiment, the position information (pic_offset_x(782), pic_offset_y(784)) of the upper left corner of the sub-region may not be used for decoding, and may be used for synthesis and display.
[0172] In the same or another embodiment, the layout information (size and position) of all or a subset of sub-regions of the input picture, and the inter-layer dependency information can be signaled in a parameter set or an SEI message. FIG. 20 is a diagram showing an example of a syntax element indicating the layout of sub-regions, the inter-layer dependency, and the relationship between a sub-region and one or more layers. In this example, the syntax element num_sub_region(791) indicates the number of (rectangular) sub-regions within the current encoded video sequence. The syntax element num_layers(792) indicates the number of layers within the current encoded video sequence. The value of num_layers(792) may be greater than or equal to the value of num_sub_region(791). If any sub-region is encoded as a single layer, the value of num_layers(792) may be equal to the value of num_sub_region(791). When one or more sub-regions are encoded as multiple layers, the value of num_layers(792) may be greater than the value of num_sub_region(791). The syntax element direct_dependency_flag[i][j](793) indicates the dependency from the j-th layer to the i-th layer. The syntax element num_layers_for_region[i](794) indicates the number of layers associated with the i-th sub-region. The syntax element sub_region_layer_id[i][j](795) indicates the layer_id of the j-th layer associated with the i-th sub-region. The syntax elements sub_region_offset_x[i](796) and sub_region_offset_y[i](797) indicate the horizontal and vertical positions of the upper left corner of the i-th sub-region, respectively. The syntax elements sub_region_width[i](798) and sub_region_height[i](799) indicate the width and height of the i-th sub-region, respectively.
[0173] In one embodiment, one or more syntax elements that specify an output layer that is set to indicate one of a plurality of layers output regardless of the presence or absence of profile tier level information may be signaled in a high-level syntax structure (e.g., VPS, DPS, SPS, PPS, APS, or SEI message). Referring to FIG. 21, a syntax element num_output_layer_sets(804) that indicates the number of output layer sets (OLS) in an encoded video sequence that refers to a VPS may be signaled in the VPS. For each output layer set, a syntax element output_layer_flag(810) may be signaled the same number of times as the number of output layers.
[0174] In the same embodiment, a syntax element output_layer_flag(810) equal to 1 specifies that the i-th layer is output. A syntax element output_layer_flag(810) equal to 0 specifies that the i-th layer is not output.
[0175] In the same or another embodiment, one or more syntax elements that specify the profile tier level information for each output layer set may be signaled in a high-level syntax structure (e.g., VPS, DPS, SPS, PPS, APS, or SEI message). Further referring to FIG. 21, a syntax element num_profile_tier_level(806) that indicates the number of profile tier level information for each OLS in an encoded video sequence that refers to a VPS may be signaled within the VPS. For each output layer set, a set of syntax elements of the profile tier level information, or an index indicating a specific profile tier level information among the entries within the profile tier level information, may be signaled the same number of times as the number of output layers.
[0176] In the same embodiment, the syntax element profile_tier_level_idx[i][j] (812) specifies an index into a list of profile_tier_level() (808) syntax structures in the VPS for the profile_tier_level() (808) syntax structure applied to the j-th layer of the i-th OLS.
[0177] Profiles, tiers, and levels (and their corresponding information) can specify restrictions on the bitstream and thus on the capabilities required to decode the bitstream. Profiles, tiers, and levels (and their corresponding information) can also be used to indicate interoperability points between individual decoder implementations. A profile may be, for example, a subset of the overall standard bitstream syntax. Each profile (and its corresponding information) can specify a subset of the algorithmic features and restrictions that can be supported by all decoders compliant with the profile. Tiers and levels may be specified within each profile, and the level of a tier may be a specified set of constraints imposed on the values of syntax elements within the bitstream. Each level of a tier (and its corresponding information) can specify a set of restrictions on the values that the syntax elements of the present disclosure can take and / or on arithmetic combinations of values. The same set of tier and level definitions can be used in all profiles, but individual implementations can support different tiers and different levels within a tier for each supported profile. For any given profile, the level of a tier can correspond to a particular decoder processing load and memory capability. Levels specified for lower tiers can be more restrictive than levels specified for higher tiers.
[0178] In the same or another embodiment, referring to FIG. 22, the syntax elements num_profile_tier_level (806) and / or num_output_layer_sets (804) can be signaled when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0).
[0179] In the same or another embodiment, referring to FIG. 22, a syntax element vps_output_layers_mode[i] (822) indicating the mode of output layer signaling for the i-th output layer set may be present within the VPS.
[0180] In the same embodiment, a syntax element vps_output_layers_mode[i] (822) equal to 0 specifies that only the topmost layer is output in the i-th output layer set. A syntax element vps_output_layers_mode[i] (822) equal to 1 specifies that all layers are output in the i-th output layer set. A syntax element vps_output_layers_mode[i] (822) equal to 2 specifies that the layers to be output are the layers having vps_output_layer_flag[i][j] equal to 1 for which the i-th output layer is set. More values may be reserved.
[0181] In the same embodiment, the syntax element output_layer_flag[i][j] (810) may or may not be signaled depending on the value of the syntax element vps_output_layers_mode[i] (822) of the i-th output layer set.
[0182] In the same or another embodiment, referring to FIG. 22, a flag vps_ptl_signal_flag[i] (824) may exist for the i-th output layer set. Depending on the value of vps_ptl_signal_flag[i] (824), the profile layer level information of the i-th output layer set may or may not be signaled.
[0183] In the same or another embodiment, referring to FIG. 23, the number of sub-pictures max_subpics_minus1 within the current CVS may be signaled in a high-level syntax structure (e.g., VPS, DPS, SPS, PPS, APS, or SEI message).
[0184] In the same embodiment, referring to FIG. 23, when the number of sub-pictures is greater than 1 (max_subpics_minus1>0), the sub-picture identifier sub_pic_id[i] (821) for the i-th sub-picture may be signaled.
[0185] In the same or another embodiment, one or more syntax elements indicating the sub-picture identifiers belonging to each layer of each output layer set may be signaled in the VPS. Referring to FIG. 23, the identifier sub_pic_id_layer[i][j][k] (826) indicates the k-th sub-picture present in the j-th layer of the i-th output layer set. The decoder can use the information of the identifier sub_pic_id_layer[i][j][k] (826) to recognize which sub-pictures can be decoded and output for each layer of a specific output layer set.
[0186] In one embodiment, a picture header (PH) is a syntax structure that includes syntax elements applied to all slices of an encoded picture. A picture unit (PU) is a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and contain exactly one encoded picture. A PU may include a picture header (PH) and one or more video coding layer (VCL) NAL units that make up the encoded picture.
[0187] In one embodiment, the SPS (RBSP) may be available for decoding processing before being referenced by being included in at least one AU having a TemporalId equal to 0 or being provided via external means.
[0188] In one embodiment, the SPS (RBSP) may be available for decoding processing before being referenced by being included in at least one AU having a TemporalId equal to 0 in a CVS that includes one or more PPSs that reference the SPS, or by being provided via external means.
[0189] In one embodiment, the SPS (RBSP) can be made available for decoding before being referenced by one or more PPSs by being included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL units that reference the SPS NAL units within the CVS that includes one or more PPSs that reference the SPS, or by being provided via external means.
[0190] In one embodiment, the SPS (RBSP) can be made available for decoding before being referenced by one or more PPSs by being included in at least one PU having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL units that reference the SPS NAL units, or by being provided via external means.
[0191] In one embodiment, the SPS (RBSP) can be made available for decoding before being referenced by one or more PPSs by being included in at least one PU having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL units within the CVS that includes one or more PPSs that reference the SPS, or by being provided via external means or being able to be provided via external means.
[0192] In the same or another embodiment, the identifier pps_seq_parameter_set_id specifies the value of the identifier sps_seq_parameter_set_id of the referenced SPS. The value of the identifier pps_seq_parameter_set_id can be the same in all PPSs referenced by coded pictures in the coded layer video sequence (CLVS).
[0193] In the same or another embodiment, all SPS NAL units having a specific value of the identifier sps_seq_parameter_set_id within the CVS can have the same content.
[0194] In the same or another embodiment, regardless of the nuh_layer_id value, SPS NAL units can share the same value space of the identifier sps_seq_parameter_set_id.
[0195] In the same or another embodiment, the nuh_layer_id value of the SPS NAL unit may be equal to the lowest nuh_layer_id value of the PPS NAL unit that refers to the SPS NAL unit.
[0196] In one embodiment, when an SPS having a nuh_layer_id equal to m is referred to by one or more PPSs having a nuh_layer_id equal to n, the layer having a nuh_layer_id equal to m can be the same as the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to n or the layer having a nuh_layer_id equal to m.
[0197] In one embodiment, the PPS (RBSP) can be made available for decoding before being referred to by being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit within the CVS that includes one or more PHs (or coded slice NAL units) referring to the PPS, or by being provided via external means.
[0198] In one embodiment, the PPS (RBSP) can be made available for decoding before being referred to by being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit within the CVS that includes the PPS and refers to one or more PHs (or coded slice NAL units), or by being provided via external means.
[0199] In one embodiment, the PPS (RBSP) can be made available for decoding prior to being referenced by one or more PHs (or coded slice NAL units) by being included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of coded slice NAL units that reference the PPS NAL unit in the CVS and that include one or more PHs (or coded slice NAL units) that reference the PPS, or by being provided via external means.
[0200] In one embodiment, the PPS (RBSP) can be made available for decoding prior to being referenced by one or more PHs (or coded slice NAL units) by being included in at least one PU having a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the lowest nuh_layer_id value of coded slice NAL units that reference the PPS NAL unit in the CVS and that include one or more PHs (or coded slice NAL units) that reference the PPS, or by being provided via external means.
[0201] In the same or another embodiment, the identifier ph_pic_parameter_set_id within the PH specifies the value of the identifier pps_pic_parameter_set_id of the reference PPS in use. The value of pps_seq_parameter_set_id can be the same in all PPSs referenced by coded pictures in the CLVS.
[0202] In the same or another embodiment, all PPS NAL units having a particular value of the identifier pps_pic_parameter_set_id within the PU can have the same content.
[0203] In the same or another embodiment, regardless of the nuh_layer_id value, PPS NAL units can share the same value space of the identifier pps_pic_parameter_set_id.
[0204] In the same or another embodiment, the nuh_layer_id value of the PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the NAL unit that references the PPS NAL unit.
[0205] In one embodiment, when a PPS having a nuh_layer_id equal to m is referenced by one or more coded slice NAL units having a nuh_layer_id equal to n, the layer having a nuh_layer_id equal to m may be the same as the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to n or the layer having a nuh_layer_id equal to m.
[0206] In one embodiment, the PPS (RBSP) can be made available for decoding before being referenced by being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit or by being provided via external means.
[0207] In one embodiment, the PPS (RBSP) can be made available for decoding before being referenced by being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit within a CVS that includes one or more PHs (or coded slice NAL units) that reference the PPS or by being provided via external means.
[0208] In one embodiment, the PPS (RBSP) can be made available for decoding prior to being referenced by one or more PHs (or coded slice NAL units) that reference the PPS, by being included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of coded slice NAL units that reference the PPS NAL unit within the CVS, where the PPS (RBSP) includes one or more PHs (or coded slice NAL units) that reference the PPS or is provided via external means.
[0209] In one embodiment, the PPS (RBSP) can be made available for decoding prior to being referenced by one or more PHs (or coded slice NAL units) that reference the PPS, by being included in at least one PU having a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the lowest nuh_layer_id value of coded slice NAL units that reference the PPS NAL unit within the CVS, where the PPS (RBSP) includes one or more PHs (or coded slice NAL units) that reference the PPS or is provided via external means.
[0210] In the same or another embodiment, the identifier ph_pic_parameter_set_id within the PH specifies the value of the identifier pps_pic_parameter_set_id of the reference PPS in use. The value of the identifier pps_seq_parameter_set_id can be the same for all PPSs referenced by coded pictures within the CLVS.
[0211] In the same or another embodiment, all PPS NAL units having a particular value of pps_pic_parameter_set_id within a PU can have the same content.
[0212] In the same or another embodiment, regardless of the nuh_layer_id value, PPS NAL units can share the same value space of the identifier pps_pic_parameter_set_id.
[0213] In the same or another embodiment, the nuh_layer_id value of the PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the NAL unit that references the PPS NAL unit.
[0214] In one embodiment, when a PPS having an nuh_layer_id equal to m is referenced by one or more coded slice NAL units having an nuh_layer_id equal to n, the layer having an nuh_layer_id equal to m can be the same as the (direct or indirect) reference layer of the layer having an nuh_layer_id equal to n or the layer having an nuh_layer_id equal to m.
[0215] The output layer may be a layer of the output layer set to be output. The output layer set (OLS) may be a set of specified layers, and one or more layers within the set of layers are specified to be output layers. The output layer set (OLS) layer index is the index of the layer within the OLS with respect to the list of layers within the OLS.
[0216] The sublayer may be a temporally scalable layer of a temporally scalable bitstream of a sublayer that includes VCL NAL units having a specific value of the TemporalId variable and related non-VCL NAL units. The sublayer representation can be a subset of the bitstream that includes the NAL units of a particular sublayer and lower sublayers.
[0217] The VPS RBSP can be made available for decoding before being referenced by being included in at least one AU where the TemporalId is 0 or by being provided via external means. All VPS NAL units having a specific value of vps_video_parameter_set_id within the CVS can have the same content.
[0218] With reference to FIGS. 24 to 25, exemplary VPS RBSP syntax elements are described below.
[0219] The syntax element vps_video_parameter_set_id(842) provides an identifier for the VPS for reference by other syntax elements. The value of the syntax element vps_video_parameter_set_id(842) may be greater than 0.
[0220] The syntax element vps_max_layers_minus1(802)+1 specifies the maximum number of layers allowed in each CVS that refers to the VPS.
[0221] The syntax element vps_max_sublayers_minus1(846)+1 specifies the maximum number of temporal sublayers that may exist in the layers within each CVS that refers to the VPS. The value of the syntax element vps_max_sublayers_minus1(846) can be in the range of 0 to 6.
[0222] The syntax element vps_all_layers_same_num_sublayers_flag(848) equal to 1 specifies that the number of temporal sublayers is the same for all layers within each CVS that refers to the VPS. The syntax element vps_all_layers_same_num_sublayers_flag(848) equal to 0 specifies that the layers within each CVS that refers to the VPS may or may not have the same number of temporal sublayers. If it does not exist, the value of vps_all_layers_same_num_sublayers_flag(848) can be presumed to be equal to 1.
[0223] The syntax element vps_all_independent_layers_flag(850) equal to 1 specifies that all layers within the CVS are independently encoded without using inter-layer prediction. The syntax element vps_all_independent_layers_flag(850) equal to 0 specifies that one or more of the layers within the CVS can use inter-layer prediction. If not present, the value of vps_all_independent_layers_flag(850) can be inferred to be equal to 1.
[0224] The syntax element vps_layer_id[i](852) specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] can be less than vps_layer_id[n].
[0225] The syntax element vps_independent_layer_flag[i](854) equal to 1 specifies that the layer with index i does not use inter-layer prediction. The syntax element vps_independent_layer_flag[i](854) equal to 0 specifies that the layer with index i can use inter-layer prediction, and the syntax element vps_direct_ref_layer_flag[i][j] for j in the range from 0 to i - 1 exists in the VPS. If not present, the value of the syntax element vps_independent_layer_flag[i](854) can be inferred to be equal to 1.
[0226] The syntax element vps_direct_ref_layer_flag[i][j] (856) equal to 0 specifies that the layer with index j is not a direct reference layer for the layer with index i. The syntax element vps_direct_ref_layer_flag[i][j] (856) equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. For i and j in the range from 0 to vps_max_layers_minus1, if the syntax element vps_direct_ref_layer_flag[i][j] (856) does not exist, that syntax element can be inferred to be equal to 0. When the syntax element vps_independent_layer_flag[i] (854) is equal to 0, there may be at least one value of j in the range from 0 to i - 1 such that the value of the syntax element vps_direct_ref_layer_flag[i][j] (856) is equal to 1.
[0227] The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] can be derived as follows. for(i = 0; i <= vps_max_layers_minus1; i++){ for(j = 0; j <= vps_max_layers_minus1; j++){ dependencyFlag[i][j] = vps_direct_ref_layer_flag[i][j] for(k = 0; k < i; k++) if(vps_direct_ref_layer_flag[i][k] && dependencyFlag[k][j]) dependencyFlag[i][j] = 1 } LayerUsedAsRefLayerFlag[i] = 0 } for(i = 0; i <= vps_max_layers_minus1; i++){ for(j = 0, d = 0, r = 0; j <= vps_max_layers_minus1; j++){(37) if(vps_direct_ref_layer_flag[i][j]){ DirectRefLayerIdx[i][d++] = j LayerUsedAsRefLayerFlag[j] = 1 } if(dependencyFlag[i][j]) RefLayerIdx[i][r++] = j } NumDirectRefLayers[i] = d NumRefLayers[i] = r }
[0228] The variable GeneralLayerIdx[i] that specifies the layer index of the layer having the nuh_layer_id equal to vps_layer_id[i] (852) can be derived as follows. for(i = 0; i <= vps_max_layers_minus1; i++)(38) GeneralLayerIdx[vps_layer_id[i]] = i
[0229] For any two different values of i and j both within the range from 0 to vps_max_layers_minus1 (846), when dependencyFlag[i][j] is equal to 1, it can be a bitstream conformity requirement that the values of chroma_format_idc and bit_depth_minus8 applied to the i-th layer are respectively equal to the values of chroma_format_idc and bit_depth_minus8 applied to the j-th layer.
[0230] The syntax element max_tid_ref_present_flag[i] (858) equal to 1 specifies that the syntax element max_tid_il_ref_pics_plus1[i] (860) exists. The syntax element max_tid_ref_present_flag[i] (858) equal to 0 specifies that the syntax element max_tid_il_ref_pics_plus1[i] (860) does not exist.
[0231] The syntax element max_tid_il_ref_pics_plus1[i] (860) equal to 0 specifies that inter-layer prediction is not used by the non-IRAP picture of the i-th layer. The syntax element max_tid_il_ref_pics_plus1[i] (860) greater than 0 specifies that for decoding the picture of the i-th layer, pictures with a TemporalId greater than max_tid_il_ref_pics_plus1[i] - 1 are not used as inter-layer reference pictures (ILRPs). If it does not exist, the value of the syntax element max_tid_il_ref_pics_plus1[i] (860) can be assumed to be equal to 7.
[0232] The syntax element each_layer_is_an_ols_flag (862) equal to 1 specifies that each OLS contains only one layer, each layer itself within the CVS referring to the VPS is an OLS, and the single layer contained is the only output layer. The syntax element each_layer_is_an_ols_flag (862) equal to 0 specifies that an OLS can contain multiple layers. If the syntax element vps_max_layers_minus1 is equal to 0, the value of the syntax element each_layer_is_an_ols_flag (862) can be assumed to be equal to 1. Otherwise, when the syntax element vps_all_independent_layers_flag (854) is equal to 0, the value of each syntax element each_layer_is_an_ols_flag (862) can be assumed to be equal to 0.
[0233] The syntax element ols_mode_idc(864) equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i, and for each OLS, only the top layer of the OLS is output.
[0234] The syntax element ols_mode_idc(864) equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i, and for each OLS, all layers within the OLS are output.
[0235] The syntax element ols_mode_idc(864) equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled, the output layer is explicitly signaled for each OLS, and the other layers are the layers that are direct or indirect reference layers of the output layer of the OLS.
[0236] The value of the syntax element ols_mode_idc(864) may be in the range from 0 to 2. The value 3 of the syntax element ols_mode_idc(864) can be reserved for future use by ITU-T|ISO / IEC.
[0237] When the syntax element vps_all_independent_layers_flag(850) is equal to 1 and each _layer_is_an_ols_flag(862) is equal to 0, it can be inferred that the value of the syntax element ols_mode_idc(864) is equal to 2.
[0238] The syntax element num_output_layer_sets_minus1(866) + 1 specifies the total number of OLSs specified by the VPS when the syntax element ols_mode_idc(864) is equal to 2.
[0239] The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS can be derived as follows. if(vps_max_layers_minus1==0) TotalNumOlss=1 else if(each_layer_is_an_ols_flag||ols_mode_idc==0||ols_mode_idc==1) TotalNumOlss=vps_max_layers_minus1+1 else if(ols_mode_idc==2) TotalNumOlss=num_output_layer_sets_minus1+1
[0240] The syntax element ols_output_layer_flag[i][j] (868) equal to 1 specifies that, when ols_mode_idc (864) is equal to 2, the layer where nuh_layer_id is equal to vps_layer_id[j] is the output layer of the i-th OLS. The syntax element ols_output_layer_flag[i][j] (868) equal to 0 specifies that, when the syntax element ols_mode_idc (864) is equal to 2, the layer where nuh_layer_id is equal to vps_layer_id[j] is not the output layer of the i-th OLS.
[0241] The variable NumOutputLayersInOls[i] that specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] that specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] that specifies whether the k-th layer is used as an output layer in at least one OLS may be derived as follows. NumOutputLayersInOls[0]=1 OutputLayerIdInOls[0][0]=vps_layer_id[0] NumSubLayersInLayerInOLS[0][0]=vps_max_sub_layers_minus1+1 LayerUsedAsOutputLayerFlag[0]=1 for(i=1,i<=vps_max_layers_minus1;i++){ if(each_layer_is_an_ols_flag||ols_mode_idc<2) LayerUsedAsOutputLayerFlag[i]=1 else / *(!each_layer_is_an_ols_flag&&ols_mode_idc==2)* / LayerUsedAsOutputLayerFlag[i]=0 } for(i=1;i<TotalNumOlss;i++) if(each_layer_is_an_ols_flag||ols_mode_idc==0){ NumOutputLayersInOls[i]=1 OutputLayerIdInOls[i][0]=vps_layer_id[i] for(j=0;j<i&&(ols_mode_idc==0);j++) NumSubLayersInLayerInOLS[i][j]=max_tid_il_ref_pics_plus1[i] NumSubLayersInLayerInOLS[i][i]=vps_max_sub_layers_minus1+1 }else if(ols_mode_idc==1){ NumOutputLayersInOls[i]=i+1 for(j=0;j<NumOutputLayersInOls[i];j++){ OutputLayerIdInOls[i][j]=vps_layer_id[j] NumSubLayersInLayerInOLS[i][j]=vps_max_sub_layers_minus1+1 } } else if(ols_mode_idc == 2) { for(j = 0; j <= vps_max_layers_minus1; j++) { layerIncludedInOlsFlag[i][j]=0 NumSubLayersInLayerInOLS[i][j]=0 } for(k = 0, j = 0; k <= vps_max_layers_minus1; k++)(40) if(ols_output_layer_flag[i][k]) { layerIncludedInOlsFlag[i][k]=1 LayerUsedAsOutputLayerFlag[k]=1 OutputLayerIdx[i][j]=k OutputLayerIdInOls[i][j++]=vps_layer_id[k] NumSubLayersInLayerInOLS[i][j]=vps_max_sub_layers_minus1+1 } NumOutputLayersInOls[i]=j for(j = 0; j < NumOutputLayersInOls[i]; j++) { idx=OutputLayerIdx[i][j] for(k = 0; k < NumRefLayers[idx]; k++) { layerIncludedInOlsFlag[i][RefLayerIdx[idx][k]]=1 if(NumSubLayersInLayerInOLS[i][RefLayerIdx[idx][k]]< max_tid_il_ref_pics_plus1[OutputLayerIdInOls[i][j]]) NumSubLayersInLayerInOLS[i][RefLayerIdx[idx][k]]= max_tid_il_ref_pics_plus1[OutputLayerIdInOls[i][j]] } } }
[0242] For each value of i in the range from 0 to vps_max_layers_minus1, the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] may not both be equal to 0. In other words, there may not be a layer that is neither an output layer of at least one OLS nor a direct reference layer of another layer.
[0243] For each OLS, there may be at least one layer that is an output layer. That is, for any i in the range from 0 to TotalNumOlss - 1, the value of NumOutputLayersInOls[i] may be set to 1 or more.
[0244] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS can be derived as follows. NumLayersInOls[0]=1 LayerIdInOls[0][0]=vps_layer_id[0] for(i=1;i<TotalNumOlss;i++){ if(each_layer_is_an_ols_flag){ NumLayersInOls[i]=1 LayerIdInOls[i][0]=vps_layer_id[i] } else if (ols_mode_idc == 0 || ols_mode_idc == 1) { NumLayersInOls[i] = i + 1 for (j = 0; j < NumLayersInOls[i]; j++) LayerIdInOls[i][j] = vps_layer_id[j] } else if (ols_mode_idc == 2) { for (k = 0, j = 0; k <= vps_max_layers_minus1; k++) if (layerIncludedInOlsFlag[i][k]) LayerIdInOls[i][j++] = vps_layer_id[k] NumLayersInOls[i] = j } }
[0245] The variable OlsLayerIdx[i][j] that specifies the OLS layer index of the layer where nuh_layer_id is equal to LayerIdInOls[i][j] is derived as follows. for (i = 0; i < TotalNumOlss; i++) for j = 0; j < NumLayersInOls[i]; j++) OlsLayerIdx[i][LayerIdInOls[i][j]] = j
[0246] The bottom layer in each OLS may be an independent layer. That is, for each i in the range from 0 to TotalNumOlss - 1, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] may be set to 1. Each layer may be included in at least one OLS specified by the VPS. In other words, for each layer having a specific value of nuh_layer_id, nuhLayerId equal to one of vps_layer_id[k] within the range from 0 to vps_max_layers_minus1, there may exist at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1 and j is in the range from 0 to NumLayersInOls[i] - 1, such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.
[0247] In one embodiment, the decoding process can operate on the current picture (e.g., sytax element CurrPic) to set the syntax element PictureOutputFlag as follows.
[0248] If one of the following conditions is true, PictureOutputFlag is set equal to 0. (1) The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1. (2) gdr_enabled_flag is equal to 1 and the current picture is a GDR picture having a NoOutputBeforeRecoveryFlag equal to 1. (3) gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture having a NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture. (4) The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU meets the following conditions: (a) PicA has a PictureOutputFlag equal to 1, (b) PicA has a nuh_layer_id nuhLid greater than the current picture, (c) PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid). All pictures (e.g., the syntax element picA) that meet all of these conditions are included. (5) The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0249] If none of the above conditions are true, the syntax element PictureOutputFlag may be set equal to the syntax element pic_output_flag.
[0250] After all slices of the current picture have been decoded, the current decoded picture may be marked as "used for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] may be marked as "used for short-term reference".
[0251] In the same or other embodiments, if each layer is an output layer set, regardless of the value of the syntax element ols_mode_idc (864), the syntax element PictureOutputFlag is set equal to pic_output_flag.
[0252] In the same or another embodiment, the syntax element PictureOutputFlag is set equal to 0 when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is equal to 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions: PicA has a PictureOutputFlag equal to 1, PicA has a nuh_layer_id nuhLid greater than that of the current picture, and PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0253] In the same or another embodiment, the syntax element PictureOutputFlag is set equal to 0 when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is equal to 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0254] An Intra Random Access Point (IRAP) picture can be an Instantaneous Decoder Refresh (IDR) picture that supports a closed group picture structure, or a Clean Random Access (CRA) picture that supports an open group picture structure, or an encoded picture for random access. A Gradual Decoder Refresh (GDR) picture may be a picture for gentle random access with partial refresh of the picture.
[0255] Embodiments of the present disclosure may include syntax elements indicating an IRAP picture or a GDR picture. For example, referring to FIG. 26, a picture header (1) may be provided. In the picture header (1), a flag ph_gdr_or_irap_pic_flag (2) may be signaled. The flag indicates that an IRAP picture or a GDR picture exists within the current PU associated with the picture header (1).
[0256] In the same or another embodiment, as shown in FIG. 26, the flag ph_no_output_of_prior_pics_flag (3) can be conditionally signaled only when ph_gdr_or_irap_pic_flag (2) is equal to 1. The value of ph_no_output_of_prior_pics_flag (3) can be used for the output and removal processing of pictures from the DPB. The value of the flag can affect the output of previously decoded pictures in the DPB after decoding of pictures in a CVSS AU that is not the first AU in the bitstream.
[0257] Since the constraint specified by the semantics of ph_gdr_or_irap_pic_flag can be "unidirectional" such that a flag equal to 1 specifies that the current picture is a GDR or IRAP picture, there is a potential problem that an IRAP picture can have a ph_gdr_or_irap_pic_flag equal to 0. A flag ph_gdr_or_irap_pic_flag equal to 0 specifies that the current picture may or may not be a GDR picture and may or may not be an IRAP picture. When the value of ph_gdr_or_irap_pic_flag of an IRAP picture is equal to 0, the value of ph_no_output_of_prior_pics_flag can be used for DPB operation without signaling or inference rules.
[0258] To address potential issues, in one embodiment, the semantics constraint of ph_gdr_or_irap_pic_flag(2) can be specified as "bidirectional", and thus, it may be necessary to signal ph_no_output_of_prior_pics_flag(3) when the current picture is an IRAP picture. The flag ph_gdr_or_irap_pic_flag(2) equal to 1 specifies that the current picture is a GDR or IRAP picture. The flag ph_gdr_or_irap_pic_flag(2) equal to 0 specifies that the current picture is neither a GDR picture nor an IRAP picture.
[0259] In the same or another embodiment, the inference rule for the ph_no_output_of_prior_pics_flag(3) value, if it does not exist, can be specified as follows: The flag ph_no_output_of_prior_pics_flag(3) affects the output of previously decoded pictures in the DPB after decoding of pictures in a CVSS AU that is not the first AU in the bitstream. If it exists, it may be a bitstream compliance requirement that the value of ph_no_output_of_prior_pics_flag(3) must be the same for all pictures within an AU.
[0260] If ph_no_output_of_prior_pics_flag(3) exists in the picture header(1) of a picture within an AU, the ph_no_output_of_prior_pics_flag(3) value of the AU is the ph_no_output_of_prior_pics_flag(3) value of the picture within the AU. If it does not exist, the value of ph_no_output_of_prior_pics_flag(3) can be presumed to be equal to 0.
[0261] Referring to FIG. 27, the AU delimiter (10) can be used to indicate the start of the AU, whether the AU is an IRAP or GDR AU, and the type of slice present in the coded picture within the AU that contains the AU delimiter NAL unit. When the bitstream contains only one layer, there may be no standard decoding process associated with the AU delimiter (10).
[0262] In the AU delimiter (10), the aud_irap_or_gdr_au_flag (12) can indicate the presence of an IRAP or GDR AU and can be signaled as shown in FIG. 27. A flag aud_irap_or_gdr_au_flag (12) with a value of 1 may specify that the AU containing the AU delimiter is an IRAP or GDR AU. Also, a flag aud_irap_or_gdr_au_flag (12) having a value of 0 may indicate that the AU containing the AU delimiter (10) is neither an IRAP nor a GDR AU.
[0263] In the same or another embodiment, the flag aud_irap_or_gdr_au_flag (12) for an IRAP or GDR AU may be present when the bitstream has multiple layers and the sps_video_parameter_set_id is greater than 0. The video coding technology or standard may require the presence of an AU delimiter for a multi-layer bitstream.
[0264] In the same or another embodiment, referring to FIGS. 26 - 27, when the aud_irap_or_gdr_au_flag (12) is present and the value of the aud_irap_or_gdr_au_flag (12) is equal to 1, the value of the ph_gdr_or_irap_pic_flag (2) may be required to be equal to 1. This is because when the aud_irap_or_gdr_au_flag (12) in the AU delimiter (10) is 1, each PU may have a GDR or IRAP picture.
[0265] In the same or another embodiment, when pps_mixed_nalu_types_in_pic_flag is 1, the value of ph_no_output_of_prior_pics_flag may not exist, and when it is determined (e.g., by a decoder) that pps_mixed_nalu_types_in_pic_flag exists, the value of ph_no_output_of_prior_pics_flag(3) may be ignored.
[0266] In the same or another embodiment, ph_no_output_of_prior_pics_flag(3) may affect the output of previously decoded pictures in the DPB after decoding of pictures in a CVSS AU that is not the first AU in the bitstream. If it does not exist, the value of ph_no_output_of_prior_pics_flag(3) may be inferred to be equal to 1.
[0267] In the same or another embodiment, to solve the problem that the value of ph_no_output_of_prior_pics_flag(3) is used without an inference rule, when ph_no_output_of_prior_pics_flag(3) does not exist and ph_gdr_or_irap_pic_flag(2) is equal to 1, as shown in FIG. 28, ph_gdr_or_irap_pic_flag(2) in the picture header(1) can be replaced with ph_irap_pic_flag(6). The flag ph_irap_pic_flag(6) equal to 1 can specify that the current picture is an IRAP picture. The flag ph_irap_pic_flag(6) equal to 0 may specify that the current picture is not an IRAP picture.
[0268] In the same or another embodiment, to solve the problem that the value of ph_no_output_of_prior_pics_flag(3) is used without inference rules, when ph_no_output_of_prior_pics_flag(3) does not exist and ph_gdr_or_irap_pic_flag(2) is equal to 1, as shown in FIG. 29, ph_no_output_of_prior_pics_flag(3) in the picture header(1) may be replaced by sh_no_output_of_prior_pics_flag(23) in the slice header(20).
[0269] In the same embodiment, sh_no_output_of_prior_pics_flag(23) may conditionally exist in the slice header(20) only when the NAL unit type of the current VCL NAL is equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT. IDR_W_RADL is a NAL unit type that includes an encoded slice segment of an IDR picture that does not have a related RASL picture present in the bitstream but may have a related RADL picture in the bitstream. IDR_N_LP may be a NAL unit type that includes an encoded slice segment of an IDR picture that does not have a related leading picture present in the bitstream. CRA_NUT is a NAL unit type that includes an encoded slice segment of a CRA picture.
[0270] In the same or another embodiment, sh_no_output_of_prior_pics_flag(23) may affect the output of previously decoded pictures in the DPB after decoding of pictures in a CVSS AU that is not the first AU in the bitstream.
[0271] In the same or another embodiment, when present, the value of sh_no_output_of_prior_pics_flag(23) should be the same for all pictures within an AU, which may be a requirement for bitstream compliance. If sh_no_output_of_prior_pics_flag(23) is present in the slice header (20) of a picture within an AU, the sh_no_output_of_prior_pics_flag(23) value of the AU may be the sh_no_output_of_prior_pics_flag(23) value of the picture within the AU.
[0272] In the same or another embodiment, when pps_mixed_nalu_types_in_pic_flag within the picture parameter set is equal to 1, the value of sh_no_output_of_prior_pics_flag(23) may not be present. If present, the value of sh_no_output_of_prior_pics_flag(23) may be ignored.
[0273] In the same or another embodiment, as shown in FIG. 30, aud_irap_au_flag(16) may be present in the AU delimiter (10). The flag aud_irap_au_flag(16) being 1 may identify that the AU containing the AU delimiter (10) is an IRAP AU. The flag aud_irap_au_flag(16) being 0 may indicate that the AU containing the AU delimiter (10) is not an IRAP AU.
[0274] Also, in this embodiment, when aud_irap_au_flag(16) exists, the value of ph_irap_pic_flag(6) in the picture header(1) may be equal to aud_irap_au_flag(16) of the AU delimiter(10). The flag ph_irap_pic_flag(7) equal to 1 may specify that the picture associated with the PH(1) is an IRAP picture. The flag ph_irap_pic_flag(6) equal to 0 may specify that the picture associated with the PH(1) is not an IRAP picture.
[0275] In the same or another embodiment, as shown in FIG. 30, aud_gdr_au_flag(17) may exist in the AU delimiter(10). The flag aud_gdr_au_flag(17) being 1 may indicate that the AU containing the AU delimiter(10) is a GDR AU. The flag aud_irap_au_flag(17) being 0 may indicate that the AU containing the AU delimiter(10) is not a GDR AU.
[0276] Also, in this embodiment, when aud_gdr_au_flag(17) exists in the AU delimiter(10), the value of ph_gdr_pic_flag(7) in the picture header(1) may be equal to aud_gdr_au_flag(17) in the AU delimiter(10). The flag ph_gdr_pic_flag(7) equal to 1 may specify that the picture associated with the PH(1) is a GDR picture. The flag ph_gdr_pic_flag(7) equal to 0 may specify that the picture associated with the PH(1) is not a GDR picture.
[0277] Gradual Decoding Refresh (GDR) can be specified by the following definition. GDR AU: An AU in which PUs exist for each layer specified in the VPS, and the encoded pictures within each existing PU are GDR pictures. GDR PU: A PU in which the encoded picture is a GDR picture. GDR picture: A picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUT. GDR sub-picture: A sub-picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUT. GDR_NUT: A NAL unit type that includes an encoded tile group of a GDR picture.
[0278] According to an embodiment, the first picture in the decoded order bitstream can be an IRAP or GDR picture. Subsequent pictures associated with the IRAP or GDR picture can also follow the IRAP or GDR picture in the decoded order. Pictures that follow the output order of the associated IRAP picture and precede the decoded order of the associated IRAP picture may not be allowed.
[0279] In one embodiment, referring to FIG. 31, the syntax element indicating GDR is signaled by PH(1) such as ph_gdr_pic_flag(7). If it does not exist, the value of ph_gdr_pic_flag(7) can be assumed to be equal to 0. When sps_gdr_enabled_flag is equal to 0, the value of ph_gdr_pic_flag(7) can be assumed to be equal to 0. The syntax element ph_recovery_poc_cnt(32) can specify the recovery point of the decoded picture in the output order.
[0280] If the current picture is a GDR picture, the variable recoveryPointPocVal may be derived as follows. recoveryPointPocVal = PicOrderCntVal + ph_recovery_poc_cnt
[0281] In the same or another embodiment, as shown in FIG. 31, since the PicOrderCntVal used to derive the recoveryPointPocVal is derived from the values of ph_pic_order_cnt_lsb(32) and ph_poc_msb_cycle_val(36), ph_recovery_poc_cnt(34) is signaled after signaling the picture order count (POC) syntax elements (e.g., ph_pic_order_cnt_lsb(32) and ph_poc_msb_cycle_val(36)).
[0282] If the current picture is a GDR picture and there is a picture picA in the CLVS that follows the current GDR picture in decoding order and has a PicOrderCntVal equal to the recoveryPointPocVal, then that picture picA can be called a recovery point picture. Otherwise, the first picture in output order that has a PicOrderCntVal greater than the recoveryPointPocVal in the CLVS can be called a recovery point picture. The recovery point picture may not be before the current GDR picture in decoding order. A picture associated with the current GDR picture and having a PicOrderCntVal less than the recoveryPointPocVal may sometimes be called a recovery picture of the GDR picture. The value of ph_recovery_poc_cnt(34) may be greater than or equal to 0 and less than or equal to MaxPicOrderCntLsb - 1.
[0283] In the same or another embodiment, the recovery point picture may not be before the current GDR picture in both decoding order and output order.
[0284] In the same or another embodiment, the recovery picture may not be before the current GDR picture in both decoding order and output order.
[0285] In the same or another embodiment, the recovery picture may precede the associated recovery point picture in both the decoding order and the output order.
[0286] In the same or another embodiment, when the current picture is a recovery picture of a GDR picture or a GDR picture, and the current picture includes a non-CTU alignment boundary between a "refresh region" (i.e., a region having an exact match of decoded sample values when starting the decoding process from the GDR picture as compared to starting the decoding process in decoding order from the previous IRAP picture, if it exists) and a "dirty region" (i.e., a region that may not have an exact match of decoded sample values when starting the decoding process from the GDR picture as compared to starting the decoding process in decoding order from the previous IRAP picture, if it exists), in order to avoid the "dirty region" and affect the decoded sample values of the "refresh region", it may be necessary to disable the chroma residual scaling of luma mapping using chroma scaling (LMCS) in the current picture.
[0287] In the same or another embodiment, the value of recoveryPointPocVal of layer A may be greater than or equal to the recoveryPointPocVal of the reference layer of layer A.
[0288] In the same or another embodiment, the value of recoveryPointPocVal of layer A with layerId equal to m may be greater than or equal to the recoveryPointPocVal of another layer B with layerId equal to n, where m is greater than n.
[0289] In the same or another embodiment, the value of recoveryPointPocVal of layer A with layerId equal to m may be greater than or equal to the recoveryPointPocVal of another layer B with layerId equal to n, where m is greater than n, and layers A and B belong to the same output layer set.
[0290] In the same or another embodiment, the value of recoveryPointPocVal of layer A may be equal to the recoveryPointPocVal of the reference layer of layer A.
[0291] In the same or another embodiment, the value of recoveryPointPocVal of layer A where layerId is equal to m may be equal to the recoveryPointPocVal of another layer B where layerId is equal to n, and m is greater than n.
[0292] In the same or another embodiment, the value of recoveryPointPocVal of layer A where layerId is equal to m may be equal to the recoveryPointPocVal of another layer B where layerId is equal to n, m is greater than n, and layers A and B belong to the same output layer set.
[0293] In the same or another embodiment, when pps_mixed_nalu_types_in_pic_flag is equal to 1, the following may apply. (1) Assume that the picture has at least two sub-pictures. (2) Assume that the VCL NAL units of the picture have two or more different nal_unit_type values. (3) Assume that there is no VCL NAL unit of the picture having a nal_unit_type equal to GDR_NUT. (4) Assume that the picture is not a restored or restored picture associated with a GDR picture. (5) When at least one sub-picture's VCL NAL unit of the picture has a specific value of nal_unit_type equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT, assume that all the VCL NAL units of the other sub-pictures in the picture have a nal_unit_type equal to TRAIL_NUT.
[0294] TRAIL_NUT is a NAL unit type that includes an encoding tile group of VCL non-STSA trailing pictures.
[0295] At the start of the decoding process for each slice of a picture, a decoding process for reference picture list construction may be called for the derivation of reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]). After constructing one or more reference picture lists, a decoding process for reference picture marking may be called, and the reference pictures may be marked as "not used for reference", "used for short-term reference", or "used for long-term reference".
[0296] Referring to FIG. 32, according to an embodiment, the decoding process (40) can be executed by a decoder. In the decoding process (40), one or more reference picture lists (RPLs) can be constructed (42). When one or more reference picture lists are constructed, one or more pictures may not be available for random access or due to unintentional picture loss. The decoder can determine whether the reference pictures in the RPL are available in the DPB (44). If it is determined that a reference picture is not available, the unavailable reference picture can be marked as "no reference picture". To avoid decoder crashes or unintentional behavior, the unavailable reference picture can be immediately generated using initial setting values of pixels and parameters (46). After generating the unavailable reference picture (and / or after it is determined that the reference picture is available), the decoder can confirm the verification of all reference pictures including the generated picture in the reference picture list (48).
[0297] In the same or another embodiment, one or more reference picture lists are constructed (42) by parsing the RPL syntax elements in the SPS, PH, and / or SH. After the construction step (42), the random access skipped leading picture (RASL) associated with the CRA picture may be discarded by the decoder or system, or one or more reference pictures in the RPL list may not be available because they cannot be correctly decoded if random access occurs to the CRA picture. Unavailable reference pictures can be generated using the initial setting values of pixels and parameters (46).
[0298] In the same or another embodiment, if the current picture is an IDR picture having a sps_idr_rpl_present_flag equal to 1, or a pps_rpl_info_in_ph_flag equal to 1, a CRA picture having a NoOutputBeforeRecoveryFlag equal to 1, or a GDR picture having a NoOutputBeforeRecoveryFlag equal to 1, at least one of the following decoding processes for generating unavailable reference pictures (46) is called, which may need to be called only for the first slice of the picture.
[0299] A. General decoding process for generating unavailable reference pictures This process can be called once per coded picture if the current picture is an IDR picture having a sps_idr_rpl_present_flag equal to 1 or a pps_rpl_info_in_ph_flag equal to 1, a CRA picture having a NoOutputBeforeRecoveryFlag equal to 1, or a GDR picture having a NoOutputBeforeRecoveryFlag equal to 1. When this process is called, the following can apply: for i in the range of 0 to 1 and j in the range of 0 to num_ref_entries[i][RplsIdx[i]] - 1, that is, for each RefPicList[i][j] equal to "no reference picture", a picture is generated as described below in "generation of one unavailable picture", and the following applies. (1) The value of nuh_layer_id of the generated picture is set equal to the nuh_layer_id of the current picture. (2) If st_ref_pic_flag[i][RplsIdx[i]][j] is equal to 1 and inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0, the value of PicOrderCntVal of the generated picture is set equal to RefPicPocList[i][j], and the generated picture is marked as "used for short-term reference". (3) Otherwise, if st_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0 and inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0, the value of PicOrderCntVal for the generated picture is set equal to RefPicLtPocList[i][j], the value of ph_pic_order_cnt_lsb for the generated picture is presumed to be equal to (RefPicLtPocList[i][j] & (MaxPicOrderCntLsb - 1)), and the generated picture is marked as "used for long-term reference". (4) The value of PictureOutputFlag of the generated reference picture is set to 0. (5) The generated reference picture has RefPicList[i][j] set. (6) The value of the TemporalId of the generated picture is set equal to the TemporalId of the current picture. (7) The value of ph_non_ref_pic_flag of the generated picture is set equal to 0. (8) The value of ph_pic_parameter_set_id of the generated picture is set equal to the ph_pic_parameter_set_id of the current picture.
[0300] The flag ph_non_ref_pic_flag equal to 1 may specify that the picture associated with PH is never used as a reference picture. The flag ph_non_ref_pic_flag equal to 0 specifies that the picture associated with PH may or may not be used as a reference picture.
[0301] B. Generation of an unusable picture When this process is called, an unusable picture is generated as follows. (1) The value of each element in the sample array S of the picture L is set equal to 1<<(BitDepth - 1). (2) When sps_chroma_format_idc is not 0, for the sample array S of that picture Cb , S Cr the value of each element is set equal to 1<<(BitDepth - 1). (3) The prediction mode CuPredMode[0][x][y] is set equal to MODE_INTRA when x is greater than or equal to 0 and less than or equal to pps_pic_width_in_luma_samples - 1, and y is greater than or equal to 0 and less than or equal to pps_pic_height_in_luma_samples - 1.
[0302] In the same or another embodiment, after generating unavailable reference pictures in the RPL list, a bitstream compliance check of all or active reference pictures in the RPL list can be invoked, for example, by a decoder. For example, the decoder can confirm that the following constraints are applied to the bitstream compliance. (1) For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i]. (2) Assume that the pictures referred to by each active entry of RefPicList[0] or RefPicList[1] exist in the DPB and are below the TemporalId of the current picture. (3) Assume that the pictures referred to by each entry in RefPicList[0] or RefPicList[1] are not the current picture and have a ph_non_ref_pic_flag equal to 0. (4) The short-term reference picture (STRP) entry in RefPicList[0] or RefPicList[1] of a slice of a picture and the long-term reference picture (LTRP) entry in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture shall not refer to the same picture. (5) There shall be no LTRP entry in RefPicList[0] or RefPicList[1] where the difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referred to by the entry is 2 24 or more. Let the setOfRefPics be the set of unique pictures referenced by all entries in RefPicList[0] that have the same nuh_layer_id as the current picture, and all entries in RefPicList[1] that have the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics must be less than or equal to MaxDpbSize - 1, where MaxDpbSize is [not provided in the original], and setOfRefPics must be the same for all slices of the pictures. (7) When the current slice has a nal_unit_type equal to STSA_NUT, assume that there are no active entries in RefPicList[0] or RefPicList[1] that have a TemporalId equal to that of the current picture and a nuh_layer_id equal to that of the current picture. (8) If the current picture is a picture that follows, in decoding order, a hierarchical temporal sub-layer access (STSA) picture that has the same TemporalId as the current picture and the same nuh_layer_id as the current picture, assume that there is no picture that precedes the STSA picture in decoding order, has the same TemporalId as the current picture, and has the same nuh_layer_id as the current picture and is included as an active entry in RefPicList[0] or RefPicList[1]. If the current subpicture having a TemporalId equal to a specific value tId, a nuh_layer_id equal to a specific value layerId, and a subpicture index equal to a specific value subpicIdx is a subpicture that follows, in decoding order, an STSA subpicture having a TemporalId equal to tId, a nuh_layer_id equal to layerId, and a subpicture index equal to subpicIdx, then it is assumed that there is no picture having a TemporalId equal to tId and a nuh_layer_id equal to layerId that precedes, in decoding order, the picture containing the STSA subpicture included as an active entry in RefPicList[0] or RefPicList[1]. When the current picture having a nuh_layer_id equal to a specific value layerId is an IRAP picture, it is assumed that there is no picture referenced by an entry within RefPicList[0] or RefPicList[1] that precedes, in output order or decoding order, (if it exists) a preceding IRAP picture having a nuh_layer_id equal to layerId in decoding order. When the current subpicture having a nuh_layer_id equal to a specific value layerId and a subpicture index equal to a specific value subpicIdx is an IRAP subpicture, it is assumed that there is no picture referenced by an entry within RefPicList[0] or RefPicList[1] that precedes, in output order or decoding order, any preceding picture containing an IRAP subpicture having a nuh_layer_id equal to layerId and a subpicture index equal to subpicIdx (if it exists) in decoding order. If the current picture is not a RASL picture associated with a CRA picture where the NoOutputBeforeRecoveryFlag is 1, then there shall be no picture referenced by an active entry in RefPicList[0] or RefPicList[1] that is generated by the decoding process to generate unavailable reference pictures for the CRA picture associated with the current picture. (13) If the current sub-picture is not a RASL sub-picture associated with a CRA sub-picture within a CRA picture where the NoOutputBeforeRecoveryFlag is 1, then there shall be no picture referenced by an active entry in RefPicList[0] or RefPicList[1] that is generated by the decoding process to generate unavailable reference pictures for the CRA picture including the CRA sub-picture associated with the current sub-picture. (14) If the current picture with nuh_layer_id equal to a specific value layerId is not any of the following, then there should be no picture referenced by an entry in RefPicList[0] or RefPicList[1] that is generated by the decoding process to generate unavailable reference pictures for the IRAP picture or GDR picture associated with the current picture. (a) An IDR picture where sps_idr_rpl_present_flag is 1 or pps_rpl_info_in_ph_flag is 1. (b) A CRA picture where the NoOutputBeforeRecoveryFlag is 1. (c) A picture associated with a CRA picture where the NoOutputBeforeRecoveryFlag is 1, and the decoding order of which is earlier than that of the first picture associated with the same CRA picture. (d) The first picture associated with a CRA picture where the NoOutputBeforeRecoveryFlag is 1 (e) A GDR picture where the NoOutputBeforeRecov.eryFlag is 1. (f) The reconstructed picture of a GDR picture where NoOutputBeforeRecoveryFlag is 1 and nuh_layer_id is layerId. (15) When the current subpicture where nuh_layer_id is equal to a specific value layerId and the subpicture index is equal to a specific value subpicIdx is not any of the following, there should be no picture referenced by an entry in RefPicList[0] or RefPicList[1] generated by the decoding process for generating a non - usable reference picture for an IRAP or GDR picture that includes the current subpicture or an IRAP or GDR subpicture associated with the current subpicture. (a) An IDR subpicture within an IDR picture where sps_idr_rpl_present_flag is 1 or pps_rpl_info_in_ph_flag is 1. (b) A CRA subpicture within a CRA picture where NoOutputBeforeRecoveryFlag is 1. (c) A subpicture associated with a CRA subpicture within a CRA picture where NoOutputBeforeRecoveryFlag is 1, and the decoding order of which is earlier than that of the leading picture associated with the same CRA picture. (d) The leading subpicture associated with a CRA subpicture within a CRA picture where NoOutputBeforeRecoveryFlag is 1. (e) A GDR subpicture within a GDR picture where NoOutputBeforeRecoveryFlag is 1. (f) A subpicture within the reconstructed picture of a GDR picture where NoOutputBeforeRecoveryFlag is 1 and nuh_layer_id is layerId. (16) When the current picture follows an IRAP picture with the same value of nuh_layer_id in both the decoding order and the output order, it is assumed that there is no picture referred to by an active entry in RefPicList[0] or RefPicList[1] that precedes the IRAP picture in the output order or the decoding order. (17) When the current sub-picture follows an IRAP sub-picture with the same value of nuh_layer_id and the same value of sub-picture index in both the decoding order and the output order, it is assumed that there is no picture referred to by an active entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the IRAP sub-picture in the output order or the decoding order. (18) When tracing an IRAP picture with the same value of nuh_layer_id and, if any, the leading picture associated with the IRAP picture in both the decoding order and the output order from the current picture, it is assumed that there is no picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes the IRAP picture in the output order or the decoding order. (19) When the current sub-picture follows an IRAP sub-picture with the same value of nuh_layer_id and the same value of sub-picture index and, if any, the leading sub-picture associated with the IRAP sub-picture in both the decoding order and the output order, it is assumed that there is no picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the IRAP sub-picture in the output order or the decoding order. (20) When the current picture is a Random Access Decodable Leading (RADL) picture, it is assumed that there is no active entry in either RefPicList[0] or RefPicList[1]. (a) The RASL picture of pps_mixed_nalu_types_in_pic_flag is 0. This means that the active entry of the RPL of the RADL picture can refer to a RASL picture having pps_mixed_nalu_types_in_pic_flag equal to 1. However, the RADL picture is subject to the following constraint that does not permit a RADL subpicture to refer to a RASL subpicture. So as to refer only to the RADL subpictures within the referred RASL picture, since the RADL subpictures within the referred RASL picture are correctly decoded, when decoding starts from the associated CRA picture, such a RADL picture can still be correctly decoded. (b) A picture preceding the associated IRAP picture in the decoding order. (21) When the current subpicture having nuh_layer_id equal to a specific value layerId and a subpicture index equal to a specific value subpicIdx is a RADL subpicture, assume that there is no active entry in either RefPicList[0] or RefPicList[1]. (a) The picture with nuh_layer_id equal to layerId includes a RASL subpicture having a subpicture index equal to subpicIdx. (b) A picture preceding the picture including the associated IRAP subpicture in the decoding order. (22) The following constraints apply to the picture referred to by each ILRP entry if it exists in RefPicList[0] or RefPicList[1] of the slice of the current picture. (a) Assume that the picture is in the same AU as the current picture. (b) Assume that the picture exists in the DPB. (c) Assume that the picture has a refPicLayerId of nuh_layer_id smaller than that of the current picture. (d) Any of the following constraints shall apply: the picture shall be an IRAP picture; or the picture shall have a TemporalId less than or equal to Max(0, vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx] - 1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively. (23) If present in RefPicList[0] or RefPicList[1] of a slice, each ILRP entry shall be an active entry. (24) If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and sps_num_subpics_minus1 is greater than 0, then either (but not both) of the following two conditions shall be true. (a) The picture referred to by each active entry in RefPicList[0] or RefPicList[1] shall have the same sub-picture layout as the current picture (i.e., the SPSs referred to by that picture and the current picture shall have the same value of sps_num_subpics_minus1, and the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], and sps_subpic_height_minus1[j] for each value of j in the range from 0 to sps_num_subpics_minus1). (b) The picture referred to by each active entry in RefPicList[0] or RefPicList[1] shall be an ILRP with a value of sps_num_subpics_minus1 equal to 0.
[0303] C. Decoding Process for Reference Picture Marking In the same or another embodiment, when a reference picture is marked, the following decoding process may be invoked.
[0304] This process may be invoked once per picture after decoding the slice header and the decoding process for constructing the reference picture list of the slice, but before decoding the slice data. By this process, one or more reference pictures in the DPB may be marked as "not used for reference" or "used for long-term reference".
[0305] The decoded pictures in the DPB can be marked as "not used for reference", "used for short-term reference", or "used for long-term reference", but only one of these three can be marked at any given moment during the operation of the decoding process. Assigning one of these markings to a picture can implicitly remove another of these markings if applicable. When a picture is referred to as being marked as "used for reference", this generically refers to a picture marked as "used for short-term reference" or "used for long-term reference" (not both).
[0306] STRP and ILRP can be identified by their nuh_layer_id and PicOrderCntVal values. LTRP can be identified by their nuh_layer_id values and the Log 2(MaxLtPicOrderCntLsb) LSB of their PicOrderCntVal values.
[0307] If the current picture is a CLVSS picture, all current reference pictures in the DPB having the same nuh_layer_id as the current picture (if any) may be marked as "not used for reference".
[0308] Otherwise, the following may apply. (1) For each LTRP entry in RefPicList[0] or RefPicList[1], if the picture is an LTRP that has the same nuh_layer_id as the current picture, then that picture is marked as "used for long-term reference". (2) A reference picture with the same nuh_layer_id as the current picture of a DPB that is not referenced by any entry in RefPicList[0] or RefPicList[1] is considered "not used for reference". (3) For each ILRP entry in RefPicList[0] or RefPicList[1], the picture is marked as "used for long-term reference".
[0309] In the same or another embodiment, for each LTRP entry in RefPicList[0] or RefPicList[1], if a picture that has the same nuh_layer_id as the current picture is marked as "used for short-term reference", then that picture is marked as "used for long-term reference".
[0310] In the same or another embodiment, when one or more RPL lists are constructed, the following may apply. For each RefPicList[i][j], i is in the range from 0 to 1, and j is in the range from 0 to num_ref_entries[i][RplsIdx[i]] - 1, that is, a picture equal to "no reference picture" is generated.
[0311] In the same or another embodiment, when one or more RPL lists are constructed, the following may apply. For each RefPicList[i][j], i is in the range from 0 to 1, and j is in the range from 0 to NumRefIdxActive[i] - 1, that is, a picture equal to "no reference picture" is generated.
[0312] In the decoding process for constructing the reference picture list, reference pictures missing in the DPB can be set equal to "no reference picture". Unusable pictures equal to "no reference picture" are generated through the decoding process for generating unusable reference pictures for compliance checking purposes. The problem is that, as follows, not only unusable reference pictures within the same layer but also inter-layer reference pictures that are unusable within the reference layer are set equal to "no reference picture". if(!inter_layer_ref_pic_flag[i][RplsIdx[i]][j]){ if(st_ref_pic_flag[i][RplsIdx[i]][j]){ RefPicPocList[i][j]=pocBase+DeltaPocValSt[i][RplsIdx[i]][j] if(there is areference picture picA in the DPB with the same nuh_layer_id as the current picture and PicOrderCntVal equal to RefPicPocList[i][j]) RefPicList[i][j]=picA else RefPicList[i][j]=“no reference picture” pocBase=RefPicPocList[i][j] }else{ if(!delta_poc_msb_cycle_present_flag[i][k]){ if(there is areference picA in the DPB with the same nuh_layer_id as the current picture and PicOrderCntVal&(MaxPicOrderCntLsb-1)equal to PocLsbLt[i][k]) RefPicList[i][j]=picA else RefPicList[i][j] = "no reference picture" RefPicLtPocList[i][j] = PocLsbLt[i][k] } else { if (there is a reference picA in the DPB with the same nuh_layer_id as the current picture and PicOrderCntVal equal to FullPocLt[i][k]) RefPicList[i][j] = picA else RefPicList[i][j] = "no reference picture" RefPicLtPocList[i][j] = FullPocLt[i][k] } k++ } } else { layerIdx = DirectRefLayerIdx[GeneralLayerIdx[nuh_layer_id]][ilrp_idx[i][RplsIdx][j]] refPicLayerId = vps_layer_id[layerIdx] if (there is a reference picture picA in the DPB with nuh_layer_id equal to refPicLayerId and the same PicOrderCntVal as the current picture) RefPicList[i][j] = picA else RefPicList[i][j] = "no reference picture" }
[0313] However, in the decoding process for generating unavailable reference pictures, all unavailable pictures equal to "no reference picture" may be treated as reference pictures within the same layer. As a result, the nuh_layer_id and PicOrderCntVal values of the unavailable inter-layer reference pictures are set to accurate values. Also, the unavailable inter-layer reference pictures are not correctly marked as "used for long-term reference". These incorrect values cause errors in the subsequent decoding process of pictures and the bitstream conformity check. Therefore, according to the embodiment, the unavailable inter-layer reference pictures equal to "no reference picture" should be correctly generated as follows.
[0314] D. Improved Decoding Process for Generating Unavailable Reference Pictures This process is called once for each encoded picture when the current picture is an IDR picture with sps_idr_rpl_present_flag equal to 1 or pps_rpl_info_in_ph_flag equal to 1, a CRA picture with NoOutputBeforeRecoveryFlag equal to 1, or a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0315] When this process is called, for each RefPicList[i][j], with i in the range of 0 to 1 and j in the range of 0 to num_ref_entries[i][RplsIdx[i]] - 1, that is, using i equal to "no reference picture", pictures can be generated as described above in the present disclosure, and the following can be applied. (1) When inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0, the value of the nuh_layer_id of the generated picture is set equal to the nuh_layer_id of the current picture. (2) Otherwise (when inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 1), the value of nuh_layer_id of the generated picture is set to be equal to vps_layer_id[DirectRefLayerIdx[GeneralLayerIdx[nuh_layer_id]][ilrp_idx[i][RplsIdx][j]]]. (3) When st_ref_pic_flag[i][RplsIdx[i]][j] is equal to 1 and inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0, the value of PicOrderCntVal of the generated picture is set to be equal to RefPicPocList[i][j], and the generated picture is marked as "used for short-term reference". (4) When st_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0 and inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 0, the value of PicOrderCntVal of the generated picture is set to be equal to RefPicLtPocList[i][j], the value of ph_pic_order_cnt_lsb of the generated picture is presumed to be equal to (RefPicLtPocList[i][j] & (MaxPicOrderCntLsb - 1)), and the generated picture is marked as "used for long-term reference". (5) Alternatively, when inter_layer_ref_pic_flag[i][RplsIdx[i]][j] is equal to 1, the value of PicOrderCntVal of the generated picture is set to be equal to the PicOrderCntVal of the current picture, the value of ph_pic_order_cnt_lsb of the generated picture is presumed to be equal to the ph_pic_order_cnt_lsb of the current picture, and the generated picture is marked as "used for long-term reference". (6) The value of PictureOutputFlag of the generated reference picture is set to 0. (7) RefPicList[i][j] is set for the generated reference picture. (8) The value of the TemporalId of the generated picture is set equal to the TemporalId of the current picture. (9) The value of the ph_non_ref_pic_flag of the generated picture is set equal to 0. (10) The value of the ph_pic_parameter_set_id of the generated picture is set equal to the ph_pic_parameter_set_id of the current picture.
[0316] In the decoding process for the reference picture marking described above, in order to clarify the marking process, the following sentence is modified: For each LTRP entry in RefPicList[0] or RefPicList[1], if the picture is an STRP having the same nuh_layer_id as the current picture, then that picture is marked as "used for long-term reference".
[0317] As a first option, the original sentence is modified as follows: For each LTRP entry in RefPicList[0] or RefPicList[1], if the picture is currently marked as "used for short-term reference" with the same nuh_layer_id as the current picture, then the picture is marked as "used for long-term reference".
[0318] As a second option, the original sentence is modified as follows: For each LTRP entry in RefPicList[0] or RefPicList[1], if the picture is an LTRP having the same nuh_layer_id as the current picture, then the picture is marked as "used for long-term reference".
[0319] Since STRP is marked as "used for long-term reference", the meaning of the original text is a bit unclear. If the original intention of the sentence is to mark the LTRP picture of the current picture in RefPicList[0] or RefPicList[1] as "used for long-term reference", Option 1 can be used when the LTRP is currently marked as "used for short-term reference". Otherwise, Option 2 can be used.
[0320] E. Improved Decoding Process for Reference Picture List Construction In HEVC, the reference picture list is constructed after generating the unavailable reference pictures with the RPS. In VVC (current draft), the reference picture list is constructed before generating the unavailable reference pictures with the RPS. If the bitstream compliance check is performed before generating the unavailable reference pictures as previously specified, the bitstream may not be able to comply with the specified constraints. To clarify the order, the decoding process can be modified as follows to have the following constraints regarding bitstream compliance. (1) For each i equal to 0 or 1, num_ref_entries[i][RplsIdx[i]] shall not be less than NumRefIdxActive[i]. (2) The pictures referred to by each active entry in RefPicList[0] or RefPicList[1] shall be in the DPB and have a TemporalId less than or equal to that of the current picture. (3) The pictures referred to by each entry in RefPicList[0] or RefPicList[1] shall not be the current picture and shall have a ph_non_ref_pic_flag equal to 0. (4) The STRP entries in RefPicList[0] or RefPicList[1] of a slice of a picture, and the LTRP entries in RefPicList[0] or RefPicList[1] of the same slice or a different slice of the same picture, shall not refer to the same picture. (5) The difference between the PicOrderCntVal of the current picture and the PicOrderCntVal of the picture referenced by the entry shall be 2 24 or more, and no LTRP entry shall exist within RefPicList[0] or RefPicList[1]. (6) Set setOfRefPics to the set of unique pictures referenced by all entries in RefPicList[0] that have the same nuh_layer_id as the current picture, and all entries in RefPicList[1] that have the same nuh_layer_id as the current picture. The number of pictures in setOfRefPics shall be less than or equal to MaxDpbSize - 1, and MaxDpbSize and setOfRefPics shall be the same for all slices of the picture. (7) When the current slice has a nal_unit_type equal to STSA_NUT, assume that there is no active entry within RefPicList[0] or RefPicList[1] that has the same TemporalId as that of the current picture and the same nuh_layer_id as that of the current picture. (8) When the current picture is a picture that follows, in decoding order, an STSA picture that has the same TemporalId as the current picture and the same nuh_layer_id as the current picture, assume that there is no picture that precedes the STSA picture in decoding order and has the same TemporalId as the current picture and the same nuh_layer_id as the current picture that is included as an active entry in RefPicList[0] or RefPicList[1]. If the current subpicture having a TemporalId equal to a specific value tId, a nuh_layer_id equal to a specific value layerId, and a subpicture index equal to a specific value subpicIdx is a subpicture that follows, in decoding order, an STSA subpicture having a TemporalId equal to tId, a nuh_layer_id equal to layerId, and a subpicture index equal to subpicIdx, then it is assumed that there is no picture having a TemporalId equal to tId and a nuh_layer_id equal to layerId that precedes, in decoding order, a picture containing the STSA subpicture included as an active entry in RefPicList[0] or RefPicList[1]. (10) When the current picture having a nuh_layer_id equal to a specific value layerId is an IRAP picture, it is assumed that there is no picture referenced by an entry within RefPicList[0] or RefPicList[1] that precedes, in output order or decoding order, (if it exists) a preceding IRAP picture having a nuh_layer_id equal to layerId in decoding order. (11) When the current subpicture having a nuh_layer_id equal to a specific value layerId and a subpicture index equal to a specific value subpicIdx is an IRAP subpicture, it is assumed that there is no picture referenced by an entry within RefPicList[0] or RefPicList[1] that precedes, in output order or decoding order, (if it exists) any preceding picture containing an IRAP subpicture having a nuh_layer_id equal to layerId and a subpicture index equal to subpicIdx in decoding order. If the current picture is not a RASL picture associated with a CRA picture where NoOutputBeforeRecoveryFlag is 1, it is assumed that there is no picture referenced by an active entry in RefPicList[0] or RefPicList[1] that is generated by a decoding process for generating an unavailable reference picture for the CRA picture associated with the current picture. (13) If the current sub - picture is not a RASL sub - picture associated with a CRA sub - picture within a CRA picture where NoOutputBeforeRecoveryFlag is 1, it is assumed that there is no picture referenced by an active entry in RefPicList[0] or RefPicList[1] that is generated by a decoding process for generating an unavailable reference picture for the CRA picture that includes the CRA sub - picture associated with the current sub - picture. (14) If the current picture with nuh_layer_id equal to a specific value layerId is not any of the following, it should be the case that there is no picture referenced by an entry in RefPicList[0] or RefPicList[1] that is generated by a decoding process for generating an unavailable reference picture for the IRAP picture or GDR picture associated with the current picture. (a) An IDR picture where sps_idr_rpl_present_flag is 1 or pps_rpl_info_in_ph_flag is 1. (b) A CRA picture where NoOutputBeforeRecoveryFlag is 1. (c) A picture associated with a CRA picture where NoOutputBeforeRecoveryFlag is 1 and whose decoding order is earlier than that of the first picture associated with the same CRA picture. (d) The first picture associated with a CRA picture where NoOutputBeforeRecoveryFlag is 1. (e) A GDR picture where NoOutputBeforeRecoveryFlag is 1. (f) The reconstructed picture of a GDR picture where NoOutputBeforeRecoveryFlag is 1 and nuh_layer_id is layerId. (15) When the current subpicture where nuh_layer_id is equal to a specific value layerId and the subpicture index is equal to a specific value subpicIdx is not any of the following, there should be no picture referenced by an entry in RefPicList[0] or RefPicList[1] generated by the decoding process for generating a non - usable reference picture for an IRAP or GDR picture that includes the current subpicture or an IRAP or GDR subpicture associated with the current subpicture. (a) An IDR subpicture within an IDR picture where sps_idr_rpl_present_flag is 1 or pps_rpl_info_in_ph_flag is 1. (b) A CRA subpicture within a CRA picture where NoOutputBeforeRecoveryFlag is 1. (c) A subpicture associated with a CRA subpicture within a CRA picture where NoOutputBeforeRecoveryFlag is 1, and the decoding order is earlier than that of the leading picture associated with the same CRA picture. (d) The leading subpicture associated with a CRA subpicture within a CRA picture where NoOutputBeforeRecoveryFlag is 1. (e) A GDR subpicture within a GDR picture where NoOutputBeforeRecoveryFlag is 1. (f) A subpicture within the reconstructed picture of a GDR picture where NoOutputBeforeRecoveryFlag is 1 and nuh_layer_id is layerId. (16) When the current picture follows an IRAP picture with the same value of nuh_layer_id in both the decoding order and the output order, it is assumed that there is no picture referred to by an active entry in RefPicList[0] or RefPicList[1] that precedes the IRAP picture in the output order or the decoding order. (17) When the current sub-picture follows an IRAP sub-picture having the same value of nuh_layer_id and the same value of sub-picture index in both the decoding order and the output order, it is assumed that there is no picture referred to by an active entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the IRAP sub-picture in the output order or the decoding order. (18) When tracing an IRAP picture with the same value of nuh_layer_id and, if any, the leading picture associated with the IRAP picture in both the decoding order and the output order from the current picture, it is assumed that there is no picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes the IRAP picture in the output order or the decoding order. (19) When the current sub-picture follows an IRAP sub-picture having the same value of nuh_layer_id and the same value of sub-picture index and, if any, the leading sub-picture associated with the IRAP sub-picture in both the decoding order and the output order, it is assumed that there is no picture referred to by an entry in RefPicList[0] or RefPicList[1] that precedes the picture containing the IRAP sub-picture in the output order or the decoding order. (20) When the current picture is a RADL picture, it is assumed that there is no active entry in either RefPicList[0] or RefPicList[1] that meets any of the following conditions. (a) The RASL picture of pps_mixed_nalu_types_in_pic_flag is 0. This means that the active entry of the RPL of the RADL picture can refer to a RASL picture having pps_mixed_nalu_types_in_pic_flag equal to 1. However, the RADL picture is subject to the following constraint that does not permit a RADL subpicture to refer to a RASL subpicture, so as to refer only to the RADL subpictures within the referred RASL picture. Since the RADL subpictures within the referred RASL picture are correctly decoded, when decoding starts from the associated CRA picture, such RADL pictures can still be correctly decoded. (b) A picture preceding in the decoding order of the associated IRAP picture. (21) When the current subpicture having nuh_layer_id equal to a specific value layerId and a subpicture index equal to a specific value subpicIdx is a RADL subpicture, assume that there is no active entry in either RefPicList[0] or RefPicList[1]. (a) The picture with nuh_layer_id equal to layerId includes a RASL subpicture having a subpicture index equal to subpicIdx. (b) A picture preceding in the decoding order of the picture including the associated IRAP subpicture. (22) The following constraints apply to the picture referred to by each ILRP entry if it exists in RefPicList[0] or RefPicList[1] of the slice of the current picture. (a) Assume that the picture is in the same AU as the current picture. (b) Assume that the picture exists in the DPB. (c) Assume that the picture has a refPicLayerId of nuh_layer_id smaller than that of the current picture. (d)Any of the following constraints shall apply: the picture shall be an IRAP picture; the picture shall have a TemporalId less than or equal to Max(0, vps_max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx] - 1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively. (23)If present in RefPicList[0] or RefPicList[1] of a slice, each ILRP entry shall be an active entry. (24)If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 0 and sps_num_subpics_minus1 is greater than 0, then either (but not both) of the following two conditions shall be true. (a)The picture referenced by each active entry in RefPicList[0] or RefPicList[1] shall have the same sub-picture layout as the current picture (i.e., the SPSs referenced by that picture and the current picture shall have the same value of sps_num_subpics_minus1, and the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], and sps_subpic_height_minus1[j] for each value of j in the range from 0 to sps_num_subpics_minus1). (b)The picture referenced by each active entry in RefPicList[0] or RefPicList[1] shall be an ILRP with a value of sps_num_subpics_minus1 equal to 0.
[0321] According to an embodiment, the decoder can check whether the above constraints are satisfied regarding bitstream compatibility. The compatibility check can be performed after calling a decoding process for generating an unavailable reference picture, as described above.
[0322] Embodiments of the present disclosure can include at least one processor and a memory that stores computer code. The computer code may be configured to cause the at least one processor to perform the functions of the embodiments of the present disclosure when executed by the at least one processor.
[0323] For example, referring to FIG. 33, the decoder of the present disclosure can include at least one processor and a memory that stores computer code (80). The decoder can be configured to receive a bitstream including at least one encoded picture and parameter sets (e.g., SPS and VPS), headers (e.g., picture headers and slice headers), and AU delimiters. This computer code can be configured to cause the at least one processor to perform any number of decoding processes (e.g., construction of a reference picture list, generation of an unavailable reference picture, and confirmation of bitstream compatibility) and aspects related to decoding (e.g., signaling of flags and other syntax elements in picture headers, slice headers, and access unit delimiters) as described in the present disclosure. For example, the computer code (80) can include a plurality of signaling codes (81) and decoding codes (82).
[0324] The plurality of signaling codes (81) can include various signaling codes configured to cause the at least one processor to signal (and / or infer) flags and other syntax elements in picture headers, slice headers, and access unit delimiters.
[0325] The decoding code (82) can be configured to cause at least one processor to decode one or more pictures. According to an embodiment, the decoding code (82) may include a construction code (83), a generation code (84), and a verification code (85). The construction code (83) can be configured to cause at least one processor to construct a reference picture list. The generation code (84) can be configured to cause at least one processor to generate unavailable reference pictures in the reference picture list. The verification code (85) can be configured to cause at least one processor to check the bitstream compliance for the reference pictures in the reference picture list with respect to the following constraints, namely, (a) the number of entries indicated as being in the reference picture list is greater than or equal to the number of active entries indicated as being in the reference picture list, (b) each picture referenced by an active entry in the reference picture list exists in the decoded picture buffer (DPB) and has a time identifier value less than or equal to the time identifier value of the current picture, and (c) each picture referenced by an entry in the reference picture list is indicated by a picture - header flag indicating that it may be a reference picture rather than the current picture.
[0326] The above - described technology can be implemented as computer software physically stored on one or more computer - readable media using computer - readable instructions. For example, FIG. 34 shows a computer system (900) suitable for implementing an embodiment of the disclosed subject matter.
[0327] The computer software can be encoded using any suitable machine code or computer language and be the subject of assembly, compilation, linking, or similar mechanisms to create code that includes instructions executable directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0328] The command can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, and the like.
[0329] The components shown in FIG. 34 for the computer system (900) are illustrative in nature and are not intended to suggest any limitation with respect to the use or functionality of the computer software implementing the embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement on any one or combination of the components shown in the exemplary embodiments of the computer system (900).
[0330] The computer system (900) may include a specific human interface input device. Such a human interface input device can respond to input by one or more users, for example, tactile input (keystrokes, swipes, movement of a data glove, etc.), voice input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown), etc. Using a human interface device, it is also possible to capture specific media that is not necessarily directly related to conscious human input, such as sound (speech, music, ambient sound, etc.), images (scanned images, picture images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).
[0331] The input human interface device may include one or more of a keyboard 901, a mouse 902, a trackpad 903, a touch screen 910, a data glove, a joystick 905, a microphone 906, a scanner 907, and a camera 908 (only one of each shown in the figure).
[0332] The computer system (900) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, by tactile output, sound, light, and smell / taste. Such human interface output devices can include tactile output devices (e.g., tactile feedback by a touch screen (910), data glove, or joystick (905), although there may also be tactile feedback devices that do not function as input devices). For example, such devices can include audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including CRT screens, LCD screens, plasma screens, OLED screens, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of which may be able to output two-dimensional visual output or three-dimensional immersive output via means such as stereo output, screens, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0333] The computer system (900) can also include memory devices accessible by humans, and related media such as optical media (921) such as CD / DVD ROM / RW (920) including CDs / DVDs, thumb drives (922), removable hard drives or solid state drives (923), legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0334] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, carrier wave, or other transient signal.
[0335] The computer system (900) may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, Wi-Fi, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial such as CANBus, etc. In certain networks, generally, an external network interface adapter connected to a specific general-purpose data port or peripheral bus (949) (such as a USB port of the computer system (900)) is required, and others are generally integrated into the core of the computer system 900 by connecting to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (900) can communicate with other entities. Such communication can be unidirectional, receive only (such as broadcast TV), transmit only unidirectional (such as from a CANbus to a specific CANbus device), or bidirectional, for example, communication to another computer system using a local area digital network or a wide area digital network. Such communication can include communication to a cloud computing environment (955). Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.
[0336] The aforementioned human interface device, human-accessible storage device, and network interface (954) can be connected to the core (940) of the computer system (900).
[0337] The core (940) can include special programmable processing devices in the form of one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), field programmable gate arrays (FPGAs) (943), hardware accelerators for specific tasks (944), etc. These devices can be connected via a system bus (948) together with a read-only memory (ROM) (945), a random access memory (946), internal mass storage devices such as internal hard drives and SSDs that are not accessible to the user (947). In some computer systems, access can be made to the system bus (948) in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the system bus of the core (948) or via a peripheral bus (949). Architectures of peripheral buses include PCI, USB, etc. The graphics adapter 950 may be included in the core 940.
[0338] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute specific instructions that can together constitute the aforementioned computer code. The computer code can be stored in the ROM (945) or the RAM (946). Migration data can also be stored in the RAM (946), while persistent data can be stored, for example, in the internal mass storage device (947). By using cache memory that can be closely associated with one or more CPUs (941), GPUs (942), mass storage devices (947), ROM (945), RAM (946), etc., fast storage and reading for any memory device becomes possible.
[0339] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or may be of the kind well-known and available to those having skill in the art of computer software technology.
[0340] By way of example and not limitation, a computer system having an architecture (900), particularly a core (940), can provide functionality as a result of software executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.) incorporated in one or more tangible computer-readable media. Such computer-readable media can be associated with the mass storage devices accessible to the user introduced above, and specific storage devices of the core (940) having a non-transitory nature such as the on-core mass storage device (947) and ROM (945). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (940). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (940), particularly the processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM (946) and modify such data structures according to the processes defined by the software, and can execute the specific processes or specific parts of the specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic embodied in circuitry (e.g., an accelerator (944)) that is hardwired or otherwise, and can operate instead of or in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. References to software can include logic, and vice versa as appropriate. References to computer-readable media can, as needed, encompass circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0341] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents within the scope of the present disclosure. Thus, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope, even though not explicitly shown or described herein.
Explanation of Signs
[0342] 80 Computer code 81 Signaling code 82 Decoding code 83 Construction code 84 Generation code 85 Verification code 100 Communication system 110 Terminal 120 Terminal 130 Terminal 140 Terminal 150 Network 200 Streaming system 201 Video source 202 Uncompressed video sample stream 203 Encoder 204 Video bitstream 205 Streaming server 206 Streaming client 209 Video bitstream 210 Video decoder 211 Output video sample stream 212 Display 310 Receiver 312 Channel 315 Buffer memory 320 Parser 321 Symbol 351 Scaler / inverse transform 352 Intra-picture prediction 353 Motion compensation prediction unit 356 Loop filter 357 Reference Picture Memory 358 Current Picture 430 Source Encoder 432 Encoding Engine 433 Decoder 434 Reference Picture Memory 435 Predictor 440 Transmitter 445 Entropy Encoder 450 Controller 460 Channel 501 Picture Header 502 ARC Information 503 ARC (Warping Coordinates) 504 Picture Parameter Set 505 ARC Reference Information 506 Table 507 Sequence Parameter Set 508 Header 509 ARC Information 511 Adaptive Parameter Set 512 ARC Information 513 ARC Reference Information 514 Tile Group Header 515 ARC Information 600 Tile Group Header 516 Parameter Set 602 Syntax Element dec_pic_size_idx 603 Adaptive Resolution 610 Sequence Parameter Set 611 Syntax Element 613 Output Resolution 615 Reference Picture Dimension 617 Table Entry 681 I Slice 682 B Slice 683 B Slice 684 B Slice 685 B Slice 686 B Slice 702 Base Layer 712 Base Layer 704 Extended CSPS Layer 714 Extended CSPS Layer 742 Spherical 360 Picture 745 Sub-Picture 752 Sub-Picture 746 Extended Layer of Sub-Picture 762 Upper-Left Sub-Region 763 Upper-Right Sub-Region 764 Lower-Left Sub-Region 765 Lower-Right Sub-Region 941 CPU 942 GPU 943 Field-Programmable Gate Array (FPGA) 944 Hardware Accelerator 948 System Bus 950 Graphics Adapter 954 Network Interface
Claims
1. A method executed by at least one processor, comprising: receiving an encoded video stream including access units including pictures; signaling a first flag at an access unit delimiter of the encoded video stream indicating whether the access unit includes one of an Intra Random Access Point (IRAP) picture and a Gradual Decoding Refresh (GDR) picture; signaling a second flag in a picture header of the encoded video stream indicating whether the picture is the GDR picture; and decoding the picture as a current picture based on the signaling of the first flag and the second flag; The value of the first flag and the value of the second flag are matched such that when a flag used for outputting and removing a picture from a decoded picture buffer does not exist and the value of the first flag is equal to 1, the value of the second flag is required to be equal to 1; method.
2. A method executed by at least one processor, comprising: receiving an encoded video stream including access units including pictures; signaling a first flag at an access unit delimiter of the encoded video stream indicating whether the access unit includes one of an Intra Random Access Point (IRAP) picture and a Gradual Decoding Refresh (GDR) picture; signaling a second flag in a picture header of the encoded video stream indicating whether the picture is the GDR picture; and decoding the picture as a current picture based on the signaling of the first flag and the second flag; when the first flag is present in the access unit delimiter, the value of the second flag is set equal to the value of the first flag. method.
3. The first flag has a value indicating that the picture is either the IRAP picture or the GDR picture, the second flag has a value indicating that the picture is a GDR picture; The method further comprises the step of signaling a third flag in a slice header of a slice of a picture of the coded video stream indicating whether any pictures before the IRAP picture are output. The method according to claim 1 or 2.
4. The method further includes determining a network abstraction layer (NAL) unit type of the slice, the third flag is signaled based on the determined NAL unit type. The method of claim 3.
5. The method of claim 4, wherein the third flag is signaled based on the NAL unit type being determined to be equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT.
6. The step of decrypting comprises: building a reference picture list; generating unavailable reference pictures in the reference picture list; checking for bitstream conformance for reference pictures in the reference picture list, subject to the following constraints: the number of entries indicated to be in the reference picture list is greater than or equal to the number of active entries indicated to be in the reference picture list; each picture referenced by an active entry in the reference picture list is present in a decoded picture buffer (DPB) and has a temporal identifier value less than or equal to a temporal identifier value of the current picture; each picture referenced by an entry in the reference picture list is not the current picture and is indicated as a possible reference picture by a picture header flag; Steps and The method according to any one of claims 1 to 5, comprising:
7. The method described in claim 6, wherein the step of verifying bitstream conformance is performed based on a determination that the current picture is an independent decoder refresh (IDR) picture, a clean random access (CRA) picture, or a gradual decoding refresh (GDR) picture.
8. An apparatus configured to perform a method according to any one of claims 1 to 7.
9. A computer program product for causing at least one processor to carry out a method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-layer video coding method for random access and device therefor, and multi-layer video decoding method for random access and device therefor
US20160044309A1
Reference picture management in video coding
WO2020037272A1