Method for output layer set mode in multilayered video stream
Patent Information
- Application Number
- JP2025035576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-08-21
- Filing Date
- 2025-03-06
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2040-11-09
AI Technical Summary
Existing video coding technologies struggle with efficiently signaling adaptive picture sizes within video bitstreams, particularly in next-generation codecs like Versatile Video Coding (VVC), which require advanced techniques for inter and intra prediction.
The proposed solution involves a method for decoding video bitstreams that includes parsing an output layer set mode indicator from the video parameter set (VPS), identifying output layer set signaling based on this indicator, and determining the picture output layers for decoding. This method allows for adaptive picture size signaling by specifying which layers are output, enabling more flexible and efficient video coding.
This approach enhances the flexibility and efficiency of video coding by allowing for adaptive picture size signaling, which can improve compression efficiency and adapt to varying content requirements in next-generation video codecs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] [Related Applications] This application claims priority to U.S. Provisional Application No. 63 / 001,045, filed Mar. 27, 2020, and U.S. Patent Application No. 17 / 000,018, filed Aug. 27, 2020, which are hereby incorporated by reference in their entirety.
[0002] [Technical Field] The present disclosure relates to video compression techniques and inter and intra prediction in advanced video codecs. In particular, the present disclosure relates to next-generation video coding technologies including video coding / decoding technologies subsequent to High Efficiency Video Coding (HEVC) such as Versatile Video Coding (VVC). More specifically, aspects of the present disclosure are directed to methods, apparatuses, and computer-readable media that provide a set of output layer derivations designed by advanced video coding techniques within a coded video stream having multiple layers.
Background Art
[0003] Video coding and decoding using interpicture or intrapicture prediction along with motion compensation has been known for decades. Uncompressed digital video can be composed of a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate, for example 60 pictures per second or 60 Hz (also known colloquially as the frame rate). Uncompressed video has significant bitrate requirements. For example, 8-bit / sample 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth close to 1.5 Gbit / s. One hour of such video may require more than 600 Gbytes of storage space.
[0004] One purpose of video coding and decoding can be the reduction of redundancy in an input video signal through compression. Compression can, in some cases, help reduce the aforementioned bandwidth or storage space requirements by more than an order of magnitude. Both lossy and lossless compression, and combinations thereof, are available. Lossless compression represents techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal is not identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to produce a reconstructed signal useful for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion depends on the application. For example, a user of a particular consumer streaming application may tolerate higher distortion than a user of a television broadcast application. The achievable compression ratio can reflect that the higher the acceptable / tolerable distortion, the higher the achievable compression ratio.
[0005] Video encoders and video decoders can utilize techniques from several broad classifications, including, for example, motion compensation, transformation, quantization, and entropy coding. Some of these are introduced below.
[0006] Historically, video encoders and decoders have tended to operate at a given picture size that is defined for and remains constant for a coding video sequence (CVS), a Group of Pictures (GOP), or a similar multi-picture temporal frame. For example, in the Moving Picture Experts Group (MPEG)-2, the system design is known to change the horizontal resolution (and thus the picture size) not only for intra-frames (or I-frames, or I-pictures), but also, and thus typically for GOPs, depending on factors such as the activity of the scene. Resampling of reference pictures for the use of different resolutions within a CVS is known, for example, from ITU-T Rec. H.263 Annex P. However, here, since the picture size does not change, only the reference pictures are resampled, and as a result, only parts of the picture canvas may be used (in the case of downsampling), or only parts of the scene may be captured (in the case of upsampling). Further, H.263 Annex Q allows resampling of individual macroblocks by a factor of two (in each dimension) either up or down. Again, the picture size remains the same. In H.263, the size of the macroblocks is fixed and thus does not need to be signaled.
[0007] Changes in picture size in predictive pictures have become more mainstream in recent video coding. For example, VP9 allows reference picture resampling and changes in the resolution of the entire picture. Similarly, certain proposals targeting Versatile Video Coding (VVC) (e.g., including Hendry, et. al, “On adaptive resolution change )(ARC) for VVC”, Joint Video Team document JVET-M0135-v1, Jan9-19, 2019, which is incorporated herein by reference in its entirety) allow resampling of the entire reference picture to a different - higher or lower - resolution. In Hendry, different candidate resolutions are proposed that should be referenced by syntax elements for each picture in the picture parameter set coded in the sequence parameter set.
Summary of the Invention
[0008] Techniques for signaling adaptive picture sizes in video bitstreams according to various embodiments are disclosed.
[0009] According to an aspect of the present disclosure, a decoding method may include: receiving a bitstream including compressed video / image data, the bitstream including a plurality of layers; parsing or deriving an output layer set mode indicator within a video parameter set (VPS) from the bitstream; identifying output layer set signaling based on the output layer set mode indicator; and identifying one or more picture output layers based on the identified output layer set signaling; decoding the identified one or more picture output layers. may be included.
[0010] The step of identifying the output layer set signaling based on the output layer set mode indicator comprises: When the output layer set mode indicator in the VPS has a first value, identifying the highest layer in the bitstream as the one or more picture output layers; When the output layer set mode indicator in the VPS has a second value, identifying all layers in the bitstream as the one or more picture output layers; When the output layer set mode indicator in the VPS has a third value, identifying the one or more picture output layers based on explicit signaling in the VPS; and may include.
[0011] The first value may be different from the second value and different from the third value, and the second value may be different from the third value.
[0012] The first value may be 0, the second value may be 1, and the third value may be 2. However, other values may be used, and the present disclosure is not limited to the use of 0, 1, and 2 as described above.
[0013] The step of identifying the one or more picture output layers based on explicit signaling in the VPS may include: (i) parsing or deriving an output layer flag from the VPS; and (ii) setting the layer having the output layer flag equal to 1 as the one or more picture output layers.
[0014] The step of identifying the output layer set signaling based on the output layer set mode indicator is: When the output layer set mode indicator in the VPS has a predetermined value, the output layer set signaling may include the step of identifying the one or more picture output layers based on explicit signaling in the VPS.
[0015] The step of identifying the one or more picture output layers based on the explicit signaling within the VPS includes: (i) parsing or deriving an output layer flag from the VPS; and (ii) setting a layer having the output layer flag equal to 1 as the one or more picture output layers, wherein the number of the plurality of layers is greater than 2.
[0016] The output layer set signaling may include a step of identifying the one or more picture output layers based on the explicit signaling within the VPS when the output layer set mode indicator is equal to 2 and the number of the plurality of layers is greater than 2.
[0017] The output layer set signaling may include a step of identifying the highest layer within the bitstream or all layers within the bitstream as the one or more picture output layers when the output layer set mode indicator is less than 2 and the number of the plurality of layers is 2. The output layer set mode indicator is actually less than 2, and the number of the plurality of layers is actually 2.
[0018] The output layer set number - 1 indicator within the VPS indicates the number of the output layers.
[0019] According to an embodiment, the VPS maximum layer - 1 indicator within the VPS indicates the number of layers within the bitstream.
[0020] According to an embodiment, the output layer set mode flag [i][j] within the VPS indicates whether the j - th layer of the i - th output layer set is an output layer.
[0021] According to an embodiment, when the plurality of layers are independent layers and the VPS all independent layer flag of the VPS is equal to 1, the output layer set mode indicator is not signaled, and the value of the output layer set mode indicator is presumed to be the second value.
[0022] According to an embodiment, when each layer is an output layer set, regardless of the value of the output layer set mode indicator, the picture output flag of the VPS is set equal to the picture output flag signaled in the picture header.
[0023] Note: A picture in an output layer may or may not have a PictureOutputFlag equal to 1. A picture in a non-output layer has a PictureOutputFlag equal to 0. A picture having a PictureOutputFlag equal to 1 is output for display. A picture having a PictureOutputFlag equal to 0 is not output for display.
[0024] According to an embodiment, when the sequence parameter set (SPS) VSP identifier is greater than 0 and indicates that more than one layer is present in the bitstream, the picture output flag is set equal to 0. When each layer has the output layer set mode flag of the VPS equal to 0 and indicates that the plurality of layers in the bitstream are not all independent, the output layer set mode indicator is equal to 0, and the current access unit includes a picture that satisfies all of the following conditions: having a picture output flag equal to 1, and having a nuh layer identifier belonging to the output layer of the output layer set that is greater than that of the current picture.
[0025] According to an embodiment, when the sequence parameter set (SPS) of the VPS is greater than 0, the picture output flag of the VPS is set equal to 0, each layer has the output layer set flag equal to 0, the output layer set mode indicator is equal to 2, and the output layer set output layer flag [Target OLS Index][General Layer Index [nuh layer identifier]] is equal to 0.
[0026] According to an embodiment, the method may further include controlling a display to display the decoded one or more picture output layers.
[0027] According to an aspect of the present disclosure, a non-transitory computer-readable storage medium storing instructions, which when executed, cause a system or apparatus including one or more processors to receive a bitstream including compressed video / image data, the bitstream including a plurality of layers, parse or derive an output layer set mode indicator in a video parameter set (VPS) from the bitstream, identify output layer set signaling based on the output layer set mode indicator, and identify one or more picture output layers based on the identified output layer set signaling, decode the identified one or more picture output layers, non-transitory computer-readable storage medium.
[0028] According to an embodiment, the instructions are further configured to cause the system or apparatus including the one or more processors to control a display to display the decoded one or more picture output layers.
[0029] According to an aspect of the present disclosure, a device may include at least one memory storing computer program code, and at least one processor configured to access the at least one memory and operate according to the computer program code. According to an embodiment, the computer program code is reception code configured to cause the at least one processor to receive a bitstream including compressed video / image data, the bitstream including a plurality of layers, parse or derive code configured to cause the at least one processor to parse or derive an output layer set mode indicator within a video parameter set (VPS) from the bitstream; output layer signaling identification code configured to cause the at least one processor to identify output layer set signaling based on the output layer set mode indicator; picture output layer identification code configured to cause the at least one processor to identify one or more picture output layers based on the identified output layer set signaling; decoding code configured to cause the at least one processor to decode the identified one or more picture output layers; may include.
[0030] According to an embodiment, the computer program code may further include display control code configured to cause the at least one processor to display the one or more picture output layers.
[0031] According to an aspect of the present disclosure, a method for signaling an adaptive picture size within a video bitstream includes: receiving a bitstream comprising compressed video / image data, the bitstream including a plurality of layers; identifying a background region and one or more foreground sub-pictures; determining whether a particular sub-picture region is selected; generating an inverse quantized block corresponding to the selected sub-picture region by a process including parsing the bitstream, decoding the entropy-coded bitstream, and inverse quantizing corresponding blocks based on determining that the particular sub-picture region is selected; may include.
[0032] The method may further include decoding and displaying the background region when the specific sub-picture region is not selected.
[0033] The bitstream may include a syntax element that specifies which layer can be output on the decoder side.
[0034] The syntax element may include a picture header that includes a variable-length Exp-Golomb coded syntax element.
[0035] The method may further include determining whether an adaptive resolution is being used for a picture or a portion thereof based on signaling in the sequence parameters.
[0036] The step of determining whether an adaptive resolution is being used for a picture or a portion thereof may include determining whether a first syntax element, which is a flag, indicates the use of an adaptive resolution.
[0037] The method may include instructing the decoder to use a specific reference picture size, rather than implicitly assuming that the size is the output picture size, by using a flag that controls the conditional presence of the reference picture dimensions.
[0038] The syntax element may include a table of possible decoded picture widths and heights.
[0039] According to an embodiment, the value in the Network Abstraction Layer (NAL) unit header may be used to indicate not only time but also spatial layers.
[0040] The value in the NAL unit header may be a Temporal Identifier (ID) field.
[0041] The method may further include, for example, in an extensible environment, using, without modification, existing selected forwarding units (SFUs) that are generated and optimized for temporal layer selection forwarding based on the NAL unit header Temporal ID value.
[0042] The method may further include mapping between the coding picture size and the temporal layer indicated by the Temporal ID field in the NAL unit header.
[0043] The method receiving additional data along with the encoded video, the additional data being included as part of the coding video sequence, and further using the additional data to accurately decode the data and / or more accurately reconstruct the original video data.
[0044] The additional data may be in one or more forms of temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes.
[0045] According to an aspect, a non-transitory computer-readable storage medium storing instructions that, when executed, cause a system or apparatus including one or more processors to receive a bitstream composed of compressed video / image data, the bitstream including a plurality of layers, identify a background region and one or more foreground sub-pictures, determine whether a particular sub-picture region is selected, and based on determining that a particular sub-picture region is selected, generate an inverse quantized block corresponding to the selected sub-picture region by a process including parsing the bitstream, decoding the entropy-coded bitstream, and inverse quantizing corresponding blocks. Non-transitory computer-readable storage medium.
[0046] According to an embodiment, a device includes at least one memory configured to store computer program code, and at least one processor configured to access the at least one memory and operate according to the computer program code. The computer program code includes: Receiving code configured to cause the at least one or more processors to receive a bitstream composed of compressed video / image data, the bitstream including a plurality of layers; Identification code configured to cause the at least one or more processors to identify a background area and one or more foreground sub-pictures; Determination code configured to cause the at least one or more processors to determine whether a specific sub-picture area has been selected; Generation code configured to cause the at least one or more processors to generate an inverse quantized block corresponding to the selected sub-picture by a process including, but not limited to: parsing the bitstream, decoding an entropy-coded bitstream, and inverse quantizing a corresponding block, based on a determination that the specific sub-picture area has been selected. A device including the above. Brief Description of the Drawings
[0047] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0048]
Figure 1
[0049]
Figure 2
[0050]
Figure 3
[0051]
Figure 4
[0052]
Figure 5
[0053]
Figure 6
[0054]
Figure 7
[0055]
Figure 8
[0056]
Figure 9
[0057]
Figure 10
[0058]
Figure 11
[0059]
Figure 12
[0060]
Figure 13
[0061]
Figure 14
[0062]
Figure 15
[0063]
Figure 16
[0064]
Figure 17
[0065]
Figure 18
[0066]
Figure 19
[0067]
Figure 20
[0068]
Figure 21
[0069]
Figure 22
[0070]
Figure 23
[0071]
Figure 24
[0072]
Figure 25A
Figure 25B
Figure 25C
Best Mode for Carrying Out the Invention
[0073] When a picture is encoded into a bitstream composed of multiple layers having different qualities, the bitstream may have a syntax element that specifies which layer may be output on the decoder side. The set of layers to be output is defined as the output layer set. In the latest video codecs that support multiple layers and scalability, one or more output layer sets are signaled in the video parameter set. The syntax elements that specify the output layer sets and their dependencies, profile / tier / level, and virtual decoder reference model parameters need to be signaled efficiently in the parameter set.
[0074] Embodiments of the present disclosure solve one or more problems of the related art.
[0075] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In one-way data transmission, the first terminal (110) may code video data at a local location for transmission to another terminal (120) via the network (150). The second terminal (120) may receive the coded video data of the other terminal from the network (150), decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0076] FIG. 1 shows a second terminal pair (130, 140) applied to support two-way transmission of coded video, which may occur, for example, during a video conference. In two-way data transmission, each terminal (130, 140) may code locally captured video data for transmission to the other terminal via the network (150). Each terminal 130, 140 may also receive the coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.
[0077] In FIG. 1, the terminals (110 - 140) are shown as laptop 110, personal computer (PC) 120, and mobile terminals 130 and 140, but the terminals (110 - 140) are not so limited, and the terminals (110 - 140) may correspond to one or more or any combination of servers, personal computers, mobile devices, tablets, and smartphones. Embodiments of the present disclosure have applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing facilities. The network (150) represents any number of networks that carry coded video data among the terminals (110 - 140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit-switching and / or packet-switching channels. Representative networks include electronic communication networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of the discussion of the present invention, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise specifically noted hereinafter.
[0078] FIG. 2 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of the application of an embodiment of the disclosure. The present disclosure is equally applicable to, for example, the storage of compressed video on digital media including video conferencing, digital television (TV), compact discs (CDs), digital versatile discs (DVDs), memory sticks, etc., and other video-enabled applications, etc.
[0079] A streaming system may include a video source (201) configured to generate, for example, an uncompressed video sample stream (202), and a capture subsystem (213) that may include, for example, a digital camera. The sample stream (202) is shown in thick lines in FIG. 2 to emphasize its high data volume when compared to an encoded video bitstream, and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) includes hardware, software, or a combination thereof, and can enable or implement aspects of the disclosed subject matter as detailed below. The encoded video bitstream (204) is shown in thin lines in FIG. 2 to emphasize its low data volume when compared to the sample stream (202), and can be stored in a streaming server (205) for future use. One or more streaming clients (206, 208) can access the streaming server (205) to read copies (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210). The video decoder (310) decodes an incoming copy of the encoded bitstream (207) and generates an output video sample stream (211) that can be rendered on a display (212) or other rendering device. In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A video coding standard under development is known informally as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.
[0080] FIG. 3 may be a functional block diagram of a video decoder (210) according to an embodiment of the present disclosure.
[0081] The receiver (310) may receive one or more coding video sequences to be decoded by the video decoder (210), or in the same or another embodiment, one coding video sequence at a time. Here, the decoding of each coding video sequence is independent of other coding video sequences. The coding video sequence may be received from a channel (312) which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data together with other data, such as coded audio data and / or an auxiliary data stream that may be transferred to respective using entities (not shown). The receiver (310) may separate the coding video sequence from the other data. To remove network jitter, a buffer memory (315) may be connected between the receiver (310) and the entropy decoder / parser (320) (hereinafter, "parser"). When the receiver (310) is receiving data controllably from a storage / transfer device with sufficient bandwidth or from an isosynchronous network, the buffer (315) may not be necessary or may be made small. When used in a best-effort packet network such as the Internet, the buffer (315) may be necessary and may be made relatively large and advantageously adaptable in size.
[0082] The video decoder (210) may include a parser (320) to reconstruct symbols (321) from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder 210 and, optionally, information for controlling a rendering device such as a display 212 that is not part of the integrated portion of the decoder but can be coupled to the decoder as shown in FIG. 2. The control information for the rendering device may be in the form of an SEI (Supplementary Enhancement Information) message or a VUI (Video Usability Information) parameter set fragment. The parser (320) may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser 320 may extract a set of subgroup parameters from the coded video sequence based on at least one parameter corresponding to at least one subgroup of pixels in the video decoder. Subgroups can include GOP (Groups of Picture), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0083] The parser (320) may perform an entropy decoding / parsing operation on the video sequence received from the buffer (315) to generate symbols (321).
[0084] The reconstruction of symbol (321) may include multiple different units depending on the type of the coded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is included can be controlled by subgroup control information parsed from the coded video sequence by the parser (320). The subgroup control information may flow between the parser (320) and the multiple units.
[0085] Beyond the function blocks already mentioned, the decoder (210) can be conceptually subdivided into a number of functional units, as will be described later. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0086] The first unit is the scaler / inverse transform unit 351. The scaler / inverse transform unit (351) receives, as symbols (321) from the parser (320), the quantized transform coefficients and control information including which transform should be used, block size, quantization coefficients, quantization scaling matrix, etc. It can output a block including sample values that can be input to the aggregator (355).
[0087] In some examples, the output samples of the scaler / inverse transform unit (351) can belong to an intra-coding block, i.e., a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current (partially already reconstructed) picture (358). The aggregator (355) in some cases adds, sample by sample, the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0088] In other cases, the output samples of the scaler / inverse transform unit (351) can be related to inter-coded, possibly motion-compensated blocks. In such cases, the motion-compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbol (321) associated with the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit to generate the output sample information (in this case, called the residual samples or residual signal). The address in the reference picture memory from which the motion-compensation prediction unit fetches the prediction samples can be controlled by the available motion vectors of the motion-compensation prediction unit, in the form of a symbol (321) that can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory when an exact motion vector of sub-samples is in use, a motion vector prediction mechanism, etc.
[0089] The output samples of the aggregator (355) may undergo various loop filtering techniques in the loop filter unit (356). The video compression technique is controlled by parameters included in the coded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but also responds to meta information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, and may include in-loop filtering techniques that can also respond to previously reconstructed and loop-filtered sample values.
[0090] The output of the loop filter unit (356) can be an output sample stream that can be output to the rendering device (212) and stored in the reference picture memory (356) for use in future inter-picture prediction.
[0091] Once a particular coded picture is completely reconstructed, it can be used as a reference picture for future prediction. When the coded picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser 320), the current reference picture 356 can become part of the reference picture memory 357, and a fresh current picture memory can be reallocated before starting the reconstruction of subsequent coded pictures.
[0092] The video decoder (parser) 320 may perform a decoding operation in accordance with a predetermined video compression technique that may be defined by a standard such as ITU-T Rec H.265. The coding video sequence may conform to the syntax of the video compression technique or standard, specifically in the profile document therein, in the sense that the coding video sequence conforms to the syntax specified by the video compression technique or standard in use. Also, what may be necessary for compliance is that the complexity of the coding video sequence be within the limits defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the HDR (Hypothetical Reference Decoder) specification and metadata for HDR buffer management signaled in the coding video sequence.
[0093] In an embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coding video sequence. The additional data may be used by the video decoder 320 for correctly decoding the data and / or for more accurately reconstructing the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0094] FIG. 4 may be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0095] The encoder (203) may receive video samples from a video source (201) (not part of the encoder) that may capture a video image to be coded by the encoder (203).
[0096] The video source (201) may provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCb, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media providing system, the video source (201) may be a storage device storing previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may subsequently be provided as a plurality of individual pictures that give motion when viewed in succession. Each picture itself may be organized as a spatial array of pixels. Each pixel may include one or more samples depending on the sampling structure, color space, etc. in use. A person skilled in the art can immediately understand the relationship between a pixel and a sample. The following description focuses on samples.
[0097] According to an embodiment, the encoder (203) may code and compress pictures of the source video sequence into a coded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate coding speed is one function of the control unit (450). The control unit (450) may control other functional units as will be described later and is functionally coupled to other functional units. It may include parameters set by the control unit (450), rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. A person skilled in the art can immediately identify other functions of the control unit 450 when it may be related to a video encoder (203) optimized for a specific system design.
[0098] Some video encoders operate in what those skilled in the art would immediately recognize as a “coding loop.” As a very simplified explanation, the coding loop can include an encoder (203) (hereinafter, the “source coder”) that generates symbols based on an input picture and a reference picture to be coded, and an (local) decoder (433) incorporated within the encoder (203) that reconstructs the symbols to generate sample data that a (remote) decoder could generate when any compression between the symbols and the coded video bitstream is lossless within the video compression techniques contemplated by the subject matter of the disclosure. The reconstructed sample stream is input into the reference picture memory 434. When the decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents of the reference picture buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder “sees” exactly the same sample values as the decoder would “see” as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0099] The operation of the “local” decoder (433) can be the same as that of the “remote” decoder (210) detailed above in connection with FIG. 3. Briefly referring also to FIG. 3, however, since symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy encoder (445) and the parser (320) can be lossless, the entropy decoding portion of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) need not be fully implemented in the local decoder (433).
[0100] The consideration made in this regard is that any decoder technology other than the parse / entropy decoding existing in the decoder needs to exist in substantially the same functional form as that in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be omitted because they are the reverse of the decoder technologies that are comprehensively described. More detailed descriptions are needed only in specific areas and are provided below.
[0101] During operation, in some examples, the source coder (203) may perform motion-compensated predictive coding. This predictively codes the input frame by referring to one or more previously coded frames from the video sequence designated as the "reference frame". In this method, the coding engine (432) codes the difference between the pixel block of the input frame and the pixel block of the reference frame that may be selected as the prediction criterion for the input frame.
[0102] The local video decoder (433) may decode the coded video data of the frame that may be designated as the reference frame based on the symbols generated by the source coder (430). The operation of the coding engine (432) may advantageously be a lossy process. When the coded video data can be decoded in a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) replicates the decoding process that may be performed by the video decoder for the reference frame, resulting in a reconstructed reference frame to be stored in the reference picture cache (434). Thus, the encoder (203) may locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by the remote video decoder (if there are no transmission errors).
[0103] Predictor (435) may perform predictive search for the coding engine (432). That is, for a new frame to be coded, predictor (435) may search the reference picture memory (434) for sample data (such as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as appropriate prediction criteria for the new picture. Predictor (435) may operate on a sample block - pixel block basis to find appropriate prediction criteria. In some examples, the input picture may have prediction criteria drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search results obtained by predictor (435).
[0104] Control unit (450) may manage the coding operations of video coder (203), including, for example, setting parameters and subgroup parameters used for encoding video data.
[0105] Outputs of all the aforementioned functional units may undergo entropy coding in entropy coder (445). The entropy coder converts symbols generated by various functional units into a coded video sequence by losslessly compressing the symbols according to techniques well known to those skilled in the art, such as Huffman coding, variable - length coding, arithmetic coding, etc.
[0106] Transmitter (440) may buffer the coded video sequence generated by entropy coder (445) for transmission via communication channel (460), which may be a hardware / software link to a storage device capable of storing the coded video data. Transmitter (440) may merge the coded video data from video coder (430) with other data to be transmitted, such as coded audio data and / or auxiliary data streams / sources.
[0107] The control unit (450) may manage the operation of the encoder (203). During coding, the control unit (450) may assign a specific coding picture type that can affect the coding technique applicable to each picture to each coding picture. For example, a picture may often be assigned as one of the following picture types.
[0108] An intra picture (I picture) may be a picture that can be coded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example, IDR (Independent Decoder Refresh) pictures. Those skilled in the art recognize the variations of I pictures and their individual applications and characteristics.
[0109] A predicted picture (P picture) may be a picture that can be coded and decoded using intra prediction or inter prediction, typically using one motion vector and a reference index to predict the sample values of each block.
[0110] A bi - directionally predicted picture (B picture) may be a picture that can be coded and decoded using intra prediction or inter prediction, typically using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi - predicted picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0111] Source pictures may be spatially subdivided, in common, into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each), and may be coded block by block. Blocks may be coded predictively by reference to other (already coded) blocks determined by a coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively by reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be coded predictively via spatial prediction or via temporal prediction by reference to one previously coded reference picture. Blocks of a B picture may be coded non-predictively via spatial prediction or via temporal prediction by reference to one or two previously coded reference pictures.
[0112] Video coder (203) may perform coding operations in accordance with a predetermined video coding technology or standard such as ITU-T Rec. H.265. In that operation, video coder (203) may perform various compression operations including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. The coded video data may thus conform to a syntax specified by the video coding technology or standard being used.
[0113] In one embodiment, transmitter (440) may transmit additional data along with the coded video. Video coder (430) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI (Supplementary Enhancement Information) messages, VUI (Visual Usability Information) parameter set fragments, etc.
[0114] Before describing certain aspects of the disclosed subject matter in further detail, it is necessary to introduce some terms that are referred to in the remainder of this description.
[0115] A sub-picture may hereinafter represent, in some cases, a rectangular configuration of samples, blocks, macroblocks, coding units, or similar entities that may be independently coded at a semantically grouped and modified resolution. One or more sub-pictures may be for a picture. One or more coding sub-pictures may form a coding picture. One or more sub-pictures may be assembled into a picture, and one or more sub-pictures may be extracted from a picture. In certain environments, one or more coding sub-pictures may be assembled into a coding picture in the compression domain without conversion to the sample level, and in the same or certain other cases, one or more coding sub-pictures may be extracted from a coding picture in the compression domain.
[0116] Adaptive Resolution Change (ARC) hereinafter represents a mechanism that allows for the change of the resolution of a picture or sub-picture within a coded video sequence, for example, by reference picture resampling. ARC parameters hereinafter represent the control information necessary to perform adaptive resolution change. This may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, etc.
[0117] The above description has focused on coding and decoding a single semantically independent coded video picture according to various embodiments. The signaling of ARC parameters may be described before explaining the meaning of coding / decoding multiple sub-pictures with independent ARC parameters and the implied additional complexity.
[0118] Referring to FIG. 5, several new options for signaling ARC parameters are shown. As noted with each of the options, they are video coding standards or techniques that have specific advantages and specific disadvantages in terms of coding efficiency, complexity, and architecture. To signal ARC parameters, one or more of these options, or options known from the prior art, may be selected. The options may not be mutually exclusive, or may be exchanged based on application needs, technically related standards, or encoder selection.
[0119] The class of ARC parameters may include the following:
[0120] - Upsampling / downsampling factors, separate or combined in the X and Y dimensions;
[0121] - Upsampling / downsampling factors indicating a constant speed zoom in / out for a given number of pictures with the addition of the temporal dimension;
[0122] - Coding of one or more possibly short syntax elements, where either of the above two may refer to a table containing factors;
[0123] - Resolution within a unit of samples, blocks, macroblocks, CUs, or any other suitable granularity of the input picture, output picture, reference picture, coded picture, combined or separate in the X or Y dimension. If there is more than one resolution (e.g., one for the input picture and one for the reference picture), in certain cases, one set of values may be estimated from another set of values. This can be controlled, for example, by the use of flags. For more detailed examples, see the following:
[0124] - The "warping" coordinates, here too at an appropriate granularity as described above, include those used in H.263 Annex P. H.263 Annex P defines one efficient way to code such warping coordinates, but other potentially more efficient ways may be devised. For example, the variable-length reversible Huffman-type coding of the warping coordinates of Annex P is replaced by a binary coding of appropriate length. Here, the length of the binary codeword is derived, for example, from the maximum picture size and may in some cases be multiplied by a specific factor and offset by a specific value, thus enabling "warping" outside the boundaries of the maximum picture size.
[0125] - Up or downsampling filter parameters. In the simplest case, there may be only a single filter for upsampling and / or downsampling. However, in certain cases, it may be advantageous to allow for more flexibility in filter design, which may require signaling of filter parameters. Such parameters may be selected through an index within a list of possible filter designs. The filter may be fully specified (e.g., through a list of filter coefficients, using appropriate entropy coding techniques), the filter may be implicitly selected through the up / downsample ratio, which is signaled according to one of the mechanisms described above according to the up / downsample ratio, etc.
[0126] In the following, it is assumed that there is coding of a finite set of up / downsample factors (the same factor to be used in both the X and Y dimensions) indicated through codewords. The codewords may advantageously be variable-length codewords that use a common Ext-Golomb code for specific syntax elements in video coding specifications such as, for example, H.264 and H.265.
[0127] One suitable mapping of values to up and / or downsample factors can follow, for example, Table 1.
Table 1
[0128] Many similar mappings can be devised according to the requirements of the applications available in video compression technologies or standards and the capabilities of the up and downscale mechanisms. The table can be extended to more values. The values may be represented, for example, using binary coding by an entropy coding mechanism other than the Ext-Bolomb code. It may have certain advantages when the resampling factor is targeted outside the video processing engine (mainly the encoder and decoder) itself, for example, by a MANET (Mobile ad hoc network). It should be noted that in the most common situation where a change in resolution is not required (presumably), the Ext-Golomb code can be selected to be short, and in the above table, it can be selected to be only a single bit. It may have an advantage in coding efficiency over using binary coding in the most common cases.
[0129] The number of entries in the table, along with their meanings, may be fully or partially configurable. For example, the basic outline of the table may be transmitted in a "high" parameter set such as a sequence or decoder parameter set. Alternatively or in an embodiment, one or more such tables may be defined in the video coding technology or standard and may be selected, for example, through a decoder or sequence parameter set.
[0130] The following describes how the upsampling / downsampling factors (ARC information) coded as described above are included in video coding technologies or standard syntax. Similar considerations may apply to one or several codeword-controlled up / downsampling filters. For discussions when relatively large amounts of data are required for filters or other data structures, see the following.
[0131] H.263 Annex P includes the ARC information (502) in the form of four warping coordinates in the picture header (501), specifically in the H.263 PLUSPTYPE (503) header extension. This can be a wise design choice when (a) there is an available picture header and (b) frequent changes to the ARC information are expected. However, the overhead when using H.263-type signaling can be very large, and since the picture header can be of a transient nature, the scaling coefficients may not belong between picture boundaries.
[0132] The previously cited JVCET-M135-v1 includes the ARC reference information (505) (index) located within the picture parameter set (504) and an index table (506) containing the target resolution located within the sequence parameter set (SPS) (507). The arrangement of possible resolutions within the table (506) within the sequence parameter set (507) can, according to the words of the author, be justified using the SPS as an interoperability negotiation point during capability exchange. The resolution can be changed for each picture, within the limits set by the values in the table (506), by referring to the appropriate picture parameter set (504).
[0133] Returning to Figure 5, there may be the following additional options for carrying the ARC information within the video bitstream. Each of these options has certain advantages over the prior art as described above. The options may exist simultaneously within the same coding technology or standard.
[0134] In an embodiment, ARC information (509), such as a resampling (zoom) factor, may be present in a slice header, a GOB (group of block) header, a tile header, or a tile group header (hereinafter, tile group header) (508). This is sufficient when the ARC information is small, for example, like a single variable length ue(v) or a fixed length codeword of several bits as described above. Having the ARC information directly within the tile group header has the additional advantage that the ARC information is applicable to a sub-picture represented by, for example, a tile group rather than the entire picture. See also below. Further, even when a video compression technology or standard assumes an adaptive resolution change for the entire picture (e.g., as opposed to a tile group based on an adaptive resolution change), putting the ARC information in the tile group header has certain advantages from the perspective of error recovery compared to putting it in the picture header in the H.263 format.
[0135] In the same or another embodiment, the ARC information (512) itself may be present within a suitable parameter set (511), such as a picture parameter set (PPS), a header parameter set, a tile parameter set, an adaptive parameter set, etc. (the adaptive parameter set shown). Advantageously, half of that parameter set is not larger than a picture, for example, a tile group. The use of the ARC information is implicitly indicated through the activation of the relevant parameter set. For example, when a video coding technology or standard assumes picture-based ARC, a picture parameter set or equivalent may be appropriate.
[0136] In the same or another embodiment, the ARC reference information (513) may be present within a tile group header (514) or a similar data structure. The reference information (513) may represent a subset (515) of the ARC information available within a parameter set (516) having a scope that extends beyond a single picture, such as a sequence parameter set, or a decoder parameter set.
[0137] As used in JVET-M0135-v1, the additional level of activation indirectly suggested from the PPS, PPS, and SPS in the tile group header seems unnecessary as it can be used for capability negotiation declaration (as in specific standards like RFC3984), just like the picture parameter set and just like the sequence parameter set. However, if the ARC information must also apply to sub-pictures represented by, for example, a tile group, a parameter set with an activation range limited to the tile group, such as an adaptive parameter set or a header parameter set, could be a suitable choice. Also, if the ARC information is of a size where it can be ignored and contains filter control information such as a large number of filter coefficients, from the perspective of coding efficiency, a parameter set could be a more suitable choice than directly using the header (508). This is because these settings can be reused by future pictures or sub-pictures by referring to the same parameter set.
[0138] When using a sequence parameter set or another higher-level parameter set with a range spanning multiple pictures, specific considerations apply:
[0139] 1. The parameter set storing the ARC information table (516) can, in some cases, be the sequence parameter set, but in other cases, advantageously, it is the decoder parameter set. The decoder parameter set can have a plurality of CVSs, that is, the activation range of the coded video stream, that is, all the coded video bits from the start of the session to the end of the session. Such a range is a characteristic of the decoder where possible ARC factors are implemented in hardware in some cases, and hardware characteristics are CVSs and tend not to change depending on the length (less than 1 second). So, it can be more appropriate (at least in some entertainment systems, Group of Pictures). That is, putting the table in the sequence parameter set is explicitly included in the choice of arrangement described here, especially in relation to the following 2.
[0140] 2. The ARC information (513) may be advantageously placed directly in the picture / slice / tile / GOB / tile group header (hereinafter the tile group header) (514), rather than in a picture parameter set as in JVCET-M0135-v1. The reason is as follows: When the encoder wants to change a single value within the picture parameter set, such as ARC reference information, the encoder has to generate a new PPS and has to refer to that new PPS. Assume that only the ARC reference information changes and other information, such as quantization matrix information within the PPS, remains the same. Such information is of a considerable size and may need to be resent to complete the new PPS.
[0141] The ARC reference information may be a single codeword, such as an index in a table (513), and the only value that changes. So, for example, it is cumbersome and wasteful to resend all of the quantization matrix information. Therefore, from the perspective of coding efficiency, contrary to JVET-M0135-v1, it may be highly appropriate to avoid the roundabout through the PPS. Similarly, putting the ARC reference information in the PPS has the additional drawback that, since the scope of picture parameter set activation is the picture, the ARC information referred to by the ARC reference information (513) is necessarily applied to the whole picture rather than to sub-pictures.
[0142] In the same or another embodiment, as outlined in FIG. 6, the signaling of the ARC parameters is described in detail below. FIG. 6 shows a syntax diagram in a representation such as has been used in video coding standards since at least 1993. The notation of such a syntax diagram generally follows C-style programming. In FIG. 6, the boldface lines indicate syntax elements that appear in the bitstream. Lines that are not in boldface may indicate control flow or variable settings.
[0143] As an exemplary syntax structure of a header applicable to a (possibly rectangular) picture part, a tile group header (601) may conditionally include a variable-length Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold). The presence of this syntax element within the tile group header can be controlled in the use of an adaptive resolution (603), here the value of a flag not shown in bold. This means that the flag is present at the point in the bitstream where it occurs in the syntax diagram.
[0144] Whether an adaptive resolution is used for this picture or part can be signaled within a high-level syntax structure either inside or outside the bitstream. In the example shown in FIG. 6, it is signaled within the sequence parameter set outlined below.
[0145] Referring further to FIG. 6, an excerpt of the sequence parameter set (610) is also shown. The first syntax element shown is adaptive_pic_resolution_change_flag (611). When true, the flag can indicate the use of an adaptive resolution, which may require specific control information. In the example, such control information is conditionally present based on the value of the flag and the tile group header (601) in an if() statement within the parameter set (612).
[0146] When adaptive resolution is used, in this example, the output resolution is coded (613) within the unit of samples. The reference sign 613 represents both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which may together define the resolution of the output picture. In other cases, in video coding technologies or standards, specific restrictions can be defined for any value. For example, the level definition may limit the total number of output samples, which may be the product of the values of those two syntax elements. Also, a specific video coding technology or standard, or an external technology or standard such as a system standard for example, may limit the numbering range (for example, one or both dimensions must be divisible by a power of 2), or the aspect ratio (for example, the width and height must be in a relationship such as 4:3 or 16:9). Such restrictions may be introduced to implement a hardware implementation or for other reasons and are well known in the art.
[0147] In a specific application, instead of implicitly assuming that the decoder assumes that the size is the output picture size, the encoder may instruct the decoder to use a specific reference picture size. In this example, the syntax element reference_pic_size_present_flag (614) controls the conditional presence of the reference picture dimensions (615) (where the reference sign also represents both the width and the height here).
[0148] Finally, FIG. 6 shows a table of possible decoded pictures having width and height. Such a table can be represented, for example, by a table indication (num_dec_pic_size_in_luma_samples_minus1) (616). "minus1" may represent the interpretation of the value of the syntax element. For example, if the coded value is 0 (zero), there is one table entry. If the value is 5, there are six table entries. For each "row" in the table, the width and height of the decoded picture are included in the syntax (617).
[0149] The existing table entries (617) can be indexed using the syntax element dec_pic_size_idx (602) within the tile group header. Thereby, different decoding sizes, in effect different zoom factors, are possible for each tile group.
[0150] Certain video coding techniques or standards, such as VP9, support spatial scalability by performing a specific form of reference picture resampling (signaled in a way completely different from the subject matter of the disclosure) in relation to temporal scalability in order to enable spatial scalability. In particular, certain reference pictures may be upsampled to a higher resolution using ARC-type techniques to form the basis of the spatial enhancement layer. These upsampled pictures can be refined using the normal prediction mechanism at the high resolution in order to add details.
[0151] The subject matter of the disclosure can be used in such an environment. In certain cases, in the same or another embodiment, values within the Network Abstraction Layer (NAL) unit header, such as the Temporal Identifier (ID) field, can be used to indicate not only the temporal layer but also the spatial layer. By doing so, certain advantages may be brought to a particular system design. For example, existing Selected Forwarding Units (SFU) generated and optimized for temporal layer selection forwarding based on the NAL unit header Temporal ID value can be used without modification in an extensible environment. In order to enable this, there may be a requirement that the mapping between the coding picture size and the temporal layer is indicated by the Temporal ID field within the NAL unit header.
[0152] In some video coding techniques, an access unit (AU) can represent a coding picture, slice, tile, NAL unit, etc., which are captured and formed into their respective picture / slice / tile / NAL unit bitstreams at a given point in time. The point in time can be the composition time.
[0153] In HEVC and certain other video coding techniques, a picture order count (POC) value can be used to indicate a selected reference picture among a plurality of reference pictures stored in a decoded picture buffer (DPB). When an access unit (AU) contains one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU may carry the same POC value, from which it can be derived that they are generated from content of the same composition time. In other words, in a scenario where two pictures / slices / tiles carry the same given POC value, it can indicate two pictures / slices / tiles belonging to the same AU and having the same composition time. Conversely, two pictures / tile / slices having different POC values can indicate that they belong to different AUs and have different composition times.
[0154] In embodiments of the disclosed subject matter, the above strict relationship can be relaxed, and an access unit can contain pictures, slices, or tiles having different POC values. By allowing different POC values within an AU, it becomes possible to use the POC value to identify picture / slice / tile that can be decoded independently in some cases having the same presentation time. This, on the one hand, can enable support for multiple scalable layers without changing the reference picture selection signaling (e.g., reference picture set signaling, or reference picture list signaling) as will be described in more detail below.
[0155] However, it is still desirable to be able to identify the access unit (AU) to which a picture / slice / tile having a different picture order count (POC) value belongs, based only on the POC value. This can be achieved as described below.
[0156] In the same or other embodiments, the access unit count (AUC) may be signaled within a higher syntax structure such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter. The value of the AUC may be used to identify which NAL unit, picture, slice, or tile belongs to a given AU. The value of the AUC may correspond to different configuration times. The AUC value may be equal to a multiple of the POC value. The AUC value may be calculated by dividing the POC value by an integer value. In certain cases, the division operation may impose a particular load on the decoder implementation. In such cases, a small constraint in the numbering space of the AUC value enables the division operation to be replaced by a shift operation. For example, the AUC value may be equal to the most significant bit (MSB) value of the POC value range.
[0157] In the same embodiment, the value of the POC cycle per AU (poc_cycle_au) may be signaled within a higher syntax structure such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or AU delimiter. The poc_cycle_au may indicate how many different consecutive POC values can be associated with the same AU. For example, if the value of the poc_cycle_au is equal to 4, pictures, slices, or tiles having POC values equal to 0 to 3, inclusive, are associated with an AU having an AUC value equal to 0, and pictures, slices, or tiles having POC values equal to 4 to 7, inclusive, are associated with an AU having an AUC value equal to 1. Thus, the AUC value may be estimated by dividing the POC value by the value of the poc_cycle_au.
[0158] In the same or another embodiment, the value of poc_cyle_au may be derived from information identifying the number of spatial or SNR layers within the coded video sequence, for example located within a video parameter set (VPS). Such possible relationships are briefly described below. Such derivation as described above can save a few bits within the VPS and thus improve coding efficiency, but it is advantageous to explicitly code poc_cyle_au within an appropriate upper syntax structure hierarchically below the video parameter set so that poc_cycle_au can be minimized for a given small portion of the bitstream such as a picture. This optimization can save more bits than can be saved through the above-described derivation process because the POC value (and / or the value of the syntax element that indirectly references the POC) can be coded in the lower syntax structure.
[0159] In the same or another embodiment, FIG. 9 shows the syntax element vps_poc_cycle_au in the VPS (or SPS) indicating poc_cycle_au used for all pictures / slices in the coding video sequence, and an example of a syntax table for signaling the syntax element slice_poc_cycle_au indicating the poc_cycle_au of the current slice in the slice header. When the POC value increases uniformly for each AU, vps_contant_poc_cycle_per_au in the VPS is set to 1, and vps_poc_cycle_au is signaled in the VPS. In this case, slice_poc_cycle_au is not explicitly signaled, and the AUC value of each AU is calculated by dividing the POC value by vps_poc_cycle_au. When the POC value does not increase uniformly for each AU, vps_contant_poc_cycle_per_au in the VPS is set to 0. In this case, vps_access_unit_cnt is not signaled, but slice_access_unit_cnt is signaled in the slice header for each slice or picture. Each slice or picture may have a different value of slice_access_unit_cnt. The AUC value of each AU is calculated by dividing the POC value by slice_poc_cycle_au. FIG. 10 shows a block diagram illustrating the related workflow.
[0160] In the same or another embodiment, even if the POC values of pictures, slices, or tiles may be different, pictures, slices, or tiles corresponding to AUs having the same AUC value may be associated with the same decoding or output time. Thus, all or part of the pictures, slices, or tiles associated with the same AU may be decoded in parallel and output at the same time without having an interlacing / decoding dependency across the pictures, slices, or tiles within the same AU.
[0161] In the same or another embodiment, pictures, slices, or tiles corresponding to the same AU having the same AUC value may be associated with the same configuration / display time point even if the POC values of the pictures, slices, or tiles are different. When the configuration time point is included in the container format, pictures can be displayed at the same time point if they have the same configuration time point even if they correspond to different AUs.
[0162] In the same or other embodiments, each picture, slice, or tile may have the same temporal identifier (temporal_id) within the same AU. All or part of the pictures, slices, or tiles corresponding to a certain time point may be associated with the same temporal sublayer. In the same or other embodiments, each picture, slice, or tile may have different spatial layer identifiers (layer_id) within the same AU. All or part of the pictures, slices, or tiles corresponding to a certain time point may be associated with the same or different spatial layers.
[0163] POC diagram 8 shows an example of a video sequence structure having a combination of temporal_id, layer_id, POC, and AUC due to adaptive resolution change. In this example, pictures, slices, or tiles in the first AU having AUC = 0 may have temporal_id = 0 and layer_id = 0 or 1, while pictures, slices, or tiles in the second AU having AUC = 1 may have temporal_id = 1 and layer_id = 0 or 1 respectively. The value of POC increases by only 1 per picture regardless of the values of temporal_id and layer_id. In this example, the value of poc_cycle_au is equal to 2. Desirably, the value of poc_cycle_au may be set equal to the number of (spatial scalability) layers. In this example, therefore, the value of POC is increased by 2 and the value of AUC is increased by 1.
[0164] In the above-described embodiments, all or part of the reference picture indication and the inter-picture or inter-layer prediction structure may be supported using the existing reference picture set (RPS) signaling or reference picture list (RPL) in HEVC. In the RPS or RPL, the selected reference picture is indicated by signaling the value of the POC or the POC delta value between the current picture and the selected reference picture. In the disclosed subject matter, the RPS and RPL can be used to indicate the inter-picture or inter-layer prediction structure without changing the signaling, but with the following constraints. If the value of the temporal_id of the reference picture is greater than the value of the temporal_id of the current picture, the current picture may not use the reference picture for motion compensation or other prediction. If the value of the layer_id of the reference picture is greater than the value of the layer_id of the current picture, the current picture may not use the reference picture for motion compensation or other prediction.
[0165] In the same or other embodiments, motion vector scaling based on the POC difference for temporal motion vector prediction may be disabled across multiple pictures within an access unit. Thus, each picture may have a different POC value within the access unit, but the motion vectors are not scaled and used for temporal motion vector prediction within the access unit. This is because reference pictures with different POCs within the same AU are considered to have the same time point. Thus, in an embodiment, the motion vector scaling function may return 1 when the reference picture belongs to the AU associated with the current picture.
[0166] In the same and other embodiments, motion vector scaling based on the POC difference for temporal motion vector prediction may optionally be disabled across multiple pictures when the spatial resolution of the reference picture is different from the spatial resolution of the current picture. When motion vector scaling is permitted, the motion vectors are scaled based on the POC difference and the spatial resolution ratio between the current picture and the reference picture.
[0167] In the same or another embodiment, particularly when poc_cycle_au has a non-uniform value (when vps_contant_poc_cycle_per_au == 0), the motion vectors may be scaled based on the AUC difference instead of the POC difference for temporal motion vector prediction. In other cases (when vps_contant_poc_cycle_per_au == 1), the motion vector scaling based on the AUC difference may be the same as the motion vector scaling based on the POC difference.
[0168] In the same or another embodiment, when the motion vectors are scaled based on the AUC difference, the reference motion vectors within the same AU (having the same AUC value) as the current picture are not scaled based on the AUC difference for motion vector prediction and may or may not involve scaling based on the spatial resolution ratio between the current picture and the reference picture.
[0169] In the same and other embodiments, the AUC value is used to identify the boundaries of the AU and is used for the operation of a hypothetical reference decoder (HDR) that requires the input and output timings of the AU granularity. In many cases, the decoded picture having the top layer within the AU may be output for display. The AUC value and the layer_id value can be used to identify the output picture.
[0170] In an embodiment, a picture may be composed of one or more sub-pictures. Each sub-picture may cover a local area or the entire area of the picture. The area supported by a sub-picture may or may not overlap with the area supported by another sub-picture. The area composed of one or more sub-pictures may or may not cover the entire area of the picture. When a picture is composed of sub-pictures, the area supported by a sub-picture is the same as the area supported by the picture.
[0171] In the same embodiment, a sub-picture may be coded by a coding method similar to the coding method used for the coding picture. A sub-picture may be coded independently, or may be coded depending on another sub-picture or the coding picture. A sub-picture may or may not have a parsing dependency relationship from another sub-picture or the coding picture.
[0172] In the same embodiment, a coding sub-picture may be included in one or more layers. The coding sub-pictures in a layer may have different spatial resolutions. The original sub-picture may be spatially resampled (upsampled or downsampled), coded with different spatial resolution parameters, and included in the bitstream corresponding to the layer.
[0173] In the same or another embodiment, a sub-picture having (W, H) may be coded and included in the coding bitstream corresponding to layer 0. Here, W represents the width of the sub-picture, and H represents the height of the sub-picture respectively. On the other hand, a sub-picture upsampled (or downsampled) from a sub-picture having the original spatial resolution and having (W*S w,k , H*S h,k ) may be coded and included in the coding bitstream corresponding to layer k. Here, S w,k , S h,kindicates resampling ratios in the horizontal and vertical directions. S w,k ,S h,k If the value of is greater than 1, the resampling is equal to upsampling. On the other hand, S w,k ,S h,k if the value of is less than 1, the resampling is equal to downsampling.
[0174] In the same or another embodiment, the coding sub-pictures within a layer may have different visual qualities from the coding sub-pictures within another layer of the same sub-picture or a different sub-picture. For example, sub-picture i within layer n is coded by quantization parameter Q i,n , and sub-picture j within layer m is coded by quantization parameter Q j,m .
[0175] In the same or another embodiment, the coding sub-pictures within a layer may be independently decodable and have no parsing or decoding dependencies from the coding sub-pictures within another layer of the same local region. A sub-picture layer that is independently decodable without referring to another sub-picture layer of the same local region is an independent sub-picture layer. The coding sub-pictures within an independent sub-picture layer may or may not have decoding or parsing dependencies from previous coding sub-pictures within the same sub-picture layer, but the coding sub-pictures need not have dependencies from coding pictures within another sub-picture layer.
[0176] In the same or another embodiment, the coding sub-pictures within a layer may be dependently decodable and have a parsing or decoding dependency from coding sub-pictures within another layer of the same local region. A sub-picture layer that is dependently decodable by referring to another sub-picture layer of the same local region is a dependent sub-picture layer. The coding sub-pictures within a dependent sub-picture may refer to coding sub-pictures belonging to the same sub-picture, previous coding sub-pictures within the same sub-picture layer, or both reference sub-pictures.
[0177] In the same or another embodiment, a coding sub-picture is composed of one or more independent sub-picture layers and one or more dependent sub-picture layers. However, for a coding sub-picture, there may be at least one independent sub-picture layer. An independent sub-picture layer may have a value of layer identifier (layer_id) equal to 0, which may exist within the NAL unit header or another upper syntax structure. A sub-picture layer having a layer_id equal to 0 is a basic sub-picture layer.
[0178] In the same or another embodiment, a picture is composed of one or more foreground sub-pictures and one or more background sub-pictures. The region supported by the background sub-picture may be equal to the region of the picture. The region supported by the foreground sub-picture may overlap with the region supported by the background sub-picture. The background sub-picture may be a basic sub-picture layer, while the foreground sub-picture may be a non-basic (extended) sub-picture layer. One or more non-basic sub-picture layers may refer to the same basic layer for decoding. Each non-basic sub-picture layer having a layer_id equal to a may refer to a non-basic sub-picture layer having a layer_id equal to b. Here, a is greater than b.
[0179] In the same or another embodiment, the picture may be composed of one or more foreground sub-pictures with or without a background sub-picture. Each sub-picture may have its own basic sub-picture layer and one or more non-basic (extended) layers. Each basic sub-picture layer may be referenced by one or more non-basic sub-picture layers. Each non-basic sub-picture layer having a layer_id equal to a may reference a non-basic sub-picture layer having a layer_id equal to b. Here, a is greater than b.
[0180] In the same or another embodiment, the picture may be composed of one or more foreground sub-pictures with or without a background sub-picture. Each coding sub-picture within a (basic or non-basic) sub-picture layer may be referenced by sub-pictures of one or more non-basic layers belonging to the same sub-picture and sub-pictures of one or more non-basic layers not belonging to the same sub-picture.
[0181] In the same or another embodiment, the picture may be composed of one or more foreground sub-pictures with or without a background sub-picture. The sub-pictures within layer a may be further partitioned into a plurality of sub-pictures within the same layer. One or more coding sub-pictures within layer b may reference the partitioned sub-pictures within layer a.
[0182] In the same or another embodiment, the coded video sequence (CVS) may be a group of coded pictures. The CVS may be composed of one or more coded sub-picture sequences (CSPS). Here, the CSPS may be a group of coding sub-pictures covering the same local area of the picture. The CSPS may have the same or different temporal resolutions as the coded video sequence.
[0183] In the same or another embodiment, the CSPS may be included in one or more coded layers. The CSPS may be composed of one or more CSPS layers. Decoding one or more CSPS layers corresponding to the CSPS may reconstruct a sequence of sub-pictures corresponding to the same local area.
[0184] In the same or another embodiment, the number of CSPS layers corresponding to the CSPS may be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0185] In the same or another embodiment, the CSPS layer may have a different temporal resolution (e.g., frame rate) from another CSPS layer. The original (uncompressed) sub-picture sequence may be temporally resampled (upsampled or downsampled), coded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.
[0186] In the same or another embodiment, a sub-picture sequence having a frame rate F may be coded and included in the coding bitstream corresponding to layer 0. On the other hand, a sub-picture sequence temporally upsampled (or downsampled) from the original sub-picture sequence by F*S t,k may be coded and included in the coding bitstream corresponding to layer k. Here, S t,k represents the temporal sampling ratio of layer k. When the value of S t,k is greater than 1, the temporal resampling process is equivalent to frame rate up-conversion. On the other hand, when the value of S t,k is less than 1, the temporal resampling process is equivalent to frame rate down-conversion.
[0187] In the same or another embodiment, when a sub-picture having a CSPS layer a is referenced by a sub-picture having a CSPS layer b for motion compensation or any inter-layer prediction, if the spatial resolution of the CSPS layer a is different from the spatial resolution of the CSPS layer b, the decoded pixels of the CSPS layer a are resampled and used for reference. The resampling process may require upsampling filtering or downsampling filtering.
[0188] FIG. 11 shows an exemplary video stream including a background video CSPS having a layer_id equal to 0 and a plurality of foreground CSPS layers. The coding sub-picture may be composed of one or more CSPS layers, and the background area that does not belong to any foreground CSPS layer may constitute the base layer. The base layer may include the background area and the foreground area, and the extended CSPS layer may include the foreground area. The extended CSPS layer may have better visual quality than the base layer in the same area. The extended CSPS layer may refer to the reconstructed pixels and motion vectors of the base layer corresponding to the same area.
[0189] In the same or another embodiment, the video bitstream corresponding to the base layer is included in a track, while the CSPS layer corresponding to each sub-picture is included in another track within the video file.
[0190] In the same or another embodiment, the video bitstream corresponding to the base layer is included in a track, while the CSPS layer corresponding to the same layer_id is included in another track. In this example, the track corresponding to layer k includes only the CSPS layer corresponding to layer k.
[0191] In the same or another embodiment, each CSPS layer of each sub-picture is stored in a separate track. Each track may or may not have a parsing or decoding dependency from one or more other tracks.
[0192] In the same or another embodiment, each track may include a bitstream corresponding to layers i to j of all or part of the CSPS layers of the subpicture. Here, 0 < i <= j <= k, and k is the highest layer of CSPS.
[0193] In the same or another embodiment, a picture is composed of one or more associated media data including a depth map, an alpha map, 3D geometry data, an occupancy map, etc. Such associated temporal media data can be divided into one or more data substreams, and each data substream corresponds to one subpicture.
[0194] In the same or another embodiment, FIG. 12 shows an example of a video conference based on the multi-layer subpicture method. The video stream includes one basic layer video bitstream corresponding to the background picture and one or more enhancement layer video bitstreams corresponding to the foreground subpicture. Each enhancement layer video bitstream may correspond to a CSPS layer. In the display, the picture corresponding to the basic layer is displayed by default. It includes the pictures of one or more users within the picture (picture in a picture (PIP)). When a specific user is selected under the control of the client, the enhanced CSPS layer corresponding to the selected user is decoded and displayed according to the enhanced quality or spatial resolution. FIG. 13 shows a block diagram of decoding and display processing of a video bitstream including a multi-layer sub-picture according to an embodiment. For example, the processing may include one or more of the following operations. For example, in operation 1301, decoding of a video bitstream having multiple layers may occur. Operation 1302 may include the step of identifying a background region and one or more foreground sub-pictures. Operation 1303 may include the step of determining whether a particular sub-picture region is selected. Operation 1304 may include the step of decoding and displaying an extended sub-picture if a particular sub-picture region is selected (i.e., 1303 = Yes). Operation 1305 may include the step of decoding and displaying the background region if a particular sub-picture region is not selected (i.e., 1303 = No).
[0195] In the same or another embodiment, a network intermediate box (e.g., a router) may select a subset of layers to be transmitted to a user depending on the bandwidth. Picture / sub-picture composition may be used for bandwidth adaptation. For example, if the user has no bandwidth, the router may delete layers or select some sub-pictures based on their importance or user settings. This can be done dynamically to adapt to the bandwidth.
[0196] FIG. 14 shows an example of the use of 360-degree video. When a 360-degree picture of a sphere is projected onto a planar picture, the projected 360-degree picture may be partitioned into a plurality of sub-pictures such as a base layer. The extended layer of a particular sub-picture may be coded and transmitted to a client. The decoder may be able to decode both the base layer including all sub-pictures and the extended layer of the selected sub-pictures. When the current viewpoint is the same as the selected sub-picture, the displayed picture may have higher quality with the decoded sub-picture having the extended layer. Alternatively, the decoded picture having the base layer may be displayed with low quality.
[0197] In the same or another embodiment, the layout information for display may exist in a file as auxiliary information (e.g., SEI message or metadata). One or more decoded sub-pictures may be rearranged and displayed according to the signaled layout information. The layout information may be signaled by a streaming server or broadcaster, or may be regenerated by a network entity or cloud server, or may be determined by the user's customized settings.
[0198] In an embodiment, an input picture may be divided into one or more (rectangular) sub-regions, and each sub-region may be coded as an independent layer. Each independent layer corresponding to a local region may have a unique layer_id value. For each independent layer, sub-picture size and position information may be signaled. For example, the picture size (width, height), and the offset information (x_offset, y_offset) of the upper left corner. FIG. 15 shows an example of the layout of the divided sub-pictures, their sub-picture size and position information, and their corresponding picture prediction structure. The layout information including the sub-picture size and sub-picture position may be signaled in a parameter set, slice or group header, or a higher syntax structure such as an SEI message.
[0199] In the same embodiment, each sub-picture corresponding to an independent layer may have its own unique POC value within an AU. When the reference pictures in the pictures stored in the DBP are indicated using the syntax elements in the RPS or RPL structure, the POC value of each sub-picture corresponding to the layer may be used.
[0200] In the same or another embodiment, in order to indicate an (inter-layer) prediction structure, the layer_id may not be used, and the POC (delta) value may be used.
[0201] In the same embodiment, a sub-picture having a POC value equal to N corresponding to a layer (or local region) may or may not be used as a reference picture of a sub-picture having a POC value equal to N + K corresponding to the same layer (or the same local region) for motion compensation prediction. In most cases, the value of the numerical value K may be equal to the number of sub-regions or may be equal to the maximum number of (independent) layers.
[0202] In the same or another embodiment, FIG. 16 shows an extended case of FIG. 15. When the input picture is divided into a plurality of (e.g., four) sub-regions, each local region may be coded by one or more layers. In this case, the number of independent layers may be equal to the number of sub-regions, and one or more layers may correspond to the sub-regions. Accordingly, each sub-region may be coded by one or more independent layers and zero or more dependent layers.
[0203] In the same embodiment, in FIG. 16, the input picture may be divided into four sub-regions. The upper right sub-region may be coded as two layers, namely layer 1 and layer 4. On the other hand, the lower right sub-region may be coded as two layers, namely layer 3 and layer 5. In this case, layer 4 may refer to layer 1 for motion compensation prediction, and layer 5 may refer to layer 3 for motion compensation.
[0204] In the same or another embodiment, an in-loop filter (e.g., a deblocking filter, an adaptive in-loop filter, a reshaper, a bilateral filter, or any deep learning-based filter) spanning a layer boundary may be (optionally) disabled.
[0205] In the same or another embodiment, motion compensation prediction or intra-block copy spanning a layer boundary may be (optionally) disabled.
[0206] In the same or another embodiment, the boundary padding for motion compensation prediction or in-loop filtering at the boundary of a sub-picture may be optionally processed. A flag indicating whether the boundary padding is processed may be signaled in a parameter set (VPS, SPS, PPS, or APS) or a top-level syntax structure such as a slice or tile group header, or an SEI message.
[0207] In the same or another embodiment, the layout information of a sub-region (or sub-picture) may be signaled within a VPS or MPS. FIG. 17 shows an example of syntax elements within a VPS and an SPS. In this example, vps_sub_picture_dividing_flag is signaled within the VPS. The flag may indicate whether the input picture is divided into a plurality of sub-regions.
[0208] When the value of vps_sub_picture_dividing_flag is equal to 0, the input picture in the coding video sequence corresponding to the current VPS may not be divided into a plurality of sub-regions. In this case, the input picture size may be equal to the coding picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled within the SPS.
[0209] When the value of vps_sub_picture_dividing_flag is equal to 1, the input picture may be divided into a plurality of sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled within the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be equal to the width and height of the input picture, respectively.
[0210] In the same embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding, but may be used for configuration and display.
[0211] In the same embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntax elements pic_offset_x and pic_offset_y may be signaled within the SPS corresponding to a particular layer. In this case, the coded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled within the SPS may be equal to the width and height of the sub-region corresponding to a particular layer. Also, the position (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may be signaled within the SPS.
[0212] In the same embodiment, the position (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may not be used for decoding, but may be used for configuration and display.
[0213] In the same or another embodiment, the layout information (size and position) of all or part of the sub-region of the input picture, and the dependency information between layers may be signaled within a parameter set or SEI message.
[0214] FIG. 18 shows an example of syntax elements for indicating information on the layout of sub-regions, dependencies between layers, and relationships between sub-regions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) sub-regions in the currently coded video sequence. According to an embodiment, the syntax element num_layers may indicate the number of layers in the currently coded video sequence. The value of num_layers may be equal to or greater than the value of num_sub_region. When any sub-region is coded as a single layer, the value of num_layers may be equal to the value of num_sub_region. When one or more sub-regions are coded as multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dependency_flag[i][j] indicates the dependency from the j-th layer to the i-th layer. num_layers_for_region[i] indicates the number of layers associated with the i-th sub-region. sub_region_layer_id[i][j] indicates the layer_id of the j-th layer associated with the i-th sub-region. sub_region_offset_x[i] and sub_region_offset_y[i] indicate the horizontal and vertical positions of the upper left corner of the i-th sub-region, respectively. sub_region_width[i] and sub_region_height[i] indicate the width and height of the i-th sub-region, respectively.
[0215] In one embodiment, one or more syntax elements that specify whether an output layer, which is set to indicate one or more layers, is output with or without profile tier level information may be signaled within a higher-level syntax structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message. Referring to FIG. 19, a syntax element num_output_layer_sets indicating the number of output layer sets (OLSs) within a coded video sequence that refers to a VPS may be signaled within the VPS. For each output layer set, an output_layer_flag may be signaled as many times as the number of output layers.
[0216] In the same embodiment, output_layer_flag[i] being equal to 1 specifies that the i-th layer is output. vps_output_layer_flag[i] being equal to 0 specifies that the i-th layer is not output.
[0217] In the same or another embodiment, one or more syntax elements that specify profile tier level information for each output layer set may be signaled within a higher-level syntax structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message. Further referring to FIG. 19, a syntax element num_profile_tile_level indicating the number of profile tier level information per OLS within a coded video sequence that refers to a VPS may be signaled within the VPS. For each output layer set, a set of syntax elements of the profile tier level information, or an index indicating a specific profile tier level information among the entries within the profile tier level information, may be signaled as many times as the number of output layers.
[0218] In the same embodiment, profile_tier_level_idx[i][j] specifies the index of the profile_tier_level() syntax structure applied to the j-th layer of the i-th OLS into the list of profile_tier_level() syntax structures within the VPS.
[0219] In the same or another embodiment, referring to FIG. 20, the syntax elements num_profile_tile_level and / or num_output_layer_sets may be signaled when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0).
[0220] In the same or another embodiment, referring to FIG. 20, the syntax element vps_output_layers_mode[i] indicating the mode of output layer signaling for the i-th output layer set may exist within the VPS.
[0221] In the same embodiment, vps_output_layers_mode[i] being equal to 0 specifies that only the top layer is output together with the i-th output layer set. vps_output_layer_mode[i] equal to 1 specifies that all layers are output together with the i-th output layer set. vps_output_layer_mode[i] equal to 2 specifies that the layers to be output are the layers having vps_output_layer_flag[i][j] equal to 1 together with the i-th output layer set. More values may be reserved.
[0222] In the same embodiment, output_layer_flag[i][j] may or may not be signaled depending on the value of vps_output_layers_mode[i] of the i-th output layer set.
[0223] In the same or another embodiment, referring to FIG. 20, flagvps_ptl_signal_flag[i] may exist for the i-th output layer set. Depending on the value of vps_ptl_signal_flag[i], the profile tier level information of the i-th output layer set may or may not be signaled.
[0224] In the same or another embodiment, referring to FIG. 21, the number of sub-pictures max_subpics_minus1 in the current CVS may be signaled in a higher-level syntax structure, such as in a VPS, DPS, SPS, PPS, APS, or SEI message.
[0225] In the same or another embodiment, referring to FIG. 21, the sub-picture identifier sub_pic_id[i] of the i-th sub-picture may be signaled when the maximum layer sub-picture number is greater than 1 (max_subpics_minus1>0).
[0226] In the same or another embodiment, one or more syntax elements indicating the sub-picture identifiers belonging to each layer of each output layer set may be signaled in the VPS. Referring to FIG. 22, sub_pic_id_layer[i][j][k] indicating the k-th sub-picture exists within the j-th layer of the i-th output layer set. With this information, the decoder may recognize which sub-pictures can be decoded and output for each layer of a specific output layer set.
[0227] In an embodiment, a picture header (PH) is a syntax structure that includes syntax elements applied to all slices of a coded picture. A picture unit (PU) is a set of NAL units that contains exactly one coded picture, consecutive in decoding order, and associated with each other according to specified classification rules. A PU may include a picture header (PH) and one or more VCL NAL units corresponding to the coded picture.
[0228] In an embodiment, an SPS (RBSP) may be available for decoding before being included in at least one AU having a TemporalId equal to 0 and being referenced, or provided through external means.
[0229] In an embodiment, an SPS (RBSP) may be available for decoding before being included in at least one AU having a TemporalId equal to 0 within a CVS that includes one or more PPSs referencing the SPS, or provided through external means.
[0230] In an embodiment, an SPS (RBSP) may be available for decoding before being included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the SPS NAL unit within a CVS that is referenced by one or more PPSs and includes one or more PPSs referencing the SPS, or provided through external means.
[0231] In an embodiment, an SPS (RBSP) may be available for decoding before being included within one or more PUs having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of a PPS NAL unit referencing the SPS NAL unit, and being referenced by one or more PPSs, or provided through external means.
[0232] In an embodiment, the SPS (RBSP) is referred to by one or more PPSs and is included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the SPS NAL unit within the CVS, or may be made available for decoding through external means or before being provided through external means, and includes one or more PPSs that refer to a TemporalId and SPS equal to 0.
[0233] In the same or another embodiment, the pps_seq_parameter_set_id specifies the value of the sps_seq_parameter_set_id of the SPS being referred to. The value of the pps_seq_parameter_set_id may be the same among all PPSs referred to by the coding picture within the CLVS.
[0234] In the same or another embodiment, all SPS NAL units having a sps_seq_parameter_set_id of a particular value within the CVS may have the same content.
[0235] In the same or another embodiment, regardless of the value of the nuh_layer_id, the SPS NAL units may share the same value space of the sps_seq_parameter_set_id.
[0236] In the same or another embodiment, the nuh_layer_id value of the SPS NAL unit may be equal to the lowest nuh_layer_id value of the PPS NAL unit that refers to the SPS NAL unit.
[0237] In an embodiment, when an SPS having a nuh_layer_id equal to m is referred to by one or more PPSs having a nuh_layer_id equal to n, the layer having a nuh_layer_id equal to m may be the same as the layer having a nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to m.
[0238] In an embodiment, the PPS (RBSP) may be available for decoding before being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit and being referenced, or provided through external means.
[0239] In an embodiment, the PPS (RBSP) may be available for decoding before being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS, which is referenced and includes one or more PHs (or coding slice NAL units) that reference the PPS, or provided through external means.
[0240] In an embodiment, the PPS (RBSP) may be available for decoding before being included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the coding slice NAL units that reference the PPS NAL unit in the CVS, which is referenced by one or more PHs (or coding slice NAL units) and includes one or more PHs (or coding slice NAL units) that reference the PPS, or provided through external means.
[0241] In an embodiment, the PPS (RBSP) may be available for decoding before being included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the coding slice NAL units that reference the PPS NAL unit in the CVS and a TemporalId equal to the TemporalId of the PPS NAL unit, which is referenced by one or more PHs (or coding slice NAL units) and includes one or more PHs (or coding slice NAL units), or provided through external means or before being provided through external means.
[0242] In the same or another embodiment, the ph_pic_parameter_set_id in the PH specifies the value of the pps_pic_parameter_set_id of the PPS being referred to in use. The value of the pps_seq_parameter_set_id may be the same among all the PPSs referred to by the coding pictures in the CLVS.
[0243] In the same or another embodiment, all the PPS NAL units having a pps_pic_parameter_set_id with a specific value within a PU may have the same content.
[0244] In the same or another embodiment, regardless of the value of the nuh_layer_id, the PPS NAL units may share the same value space of the pps_pic_parameter_set_id.
[0245] In the same or another embodiment, the nuh_layer_id value of the PPS NAL unit may be equal to the lowest nuh_layer_id value of the coding slice NAL unit that refers to the NAL unit that refers to the PPS NAL unit.
[0246] In an embodiment, when a PPS having a nuh_layer_id equal to m is referred to by one or more coding slice NAL units having a nuh_layer_id equal to n, the layer having a nuh_layer_id equal to m may be the same as the layer having a nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to m.
[0247] In an embodiment, the PPS (RBSP) may be available for decoding before being included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit or being provided through external means.
[0248] In an embodiment, the PPS (RBSP) is referenced and included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS, which includes one or more PHs (or coding slice NAL units) that reference the PPS, or may be available for decoding before being provided through external means.
[0249] In an embodiment, the PPS (RBSP) is referenced by one or more PHs (or coding slice NAL units) and is included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the coding slice NAL units that reference the PPS NAL unit in the CVS and include one or more PHs (or coding slice NAL units) that reference the PPS, or may be available for decoding before being provided through external means.
[0250] In an embodiment, the PPS (RBSP) is referenced by one or more PHs (or coding slice NAL units) and is included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the coding slice NAL units that reference the PPS NAL unit in the CVS and a TemporalId equal to the TemporalId of the PPS NAL unit, which includes one or more PHs (or coding slice NAL units), or may be available for decoding before being provided through external means or before being provided through external means.
[0251] In the same or another embodiment, the ph_pic_parameter_set_id in the PH specifies the value of the pps_pic_parameter_set_id of the referenced PPS in use. The value of the pps_seq_parameter_set_id may be the same among all the PPSs referenced by the coding pictures in the CLVS.
[0252] In the same or another embodiment, all PPS NAL units having a particular value of pps_pic_parameter_set_id within the PU may have the same content.
[0253] In the same or another embodiment, regardless of the value of nuh_layer_id, PPS NAL units may share the same value space of pps_pic_parameter_set_id.
[0254] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit may be equal to the lowest nuh_layer_id value of the coding slice NAL units that reference the NAL unit that references the PPS NAL unit.
[0255] In an embodiment, a PPS having a nuh_layer_id equal to m is referenced by one or more coding slice NAL units having a nuh_layer_id equal to n. The layer having a nuh_layer_id equal to m may be the same as the layer having a nuh_layer_id equal to n, or the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to m.
[0256] The output layer indicates the layer of the output layer set to be output. The output layer set (OLS) indicates a layer set composed of a specified layer set, and one or more layers in the layer set are specified as output layers. The output layer set (OLS) layer index is the index of the layer within the OLS to the list of layers within the OLS.
[0257] A sublayer represents a temporally scalable layer of a temporally scalable bitstream composed of VCL NAL units having a TemporalId variable of a specific value and associated non-VCL NAL units. A sublayer representation represents a subset of a bitstream composed of NAL units of a specific sublayer and lower sublayers.
[0258] The VPS RBSP may be available for decoding before being included in at least one AU having a TemporalId equal to 0 and referenced, or provided through external means. All VPS NAL units having a vps_video_parameter_set_id of a specific value within a CVS may have the same content.
[0259] vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id may be greater than 0.
[0260] One plus vps_max_sublayers_minus1 specifies the maximum number of temporal sublayers that may exist within each CVS that references the VPS.
[0261] One plus vps_max_sublayers_minus1 specifies the maximum number of temporal sublayers that may exist within each CVS that references the VPS. The value of vps_max_sublayers_minus1 may be in the range of 0 to 6, inclusive.
[0262] vps_all_layers_same_num_sublayers_flag equal to 1 specifies that the number of temporal sublayers is the same for all layers within each CVS that references the VPS.
[0263] The vps_all_layers_same_num_sublayers_flag equal to 0 specifies whether the layers within each CVS referring to the VPS may have the same number of temporal sublayers or not. When it does not exist, the value of the vps_all_layers_same_num_sublayers_flag is presumed to be equal to 1.
[0264] The vps_all_independent_layers_flag equal to 1 specifies that all the layers within the CVS are coded independently without using inter-layer prediction.
[0265] The vps_all_independent_layers_flag equal to 0 specifies that one or more of the layers within the CVS may use inter-layer prediction. When it does not exist, the value of the vps_all_independent_layers_flag is presumed to be equal to 1.
[0266] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is smaller than n, the value of vps_layer_id[m] may be smaller than vps_layer_id[n].
[0267] That vps_independent_layer_flag[i] is equal to 1 specifies that the layer with index i does not use inter-layer prediction. That vps_independent_layer_flag[i] is equal to 0 specifies that the layer with index i may use inter-layer prediction and syntactic elements.
[0268] For j in the range of 0 to i - 1 including both ends, vps_direct_ref_layer_flag[i][j] exists within the VPS. When it does not exist, the value of vps_independent_layer_flag[i] is presumed to be equal to 1. vps_direct_ref_layer_flag[i][j] equal to 0 specifies that the layer with index j is not a direct reference layer of the layer with index i. vps_direct_ref_layer_flag[i][j] equal to 1 specifies that the layer with index j is a direct reference layer of the layer with index i.
[0269] When vps_direct_ref_layer_flag[i][j] does not exist for i and j in the range of 0 to vps_max_layers_minus1 including both ends, it is presumed to be equal to 0. When vps_independent_layer_flag[i] is equal to 0, at least one value of j in the range of 0 to i - 1 including both ends may exist, and as a result, the value of vps_direct_ref_layer_flag[i][j] is equal to 1.
[0270] The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are derived as follows:
Number
[0271] The variable GeneralLayerIdx[i] that specifies the layer index of the layer having nuh_layer_id equal to vps_layer_id[i] is derived as follows:
Number
[0272] For any two different values of both i and j within the range of 0 to vps_max_layers_minus1 including both ends, when dependencyFlag[i][j] is equal to 1, it is a requirement for bitstream compliance that the values of chroma_format_idc and bit_depth_minus8 applied to the i-th layer can be equal to the values of chroma_format_idc and bit_depth_minus8 applied to the j-th layer, respectively.
[0273] max_tid_ref_present_flag[i] equal to 1 specifies that the syntax element max_tid_il_ref_pics_plus1[i] exists. max_tid_ref_present_flag[i] equal to 0 specifies that the syntax element max_tid_il_ref_pics_plus1[i] does not exist.
[0274] max_tid_il_ref_pics_plus1[i] equal to 0 specifies that inter-layer prediction is not used by non-IRAP pictures of the i-th layer. max_tid_il_ref_pics_plus1[i] greater than 0 specifies that for decoding pictures of the i-th layer, pictures with a TemporalId greater than max_tid_il_ref_pics_plus1[i] - 1 are not used as ILRP. When it does not exist, the value of max_tid_il_ref_pics_plus1[i] is presumed to be equal to 7.
[0275] each_layer_is_an_ols_flag equal to 1 specifies that each OLS contains only one layer and that each layer itself within the CVS referring to the VPS is an OLS where the single layer it contains is only the output layer. each_layer_is_an_ols_flag equal to 0 allows for more than one layer. If vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be 1. In other cases, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be 0.
[0276] ols_mode_idc equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains the layers with layer indices from 0 to i including both ends, and for each OLS, only the topmost layer within the OLS is output.
[0277] ols_mode_idc equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS contains the layers with layer indices from 0 to i including both ends, and for each OLS, all the layers within the OLS are output.
[0278] ols_mode_idc equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled, and for each OLS, the output layer is explicitly signaled and the other layers are layers that are direct or indirect reference layers to the output layer of the OLS.
[0279] The value of ols_mode_idc may be in the range from 0 to 2 including both ends. ols_mode_idc with a value of 3 is reserved for future use by ITU-T / ISO / IEC.
[0280] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is presumed to be equal to 2.
[0281] Adding 1 to num_output_layer_sets_minus1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0282] The variable TotalNumOlss that specifies the total number of PLSs specified by the VPS is derived as follows:
Number
[0283] vps_all_layers_same_num_sublayers_flag equal to 0 specifies whether the layers within each CVS referring to the VPS may have the same number of temporal sublayers or not. When it does not exist, the value of vps_all_layers_same_num_sublayers_flag is presumed to be equal to 1.
[0284] vps_all_independent_layers_flag equal to 1 specifies that all the layers within the CVS are independently coded without using inter-layer prediction.
[0285] ols_output_layer_flag[i][j] equal to 1 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS when ols_mode_idc is equal to 2. ols_output_layer_flag[i][j] equal to 0 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS when ols_mode_idc is equal to 2.
[0286] The variable NumOutputLayersInOls[i] that specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] that specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] that specifies whether the k-th layer is used as an output layer in at least one OLS are derived as follows:
Number
[0287] For each value of i in the range of 0 to vps_max_layers_minus1 including both ends, the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] do not both have to be equal to 0. In other words, there does not have to be a layer that is neither an output layer of at least one OLS nor a direct reference layer of any other layer.
[0288] For each OLS, there may be at least one layer that is an output layer. In other words, for any value of i in the range of 0 to TotalNumOlss - 1 including both ends, the value of NumOutputLayersInOls[i] may be 1 or more.
[0289] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS are derived as follows:
Number
[0290] The variable OlsLayerIdx[i][j] that specifies the OLS layer index of the layer having nuh_layer_idequal equal to LayerIdInOls[i][j] is derived as follows: [Number]
[0291] The lowest layer within each OLS may be an independent layer. In other words, for each i in the range of 0 to TotalNumOlss - 1 including both ends, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] may be equal to 1.
[0292] Each layer may be included in at least one OLS specified by the VPS. In other words, for each layer having a specific value of nuh_layer_idnuhLayerId equal to one of vps_layer_id[k] in the range of 0 to vps_max_layers_minus1 including both ends, there may exist at least one pair of values of i and j. Here, i is in the range of 0 to TotalNumOlss - 1 including both ends, and j is in the range of 0 to NumLayersInOls[i] - 1 including the ends, such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.
[0293] In an embodiment, the decoding process operates as follows for the current picture CurrPic: (1) PictureOutputFlag is set as follows: i) If one or more of the following conditions are true, PictureOutputFlag is set equal to 0: a) The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the related IRAP picture is equal to 1. b) The gdr_enabled_flag is equal to 1, and the current picture is a GDR picture having a NoOutputBeforeRecoveryFlag equal to 1. d) The gdr_enabled_flag is equal to 1, the current picture is a GDR picture having a NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is smaller than the RpPicOrderCntVal of the related GDR picture. e) The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions: · PicA has a PictureOutputFlag equal to 1. · PicA has a nuh_layer_id nuhLid greater than that of the current picture. · PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid). f) The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0. ii) In other cases, the PictureOutputFlag is set equal to pic_output_flag.
[0294] After all slices of the current picture have been decoded, the current decoded picture is marked as "used for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "used for short-term reference".
[0295] In some or other embodiments, when each layer is an output layer set, regardless of the value of valueofols_mode_idc, PictureOutputFlag is set equal to pic_output_flag.
[0296] In the same or another embodiment, when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is equal to 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions, PictureOutputFlag is set equal to 0: PicA has a PictureOutputFlag equal to 1, PicA has a nuh_layer_id nuhLid greater than that of the current picture, and PicA belongs to an output layer of the OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0297] In the same or another embodiment, when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is equal to 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0, PictureOutputFlag is set equal to 0.
[0298] FIG. 23 shows an example of a syntax table showing an output layer set having an output layer set mode indicator according to an embodiment.
[0299] FIG. 24 shows a block diagram of the decoding process of a bitstream according to an embodiment of the present disclosure. In particular, FIG. 24 shows a decoder-side flowchart showing an output layer set having an output layer set mode according to an embodiment.
[0300] According to an aspect of the present disclosure, the decoding method may include receiving a bitstream including compressed video / image data (operation 1001 in FIG. 24). The bitstream may have a plurality of layers.
[0301] The decoding method may further include operation 1002 of parsing or deriving an output layer set mode indicator (e.g., ols_mode_idc) within a video parameter set (VPS) from the bitstream.
[0302] The decoding method may further include operation 1003 of identifying output layer set signaling based on the output layer set mode indicator.
[0303] The decoding method may further include operation 1004 (e.g., operation 1004A, 1004B, or 1004C) of identifying one or more picture output layers based on the identified output layer set signaling.
[0304] The decoding method may further include operation 1005 of decoding the identified one or more picture output layers. The decoded one or more picture output layers may be displayed.
[0305] The step of identifying output layer set signaling based on the output layer set mode indicator may include identifying the highest layer in the bitstream as one or more picture output layers when the output layer set mode indicator in the VPS is a first value (see, for example, FIG. 25A), identifying all layers in the bitstream as one or more picture output layers when the output layer set mode indicator in the VPS is a second value (see, for example, FIG. 25B), and identifying one or more picture output layers based on explicit signaling in the VPS when the output layer set mode indicator in the VPS is a third value (see, for example, FIG. 25C).
[0306] Figures 25A to 25C show the output pictures with underlines. As shown in Figures 25A to 25C, there may be 5 layers in the bitstream, and only a specific layer may be output to the display. For example, layer 3 may be output (see Figure 25C for example).
[0307] According to an embodiment, the Ols_mode_idc in the VPS may indicate a method (mechanism) for signaling the output layer set. For example, when it is equal to 0, the highest layer in the bitstream may be the only output layer, when it is equal to 1, all the layers in the bitstream may be output layers, and when it is equal to 2, one or more output layers may be explicitly signaled in the VPS. That is, the picture output may be determined by the output layer set signaling.
[0308] For example, the picture output of each layer may be determined by the output layer set signaling, and the method of the output layer signaling may be determined by the ols_mode_idc.
[0309] According to an embodiment, each bitstream may signal an output layer mode, and the output layer mode may change over time.
[0310] According to an embodiment, as shown in Figure 25A, there may be 5 layers (layer 4 to layer 0) in the bitstream. As shown in Figure 25A, at time = K, the output layer set mode indicator (ols_mode_idc) = 0, and the highest layer may be output. Therefore, as shown in Figure 25A, at time = K, the output picture of layer 4 is output. At time = K + 1, layer 3, which is the highest layer, may be output, and the same applies to K + 2 and K + 3.
[0311] As shown in Figure 25B, the output layer set mode indicator (ols_mode_idc) is equal to 1, and all the layers may be output as in each of K to K + 3.
[0312] As shown in FIG. 25C, explicit signaling may be performed. For example, as shown in FIG. 25C, the picture layer 3 (shown as an output by an underline) is an output layer, for example, during times K to K+2.
[0313] The first value may be different from the second value, may be different from the third value, and the second value may be different from the third value.
[0314] The first value may be 0, the second value may be 1, and the third value may be 2. However, other values may be used, and the present disclosure is not limited to the use of 0, 1, and 2 as described above.
[0315] The step of identifying one or more picture output layers by explicit signaling within the VPS may include: (i) a step of parsing or deriving an output layer flag from the VPS; and (ii) a step of setting a layer having an output layer flag equal to 1 as one or more picture output layers.
[0316] The step of identifying output layer set mode signaling based on an output layer set mode indicator may include, when the output layer set mode indicator within the VPS is a predetermined value, the step of identifying one or more picture output layers by the output layer set mode signaling based on explicit signaling within the VPS.
[0317] The step of identifying one or more picture output layers based on explicit signaling within the VPS includes: (i) a step of parsing or deriving an output layer flag from the VPS; and (ii) a step of setting a layer having an output layer flag equal to 1 as one or more picture output layers, and the number of a plurality of layers is greater than 2.
[0318] The output layer set mode signaling may include, when the output layer set mode indicator is equal to 2 and the number of a plurality of layers is greater than 2, the step of identifying one or more picture output layers based on the explicit signaling within the VPS.
[0319] Output layer set mode signaling may include identifying the highest layer in the bitstream or all layers in the bitstream as one or more picture output layers when the output layer set mode indicator is less than 2 and the number of multiple layers is 2, where the output layer set mode indicator is actually less than 2 and the number of multiple layers is actually 2.
[0320] The output layer set number - 1 indicator in the VPS indicates the number of output layers.
[0321] According to an embodiment, the VPS maximum layer - 1 indicator in the VPS indicates the number of layers in the bitstream.
[0322] According to an embodiment, the output layer set mode flag [i][j] in the VPS indicates whether the j - th layer of the i - th output layer set is an output layer.
[0323] According to an embodiment, when multiple layers are independent layers and the VPS all independent layer flag of the VPS is equal to 1, the output layer set mode indicator is not signaled and the value of the output layer set mode indicator is presumed to be a second value.
[0324] According to an embodiment, when each layer is an output layer set, regardless of the value of the output layer set mode indicator, the picture output flag of the VPS is set equal to the picture output flag signaled in the picture header.
[0325] Note: A picture in the output layer may or may not have a PictureOutputFlag equal to 1. A picture in a non - output layer has a PictureOutputFlag equal to 0. A picture with a PictureOutputFlag equal to 1 is output for display. A picture with a PictureOutputFlag equal to 0 is not output for display.
[0326] According to an embodiment, when the sequence parameter set (SPS) VPS identifier is greater than 0 and indicates that more than 1 layer is present in the bitstream, the picture output flag is set equal to 0, and for each layer, when the output layer set mode flag of the VPS is equal to 0 and indicates that not all of the plurality of layers in the bitstream are independent, the output layer set mode indicator is equal to 0, and the current access unit includes a picture that satisfies all of the following conditions: having a picture output flag equal to 1; having a nuh layer identifier belonging to the output layer of the output layer set that is greater than that of the current picture.
[0327] According to an embodiment, when the sequence parameter set (SPS) of the VPS is greater than 0, the picture output flag of the VPS is set equal to 0, for each layer, the output layer set flag is equal to 0, the output layer set mode indicator is equal to 2, and the output layer set output layer flag [Target OLS Index][General Layer Index [nuh layer identifier]] is equal to 0.
[0328] According to an embodiment, the method may further include controlling a display to display one or more decoded picture output layers.
[0329] According to an aspect of the present disclosure, a non-transitory computer-readable storage medium storing instructions which, when executed, cause a system or apparatus including one or more processors to receive a bitstream including compressed video / image data, the bitstream including a plurality of layers, parse or derive an output layer set mode indicator in a video parameter set (VPS) from the bitstream, identify output layer set mode signaling based on the output layer set mode indicator, identify one or more picture output layers based on the identified output layer set mode signaling, and decode the identified one or more picture output layers.
[0330] According to an embodiment, the instructions are further configured to cause a display to be controlled to display the decoded one or more picture output layers by a system or apparatus including one or more processors.
[0331] According to an aspect of the present disclosure, the apparatus may include at least one memory storing computer program code, and at least one processor configured to access the at least one memory and operate according to the computer program code. According to an embodiment, the computer program code may include reception code configured to cause the at least one processor to receive a bitstream including compressed video / image data, the bitstream including a plurality of layers, parse or derive code configured to cause the at least one processor to parse or derive an output layer set mode indicator within a video parameter set (VPS) from the bitstream, output layer signaling identification code configured to cause the at least one processor to identify output layer set mode signaling based on the output layer set mode indicator, picture output layer identification code configured to cause the at least one processor to identify one or more picture output layers based on the identified output layer set mode signaling, and decoding code configured to cause the at least one processor to decode the identified one or more picture output layers.
[0332] According to an embodiment, the computer program code may further include display control code configured to cause the at least one processor to display one or more picture output layers.
[0333] The techniques for decoding, displaying, and signaling the adaptive resolution parameters described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 7 shows a computer system 700 suitable for implementing a particular embodiment of the subject matter of the present disclosure.
[0334] Computer software can be coded using any suitable machine code or computer language that can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed directly or through interpretation, microcode execution, etc. by a computer central processing unit (CPU), a graphics processing unit (GPU), etc.
[0335] Instructions can be executed on various computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.
[0336] The components shown in FIG. 7 of computer system 700 are illustrative in nature and do not imply any limitation on the use or functionality of the computer software implementing the embodiments of the present disclosure. Further, the component configuration should not be construed as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiments of computer system 700.
[0337] Computer system 700 may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, sensory input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by a human, such as audio (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a digital camera), video (including, for example, 2D video, 3D video, stereoscopic video).
[0338] The input human interface device may include one or more of a keyboard 701, a mouse 702, a trackpad 703, a touch screen 710, a data glove 704, a joystick 705, a microphone 706, a scanner 707, a camera 708 (only one of which is shown).
[0339] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., sensory feedback by the touch screen 710, data glove 704, or joystick 705, but there may also be sensory feedback devices that do not function as input devices), audio output devices (e.g., speakers 709, headphones (not shown)), visual output devices (e.g., including the screen 710, CRT screen, LCD screen, plasma screen, OLED screen, each with or without touch screen input capability, each with or without sensory feedback capability, some of which may output two-dimensional visual output or three-dimensional or higher-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown)).
[0340] The computer system 700 may also include a human-accessible memory device and associated media such as an optical medium like a CD / DVDROM / RW 720 with a medium 721 such as a CD / DVD, a thumb drive 722, a removable hard drive or solid-state drive 723, legacy magnetic media such as tapes and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown).
[0341] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0342] Computer system 700 may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network generally requires an external network interface attached to a particular general-purpose data port or peripheral device bus (749) (e.g., a USB port of computer system 700). Others are generally integrated into the core of computer system 700 by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using these networks, computer system 700 can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANBus to a particular CANBus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used with each of the networks and network interfaces described above.
[0343] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core 740 of computer system 700.
[0344] The core 740 may include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a dedicated programmable processing unit 743 in the form of an FPGA, a hardware accelerator 744 for specific tasks, and the like. These devices may be connected through a system bus 748 together with a built-in mass storage device 747 such as a read-only memory (ROM) 745, a random access memory 746, an internal hard drive that is not accessible to users, an SSD, etc. In some computer systems, the system bus 748 is accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core system bus 748 or through a peripheral device bus 749. Architectures of peripheral device buses include PCI, USB, etc.
[0345] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can be combined to execute specific instructions that can generate the aforementioned computer code. The computer code can be stored in the ROM 745 or the RAM 746. Temporary data can also be stored in the RAM 746, while permanent data can be stored, for example, in the built-in mass storage device 747. Fast storage and reading to any of the memory devices can be enabled through the use of a cache memory that can be closely associated with one or more CPUs 741, GPUs 742, the mass storage device 747, the ROM 745, the RAM 746, etc.
[0346] A computer-readable medium may have computer code for performing operations implemented by various computers. The medium and the computer code may be specially designed and configured for the purposes of the present disclosure, or may be of the kind well known and available to those skilled in the field of computer software.
[0347] By way of example and not limitation, a computer system 700 having an architecture, and specifically a core 740, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied within one or more tangible computer-readable media. Such computer-readable media can be associated with a specific storage device of the core 740 having non-transitory characteristics such as the core built-in mass storage device 747 or ROM 745, and media associated with a user-accessible mass storage device as described above. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 740. The computer-readable media can include one or more memory devices or chips as required by a particular need. The software can cause the core 740 and specifically the processor (including a CPU, GPU, FPGA, etc.) therein to execute a specific process or a specific part of a specific process described herein, including the definition of a data structure stored in the RAM 746 and the modification of the data structure according to the process defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of an implementation (e.g., accelerator 744) within logic hardwired or other circuitry operable with or instead of the software to execute a specific process or a specific part of a specific process described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can, where appropriate, include circuitry (such as an integrated circuit (IC)) for storing software for execution, circuitry for implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0348] The present disclosure has described several exemplary embodiments, but there are alternatives, substitutions, and various equivalents that are encompassed by the scope of the present disclosure. It will be apparent to those skilled in the art that many systems and methods can be devised that, while not explicitly shown or described herein, implement the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method, executed by at least one processor, for encoding video data into a coded bitstream, the coded bitstream including a plurality of layers; generating a syntax element ols_mode_idc in a video parameter set (VPS); when the syntax element sps_video_parameter_set_id is greater than 0, the syntax element each_layer_is_an_ols_flag is equal to 0, and the syntax element ols_mode_idc is equal to 2, setting the variable PictureOutputFlag of the current picture equal to 0; otherwise, setting the variable PictureOutputFlag of the current picture to the value of the syntax element pic_output_flag signaled in a picture header; A method comprising:
2. A method executed by at least one processor, comprising: generating a coded bitstream, the coded bitstream including a plurality of layers, the generating the coded bitstream comprising: generating a syntax element ols_mode_idc in a video parameter set (VPS); when the syntax element sps_video_parameter_set_id is greater than 0, the syntax element each_layer_is_an_ols_flag is equal to 0, and the syntax element ols_mode_idc is equal to 2, setting the variable PictureOutputFlag of the current picture equal to 0; otherwise, setting the variable PictureOutputFlag of the current picture to the value of the syntax element pic_output_flag signaled in a picture header; generating the coded bitstream, transmitting the coded video bitstream or storing the coded video bitstream on a storage medium; A method comprising:
3. A decoding method executed by at least one processor, the method comprising: receiving a bitstream, the bitstream including a plurality of layers; determining a syntax element ols_mode_idc in a video parameter set (VPS) from the bitstream; decoding one or more layers of the current picture; wherein the decoding step comprises: when a syntax element sps_video_parameter_set_id is greater than 0, a syntax element each_layer_is_an_ols_flag is equal to 0, and the syntax element ols_mode_idc is equal to 2, setting a variable PictureOutputFlag of the current picture equal to 0; otherwise, setting the variable PictureOutputFlag of the current picture to the syntax element pic_output_flag signaled in a picture header; A method comprising: