Encoding method, encoding device, decoding method, decoding device, and program

Adaptive resolution change signaling in video encoding and decoding addresses the challenge of handling varying scene activities by enabling flexible resolution adjustments within a single picture, enhancing efficiency in applications like 360-degree coding and surveillance.

JP7897401B2Active Publication Date: 2026-07-29TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2025-08-14
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to efficiently handle adaptive resolution changes within a single picture, particularly in applications like 360-degree coding and surveillance, where multiple semantically independent source pictures require separate resolution settings due to varying scene-specific activities.

Method used

The implementation of adaptive resolution change (ARC) signaling mechanisms that allow for resampling reference pictures and signaling output layer sets in coded video data, enabling flexible resolution adjustments within a single picture.

Benefits of technology

Enhances video coding and decoding efficiency by allowing adaptive resolution changes, improving the handling of diverse scene activities in applications such as 360-degree coding and surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897401000001
    Figure 0007897401000001
  • Figure 0007897401000002
    Figure 0007897401000002
  • Figure 0007897401000003
    Figure 0007897401000003
Patent Text Reader

Abstract

To provide a method, a system, and a computer program for signaling output layer sets in coded video data.SOLUTION: A method, a computer program, and a computer system are provided for signaling an output layer set in an encoded video stream. Video data having multiple layers is received. One or more syntactic elements are identified. The syntactic elements specify one or more output layer sets that correspond to output layers from among multiple layers of received video data. One or more output layers corresponding to the specified output layer set are decoded and displayed.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 903660, filed on September 20, 2019, and U.S. Patent Application No. 17 / 021243, filed on September 15, 2020, which are incorporated in their entirety herein.

[0002] This disclosure generally relates to the field of video coding and decoding, and more particularly to parameter set references and ranges in coded video streams. [Background technology]

[0003] Video coding and decoding using picture-to-picture prediction with motion compensation has been known for decades. Uncompressed digital video can consist of a series of pictures, each having spatial dimensions of, for example, 1920 x 1080 luminance samples and associated chromaticity samples. The series of pictures can have a fixed or variable picture rate (informally also known as frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has considerable bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920 x 1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GByte of storage space.

[0004] One purpose of video encoding and decoding may be to reduce the redundancy of the input video signal through compression. Compression can help reduce the aforementioned bandwidth or storage space requirements by more than two orders of magnitude, in some cases. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique that allows an exact copy of the original signal to be reconstructed from the compressed original signal. With lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for its intended purpose. For video, lossy compression is widely used. The amount of distortion that is acceptable depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television broadcast applications. The achievable compression ratio may reflect that a higher acceptable / tolerable distortion allows for a higher compression ratio.

[0005] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding, some of which are described below.

[0006] Historically, video encoders and decoders have almost always tended to operate with a given picture size defined for encoded video sequences (CVS), Groups of Pictures (GOP), or similar multi-picture timeframes, and which remain constant for these. For example, in MPEG-2, system design is known to change the horizontal resolution (and thus the picture size) depending on factors such as scene activity, but only for individual pictures, and therefore typically for GOPs. Resampling of a reference picture to use different resolutions within a CVS is known, for example, from ITU-T Recommendation H.263 Annex P. However, in this case the picture size does not change, only the reference picture is resampled, and potentially only a portion of the picture canvas is used (in the case of downsampling) or only a portion of the scene is captured (in the case of upsampling). Furthermore, H.263 Annex Q allows for resampling individual macroblocks upward or downward by a factor of 2 (in each dimension). Again, the picture size remains the same. The size of macroblocks is fixed in H.263 and therefore does not need to be signaled.

[0007] Changing the picture size of a predicted picture has become more common in modern video encoding. For example, VP9 allows for the resampling of a reference picture and the changing of the resolution of the entire picture. Similarly, several proposals made for VVC (including, for example, Hendry, et al., “On adaptive resolution change (ARC) for VVC”, Joint Video Team document JVET-M0135-v1, January 9-19, 2019, which is incorporated herein in its entirety) allow for the resampling of the entire reference picture to different, i.e., higher or lower resolutions. In that document, it is proposed that different candidate resolutions be encoded within a sequence parameter set and referenced by per-picture syntactic elements within a picture parameter set. [Overview of the project] [Means for solving the problem]

[0008] Embodiments relate to methods, systems, and computer-readable media for signaling output layer sets in coded video data. According to one embodiment, a method is provided for signaling output layer sets in coded video data. This method may include the steps of receiving video data having multiple layers; identifying one or more syntactic elements; specifying one or more output layer sets corresponding to output layers from among the multiple layers of the received video data; decoding and displaying one or more output layers corresponding to the specified output layer sets.

[0009] In another embodiment, a computer system is provided for signaling output layer sets in coded video data. This computer system may include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, thereby enabling the computer system to perform a method. This method may include the steps of receiving video data having multiple layers; one or more syntactic elements being identified; the syntactic elements specifying one or more output layer sets corresponding to output layers from among the multiple layers of the received video data; one or more output layers corresponding to the specified output layer sets being decoded and displayed.

[0010] In yet another embodiment, a computer-readable medium is provided for signaling output layer sets in coded video data. This computer-readable medium may include one or more computer-readable storage devices and program instructions, which are executable by a processor and stored in at least one of one or more tangible storage devices. The program instructions are executable by the processor to perform a method which may include the step of receiving video data having a plurality of corresponding layers. One or more syntactic elements are identified. The syntactic elements specify one or more output layer sets corresponding to output layers from among the plurality of layers of the received video data. One or more output layers corresponding to the specified output layer sets are decoded and displayed.

[0011] The aforementioned and other purposes, features, and advantages will become apparent from the following detailed description of exemplary embodiments, which should be read in conjunction with the accompanying drawings. Since the examples are for clarity to facilitate understanding for those skilled in the art, together with the detailed description, various features in the drawings are not to exact scale. [Brief explanation of the drawing]

[0012] [Figure 1] This is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 2] This is a schematic diagram of a simplified block diagram of a communication system according to one embodiment. [Figure 3] This is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 4] This is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 5] This is a schematic diagram of options for signaling ARC parameters according to one embodiment. [Figure 6] This is a diagram illustrating an example of a syntax table according to one embodiment. [Figure 7] This is a schematic diagram of a computer system according to one embodiment. [Figure 8] FIG. is an example of a prediction structure for scalability using adaptive resolution change. [Figure 9] FIG. is an example of a syntax table according to an embodiment. [Figure 10] FIG. is a schematic diagram of a simplified block diagram of parsing and decoding poc cycles and access unit count values per access unit. [Figure 11] FIG. is a schematic diagram of a video bitstream structure including a multi-layer subpicture. [Figure 12] FIG. is a schematic diagram of the display of a selected subpicture having an enhanced resolution. [Figure 13] FIG. is a block diagram of a decoding and display process of a video bitstream including a multi-layer subpicture. [Figure 14] FIG. is a schematic diagram of a 360-degree video display having an enhancement layer of a subpicture. [Figure 15] FIG. is an example of a diagram of subpicture layout information and its corresponding layer and picture prediction structure. [Figure 16] FIG. is an example of a diagram of subpicture layout information and its corresponding layer and picture prediction structure by a spatial scalability modality of a local area. [Figure 17] FIG. is an example of a syntax table of subpicture layout information. [Figure 18] FIG. is an example of a syntax table of an SEI message of subpicture layout information. [Figure 19] FIG. is an example of a syntax table showing output layers and profile / tiers / levels information for each output layer set. [[ID=三十七]] [Figure 20] FIG. is an example of a syntax table showing an output layer mode for each output layer set. [Figure 21] FIG. is an example of a syntax table showing the current subpicture of each layer for each output layer set.

BEST MODE FOR CARRYING OUT THE INVENTION

[0013] This specification discloses detailed embodiments of the claimed structures and methods. However, it should be understood that the disclosed embodiments are merely illustrative of the claimed structures and methods, which may be embodied in various forms. These structures and methods may, however, be embodied in many different forms and should not be construed as being limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided to ensure that this disclosure is detailed and complete and fully conveys its scope to those skilled in the art. In this specification, well-known features and technical details may be omitted to avoid unnecessarily obscuring the embodiments presented.

[0014] The embodiments generally relate to the field of data processing, and more specifically to media processing. The exemplary embodiments described below provide, in particular, systems, methods, and computer programs that enable signaling of output layer sets of coded video data. Thus, some embodiments have the ability to improve the computing field through improved video coding and decoding.

[0015] As mentioned earlier, video encoders and decoders have generally tended to operate with a given picture size defined for encoded video sequences (CVS), Groups of Pictures (GOP), or similar multi-picture timeframes, which remain constant for these. For example, in MPEG-2, the system design is known to change the horizontal resolution (and thus the picture size) depending on factors such as scene activity, but only for individual pictures, and therefore typically for GOPs. Resampling of a reference picture to use different resolutions within a CVS is known, for example, from ITU-T Recommendation H.263 Annex P. However, in this case the picture size does not change, only the reference picture is resampled, and potentially only a portion of the picture canvas is used (in the case of downsampling) or only a portion of the scene is captured (in the case of upsampling). Furthermore, H.263 Annex Q allows for resampling individual macroblocks upward or downward by a factor of 2 (in each dimension). Again, the picture size remains the same. The size of macroblocks is fixed in H.263 and therefore does not need to be signaled.

[0016] However, in the context of 360-degree coding or certain surveillance applications, for example, multiple semantically independent source pictures (e.g., six cubic surfaces of a cubically projected 360-degree scene, or individual camera inputs in the case of a multi-camera surveillance system) may require separate adaptive resolution settings to handle different scene-specific activities at a given point in time. In other words, an encoder may choose to use different resampling factors for different semantically independent pictures that make up the entire 360-degree or surveillance scene at a given point in time. When combined into a single picture, it further requires that the reference picture be resampled and that adaptive resolution coding signaling be available for each portion of the coded picture. Therefore, it may be advantageous to use available adaptive resolution coding signaling data for better signaling, coding, decoding, and display of the video layer.

[0017] Figure 1 shows a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The system (100) may include at least two terminals (110-120) interconnected via a network (150). For unidirectional data transmission, a first terminal (110) may code video data at its local location for transmission to another terminal (120) via the network (150). A second terminal (120) may receive the coded video data from the other terminal via the network (150), decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.

[0018] Figure 1 shows a second pair of terminals (130, 140) provided to support the bidirectional transmission of coded video, for example, during a video conference. In bidirectional data transmission, each terminal (130, 140) can encode video data captured at its local location for transmission to the other terminal over the network (150). Each terminal (130, 140) can also receive coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.

[0019] In Figure 1, terminals (110-140) may be exemplified as servers, personal computers, and smartphones, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure apply to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (150) represents any number of networks that transmit coded video data between terminals (110-140), including, for example, wired and / or wireless communication networks. Communication network (150) may exchange data over circuit-switched channels and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. In this consideration, the architecture and topology of network (150) may not be important to the operation of this disclosure unless described below herein.

[0020] Figure 2 shows an example of an application of the subject matter of disclosure, illustrating the arrangement of a video encoder and decoder in a streaming environment. The subject matter of disclosure may also be equally applicable to other video-enabled applications, such as video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, and memory sticks.

[0021] The streaming system may include a capture subsystem (213) which may include, for example, a digital camera that creates a video source (201), for example, an uncompressed video sample stream (202). The sample stream (202) is shown in thick lines to highlight its higher data volume compared to the encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the subject of disclosure as described in more detail below. The encoded video bitstream (204) is shown in thin lines to highlight its lower data volume compared to the sample stream and can be stored in a streaming server (205) for future use. One or more streaming clients (206, 208) may access the streaming server (205) to obtain a copy (207, 209) of the encoded video bitstream (204). A client (206) may include a video decoder (210) that decodes an incoming copy of an encoded video bitstream (207) and creates an outgoing video sample stream (211) that can be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstream (204, 207, 209) can be encoded according to a certain video encoding / compression standard. An example of such a standard is ITU-T Recommendation H.265. A video encoding standard informally known as Versatile Video Coding or VVC is under development. The subject of this disclosure may be used in the context of VVC.

[0022] Figure 3 may be a functional block diagram of a video decoder (210) according to one or more embodiments.

[0023] The receiver (310) may receive one or more codec video sequences to be decoded by the decoder (210), which in the same or different embodiments may be one encoded video sequence at a time, with the decoding of each encoded video sequence being independent of other encoded video sequences. The encoded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device storing coded video data. The receiver (310) may receive coded video data with other data, e.g., coded audio data and / or auxiliary data streams, which may be transferred to each other using entities (not shown). The receiver (310) may isolate the coded video sequences from other data. To counteract network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter "Parser"). When the receiver (310) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may not be necessary or can be made small. For use in best-effort packet networks such as the Internet, the buffer (315) may be necessary and can be made relatively large, or advantageously, adaptively sized.

[0024] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-encoded video sequence. These categories of symbols may include information used to manage the operation of the decoder (210), and potentially information for controlling rendering devices such as a display (212), which is not an integral part of the decoder but can be coupled to it, as shown in Figure 2. Control information for (one or more) rendering devices may take the form of Supplementary Enhancement Information (SEI messages) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (320) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may conform to video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, context-sensitive or non-context-sensitive arithmetic coding, etc. The parser (320) may extract from the encoded video sequence a set of at least one subgroup parameters of a subgroup of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), predictive unit (PU), etc. The entropy decoder / parser may also extract encoded video sequence information such as transform coefficients, quantizer parameter values, and motion vectors.

[0025] The parser (320) may perform an entropy decoding / parse operation on the video sequence received from the buffer (315) in order to create a symbol (321).

[0026] The reconstruction of symbol (321) may involve multiple different units depending on the type of coded video picture or part thereof (such as interpictures and intrapictures, or interblocks and intrablocks), as well as other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (320). The flow of such subgroup control information between the parser (320) and the following multiple units is not illustrated for clarity.

[0027] Beyond the functional blocks already described, the decoder 210 can be conceptually subdivided into several functional units, as described below. In actual implementations operating under commercial constraints, many of these units can interact closely with each other and, at least partially, be integrated with one another. However, for the purpose of illustrating the subject of disclosure, the following conceptual subdivision into functional units is appropriate.

[0028] The first unit is the scaler / inverse unit (351). The scaler / inverse unit (351) receives control information from the parser (320) as one or more symbols (321), including the quantization transformation coefficients, which transformation to use, the block size, the quantization coefficients, and the quantization scaling matrix. The scaler / inverse unit (351) can output a block containing sample values ​​that can be input to the aggregator (355).

[0029] In some cases, the output samples of the scaler / inverse transform (351) may relate to intra-encoded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information extracted from the current (partially reconstructed) picture (356). The aggregator (355) may optionally append the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351), sample by sample.

[0030] In other cases, the output samples of the scaler / inverse unit (351) may relate to intercoded, potentially motion-compensated blocks. In such cases, the motion-compensated prediction unit (353) can access the reference picture memory (357) to retrieve samples to be used for prediction. After motion-compensating the retrieved samples according to the symbols (321) related to the blocks, these samples can be appended by the aggregator (355) to the output of the scaler / inverse unit (in this case, called residual samples or residual signals) to generate output sample information. The address of the reference picture memory from which the motion-compensated unit retrieves the prediction samples can be controlled by a motion vector, which is available to the motion-compensated unit in the form of a symbol (321) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory when the exact motion vectors of the subsamples are used, a motion vector prediction mechanism, and so on.

[0031] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technique may include in-loop filtering techniques, which are controlled by parameters included in the encoded video bitstream and provided to the loop filter unit (356) as symbols (321) from the parser (320), but may also depend on metadata obtained during decoding of earlier parts (in decoding order) of the coded picture or coded video sequence, or on previously reconstructed and loop-filtered sample values.

[0032] The output of the loop filter unit (356) can be output to the rendering device (212) and can also be a sample stream that can be stored in the reference picture memory (356) for use in future picture-to-picture predictions.

[0033] A coded picture, once fully reconstructed, can be used as a reference picture for future predictions. Once a coded picture is fully reconstructed and identified as a reference picture (for example, by a parser (320)), the current reference picture (356) can become part of the reference picture buffer (357), allowing the memory of the new current picture to be reallocated before the reconstruction of subsequent coded pictures begins.

[0034] The video decoder 320 may perform decoding operations according to a specified video compression technique that may be documented in a standard, such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified in the video compression technique or standard being used, in the sense that it conforms to the syntax of the video compression technique or standard, as specified in the video compression technique documentation or standard, particularly the profile documentation within it. Compliance may also require that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, in some cases, be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the encoded video sequence.

[0035] In one embodiment, the receiver (310) may receive additional (redundant) data having coded video. The additional data may be included as part of (one or more) coded video sequences. The additional data may be used by the video decoder (320) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, time, space, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0036] Figure 4 may be a functional block diagram of a video encoder (203) according to one embodiment of the present disclosure.

[0037] The encoder (203) may receive video samples from a video source (201) that is not part of the encoder and can capture (one or more) video images to be encoded by the encoder (203).

[0038] The video source (201) may provide a source video sequence to be encoded by the encoder (203) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a series of individual pictures that convey motion when viewed sequentially. The picture itself may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description will focus on samples.

[0039] According to one embodiment, the encoder (203) can encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. One function of the controller (450) is to ensure that an appropriate encoding rate is maintained. The controller controls and is functionally coupled to other functional units, as described below. The couplings are not illustrated for clarity. Parameters set by the controller may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will readily identify other functions of the controller (450) that may be relevant to a video encoder (203) optimized for a particular system design.

[0040] Some video encoders operate in what is readily recognizable to those skilled in the art as a “coding loop.” In an overly simplified explanation, the coding loop may consist of an encoding portion of the encoder (430) (hereinafter referred to as the “source coder”) (responsible for creating symbols based on the input picture to be coded and one or more reference pictures) and a (local) decoder (433) incorporated into the encoder (203) that reconstructs the symbols to create sample data, which will also create a (remote) decoder (since any compression between the symbols and the encoded video bitstream is reversible in the video compression techniques considered in the subject of the disclosure). The reconstructed sample stream is input to the reference picture memory (434). Since decoding the symbol stream yields bit-strict results regardless of the decoder location (local or remote), the contents of the reference picture buffer are also bit-strict between the local and remote encoders. In other words, the prediction portion of the encoder “sees” the exact same sample values ​​as reference picture samples that the decoder “sees” when using predictions during decoding. This basic principle of the synchronization of a reference picture (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.

[0041] The operation of the “local” decoder (433) may be the same as the operation of the “remote” decoder (210), which has already been described in detail in relation to Figure 3. However, as also briefly referring to Figure 3, since symbols are available and the encoding / decoding of symbols to an encoded video sequence by the entropy coder (445) and parser (320) may be reversible, the entropy decoding portion of the decoder (210), including the channel (312), receiver (310), buffer (315), and parser (320), may not be fully implemented in the local decoder (433).

[0042] An observation that can be made at this point is that any decoder technique present in the decoder, excluding parse / entropy decoding, must also exist in the corresponding encoder in substantially the same functional form. For this reason, the subject of the disclosure focuses on decoder operation. The description of encoder techniques can be omitted as they are the inverse of the decoder techniques that have been comprehensively described. More detailed explanations are necessary only in specific areas, which are shown below.

[0043] As part of this operation, the source coder (430) may perform motion-compensated predictive coding, predictively coding the input frame by referencing one or more previously coded frames from a video sequence designated as “reference frames”. In this way, the coding engine (432) codes the difference between the pixel blocks of the input frame and the pixel blocks of one or more reference frames that can be selected as prediction criteria (one or more) for the input frame.

[0044] The local video decoder (433) can decode coded video data of a frame that may be designated as a reference frame based on symbols created by the source coder (430). The operation of the coding engine (432) may be advantageously a lossy process. If coded video data can be decoded by a video decoder (not shown in Figure 4), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process that may be performed by the video decoder on the reference frame, which may cause the reconstructed reference frame to be stored in the reference picture cache (434). Thus, the encoder (203) may locally store a copy of the reconstructed reference frame that has common content as the reconstructed reference frame acquired by the far-end video decoder (without transmission errors).

[0045] The predictor (435) may perform a predictive search for the coding engine (432). That is, for a new frame to be coded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or metadata such as motion vectors and block shapes of reference pictures that can function as appropriate prediction criteria for the new picture. The predictor (435) may operate on a sample block pixel block-by-block basis to find appropriate prediction criteria. In some cases, the input picture may have prediction criteria drawn from multiple reference pictures stored in the reference picture memory (434), as determined by the search results obtained by the predictor (435).

[0046] The controller (450) may manage the coding operations of the video coder (430), including, for example, setting parameters and subgroup parameters used to encode video data.

[0047] The outputs of all the aforementioned functional units can undergo entropy coding in the entropy coder (445). The entropy coder converts the symbols generated by the various functional units into coded video sequences by lossless compression of the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, and arithmetic coding.

[0048] A transmitter (440) may buffer (one or more) encoded video sequences created by the entropy coder (445) in preparation for transmission over a communication channel (460), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) may merge the encoded video data from the video coder (430) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0049] The controller (450) may manage the operation of the encoder (203). During encoding, the controller (450) may assign each coded picture a coded picture type that may affect the coding technique that can be applied to that picture. For example, a picture may often be assigned as one of the following frame types:

[0050] An intra-picture (I-picture) can be a picture that can be encoded and decoded without using other frames in a sequence as a source of prediction. Some video codecs enable different types of intra-pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are familiar with their variations of I-pictures and their respective uses and characteristics.

[0051] A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-prediction or inter-prediction, which predicts the sample value of each block using at most one motion vector and reference index.

[0052] A bidirectional predictive picture (B-picture) can be a picture that can be encoded and decoded using intra-prediction or inter-prediction, which predicts the sample values ​​of each block using up to two motion vectors and reference indices. Similarly, a multi-predictive picture can use three or more reference pictures and associated metadata to reconstruct a single block.

[0053] A source picture can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block can be coded. Blocks can be coded predictively by referencing other (already coded) blocks, as determined by the coding assignment applied to each picture in the block. For example, blocks of picture I can be coded unpredictably or predictively by referencing already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of picture P can be coded unpredictably by spatial prediction or temporal prediction by referencing one previously coded reference picture. Blocks of picture B can be coded unpredictably by spatial prediction or temporal prediction by referencing one or two previously coded reference pictures.

[0054] The videocoder (203) may perform encoding operations in accordance with a specified video encoding technique or standard, such as ITU-T Recommendation H.265. In doing so, the videocoder (203) may perform various compression operations, including predictive encoding operations that utilize temporal and spatial redundancy in the input video sequence. The encoded video data may therefore conform to the syntax specified by the video encoding technique or standard being used.

[0055] In one embodiment, the transmitter (440) may transmit additional data along with the coded video. The video coder (430) may include such data as part of the coded video sequence. The additional data may include a time / space / SNR enhancement layer, other forms of redundant data such as redundant pictures and redundant slices, supplemental enhancement information (SEI) messages, and fragments of visual usability information (VUI) parameter sets.

[0056] Before describing in more detail specific aspects of the subject matter of the disclosure, it is necessary to explain some terms that will be referenced in the remainder of this specification.

[0057] A subpicture, in the following context, refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that, in some cases, are semantically grouped and can be independently coded at a modified resolution. A picture may have one or more subpictures. One or more coded subpictures may form a coded picture. One or more subpictures may be assembled into a picture, and one or more subpictures may be extracted from a picture. In some environments, one or more coded subpictures may be assembled into a coded picture in a compressed region without code conversion to the sample level, and in the same case, or in some other cases, one or more coded subpictures may be extracted from a coded picture in a compressed region.

[0058] Adaptive Resolution Change (ARC) refers, below, to a mechanism that allows for changes in the resolution of a picture or sub-picture within an encoded video sequence, for example, by resampling a reference picture. ARC parameters refer to the control information required to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, the resolution of the output picture and / or reference picture, and various control flags.

[0059] The above explanation focuses on encoding and decoding a single, semantically independent coded video picture. Before discussing the implications of encoding / decoding multiple subpictures with independent ARC parameters and the additional complexities that this implies, we will describe the options for signaling ARC parameters.

[0060] Referring to Figure 5, several novel options for signaling ARC parameters are shown. As can be seen from each option, they have some advantages and some disadvantages in terms of encoding efficiency, complexity, and architecture. A video encoding standard or technology may select one or more of these options, or options known from prior art, to signal ARC parameters. These options may not be mutually exclusive and, to the extent possible, may be interchangeable based on the needs of the application, the standards technology involved, or the choice of encoder.

[0061] The classes of ARC parameters may include the following: Upsampling / downsampling coefficients separated or combined in the X and Y dimensions. Upsampling / downsampling coefficients with added time dimension, representing constant-speed zoom in / out of a given number of pictures. Either of the above two may involve the coding of one or more possibly short syntactic elements that can point to a table containing (one or more) those coefficients. The X-dimensional or Y-dimensional resolution of the combined or separate input picture, output picture, reference picture, and coded picture, in units of samples, blocks, macroblocks, CUs, or any other appropriate granularity. If there are two or more resolutions (e.g., the resolution of the input picture, the resolution of the reference picture, etc.), then in some cases one set of values ​​may be inferred from another set of values. This can be gated, for example, by using flags. See below for more detailed examples. Again, the appropriate granularity of "warping" coordinates, similar to those used in H.263 Annex P, as described above. While H.263 Annex P defines one efficient method for encoding such warping coordinates, other potentially more efficient methods may be devised. For example, the variable-length, reversible "Huffman" style coding of warping coordinates in Annex P could be replaced with a binary coding of appropriate length, where the length of the binary codeword could be derived, for example, from the maximum picture size, possibly multiplied by a coefficient and offset by a certain value to enable "warping" beyond the boundaries of the maximum picture size. Upsampling or downsampling of filter parameters. In the simplest case, only a single filter for upsampling and / or downsampling may exist. However, in some cases, it may be advantageous to allow more flexibility in filter design, which may require signaling of filter parameters. Such parameters may be selected via an index in a list of possible filter designs, the filter may be fully specified (e.g., via a list of filter coefficients using appropriate entropy coding techniques), the filter may be implicitly selected via an upsampling / downsampling ratio that matches what is signaled according to one of the mechanisms described above, and so on.

[0062] In the following explanation, we assume the coding of a finite set of upsampling / downsampling coefficients (the same coefficients to be used in both the X and Y dimensions) represented by a codeword. This codeword can, advantageously, be variable-length coded using Ext-Golomb coding, which is common to several syntactic elements in video coding specifications such as H.264 and H.265.

[0063] Many similar mappings can be devised depending on the application requirements and the capabilities of the upscaling and downscaling mechanisms available in the video compression technology or standard. The table can be extended to more values. The values ​​may also be represented by an entropy coding mechanism other than Ext-Golomb coding, for example, using binary coding. This may have some advantages when the resampling coefficient is an external object of the video processing engine (firstly the encoder and decoder) itself, as can be seen with MANE, for example. In the (perhaps) most common case where resolution change is not required, a short Ext-Golomb code can be chosen, and it should be noted that in the table above, this is only 1 bit. This may have the advantage of coding efficiency over using binary code in the most common case.

[0064] The number of entries in a table, as well as their semantics, can be fully or partially configurable. For example, the basic outline of a table may be conveyed by a “high” parameter set, such as a sequence parameter set or a decoder parameter set. Alternatively or additionally, one or more such tables may be defined in a video coding technique or standard, and may be selected, for example, via a decoder parameter set or a sequence parameter set.

[0065] The following explains how the upsampling / downsampling coefficients (ARC information) encoded as described above may be incorporated into the syntax of video coding techniques or standards. Similar considerations may apply to one or more codewords that control upsampling / downsampling filters. For considerations when filters or other data structures require relatively large amounts of data, see below.

[0066] H.263 Annex P includes ARC information 502 in the form of four warping coordinates within the picture header 501, specifically within the H.263 PLUSPTYPE(503) header extension. This can be a sensible design choice when a) there is a picture header available and b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be very high, and because picture headers can be transient, scaling factors may not be relevant across picture boundaries.

[0067] The JVCET-M135-v1 cited above includes ARC reference information (505) (index) located in the picture parameter set (504), which points to a table (506) containing target resolutions located in the sequence parameter set (507). The arrangement of possible resolutions in table (506) within the sequence parameter set (507) can be justified, according to the authors, by using SPS as a point of negotiation for interoperability during capability exchange. Resolutions can vary for each picture within the limits set by the values ​​in table (506) by referring to the appropriate picture parameter set (504).

[0068] Referring again to Figure 5, the following additional options may exist for transmitting ARC information in a video bitstream. Each of these options has several advantages over existing technologies, as mentioned above. These options may coexist simultaneously in the same video encoding technology or standard.

[0069] In one embodiment, ARC information (509), such as a resampling (zoom) coefficient, may reside in a slice header, GOB header, tile header, or tile group header (hereinafter, tile group header) (508). This may be sufficient if the ARC information is small, such as a single variable-length ue(v) or a fixed-length codeword of a few bits, as shown above. Having the ARC information directly within the tile group header has the additional advantage that the ARC information may be applicable to, for example, the subpictures represented by that tile group, rather than the entire picture. See also below. In addition, even if the video compression technology or standard assumes only whole-picture adaptive resolution change (as opposed to, for example, tile group-based adaptive resolution change), there are several advantages from the standpoint of fault tolerance to placing the ARC information in the tile group header compared to placing it in an H.263-style picture header.

[0070] In the same or a different embodiment, the ARC information (512) itself may reside in a suitable parameter set (511), such as a picture parameter set, a header parameter set, a tile parameter set, an adaptive parameter set, etc. (the adaptive parameter set is illustrated). The scope of the parameter set may, advantageously, be below the picture, for example, a tile group. The use of the ARC information is implicit through the activation of the relevant parameter set. For example, if the video coding technology or standard intends only picture-based ARC, then a picture parameter set or equivalent may be appropriate.

[0071] In the same or a different embodiment, the ARC reference information (513) may reside in a tile group header (514) or a similar data structure. The reference information (513) may refer to a portion of the ARC information (515) available in a parameter set (516) that has a range beyond a single picture, for example, a sequence parameter set or a decoder parameter set.

[0072] The indirectly suggested activation of additional levels of PPS from tile group headers, PPS, and SPS, as used in JVET-M0135-v1, seems unnecessary since picture parameter sets can be used for capability negotiation or notification, similar to sequence parameter sets (and have in some standards such as RFC3984). However, if the ARC information should be applicable to subpictures that are also represented by tile groups, for example, then parameter sets with activation ranges limited to tile groups, such as adaptive parameter sets or header parameter sets, may be a better choice. Also, if the ARC information is of a significant size, for example, if it includes filter control information such as a large number of filter coefficients, then parameters may be a better choice than directly using the header (508) from an encoding efficiency standpoint, since these settings can be reused by future pictures or subpictures by referring to the same parameter set.

[0073] When using a sequence parameter set or another higher-level parameter set that spans multiple pictures, several considerations may apply.

[0074] The parameter set for storing the ARC information table (516) may, in some cases, be a sequence parameter set, but in other cases, it is preferable to be a decoder parameter set. The decoder parameter set may have multiple CVSs, i.e., the activation range of the encoded video stream, i.e., all coded video bits from the start to the end of the session. Such a range may be more appropriate because the possible ARC factors may be decoder functions implemented in hardware, and hardware functions tend not to change with any CVS (at least in some entertainment systems, a Group of Pictures of less than 1 second in length). Nevertheless, placing the table in a sequence parameter set is explicitly included among the arrangement options described herein.

[0075] ARC reference information (513) can, advantageously, be placed directly within the picture / slice tile / GOB / tile group header (hereinafter, tile group header) (514) rather than within the picture parameter set, as in JVCET-M0135-v1. The reason is as follows: If the encoder wants to change a single value in the picture parameter set, such as ARC reference information, it needs to create a new PPS and reference that new PPS. Assume that only the ARC reference information changes, while other information, such as quantization matrix information in the PPS, remains unchanged. Such information can be quite large and would need to be retransmitted to complete the new PPS. Since ARC reference information can be a single codeword, such as an index to the table (513), and that is the only value that changes, retransmitting all the information, such as quantization matrix information, would be cumbersome and wasteful. To that extent, avoiding the detour through the PPS, as proposed in JVET-M0135-v1, can be quite good from the standpoint of coding efficiency. Similarly, including ARC reference information in the PPS has a further drawback: because the scope of picture parameter set activation is the picture, the ARC information referenced by the ARC reference information (513) must always be applied to the entire picture, not just to subpictures.

[0076] In the same or a different embodiment, the signaling of ARC parameters may follow the detailed example outlined in Figure 6. Figure 6 shows a syntax diagram of a representation used in video coding standards since at least 1993. The notation of such a syntax diagram broadly follows C-style programming. Lines in bold indicate syntactic elements present in the bitstream, while lines without bold often indicate control flow or variable settings.

[0077] A tile group header (601), as an exemplary syntactic structure for a header applicable to a portion of a picture (possibly a rectangular portion), can conditionally contain a variable-length Exp-Golomb coded syntactic element dec_pic_size_idx (602) (shown in bold). The presence of this syntactic element in the tile group header can be gated with respect to the use of an adaptive resolution (603), a flag value not shown in bold here, meaning that the flag is present in the bitstream at the point where it occurs in the syntactic diagram. Whether or not the adaptive resolution is used in this picture or portion of the picture can be signaled in any high-level syntactic structure inside or outside the bitstream. In the illustrated example, it is signaled by a sequence parameter set as outlined below.

[0078] Referring again to Figure 6, an excerpt of the sequence parameter set (610) is also illustrated. The first syntactic element illustrated is adaptive_pic_resolution_change_flag (611). If true, the flag can indicate the use of adaptive resolution, and the use of adaptive resolution may require some control information. In this example, such control information exists conditionally based on the value of the flag based on the parameter set (612) and the if() statement in the tile group header (601).

[0079] When adaptive resolution is used, in this example, the output resolution coded is in samples (613). The number 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which together can define the resolution of the output picture. Other parts of the video coding technique or standard may define some restrictions on either value. For example, the level definition may limit the total number of output samples, which can be the product of the values ​​of those two syntactic elements. Also, some video coding techniques or standards, or external technologies or standards such as system standards, may restrict number ranges (e.g., one or both dimensions must be divisible by a power of 2) or aspect ratios (e.g., width and height must be in a relationship such as 4:3 or 16:9). Such restrictions may be introduced to facilitate hardware implementation or for other reasons and are well known in the art.

[0080] For certain applications, it may be desirable to instruct the decoder to use a specific reference picture size rather than implicitly assuming that the size is the output picture size. In this example, the syntax element reference_pic_size_present_flag(614) gates the conditional existence of reference picture dimensions(615) (again, the numbers refer to both width and height).

[0081] Finally, a table of possible decoded picture widths and heights is shown. Such a table can be represented, for example, by the table directive (num_dec_pic_size_in_luma_samples_minus1)(616), where "minus1" can refer to the interpretation of the value of its syntactic element. For example, if the encoded value is 0, there is one table entry. If the value is 5, there are six table entries. For each "row" in the table, the decoded picture width and height are contained in the syntax (617).

[0082] The table entries (617) presented can be indexed using the syntax element dec_pic_size_idx (602) in the tile group header, thereby allowing for different decoded sizes, or in other words, zoom ratios, for each tile group.

[0083] The techniques for signaling the adaptive resolution parameters described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 7 shows a computer system 700 suitable for carrying out a particular embodiment of the subject matter of the disclosure.

[0084] Computer software can be coded using any suitable machine language or computer language that can undergo assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed directly or by interpretation, microcode execution, etc., by the computer's central processing unit (CPU), graphics processing unit (GPU), etc.

[0085] Instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.

[0086] The components shown in Figure 7 for the computer system 700 are essentially illustrative and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependencies or requirements relating to any one or combination of components shown in the exemplary embodiments of the computer system 700.

[0087] The computer system 700 may include a human interface input device. Such a human interface input device may respond to input from one or more human users via, for example, tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, applause, etc.), visual input (gestures, etc.), or olfactory input (not shown). The human interface device may be used to capture a medium that is not necessarily directly related to conscious human input, such as audio (voices, music, ambient sounds, etc.), images (scanned images, photographic images acquired from still image cameras, etc.), or video (2D video, 3D video including stereoscopic video, etc.).

[0088] The input human interface device may include one or more of the following (only one of each is illustrated): a keyboard 701, a mouse 702, a trackpad 703, a touchscreen 710, a data glove 704, a joystick 705, a microphone 706, a scanner 707, and a camera 708.

[0089] The computer system 700 may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen 710, data glove 704, or joystick 705, although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speaker 709, headphones (not shown)), visual output devices (e.g., screen 710 including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input functionality, each with or without tactile feedback functionality, some of which may be able to output output beyond three dimensions via means such as two-dimensional visual output or stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0090] The computer system 700 may also include human-accessible storage devices and media associated with storage devices such as CD / DVD ROM / RW 720 with media such as CD / DVD 721, thumb drives 722, removable hard drives or solid-state drives 723, legacy magnetic media such as tapes and floppy disks (not shown), and optical media such as dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0091] Furthermore, those skilled in the art will understand that the term “computer-readable medium” as used in relation to the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0092] The computer system 700 may also include interfaces to one or more communication networks. These networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicle and industrial, real-time, or latency-tolerant. Examples of networks include local area networks such as Ethernet and Wi-Fi; cellular networks including GSM, 3G, 4G, 5G, and LTE; wired or wireless wide-area digital television networks including cable television, satellite television, and terrestrial broadcast television; and vehicle and industrial networks including CANBus. Some networks generally require an external network interface adapter connected to several general-purpose data ports or peripheral buses (749) (e.g., USB ports on the computer system 700), while other networks are generally integrated into the core of the computer system 700 by connections to system buses (e.g., Ethernet interfaces to PC computer systems or cellular network interfaces to smartphone computer systems), as described below. Using any of these networks, the computer system 700 can communicate with other entities. Such communication can be unidirectional, receive only (e.g., television broadcasting), unidirectional transmit only (e.g., from CANbus to several CANbus devices), or bidirectional to other computer systems using local or wide-area digital networks, for example. Several protocols and protocol stacks may be used on each of these networks and network interfaces, as described above.

[0093] The aforementioned human interface devices, human-accessible memory devices, and network interfaces may be attached to the core 740 of the computer system 700.

[0094] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, specialized programmable processing units in the form of field-programmable gate areas (FPGAs) 743, and hardware accelerators 744 for certain tasks. These devices may be connected via a system bus 748, along with internal mass storage units 747 such as read-only memory (ROM) 745, random access memory 746, and internal non-user-accessible hard drives, SSDs, etc. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 748 or via a peripheral bus 749. Peripheral bus architectures include PCI, USB, etc.

[0095] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can execute several instructions that, in combination, can constitute the aforementioned computer code. This computer code can be stored in ROM 745 or RAM 746. Transient data can also be stored in RAM 746, while persistent data can be stored, for example, in the internal mass storage unit 747. High-speed storage and retrieval to any of the memory devices can be enabled by the use of cache memory, which may be closely associated with one or more CPUs 741, GPUs 742, mass storage units 747, ROM 745, RAM 746, etc.

[0096] Computer-readable media may contain computer code for performing various computer implementations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to those skilled in the computer software technology.

[0097] For example, but not limited to, a computer system 700 having an architecture, specifically a core 740, can provide functionality as a result of (one or more) processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, as well as media associated with some storage of the core 740 that is non-transient in nature, such as the core's internal mass storage 747 or ROM 745. Software implementing various embodiments of this disclosure can be stored in such devices and executed by the core 740. The computer-readable media can include one or more memory devices or chips, depending on the specific needs. The software can cause the core 740, specifically the processors in the core (including CPUs, GPUs, FPGAs, etc.), to execute certain processes or specific parts of certain processes as described herein, including defining data structures stored in RAM 746 and modifying such data structures according to processes defined by the software. In addition, or as an alternative, a computer system may also provide functionality resulting from logic (e.g., accelerator 744) wired or otherwise embodied in a circuit, which can operate in place of or in conjunction with software to perform a particular process or a particular part of a particular process described herein. Where referring to software, it may, where appropriate, encompass logic, and vice versa. Where referring to a computer-readable medium, it may, where appropriate, encompass circuitry that houses software for execution (such as an integrated circuit (IC)), circuitry that embodies logic for execution, or both. This disclosure encompasses any appropriate combination of hardware and software.

[0098] Some video encoding techniques or standards, such as VP9, ​​support spatial scalability by implementing resampling of a certain form of reference picture (signaled entirely differently from the subject of disclosure) in conjunction with temporal scalability to enable spatial scalability. In particular, a reference picture may be upsampled to a higher resolution using ARC-style techniques to form the base of a spatial enhancement layer. Those upsampled pictures can then be refined at high resolution using conventional predictive mechanisms to add detail.

[0099] The subject matter of the disclosure can be used in such environments. In some cases, in the same or different embodiments, spatial layers as well as temporal layers can be indicated using the value in the NAL unit header, for example, the time ID field. Doing so offers several advantages in a given system design. For example, an existing selected transfer unit (SFU) created and optimized for selected transfers of temporal layers based on the time ID value in the NAL unit header can be used in a scalable environment without modification. To enable this, there may be requirements for a mapping between coded picture sizes and temporal layers, indicated by the time ID field in the NAL unit header.

[0100] In some video encoding techniques, an access unit (AU) can refer to one or more coded pictures, slices, tiles, or NAL units that are incorporated into and configured within the respective picture / slice / tile / NAL unit bitstream in a given temporal instance. This temporal instance may be composite time.

[0101] In HEVC and some other video encoding techniques, a picture order count (POC) value may be used to indicate a selected reference picture from among multiple reference pictures stored in a decoded picture buffer (DPB). If an access unit (AU) contains one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU may have the same POC value, from which it can be deduced that they were created from content of the same composition time. In other words, in a scenario where two pictures / slices / tiles have the same given POC value, it can indicate two pictures / slices / tiles belonging to the same AU and having the same composition time. Conversely, two pictures / tiles / slices with different POC values ​​can indicate those pictures / slices / tiles belonging to different AUs and having different composition times.

[0102] In one embodiment of the subject matter of the disclosure, the aforementioned strict relationship may be relaxed in that an access unit may include pictures, slices, or tiles having different POC values. By allowing different POC values ​​within an AU, it becomes possible to use the POC values ​​to identify potentially independently decodeable pictures / slices / tiles having the same presentation time. This makes it possible to support multiple scalable layers without changing the reference picture selection signaling (e.g., reference picture set signaling or reference picture list signaling), as will be described in more detail below.

[0103] However, it is still desirable to be able to identify the AU to which a picture / slice / tile belongs, relative to other pictures / slices / tiles with different POC values, solely from the POC value. This can be achieved as described below.

[0104] In the same or other embodiments, the access unit count (AUC) may be signaled by high-level syntactic structures such as NAL unit headers, slice headers, tile group headers, SEI messages, parameter sets, or AU delimiters. The value of the AUC may be used to identify which NAL unit, picture, slice, or tile belongs to a given AU. The value of the AUC may correspond to a separate synthetic time instance. The AUC value may be equal to a multiple of the POC value. The AUC value may be calculated by dividing the POC value by an integer value. In some cases, the division operation may impose a certain burden on the decoder implementation. In such cases, the division operation can be replaced by a shift operation because the number space constraint of the AUC value is small. For example, the AUC value may be equal to the most significant bit (MSB) value of the POC value range.

[0105] In the same embodiment, the POC cycle value for each AU (poc_cycle_au) may be signaled in high-level syntactic structures such as NAL unit headers, slice headers, tile group headers, SEI messages, parameter sets, or AU delimiters. poc_cycle_au may indicate how many different consecutive POC values ​​can be associated with the same AU. For example, if the value of poc_cycle_au is equal to 4, then pictures, slices, or tiles with POC values ​​equal to 0-3, including 0 and 3, are associated with AUs with AUC values ​​equal to 0, and pictures, slices, or tiles with POC values ​​equal to 4-7, including 4 and 7, are associated with AUs with AUC values ​​equal to 1. Thus, the AUC value can be inferred by dividing the POC value by the value of poc_cycle_au.

[0106] In the same or another embodiment, the value of poc_cycle_au may be derived from information located, for example, in the video parameter set (VPS), which identifies the number of spatial or SNR layers in the encoded video sequence. Such possible relationships are briefly described below. While the derivation described above may save a few bits in the VPS and thus improve encoding efficiency, it may be advantageous to explicitly encode poc_cycle_au in a suitable high-level syntactic structure hierarchically lower in the video parameter set so that poc_cycle_au can be minimized for a given small portion of the bitstream, such as a picture. This optimization may save more bits than can be saved by the derivation process described above, because the POC value (and / or the value of syntactic elements that indirectly reference the POC) can be encoded in a low-level syntactic structure.

[0107] Figure 8 shows an example of a video sequence structure with a combination of temporal_id, layer_id, POC, and AUC values, accompanied by adaptive resolution changes. In this example, a picture, slice, or tile in the first AU with AUC=0 may have temporal_id=0 and layer_id=0 or 1, while a picture, slice, or tile in the second AU with AUC=1 may have temporal_id=1 and layer_id=0 or 1, respectively. Regardless of the values ​​of temporal_id and layer_id, the POC value increases by 1 for each picture. In this example, the value of poc_cycle_au can be equal to 2. Preferably, the value of poc_cycle_au can be set to be equal to the number of (spatial scalability) layers. Thus, in this example, the POC value increases by 2 and the AUC value increases by 1.

[0108] In the embodiments described above, all or part of the picture- or layer-based prediction structures and reference picture indications may be supported by using existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC. In RPS or RPL, a selected reference picture is indicated by signaling a POC value or a delta value of the POC between the current picture and the selected reference picture. In the subject matter of the disclosure, RPS and RPL can be used to indicate picture- or layer-based prediction structures without modifying the signaling, but with the following limitations: If the value of the reference picture's temporal_id is greater than the value of the current picture's temporal_id, the current picture may not use that reference picture for motion compensation or other predictions. If the value of the reference picture's layer_id is greater than the value of the current picture's layer_id, the current picture may not use that reference picture for motion compensation or other predictions.

[0109] In this embodiment and other embodiments, motion vector scaling based on POC differences for temporal motion vector prediction may be disabled across multiple pictures within an access unit. Thus, each picture may have a different POC value within the access unit, but the motion vector is not scaled and is not used for temporal motion vector prediction within the access unit. This is because reference pictures with different POCs within the same AU are considered to be reference pictures with the same time instance. Therefore, in this embodiment, the motion vector scaling function may return 1 if the reference picture belongs to the AU associated with the current picture.

[0110] In the same and other embodiments, if the spatial resolution of the reference picture differs from that of the current picture, motion vector scaling based on the POC difference for temporal motion vector prediction may be optionally disabled across multiple pictures. When motion vector scaling is permitted, the motion vectors are scaled based on both the POC difference and the spatial resolution ratio between the current picture and the reference picture.

[0111] In the same or a different embodiment, particularly when poc_cycle_au has non-uniform values ​​(vps_contant_poc_cycle_per_au == 0), the motion vector may be scaled based on the AUC difference rather than the POC difference for temporal motion vector prediction. Otherwise (vps_contant_poc_cycle_per_au == 1), the scaling of the motion vector based on the AUC difference may be identical to the scaling of the motion vector based on the POC difference.

[0112] In the same or another embodiment, if the motion vector is scaled based on the AUC difference, a reference motion vector in the same AU having the current picture (having the same AUC value) is not scaled based on the AUC difference and is used for motion vector prediction with or without scaling based on the spatial resolution ratio between the current picture and the reference picture.

[0113] In the same and other embodiments, the AUC value is used to identify the boundaries of the AU and for virtual reference decoder (HRD) operation that requires input and output timings with AU granularity. In most cases, the decoded picture of the top layer within the AU may be output for display. The AUC value and layer_id value can be used to identify the output picture.

[0114] In one embodiment, a picture may consist of one or more subpictures. Each subpicture may cover a local or entire area of ​​the picture. The area supported by a subpicture may or may not overlap with the area supported by another subpicture. The area composed of one or more subpictures may or may not cover the entire area of ​​the picture. If the picture consists of subpictures, the area supported by a subpicture is the same as the area supported by the picture.

[0115] In the same embodiment, a subpicture may be encoded by an encoding method similar to the encoding method used for the coded picture. A subpicture may be encoded independently or may be encoded in dependence of another subpicture or coded picture. A subpicture may or may not have parsing dependencies from another subpicture or coded picture.

[0116] In the same embodiment, coded subpictures may be contained in one or more layers. Coded subpictures within a layer may have different spatial resolutions. The original subpicture may be spatially resampled (upsampled or downsampled), encoded with different spatial resolution parameters, and contained in the bitstream corresponding to the layer.

[0117] In the same or different embodiments, a subpicture of (W,H) may be encoded and contained in a coded bitstream corresponding to layer 0, where W represents the width of the subpicture and H represents the height of the subpicture, respectively (W*S w,k ,H*S h,k A subpicture that has been upsampled (or downsampled) from a subpicture with the original spatial resolution may be encoded and contained in the coded bitstream corresponding to layer k, S w,k S h,krepresents the resampling ratio between the horizontal and vertical directions. S w,k , S h,k If the value of S is greater than 1, the resampling is equal to upsampling. On the other hand, S w,k , S h,k If the value of S is less than 1, the resampling is equal to downsampling.

[0118] In the same or another embodiment, the coded sub - pictures within a layer may have a different visual quality from the coded sub - pictures within another layer, whether the same or different sub - pictures. For example, sub - picture i within layer n is coded with quantization parameter Q i,n and sub - picture j within layer m is coded with quantization parameter Q j,m .

[0119] In the same or another embodiment, the coded sub - pictures within a layer may be independently decodable without any parsing - dependence or decoding - dependence from the coded sub - pictures within another layer of the same local region. A sub - picture layer that can be independently decoded without referring to another sub - picture layer of the same local region is an independent sub - picture layer. The coded sub - pictures within an independent sub - picture layer may or may not have decoding - dependence or parsing - dependence from previously coded sub - pictures within the same sub - picture layer, but the coded sub - pictures cannot have any dependence from the coded pictures within another sub - picture layer.

[0120] In the same or a different embodiment, coded subpictures within a layer may be dependently decodeable, with parsing or decoding dependencies from coded subpictures in another layer in the same local region. A subpicture layer that may be dependently decodeable by referencing another subpicture layer in the same local region is a dependent subpicture layer. Coded subpictures within a dependent subpicture may reference coded subpictures belonging to the same subpicture, previously coded subpictures within the same subpicture layer, or both reference subpictures.

[0121] In the same or a different embodiment, a coded subpicture consists of one or more independent subpicture layers and one or more dependent subpicture layers. However, there may be at least one independent subpicture layer for a coded subpicture. An independent subpicture layer may have a layer identifier (layer_id) value equal to 0, which may be present in the NAL unit header or another high-level syntactic structure. A subpicture layer with a layer_id equal to 0 is the base subpicture layer.

[0122] In the same or another embodiment, a picture may consist of one or more foreground subpictures and one background subpicture. The region supported by the background subpicture may be equal to the region of the picture. The region supported by the foreground subpicture may overlap with the region supported by the background subpicture. The background subpicture may be a base subpicture layer, and the foreground subpicture may be a non-base (enhancement) subpicture layer. One or more non-base subpicture layers may refer to the same base layer for decoding. Each non-base subpicture layer having a layer_id equal to a may refer to a non-base subpicture layer having a layer_id equal to b, where a is greater than b.

[0123] In the same or another embodiment, a picture may consist of one or more foreground subpictures with or without a background subpicture. Each subpicture may have its own base subpicture layer and one or more non-base (enhancement) layers. Each base subpicture layer may be referenced by one or more non-base subpicture layers. Each non-base subpicture layer having a layer_id equal to a may reference a non-base subpicture layer having a layer_id equal to b, where a is greater than b.

[0124] In the same or another embodiment, a picture may consist of one or more foreground subpictures with or without a background subpicture. Each coded subpicture in a (base or non-base) subpicture layer may be referenced by one or more non-base layer subpictures belonging to the same subpicture and one or more non-base layer subpictures not belonging to the same subpicture.

[0125] In the same or another embodiment, a picture may consist of one or more foreground subpictures with or without background subpictures. A subpicture in layer a may be further divided into multiple subpictures within the same layer. One or more coded subpictures in layer b may refer to divided subpictures in layer a.

[0126] In the same or different embodiments, an encoded video sequence (CVS) may be a group of encoded pictures. The CVS may consist of one or more encoded subpicture sequences (CSPS), which may be a group of encoded subpictures covering the same local region of a picture. The CSPS may have the same or different temporal resolution as the encoded video sequence.

[0127] In the same or a different embodiment, the CSPS may be encoded and contained in one or more layers. The CSPS may consist of one or more CSPS layers. By decoding one or more CSPS layers corresponding to the CSPS, a sequence of subpictures corresponding to the same local region can be reconstructed.

[0128] In the same or a different embodiment, the number of CSPS layers corresponding to one CSPS may be the same as, or different from, the number of CSPS layers corresponding to another CSPS.

[0129] In the same or a different embodiment, a CSPS layer may have a different temporal resolution (e.g., frame rate) than another CSPS layer. The original (uncompressed) subpicture sequence may be resampled in time (upsampled or downsampled), encoded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.

[0130] In the same or another embodiment, a subpicture sequence having a frame rate F may be encoded and contained in a coded bitstream corresponding to layer 0, F*S t,k Then, the sub-picture sequence that has been temporally upsampled (or downsampled) from the original sub-picture sequence may be encoded and contained in the coded bitstream corresponding to layer k, S t,k This indicates the time sampling ratio of layer k. t,k If the value of is greater than 1, the time resampling process is equivalent to frame rate upconversion. On the other hand, S t,k If the value is less than 1, the time resampling process is equivalent to a downconversion of the frame rate.

[0131] In the same or another embodiment, when a subpicture having CSPS layer a is referenced by a subpicture having CSPS layer b for motion compensation or arbitrary inter-layer prediction, if the spatial resolution of CSPS layer a differs from that of CSPS layer b, the decoded pixels in CSPS layer a are resampled and used for reference. The resampling process may require upsampling filtering or downsampling filtering.

[0132] Figure 9 shows an example syntax table for signaling the `vps_poc_cycle_au` syntax element in the VPS (or SPS), which indicates the `poc_cycle_au` used for all pictures / slices in the encoded video sequence, and the `slice_poc_cycle_au` syntax element, which indicates the `poc_cycle_au` of the current slice in the slice header. If the POC value increases uniformly per AU, `vps_contant_poc_cycle_per_au` in the VPS is set to equal 1, and `vps_poc_cycle_au` is signaled in the VPS. In this case, `slice_poc_cycle_au` is not explicitly signaled, and the AUC value per AU is calculated by dividing the POC value by `vps_poc_cycle_au`. If the POC value does not increase uniformly per AU, `vps_contant_poc_cycle_per_au` in the VPS is set to equal 0. In this case, vps_access_unit_cnt is not signaled, but slice_access_unit_cnt is signaled in the slice header for each slice or picture. Each slice or picture may have a different slice_access_unit_cnt value. The AUC value per AU is calculated by dividing the POC value by slice_poc_cycle_au. Figure 10 shows a block diagram illustrating the relevant workflow.

[0133] Even if the POC values ​​of the pictures, slices, or tiles may differ in the same or other embodiments, pictures, slices, or tiles corresponding to an AU with the same AUC value can be associated with the same decoding or output time instance. Thus, all or part of the pictures, slices, or tiles associated with the same AU can be decoded in parallel and output in the same time instance without inter-parse / decoding dependencies across pictures, slices, or tiles within the same AU.

[0134] Even if the POC values ​​of a picture, slice, or tile may differ in the same or other embodiments, a picture, slice, or tile corresponding to an AU with the same AUC value can be associated with the same composite / display time instance. If the composite time is included in the container format, even if the pictures correspond to different AUs, if the pictures have the same composite time, those pictures can be displayed in the same time instance.

[0135] In the same or other embodiments, each picture, slice, or tile may have the same temporal identifier (temporal_id) within the same AU. All or some of the pictures, slices, or tiles corresponding to a time instance may be associated with the same time sublayer. In the same or other embodiments, each picture, slice, or tile may have the same or different spatial layer ID (layer_id) within the same AU. All or some of the pictures, slices, or tiles corresponding to a time instance may be associated with the same or different spatial layer.

[0136] Figure 11 shows an exemplary video stream including a background video CSPS having a layer_id equal to 0 and a plurality of foreground CSPS layers. The coded subpictures can consist of one or more CSPS layers, but the background regions that do not belong to any foreground CSPS layer can consist of the base layer. The base layer can include background regions and foreground regions, while the enhancement CSPS layers include foreground regions. The enhancement CSPS layers can have better visual quality than the base layer in the same region. The enhancement CSPS layers can refer to the reconstructed pixels and the motion vectors of the base layer corresponding to the same region.

[0137] In the same or another embodiment, the video bitstream corresponding to the base layer is included in a track, and the CSPS layers corresponding to each subpicture are included in separate tracks within the video file.

[0138] In the same or another embodiment, the video bitstream corresponding to the base layer is included in a track, and the CSPS layers having the same layer_id are included in separate tracks. In this example, the track corresponding to layer k includes only the CSPS layer corresponding to layer k.

[0139] In the same or another embodiment, each CSPS layer of each subpicture is stored in a separate track. Each track (trach) may or may not have a parsing dependency or decoding dependency from one or more other tracks.

[0140] In the same or another embodiment, each track can include the bitstream corresponding to the CSPS layers from layer i to layer j of all or part of the subpicture, where 0 < i ≤ j ≤ k, and k is the top layer of the CSPS.

[0141] In the same or another embodiment, the picture consists of one or more associated media data, including a depth map, an alpha map, 3D shape data, an occupancy map, and the like. Such associated time-limited media data can be divided into one or more data substreams, each corresponding to a single subpicture.

[0142] In the same or a different embodiment, Figure 12 shows an example of video conferencing based on a multi-layer subpicture method. The video stream includes one base layer video bitstream corresponding to the background picture and one or more enhancement layer video bitstreams corresponding to the foreground subpictures. Each enhancement layer video bitstream corresponds to a CSPS layer. On the display, the picture corresponding to the base layer is displayed by default. This includes picture-in-picture (PIP) for one or more users. When a user is selected by client control, the enhancement CSPS layer corresponding to the selected user is decoded and displayed with enhanced quality or spatial resolution. Figure 13 shows a diagram for operation.

[0143] In the same or a different embodiment, a network middlebox (such as a router) may select which layers to send to a user depending on its bandwidth. Picture / subpicture organization may be used for bandwidth adaptation. For example, if a user does not have the bandwidth, the router may dynamically remove layers or select some subpictures based on their importance or the setup used, and adopt this to accommodate the bandwidth.

[0144] Figure 14 illustrates a use case for 360-degree video. When a spherical 360-degree picture is projected onto a planar picture, the projected 360-degree picture may be divided into multiple subpictures as a base layer. The enhancement layer of a specific subpicture may be encoded and sent to the client. The decoder can decode both the base layer containing all subpictures and the enhancement layer of the selected subpicture. If the current viewport is identical to the selected subpicture, the displayed picture may have higher quality with the decoded subpicture that has the enhancement layer. Otherwise, the decoded picture with the base layer may be displayed at lower quality.

[0145] In the same or a different embodiment, any layout information for display may be present in the file as supplementary information (such as SEI messages or metadata). One or more decoded subpictures may be rearranged and displayed according to the signaled layout information. The layout information may be signaled by a streaming server or broadcaster, regenerated by a network entity or cloud server, or determined by user-customized settings.

[0146] In one embodiment, if an input picture is divided into one or more (rectangular) subregions, each subregion may be encoded as an independent layer. Each independent layer corresponding to a local region may have a unique layer_id value. For each independent layer, subpicture size and position information may be signaled. For example, picture size (width, height) and upper-left corner offset information (x_offset, y_offset). Figure 15 shows an example of the layout of the divided subpictures, their subpicture size and position information, and their corresponding picture prediction structure. Layout information, including (one or more) subpicture sizes and (one or more) subpicture positions, may be signaled by (one or more) parameter sets, slice or tile group headers, or high-level syntactic structures such as SEI messages.

[0147] In the same embodiment, each subpicture corresponding to an independent layer may have its own POC value within the AU. If a reference picture among the pictures stored in the DPB is indicated using one or more syntactic elements in the RPS or RPL structure, the one or more POC values ​​of each subpicture corresponding to the layer may be used.

[0148] In the same or a different embodiment, layer_id may not be used to indicate the (inter-layer) prediction structure, and POC(delta) values ​​may be used instead.

[0149] In the same embodiment, a subpicture having a POC value equal to N corresponding to a layer (or local region) may or may not be used as a reference picture for a subpicture having a POC value equal to N+K corresponding to the same layer (or the same local region) for motion compensation prediction. In most cases, the value of K may be equal to the maximum number of (independent) layers, which may be the same as the number of subregions.

[0150] In the same or a different embodiment, Figure 16 shows an extended case of Figure 15. When the input picture is divided into multiple (e.g., four) sub-regions, each local region may be encoded with one or more layers. In this case, the number of independent layers may be equal to the number of sub-regions, and one or more layers may correspond to a sub-region. Thus, each sub-region may be encoded using one or more independent layers and zero or more dependent layers.

[0151] In the same embodiment, in Figure 16, the input picture may be divided into four sub-regions. The upper right sub-region may be encoded as two layers, Layer 1 and Layer 4, and the lower right sub-region may be encoded as two layers, Layer 3 and Layer 5. In this case, Layer 4 may refer to Layer 1 for motion compensation prediction, and Layer 5 may refer to Layer 3 for motion compensation.

[0152] In the same or another embodiment, in-loop filtering across layer boundaries (such as deblocking filtering, adaptive in-loop filtering, reshapers, bilateral filters, or any deep learning-based filtering) may be (optionally) disabled.

[0153] In the same or another embodiment, motion compensation prediction or intrablock copying across layer boundaries may be (optionally) disabled.

[0154] In the same or a different embodiment, boundary padding for motion compensation prediction or in-loop filtering at the boundaries of subpictures may be optionally processed. A flag indicating whether boundary padding is processed may be signaled in a parameter set (VPS, SPS, PPS, or APS), a slice or tile group header, or a high-level syntactic structure such as an SEI message.

[0155] In the same or a different embodiment, layout information for (one or more) subregions (or (one or more) subpictures) may be signaled in the VPS or SPS. Figure 17 shows examples of VPS and SPS syntactic elements. In this example, vps_sub_picture_dividing_flag is signaled in the VPS. This flag may indicate whether (one or more) input pictures are divided into multiple subregions. When the value of vps_sub_picture_dividing_flag is equal to 0, (one or more) input pictures in (one or more) encoded video sequences corresponding to the current VPS may not be divided into multiple subregions. In this case, the input picture size may be equal to the coded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS. When the value of vps_sub_picture_dividing_flag is equal to 1, (one or more) input pictures may be divided into multiple subregions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled in the VPS. The values ​​of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples can be equal to the width and height of (one or more) input pictures, respectively.

[0156] In the same embodiment, the values ​​of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding but for synthesis and display.

[0157] In the same embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntactic elements pic_offset_x and pic_offset_y may (a) be signaled in the SPS corresponding to (one or more) specific layers. In this case, the coded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS may be equal to the width and height of the sub-region corresponding to the specific layer. The position of the upper-left corner of the sub-region (pic_offset_x, pic_offset_y) may also be signaled in the SPS.

[0158] In the same embodiment, the position information of the upper left corner of the subregion (pic_offset_x, pic_offset_y) may not be used for decoding but may be used for composition and display.

[0159] In the same or a different embodiment, layout information (size and position) of all or part of (one or more) subregions of an input picture, and dependency information between (one or more) layers, may be signaled by a parameter set or SEI message. Figure 18 shows an example of syntax elements that indicate information about the layout of subregions, dependencies between layers, and relationships between subregions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) subregions in the current encoded video sequence. The syntax element num_layers indicates the number of layers in the current encoded video sequence. The value of num_layers can be greater than or equal to the value of num_sub_region. If any subregion is encoded as a single layer, the value of num_layers may be equal to the value of num_sub_region. If one or more subregions are encoded as multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dependency_flag[i][j] indicates the dependency from the j-th layer to the i-th layer. num_layers_for_region[i] indicates the number of layers associated with the i-th subregion. sub_region_layer_id[i][j] indicates the layer_id of the j-th layer associated with the i-th subregion. The sub_region_offset_x[i] and sub_region_offset_y[i] indicate the horizontal and vertical positions of the top-left corner of the i-th subregion, respectively. sub_region_width[i] and sub_region_height[i] indicate the width and height of the i-th subregion, respectively.

[0160] In one embodiment, one or more syntactic elements that specify an output layer set to indicate one of several layers to be output with or without profile tier level information may be signaled in a high-level syntactic structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message. Referring to Figure 19, the syntactic element num_output_layer_sets, which indicates the number of output layer sets (OLS) in an encoded video sequence referencing a VPS, may be signaled in a VPS. For each output layer set, output_layer_flag may be signaled as many times as there are output layers.

[0161] In the same embodiment, output_layer_flag[i] equal to 1 specifies that the i-th layer is output. vps_output_layer_flag[i] equal to 0 specifies that the i-th layer is not output.

[0162] In the same or another embodiment, one or more syntactic elements specifying profile tier level information for each output layer set may be signaled in a high-level syntactic structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message. Referring further to Figure 19, the syntactic element num_profile_tile_level, which indicates the number of profile tier level entries per OLS in an encoded video sequence referencing a VPS, may be signaled in a VPS. For each output layer set, a set of syntactic elements for profile tier level information, or an index indicating a specific profile tier level information entry within the profile tier level information, may be signaled as many times as there are output layers.

[0163] In the same embodiment, profile_tier_level_idx[i][j] specifies an index to a list of profile_tier_level() syntax structures in the VPS for the profile_tier_level() syntax structure applied to the j-th layer of the i-th OLS.

[0164] In the same or another embodiment, referring to Figure 20, the syntax elements num_profile_tile_level and / or num_output_layer_sets may be signaled when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0).

[0165] In the same or another embodiment, referring to Figure 20, a syntactic element vps_output_layers_mode[i] may exist in the VPS indicating the mode of output layer signaling for the i-th output layer set.

[0166] In the same embodiment, vps_output_layers_mode[i] equal to 0 specifies that only the top layer in the i-th output layer set is output. vps_output_layer_mode[i] equal to 1 specifies that all layers in the i-th output layer set are output. vps_output_layer_mode[i] equal to 2 specifies that the layers to be output are those in the i-th output layer set where vps_output_layer_flag[i][j] is equal to 1. More values ​​may be reserved.

[0167] In the same embodiment, output_layer_flag[i][j] may or may not be signaled depending on the value of vps_output_layers_mode[i] of the i-th output layer set.

[0168] In the same or another embodiment, referring to Figure 20, the flag vps_ptl_signal_flag[i] may exist for the i-th output layer set. Depending on the value of vps_ptl_signal_flag[i], the profile tier level information for the i-th output layer set may or may not be signaled.

[0169] In the same or another embodiment, referring to Figure 21, the number of subpictures in the current CVS, max_subpics_minus1, can be signaled in a high-level syntactic structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message.

[0170] In the same embodiment, referring to Figure 21, if the number of subpictures is greater than 1 (max_subpics_minus1>0), the subpicture identifier sub_pic_id[i] of the i-th subpicture may be signaled.

[0171] In the same or a different embodiment, one or more syntactic elements indicating subpicture identifiers belonging to each layer of each output layer set may be signaled in the VPS. Referring to Figure 22, sub_pic_id_layer[i][j][k] indicates the k-th subpicture that resides in the j-th layer of the i-th output layer set. Using this information, the decoder may know which subpictures for each layer of a particular output layer set can be decoded and output.

[0172] In one embodiment, a picture header (PH) is a syntactic structure containing syntactic elements that apply to all slices of a coded picture. A picture unit (PU) is a set of NAL units that are related to one another according to a specified classification rule, are consecutive in decoding order, and contain exactly one coded picture. A PU may include a picture header (PH) and one or more VCL NAL units that constitute a coded picture.

[0173] In one embodiment, the SPS(RBSP) may be available to the decoding process before being referenced and may be contained in at least one AU having a TemporalId equal to 0, or may be provided via external means.

[0174] In one embodiment, the SPS(RBSP) may be available to the decoding process before being referenced and may be contained in at least one AU having a TemporalId equal to 0 within the CVS, which includes one or more PPSs that reference the SPS, or may be provided via external means.

[0175] In one embodiment, the SPS(RBSP) may be available to the decryption process before being referenced by one or more PPSs, and may be contained in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit referencing the SPS NAL unit in the CVS, which includes one or more PPS referencing the SPS, or may be provided via external means.

[0176] In one embodiment, the SPS(RBSP) may be available to the decoding process before being referenced by one or more PPSs and may be contained in at least one PU having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit referencing the SPS NAL unit, or may be provided via external means.

[0177] In one embodiment, the SPS(RBSP) may be available to the decoding process before being referenced by one or more PPSs, and may be contained in or provided via external means, having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of a PPS NAL unit that references an SPS NAL unit in a CVS, including one or more PPSs that reference the SPS.

[0178] In the same or another embodiment, pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the referenced SPS. The value of pps_seq_parameter_set_id may be the same for all PPS referenced by the coded picture in the CLVS.

[0179] In the same or a different embodiment, all SPS NAL units having a specific value for sps_seq_parameter_set_id in CVS may have the same content.

[0180] In the same or another embodiment, regardless of the nuh_layer_id value, SPS NAL units may share the same value space for sps_seq_parameter_set_id.

[0181] In the same or another embodiment, the nuh_layer_id value of an SPS NAL unit may be equal to the lowest nuh_layer_id value of a PPS NAL unit referencing the SPS NAL unit.

[0182] In one embodiment, if an SPS having a nuh_layer_id equal to m is referenced by one or more PPSs having a nuh_layer_id equal to n, then the layer having a nuh_layer_id equal to m may be the same as the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to n or the layer having a nuh_layer_id equal to m.

[0183] In one embodiment, the PPS(RBSP) is available for the decoding process before being referenced and is contained in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit, or is provided via external means.

[0184] In one embodiment, the PPS(RBSP) may be available for the decoding process before being referenced and may be contained in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS, which includes one or more PHs (or coded slice NAL units) that reference the PPS, or may be provided via external means.

[0185] In one embodiment, the PPS (RBSP) may be available to the decoding process before being referenced by one or more PHs (or coded slice NAL units), and may be contained in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the coded slice NAL units referencing the PPS NAL units in the CVS, which includes one or more PHs (or coded slice NAL units) referencing the PPS, or may be provided via external means.

[0186] In one embodiment, the PPS (RBSP) may be available to the decoding process before being referenced by one or more PHs (or coded slice NAL units), and may be contained in or provided via external means at least one PU having a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the lowest nuh_layer_id value of the coded slice NAL units referencing the PPS NAL units in the CVS, which includes one or more PHs (or coded slice NAL units) referencing the PPS.

[0187] In the same or another embodiment, ph_pic_parameter_set_id in PH specifies the value of pps_pic_parameter_set_id of the PPS being referenced and in use. The value of pps_seq_parameter_set_id may be the same for all PPS referenced by the coded picture in CLVS.

[0188] In the same or another embodiment, all PPS NAL units having a specific value for pps_pic_parameter_set_id within the PU shall have the same content.

[0189] In the same or a different embodiment, regardless of the nuh_layer_id value, PPS NAL units may share the same value space for pps_pic_parameter_set_id.

[0190] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit may be equal to the lowest nuh_layer_id value of an encoded slice NAL unit that references a NAL unit that references the PPS NAL unit.

[0191] In one embodiment, if a PPS having a nuh_layer_id equal to m is referenced by one or more coded slice NAL units having a nuh_layer_id equal to n, then the layer having a nuh_layer_id equal to m may be the same as the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to n or the layer having a nuh_layer_id equal to m.

[0192] In one embodiment, the PPS(RBSP) is available for the decoding process before being referenced and is contained in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit, or is provided via external means.

[0193] In one embodiment, the PPS(RBSP) may be available for the decoding process before being referenced and may be contained in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS, which includes one or more PHs (or coded slice NAL units) that reference the PPS, or may be provided via external means.

[0194] In one embodiment, the PPS (RBSP) may be available to the decoding process before being referenced by one or more PHs (or coded slice NAL units), and may be contained in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the coded slice NAL units referencing the PPS NAL units in the CVS, which includes one or more PHs (or coded slice NAL units) referencing the PPS, or may be provided via external means.

[0195] In one embodiment, the PPS (RBSP) may be available to the decoding process before being referenced by one or more PHs (or coded slice NAL units), and may be contained in or provided via external means at least one PU having a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the lowest nuh_layer_id value of the coded slice NAL units referencing the PPS NAL units in the CVS, which includes one or more PHs (or coded slice NAL units) referencing the PPS.

[0196] In the same or another embodiment, ph_pic_parameter_set_id in PH specifies the value of pps_pic_parameter_set_id of the PPS being referenced and in use. The value of pps_seq_parameter_set_id may be the same for all PPS referenced by the coded picture in CLVS.

[0197] In the same or another embodiment, all PPS NAL units having a specific value for pps_pic_parameter_set_id within the PU shall have the same content.

[0198] In the same or a different embodiment, regardless of the nuh_layer_id value, PPS NAL units may share the same value space for pps_pic_parameter_set_id.

[0199] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit may be equal to the lowest nuh_layer_id value of an encoded slice NAL unit that references a NAL unit that references the PPS NAL unit.

[0200] In one embodiment, if a PPS having a nuh_layer_id equal to m is referenced by one or more coded slice NAL units having a nuh_layer_id equal to n, then the layer having a nuh_layer_id equal to m may be the same as the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to n or the layer having a nuh_layer_id equal to m.

[0201] In one embodiment, if the flag no_temporal_sublayer_switching_flag is signaled in the DPS, VPS, or SPS, the TemporalId value of the PPS referring to a parameter set containing the flag equal to 1 may be equal to 0, while the TemporalId value of the PPS referring to a parameter set containing the flag equal to 1 may be greater than or equal to the TemporalId value of the parameter set.

[0202] In one embodiment, each PPS(RBSP) may be available to the decoding process before being referenced and may be contained in or provided via external means at least one AU having a TemporalId less than or equal to the TemporalId of the coding slice NAL unit (or PH NAL unit) referencing each PPS(RBSP). If a PPS NAL unit is contained in an AU prior to the AU containing the coding slice NAL unit referencing the PPS, there may not be any VCL NAL units or VCL NAL units with a nal_unit_type equal to STSA_NUT that enable temporal upper-layer switching, indicating that the picture in the VCL NAL unit may be a Stepwise Temporal Sublayer Access (STSA) picture, after the PPS NAL unit and before the coding slice NAL unit referencing the APS.

[0203] In the same or a different embodiment, a PPS NAL unit and an encoded slice NAL unit (and its PH NAL unit) that references a PPS may be contained in the same AU.

[0204] In the same or different embodiments, the PPS NAL unit and the STSA NAL unit may be contained in the same AU preceding the coded slice NAL unit (and its PH NAL unit) that references the PPS.

[0205] In the same or a different embodiment, the STSA NAL unit, the PPS NAL unit, and the coded slice NAL unit (and its PH NAL unit) that references the PPS may reside within the same AU.

[0206] In the same embodiment, the TemporalId value of the VCL NAL unit containing PPS may be equal to the TemporalId value of the previous STSA NAL unit.

[0207] In the same embodiment, the picture sequence count (POC) value of the PPS NAL unit may be greater than or equal to the POC value of the STSA NAL unit.

[0208] In the same embodiment, the picture sequence count (POC) value of an encoded slice or PH NAL unit that references a PPS NAL unit may be greater than or equal to the POC value of the referenced PPS NAL unit.

[0209] In one embodiment, since all VCL NAL units in the AU have the same TemporalId value, the value of sps_max_sublayers_minus1 is the same across all layers in the encoded video sequence. The value of sps_max_sublayers_minus1 is the same in all SPS referenced by the coded picture in the CVS.

[0210] In one embodiment, the chroma_format_idc value of an SPS referenced by one or more coded pictures in Layer A is equal to the chroma_format_idc value in an SPS referenced by one or more coded pictures in Layer B, if Layer A is a direct reference layer of Layer B. This is because any coded picture has the same chroma_format_idc value as its reference picture. The chroma_format_idc value of an SPS referenced by one or more coded pictures in Layer A is equal to the chroma_format_idc value in an SPS referenced by one or more coded pictures in a direct reference layer of Layer A in the CVS.

[0211] In one embodiment, the subpics_present_flag and sps_subpic_id_present_flag values ​​of an SPS referenced by one or more coded pictures in Layer A are equal to the subpics_present_flag and sps_subpic_id_present_flag values ​​in an SPS referenced by one or more coded pictures in Layer B, if Layer A is a direct reference layer of Layer B. This is because the layout of the subpictures needs to be aligned or associated across layers. Otherwise, subpictures with multiple layers may not be extracted correctly. The subpics_present_flag and sps_subpic_id_present_flag values ​​of an SPS referenced by one or more coded pictures in Layer A are equal to the subpics_present_flag and sps_subpic_id_present_flag values ​​in an SPS referenced by one or more coded pictures in a direct reference layer of Layer A, within a CVS.

[0212] In one embodiment, if an STSA picture in layer A is referenced by a picture in a direct reference layer of layer A within the same AU, the picture referencing the STSA is assumed to be an STSA picture. Otherwise, the up-switching of time sublayers cannot be synchronized between layers. If an STSA NAL unit in layer A is referenced by a VCL NAL unit in a direct reference layer of layer A within the same AU, the nal_unit_type value of the VCL NAL unit referencing the STSA NAL unit is assumed to be equal to STSA_NUT.

[0213] In one embodiment, if a RASL picture in layer A is referenced by a picture in a direct reference layer of layer A within the same AU, the picture referencing RASL is assumed to be a RASL picture. Otherwise, the picture cannot be decoded correctly. If a RASL NAL unit in layer A is referenced by a VCL NAL unit in a direct reference layer of layer A within the same AU, the nal_unit_type value of the VCL NAL unit referencing the RASL NAL unit is assumed to be equal to RASL_NUT.

[0214] While this disclosure describes several exemplary embodiments, there are many modifications, substitutions, and alternative equivalents that fall within the scope of this disclosure. Those skilled in the art will therefore understand that numerous systems and methods, not expressly illustrated or described herein, can be devised to embody the principles of this disclosure and thus fall within its spirit and scope. [Explanation of Symbols]

[0215] 100 Communication Systems 110 First terminal 120 Second terminal 130 devices 140 devices 150 Networks 201 Video Sources 202 Sample Streams 203 Encoder 204 video bitstream 205 Streaming Servers 206 Streaming Clients 207 video bitstreams 208 Streaming Clients 209 video bitstreams 210 Video Decoders 211 Video Sample Streams 212 displays, rendering devices 213 Capture Subsystem 310 Receiver 312 channels 315 buffer memory 320 Entropy Decoder / Parser 321 Symbols 351 Scaler / Inverse Unit 352 IntraPicture Prediction Units 353 Motion Compensation Prediction Unit 355 Aggregator 356 Loop Filter Units, Current Reference Picture 357 Reference picture memory (buffer) 430 Source Coder 432 Coding Engine 433 Local Decoder 434 Reference Picture Memory 435 Predictor 440 Transmitters 443 Encoded Video Sequence 445 Entropy Coder 450 Controllers 460 communication channels 501 Picture Header 502 ARC information 503 H.263 PLUSPTYPE Header Extension 504 Picture Parameter Set 505 ARC Reference Information 506 Table 507 Sequence Parameter Set 508 Tile Group Header 509 ARC information 511 Parameter Set 512 ARC information 513 ARC Reference Information 514 Tile Group Header 515 ARC information 516 Parameter Set, ARC Information Table 601 Tile Group Header 602 Syntax element dec_pic_size_idx 603 Adaptive Resolution 610 Sequence Parameter Set 611 adaptive_pic_resolution_change_flag 612 Parameter Set Output resolution in units of 613 samples 614 Syntax element reference_pic_size_present_flag 615 Reference Picture Dimensions 616 Table instruction (num_dec_pic_size_in_luma_samples_minus1) 617 Syntax, Table Entries 700 Computer Systems 701 Keyboard 702 Mouse 703 Trackpad 704 Data Globe 705 Joystick 706 Microphone 707 Scanner 708 Camera 709 Speaker 710 Touchscreen, Screen 720 CD / DVD ROM / RW 721 Medium 722 Thumb Drive 723 Removable hard drive or solid state drive 740 cores 741 Central Processing Unit (CPU) 742 Graphics Processing Unit (GPU) 743 Field-Programmable Gate Area (FPGA) 744 Hardware Accelerators 745 Read-only memory (ROM) 746 random access memory 747 Core internal large-capacity storage unit 748 System Bus 749 Local Buses

Claims

1. An encoding method for encoding video data having one or more layers using a processor, A step of encoding a first syntactic element that defines multiple modes specifying one or more output layers from among the one or more layers associated with each of one or more sets of output layers, Includes, The first syntactic element is, A first mode in which only the topmost layer among the one or more layers associated with the specified output layer set is designated as the output layer. A second mode in which all of the one or more layers associated with the specified output layer set are designated as output layers, and A third mode in which a second syntactic element is signaled, associated with each of the one or more layers associated with the specified output layer set. It stipulates, The second syntactic element, when 1, designates the associated layer as the output layer, and when 0, does not designate the associated layer as the output layer. Encoding method.

2. The step of encoding the second syntactic element for each of the one or more output layer sets and for each of the one or more associated layers, when the first syntactic element is in the third mode. The encoding method according to claim 1, further comprising:

3. A step of encoding a third syntactic element that defines a subpicture identifier corresponding to each layer associated with each of the one or more output layer sets. The encoding method according to claim 1, further comprising:

4. An encoding method for encoding video data having one or more layers using a processor, The step of generating and transmitting a bitstream encoded from the aforementioned video data, The step of generating and transmitting the bitstream is: A step of encoding a first syntactic element that defines multiple modes specifying one or more output layers from among the one or more layers associated with each of one or more sets of output layers, Includes, The first syntactic element is, A first mode in which only the topmost layer among the one or more layers associated with the specified output layer set is designated as the output layer. A second mode in which all of the one or more layers associated with the specified output layer set are designated as output layers, and A third mode in which a second syntactic element is signaled, associated with each of the one or more layers associated with the specified output layer set. It stipulates, The second syntactic element, when 1, designates the associated layer as the output layer, and when 0, does not designate the associated layer as the output layer. Encoding method.

5. An encoding device configured to perform the encoding method described in any one of Claims 1 to 4.

6. A program for causing a computer to perform the encoding method described in any one of claims 1 to 4.

7. A method for decoding video data having one or more layers by a processor, Receive or retrieve a first syntax element that defines multiple modes specifying one or more output layers from among the one or more layers associated with each of one or more sets of output layers, If the first syntactic element is in the first mode, then only the topmost layer among the one or more layers associated with the specified output layer set is decoded as an output layer. If the first syntactic element is in the second mode, all of the one or more layers associated with the specified output layer set are decoded as output layers. If the first syntactic element is in third mode, the second syntactic element associated with each of the one or more layers associated with the specified output layer set is further received or obtained; if the second syntactic element is 1, the associated layer is decoded as the output layer; if the second syntactic element is 0, the associated layer is not decoded as the output layer. Decryption method.

8. A decoding device configured to perform the decoding method described in Claim 7.

9. A program for causing a computer to execute the decryption method described in Claim 7.