Technique for Extracting Bitstream of Sub-Image in Symbolic Video Stream
By extracting and applying resampling and spatial scalability parameters within the video encoding process, the method addresses the challenge of adaptive picture size and resolution changes, enhancing encoding efficiency and reducing bandwidth and storage needs.
Patent Information
- Application Number
- JP2023189274
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-01
- Filing Date
- 2023-11-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-06-07
AI Technical Summary
Existing video encoding technologies face challenges in efficiently encoding and decoding video data while allowing for adaptive changes in picture size and resolution, which affects bandwidth and storage requirements.
The method involves receiving video data with sub-images, extracting resampling and spatial scalability parameters signaled in a parameter set, and decoding the video data based on these parameters to enable adaptive resolution changes and efficient bitstream extraction.
This approach improves video encoding efficiency by allowing adaptive resolution changes, reducing bandwidth and storage requirements, and enabling improved video encoding and decoding processes.
Smart Images

Figure 0007698020000002 
Figure 0007698020000003 
Figure 0007698020000004
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority from U.S. Provisional Patent Application No. 63 / 037,202, filed on June 10, 2020, and U.S. Patent Application No. 17 / 335,600, filed on June 1, 2021, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure generally relates to the field of data processing, and more particularly to video encoding.
Background Art
[0003] Video encoding and decoding using inter - picture prediction with motion compensation has been known for decades. Uncompressed digital video can be composed of a series of images, and each image has, for example, spatial dimensions of 1920x1080 luminance samples and associated chrominance samples. A series of images can have, for example, a fixed or variable image rate (informally also called the frame rate) of 60 images per second or 60 Hz. Uncompressed video has significant bit - rate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution at a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.
[0004] One of the purposes of video encoding and decoding is to reduce the redundancy of the input video signal by compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, sometimes by more than two orders of magnitude. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. In the case of video, irreversible compression is widely adopted. The amount of distortion tolerated varies depending on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television contribution applications. The achievable compression ratio can reflect that the higher the tolerable / acceptable distortion, the higher the compression ratio can be.
[0005] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy encoding, some of which are introduced below.
[0006] Historically, video encoders and decoders have most often been defined for a coded video sequence (CVS), a group of pictures (GOP), or a similar multi-picture temporal frame, and have tended to operate at a given picture size. For example, in MPEG-2, it is known that the system design can change the horizontal resolution (and thus the picture size) according to factors such as scene activity, but this is done only in I pictures and thus typically for a GOP. Resampling reference pictures to use different resolutions in a CVS is known, for example, from ITU-T Rec. H.263 Annex P. However, here the picture size is not changed and only the reference pictures are resampled, so as a result, only a part of the picture canvas may be used (in the case of downsampling), or only a part of the scene may be captured (in the case of upsampling). Further, H.263 Annex Q enables resampling individual macroblocks by a factor of two (in each dimension) up or down. Again, the picture size remains the same. Since the size of the macroblocks is fixed in H.263, there is no need to signal it.
[0007] In modern video coding, it has become mainstream to change the picture size in predicted pictures. For example, VP9 enables resampling of reference pictures and changing the resolution of the entire picture. Similarly, certain proposals made towards VVC (e.g., including Hendry et al., "On adaptive resolution change (ARC) for VVC", Joint Video Team document JVET-M0135-v1, Jan 9 - 19, 2019, which is entirely incorporated herein) enable resampling the entire reference picture to a different resolution (higher or lower). In this document, it is proposed to code different candidate resolutions in the sequence parameter set and reference them by picture-by-picture syntactic elements in the picture parameter set.
Summary of the Invention
Means for Solving the Problems
[0008] Embodiments relate to a method, a system, and a computer-readable medium for video encoding. According to one aspect, a method for video encoding is provided. The method can include receiving video data having one or more sub-images. Resampling parameters and spatial scalability parameters corresponding to the sub-images are extracted. The resampling and spatial scalability parameters correspond to one or more flags signaled in a parameter set associated with the video data. The video data is decoded based on the extracted resampling and spatial scalability parameters.
[0009] According to another aspect, a computer system for video encoding is provided. The computer system can include program instructions stored in at least one of one or more memories for execution by at least one of one or more processors via at least one of one or more computer-readable memories, one or more computer-readable tangible storage devices, and one or more memories, whereby the computer system can execute the method. The method can include receiving video data having one or more sub-images. Resampling parameters and spatial scalability parameters corresponding to the sub-images are extracted. The resampling and spatial scalability parameters correspond to one or more flags signaled in a parameter set associated with the video data. The video data is decoded based on the extracted resampling and spatial scalability parameters.
[0010] According to yet another aspect, a computer-readable medium for video encoding is provided. The computer-readable medium can include one or more computer-readable storage devices and program instructions stored in at least one of the one or more tangible storage devices, the program instructions being executable by a processor. The program instructions can be executable by the processor to perform a method that optionally includes receiving video data having one or more sub-images. Resampling parameters and spatial scalability parameters corresponding to the sub-images are extracted. The resampling and spatial scalability parameters correspond to one or more flags signaled in a parameter set associated with the video data. The video data is decoded based on the extracted resampling and spatial scalability parameters.
[0011] These and other objects, features, and advantages will become apparent from the following detailed description of exemplary embodiments to be read in conjunction with the accompanying drawings. For the purpose of making the figures easy to understand to facilitate the understanding of those skilled in the art in conjunction with the detailed description, the various features of the drawings are not to scale.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Best Mode for Carrying Out the Invention
[0013] Although detailed embodiments of the claimed structure and method are disclosed herein, it can be understood that the disclosed embodiments are merely examples of the claimed structure and method that can be embodied in various forms. However, these structures and methods can be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure will be complete and full and will fully convey the scope thereof to those skilled in the art. In the description, well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0014] Embodiments generally relate to the field of data processing, and more particularly to video encoding. The exemplary embodiments described below provide, among other things, a system, method, and computer program for bitstream extraction of sub-images in an encoded video stream having multiple layers. Accordingly, some embodiments have the ability to improve the field of computing by enabling improved video encoding and decoding based on reference picture resampling and signaling of spatial scalability parameters in a video bitstream.
[0015] Aspects are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer-readable media according to various embodiments. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0016] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) can include at least two terminals (110-120) interconnected via a network (150). In the case of one-way data transmission, a first terminal (110) can encode video data at a local location for transmission to another terminal (120) via the network (150). A second terminal (120) can receive the encoded video data of another terminal from the network (150), decode the encoded data, and display the restored video data. One-way data playing is common in media serving applications and the like.
[0017] FIG. 1 shows a second pair of terminals (130, 140) provided to support bidirectional transmission of an encoded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal (130, 140) can encode video data captured at a local location for transmission to other terminals via a network (150). Each terminal (130, 140) can also receive encoded video data transmitted by other terminals, decode the encoded data, and display the restored video data on a local display device.
[0018] In FIG. 1, the terminals (110-140) may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (150) represents any number of networks that transmit encoded video data between terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) can exchange data in a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described herein below.
[0019] FIG. 2 shows the arrangement of video encoders and decoders in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, storage of compressed video on digital media such as CDs, DVDs, memory sticks, and the like.
[0020] A streaming system can include, for example, a video source (201) that generates an uncompressed video sample stream (202), and a capture subsystem (213) that can include, for example, a digital camera. That sample stream (202), drawn in thick lines to emphasize that it has a large amount of data when compared to an encoded video bitstream, can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bitstream (204), drawn in thin lines to emphasize that it has a small amount of data compared to the sample stream, can be stored in a streaming server (205) for future use. One or more streaming clients (206, 208) can access the streaming server (205) to retrieve a copy (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210) that decodes an incoming copy of the encoded video bitstream (207) and creates an outgoing video sample stream (211) that can be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. Among those under development is a video encoding standard informally known as Versatile Video Coding or VVC. The disclosed subject matter may be used in the context of VVC.
[0021] FIG. 3 may be a functional block diagram of a video decoder (210) according to an embodiment of the present invention.
[0022] The receiver (310) can receive one or more codec video sequences decoded by the decoder (210), and in the same or another embodiment, can receive one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (312), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) can receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, which may be transferred to their respective using entities (not shown). The receiver (310) can separate the encoded video sequence from other data. To handle network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter, "parser"). If the receiver (310) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may not be necessary or can be made small. For use in a best-effort packet network such as the Internet, the buffer (315) may be necessary, may be relatively large, and advantageously can be of an adaptable size.
[0023] The video decoder (210) can include a parser (320) for reconstructing symbols (321) from an entropy-coded video sequence. The categories of these symbols potentially include information used to manage the operation of the decoder (210) and information for controlling rendering devices such as a display (212) which, as shown in FIG. 2, is not an essential part of the decoder but can be coupled to the decoder. The control information for the rendering device may be in the form of SEI messages (Supplementary Enhancement Information) or VUI (Video Usability Information) parameter set fragments (not shown). The parser (320) can analyze / entropy-decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow principles well known to those skilled in the art such as variable length coding, Huffman coding, arithmetic coding with or without context sensitivity. The parser (320) can extract a set of subgroup parameters for at least one of the subgroups of pixels from the coded video sequence in the video decoder based on at least one parameter corresponding to the group. The subgroups may include Groups of Picture (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The entropy decoder / parser can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0024] The parser (320) can perform an entropy decoding / analysis operation on the video sequence received from the buffer (315) to create symbols (321).
[0025] The reconstruction of the symbol (321) can include a plurality of different units depending on the type of the encoded video image or a part thereof (inter - picture and intra - picture, inter - block and intra - block, etc.) and other factors. How each unit participates can be controlled by subgroup control information parsed from the encoded video sequence by the parser (320). Such a flow of subgroup control information between the parser (320) and the following plurality of units is not depicted for clarity.
[0026] In addition to the function blocks already described, the decoder 210 can be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, it is appropriate to conceptually subdivide it into the following functional units.
[0027] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives, as a symbol (321), the quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. from the parser (320). This unit can output a block containing sample values that can be input to the aggregator (355).
[0028] In some cases, the output samples of the scaler / inverse transform (351) can be related to an intra-coded block, i.e., a block that does not use prediction information from the previously reconstructed image but can use prediction information from the previously reconstructed part of the current image. Such prediction information can be provided by the intra prediction unit (352). In some cases, the intra prediction unit (352) uses the surrounding already reconstructed information fetched from the current (partially reconstructed) image (356) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) can, in some cases, add the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) for each sample.
[0029] In other cases, the output samples of the scaler / inverse transform unit (351) can be related to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit (353) can access the reference image memory (357) to fetch the samples for prediction. After motion compensating the fetched samples according to the symbols (321) related to the block, these samples can be added to the output of the scaler / inverse transform unit by the aggregator (355) (in this case called the residual samples or residual signal) to generate the output sample information. The address in the reference image memory form where the motion compensation unit fetches the prediction samples can be controlled by the motion vectors available to the motion compensation unit, for example, in the form of symbols (321) having X, Y, and reference image components. Motion compensation can also include interpolation of the sample values fetched from the reference image memory when exact sub-sample motion vectors are used, and motion vector prediction mechanisms, etc.
[0030] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technique is controlled by parameters included in the encoded video bitstream and can include in-loop filter techniques made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also respond to meta information obtained during the decoding of the previous part (in decoding order) of the encoded image or encoded video sequence, and can also respond to previously reconstructed and loop filtered sample values.
[0031] The output of the loop filter unit (356) can be not only output to the rendering device (212), but can be a sample stream that can be stored in the reference image memory (356) for use in future inter-picture prediction.
[0032] When a particular encoded image is fully reconstructed, it can be used as a reference image for future prediction. When the encoded image is fully reconstructed and the encoded image is identified as a reference image (e.g., by the parser (320)), the current reference image (356) can become part of the reference image buffer (357), and a new current image memory can be reallocated before starting the reconstruction of subsequent encoded images.
[0033] The video decoder 320 can perform a decoding operation according to a predetermined video compression technique that can be documented in standards such as ITU-T Rec.H.265. The encoded video sequence can comply with the syntax specified by the video compression technique or standard being used, in the sense that it complies with the syntax of the video compression technique or standard as specified in the document or standard of the video compression technique, specifically in the profile document therein. Also, it is necessary for compliance that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The restrictions set by the level may, in some cases, be further restricted by the specifications of the HRD (Hypothetical Reference Decoder) and the metadata for HRD buffer management signaled in the encoded video sequence.
[0034] In one embodiment, the receiver (310) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (320) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0035] FIG. 4 may be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0036] The encoder (203) can receive video samples from a video source (201) (not part of the encoder) that can capture video images to be encoded by the encoder (203).
[0037] The video source (201) can provide a source video sequence to be encoded by an encoder (203) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCB, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual images that give motion when viewed in sequence. The images themselves can be configured as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. One skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0038] According to one embodiment, the encoder (203) can encode and compress the images of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Enforcing an appropriate encoding speed is one function of the controller (450). The controller controls other functional units as described below and is functionally coupled to these units. For clarity, this coupling is not depicted. Parameters set by the controller include rate control related parameters (such as picture skip, quantizer, lambda value of rate distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) since they may be related to a video encoder (203) optimized for a specific system design.
[0039] Among video encoders, there are those that operate in an "encoding loop" that can be easily recognized by those skilled in the art. As an overly simplified explanation, the encoding loop consists of an encoder (430) (hereinafter, the "source coder") that plays a role in generating symbols based on the input image and reference image to be encoded, and a (local) decoder (433) incorporated into the encoder (203) that reconstructs the symbols, and can generate sample data that the (remote) decoder can also generate (in the video compression technology considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to the reference image memory (434). Since the decoding of the symbol stream results in a bit-exact result regardless of the position of the decoder (local or remote), the content of the reference image buffer is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference image samples that the decoder "sees" when using prediction during decoding. This basic principle of the synchronization of the reference image (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0040] The operation of the "local" decoder (433) can be made the same as the operation of the "remote" decoder (210), which was described in detail above in connection with FIG. 3. However, referring briefly to FIG. 3 as well, since the symbols are available and the symbols can be encoded / decoded into the encoded video sequence by the entropy coder (445) and the parser (320) losslessly, the entropy decoding part of the decoder (210) including the channel (312), the receiver (310), the buffer (315), and the parser (320) may not be fully implemented in the local decoder (433).
[0041] What can be said based on the observations at this point is that decoder technologies other than the parsing / entropy decoding existing in the decoder must necessarily exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. Since the description of encoder technologies is the reverse of the decoder technologies described comprehensively, it can be omitted. Only in certain areas is a more detailed description necessary and is provided below.
[0042] As part of its operation, the source coder (430) can perform motion-compensated predictive coding that predictively encodes an input frame by referring to one or more previously encoded frames from a video sequence designated as a "reference frame". In this way, the encoding engine (432) encodes the difference between a pixel block of the input frame and a pixel block of a reference frame that may be selected as a predictive reference for the input frame.
[0043] The local video decoder (433) can decode the encoded video data of a frame that may be designated as a reference frame based on the symbols generated by the source coder (430). The operation of the encoding engine (432) may advantageously be a non-invertible process. If the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (433) can replicate the decoding process that a video decoder can perform on a reference frame and store the reconstructed reference frame in the reference image cache (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame obtained by a remote video decoder (without transmission errors).
[0044] Predictor (435) can perform prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference image memory (434) for sample data (as candidates for reference pixel blocks) or specific metadata such as reference image motion vectors and block shapes that may serve as appropriate prediction references for the new image. The predictor (435) can operate on samples for each pixel block to find an appropriate prediction reference. In some cases, the input image can have prediction references drawn from a plurality of reference images stored in the reference image memory (434) as determined by the search results obtained by the predictor (435).
[0045] Controller (450) can manage the encoding operation of the video coder (430), including, for example, setting parameters and sub-group parameters used for encoding video data.
[0046] The outputs of all the above-described functional units may be entropy encoded in the entropy coder (445). The entropy coder converts the symbols generated by various functional units into an encoded video sequence by reversibly compressing them according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.
[0047] Transmitter (440) can buffer the encoded video sequence created by the entropy coder (445) and prepare it for transmission via a communication channel (460) that may be a hardware / software link to a storage device for storing the encoded video data. The transmitter (440) can merge the encoded video data from the video coder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0048] The controller (450) can manage the operation of the encoder (203). During encoding, the controller (450) can assign a specific encoding image type to each encoded image, which may affect the encoding technique that may be applied to each image. For example, an image may often be assigned as one of the following frame types.
[0049] An Intra Picture (I picture) may be one that can be encoded and decoded without using other frames in the sequence as a prediction source. Some video codecs enable different types of intra pictures, such as Independent Decoder Refresh Pictures. Those skilled in the art are aware of these variations of I pictures, as well as their respective applications and features.
[0050] A Predicted Picture (P picture) may be one that can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0051] A Bi - Directionally Predicted Picture (B picture) may be one that can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0052] Source images are typically spatially subdivided into a plurality of sample blocks (e.g., 4x4, 8x8, 4x8, or 16x16 sample blocks each), and may be encoded block by block. Blocks can be encoded predictively with reference to other (already encoded) blocks, as determined by an encoding assignment applied to each image of the block. For example, blocks of an I image may be encoded non-predictively or may be encoded predictively with reference to already encoded blocks of the same image (spatial prediction or intra prediction). Pixel blocks of a P image may be encoded non-predictively, via spatial prediction, or via temporal prediction, with reference to one previously encoded reference image. Blocks of a B image may be encoded non-predictively, via spatial prediction, or via temporal prediction, with reference to one or two previously encoded reference images.
[0053] Video coder (203) can perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec.H.265. In that operation, video coder (203) can perform various compression operations including a predictive encoding operation that utilizes temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to a syntax specified by the video encoding technique or standard being used.
[0054] In one embodiment, transmitter (440) can transmit additional data along with the encoded video. Video coder (430) can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, and the like.
[0055] Before describing certain aspects of the disclosed subject matter in more detail, it is necessary to introduce some terms that are referred to in the remainder of this specification.
[0056] Hereinafter, a sub-image refers to a rectangular arrangement of samples, blocks, macroblocks, coding units, or similar entities that can, in some cases, be semantically grouped and independently coded at a modified resolution. One or more sub-images can form an image. One or more coded sub-images can form a coded image. One or more sub-images can be assembled into an image, and one or more sub-images can be extracted from an image. In certain environments, one or more coded sub-images can be assembled into a coded image in the compressed domain without transcoding to the sample level, and in the same or certain other cases, one or more coded sub-images can be extracted from the coded image in the compressed domain.
[0057] Hereinafter, Adaptive Resolution Change (ARC) refers to a mechanism that can change the resolution of an image or sub-image within a coded video sequence, for example, by resampling a reference image. Hereinafter, an ARC parameter refers to the control information necessary to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, the resolution of the output and / or reference image, various control flags, and the like.
[0058] The above description has focused on encoding and decoding a single semantically independent coded video image. Before explaining the meaning and its implied additional complexity of encoding / decoding multiple sub-images with independent ARC parameters, the options for signaling ARC parameters can be described.
[0059] Referring to FIG. 5, several new options for signaling ARC parameters are shown. As described for each option, there are certain advantages and disadvantages from the viewpoints of encoding efficiency, complexity, and architecture. A video encoding standard or technology can select one or more of these options, or options known from the prior art, for signaling ARC parameters. The options may not be mutually exclusive and may perhaps be exchanged based on application needs, related standard technologies, or encoder selection.
[0060] The class of ARC parameters may include the following.
[0061] - Up / downsample coefficients separated or combined in the X and Y dimensions.
[0062] - Up / downsample coefficients with an added temporal dimension, which indicate a constant speed zoom-in / zoom-out for a given number of images.
[0063] - Either of the above two may involve encoding of one or more perhaps short syntax elements that may refer to a table containing coefficients.
[0064] - The resolution in the X or Y dimension in units of samples, blocks, macroblocks, CUs, or other appropriate granularity of the synthesized or separate input image, output image, reference image, encoded image. If there are two or more resolutions (e.g., one for the input image, one for the reference image, etc.), in some cases, a set of values may be inferred from another set of values. Such things may be gated, for example, using a flag. For more detailed examples, see below.
[0065] - The "warping" coordinates are similar to those used in H.263 Annex P and are also used at an appropriate granularity as described above. H.263 Annex P defines one efficient method for encoding such warping coordinates, but other, potentially more efficient methods may also be devised. For example, the variable-length reversible "Huffman"-style encoding of the warping coordinates in Annex P can be replaced with a binary encoding of appropriate length, and the length of the binary codeword can be derived, for example, from the maximum image size and, in some cases, multiplied by a specific factor and offset by a specific value to enable "warping" outside the boundaries of the maximum image size.
[0066] - Up or downsampling filter parameters. In the simplest case, there may be only one filter for upsampling and / or downsampling. However, in some cases, it may be advantageous to increase the flexibility of filter design, which may require signaling of filter parameters. Such parameters may be selected via an index of a list of possible filter designs, the filter may be fully specified (e.g., via a list of filter coefficients using appropriate entropy coding techniques), the filter may be implicitly selected via the up / downsampling ratio, and accordingly, signaled according to any of the above mechanisms, etc.
[0067] Subsequently, the description assumes the encoding of a finite set of up / downsampling coefficients (the same coefficients used in both the X and Y dimensions) indicated via a codeword. The codeword can advantageously be encoded in variable length using, for example, the Exp-Golomb code common to certain syntax elements of video coding specifications such as H.264 and H.265. One appropriate mapping of values to up / downsampling coefficients can follow, for example, the following table.
[0068] [Table 1]
[0069] Depending on the needs of the application and the capabilities of the upscaling and downscaling mechanisms available in the video compression technology or standard, many similar mappings can be devised. This table can be extended to more values. The values can also be represented using entropy coding mechanisms other than the Ext-Golomb code, for example binary coding. This has certain advantages, for example in cases where the resampling factor is important outside the video processing engine itself (with the encoder and decoder having top priority), such as by MANE. Note that in the most common case where no resolution change is needed (presumably), a short Ext-Golomb code (only a single bit in the above table) can be selected. This may be more efficient in coding than using a binary code in the most common cases.
[0070] The number of entries in the table and their semantics may be fully or partially configurable. For example, the basic outline of the table may be conveyed in a "higher-level" parameter set such as a sequence or decoder parameter set. Alternatively, or in addition, one or more such tables may be defined in the video coding technology or standard and may be selected, for example, via a decoder or sequence parameter set.
[0071] Subsequently, how the upsampling / downsampling factors (ARC information) encoded as described above are included in the video coding technology or standard syntax will be explained. Similar considerations may apply to one or several codewords that control the up / downsampling filter. For discussion in cases where a relatively large amount of data is required for the filter or other data structures, see below.
[0072] Annex P of H.263 includes the ARC information 502 in the image header 501 in the form of four warping coordinates, specifically in the H.263 PLUSPTYPE(503) header extension. This can be a wise design choice when a) there is an available image header and b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be very large, and the image header may be of a transient nature, so the scaling factor may not be related to the image boundaries.
[0073] JVCET-M135-v1 cited above includes the ARC reference information (505) (index) located within the picture parameter set (504), and indexes a table (506) that includes the target resolution located within the sequence parameter set (507). Placing the possible resolutions in the table (506) of the sequence parameter set (507) can be justified by using the SPS as an interoperability negotiation point during capability exchange according to the language statements of the creator. The resolution can be changed for each image within the limits set by the values in the table (506) by referring to the appropriate picture parameter set (504).
[0074] Referring further to Figure 5, there may be the following additional options for transmitting the ARC information in the video bitstream. Each of these options has specific advantages over the existing technologies as described above. The options may coexist in the same video coding technology or standard simultaneously.
[0075] In one embodiment, ARC information (509), such as resampling (zoom) ratio, can be present in a slice header, GOB header, tile header, or tile group header (hereinafter, tile group header) (508). This is sufficient, for example, when the ARC information, such as a single variable length ue(v) or a fixed length codeword of several bits, is small as described above. There is an additional advantage in directly having the ARC information within the tile group header in that the ARC information may be applicable to a sub-image represented by the tile group, for example, rather than to the entire image. See also below. Additionally, even when a video compression technology or standard assumes only an adaptive resolution change for the entire image (as opposed to, for example, a tile group-based adaptive resolution change), putting the ARC information in the tile group header has certain advantages from the perspective of error recovery compared to putting it in an H.263-style image header.
[0076] In the same or another embodiment, the ARC information (512) itself can be present, for example, within a suitable parameter set (511), such as an image parameter set, header parameter set, tile parameter set, adaptive parameter set (where the adaptive parameter set is shown). It may be advantageous for the scope of that parameter set not to be larger than the image, for example, the tile group. The use of the ARC information is implicitly done through the activation of the relevant parameter set. For example, when a video coding technology or standard contemplates only image-based ARC, an image parameter set or its equivalent may be appropriate.
[0077] In the same or another embodiment, the ARC reference information (513) may be present in a tile group header (514) or a similar data structure. The reference information (513) can refer to a subset of the ARC information (515) available in a parameter set (516) having a scope that extends beyond a single image, such as a sequence parameter set or a decoder parameter set.
[0078] The indirect implicit activation of the additional level of PPS from tile group headers PPS, SPS as used in JVET-M0135-v1, like the sequence parameter set, can be used for capability negotiation or announcement of the picture parameter set (as used in certain standards such as RFC3984), so it seems unnecessary. However, if the ARC information should also be applicable to sub-images represented by, for example, tile groups, a parameter set with an activation scope limited to tile groups, such as an adaptive parameter set or a header parameter set, may be a better choice. Also, when the size of the ARC information exceeds a negligible size (for example, when filter control information such as a large number of filter coefficients is included), the parameter may be reusable for future images or sub-images by referring to the same parameter set, so from the perspective of coding efficiency, it may be a better choice than directly using the header (508).
[0079] When using a sequence parameter set or another higher-level parameter set across multiple images, certain considerations may apply.
[0080] 1. The parameter set for storing the ARC information table (516) may, in some cases, be a sequence parameter set, but in other cases, may advantageously be a decoder parameter set. The decoder parameter set can have a plurality of CVSs, i.e., the active range of all the encoded video bits from the start of the session to the disconnection of the session, i.e., the encoded video stream. Such a range may be a decoder function that may be implemented in hardware, and since the hardware function tends not to change in the CVS (at least in some entertainment systems, it is a group of images with a length of 1 second or less), it may be more appropriate. However, putting the table in the sequence parameter set is explicitly included in the arrangement options described in this specification, particularly in relation to point 2 below.
[0081] 2. The ARC reference information (513) can advantageously be placed directly within the image / slice / tile / GOB / tile group header (hereinafter, tile group header) (514), rather than within the image parameter set, such as JVCET-M0135-v1. The reason is as follows. When the encoder wants to change a single value within the image parameter set, such as the ARC reference information, it is necessary to create a new PPS and refer to that new PPS. Assume that only the ARC reference information is changed and other information, such as the quantization matrix information of the PPS, remains the same. Such information can be quite large and needs to be resent to complete the new PPS. Since the ARC reference information can be a single codeword, such as an index to the table (513), and is the only value that is changed, it is cumbersome and wasteful to resent all the quantization matrix information, for example. In that regard, avoiding indirect reference via the PPS, as proposed in JVET-M0135-v1, can potentially be quite superior from the perspective of coding efficiency. Similarly, putting the ARC reference information in the PPS has the additional drawback that since the scope of the image parameter set activation is the image, the ARC information referred to by the ARC reference information (513) must necessarily be applied to the entire image rather than a sub-image.
[0082] In the same or another embodiment, the signaling of the ARC parameters can follow a detailed example as outlined in FIG. 6. FIG. 6 shows a syntax diagram in the representation used in video coding standards since at least 1993. The notation of such a syntax diagram generally follows C-style programming. The bold lines indicate the syntax elements present in the bitstream, and the non-bold lines often indicate control flow or variable settings.
[0083] As an exemplary syntax structure of a header applicable to a (presumably rectangular) portion of an image, a tile group header (601) can conditionally include a variable-length Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold). The presence of this syntax element within the tile group header, here the value of a flag not shown in bold, can gate with respect to the use of adaptive resolution (603), which means that the flag is present in the bitstream at the point where it occurs in the syntax diagram. Whether adaptive resolution is used for this image or a part thereof may be signaled in any high-level syntax structure, either internal or external to the bitstream. In the example shown, it is signaled in the sequence parameter set, as outlined below.
[0084] Referring further to FIG. 6, an excerpt of the sequence parameter set (610) is also shown. The first syntax element shown is adaptive_pic_resolution_change_flag (611). When true, that flag can indicate the use of adaptive resolution, which may require certain control information. In this example, such control information is conditionally present based on the value of the flag based on the if() statement of the parameter set (612) and the tile group header (601).
[0085] When adaptive resolution is being used, in this example, what is encoded is the output resolution in sample units (613). The number 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, which together can define the resolution of the output image. In other places in the video encoding technology or standard, specific restrictions for any value can be defined. For example, in the level definition, the total number of output samples may be restricted, which may be the product of the values of these two syntax elements. Also, a particular video encoding technology or standard, or an external technology or standard such as a system standard, may restrict the range of values (for example, one or both dimensions need to be divisible by a power of 2) or the aspect ratio (for example, the width and height need to be in a relationship such as 4:3 or 16:9). Such restrictions may be introduced to facilitate the hardware implementation or for other reasons and are well known in the art.
[0086] In certain applications, it may be advisable for the encoder to instruct the decoder to use a specific reference image size rather than implicitly assuming that the reference image size is the output image size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference image dimensions (615) (again, the numbers indicate both width and height).
[0087] Finally, a table showing the possible widths and heights of the decoded images is presented. Such a table can be represented, for example, by table display (num_dec_pic_size_in_luma_samples_minus1) (616). "minus1" can be referred to for the interpretation of the value of that syntax element. For example, if the encoded value is zero, there is one table entry. If the value is 5, there are six table entries. Then, for each "row" in the table, the width and height of the decoded image are included in the syntax (617).
[0088] The indicated table entry (617) can be indexed using the syntax element dec_pic_size_idx(602) within the tile group header, allowing for different decoding sizes (actually zoom ratios) for each tile group.
[0089] In certain video encoding techniques or standards, such as VP9, to enable spatial scalability, along with temporal scalability, spatial scalability is supported by implementing a specific form of reference picture resampling (signaled in a way completely different from the disclosed subject matter). In particular, certain reference pictures can be upsampled to a higher resolution using ARC-style techniques and can form the basis of a spatial enhancement layer. These upsampled pictures can be refined using normal prediction mechanisms at a high resolution to add details.
[0090] The disclosed subject matter can be used in such an environment. In some cases, in the same or different embodiments, values within the NAL unit header, such as the Temporal ID field, can be used to indicate not only temporal layers but also spatial layers. Doing so has certain advantages for a particular system design. For example, an existing Selected Forwarding Unit (SFU) created and optimized for the selected transmission of a temporal layer based on the Temporal ID value of the NAL unit header can be used without modification for a scalable environment. To enable this, there may be requirements for the mapping between the encoded picture size and the temporal layer, indicated by the temporal ID field within the NAL unit header.
[0091] In some video encoding techniques, an access unit (AU) may refer to an encoded image, slice, tile, NAL unit, etc., that is captured at a given time instance and synthesized into the bitstream of each image / slice / tile / NAL unit. This time instance can be the composition time.
[0092] In HEVC and other specific video encoding techniques, a picture order count (POC) value can be used to indicate a selected reference image among a plurality of reference images stored in a decoded picture buffer (DPB). When an access unit (AU) is composed of one or more images, slices, or tiles, each image, slice, or tile belonging to the same AU may carry the same POC value, from which it can be inferred that they are created from content with the same composition time. In other words, in a scenario where two images / slices / titles carry the same specific POC value, it can be shown that the two images / slices / titles belong to the same AU and have the same composition time. Conversely, two images / titles / slices with different POC values can indicate that each image / slice / tile belongs to a different AU and has a different composition time.
[0093] In an embodiment of the disclosed subject matter, the strict relationship described above can be relaxed in that an access unit can include images, slices, or tiles having different POC values. By enabling different POC values within an AU, it becomes possible to use the POC value to identify potentially independently decodable images / slices / titles at the same presentation time. This can enable support for multiple scalable layers without changing the reference image selection signaling (e.g., reference image set signaling or reference image list signaling), as will be described in more detail below.
[0094] However, for other images / slices / tiles having different POC values, it is still desirable to be able to identify the AU to which the image / slice / tiles belong from the POC value alone. This can be achieved as described below.
[0095] In the same or other embodiments, an access unit count (AUC) may be signaled in a high-level syntax structure such as a NAL unit header, a slice header, a tile group header, an SEI message, a parameter set, or an AU delimiter. The value of the AUC can be used to identify which NAL unit, image, slice, or tile belongs to a given AU. The value of the AUC may correspond to an individual composition time instance. The AUC value may be equal to a multiple of the POC value. The AUC value can be calculated by dividing the POC value by an integer value. In some cases, the division operation may impose a particular burden on the decoder implementation. In such cases, the division operation can be replaced by a shift operation by imposing a slight restriction on the numbering space of the AUC value. For example, the AUC value may be equal to the most significant bit (MSB) value of the POC value range.
[0096] In the same embodiment, the value of the POC cycle per AU (poc_cycle_au) may be signaled in a high-level syntax structure such as a NAL unit header, a slice header, a tile group header, an SEI message, a parameter set, or an AU delimiter. The poc_cycle_au can indicate how many different consecutive POC values can be associated with the same AU. For example, when the value of poc_cycle_au is equal to 4, an image, slice, or tile with POC values from 0 to 3 including both end values is associated with an AU with an AUC value of 0, and an image, slice, or tile with POC values from 4 to 7 including both end values is associated with an AU with an AUC value of 1. Therefore, the value of AUC can be inferred by dividing the POC value by the value of poc_cycle_au.
[0097] In the same or another embodiment, the value of poc_cyle_au may be derived from information located, for example, within a video parameter set (VPS) that identifies the number of spatial or SNR layers within an encoded video sequence. Such a possible relationship will be briefly described below. Although such derivation as described above may save several bits of the VPS and improve the encoding efficiency, it may be advantageous to explicitly encode poc_cycle_au in an appropriate high-level syntax structure hierarchically lower than the video parameter set so that poc_cycle_au can be minimized for a specific small portion of the bitstream such as an image. This optimization makes it possible to encode the POC value (and / or the value of the syntax element that indirectly references the POC) in a low-level syntax structure, and thus it may be possible to save more bits than the bits saved in the above derivation process.
[0098] In the same or another embodiment, FIG. 9 shows, in a VPS (or SPS), the syntax element vps_poc_cycle_au indicating the poc_cycle_au used for all pictures / slices of an encoded video sequence, and in a slice header, the syntax element slice_poc_cycle_au indicating the poc_cycle_au of the current slice, and shows an example of a syntax table for signaling them. When the POC value increases uniformly for each AU, vps_contant_poc_cycle_per_au in the VPS is set equal to 1, and in this case where vps_poc_cycle_au in the VPS is being signaled, slice_poc_cycle_au is not signaled explicitly, and the value of AUC for each AU is calculated by dividing the POC value by vps_poc_cycle_au. When the POC value does not increase uniformly for each AU, vps_contant_poc_cycle_per_au in the VPS is set to 0. In this case, vps_access_unit_cnt is not signaled, but slice_access_unit_cnt is signaled in the slice header of each slice or picture. Each slice or picture may have a different value of slice_access_unit_cnt. The value of AUC for each AU is calculated by dividing the POC value by slice_poc_cycle_au. FIG. 10 shows a block diagram illustrating the associated workflow.
[0099] In the same or other embodiments, even if the POC values of pictures, slices, or tiles may be different, pictures, slices, or tiles corresponding to AUs having the same AUC value can be associated with the same decoding or output time instance. Thus, all or a subset of the pictures, slices, or tiles associated with the same AU may be decoded in parallel and output simultaneously without any inter-analysis / decoding dependency between the pictures, slices, or tiles within the same AU.
[0100] In the same or other embodiments, even if the POC values of the images, slices, or tiles may be different, the images, slices, or tiles corresponding to the AUs having the same AUC value may be associated with the same configuration / display time instance. If the composition time is included in the container format, even if the images correspond to different AUs, the images can be displayed at the same time instance as long as the composition times of the images are the same.
[0101] In the same or other embodiments, each image, slice, or tile can have the same temporal identifier (temporal_id) within the same AU. All or a subset of the images, slices, or tiles corresponding to a time instance may be associated with the same time sublayer. In the same or other embodiments, each image, slice, or tile can have the same or different spatial layer ids (layer_id) within the same AU. All or a subset of the images, slices, or tiles corresponding to a time instance may be associated with the same or different spatial layers.
[0102] The techniques for signaling the adaptive resolution parameters described throughout may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, FIG. 7 shows a computer system 700 suitable for implementing certain embodiments of the disclosed subject matter.
[0103] The computer software can be coded using any suitable machine language or computer language, which are the subject of mechanisms such as assembly, compilation, linking, etc., and can generate code including instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.
[0104] The command can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0105] The components shown in FIG. 7 for computer system 700 are exemplary in nature and are not intended to suggest any limitation regarding the scope or functionality of the computer software implementing the embodiments of the present disclosure. Also, the component configuration should not be construed as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiment of computer system 700.
[0106] Computer system 700 can include specific human interface input devices. Such human interface input devices can respond to input by one or more human users, for example, via tactile input (keystrokes, swipes, movement of a data glove, etc.), voice input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious human input, such as voice (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).
[0107] The input human interface device can include one or more (only one of each shown) of keyboard 701, mouse 702, trackpad 703, touch screen 710, data glove 704, joystick 705, microphone 706, scanner 707, camera 708.
[0108] Computer system 700 can also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices can include tactile output devices (e.g., tactile feedback by touch screen 710, data glove 704, or joystick 705, but a tactile feedback device that does not function as an input device), audio output devices (such as speaker 709, headphones (not shown), etc.), visual output devices (including CRT screens, LCD screens, plasma screens, OLED screens, each with or without a touch screen input function, each with or without a tactile feedback function, some of which can output 2D visual output or output beyond 3D through means such as stereographic output, virtual reality glasses (not shown), holographic display, smoke tank (not shown), etc., such as screen 710), and printers (not shown)).
[0109] Computer system 700 can also include optical media such as CD / DVD ROM / RW 720 with media 721 such as CD / DVD, thumb drive 722, removable hard drive or solid state drive 723, legacy magnetic media such as tapes and floppy disks (not shown), special ROM / ASIC / PLD-based devices such as security dongles (not shown), etc., including human-accessible storage devices and their associated media.
[0110] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.
[0111] Computer system 700 can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network typically requires an external network interface adapter connected to a particular general-purpose data port or peripheral bus (749) (e.g., a USB port of computer system 700, others are usually incorporated into the core of computer system 700 by attaching to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 700 can communicate with other entities. Such communication can be unidirectional, receive only (e.g., TV broadcast), transmit only unidirectional (e.g., from CANbus to a particular CANbus device), or bidirectional (e.g., to another computer system using a local or wide area digital network). Particular protocols and protocol stacks can be used with each of these networks and network interfaces as described above.
[0112] The aforementioned human interface device, human-accessible memory device, and network interface can be attached to the core 740 of computer system 700.
[0113] The core 740 can include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a dedicated programmable processing device in the form of a field programmable gate array (FPGA) 743, a hardware accelerator 744 for specific tasks, etc. These devices can be connected via a system bus 748 together with a read-only memory (ROM) 745, a random access memory 746, an internal mass storage such as an internal hard drive inaccessible to users, an SSD 747, etc. In some computer systems, the system bus 748 can be accessible in the form of one or more physical plugs, enabling expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the core system bus 748 or via a peripheral bus 749. Architectures of peripheral buses include PCI, USB, etc.
[0114] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can be combined to execute specific instructions that can constitute the aforementioned computer code. The computer code can be stored in the ROM 745 or the RAM 746. Also, transient data can be stored in the RAM 746, while persistent data can be stored, for example, in the internal mass storage device 747. By using a cache memory that can be closely associated with one or more CPUs 741, GPUs 742, mass storage device 747, ROM 745, RAM 746, etc., fast storage and retrieval to any of the memory devices can be enabled.
[0115] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be those specifically designed and constructed for the purposes of this disclosure or may be of the available types well known to those skilled in the field of computer software.
[0116] By way of example and not limitation, a computer system 700 having an architecture, specifically a core 740, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be media associated with specific storage devices of the core 740 that are of a non-transitory nature, such as the user-accessible mass storage devices introduced above, as well as mass storage devices 747 or ROM 745 within the core. The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core 740.
[0117] The computer-readable media can include one or more memory devices or chips, depending on specific needs. The software can cause the core 740, specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM 746 and modify such data structures according to processes defined by the software, thereby executing specific processes or specific portions of specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic embodied in a circuit (e.g., accelerator 744) that is hardwired or otherwise embodied to operate instead of or in conjunction with software to execute specific processes or specific portions of specific processes described herein.
[0118] Where appropriate, references to software can include logic, and vice versa. References to computer-readable media can, where appropriate, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0119] FIG. 8 shows an example of a video sequence structure with adaptive resolution change according to a combination of temporal_id, layer_id, POC, and AUC values. In this example, an image, slice, or tile within the first AU having AUC = 0 can have temporal_id = 0 and layer_id = 0 or 1, while an image, slice, or tile within the second AU having AUC = 1 can have temporal_id = 1 and layer_id = 0 or 1 respectively. The value of POC increases by 1 for each image regardless of the values of temporal_id and layer_id. In this example, the value of poc_cycle_au can be equal to 2. Preferably, the value of poc_cycle_au may be set equal to the number of (spatial scalability) layers. Thus, in this example, the value of POC increases by only 2 and the value of AUC increases by only 1.
[0120] In the above embodiments, all or a subset of the inter-image or inter-layer prediction structure and reference image display may be supported by using the existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC. In RPS or RPL, the selected reference image is indicated by signaling the value of POC or the POC delta value between the current image and the selected reference image. For the disclosed subject matter, RPS and RPL can be used to indicate the inter-image or inter-layer prediction structure without changing the signaling, but there are the following limitations. If the value of temporal_id of the reference image is greater than the value of temporal_id of the current image, the current image may not use the reference image for motion compensation or other prediction. If the value of layer_id of the reference image is greater than the value of layer_id of the current image, the current image may not use the reference image for motion compensation or other prediction.
[0121] In the same and other embodiments, the scaling of motion vectors based on POC differences for temporal motion vector prediction may be disabled across multiple images within an access unit. Thus, each image may have a different POC value within the access unit, but the motion vectors are not scaled and are not used for temporal motion vector prediction within the access unit. This is because reference images having different POCs within the same AU are considered to be reference images having the same temporal instance. Thus, in this embodiment, when the reference image belongs to the AU associated with the current image, the motion vector scaling function can return 1.
[0122] In the same and other embodiments, the scaling of motion vectors based on POC differences for temporal motion vector prediction may optionally be disabled across multiple images when the spatial resolution of the reference image is different from the spatial resolution of the current image. When motion vector scaling is permitted, the motion vectors are scaled based on both the POC difference and the spatial resolution ratio between the current image and the reference image.
[0123] In the same or another embodiment, particularly when poc_cycle_au has a non-uniform value (when vps_contant_poc_cycle_per_au == 0), for temporal motion vector prediction, the motion vectors may be scaled based on AUC differences rather than POC differences. Otherwise (when vps_contant_poc_cycle_per_au == 1), the scaling of motion vectors based on AUC differences may be the same as the scaling of motion vectors based on POC differences.
[0124] In the same or another embodiment, when the motion vectors are scaled based on AUC differences, the reference motion vectors of the same AU (having the same AUC value) as the current image are not scaled based on AUC differences and are used for motion vector prediction regardless of whether scaling is performed, based on the spatial resolution ratio between the current image and the reference image.
[0125] In the same and other embodiments, the AUC value is used to identify the boundary of the AU and is used in the virtual reference decoder (HRD) operation that requires input and output timings at the AU granularity. In most cases, the decoded image with the top layer within the AU is output for display. The AUC value and the layer_id value can be used to identify the output image.
[0126] In one embodiment, an image can be composed of one or more sub-images. Each sub-image can cover a local area or the entire area of the image. The area supported by a sub-image may or may not overlap with the area supported by another sub-image. The area covered by one or more sub-images may or may not cover the entire area of the image. When an image is composed of sub-images, the area supported by the sub-images is the same as the area supported by that image.
[0127] In the same embodiment, the sub-images may be encoded by an encoding method similar to the encoding method used for the encoded image. The sub-images may be encoded independently, or may be encoded depending on another sub-image or the encoded image. The sub-images may or may not have an analytical dependency relationship from another sub-image or the encoded image.
[0128] In the same embodiment, the encoded sub-images may be included in one or more layers. The encoded sub-images within a layer may have different spatial resolutions. The original sub-images may be spatially resampled (upsampled or downsampled), encoded with different spatial resolution parameters, and included in the bitstream corresponding to the layer.
[0129] In the same or another embodiment, a sub-image of (W, H), where W represents the width of the sub-image and H represents the height of the sub-image, may be encoded and included in the encoded bitstream corresponding to layer 0. On the other hand, a sub-image of (W*S w,k , H*S h,k ), which is upsampled (or downsampled) from the sub-image with the original spatial resolution, may be encoded and included in the encoded bitstream corresponding to layer k, where S w,k、 S h,k represents the resampling ratio in the horizontal and vertical directions. When the value of S w,k、 S h,k is greater than 1, the resampling is equal to upsampling. On the other hand, when the value of S w,k、 S h,k is less than 1, the resampling is equal to downsampling.
[0130] In the same or another embodiment, the encoded sub-images within a layer may have a visual quality different from that of the encoded sub-images within another layer in the same or different sub-images. For example, sub-image i of layer n is encoded with quantization parameter Q i,n , while sub-image j of layer m is encoded with quantization parameter Q j,m .
[0131] In the same or another embodiment, the encoded sub-images within a layer may be independently decodable without any analytical or decoding dependency from the encoded sub-images within another layer in the same local region. A sub-image layer that can be independently decoded without referring to another sub-image layer in the same local region is an independent sub-image layer. The encoded sub-images within an independent sub-image layer may or may not have a decoding or analytical dependency from the previously encoded sub-images within the same sub-image layer, but the encoded sub-images may not have a dependency from the encoded images within another sub-image layer.
[0132] In the same or another embodiment, the encoded sub-images within a layer may be dependently decodable depending on any analysis or decoding from encoded sub-images within another layer of the same local region. A sub-image layer that is dependently decodable by referring to another sub-image layer of the same local region is a dependent sub-image layer. The encoded sub-images within a dependent sub-image may be able to refer to encoded sub-images belonging to the same sub-image, previously encoded sub-images within the same sub-image layer, or both reference sub-images.
[0133] In the same or another embodiment, an encoded sub-image is composed of one or more independent sub-image layers and one or more dependent sub-image layers. However, at least one independent sub-image layer may exist for the encoded sub-image. An independent sub-image layer can have a value of a layer identifier (layer_id) that may exist in the NAL unit header or another high-level syntax structure, and this value is equal to 0. A sub-image layer with layer_id equal to 0 is a base sub-image layer.
[0134] In the same or another embodiment, an image may be composed of one or more foreground sub-images and one background sub-image. The region supported by the background sub-image may be equal to the region of the image. The region supported by the foreground sub-images may overlap with the region supported by the background sub-image. The background sub-image may be a base sub-image layer, while the foreground sub-images may be non-base (enhancement) sub-image layers. One or more non-base sub-image layers can refer to the same base layer for decoding. Each non-base sub-image layer with layer_id equal to a can refer to a non-base sub-image layer with layer_id equal to b, where a is greater than b.
[0135] In the same or another embodiment, the image may be composed of one or more foreground sub-images with or without a background sub-image. Each sub-image can have its own base sub-image layer and one or more non-base (enhancement) layers. Each base sub-image layer may be referenced by one or more non-base sub-image layers. Each non-base sub-image layer with layer_id equal to a can reference a non-base sub-image layer with layer_id equal to b, where a is greater than b.
[0136] In the same or another embodiment, the image may be composed of one or more foreground sub-images with or without a background sub-image. Each encoded sub-image within a (base or non-base) sub-image layer may be referenced by one or more sub-images of non-base layers belonging to the same sub-image and one or more sub-images of non-base layers not belonging to the same sub-image.
[0137] In the same or another embodiment, the image may be composed of one or more foreground sub-images with or without a background sub-image. The sub-images within layer a may be further divided into a plurality of sub-images within the same layer. One or more encoded sub-images within layer b can reference the divided sub-images within layer a.
[0138] In the same or another embodiment, a coded video sequence (CVS) may be a group of coded images. The CVS may be composed of one or more coded sub-picture sequences (CSPS), where the CSPS may be a group of coded sub-images covering the same local region of the image. The CSPS can have a temporal resolution the same as or different from that of the coded video sequence.
[0139] In the same or another embodiment, the CSPS may be encoded and included within one or more layers. The CSPS may be composed of one or more CSPS layers. By decoding one or more CSPS layers corresponding to the CSPS, a sequence of sub-images corresponding to the same local region can be reconstructed.
[0140] In the same or another embodiment, the number of CSPS layers corresponding to the CSPS may be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0141] In the same or another embodiment, the CSPS layer can have a different temporal resolution (e.g., frame rate) from another CSPS layer. The original (uncompressed) sub-image sequence may be temporally resampled (upsampled or downsampled), encoded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.
[0142] In the same or another embodiment, a sub-image sequence having a frame rate F may be encoded and included in the encoded bitstream corresponding to layer 0, while the original sub-image sequence temporally upsampled (or downsampled) from the one having F*S t,k may be encoded and included in the encoded bitstream corresponding to layer k, where S t,k is the temporal sampling ratio for layer k. When the value of S t,k is greater than 1, the temporal resampling process is equivalent to an up-conversion of the frame rate. On the other hand, when the value of S t,k is less than 1, the temporal resampling process is equivalent to a down-conversion of the frame rate.
[0143] In the same or another embodiment, when a sub-image having a CSPS layer a is referenced by a sub-image having a CSPS layer b for motion compensation or any inter-layer prediction, if the spatial resolution of the CSPS layer a is different from the spatial resolution of the CSPS layer b, the decoded pixels in the CSPS layer a are resampled and used for reference. In the resampling process, upsampling filtering or downsampling filtering may be required.
[0144] FIG. 11 shows an exemplary video stream including a background video CSPS with layer_id equal to 0 and a plurality of foreground CSPS layers. The encoded sub-image may be composed of one or more CSPS layers, but the background area that does not belong to any foreground CSPS layer may be composed of a base layer. The base layer can include the background area and the foreground area, while the enhancement CSPS layer includes the foreground area. The enhancement CSPS layer may have better visual quality than the base layer in the same area. The enhancement CSPS layer can refer to the reconstructed pixels corresponding to the same area and the motion vectors of the base layer.
[0145] In the same or another embodiment, the video bitstream corresponding to the base layer is included in a track, while the CSPS layer corresponding to each sub-image is included in a separate track in the video file.
[0146] In the same or another embodiment, the video bitstream corresponding to the base layer is included in a track, while the CSPS layer having the same layer_id is included in a separate track. In this example, the track corresponding to layer k includes only the CSPS layer corresponding to layer k.
[0147] In the same or another embodiment, each CSPS layer of each sub-image is stored in a separate track. Each track may or may not have an analytical or decoding dependency from one or more other tracks.
[0148] In the same or another embodiment, each track can include a bitstream corresponding to layers i to j among all or a subset of the CSPS layers of the sub-image, where 0 < i <= j <= k and k is the topmost layer of the CSPS.
[0149] In the same or another embodiment, the image is composed of one or more associated media data including a depth map, an alpha map, 3D geometry data, an occupancy map, etc. Such associated time-limited media data can each be divided into one or more data sub-streams corresponding to one sub-image.
[0150] In the same or another embodiment, FIG. 12 shows an example of a video conference based on a multi-layer sub-image method. The video stream includes one base-layer video bitstream corresponding to the background image and one or more enhancement-layer video bitstreams corresponding to the foreground sub-images. Each enhancement-layer video bitstream corresponds to a CSPS layer. On the display, the image corresponding to the base layer is displayed by default. This includes one or more picture-in-picture (PIP) of the users. When a specific user is selected under the control of the client, the enhancement CSPS layer corresponding to the selected user is decoded and displayed with improved quality or spatial resolution. FIG. 13 is a diagram showing the operation.
[0151] In the same or another embodiment, a network middle box (such as a router) can select a subset of layers to send to the user according to its bandwidth. The composition of the image / sub-image can be used for bandwidth adaptation. For example, if the user has no bandwidth, the router can delete layers or select some sub-images according to importance or based on the setup being used, which can be done dynamically to adapt to the bandwidth.
[0152] FIG. 14 shows a use case of 360-degree video. When a spherical 360-degree image is projected onto a planar image, the projected 360-degree image may be divided into a plurality of sub-images as a base layer. Enhancement layers of specific sub-images can be encoded and sent to the client. The decoder may be able to decode both the base layer including all sub-images and the enhancement layers of the selected sub-images. If the current viewport is the same as the selected sub-image, the displayed image may be of higher quality with the decoded sub-image having the enhancement layer. Otherwise, the decoded image with the base layer may be displayed with low quality.
[0153] In the same or another embodiment, any layout information for display may exist in the file as supplementary information (such as SEI messages or metadata). One or more decoded sub-images may be rearranged and displayed according to the signaled layout information. The layout information may be signaled by a streaming server or broadcaster, may be regenerated by a network entity or cloud server, or may be determined by the user's customized settings.
[0154] In one embodiment, when the input image is divided into one or more (rectangular) sub-regions, each sub-region may be encoded as an independent layer. Each independent layer corresponding to a local region can have a unique layer_id value. For each independent layer, the size and position information of the sub-image may be signaled. For example, the image size (width, height), and the offset information (x_offset, y_offset) of the upper left corner. FIG. 15 shows an example of the layout of the divided sub-images, the size and position information of the sub-images, and its corresponding image prediction structure. The layout information including the sub-image size and sub-image position may be signaled in a high-level syntax structure such as a parameter set, a slice or tile group header, or an SEI message.
[0155] In the same embodiment, each sub-image corresponding to an independent layer can have its unique POC value within an AU. When the reference images among the images stored in the DPB are indicated using the syntax elements of the RPS or RPL structure, the POC value of each sub-image corresponding to the layer can be used.
[0156] In the same or another embodiment, POC (delta) values may be used to indicate the prediction structure (between layers) without using layer_id.
[0157] In the same embodiment, a sub-image whose POC value corresponding to a certain layer (or local region) is equal to N may or may not be used as a reference image for motion compensation prediction of a sub-image whose POC value corresponding to the same layer (or the same local region) is equal to N+K. In most cases, the value of the number K may be equal to the maximum number of (independent) layers, which may be equal to the number of sub-regions.
[0158] In the same or another embodiment, FIG. 16 shows an extended case of FIG. 15. When the input image is divided into a plurality (e.g., four) of sub-regions, each local region may be encoded with one or more layers. In this case, the number of independent layers may be equal to the number of sub-regions, or one or more layers may correspond to the sub-regions. Thus, each sub-region may be encoded with one or more independent layers and zero or more dependent layers.
[0159] In the same embodiment, in FIG. 16, the input image may be divided into four sub-regions. The upper-right sub-region may be encoded as two layers, layer 1 and layer 4, while the lower-right sub-region may be encoded as two layers, layer 3 and layer 5. In this case, layer 4 can refer to layer 1 for motion compensation prediction, while layer 5 can refer to layer 3 for motion compensation.
[0160] In the same or another embodiment, in-loop filtering (such as deblocking filtering, adaptive in-loop filtering, reshaper, bidirectional filtering, or any deep learning-based filtering, etc.) across layer boundaries may be (optionally) disabled.
[0161] In the same or another embodiment, motion compensation prediction or intra-block copy across layer boundaries may be (optionally) disabled.
[0162] In the same or another embodiment, boundary padding for motion compensation prediction or in-loop filtering at the boundaries of sub-images can be processed as desired. A flag indicating whether the boundary padding is processed may be signaled in a high-level syntax structure such as a parameter set (VPS, SPS, PPS, or APS), a slice or tile group header, or an SEI message.
[0163] In the same or another embodiment, the layout information of the sub-region (or sub-image) may be signaled in the VPS or SPS. FIG. 17 shows an example of the syntax elements of the VPS and SPS. In this example, the vps_sub_picture_dividing_flag is signaled in the VPS. The flag can indicate whether the input image is divided into a plurality of sub-regions. When the value of the vps_sub_picture_dividing_flag is equal to 0, the input image of the coded video sequence corresponding to the current VPS may not be divided into a plurality of sub-regions. In this case, the input image size may be equal to the coded image size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS. When the value of the vps_sub_picture_dividing_flag is equal to 1, the input image may be divided into a plurality of sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled in the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be equal to the width and height of the input image, respectively.
[0164] In the same embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding, but may be used for synthesis and display.
[0165] In the same embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntax elements pic_offset_x and pic_offset_y may be signaled in the SPS corresponding to a specific layer. In this case, the sizes of the coded pictures (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS may be equal to the width and height of the sub-region corresponding to a specific layer. Also, the position (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may be signaled in the SPS.
[0166] In the same embodiment, the position information (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may not be used for decoding, but may be used for synthesis and display.
[0167] In the same or another embodiment, the layout information (size and position) of the sub-regions of all or a subset of the input image, and the dependency information between layers may be signaled in a parameter set or an SEI message. FIG. 18 shows an example of a syntax element for indicating information regarding the layout of sub-regions, the dependencies between layers, and the relationship between sub-regions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) sub-regions within the current encoded video sequence, and the syntax element num_layers indicates the number of layers within the current encoded video sequence. The value of num_layers may be greater than or equal to the value of num_sub_region. If any sub-region is encoded as a single layer, the value of num_layers may be equal to the value of num_sub_region. If one or more sub-regions are encoded as multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dependency_flag[i][j] indicates the dependency from the j-th layer to the i-th layer. num_layers_for_region[i] indicates the number of layers associated with the i-th sub-region. sub_region_layer_id[i][j] indicates the layer_id of the j-th layer associated with the i-th sub-region. sub_region_offset_x[i] and sub_region_offset_y[i] indicate the horizontal and vertical positions of the upper left corner of the i-th sub-region, respectively. sub_region_width[i] and sub_region_height[i] indicate the width and height of the i-th sub-region, respectively.
[0168] In one embodiment, one or more syntax elements that specify an output layer set indicating one of more layers that are output regardless of the presence or absence of profile tier level information may be signaled in a high-level syntax structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message. Referring to FIG. 19, a syntax element num_output_layer_sets indicating the number of output layer sets (OLS) within an encoded video sequence that refers to a VPS may be signaled in the VPS. For each output layer set, an output_layer_flag may be signaled the same number of times as the number of output layers.
[0169] In the same embodiment, an output_layer_flag[i] equal to 1 specifies that the i-th layer is output. A vps_output_layer_flag[i] equal to 0 specifies that the i-th layer is not output.
[0170] In the same or another embodiment, one or more syntax elements that specify profile tier level information for each output layer set may be signaled in a high-level syntax structure, such as a VPS, DPS, SPS, PPS, APS, or SEI message. Further referring to FIG. 19, a syntax element num_profile_tile_level indicating the number of profile tier level information per OLS in an encoded video sequence that refers to a VPS may be signaled in the VPS. For each output layer set, a set of syntax elements for the profile tier level information, or an index indicating a specific profile tier level information among the entries within the profile tier level information, may be signaled the same number of times as the number of output layers.
[0171] In the same embodiment, profile_tier_level_idx[i][j] specifies an index within a list of profile_tier_level() syntax structures in the VPS for the profile_tier_level() syntax structure applied to the j-th layer of the i-th OLS.
[0172] In the same or another embodiment, referring to FIG. 20, when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0), the syntax elements num_profile_tile_level and / or num_output_layer_sets may be signaled.
[0173] In the same or another embodiment, referring to FIG. 20, a syntax element vps_output_layers_mode[i] indicating the mode of output layer signaling for the i-th output layer set may be present in the VPS.
[0174] In the same embodiment, vps_output_layers_mode[i] equal to 0 specifies outputting only the top layer having the i-th output layer set. When vps_output_layer_mode[i] is 1, it specifies outputting all layers having the i-th output layer set. When vps_output_layer_mode[i] is equal to 2, it specifies that the layers to be output are the layers for which vps_output_layer_flag[i][j] for the i-th output layer set is equal to 1. More values may be reserved.
[0175] In the same embodiment, output_layer_flag[i][j] may or may not be signaled depending on the value of vps_output_layers_mode[i] for the i-th output layer set.
[0176] In the same or another embodiment, referring to FIG. 20, the flag vps_ptl_signal_flag[i] can exist for the i-th output layer set. Depending on the value of vps_ptl_signal_flag[i], the information of the profile layer level of the i-th output layer set may or may not be signaled.
[0177] In the same or another embodiment, referring to FIG. 21, the number of sub-images max_subpics_minus1 in the current CVS may be signaled in a high-level syntax structure, such as in a VPS, DPS, SPS, PPS, APS or SEI message.
[0178] In the same embodiment, referring to FIG. 21, when the number of sub-images is greater than 1 (max_subpics_minus1>0), the sub-image identifier sub_pic_id[i] of the i-th sub-image may be signaled.
[0179] In the same or another embodiment, one or more syntax elements indicating the sub-image identifiers belonging to each layer of each output layer set may be signaled in the VPS. Referring to FIG. 22, sub_pic_id_layer[i][j][k] indicates the k-th sub-image existing in the j-th layer of the i-th output layer set. Using this information, the decoder can recognize which sub-images can be decoded and output for each layer of a specific output layer set.
[0180] In one embodiment, a picture header (PH) is a syntax structure including syntax elements applied to all slices of an encoded image. A picture unit (PU) is a set of NAL units that are associated with each other according to specified classification rules, are consecutive in decoding order, and strictly contain one encoded image. A PU can include a picture header (PH) and one or more VCL NAL units constituting the encoded image.
[0181] In one embodiment, the SPS (RBSP) may be available for use in the decoding process before being referenced, may be included in at least one AU with a TemporalId equal to 0, or may be provided through external means.
[0182] In one embodiment, the SPS (RBSP) may be available for use in the decoding process before being referenced, may be included in at least one AU with a TemporalId equal to 0 in the CVS and including one or more PPSs that reference the SPS, or may be provided through external means.
[0183] In one embodiment, the SPS (RBSP) may be available for use in the decoding process before being referenced by one or more PPSs, may be included in at least one PU with a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit within the CVS and including one or more PPSs that reference the SPS, or may be provided through external means.
[0184] In one embodiment, the SPS (RBSP) may be available for use in the decoding process before being referenced by one or more PPSs, may be included in at least one PU with a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit, or may be provided through external means.
[0185] In one embodiment, the SPS (RBSP) may be available for use in the decoding process before being referenced by one or more PPSs, may be included in at least one PU with a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit within the CVS and including one or more PPSs that reference the SPS, or may be provided through external means.
[0186] In the same or another embodiment, pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the referenced SPS. The value of pps_seq_parameter_set_id may be the same in all PPSs referenced by the encoded pictures within the CLVS.
[0187] In the same or another embodiment, all SPS NAL units having a particular value of sps_seq_parameter_set_id in the CVS may have the same content.
[0188] In the same or another embodiment, regardless of the nuh_layer_id value, SPS NAL units may share the same value space of sps_seq_parameter_set_id.
[0189] In the same or another embodiment, the nuh_layer_id value of an SPS NAL unit may be equal to the lowest nuh_layer_id value of the PPS NAL units that reference the SPS NAL unit.
[0190] In one embodiment, if an SPS with nuh_layer_id equal to m is referenced by one or more PPSs with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n, or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0191] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced, may be included in at least one AU with a TemporalId equal to the TemporalId of the PPS NAL unit, or may be provided through external means.
[0192] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced, and may be included in at least one AU whose TemporalId is equal to the TemporalId of the PPS NAL unit within the CVS, and which includes one or more PHs (or encoded slice NAL units) that reference the PPS, or may be provided through external means.
[0193] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced by one or more PHs (or encoded slice NAL units), and may be included in at least one PU whose nuh_layer_id is equal to the lowest nuh_layer_id value of the encoded slice NAL units that reference the PPS NAL unit within the CVS, and which includes one or more PHs (or encoded slice NAL units) that reference the PPS, or may be provided through external means.
[0194] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced by one or more PHs (or encoded slice NAL units), and may be included in at least one PU whose TemporalId is equal to the TemporalId of the PPS NAL unit and whose nuh_layer_id is equal to the lowest nuh_layer_id value of the encoded slice NAL units that reference the PPS NAL unit within the CVS, and which includes one or more PHs (or encoded slice NAL units) that reference the PPS, or may be provided through external means.
[0195] In the same or another embodiment, the ph_pic_parameter_set_id within the PH specifies the value of the pps_pic_parameter_set_id for the PPS referenced in use. The value of the pps_seq_parameter_set_id may be the same for all PPSs referenced by the encoded pictures within the CLVS.
[0196] In the same or another embodiment, all PPS NAL units having a specific value of pps_pic_parameter_set_id within the PU can have the same content.
[0197] In the same or another embodiment, regardless of the nuh_layer_id value, PPS NAL units can share the same value space of pps_pic_parameter_set_id.
[0198] In the same or another embodiment, the nuh_layer_id value of a PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the NAL units that reference the PPS NAL unit.
[0199] In one embodiment, if a PPS with nuh_layer_id equal to m is referenced by one or more coded slice NAL units with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n, or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0200] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced, may be included in at least one AU whose TemporalId is equal to the TemporalId of the PPS NAL unit, or may be provided through external means.
[0201] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced, may be included in at least one AU that contains one or more PHs (or coded slice NAL units) that reference the PPS and whose TemporalId is equal to the TemporalId of the PPS NAL unit within the CVS, or may be provided through external means.
[0202] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced by one or more PHs (or encoded slice NAL units), and may be included in at least one PU that includes one or more PHs (or encoded slice NAL units) that reference the PPS, where the nuh_layer_id is equal to the lowest nuh_layer_id value of the encoded slice NAL units that reference the PPS NAL unit within the CVS, or may be provided through external means.
[0203] In one embodiment, the PPS (RBSP) may be available for the decoding process before being referenced by one or more PHs (or encoded slice NAL units), and may be included in at least one PU that includes one or more PHs (or encoded slice NAL units) that reference the PPS, where the TemporalId is equal to the TemporalId of the PPS NAL unit and the nuh_layer_id is equal to the lowest nuh_layer_id value of the encoded slice NAL units that reference the PPS NAL unit within the CVS, or may be provided through external means.
[0204] In the same or another embodiment, the ph_pic_parameter_set_id within the PH specifies the value of the pps_pic_parameter_set_id for the PPS being referenced in use. The value of the pps_seq_parameter_set_id may be the same for all PPSs referenced by the encoded pictures within the CLVS.
[0205] In the same or another embodiment, all PPS NAL units having a particular value of pps_pic_parameter_set_id within the PU may have the same content.
[0206] In the same or another embodiment, regardless of the nuh_layer_id value, the PPS NAL units may share the same value space of pps_pic_parameter_set_id.
[0207] In the same or another embodiment, the nuh_layer_id value of the PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the NAL units that reference the PPS NAL unit.
[0208] In one embodiment, if a PPS with nuh_layer_id equal to m is referenced by one or more coded slice NAL units with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n, or the (direct or indirect) reference layer of the layer with nuh_layer_id equal to m.
[0209] The output layer indicates the layer of the output layer set to be output. The output layer set (OLS) indicates a set of layers composed of a specified set of layers, where one or more layers within the set of layers are specified as output layers. The layer index of the output layer set (OLS) is the index of the layer within the OLS with respect to the list of layers within the OLS.
[0210] A sublayer indicates a temporally scalable layer of a temporally scalable bitstream composed of VCL NAL units having a specific value of the TemporalId variable and associated non-VCL NAL units. The sublayer representation indicates a subset of the bitstream composed of NAL units of a specific sublayer and lower sublayers.
[0211] The VPS RBSP may be used in the decoding process before being referenced, may be included in at least one AU with TemporalId equal to 0, or may be provided through external means. All VPS NAL units having a specific value of vps_video_parameter_set_id within the CVS may have the same content.
[0212] The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id may be greater than 0.
[0213] vps_max_layers_minus1 plus 1 specifies the maximum number of layers allowed in each CVS that refers to the VPS.
[0214] vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in the layers within each CVS that refers to the VPS. The value of vps_max_sublayers_minus1 may be in the range of 0 to 6, inclusive.
[0215] A vps_all_layers_same_num_sublayers_flag equal to 1 specifies that the number of temporal sublayers is the same for all layers within each CVS that refers to the VPS. A vps_all_layers_same_num_sublayers_flag equal to 0 specifies that the layers within each CVS that refers to the VPS may or may not have the same number of temporal sublayers. If it does not exist, the value of vps_all_layers_same_num_sublayers_flag is assumed to be equal to 1.
[0216] A vps_all_independent_layers_flag equal to 1 specifies that all layers within the CVS are encoded independently without using inter-layer prediction. A vps_all_independent_layers_flag equal to 0 specifies that one or more layers within the CVS may use inter-layer prediction. If it does not exist, the value of vps_all_independent_layers_flag is assumed to be equal to 1.
[0217] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, if m is less than n, the value of vps_layer_id[m] may be less than vps_layer_id[n].
[0218] vps_independent_layer_flag[i] equal to 1 specifies that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[i] equal to 0 specifies that the layer with index i can use inter-layer prediction, and the syntax element vps_direct_ref_layer_flag[i][j] for j in the range from 0 to i - 1 including both end values exists within the VPS. If it does not exist, the value of vps_independent_layer_flag[i] is presumed to be equal to 1.
[0219] vps_direct_ref_layer_flag[i][j] equal to 0 specifies that the layer with index j is not a direct reference layer for the layer with index i. vps_direct_ref_layer_flag[i][j] equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. For i and j in the range from 0 to vps_max_layers_minus1 including both end values, if vps_direct_ref_layer_flag[i][j] does not exist, this is presumed to be equal to 0. If vps_independent_layer_flag[i] is equal to 0, there may be at least one value of j in the range from 0 to i - 1 including both end values such that the value of vps_direct_ref_layer_flag[i][j] is equal to 1.
[0220] The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are derived as follows. for(i = 0; i <= vps_max_layers_minus1; i++){ for(j = 0; j <= vps_max_layers_minus1; j++){ dependencyFlag[i][j] = vps_direct_ref_layer_flag[i][j] for(k = 0; k < i; k++) if(vps_direct_ref_layer_flag[i][k] && dependencyFlag[k][j]) dependencyFlag[i][j] = 1 } LayerUsedAsRefLayerFlag[i] = 0 } for(i = 0; i <= vps_max_layers_minus1; i++){ for(j = 0, d = 0, r = 0; j <= vps_max_layers_minus1; j++){ if(vps_direct_ref_layer_flag[i][j]){ DirectRefLayerIdx[i][d++] = j LayerUsedAsRefLayerFlag[j] = 1 } if(dependencyFlag[i][j]) RefLayerIdx[i][r++] = j } NumDirectRefLayers[i] = d NumRefLayers[i] = r }
[0221] The variable GeneralLayerIdx[i] that specifies the layer index of the layer where nuh_layer_id is equal to vps_layer_id[i] is derived as follows. for(i = 0; i <= vps_max_layers_minus1; i++) GeneralLayerIdx[vps_layer_id[i]] = i
[0222] For two different values of i and j in the range from 0 to vps_max_layers_minus1 including both end values, when dependencyFlag[i][j] is equal to 1, it is a requirement for bitstream compliance that the values of chroma_format_idc and bit_depth_minus8 applied to the i-th layer can be equal to the values of chroma_format_idc and bit_depth_minus8 applied to the j-th layer, respectively.
[0223] max_tid_ref_present_flag[i] equal to 1 specifies that the syntax element max_tid_il_ref_pics_plus1[i] exists. max_tid_ref_present_flag[i] equal to 0 specifies that the syntax element max_tid_il_ref_pics_plus1[i] does not exist.
[0224] max_tid_il_ref_pics_plus1[i] equal to 0 specifies that inter-layer prediction is not used by the non-IRAP images of the i-th layer. max_tid_il_ref_pics_plus1[i] greater than 0 specifies that for decoding the image of the i-th layer, images with a TemporalId greater than max_tid_il_ref_pics_plus1[i] - 1 are not used as ILRP. If it does not exist, the value of max_tid_il_ref_pics_plus1[i] is assumed to be equal to 7.
[0225] each_layer_is_an_ols_flag equal to 1 specifies that each OLS contains only one layer, and each layer itself within the CVS referring to the VPS is an OLS with the single contained layer as the only output layer. each_layer_is_an_ols_flag equal to 0 specifies that the OLS can contain two or more layers. When vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is presumed to be 0.
[0226] ols_mode_idc equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i inclusive, and for each OLS, only the topmost layer of the OLS is output.
[0227] ols_mode_idc equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i inclusive, and for each OLS, all layers within the OLS are output.
[0228] ols_mode_idc equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly signaled, and for each OLS, the output layer is explicitly signaled, and the other layers are layers that are direct or indirect reference layers of the output layer of the OLS.
[0229] The value of ols_mode_idc may be in the range from 0 to 2 inclusive. The value 3 of ols_mode_idc is reserved for future use by ITU-T|ISO / IEC.
[0230] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, it is presumed that the value of ols_mode_idc is equal to 2.
[0231] num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0232] The variable TotalNumOlss that specifies the total number of OLSs specified by the VPS is derived as follows. if(vps_max_layers_minus1==0) TotalNumOlss=1 else if(each_layer_is_an_ols_flag||ols_mode_idc==0||ols_mode_idc==1) TotalNumOlss=vps_max_layers_minus1+1 else if(ols_mode_idc==2) TotalNumOlss=num_output_layer_sets_minus1+1
[0233] ols_output_layer_flag[i][j] equal to 1 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS when ols_mode_idc is equal to 2. ols_output_layer_flag[i][j] equal to 0 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS when ols_mode_idc is equal to 2.
[0234] The variable NumOutputLayersInOls[i] that specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] that specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] that specifies whether the k-th layer is used as an output layer in at least one OLS are derived as follows. NumOutputLayersInOls[0]=1 OutputLayerIdInOls[0][0]=vps_layer_id[0] NumSubLayersInLayerInOLS[0][0]=vps_max_sub_layers_minus1+1 LayerUsedAsOutputLayerFlag[0]=1 for(i=1,i<=vps_max_layers_minus1;i++){ if(each_layer_is_an_ols_flag||ols_mode_idc<2) LayerUsedAsOutputLayerFlag[i]=1 else / *(!each_layer_is_an_ols_flag&&ols_mode_idc==2)* / LayerUsedAsOutputLayerFlag[i]=0 } for(i=1;i<TotalNumOlss;i++) if(each_layer_is_an_ols_flag||ols_mode_idc==0){ NumOutputLayersInOls[i]=1 OutputLayerIdInOls[i][0]=vps_layer_id[i] for(j=0;j<i&&(ols_mode_idc==0);j++) NumSubLayersInLayerInOLS[i][j]=max_tid_il_ref_pics_plus1[i] NumSubLayersInLayerInOLS[i][i]=vps_max_sub_layers_minus1+1 } else if(ols_mode_idc == 1){ NumOutputLayersInOls[i]=i+1 for(j = 0; j < NumOutputLayersInOls[i]; j++){ OutputLayerIdInOls[i][j]=vps_layer_id[j] NumSubLayersInLayerInOLS[i][j]=vps_max_sub_layers_minus1+1 } } else if(ols_mode_idc == 2){ for(j = 0; j <= vps_max_layers_minus1; j++){ layerIncludedInOlsFlag[i][j]=0 NumSubLayersInLayerInOLS[i][j]=0 } for(k = 0, j = 0; k <= vps_max_layers_minus1; k++) if(ols_output_layer_flag[i][k]){ layerIncludedInOlsFlag[i][k]=1 LayerUsedAsOutputLayerFlag[k]=1 OutputLayerIdx[i][j]=k OutputLayerIdInOls[i][j++]=vps_layer_id[k] NumSubLayersInLayerInOLS[i][j]=vps_max_sub_layers_minus1+1 } NumOutputLayersInOls[i]=j for (j = 0; j < NumOutputLayersInOls[i]; j++) { idx = OutputLayerIdx[i][j] for (k = 0; k < NumRefLayers[idx]; k++) { layerIncludedInOlsFlag[i][RefLayerIdx[idx][k]] = 1 if (NumSubLayersInLayerInOLS[i][RefLayerIdx[idx][k]] < max_tid_il_ref_pics_plus1[OutputLayerIdInOls[i][j]]) NumSubLayersInLayerInOLS[i][RefLayerIdx[idx][k]] = max_tid_il_ref_pics_plus1[OutputLayerIdInOls[i][j]] } } }
[0235] For each value of i in the range from 0 to vps_max_layers_minus1 inclusive of both end values, the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] may not both be equal to 0. In other words, there may be no layer that is neither a direct reference layer of another layer nor an output layer of at least one OLS.
[0236] For each OLS, there may be at least one layer that is an output layer. In other words, for any value of i in the range from 0 to TotalNumOlss - 1 inclusive of both end values, the value of NumOutputLayersInOls[i] may be 1 or more.
[0237] The variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS are derived as follows. NumLayersInOls[0]=1 LayerIdInOls[0][0]=vps_layer_id[0] for(i=1;i<TotalNumOlss;i++){ if(each_layer_is_an_ols_flag){ NumLayersInOls[i]=1 LayerIdInOls[i][0]=vps_layer_id[i] }else if(ols_mode_idc==0||ols_mode_idc==1){ NumLayersInOls[i]=i+1 for(j=0;j<NumLayersInOls[i];j++) LayerIdInOls[i][j]=vps_layer_id[j] }else if(ols_mode_idc==2){ for(k=0,j=0;k<=vps_max_layers_minus1;k++) if(layerIncludedInOlsFlag[i][k]) LayerIdInOls[i][j++]=vps_layer_id[k] NumLayersInOls[i]=j } }
[0238] The variable OlsLayerIdx[i][j] that specifies the OLS layer index of the layer where nuh_layer_id is equal to LayerIdInOls[i][j] is derived as follows. for(i=0;i<TotalNumOlss;i++) for j=0;j<NumLayersInOls[i];j++) OlsLayerIdx[i][LayerIdInOls[i][j]] = j
[0239] The lowest layer of each OLS may be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss - 1 including both end values, the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] may be equal to 1.
[0240] Each layer may be included in at least one OLS specified by the VPS. In other words, for each layer for which a specific value of nuh_layer_id nuhLayerId is equal to one of vps_layer_id[k] for k in the range from 0 to vps_max_layers_minus1 including both end values, there may be at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1 including both end values and j is in the range from 0 to NumLayersInOls[i] - 1 including both end values, such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.
[0241] In one embodiment, the decoding process operates as follows for the current image CurrPic. PictureOutputFlag is set as follows. If any of the following conditions is true, PictureOutputFlag is set to 0. Otherwise, PictureOutputFlag is set equal to pic_output_flag. - The current image is a RASL image and the NoOutputBeforeRecoveryFlag of the associated IRAP image is 1. - gdr_enabled_flag is equal to 1 and the current image is a GDR image for which NoOutputBeforeRecoveryFlag is equal to 1. - The -gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture where the NoOutputBeforeRecoveryFlag is equal to 1, and the PicOrderCntVal of the current picture is smaller than the RpPicOrderCntVal of the associated GDR picture. - The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions. - PicA has a PictureOutputFlag equal to 1. - PicA has a larger nuh_layer_id nuhLid than the current picture. - PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid). - The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0242] After all slices of the current picture have been decoded, the current decoded picture is marked as "used for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "used for short-term reference".
[0243] In the same or another embodiment, if each layer is an output layer set, the PictureOutputFlag is set equal to pic_output_flag regardless of the value of ols_mode_idc.
[0244] In the same or another embodiment, when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is equal to 0, ols_mode_idc is equal to 0, and the current AU contains an image picA that satisfies all of the following conditions, PictureOutputFlag is set to 0. That is, PicA has a PictureOutputFlag equal to 1, PicA has a nuh_layer_id nuhLid greater than the current image, and PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0245] In the same or another embodiment, when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is equal to 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0, PictureOutputFlag is set equal to 0.
[0246] Resampling of reference images enables adaptive resolution changes within an encoded (layered) video sequence and spatial scalability between layers that have dependencies between layers belonging to the same output layer set.
[0247] In one embodiment, as shown in FIG. 24, the sps_ref_pic_resampling_enabled_flag is signaled in a parameter set (e.g., a sequence parameter set). The sps_ref_pic_resampling_enabled_flag indicates whether resampling of reference pictures is used for adaptive resolution change within an encoded video sequence that references the SPS, or for spatial scalability between layers. A sps_ref_pic_resampling_enabled_flag equal to 1 specifies that resampling of reference pictures is enabled and that one or more picture slices within the CLVS reference reference pictures of different spatial resolutions within the active entries of the reference picture list. A sps_ref_pic_resampling_enabled_flag equal to 0 specifies that resampling of reference pictures is disabled and that no picture slice within the CLVS references a reference picture having a different spatial resolution within the active entries of the reference picture list.
[0248] In the same or another embodiment, when the sps_ref_pic_resampling_enabled_flag is equal to 1, for the current picture, reference pictures having different spatial resolutions belong to either the same layer as the layer containing the current picture or a different layer.
[0249] In another embodiment, a sps_ref_pic_resampling_enabled_flag equal to 1 specifies that resampling of reference pictures is enabled and that one or more slices of pictures within the CLVS reference reference pictures having different spatial resolutions or different scaling windows within the active entries of the reference picture list. A sps_ref_pic_resampling_enabled_flag equal to 0 specifies that resampling of reference pictures is disabled and that no slice of pictures within the CLVS references a reference picture having a different spatial resolution or different scaling window within the active entries of the reference picture list.
[0250] In the same or another embodiment, when sps_ref_pic_resampling_enabled_flag equal to 1, for the current picture, reference pictures with different spatial resolutions or different scaling windows belong to either the same layer as the layer containing the current picture or a different layer from it.
[0251] In the same or another embodiment, sps_res_change_in_clvs_allowed_flag indicates whether the picture resolution is changed in CLVS or CVS. sps_res_change_in_clvs_allowed_flag equal to 1 specifies that the spatial resolution of the picture may change within the CLVS referring to the SPS. sps_res_change_in_clvs_allowed_flag equal to 0 specifies that the spatial resolution of the picture does not change within the CLVS referring to the SPS. If it does not exist, the value of sps_res_change_in_clvs_allowed_flag is presumed to be equal to 0.
[0252] In the same or another embodiment, when sps_ref_pic_resampling_enabled_flag is equal to 1 and sps_res_change_in_clvs_allowed_flag is equal to 0, resampling of the reference picture can be used only for spatial scalability rather than for the change of adaptive resolution within the CLVS.
[0253] In the same or another embodiment, when sps_ref_pic_resampling_enabled_flag is equal to 1 and sps_res_change_in_clvs_allowed_flag is equal to 1, resampling of the reference picture can be used for both spatial scalability and adaptive resolution change within the CLVS.
[0254] When sps_ref_pic_resampling_enabled_flag is equal to 1, sps_res_change_in_clvs_allowed_flag is equal to 0, and sps_video_parameter_set_id is equal to 0, pps_scaling_window_explicit_signalling_flag may be equal to 1. This means that when the resolution of the image is constant in CLVS or CVS and resampling of the reference image is used, the scaling window parameters need to be signaled explicitly instead of inferring values from the adaptive window parameters.
[0255] In one embodiment, sps_virtual_boundaries_present_flag is signaled in the SPS as shown in FIG. 24. The flag sps_virtual_boundaries_present_flag indicates whether virtual boundary information is signaled in the SPS.
[0256] In the same or another embodiment, sps_virtual_boundaries_present_flag is conditionally signaled only when sps_res_change_in_clvs_allowed_flag is equal to 0, because when resampling of the reference image is used, virtual boundary information may not be signaled in the SPS.
[0257] In the same embodiment, sps_virtual_boundaries_present_flag equal to 1 specifies that the information of virtual boundaries is signaled in SPS. sps_virtual_boundaries_present_flag equal to 0 specifies that the information of virtual boundaries is not signaled in SPS. If there are one or more virtual boundaries signaled in SPS, the loop filtering operations across the virtual boundaries of the picture referred to by SPS are disabled. The loop filtering operations include deblocking filtering, sample adaptive offset filtering, and adaptive loop filtering operations. If not present, the value of sps_virtual_boundaries_present_flag is assumed to be equal to 0.
[0258] In one embodiment, sps_subpic_info_present_flag is signaled in SPS as shown in FIG. 24. The flag sps_subpic_info_present_flag indicates whether the sub-picture segmentation information is signaled in SPS.
[0259] In the same or another embodiment, sps_subpic_info_present_flag is conditionally signaled only when sps_res_change_in_clvs_allowed_flag is equal to 0, because when resampling of the reference picture is used, the segmentation information of the sub-picture may not be signaled in SPS.
[0260] In the same embodiment, the sps_subpic_info_present_flag equal to 1 specifies that sub-picture information exists for CLVS, and there may be one or more sub-pictures in each picture of CLVS. The sps_subpic_info_present_flag equal to 0 specifies that sub-picture information does not exist for CLVS, and there is only one sub-picture in each picture of CLVS. If it does not exist, the value of the sps_subpic_info_present_flag is presumed to be equal to 0.
[0261] In one embodiment, as shown in FIG. 25, the pps_res_change_in_clvs_allowed_flag may be signaled in the PPS. The value of the pps_res_change_in_clvs_allowed_flag in the PPS may be the same as the value of the sps_res_change_in_clvs_allowed_flag in the SPS referred to by the PPS.
[0262] In the same embodiment, the information on the width and height of the picture may be signaled in the PPS only when the value of the pps_res_change_in_clvs_allowed_flag is equal to 1. When the pps_res_change_in_clvs_allowed_flag is equal to 0, the values of the width and height of the picture are presumed to be equal to the maximum width and height values of the picture signaled in the SPS.
[0263] In the same embodiment, pps_pic_width_in_luma_samples specifies the width of each decoded picture that refers to the PPS in units of luma samples. pps_pic_width_in_luma_samples may not be equal to 0, may be an integer multiple of Max(8, MinCbSizeY), and may be less than or equal to sps_pic_width_max_in_luma_samples. If it does not exist, the value of pps_pic_width_in_luma_samples is presumed to be equal to sps_pic_width_max_in_luma_samples. When sps_ref_wraparound_enabled_flag is equal to 1, the value of (CtbSizeY / MinCbSizeY + 1) may be less than or equal to the value of (pps_pic_width_in_luma_samples / MinCbSizeY - 1). pps_pic_height_in_luma_samples specifies the height of each decoded picture that refers to the PPS in units of luma samples. pps_pic_height_in_luma_samples may not be equal to 0, may be an integer multiple of Max(8, MinCbSizeY), and may be less than or equal to sps_pic_height_max_in_luma_samples. If it does not exist, the value of pps_pic_height_in_luma_samples is presumed to be equal to sps_pic_height_max_in_luma_samples.
[0264] In the reference picture list, all active reference pictures for a picture have the same sub-picture layout as the picture itself, and all active reference pictures are inter-layer reference pictures having a single sub-picture.
[0265] In the same or another embodiment, the images referred to by each active entry in RefPicList[0] or RefPicList[1] have the same image size and the same sub - image layout as the current image (i.e., the SPSs referred to by those images and the current image have, for each value of j in the range from 0 to sps_num_subpics_minus1 inclusive, the same value of sps_num_subpics_minus1, and the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], and sps_subpic_height_minus1[j]). The images referred to by each active entry in RefPicList[0] or RefPicList[1] are ILRPs where the value of sps_num_subpics_minus1 is equal to 0.
[0266] In the same or another embodiment, when sps_num_subpics_minus1 is greater than 0 and sps_subpic_treated_as_pic_flag[i] is equal to 1, for each CLVS of the current layer that refers to an SPS, if targetAuSet is set to all AUs from the AU containing the first image of the CLVS in decoding order to the AU containing the last image of the CLVS in decoding order, inclusive of both end values, then for targetLayerSet, which is composed of the current layer and all layers that have the current layer as a reference layer, the following conditions must all be true for bit - stream conformance requirements. - For each AU in targetAuSet, all images of the layers in targetLayerSet can have the same value of pps_pic_width_in_luma_samples and the same value of pps_pic_height_in_luma_samples. - In the -targetLayerSet, all SPSs referred to by layers having the current layer as the reference layer can have the same value of sps_num_subpics_minus1, and for each value of j in the range from 0 to sps_num_subpics_minus1 including both end values, can have the same values of sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], and sps_subpic_treated_as_pic_flag[j]. - For each AU in the -targetAuSet, all images of the layers having the current layer in the -targetLayerSet as the reference layer can have the same value of SubpicIdVal[j] for each value of j in the range from 0 to sps_num_subpics_minus1 including both end values.
[0267] In the same or another embodiment, pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset specify the offsets applied to the image size for scaling ratio calculation. If they do not exist, the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset are presumed to be equal to pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset, respectively.
[0268] The value of SubWidthC * (Abs(pps_scaling_win_left_offset) + Abs(pps_scaling_win_right_offset)) may be less than pps_pic_width_in_luma_samples, and the value of SubHeightC * (Abs(pps_scaling_win_top_offset) + Abs(pps_scaling_win_bottom_offset)) may be less than pps_pic_height_in_luma_samples.
[0269] The variables CurrPicScalWinWidthL and CurrPicScalWinHeightL are derived as follows. CurrPicScalWinWidthL = pps_pic_width_in_luma_samples - SubWidthC * (pps_scaling_win_right_offset + pps_scaling_win_left_offset) CurrPicScalWinHeightL = pps_pic_height_in_luma_samples - SubHeightC * (pps_scaling_win_bottom_offset + pps_scaling_win_top_offset)
[0270] Let refPicScalWinWidthL and refPicScalWinHeightL be the CurrPicScalWinWidthL and CurrPicScalWinHeightL of the reference picture of the current picture referring to this PPS, respectively. It is a bitstream compliance requirement that all of the following conditions are met. - currPicScalWinWidthL * 2 is greater than or equal to refPicScalWinWidthL. - currPicScalWinHeightL * 2 is greater than or equal to refPicScalWinHeightL. - currPicScalWinWidthL is less than or equal to refPicScalWinWidthL * 8. - currPicScalWinHeightL is less than or equal to refPicScalWinHeightL * 8. - currPicScalWinWidthL * sps_pic_width_max_in_luma_samples is greater than or equal to refPicScalWinWidthL * (pps_pic_width_in_luma_samples - Max(8, MinCbSizeY)). - currPicScalWinHeightL * sps_pic_height_max_in_luma_samples is greater than or equal to refPicScalWinHeightL * (pps_pic_height_in_luma_samples - Max(8, MinCbSizeY)).
[0271] The value of SubWidthC * (Abs(pps_scaling_win_left_offset) + Abs(pps_scaling_win_right_offset)) may be less than pps_pic_width_in_luma_samples, and the value of SubHeightC * (Abs(pps_scaling_win_top_offset) + Abs(pps_scaling_win_bottom_offset)) may be less than pps_pic_height_in_luma_samples.
[0272] In the same or another embodiment, the value of SubWidthC * (pps_scaling_win_left_offset + pps_scaling_win_right_offset) may be greater than or equal to -pps_pic_width_in_luma_samples * 15 and less than pps_pic_width_in_luma_samples, and the value of SubHeightC * (pps_scaling_win_top_offset + pps_scaling_win_bottom_offset) may be greater than or equal to -pps_pic_height_in_luma_samples * 15 and less than pps_pic_height_in_luma_samples.
[0273] In the same or another embodiment, the value of SubWidthC * (pps_scaling_win_left_offset + pps_scaling_win_right_offset) may be greater than or equal to -pps_pic_width_in_luma_samples * 7 and less than pps_pic_width_in_luma_samples, and the value of -SubHeightC * (pps_scaling_win_top_offset + pps_scaling_win_bottom_offset) may be greater than or equal to -pps_pic_height_in_luma_samples * 7 and less than pps_pic_height_in_luma_samples.
[0274] In the same or another embodiment, when sps_ref_pic_resampling_enabled_flag is equal to 1, sps_res_change_in_clvs_allowed_flag is equal to 0, and sps_subpic_info_present_flag is equal to 1, the value of SubWidthC*(Abs(pps_scaling_win_left_offset)+Abs(pps_scaling_win_right_offset)) may be smaller than the minimum value of sps_subpic_width_minus1[i]+1 for i in the range from 0 to sps_num_subpics_minus1, and the value of SubHeightC*(Abs(pps_scaling_win_top_offset)+Abs(pps_scaling_win_bottom_offset)) may be smaller than the minimum value of sps_subpic_height_minus1[i]+1 for i in the range from 0 to sps_num_subpics_minus1.
[0275] In the same or another embodiment, when the value of sps_res_change_in_clvs_allowed_flag of the current layer is equal to 1, the value of sps_subpic_info_present_flag may be equal to 0.
[0276] In the same or another embodiment, when the value of sps_res_change_in_clvs_allowed_flag of the current layer is equal to 1, the value of sps_subpic_info_present_flag of the layer referring to the current layer may be equal to 0.
[0277] In the same or another embodiment, when the value of sps_res_change_in_clvs_allowed_flag of the current layer is equal to 1, the values of sps_subpic_info_present_flag of the current layer and all the layers referring to the current layer may be equal to 0.
[0278] In the same or another embodiment, when the value of sps_subpic_info_present_flag of the current layer may be equal to 1, the value of sps_res_change_in_clvs_allowed_flag of the current layer may be equal to 0.
[0279] In the same or another embodiment, when the value of sps_subpic_info_present_flag of the current layer may be equal to 1, the value of sps_res_change_in_clvs_allowed_flag of the reference layer of the current layer may be equal to 0.
[0280] In the same or another embodiment, when the value of sps_subpic_info_present_flag of the current layer may be equal to 1, the values of sps_res_change_in_clvs_allowed_flag of the current layer and all reference layers of the current layer may be equal to 0.
[0281] In the same or another embodiment, when the value of sps_subpic_info_present_flag of the current layer may be equal to 1 and the value of sps_ref_pic_resampling_enabled_flag may be equal to 1, the value of sps_ref_pic_resampling_enabled_flag of the reference layer of the current layer may be equal to 1.
[0282] In one embodiment, the encoded image in layer k may be divided into one or more sub-images as shown in FIG. 26, and one or more reference images in the same layer can be referred to, and one or more reference images in the reference layer of layer k can be referred to. In the example of FIG. 26, when the value of sps_ref_pic_resampling_enabled_flag (of FIG. 24) is equal to 1, even if the image sizes are the same, the current image and each reference image can have different scaling windows.
[0283] In the same or another embodiment, when a sub-image is extracted, the size of the scaling window used for calculating the scaling ratio for reference image resampling and its offset value may be updated according to the size and position of the sub-image. When the size of the scaling window of the current image and its offset value are updated, the size of the scaling window of one or more reference images of the current image and its offset value may be updated accordingly. FIG. 27 shows an example of the update of the scaling window of the reference image within the same layer and the inter-layer reference image between different layers.
[0284] In the same or another embodiment, when an image within layer k is divided into one or more sub-images with the same split layout, the image in another layer that layer k refers to as a reference layer may not be divided into a plurality of sub-images.
[0285] In the same or another embodiment, when the size and offset value of the scaling window are updated as in the example of FIG. 27, the size of the scaling window may be rescaled in relation to the scaling ratio between the original image size and the extracted sub-image size. When the size of the scaling window is updated, depending on the original image size and the sub-image size, the updated size of the scaling window may have fractional pixel values that cannot be represented by the scaling window offset values (pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, pps_scaling_win_bottom_offset) signaled in the PPS as shown in FIG. 25. Also, in this method, all the scaling offset values of the reference image may be updated. This is a fairly large burden.
[0286] In the same or another embodiment, as shown in FIG. 28, when a sub-image is extracted, in order to signal the same scaling ratio between the current image and the reference image, the size of the scaling window may not be changed compared to the original scaling window before extraction, but only the position of the scaling window is shifted by updating the values of the offset values (pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, pps_scaling_win_bottom_offset) of the scaling window signaled in the PPS.
[0287] In the same embodiment, if the offset values of the scaling window of the original image are orgScalingWinLeft, orgScalingWinRight, orgScalingWinTop, and orgScalingWinBottom, which are equal to the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the original image respectively, the position and size of the extracted sub-image are represented by SubpicLeftBoundaryPos, SubpicRightBoundaryPos, SubpicTopBoundaryPos, SubpicBotBoundaryPos, where these values are derived as follows. SubpicLeftBoundaryPos = sps_subpic_ctu_top_left_x[CurrSubpicIdx] * CtbSizeY SubpicRightBoundaryPos = Min(sps_pic_width_max_in_luma_samples - 1, (sps_subpic_ctu_top_left_x[CurrSubpicIdx] + ((sps_subpic_width_minus1[CurrSubpicIdx]+1)*CtbSizeY-1) SubpicTopBoundaryPos = sps_subpic_ctu_top_left_y[CurrSubpicIdx]*CtbSizeY SubpicBotBoundaryPos = Min(sps_pic_height_max_in_luma_samples - 1 (sps_subpic_ctu_top_left_y[CurrSubpicIdx]+ (sps_subpic_height_minus1[CurrSubpicIdx]+1)*CtbSizeY-1)
[0288] The values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the extracted sub - picture are derived as follows. pps_scaling_win_left_offset = orgScalingWinLeft-(SubpicLeftBoundaryPos / SubWidthC); pps_scaling_win_right_offset = orgScalingWinRight-(SubpicLeftBoundaryPos / SubWidthC); pps_scaling_win_top_offset = orgScalingWinTop-(SubpicTopBoundaryPos / SubWidthC); pps_scaling_win_bottom_offset = orgScalingWinBottom-(SubpicTopBoundaryPos / SubWidthC).
[0289] In the same or another embodiment, the sub-image sub-bitstream extraction process is as follows. The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, the target maximum TemporalId value tIdTarget, and an array subpicIdxTarget[] of target sub-image index values for each layer. The output of this process is the sub-bitstream outBitstream.
[0290] The requirements for bitstream compliance with respect to the input bitstream are that the output sub-bitstream that satisfies all of the following conditions is a compliant bitstream. The output sub-bitstream is the output of the process specified in this section that uses as inputs the bitstream, the targetOlsIdx equal to the index into the list of OLSs specified by the VPS, and the subpicIdxTarget[] equal to the sub-image indices present in the OLS. The output sub-bitstream includes at least one VCL NAL unit for which the nuh_layer_id is equal to each of the nuh_layer_id values of LayerIdInOls[targetOlsIdx]. The output sub-bitstream includes at least one VCL NAL unit for which the TemporalId is equal to tIdTarget. The compliant bitstream includes one or more coded slice NAL units for which the TemporalId is equal to 0, but need not include a coded slice NAL unit for which the nuh_layer_id is equal to 0. The output sub-bitstream includes at least one VCL NAL unit for which the nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and the sh_subpic_id is equal to the value of SubpicIdVal[subpicIdxTarget[i]] for each i in the range from 0 to NumLayersInOls[targetOlsIdx] - 1 inclusive.
[0291] The output sub-bitstream outBitstream is derived as follows. The sub-bitstream extraction process is called with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the output of the process is assigned to outBitstream. If any external means not specified in this document are available to provide a replacement parameter set for the sub-bitstream outBitstream, all parameter sets are replaced with the replacement parameter set.
[0292] Instead, if the SEI message of sub - picture level information exists in inBitstream, the following applies. The variable subpicIdx is set equal to the value of subpicIdxTarget[[NumLayersInOls[targetOlsIdx]-1]]. Rewrite the value of general_level_idc of the vps_ols_ptl_idx[targetOlsIdx] - th entry in the list of profile_tier_level() syntax structures of all referenced VPS NAL units to be equal to SubpicLevelIdc for the set of sub - pictures composed of sub - pictures having a sub - picture index equal to subpicIdx. If there are VCL HRD parameters or NAL HRD parameters, in the ols_hrd_parameters() syntax structure of the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] - th in all referenced VPS NAL units and in the ols_hrd_parameters() syntax structure of all SPS NAL units referenced by the i - th layer, rewrite the respective values of cpb_size_value_minus1[tIdTarget][j] and bit_rate_value_minus1[tIdTarget][j] of the j - th CPB so that they correspond to SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx] and SubpicCpbSizeNal[SubpicSetLevelIdx][subpicIdx].
[0293] SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx] and SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx] for the sub - picture with the sub - picture index equal to subpicIdx, j is in the range from 0 to hrd_cpb_cnt_minus1 inclusive, and i is in the range from 0 to NumLayersInOls[targetOlsIdx]-1 inclusive.
[0294] For the i-th layer where i ranges from 0 to NumLayersInOls[targetOlsIdx] - 1, the following applies. The variable subpicIdx is set equal to the value of subpicIdxTarget[i]. Rewrite the value of general_level_idc in the profile_tier_level() syntax structure of all reference SPS NAL units where sps_ptl_dpb_hrd_params_present_flag is equal to 1 to be equal to SubpicLevelIdc for the set of subpictures composed of subpictures having a subpicture index equal to subpicIdx.
[0295] The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are derived as follows. subpicWidthInLumaSamples = min((sps_subpic_ctu_top_left_x[subpicIdx] + sps_subpic_width_minus1[subpicIdx] + 1) * CtbSizeY, pps_pic_width_in_luma_samples) - sps_subpic_ctu_top_left_x[subpicIdx] * CtbSizeY subpicHeightInLumaSamples = min((sps_subpic_ctu_top_left_y[subpicIdx] + sps_subpic_height_minus1[subpicIdx] + 1) * CtbSizeY, pps_pic_height_in_luma_samples) - sps_subpic_ctu_top_left_y[subpicIdx] * CtbSizeY
[0296] Rewrite the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples of all referenced SPS NAL units, and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples of all referenced PPS NAL units to be equal to subpicWidthInLumaSample and subpicHeightInLumaSamples, respectively. Rewrite the values of sps_num_subpics_minus1 of all referenced SPS NAL units and the values of pps_num_subpics_minus1 of all referenced PPS NAL units to 0. In all referenced SPS NAL units, if the syntax elements sps_subpic_ctu_top_left_x[subpicIdx] and sps_subpic_ctu_top_left_y[subpicIdx] exist, rewrite them to 0. In all referenced SPS NAL units, for each j not equal to subpicIdx, delete the syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j], and sps_subpic_id[j]. Rewrite the syntax elements of all referenced PPSs for tile and slice signaling, and delete all tile rows, tile columns, and slices not associated with the subpicture whose subpicture index is equal to subpicIdx.
[0297] The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are derived as follows. subpicConfWinLeftOffset = sps_subpic_ctu_top_left_x[subpicIdx] == 0 sps_conf_win_left_offset: 0 subpicConfWinRightOffset = (sps_subpic_ctu_top_left_x[subpicIdx] + sps_subpic_width_minus1[subpicIdx] + 1) * CtbSizeY >= sps_pic_width_max_in_luma_samples sps_conf_win_right_offset: 0 subpicConfWinTopOffset = sps_subpic_ctu_top_left_y[subpicIdx] == 0 sps_conf_win_top_offset: 0 subpicConfWinBottomOffset = (sps_subpic_ctu_top_left_y[subpicIdx] + sps_subpic_height_minus1[subpicIdx] + 1) * CtbSizeY >= sps_pic_height_max_in_luma_samples sps_conf_win_bottom_offset: 0
[0298] Rewrite the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset of all referenced SPS NAL units, and the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset of all referenced PPS NAL units to be equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively. Rewrite the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset as follows. pps_scaling_win_left_offset = orgScalingWinLeft - (SubpicLeftBoundaryPos / SubWidthC) pps_scaling_win_right_offset = orgScalingWinRight - (SubpicLeftBoundaryPos / SubWidthC) pps_scaling_win_top_offset = orgScalingWinTop - (SubpicTopBoundaryPos / SubWidthC) pps_scaling_win_bottom_offset = orgScalingWinBottom - (SubpicTopBoundaryPos / SubWidthC), Here, orgScalingWinLeft, orgScalingWinRight, orgScalingWinTop, and orgScalingWinBottom are equal to the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the original encoded picture. Delete all VCL NAL units from outBitstream where nuh_layer_id is equal to the nuh_layer_id of the i-th layer and sh_subpic_id is not equal to SubpicIdVal[subpicIdx].
[0299] If sli_cbr_constraint_flag is equal to 1, delete all NAL units where nal_unit_type is equal to FD_NUT and all filler payload SEI messages not associated with the VCL NAL units of the sub-pictures within subpicIdTarget[]. In the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]]-th ols_hrd_parameters() syntax structure of all referenced VPS NAL units and SPS NAL units, set cbr_flag[tIdTarget][j] to 1 for the j-th CPB, where j ranges from 0 to hrd_cpb_cnt_minus1. Otherwise (when sli_cbr_constraint_flag is equal to 0), delete all NAL units where nal_unit_type is equal to FD_NUT and all filler payload SEI messages, and set cbr_flag[tIdTarget][j] to 0.
[0300] If the outBitstream contains an SEI NAL unit that contains a scalable nested SEI message applicable to the outBitstream, where the sn_ols_flag is equal to 1 and the sn_subpic_flag is equal to 1, extract an appropriate non-scalable nested SEI message where the payloadType is equal to 1 (PT), 130 (DUI), or 132 (decoded image hash) from the scalable nested SEI message, and place the extracted SEI message in the outBitstream.
[0301] Some embodiments may relate to a system, method, and / or computer-readable medium in any possible technical detail integration. The computer-readable medium can include a computer-readable non-transitory storage medium having computer-readable program instructions for causing a processor to execute operations.
[0302] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction-executing device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed as being a signal per se that is transient, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0303] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each respective computing / processing device.
[0304] The computer-readable program code / instructions for performing the operations may be in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in an object-oriented programming language such as Smalltalk(R), C++, and a procedural programming language such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit for performing the aspects or operations.
[0305] These computer-readable program instructions are provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium contains instructions for implementing the aspects of the functions / acts specified in the blocks of the flowchart and / or block diagram.
[0306] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to produce a series of operational steps to be executed on the computer, other programmable apparatus, or other device such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the blocks of the flowchart and / or block diagram, thereby producing a computer-implemented process.
[0307] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function. The present methods, computer systems, and computer-readable media can include additional blocks, fewer blocks, different blocks, or blocks arranged in a different configuration than those depicted in the figures. In some alternative embodiments, the functions represented by the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may actually be executed simultaneously, or substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks of the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified function or action, or by a combination of dedicated hardware and computer instructions.
[0308] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not limiting. Thus, the operation and behavior of the systems and / or methods have been described herein without reference to specific software code, but it should be understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0309] No element, act, or instruction used in this book should be construed as important or essential unless explicitly described. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar language is used. Also, as used herein, terms such as "has," "have," "having," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified.
[0310] The descriptions of various aspects and embodiments are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Even if a combination of functions is recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible embodiments includes each dependent claim combined with all other claims in the claim set. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application, or a technical improvement to technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
Description of Reference Numerals
[0311] 100 Communication system 110 Terminal 120 Terminal 130 Terminal 140 Terminal 150 Communication network 201 Video source, camera 202 Data stream 203 Encoder 204 Video bitstream 205 Streaming server 206 Streaming client 207 Copy of video bitstream 208 Streaming client 209 Copy of video bitstream 210 Decoder 211 Video sample stream 212 Display, rendering device 213 Capture subsystem 310 Receiver 312 Channel 315 Buffer memory 320 Parser, video decoder 321 Symbol 351 Scaler / inverse transform unit 352 Inter-picture prediction unit, intra prediction unit 353 Motion compensation prediction unit 355 Aggregator 356 Current picture or loop filter unit 357 Reference picture memory, reference picture buffer 430 Encoder 432 Encoding engine 433 Local decoder 434 Reference picture memory 435 Predictor 440 Transmitter 443 Video sequence 445 Entropy coder 450 Controller 460 Communication Channel 501 Image Header 502 ARC Information 504 Image Parameter Set 505 ARC Reference Information 506 Table 507 Sequence Parameter Set 508 Tile Group Header 509 ARC Information 511 Parameter Set 512 ARC Information 513 ARC Reference Information, Table 514 Tile Group Header 515 ARC Information 516 ARC Information Table, Set 601 Tile Group Header 602 Syntax Element 603 Syntax Element 610 Sequence Parameter Set 611 Syntax Element 612 Parameter Set 613 Syntax Element 614 Syntax Element 615 Reference Image Dimension 616 Syntax Element 617 Syntax Element 700 Computer System 701 Keyboard 702 Mouse 703 Track Pad 704 Data Glove 705 Joystick 706 Microphone 707 Scanner 708 Camera 709 Speaker 710 Screen 720 CD / DVD ROM / RW 721 Medium 722 Thumb Drive 723 Solid State Drive 740 Core 743 FPGA 741 CPU 742 GPU 743 FPGA 744 Accelerator 745 ROM 746 RAM 747 Internal Mass Storage Device 748 System Bus 749 Peripheral Bus
Claims
1. A method for video decoding executed by a processor, comprising: receiving video data having one or more sub-images; generating a bitstream including a sub-bitstream configured to be extracted based on an array of a target output layer set index, a target maximum time identification value, and a target sub-image index value, the sub-bitstream being associated with resampling parameters and spatial scalability parameters corresponding to the sub-image; In the decoding of the video data, the left and right offset values of the scaling window of the extracted sub-image are updated based on the left boundary position of the extracted sub-image, the upper and lower offset values of the scaling window are updated based on the upper boundary position of the extracted sub-image, and the resampling parameters and the spatial scalability parameters are configured to shift the position of the scaling window based on the updated offset values for use in scaling the video data.
2. The method according to claim 1, wherein the left and right offset values of the scaling window and the upper and lower offset values of the scaling window are updated according to the size and position of the sub-image.
3. The method according to claim 1, wherein the left and right offset values of the scaling window are updated based on the width of the sub-image and the left boundary position of the sub-image, and the upper and lower offset values of the scaling window are updated based on the width of the sub-image and the upper boundary position of the sub-image.
4. The method according to claim 1, enabling adaptive resolution change of the received video data based on the resampling parameters.
5. The method according to claim 1, wherein the resampling parameters correspond to one or more flags signaled in a parameter set associated with the video data.
6. The method according to claim 1, wherein the spatial scalability parameters correspond to one or more flags signaled in a parameter set associated with the video data.
7. The method according to claim 1, wherein resampling of the video data being decoded is disabled based on the resampling parameter. **Claim 8** A computer system for video decoding, comprising: One or more computer-readable non-transitory storage media configured to store computer program code; and One or more computer processors configured to access the computer program code and operate as instructed by the computer program code, wherein the computer program code is configured to cause the one or more computer processors to execute the method according to any one of claims 1 to 7. **Claim 9** A computer program for video decoding, wherein the computer program is configured to cause one or more computer processors to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Reference Picture Resampling
JP2023518440A
Manipulating coded video in the sub-bitstream extraction process
JP2023526373A