Method for setting output levels in multi-level video stream
By signaling adaptive picture sizes in a video bitstream using a VPS, the method addresses the challenge of varying picture sizes in video coding, enhancing compression efficiency and display flexibility.
Patent Information
- Authority / Receiving Office
- RU · RU
- Patent Type
- Patents
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2020-11-09
- Publication Date
- 2026-07-09
AI Technical Summary
Existing video coding technologies struggle with efficiently adapting to varying picture sizes and resolutions within a coded video sequence, limiting the flexibility and efficiency of video compression.
The method involves signaling an adaptive picture size in a video bitstream by using a video parameter set (VPS) to determine output levels, allowing for flexible image output levels based on mode indicators and explicit signaling, and controlling display of decoded image layers.
Enables efficient compression and display of video data with adaptive resolution changes, reducing bandwidth and storage requirements while maintaining video quality.
Smart Images

Figure 00000009_ABST
Abstract
Description
Cross-references to related applications
[0001] This application is separated from Russian Federation Patent Application No. 2021131317, filed October 26, 2021, which claims priority from US Provisional Patent Application No. 63 / 001045, filed March 27, 2020, and US Patent Application No. 17 / 000018, filed August 21, 2020, which applications are fully incorporated herein by reference.Background of the InventionField of Technology
[0002] The present invention relates to video compression technologies, as well as to inter- and intra-prediction in advanced video codecs. Specifically, the present invention relates to next-generation video coding / decoding technologies replacing High Efficiency Video Coding (HEVC), such as Versatile Video Coding (VVC). In particular, one aspect of the present invention relates to a method, device, and computer-readable medium providing a set of advanced video coding technologies for determining output layers in an encoded video stream with multiple layers. Description of the Prior Art
[0003] Video coding and decoding techniques using inter or intra image prediction with motion compensation have existed for decades. Uncompressed digital video may contain a sequence of images, each of which has a given spatial dimension, for example, 1920x1080 luminance samples and corresponding chrominance samples. Such an image sequence may have a fixed or variable image rate (informally also called a frame rate), equal to, for example, 60 images per second, or 60 Hz. Uncompressed video places high demands on the bit rate. For example, 1080p60 4:2:0 video with 8-bit sample depth (a resolution of 1920x1080 luminance samples at a frame rate of 60 Hz) requires a bandwidth of approximately 1.5 Gbps. An hour of such video can require over 600 GB of storage.
[0004] One of the goals of video coding and decoding is to reduce redundancy in the input video signal through compression. Compression can reduce bandwidth or storage requirements by up to two orders of magnitude or more in some cases. Both lossless and lossy compression, as well as combinations of both, can be used. Lossless compression refers to methods that allow an exact copy of the original signal to be reconstructed from the compressed signal. When lossy compression is used, the reconstructed signal may not be identical to the original, but the discrepancy between the original and reconstructed signals is small enough that the reconstructed signal is suitable for the intended use. Lossy compression is widely used for video. The amount of acceptable distortion depends on the specific application; for example, users of commercial streaming applications may be more tolerant of distortion than users of broadcast applications.The degree of compression is subject to the following rule: the greater the permissible distortion, the greater the achievable degree of compression.
[0005] A video encoder and video decoder may apply a number of techniques under different categories, including, for example, motion compensation, transforms, quantization, and entropy coding, some of which will be discussed below.
[0006] Historically, video encoders and decoders have worked with a given picture size, which in most cases was known and remained constant throughout a coded video sequence (CVS), Group of Pictures (GOP), or similar time frames containing multiple pictures. For example, the Motion Picture Experts Group (MPEG-2) standard is designed to allow the horizontal resolution (and hence the picture size) to vary depending on certain factors, such as activity in a video scene, but only within intra-predicted pictures (I-frames or I-pictures), and therefore typically only for a GOP. ITU-T Rec. H.263 Annex P specifies resampling of reference pictures to allow for different resolutions within a CVS.However, in this case, only the size of the reference images changed, while the size of the images themselves remained unchanged. This resulted in the ability to use only fragments of the entire image surface (downsampling) or capture only part of the scene (upsampling). H.263 Annex Q also allows for upsampling and downsampling of individual macroblocks by a power of two (along any axis). However, again, the image size remains unchanged. The macroblock size in H.263 is fixed and therefore does not need to be signaled.
[0007] In modern video coding, changing the size of predicted pictures is more common. For example, the VP9 standard allows resampling of reference pictures and changing the resolution of a picture as a whole. Similarly, corresponding proposals have been made for the Universal Video Coding (VVC) standard (including, for example, the Joint Video Experts Group document JVET-M0135-v1, “On Adaptive Resolution Change (ARC) for VVC”, January 9-19, 2019, Hendry et al., incorporated herein by reference in its entirety), allowing resampling of reference pictures as a whole, to different, higher or lower, resolutions. In the mentioned document (Hendry et al.), it is proposed to encode different candidate resolutions in a sequence parameter set and to reference them in syntax elements, for each picture, in the picture parameter set. Summary of the invention
[0008] According to various embodiments of the present invention, methods are provided for signaling an adaptive picture size in a video bitstream.
[0009] According to one aspect of the present invention, a decoding method may include: receiving a bit stream including compressed video / image data, wherein said bit stream has a plurality of levels; obtaining, by analyzing or deriving, from the bit stream, an indicator of a mode of a set of output levels in a video parameter set (VPS); identifying a signaling of a set of output levels based on the mode indicator of a set of output levels; identifying one or more image output levels based on the identified signaling of a set of output levels; and decoding one or more identified image output levels.
[0010] Identifying the signaling of a set of output levels based on the output level set mode indicator may include: if the output level set mode indicator in the VPS has a first value, determining that the uppermost level in the bit stream is said one or more image output levels; if the output level set mode indicator in the VPS has a second value, determining that all levels in the bit stream are said one or more image output levels; and if the output level set mode indicator in the VPS has a third value, identifying one or more image output levels based on explicit signaling in the VPS.
[0011] The first value may be different from the second value and may be different from the third value, and the second value may be different from the third value.
[0012] The first value may be 0, the second value may be 1, and the third value may be 2. However, other values may be used, without limiting the present invention to the use of the values 0, 1, and 2 described above.
[0013] Identifying one or more image output levels by explicit signaling in a VPS may include: (i) obtaining from the VPS, by analysis or deduction, an output level flag, and (ii) assigning levels whose output level flag is equal to 1 to said one or more image output levels.
[0014] The output level set signaling identification based on the output level set mode indicator may include the following: in the case where the output level set mode indicator in the VPS has a preset value, the output level set signaling includes the identification of one or more image output levels based on the explicit signaling in the VPS.
[0015] Identifying one or more image output levels by explicit signaling in a VPS includes: (i) obtaining from the VPS, by analyzing or deducing, an output level flag, and (ii) assigning levels whose output level flag is equal to 1 to said one or more image output levels, wherein the number of levels in said plurality of levels is greater than 2.
[0016] The output level set mode signaling may include identifying one or more image output levels based on explicit signaling in the VPS when the output level set mode indicator is 2 and when the number of levels in said plurality of levels is greater than 2.
[0017] The output level set signaling includes determining that the highest level in the bit stream or all levels in the bit stream are said one or more image output levels, by logically inferring one or more image output levels when the output level set mode indicator is less than 2, and the number of levels in said plurality of levels is 2, and the output level set mode indicator is actually less than 2, and the number of levels in said plurality of levels is actually 2.
[0018] According to one embodiment of the present invention, the indicator of the number of sets of output levels minus 1 in the VPS indicates the number of output levels.
[0019] According to one embodiment of the present invention, the maximum level indicator in the VPS minus 1 in the VPS indicates the number of levels in the bit stream.
[0020] According to one embodiment of the present invention, the output level set flag [i][j] in the VPS indicates whether the j-th level of the i-th set of output levels is an output level or not.
[0021] According to one embodiment of the present invention, if all of said plurality of levels are independent, and the flag "all levels in the VPS are independent" in the VPS is equal to 1, the output level set mode indicator is not signaled, and the value of the output level set mode indicator is taken to be equal to said second value.
[0022] According to one embodiment of the present invention, when each level is an output level set, the image output flag in the VPS is set equal to the image output flag signaled in the image header, regardless of the value of the output level set mode indicator.
[0023] Note: Images in the output layer may or may not have PictureOutputFlag set to 1. Images in the non-output layer have PictureOutputFlag set to 0. Images with PictureOutputFlag set to 1 are output for display. Images with PictureOutputFlag set to 0 are not output for display.
[0024] According to one embodiment of the present invention, the image output flag is set to 0 when the VPS identifier in the sequence parameter set (SPS) is greater than 0, which indicates: there is more than one level in the bitstream, the "each level is an output level set" flag in the VPS is 0, which indicates: not all levels among the plurality of levels in the bitstream are independent, the output level set mode indicator is 0, and the current access unit contains an image that satisfies all of the following conditions, including: the presence of an image output flag equal to 1, the presence of a level identifier nuh greater than that of the current image, and belonging to an output level from the output level set.
[0025] According to one embodiment of the present invention, the image output flag in the VPS is set to 0 when the sequence parameter set (SPS) in the VPS is greater than 0, the "each layer is an output layer set" flag is 0, the output layer set mode indicator is 2, and the "output layer of the output layer set" flag [target OLS index] [general layer index [NUH layer identifier]] is 0.
[0026] According to one embodiment of the present invention, the method may further include controlling the display to display the decoded one or more image output layers.
[0027] According to one aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored that, when executed, cause a system or device including one or more processors to perform the following: receiving a bit stream including compressed video / image data, wherein said bit stream has a plurality of levels; obtaining, by analyzing or deriving, from the bit stream, an indicator of a mode of a set of output levels in a video parameter set (VPS); identifying a signaling of a set of output levels based on the mode indicator of a set of output levels; identifying one or more image output levels based on the identified signaling of a set of output levels; and decoding one or more identified image output levels.
[0028] According to one embodiment of the present invention, said instructions are also configured to ensure that a system or device including one or more processors performs the following: controlling a display to display decoded one or more image output levels.
[0029] According to one aspect of the present invention, an apparatus may include: at least one memory configured to store a computer program code; at least one processor configured to access said at least one memory and to perform operations in accordance with the computer program code. According to one embodiment of the present invention, the computer program code may include: a receiving code configured to cause said at least one processor to receive a bitstream including compressed video / image data, wherein said bitstream has a plurality of layers; an obtaining code configured to cause said at least one processor to obtain, by analyzing or deriving, from the bitstream, an indicator of a mode of a set of output layers in a video parameter set (VPS);an output level signaling identification code configured to cause said at least one processor to identify a signaling of a set of output levels based on an output level set mode indicator; an image output level identification code configured to cause said at least one processor to identify one or more image output levels based on the identified signaling of the set of output levels; and a decoding code configured to cause said at least one processor to perform the following: decoding the one or more identified image output levels.
[0030] According to one embodiment of the present invention, the computer program code may further include display control code configured to control, by said at least one processor, the display for displaying one or more image output levels.
[0031] According to one aspect of the present invention, a method for signaling an adaptive picture size in a bitstream may include: receiving a bitstream consisting of compressed video / image data, wherein said bitstream has a plurality of layers; identifying a background region and one or more foreground sub-images; determining whether a specific region of the sub-image is selected; and if it is determined that a specific region of the sub-image is selected: generating dequantized blocks corresponding to the selected sub-image, using a procedure including, but not limited to, analyzing the bitstream, decoding the entropy-coded bitstream and dequantizing the corresponding blocks.
[0032] The method may further include: decoding and displaying a background region if a specific region of the sub-image has not been selected.
[0033] The bitstream may include syntax elements that define which layers can be output on the decoder side.
[0034] These syntax elements may include an image header containing a variable-length syntax element encoded with the Exponential Golomb (Exp-Golomb) code.
[0035] The method may further include: determining whether or not adaptive resolution is used for the image or parts thereof based on the signaling in the sequence parameter.
[0036] Determining whether or not adaptive resolution is used for an image or parts of an image may include: determining whether the first syntax element, which is a flag, indicates the use of adaptive resolution.
[0037] The method may include instructing the encoder to use a specific reference image size, instead of unconditionally using a size equal to the output image size, using a flag that specifies the conditional presence of the reference image sizes.
[0038] The mentioned syntax elements may include a table of possible values of the width and height of the decoded image.
[0039] According to one embodiment of the present invention, a value in a Network Abstraction Layer (NAL) block header may be used to indicate not only a temporal level but also a spatial level.
[0040] This value in the NAL unit header can be the Temporal ID field.
[0041] The method may further include: using, without modification, for environments where scalability is applied, existing Selected Forwarding Units (SFUs) created and optimized for selectively forwarding temporal layers, depending on the value of the Temporal ID in the NAL unit header.
[0042] The method may further include: establishing a correspondence between the size of the encoded picture and a temporal level referenced by a Temporal ID field in a NAL unit header.
[0043] The method may further include: receiving additional data together with the encoded video, wherein the additional data is included as part of the encoded video sequence (or sequences); and using the additional data to correctly decode the data and / or to more accurately reconstruct the original video data.
[0044] Additional data may be in the form of one or more of the following: temporal, spatial, or SNR refinement layers, redundant slices, redundant pictures, or forward error correction codes.
[0045] According to one embodiment of the present invention, a machine-readable storage medium is provided, on which instructions are stored that, when executed, can cause a system or device including one or more processors to perform the following: receiving a bit stream consisting of compressed video / image data, wherein said bit stream has a plurality of layers; identifying a background region and one or more foreground sub-images; determining whether a particular region of a sub-image has been selected; and if it is determined that a particular region of a sub-image has been selected: generating dequantized blocks of samples corresponding to the selected sub-image, using a procedure including, but not limited to, analyzing the bit stream, decoding the entropy-coded bit stream and dequantizing the corresponding blocks of samples.
[0046] According to one embodiment of the present invention, the device may include: at least one memory configured to store a computer program code; and at least one processor configured to access said at least one memory and to perform operations in accordance with the computer program code, wherein the computer program code includes: a receiving code configured to cause said at least one processor to receive a bitstream consisting of compressed video / image data, wherein said bitstream has a plurality of layers; an identification code configured to identify, by said at least one processor, a background region and one or more foreground sub-images; a determining code configured to cause said at least one processor to determine whether a specific region of the sub-image is selected;and a generating code configured to ensure that said at least one processor performs the following: if it is determined that a particular region of the sub-image has been selected: generating dequantized blocks of samples corresponding to the selected sub-image using a procedure including, but not limited to, analyzing a bit stream, decoding the entropy encoded bit stream, and dequantizing the corresponding blocks of samples.Brief description of the drawings;
[0047] Further features, nature and various advantages of the proposed invention may be understood in more detail with the help of the following detailed description and the attached drawings, where:
[0048] Fig. 1 shows a schematic diagram of a simplified block diagram of a communication system in accordance with one embodiment of the present invention;
[0049] Fig. 2 is a schematic diagram of a simplified block diagram of a communication system in accordance with one embodiment of the present invention;
[0050] Fig. 3 is a schematic diagram of a simplified block diagram of a decoder in accordance with one embodiment of the present invention;
[0051] Fig. 4 is a schematic diagram of a simplified block diagram of an encoder in accordance with one embodiment of the present invention;
[0052] Fig. 5 schematically illustrates embodiments of signaling ARC parameters in accordance with the existing art or one of the embodiments of the present invention, as indicated in the drawing;
[0053] Fig. 6 shows an example of a syntax table according to one embodiment of the present invention;
[0054] Fig. 7 is a schematic illustration of a computer system in accordance with one embodiment of the present invention;
[0055] Fig. 8 shows an example of a prediction structure for adaptive resolution scaling;
[0056] Fig. 9 shows an example of a syntax table according to one embodiment of the present invention;
[0057] Fig. 10 shows a simplified block diagram of the analysis and decoding of the ROS cycle for each access block and each value of the access block serial number;
[0058] Fig. 11 is a schematic illustration of a video bitstream structure including multi-level sub-images;
[0059] Fig. 12 shows a sketch of the display of the selected sub-image with increased resolution;
[0060] Fig. 13 shows a block diagram of a procedure for decoding and displaying a video bitstream including multi-level sub-images;
[0061] Fig. 14 is a schematic illustration of the display of a 360-degree video with a sub-image refinement level;
[0062] Fig. 15 shows an example of sub-image arrangement information and a corresponding layer and image prediction structure according to one embodiment of the present invention;
[0063] Fig. 16 shows an example of sub-image layout information and the corresponding level and image prediction structure, with the possibility of spatial scaling of a local area;
[0064] Fig. 17 shows an example of a syntax table for sub-image layout information;
[0065] Fig. 18 shows an example of a syntax table of an SEI message for sub-image layout information;
[0066] Fig. 19 shows an example of a syntax table for specifying output levels and profile / tier / standard level information for each set of output levels;
[0067] Fig. 20 shows an example of a syntax table for specifying the output level mode for each set of output levels;
[0068] Fig. 21 shows an example of a syntax table for indicating the presence of a sub-image for each level, for each set of output levels;
[0069] Fig. 22 shows an example of a syntax table for the RBSP video parameter set;
[0070] Fig. 23 shows an example of a syntax table for indicating a set of output levels using an output level set mode indicator;
[0071] Fig. 24 shows a block diagram of the bit stream decoding procedure, indicating the output level set using the output level set mode indicator;
[0072] Fig. 25A-25C illustrate information related to specifying a set of output levels using the output level set mode indicator. Detailed description
[0073] When images are encoded into a bitstream consisting of multiple layers of varying quality, this bitstream may include syntax elements defining which layers can be output at the decoder. The set of layers to be output is called an output layer set. In newer video codecs supporting multi-layer video and scaling, one or more output layer sets are signaled in a video parameter set. These syntax elements, which define output layer sets and their dependencies, the profile / tier / level of the standard, and the parameters of the reference model of a hypothetical reference decoder, must be effectively signaled in one of the parameter sets.
[0074] Embodiments of the present invention provide solutions to one or more problems existing in the prior art.
[0075] Fig. 1 illustrates a simplified block diagram of a communication system (100) according to an embodiment of the present invention. The system (100) may include at least two terminals (110, 120) connected to each other by a network (150). In unidirectional data transmission, the first terminal (100) may encode video data locally for transmission to the second terminal (120) over the network (150). The second terminal (120) may receive encoded video data from the first terminal from the network (150), decode the encoded data, and display the reconstructed video data. Unidirectional data transmission is widely used in media service applications and similar applications.
[0076] Fig. 1 illustrates a second pair of terminals (130, 140) configured to support bidirectional transmission of encoded video, which may be required, for example, in videoconferencing. In bidirectional data transmission, both terminals (130, 140) can encode locally captured video data for transmission to another terminal over a network (150). Both terminals (130, 140) can also receive encoded video data transmitted by the other terminal, can decode the encoded data, and display the reconstructed video data on a local display device.
[0077] In Fig. 1, the terminals (110-140) are illustrated as a laptop computer 110, a personal computer (PC) 120, and mobile terminals 130 and 140. However, the terminals (110-140) are not limited to such embodiments and may correspond to one or more, or any combination, of the following: servers, personal computers, mobile devices, tablet computers, and smartphones. Embodiments of the present invention can be applied to laptop computers, tablet computers, media players, and / or specialized equipment for videoconferencing. The network (150) can be any number of networks transmitting encoded video data between the terminals (110 140), including, for example, wired and / or wireless communication networks. The communication network (150) can provide data exchange over circuit-switched and / or packet-switched communication lines.Examples of such networks may include telecommunication networks, local area networks (LAN), wide area networks (WAN), and / or the Internet. In the present description, the architecture and topology of the network (150) do not play any role in the operation of the proposed invention unless expressly stated.
[0078] Fig. 2 illustrates, as an example of the application of one of the embodiments of the present invention, the placement of a video encoder and a video decoder in a streaming environment. The proposed invention can be applied with equal efficiency in other areas where video is used, including, for example, videoconferencing, digital television (TV), storage of compressed video on digital media, including compact discs (CDs), digital versatile discs (DVDs), memory cards, etc.
[0079] The streaming system may include a capture subsystem (213), which may include a video source (201), such as a video camera, configured, for example, to generate a stream (202) of uncompressed samples. The sample stream (202), shown in bold in Fig. 2 to emphasize the larger amount of data compared to encoded video streams, may be processed by an encoder (303) associated with the camera (201). The encoder (203) may include hardware, software, or a combination thereof that enable aspects of the proposed invention to be implemented, in accordance with the following more detailed description. The coded video bitstream (204), shown in thin in Fig. 2 to emphasize the smaller amount of data compared to the sample stream (202), may be stored on the streaming server (205) for subsequent use.One or more streaming clients (206, 208) may access a streaming server (205) to obtain copies (207, 209) of an encoded video bitstream (204). A client device (206) may include a video decoder (210) that decodes a received copy of the encoded video bitstream (207) and generates an output stream (211) of video samples that may be displayed on a display (212) or other display device. In some streaming systems, the video bitstreams (204, 207, 209) may be encoded in accordance with specified video coding / video compression standards. An example of such standards may be ITU-T Recommendation H.265. A video coding standard that is informally referred to as Versatile Video Coding (VVC) is currently being developed. The invention described in this document can be applied in the context of the VVC standard.
[0080] Fig. 3 shows a functional block diagram of a video decoder (210) according to one embodiment of the present invention.
[0081] The receiver (310) can receive one or more encoded video sequences for decoding by the decoder (210). In the same embodiment or in alternative embodiments of the present invention, the reception of video sequences can be performed one by one, wherein the decoding of each of the encoded video sequences is independent of the decoding of the remaining video sequences. The encoded video sequence can be received from the channel (312), which can be a hardware and / or software communication line with a memory device where the encoded video data is stored. The receiver (310) can receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to the corresponding elements using them (not shown in the drawing). The receiver (310) can separate the encoded video sequence from the remaining data.To combat network jitter, a buffer memory (315) can be installed between the receiver (310) and the entropy decoder / analyzer (320) (hereinafter, the "analyzer"). When the receiver (310) receives data from a storage or transmission device with sufficient bandwidth and controllability, or from a network with isosynchronous transmission, the buffer (315) may not be used or have a small volume. In the case of using packet networks with non-guaranteed delivery, such as the Internet, the buffer (315) is necessary, can be relatively large and, preferably, have an adaptive size.
[0082] The video decoder (210) may include an analyzer (320) for recovering symbols (321) from the entropy-coded video sequence. The types of these symbols may include, for example, information used to control the operation of the decoder (210), and also, possibly, information for controlling a display device, such as a display (212), which may be associated with the decoder, but is not an integral part of it, as shown in Fig. 2. The control information for display devices may, for example, have the form of additional clarifying information (Supplementary Enhancement Information, SEI) messages or fragments of video usability information (VUI) parameter sets. The analyzer (320) may perform analysis / entropy decoding of the received encoded video sequence.The encoded video sequence may be encoded in accordance with some video coding technology or standard, and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, context-dependent or context-independent arithmetic coding, and the like. The analyzer (320) may extract from the encoded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder, based on at least one parameter corresponding to the group. Subgroups may include Groups of Pictures (GOP), images, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), and the like.The analyzer can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0083] The analyzer (320) can perform entropy decoding / analysis operations on the video sequence received from the buffer (315) and generate symbols (321).
[0084] When restoring symbols (321), different device units may be used depending on the type of encoded video images or parts thereof (e.g., intra- and inter-predicted images, intra- and inter-predicted sample blocks), as well as other factors. Which device units will be used and how can be determined by the subgroup control information extracted from the encoded video sequence by the analyzer (320). The subgroup control information stream is transmitted between the analyzer (320) and a plurality of device units.
[0085] In addition to the functional blocks already mentioned, decoder 210 can be conceptually subdivided into a set of functional blocks described below. In practical implementations used in commercial settings, many of these blocks closely interact with each other and can be, at least partially, mutually integrated. However, for the purposes of describing the proposed invention, the subdivision into functional blocks described below is suitable.
[0086] The first of such blocks may be the scaling / inverse transform block (351). The scaling / inverse transform block (351) receives quantized transform coefficients, as well as control information, including information on which transform to use, the block size, the quantization coefficient, the quantization scaling matrices, etc., in the form of symbols (321) from the analyzer (320). It outputs sample blocks, including sample values, which can be input into the aggregator (355).
[0087] In some cases, the samples at the output of the scaling / inverse transform unit (351) may belong to an intra-coded block of samples, that is: a block of samples for which prediction information from previously reconstructed images is not used, but information from previously reconstructed parts of the current image may be used. Such prediction information may be provided by the intra-image prediction unit (352). In some cases, the intra-image prediction unit (352) may form blocks of samples of the same size and shape as the block of samples being reconstructed, using already reconstructed information of its surroundings obtained from the current (partially reconstructed) image in the current image memory (358).The aggregator (355) may, in some cases, add the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaling / inverse transform unit (351), individually for each sample.
[0088] In other cases, the samples at the output of the scaling / inverse transform block (351) may belong to a block with inter prediction, possibly with motion compensation. In such cases, the motion-compensated prediction block (353) may access the reference image memory (357) and obtain the samples used for prediction. After motion compensation, the obtained samples, in accordance with the symbols (321) related to this block of samples, may be added by the aggregator (355) to the output data of the scaling / inverse transform block (in this case, they are called difference samples or a difference signal), thereby forming the output sample information.The addresses in the reference image memory from which the motion-compensated prediction unit obtains the predicted samples may be determined by motion vectors accessible to the motion-compensated prediction unit in the form of symbols (321), which may have, for example, an X-component, a Y-component, and a reference image component. Motion compensation may also include interpolation of sample values obtained from the reference image memory (457), when motion vectors, motion vector prediction mechanisms, etc., with subpixel accuracy are used.
[0089] The samples at the output of the aggregator (355) may be processed using various loop filtering methods in the loop filtering unit (356). Video compression technologies may include in-loop filtering technologies that are controlled by parameters contained in the encoded video bitstream and provided to the loop filtering unit (356) in the form of symbols (321) from the analyzer (320). They may also depend on metainformation obtained during decoding of preceding (in decoding order) parts of the encoded image or encoded video sequence, as well as on previously reconstructed and loop filtered sample values.
[0090] The output data of the loop filtering unit (356) may be a stream of samples that is fed to the display device (212) and also stored in the reference image memory (356) for use in future external image prediction.
[0091] Individual encoded images, after their complete restoration, can be used as reference images for future prediction. After the encoded image is completely restored, and if it has been determined to be a reference image (e.g., by the analyzer (320)), the current reference image (356) can be placed in the reference image buffer (357), and new memory for the current images can be allocated before the next encoded image is restored.
[0092] Video decoder (analyzer) 320 can perform decoding operations in accordance with a predetermined video compression technology, which can be documented in a standard, for example, ITU-T Recommendation H.265. The encoded video sequence can satisfy the syntax specified by the applied video compression technology or standard, in the sense that it satisfies the syntax specified in the document or standard of the video compression technology, and in particular, the syntax of the specified profiles of the standard. Moreover, in order to comply with some of the video compression technologies or standards, the complexity of the encoded video sequence must be within the limitations determined by the level of the video compression technology or standard.In some cases, the standard's levels limit the maximum image size, maximum frame rate, maximum sample rate (measured, for example, in millions of samples per second), maximum reference image size, etc. The limitations imposed by the levels can in some cases be further limited by the specifications of the Hypothetical Reference Decoder (HRD) and the metadata for managing the HRD decoder's buffer, signaled in the encoded video sequence.
[0093] In one embodiment of the present invention, the receiver (310) may receive additional (redundant) data along with the encoded video. This additional data may be a component of the encoded video sequence (or video sequences). The additional data may be used by the video decoder (320) to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, refinement temporal, spatial, or SNR layers, redundant slices, redundant images, forward error correction codes, etc.
[0094] Fig. 4 shows a functional block diagram of a video encoder (203) according to one embodiment of the present invention.
[0095] The encoder (203) can receive video samples from a video source (201) (which is not part of the encoder), capturing video images for encoding by the encoder (203).
[0096] The video source (201) may supply a source video sequence for encoding by the video encoder (203) in the form of a digital stream of video samples having any suitable bit depth (e.g. 8 bit, 10 bit, 12 bit, …), any color space (e.g. BT.601 Y CrCb, RGB, …) and any suitable report structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (201) may be a storage device storing pre-prepared video. In a videoconferencing system, the video source (203) may be a video camera that captures image information locally in the form of a video sequence. The video data may be in the form of a plurality of individual images that convey a sense of motion when viewed sequentially.The images themselves can be organized as a spatial matrix of pixels, where each pixel can include one or more samples, depending on the sample structure used, the color space, etc. Those skilled in the art will understand the relationship between pixels and samples. Further in this description, samples will be discussed.
[0097] According to one embodiment of the present invention, the video encoder (203) can encode and compress images of the original video sequence into the encoded video sequence (443) in real time, or in accordance with other time constraints imposed by practical application. One of the functions of the controller (450) can be to ensure a suitable encoding rate. The controller (450) can control the remaining functional blocks described below and be operatively connected with these blocks. The parameters set by the controller (450) can include parameters related to rate control (picture skipping, quantizer, λ value for rate-distortion optimization methods, ...), picture size, group of pictures (GOP) arrangement, maximum search range of motion vectors, etc.Those skilled in the art will also be aware of other functions of the controller (450) that may correspond to a video encoder (203) optimized for a particular system design.
[0098] Some video encoders operate in a configuration that those skilled in the art call a "coding loop." In very simplified terms, the coding loop may consist of a coding subsystem in the encoder (203) (hereinafter referred to as the "source encoder"), responsible for generating symbols based on the input images to be encoded, as well as reference images, and a (local) decoder (433), built into the encoder (203) and restoring the symbols, generating sample data that would be identically generated by a (remote) decoder (since the compression of symbols into an encoded video bitstream is lossless compression in the video compression technologies considered in the present invention). This restored sample stream is input into a reference image memory (434).Since decoding a symbol stream yields the same bit-perfect results regardless of the decoder (local or remote), the contents of the reference picture buffer are also bit-perfectly identical in both the local and remote encoder. In other words, the prediction subsystem in the encoder "sees" as reference picture samples exactly the same sample values that the decoder would see using prediction during decoding. This fundamental principle of reference picture synchronicity (and the resulting drift if synchronicity cannot be ensured, for example, due to channel errors) should be well known to those skilled in the art.
[0099] The operation of the "local" decoder (433) is essentially identical to the "remote" decoder (210), which has already been described in detail above in connection with Fig. 3. However, returning to Fig. 3, since the symbols are available, and the encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (445) and the analyzer (320) can be performed without loss, the local decoder (433) may not fully implement the entropy decoding subsystems of the decoder (210), including the channel (312), the receiver (310), the buffer (315), and the analyzer (320).
[0100] It should be noted here that any decoding technology other than analysis / entropy decoding present in a decoder must be present in a substantially identical functional form in the corresponding encoder. For this reason, the description of the present invention focuses on the operation of the decoder. A description of the encoding technologies may be omitted, as they may be inverse to the decoding technologies described in detail. Only in a few places is a more detailed description necessary, and this will be provided below.
[0101] Among its operations, the source encoder (203) may perform motion-compensated predictive coding, in which input frames are predictively encoded based on one or more previously encoded frames of a video sequence that have been marked as "reference frames." Thus, the coding subsystem (432) encodes differences between blocks of pixels in an input frame and blocks of pixels in a reference frame (or frames), which can be selected as reference for predicting the input frame.
[0102] The local video decoder (433) can decode the encoded video data of the frames marked as reference frames, depending on the symbols generated by the source encoder (430). The operations of the coding subsystem (432) are preferably lossy data processing. When the encoded video data is decoded in the video decoder (not shown in Fig. 4), the reconstructed video sequence is typically a replica of the original video sequence with some errors. The local video decoder (433) exactly reproduces the decoding process that could be performed by the remote video decoder on the reference frames and places the reconstructed reference frames in the reference picture cache (434). In this way, the encoder (203) can locally store copies of the reconstructed reference frames, the contents of which match the reconstructed reference frames received by the remote video decoder (in the absence of transmission errors).
[0103] The predictor (435) can search for predictions for the coding subsystem (432). That is, for a new frame to be coded, the predictor (435) can search the reference image memory (434) to find sample data (as candidate reference pixel blocks) or metadata, such as motion vectors of reference images, block shapes, etc., that can serve as reference for new images. The predictor (435) can find suitable reference data for each individual pixel block. In some cases, depending on the search results obtained by the predictor (435), the reference data for predicting the input image can be extracted from multiple reference images stored in the reference image memory (434).
[0104] The controller (450) may control encoding operations in the video encoder (203), including, for example, setting parameters and subgroup parameters used to encode video data.
[0105] The output data of all the above-described functional blocks can be entropy-encoded in the entropy encoder (445). The entropy encoder converts the symbols generated by various functional blocks into an encoded video sequence by losslessly compressing these symbols in accordance with technologies known to those skilled in the art, such as Huffman coding, variable-length coding, arithmetic coding, and the like.
[0106] The transmitter (440) may buffer the encoded video sequence (or video sequences) generated by the entropy encoder (445) to prepare it for transmission over the communication channel (460), which may be a hardware and / or software communication line with a storage device where the encoded video data is stored. The transmitter (440) may combine the encoded video data from the video encoder (430) with other transmitted data, such as streams of encoded audio data and / or service data.
[0107] The controller (450) can control the operation of the encoder (203). During encoding, the controller (450) can assign each encoded image a certain encoded image type, which can influence the encoding methods applied to it. For example, images can be assigned one of the frame types described below.
[0108] An intra-predicted picture (I-picture) is a picture that is encoded and decoded without using any other frames in the video sequence as a source for prediction. Some video codecs support various types of intra-predicted pictures, such as Independent Decoder Refresh (IDR) pictures. Those skilled in the art should be aware of these types of I-pictures, as well as their properties and applicability.
[0109] A predictable image (P-image) is an image that can be encoded and decoded using intra-prediction or inter-prediction using at most one motion vector and a pointer to a reference image to predict the sample values of each block.
[0110] A bidirectionally predicted image (B-image) is an image that can be encoded and decoded using intra- or inter-prediction using a maximum of two motion vectors and reference image pointers to predict the sample values of each block. Similarly, in the case of multidirectionally predicted images, more than two reference images and corresponding metadata can be used to reconstruct one block of samples.
[0111] Source images are typically spatially partitioned into multiple blocks (e.g., 4x4, 8x8, 4x8, or 16x16 sample blocks) and encoded block-by-block. Sample blocks can be coded with prediction based on other (already coded) blocks, depending on the coding types assigned to the images corresponding to these blocks. For example, blocks in I-pictures can be coded without prediction or with prediction based on already coded blocks of the same image (“spatial prediction” or “intra prediction”). Blocks of pixels in P-pictures can be coded without prediction, using spatial prediction, or using temporal prediction based on one previously coded reference image. Blocks in B-pictures can be coded without prediction, using spatial prediction, or using temporal prediction based on one or two previously coded reference images.
[0112] The video encoder (203) may perform encoding operations in accordance with a predetermined video coding technology or standard, which may be documented in a standard such as ITU-T Recommendation H.265. During operation, the video encoder (203) may perform various compression operations, including predictive encoding operations that utilize temporal and spatial redundancy in the input video sequence. The encoded video data may accordingly satisfy the syntax specified by the video coding technology or standard used.
[0113] In one embodiment of the present invention, the transmitter (440) may transmit additional data along with the encoded video. The video encoder (430), for example, may provide such data as a fragment of the encoded video sequence. The additional data may include data on temporal, spatial, or SNR refinement levels, redundant pictures or slices, additional refinement information (SEI) messages, or fragments of video usability information (VUI) parameter sets, etc.
[0114] Before aspects of the present invention are discussed in more detail, several terms used in the remainder of this description are introduced.
[0115] A sub-image, in some cases, is further understood as a rectangular structure of samples, blocks, macroblocks, coding units, or similar elements, grouped semantically, which can be coded independently at different resolutions. One or more sub-images can form an image. One or more coded sub-images can form a coded image. One or more sub-images can be combined into an image, and one or more sub-images can be extracted from the image. In some embodiments of the present invention, one or more sub-images can be combined in compressed form, without re-coding at the sample level, into a single coded image, and in this case, as well as in other cases, one or more coded sub-images can be extracted from the coded image in compressed form.
[0116] Adaptive Resolution Change (ARC) is a mechanism that allows for changing the resolution of an image or subimage within an encoded video sequence, for example, by resampling reference images. ARC parameters are the control information required to perform adaptive resolution change, which may include, for example, filter parameters, scaling factors, output and / or reference image resolutions, various control flags, etc.
[0117] The above description relates to the encoding and decoding of individual semantically independent video images, in accordance with various embodiments of the present invention. Before describing the essence of encoding / decoding multiple sub-images with independent ARC parameters and the additional complications they introduce, we will consider options for signaling ARC parameters.
[0118] FIG. 5 illustrates several embodiments of ARC parameter signaling proposed in the present invention. As noted in each of the embodiments, they have their own advantages and disadvantages in terms of coding efficiency and complexity, as well as in terms of architecture. In a particular video coding standard or technology, one or more of the proposed embodiments of ARC parameter signaling, as well as embodiments known in the existing art, may be selected. These embodiments are not necessarily mutually exclusive and may be interchangeable depending on the requirements of a particular application, the technologies and standards used, and the choices made by the encoder.
[0119] ARC parameter classes may include:
[0120] - upsampling or downsampling coefficients, separately for the X and Y axes, or in combination;
[0121] - upsampling or downsampling coefficients, with the addition of a time dimension indicating a constant rate of upscaling / downscaling for a given number of images;
[0122] - either of the previous two may include encoding one or more as short as possible syntactic elements that may point to a table containing the mentioned coefficients;
[0123] - The resolution, along the X or Y axis, measured in samples, blocks, macroblocks, coding units (CUs), or any other suitable precision, for the input image, output image, reference image, and encoded image, in combination or individually. If more than one resolution is present (for example, one for the input image and one for the reference image), then in some cases one set of values may be calculated based on the other set of values. This can be controlled, for example, using special flags. A more detailed example will be discussed below.
[0124] - warping coordinates similar to those used in H.263 Annex P, again, as mentioned above, with suitable precision. H.263 Annex P defines one efficient method for encoding warping coordinates, but more efficient methods could potentially be proposed. For example, the reversible coding, such as Huffman coding, with a variable-length codeword of the warping coordinates in Annex P could be replaced by binary coding with a codeword of appropriate length, which could, for example, be calculated based on the maximum image size, possibly multiplied by a coefficient and offset by a given value to allow "warping" outside the boundaries defined by the maximum image size; and
[0125] - upsampling and downsampling filter parameters. In the simplest case, there may be only one filter for upsampling and / or downsampling. However, in other cases, greater flexibility in filter design may be desirable, which may require signaling of filter parameters. Such parameters can be selected using a pointer in a list of possible filter constructions, the filter can be fully specified (e.g., using a list of filter coefficients, using suitable entropy encoding methods), the filter can be implicitly selected using upsampling or downsampling coefficients, which, in turn, are signaled according to the mechanisms described above, etc.
[0126] In the following description, we assume encoding a finite set of up- or downsampling coefficients (the same coefficient is used for both the X-axis and Y-axis), specified using a codeword. Such a codeword may preferably have a variable length, for example, using Exponential Golomb coding (Exp-Golomb) for the corresponding syntax elements in video coding standards such as H.264 and H.265.
[0127] One of the possible tables of correspondence between the values of upsampling and downsampling coefficients is illustrated in the table below.
[0128] Multiple similar lookup tables can be proposed, depending on application requirements and the capabilities of the up- and downsampling mechanisms available in a video coding standard or technology. The table can be expanded to include more values. Values can also be represented using coding mechanisms other than Exp-Golomb codes, such as binary encoding. This provides certain advantages when resampling coefficients are used outside of the video processing subsystems themselves (primarily in the encoder and decoder), for example, in mobile ad hoc networks (MANETs). It should be noted that in the (presumably) most common case, when resolution changes are not required, a short Exp-Golomb code, such as the one in the table above, which has a length of only one bit, can be selected.This allows for an advantage in coding efficiency, in general, compared to using binary codes.
[0129] The number of table entries, as well as their semantics, can be partially or fully configurable. For example, the basic table framework can be conveyed in a high-level parameter set, such as a sequence parameter set or a decoder parameter set. Alternatively, or in addition, a video coding technology or standard may define several similar tables, selected, for example, using a decoder parameter set or sequence.
[0130] Below, we will discuss how up- or downsampling coefficients (ARC information), encoded as described above, can be incorporated into the syntax of a video coding technology or standard. Similar considerations can be applied to one or more codewords controlling up- or downsampling filters. The case where filters or other data structures require relatively large amounts of data will be discussed below.
[0131] Annex P of the H.263 specification (H.263 Annex P) includes the 502 ARC information in the form of four warp coordinates in the 501 image header, namely, in the PLUSPTYPE 503 header extension of the H.263 standard. This seems to be a rational solution for cases where a) the image header is available, and b) frequent changes to the ARC information are expected. However, the amount of overhead information when using the signaling type proposed in H.263 can be quite high, and the scaling factors between image boundaries may be inconsistent, since image headers are not necessarily stored permanently.
[0132] The JVCET-M135-v1 specification cited above includes ARC reference information (505) (index) located in the sequence parameter set (504), which points to a table (506) including target resolutions, which, in turn, is located inside the sequence parameter set (SPS) (507). The placement of the available resolutions in the table (506) in the sequence parameter set (507) (507), according to the authors, is due to the use of the SPS set as a compatibility agreement point in the exchange of capability information. The resolution can vary within the range specified by the values of the table (506) from image to image by referencing the appropriate image parameter set (504).
[0133] Returning to Fig. 5, we also see additional possible options for transmitting ARC information in the encoded video bitstream, described below. Each of these options offers advantages over the existing technology described above. The proposed options may be present simultaneously in a video coding technology or standard.
[0134] In one embodiment of the present invention, the ARC information (509), such as a resampling (scaling) factor, may be present in a slice header, a group of block (GOB) header, a tile header, or a tile group header (508) (hereinafter, in the tile group header). This may be an adequate solution if the ARC information has a small volume, such as a single variable-length ue(v) value or a fixed-length codeword of several bits, as shown above. Placing the ARC information directly in the tile group header has the additional advantage that the ARC information can be applied to a sub-image represented, for example, by this tile group, and not only to the entire image. This will be described in more detail below.Furthermore, even if a video coding technology or standard only allows for adaptive resolution change for images as a whole (as opposed to, for example, adaptive resolution change at the tile group level), placing ARC information in the tile group header, as opposed to placing it in the image header, like H.263, provides certain advantages in terms of error resilience.
[0135] In this or other embodiments of the present invention, for example, the ARC information (512) may be directly represented in a corresponding parameter set (511), such as a picture parameter set (PPS), a header parameter set, a tile parameter set, an adaptation parameter set (an adaptation parameter set is shown in the figure), etc. The scope of this parameter set may preferably be no larger than an image, for example, a group of tiles. The use of the ARC information is specified implicitly by activating the corresponding parameter set. For example, if only picture-level ARC is defined in a video coding technology or standard, then a picture parameter set or an equivalent set may be suitable.
[0136] In this or other embodiments of the present invention, for example, the ARC reference information (513) may be present in a tile group header (514) or in a similar data structure. Such reference information (513) may reference a subset of the ARC information (515) available in a parameter set (516) with a scope exceeding an individual picture, for example, in a sequence parameter set or a decoder parameter set.
[0137] The additional layer of implicit activation of the PPS set from the tile group header, PPS, SPS, used in JVET-M0135-v1 seems redundant, since image parameter sets, like sequence parameter sets, can (and should in some standards, such as RFC3984) be used for interoperability agreements or declarations. However, if the ARC information should also apply to a subimage, such as that represented by a tile group, a better choice might be a parameter set with a scope limited to the tile group, such as an adaptation parameter set or a header parameter set.Also, if the ARC information is non-negligible in size, such as containing filter control information, i.e., a set of filter coefficients, then from a coding efficiency standpoint, a parameter set may be a better solution than using the header (508) directly, since these parameters can be reused for future images or sub-images by referencing the same parameter set.
[0138] When using a sequence parameter set or higher-level parameter sets with a scope spanning multiple images, the considerations described below may be relevant.
[0139] 1. The parameter set for storing the table (516) with the ARC information may in some cases be a sequence parameter set, but in other cases, it is preferably a decoder parameter set. The scope of the decoder parameter set may cover several CVSs, namely, an encoded video stream, i.e., all bits of the encoded video, from the beginning to the end of the session. Such a scope may be more suitable since the ARC coefficients can be programmed in the decoder, and possibly implemented in hardware, so the hardware parameters remain largely constant within a single CVS (which, at least in some multimedia systems, is a group of pictures of one second or less in length). At the same time, placing the table in a sequence parameter set is explicitly listed among the placement options discussed in this document, namely, in connection with paragraph 2 below.
[0140] 2. The ARC reference information (513) can be preferably placed directly in the picture / slice / tile / GOP / tile group header (hereinafter referred to as the tile group header) (514) instead of the picture parameter set as in JVCET-M0135-v1. This is preferable for the following reasons: When an encoder needs to change one value in the picture parameter set, such as the ARC reference information, it needs to create a new PPS and reference this new PPS. Suppose that only the ARC reference information is changed, and the rest of the information, such as the quantization matrix information, in the PPS remains unchanged. This information may be quite large, and it needs to be retransmitted to make the new PPS complete.
[0141] Since the ARC reference information (513) can be a single codeword, such as a table pointer, and this is the only value that changes, it is inconvenient and irrational to retransmit all the quantization matrix information. Therefore, from the perspective of coding efficiency, it would be significantly better to eliminate the indirect reference via the PPS set proposed in JVET-M0135-v1. Also, placing the ARC reference information in the PPS set has an additional inconvenience related to the fact that the ARC information referenced by the ARC reference information (513) will need to be applied to the entire image rather than to sub-images, since the scope of the image parameter set is the entire image.
[0142] In the same or an alternative embodiment of the present invention, the signaling of ARC parameters may correspond to the detailed example shown in Fig. 6. Fig. 6 shows syntax trees in the notation used in video coding standards since at least 1993. This notation roughly corresponds to the C programming language. The lines in bold in Fig. 6 indicate syntactic elements present in the bitstream. Lines not in bold generally indicate control commands or the assignment of variable values.
[0143] The tile group header (601), as an example of the syntactic structure of a header applicable to a (possibly rectangular) portion of an image, may, by convention, contain the variable-length, Exp-Golomb-encoded dec_pic_size_idx syntax element (602) (highlighted in bold). The presence of this syntax element in the tile group header may be determined by the use of adaptive resolution (603). In this example, the flag value is not highlighted in bold, meaning that the flag is present in the bitstream at the time it appears in the syntax diagram.
[0144] Whether adaptive resolution is applied to a given image or part of it can be signaled in any high-level syntactic structure within or outside the bitstream. In the example illustrated in Fig. 6, this is signaled in the sequence parameter set, as will be shown below.
[0145] Fig. 6 also shows a fragment of the sequence parameter set (610). The first syntax element shown is the adaptive_pic_resolution_change_flag flag (611). When its value is "TRUE," the flag indicates the use of adaptive resolution, which in turn may require corresponding control information. In this example, such control information is conditionally included depending on the flag value, using an if() expression in the parameter set (612) and in the tile group header (601).
[0146] When using adaptive resolution, in this example, the output resolution is encoded in samples (613). The number denoted by 613 refers to both output_pic_width_in_luma_samples and output_pic_width_in_luma_samples, which together define the output image resolution. A specific video coding technology or standard may elsewhere define some restrictions on any of these values. For example, the standard-level resolution may limit the total number of output samples, which may be equal to the product of the values of these two syntax elements. Also, specific video coding technologies or standards, or external technologies or standards, such as system standards, may restrict the numerical range (for example, one or both dimensions must be a multiple of a power of two) or the aspect ratio (for example, the height and width must have a specified ratio, such as 4:3 or 16:9).Such limitations may be imposed by hardware implementation or other reasons and should be known to those skilled in the art.
[0147] In some applications, the encoder may instruct the decoder to use a specific image size, rather than unconditionally using a size equal to the output image size. In this example, the reference_pic_size_present_flag (614) specifies the conditional presence of the reference image dimensions (615) (again, the notations shown refer to both the width and height).
[0148] Finally, Fig. 6 shows a table of possible values of the width and height of the decoded image. Such a table can be specified, for example, using a pointer (616) to a table, (num_dec_pic_size_in_luma_samples_minus1). "minus1" can indicate the interpretation of the value of this syntax element. For example, if the encoded value is zero, there is only one table entry. If the value is five, there are six table entries. For each "row" of the table, the width and height of the decoded image are then included in the syntactic structure (617).
[0149] The table entries designated by (617) can be referenced by the syntax element (602) dec_pic_size_idx in the tile group header, which allows for different decoded image sizes, essentially a zoom effect, in each individual tile group.
[0150] Some video coding technologies or standards, such as VP9, support spatial scaling by performing some form of reference image resampling (the corresponding signaling differs significantly from that proposed in this document) in combination with temporal scaling to enable spatial scalability. Specifically, individual reference images can be upsampled to higher resolutions using methods such as ARC, forming the basis of a spatial refinement layer. The upsampled images can be refined using standard high-resolution prediction mechanisms, allowing for greater detail.
[0151] The present invention can be applied in such an environment. In some cases, in the same or an alternative embodiment of the present invention, a value in the Network Abstraction Layer (NAL) unit header, for example, the Temporal ID field, can be used to indicate not only the temporal but also the spatial layer. This can provide advantages for some types of systems; for example, for environments where scalability is applied, existing Selected Forwarding Units (SFUs) created and optimized for selectively forwarding temporal layers, depending on the Temporal ID value in the NAL unit header, can be used without modification. To implement this functionality, a correspondence between the encoded picture size and the temporal layer referenced by the Temporal ID field in the NAL unit header may be necessary.
[0152] In some video coding technologies, an Access Unit (AU) may refer to coded images, slices, tiles, NAL units, etc., that were captured and composed into the corresponding image / slice / tile / NAL unit bitstream at a given point in time. Such a point in time could be, for example, composition time.
[0153] In the HEVC standard and some other video coding technologies, the picture order count (POC) value can be used to indicate a selected reference picture among multiple reference pictures stored in the decoded picture buffer (DPB). When an access unit (AU) contains one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU can have the same POC value, which indicates that they were formed based on content with the same composition time. In other words, if two pictures / slices / tiles contain the same specified POC value, these two pictures / slices / tiles belong to the same AU and have the same composition time.Conversely, two images / slices / tiles with different POC values indicate that these images / slices / tiles belong to different access blocks and have different composition times.
[0154] In one embodiment of the present invention, this strict requirement may be relaxed, i.e., an access unit may include pictures, slices, or tiles with different POC values. By allowing different POC values in a single access unit, the POC value can be used to identify potential independently decodable pictures / slices / tiles with the same display time. This, in turn, allows for support of multiple zoom levels without changing the reference picture selection signaling (e.g., signaling reference picture sets or a reference picture list), described in more detail below.
[0155] However, it is still desirable to be able to identify the access unit to which an image / slice / tile belongs among other images / slices / tiles with different POC values, based solely on the POC value. This can be achieved as described below.
[0156] In this or other embodiments, the access unit count (AUC) may be signaled in a high-level syntax structure, such as a NAL unit header, slice header, tile group header, SEI message, parameter set, or access unit delimiter. The AUC value may be used to determine which NAL units, pictures, slices, or tiles belong to a given access unit. The AUC value may correspond to a specific point in time of the composition. The AUC value may be a multiple of the POC value. The AUC value may be calculated by dividing the POC value by an integer value. In some embodiments, division operations may impose a significant load on the decoder. In such cases, a slight limitation in the number space for AUC values allows replacing the division operation with a shift operation.For example, the AUC value may be equal to the value of the Most Significant Bit (MSB) in the range of POC values.
[0157] In this embodiment, the POC cycle value for an access unit (poc_cycle_au) may be signaled in a top-level syntax structure, such as a NAL unit header, a slice header, a tile group header, an SEI message, a parameter set, or an access unit delimiter. The poc_cycle_au value may indicate how many consecutive distinct POC values may be associated with one access unit. For example, if the poc_cycle_au value is 4, pictures, slices, or tiles with a POC value of 0-3, inclusive, belong to an access unit with an AUC value of 0, and pictures, slices, or tiles with a POC value of 4-7, inclusive, belong to an access unit with an AUC value of 1. Therefore, the AUC value can be calculated by dividing the POC value by the poccycleau value.
[0158] In this or another embodiment, the poc_cycle_au value may be calculated based on information located, for example, in a video parameter set (VPS), which specifies the number of spatial or SNR levels in the encoded video sequence. One example of such a potential correspondence is briefly described below. The calculations described above allow for saving several bits in the VPS, and therefore increasing coding efficiency. However, it may be preferable to explicitly encode poc_cycle_au in an appropriate top-level syntax structure hierarchically subordinate to the VPS, which allows for minimizing poc_cycle_au for a given small portion of the bitstream, such as an image.This optimization allows saving more bits than the calculation procedure described above, since the values of the POC (and / or the values of syntactic elements that indirectly refer to the POC) can be encoded in lower-level syntactic structures.
[0159] Fig. 9, according to the same or an alternative embodiment, shows an example of syntax tables for signaling the vps_poc_cycle_au syntax element in the VPS (or SPS), which indicates the poc_cycle_au used for all pictures / slices in the coded video sequence in the slice header. If the POC value increases uniformly in each access unit, vps_contant_poc_cycle_per_au in the VPS can be set to 1, and vps_poc_cycle_au is signaled in the VPS. In this case, slice_poc_cycle_au is not explicitly signaled, and the AUC value for each access unit is calculated by dividing the POC value by vps_poc_cycle_au. If the POC value in each access unit does not increase uniformly, then vps_contant_poc_cycle_per_au in VPS can be set to 0. In this case, vps_access_unit_cnt is not signaled, while Slice_access_unit_cnt is signaled in the slice header for each image slice.Each image slice may have a different slice_access_unit_cnt value. The AUC value for each access unit is calculated by dividing the POC value by slice_poc_cycle_au. Figure 10 shows a flowchart illustrating the corresponding sequence of operations.
[0160] In this or other embodiments, although the AUC values in images, slices, or tiles may be different, these images, slices, or tiles corresponding to an access unit with the same AUC values may be assigned to the same decoding or output time. Therefore, in the absence of any analysis or decoding dependencies between images, slices, or tiles in a single access unit, all images, slices, or tiles assigned to a single access unit, or a subset thereof, may be decoded in parallel and output at the same time.
[0161] In this or other embodiments, although the POC values in images, slices, or tiles may be different, these images, slices, or tiles corresponding to an access unit with the same AUC values may belong to the same composition / display time. When the composition time is contained in a container format, even if the images correspond to different access units but have the same composition time, these images may be displayed at the same time.
[0162] In this or other embodiments, each image, slice, or tile may have the same temporal identifier (temporal_id) in a single access unit. All images, slices, or tiles corresponding to a single moment in time, or a subset thereof, may correspond to a single temporal sublevel. In this or other embodiments, all images, slices, or tiles may have the same or different spatial layer identifiers (layer_id) in a single access packet. All images, slices, or tiles corresponding to a single moment in time, or a subset thereof, may correspond to a single or different spatial levels.
[0163] Fig. 8 shows an example of the structure of a video sequence with a combination of temporal_id, layer_id, POC, and AUC values during adaptive resolution change. In this example, an image, slice, or tile in the first access unit with AUC=0 may have temporal_id=0 and layer_id=0 or 1, while an image, slice, or tile in the second access unit with AUC=1 may have temporal_id=1 and layer_id=0 or 1, respectively. The POC value is increased by 1 in each image regardless of the values of temporal_id and layer_id. In this example, the poc_cycle_au value may be equal to 2. Preferably, the poc_cycle_au value may be set equal to the number of levels (spatial scaling). In this example, accordingly, the POC value is increased by 2, while the AUC value is increased by 1.
[0164] In the above embodiments, the entire inter-layer prediction and reference picture indication structure, or a subset thereof, may be supported using the existing reference picture set (RPS) signaling in HEVC or using the reference picture list (RPL) signaling. In the RPS or RPL, the indication of a selected reference picture is performed by signaling a POC value or a POC difference value between the current picture and the selected reference picture. In the present invention, the RPS and RPL can be used to indicate the inter-picture prediction or inter-layer prediction structure without changing the signaling, but with the limitations described below. If the temporal_id value of a reference picture is greater than the temporal_id value of the current picture, this reference picture cannot be used for motion compensation in the current picture or other prediction.If the layer_id value of the reference image is greater than the layer_id value of the current image, then this reference image cannot be used for motion compensation in the current image or other prediction.
[0165] In this or other embodiments, scaling motion vectors based on the difference in POC for temporal motion vector prediction may be prohibited between different images within an access unit. Therefore, although each value may have different POC values within an access unit, the motion vector is not scaled and is used for temporal motion vector prediction within the access unit. This is necessary because a reference image with a different POC in the same access unit is considered a reference image of the same time instance. Accordingly, in this embodiment, the motion vector scaling function may return 1 when the reference image belongs to the access unit to which the current image belongs.
[0166] In this or other embodiments, motion vector scaling based on the POC difference for temporal motion vector prediction may optionally be disabled over several images when the spatial resolution of the reference image differs from the spatial resolution of the current image. When motion vector scaling is enabled, it is performed based on both the POC difference and the spatial resolution ratio between the current image and the reference image.
[0167] In the same or other embodiments, motion vector scaling may be performed based on the AUC difference, instead of the POC difference, for temporal prediction of motion vectors, in particular when poc_cycle_au has non-uniform values (when vps_contant_poc_cycle_per_au=0). Otherwise (when vps_contant_poc_cycle_per_au=1), motion vector scaling based on the AUC difference may be identical to motion vector scaling based on the POC difference.
[0168] In the same or other embodiments, when a motion vector is scaled based on the AUC difference, a reference motion vector in the same access block (with the same AUC value) as the current image is not scaled based on the AUC difference and is used to predict motion vectors without scaling or with scaling based on the spatial resolution ratio between the current image and the reference image.
[0169] In this or other embodiments, the AUC value is used to determine the boundaries of an access unit, as well as for the operation of a hypothetical reference decoder (HRD), which requires input and output times accurate to the access unit. In most cases, the decoded image in the upper layer of the access unit can be output for display. The AUC value and the layer_id value can be used to identify the output image.
[0170] In one embodiment, an image may consist of one or more sub-images. Each sub-image may cover a local region or the entire image region. The region occupied by a sub-image may or may not overlap with regions of other images. A region consisting of one or more sub-images may cover either the entire image region or a portion thereof. If an image consists of a sub-image, the region occupied by the sub-image may be identical to the region occupied by the image.
[0171] In the same embodiment, a sub-image may be encoded using a coding method similar to the coding method used for the encoded image. A sub-image may be encoded independently or coding dependent on another sub-image or the encoded image. A sub-image may be either dependent, with respect to its syntactic analysis, on another sub-image or the encoded image, or independent.
[0172] In this embodiment, an encoded sub-image may be contained in one or more layers. The encoded sub-images in different layers may have different spatial resolutions. The original sub-image may be spatially resampled (up- or down-sampled), encoded using different spatial resolution parameters, and included in the bitstream corresponding to one of the layers.
[0173] In the same or other embodiments, a sub-image with parameters (W, H), where W is the width of the image and H is the height of the sub-image, respectively, may be encoded and included in the encoded bitstream corresponding to level 0, while a sub-image with increased (or decreased) resolution, obtained based on the sub-image with the original spatial resolution, with parameters (W*S w,k , H*S h,k ), can be encoded and included in the encoded bitstream corresponding to level k, where S w,k , S h,k - horizontal and vertical resampling coefficients. If the value of S w,k , S h,k greater than 1, resampling will be upward. Whereas if the values of S w,k , S h,k less than 1, resampling will be downsampling.
[0174] In this or other embodiments, an encoded image in one of the layers may have a different visual quality than an encoded sub-image in another layer for the same sub-image or a different sub-image. For example, sub-image i in layer n may be encoded with a parameter Q i,n quantization, while subimage j in level m can be encoded with parameter Q j,m quantization.
[0175] In the same or another embodiment, an encoded sub-picture in one of the layers can be decoded independently, without any analysis or decoding dependency on encoded sub-pictures in other layers for the same local region. A sub-picture layer that can be independently decoded without reference to other sub-picture layers of the same local region is called an independent sub-picture layer. An encoded sub-picture in an independent sub-picture layer may or may not have a decoding or analysis dependency on previously encoded sub-pictures in the same sub-picture layer; however, an encoded sub-picture cannot have any dependencies on encoded images in another sub-picture layer.
[0176] In the same or another embodiment, an encoded sub-picture in one of the layers can be decoded dependently, with any dependencies, analysis or decoding, on encoded sub-pictures in other layers for the same local area. A sub-picture layer that can be decoded dependently, with references to other sub-picture layers of the same local area, is called a dependent sub-picture layer. An encoded sub-picture in a dependent sub-picture layer can reference an encoded sub-picture belonging to the same sub-picture, a previously encoded sub-picture in the same sub-picture layer, or both of these reference sub-pictures.
[0177] In this or an alternative embodiment, a coded sub-picture consists of one or more independent sub-picture layers and one or more dependent sub-picture layers. However, at least one independent sub-picture layer must be present in each coded sub-picture. The layer identifier (layer_id), which may be located in the NAL unit header or other high-level syntax structure, may be equal to 0 for an independent sub-picture layer. A sub-picture layer with a layer_id of 0 may be the base sub-picture layer.
[0178] In this or an alternative embodiment, a picture may consist of one or more background subpictures and one foreground subpicture. The area covered by a background subpicture may be equal to the picture area. The area occupied by a foreground subpicture may overlap the area covered by a background subpicture. A background subpicture may be a base subpicture layer, while a foreground subpicture may be a non-base (refinement) subpicture layer. One or more non-base subpicture layers may reference the same base subpicture layer for decoding. Each non-base subpicture layer with a layer_id equal to a may reference a non-base subpicture layer with a layer_id equal to b, where a is greater than b.
[0179] In this or an alternative embodiment, an image may consist of one or more background subimages, with or without a foreground subimage. Each subimage may have its own base subimage layer and one or more non-base (refinement) layers. Each base subimage layer may be referenced by one or more non-base subimage layers. Each non-base subimage layer with a layer_id equal to a may reference a non-base subimage layer with a layer_id equal to b, where a is greater than b.
[0180] In the same or an alternative embodiment, a picture may consist of one or more background subpictures, with or without a foreground subpicture. Each encoded subpicture in a (base or non-base) subpicture layer may be referenced by one or more non-base subpicture layers belonging to the same subpicture, and one or more non-base subpicture layers not belonging to the same subpicture.
[0181] In this or an alternative embodiment, the image may consist of one or more background sub-images, with or without a foreground sub-image. A sub-image in one of the layers may be further split into multiple sub-images in the same layer. One or more encoded sub-images in layer b may reference a split sub-image in layer a.
[0182] In this or an alternative embodiment, a coded video sequence (CVS) may be a group of coded pictures. A CVS may consist of one or more coded sub-picture sequences (CSPS), where a CSPS may be a group of coded sub-pictures covering the same local image region. A CSPS may have the same or different temporal resolution as the coded video sequence.
[0183] In this or an alternative embodiment, a CSPS may be encoded and included in one or more layers. A CSPS may consist of one or more CSPS layers. Decoding one or more CSPS layers corresponding to the CSPS allows for the reconstruction of a sequence of sub-images corresponding to the same local region.
[0184] In the same or an alternative embodiment, the number of CSPS levels corresponding to a CSPS may be equal to or different from the number of CSPS levels corresponding to another CSPS.
[0185] In this or an alternative embodiment, a CSPS layer may have a temporal resolution (e.g., frame rate) different from another CSPS layer. The original (uncompressed) sub-picture sequence may be temporally resampled (e.g., upsampled or downsampled), encoded using different temporal resolution parameters, and included in the bitstream corresponding to one of the layers.
[0186] In the same or an alternative embodiment, a sub-picture sequence with a frame rate of F may be encoded and included in the encoded bitstream corresponding to level 0, while a sub-picture sequence with increased (or decreased) temporal resolution, compared to the original sub-picture sequence, with F*S t,k , can be encoded and included in the encoded bitstream corresponding to level k, where S t,k - the temporal resampling coefficient for level k. If the value of S t,k greater than 1, the temporal resampling procedure can be a frame rate upconversion. Whereas if the value of S t,k less than 1, the temporal resampling procedure may be a frame rate reduction conversion.
[0187] In this or an alternative embodiment, when a sub-picture in a CSPS layer is referenced by a sub-picture from layer b, for motion compensation or any inter-layer prediction, if the spatial resolution of this CSPS layer differs from the spatial resolution of CSPS layer b, then the decoded pixels in the CSPS layer are resampled and used as reference pixels. The resampling procedure may require up- or down-resampling filtering.
[0188] Fig. 11 shows an example of a video stream including a background video CSPS with a layer_id of 0 and multiple foreground CSPS layers. While an encoded subimage may consist of one or more CSPS layers, a background region that does not belong to any of the foreground CSPS layers may consist of a base layer. The base layer may include the background region and the foreground region, while a refinement CSPS layer may include the foreground region. The refinement CSPS layer may have higher visual quality than the base layer in the same region. The refinement CSPS layer may reference the reconstructed pixels and motion vectors of the base layer corresponding to the same region.
[0189] In the same or an alternative embodiment, the video bitstream corresponding to the base layer is contained in one track in the video file, while the CSPS layers corresponding to each sub-picture are contained in a separate track.
[0190] In this or an alternative embodiment, the video bitstream corresponding to a layer is contained in a single track, while the CSPS layers with the same layer_id are contained in a separate track. In this example, the track corresponding to layer k includes only the CSPS layers corresponding to layer k.
[0191] In this or an alternative embodiment, each CSPS level of each subimage is stored in a separate track. Each track may or may not have an analysis or decoding dependency on one or more other tracks.
[0192] In the same or an alternative embodiment, each track may contain bitstreams corresponding to levels i through j of the CSPS levels of all sub-images, or a subset thereof, where 0 <i=<j=<k, а k - верхний уровень CSPS.
[0193] In the same or an alternative embodiment, the image consists of corresponding media data, including one or more of the following: a depth map, an alpha channel value map, 3D geometry data, an occupancy map, etc. These time-stamped media data may be divided into one or more sub-data streams, each of which corresponds to one sub-image.
[0194] Fig. 12, according to the same or an alternative embodiment, shows an example of videoconferencing based on the multi-level sub-image method. The video stream contains one base-layer video bitstream corresponding to the background image and one or more refinement-layer video bitstreams corresponding to the foreground sub-images. Each refinement-layer video bitstream may correspond to a CSPS level. By default, the display shows an image corresponding to the base layer. It contains images of one or more users placed within the image using the picture-in-picture (PIP) method. When the client selects a specific user, the refinement-layer CSPS corresponding to the selected user and having an increased quality or spatial resolution is decoded and displayed.
[0195] Fig. 13 shows a flow chart of a procedure for decoding and displaying a video bitstream including multi-layer sub-pictures according to one embodiment. For example, the procedure may include one or more of the operations described below. For example, in operation 1301, a video bitstream with a plurality of layers may be decoded. Operation 1302 may include identifying a background region and one or more background sub-pictures. Operation 1303 may include determining whether a region of a specific sub-picture is selected. Operation 1304 may include decoding and displaying the refined sub-picture if the region of the specific sub-picture has been selected (i.e., 1303=Yes). Operation 1305 may include decoding and displaying the background region if the region of the specific sub-picture has not been selected (i.e., 1303=No).
[0196] In this or an alternative embodiment, network intermediate equipment (e.g., a router) may select a subset of layers to transmit to a user based on their bandwidth. To adapt to bandwidth, images may be split into sub-images. For example, if the user lacks bandwidth, the router may discard layers or select sub-images based on their importance or the applicable configuration, and this may be done dynamically to adapt to the user's bandwidth.
[0197] Fig. 14 shows an application scenario corresponding to 360-degree video. When a spherical 360-degree image is projected onto a flat image, the projected 360-degree image can be split into multiple sub-images as a base layer. A refinement layer for a specific sub-image, for example, the front sub-image, can be encoded and transmitted to the client. A decoder can decode both the base layer, which includes all sub-images, and the refinement layer of a selected sub-image. When the current viewing port is identical to the selected sub-image, the displayed image can have a higher quality due to the decoded sub-image in the refinement layer. Otherwise, the decoded image of the base layer, which has lower quality, can be displayed.
[0198] In this or an alternative embodiment, display arrangement information may be present in the file as auxiliary information (e.g., an SEI message or metadata). One or more decoded sub-pictures may be moved and displayed depending on the signaled arrangement information. The arrangement information may be signaled by a streaming server or a broadcast station, or may be re-generated by a network entity or a cloud server, or may be determined using user-selectable settings.
[0199] In one embodiment, when the input image is divided into one or more (rectangular) subregions, each such subregion may be encoded as an independent layer. Each independent layer corresponding to a local region may have a unique layer_id value. For each independent layer, the size of the sub-image and the location information of the sub-image may be signaled. For example, the size of the image (width, height), the offset information relative to the upper left corner (x_offset, y_offset). Fig. 15 shows an example of an arrangement of sub-images, the corresponding size and position information of the sub-images, and the corresponding image prediction structure. The arrangement information, including the sizes and positions of the sub-images, may be signaled in a high-level syntactic structure, for example, parameter sets, a slice or tile group header, or in an SEI message.
[0200] In this embodiment, each sub-picture corresponding to an independent layer may have a unique POC value within the access block. When a reference picture among the pictures stored in the DPB is referenced using a syntax element(s) in the RPS or RPL structure, the POC value(s) of each sub-picture corresponding to one of the layers may be used.
[0201] In this or an alternative embodiment, to indicate the (inter-layer) prediction structure, the layer_id may not be used, but the value of the (difference) of the POC may be used.
[0202] In this embodiment, a sub-image with a POC value of N corresponding to a level (or local region) may (but need not) be used as a reference image for a sub-image with a POC value of N+K corresponding to the same level (or the same local region) for motion-compensated prediction. In most cases, the value of K may be equal to the maximum number of (independent) levels, which can coincide with the number of sub-regions.
[0203] Fig. 16 illustrates an extended case of Fig. 15, according to the same or an alternative embodiment. When the input image is divided into a plurality (e.g., four) subregions, each local region can be encoded using one or more layers. In such a case, the number of independent layers may be equal to the number of subregions, and each subregion will correspond to one or more layers. Thus, each subregion can be encoded using one or more independent layers and zero or more dependent layers.
[0204] In this embodiment, according to Fig. 16, the input image can be divided into four sub-regions. The upper right sub-region can be encoded as two layers, layer 1 and layer 4, while the lower right sub-region can be encoded as two layers, layer 3 and layer 5. In this case, layer 4 can refer to layer 1 for motion-compensated prediction, while layer 5 can refer to layer 3 for motion-compensated prediction.
[0205] In this or an alternative embodiment, in-loop filtering (e.g., deblocking filtering, adaptive in-loop filtering, shape restoration, bilateral filtering, or any deep learning-based filtering) crossing layer boundaries may be (optionally) prohibited.
[0206] In this or an alternative embodiment, motion-compensated prediction or intra-block copying that crosses block boundaries may be (optionally) disabled.
[0207] In this or an alternative embodiment, boundary padding or in-loop filtering at sub-image boundaries may be optionally performed. A flag indicating whether or not to perform boundary padding may be signaled in a high-level syntax structure, such as a parameter set (or sets) (VPS, SPS, PPS, or APS), a slice or tile group header, or an SEI message.
[0208] In this or an alternative embodiment, the sub-region (or sub-picture) layout information may be signaled in the VPS or SPS. Fig. 17 shows examples of syntax elements in the VPS and SPS sets. In this example, the vps_sub_picture_dividing_flag flag is signaled in the VPS. This flag may indicate whether the input image (or images) is divided into multiple sub-regions or not.
[0209] When the vps_sub_picture_dividing_flag flag is 0, the input image (or images) in the encoded video sequence (or sequences) corresponding to the current VPS set cannot be divided into multiple subregions. In this case, the input image size may be equal to the encoded image size (pic_width_in_luma_samples, pic_height_in_luma_samples), which is signaled in the SPS set.
[0210] When the value of the vps_sub_picture_dividing_flag flag is 1, the input image (or images) can be divided into multiple subregions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples signal in the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples can be equal to the width and height of the input image (or images), respectively.
[0211] In this embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may not be used for decoding, but may be used for layout and display.
[0212] In this embodiment, when the value of the vps_sub_picture_dividing_flag flag is 1, the syntax elements Pic_offset_x and Pic_offset_y may be signaled in the SPS corresponding to a specific level (or levels). In this case, the coded picture size (Pic_width_in_luma_samples, Pic_height_in_luma_samples) signaled in the SPS may be equal to the width and height of the sub-region corresponding to a specified level. The position (pic_offset_x, pic_offset_y) of the upper-left corner of the sub-region may also be signaled in the SPS.
[0213] In this embodiment, the information (pic_offset_x, pic_offset_y) about the position of the upper left corner of the sub-region may not be used for decoding, but may be used for layout and display.
[0214] In this embodiment, the layout information (size and position) of all sub-regions or a subset of sub-regions of the input image (or images) and the dependency information between layers may be signaled in a parameter set or in an SEI message.
[0215] Fig. 18 shows examples of syntax elements for indicating information about the arrangement of subregions, about dependencies between layers, and about the relationship between subregions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) subregions in the current coded video sequence. According to one embodiment, the syntax element num_layers may indicate the number of layers in the current coded video sequence. The value of num_layers may be greater than or equal to the value of num_sub_region. When all subregions are coded as single layers, the value of num_layers may be equal to the value of num_sub_region. When one or more subregions are coded as a plurality of layers, the value of num_layers must be greater than the value of num_sub_region. The syntax element direct_dependency_flag[i][j] indicates the dependence of the i-th layer on the j-th layer.The num_layers_for_region[i] value specifies the number of layers belonging to the i-th subregion. The sub_region_layer_id[i][j] value specifies the layer_id of the j-th layer belonging to the i-th subregion. The sub_region_offset_x[i] and sub_region_offset_y[i] values specify the horizontal and vertical locations of the upper-left corner of the i-th subregion, respectively. The sub_region_width[i] and sub_region_height[i] values specify the width and height of the i-th subregion, respectively.
[0216] In one embodiment, one or more syntax elements that define a set of output layers for indicating one or more layers for output, with or without information on the level, tier, and profile of the standard, may be signaled in a high-level syntax structure, for example, VPS, DPS, SPS, PPS, APS, or in an SEI message. According to the illustration of Fig. 19, the syntax element num_output_layer_sets, indicating the number of output layer sets (OLS) in the coded video sequence referring to the VPS, may be signaled in the VPS. For each set of output layers, the flag output_layer_flag may be signaled as many times as there are output layers.
[0217] In this embodiment, the output_layer_flag[i] flag, equal to 1, specifies that the i-th layer should be output. The vps_output_layer_flag[i] flag, equal to 0, specifies that the i-th layer should not be output.
[0218] In the same or an alternative embodiment, one or more syntax elements that define the level, tier, and profile information for each set of output levels may be signaled in a high-level syntax structure, such as a VPS, DPS, SPS, PPS, APS, or in an SEI message. Also, in accordance with the illustration of Fig. 19, a num_profile_tile_level syntax element indicating the number of the level, tier, and profile information for each OLS in the coded video sequence referencing the VPS may be signaled in the VPS. For each set of output levels, a set of syntax elements for the level, tier, and profile information or a pointer pointing to the level, tier, and profile information among entries in the level, tier, and profile information may be signaled (as many times as there are output levels).
[0219] In this embodiment, the value of profile_tier_level_idx[i][j] specifies a pointer, in the list of profile_tier_level() syntax structures in the VPS, to the profile_tier_level() syntax structure applicable to the j-th level of the i-th OLS set.
[0220] In the same or an alternative embodiment, as illustrated in Fig. 20, the num_profile_tile_level and / or num_output_layer_sets syntax element may be signaled when the maximum number of layers is greater than 1 (vps_max_layers_minus 1>0).
[0221] In the same or an alternative embodiment, as illustrated in Fig. 20, a syntax element vps_output_layers_mode[i] indicating the output layer signaling mode for the i-th output layer may be present in the VPS.
[0222] In this embodiment, a vps_output_layers_mode[i] value of 0 specifies that only the top layer is output using the i-th set of output layers. A vps_output_layer_mode[i] value of 1 specifies that all layers are output using the i-th set of output layers. A vps_output_layer_mode[i] value of 2 specifies that the output layers are the layers with the vps_output_layer_flag[i][j] flag set to 1, using the i-th set of output layers. More values may be reserved.
[0223] In this same embodiment, output_layer_flag[i][j] may be signaled or not signaled depending on the value of vps_output_layers_mode[i] for the i-th set of output layers.
[0224] In the same or an alternative embodiment, as illustrated in Fig. 20, the flag vps_ptl_signal_flag[i] may be present for the i-th set of output levels. Depending on the value of vps_ptl_signal_flag[i], the level, tier, and profile information of the standard for the i-th set of output levels may or may not be signaled.
[0225] In the same or an alternative embodiment, as illustrated in Fig. 21, the number of subpictures, max_subpics_minus1, in the current CVS may be signaled in a high-level syntax structure, such as VPS, DPS, SPS, PPS, APS, or in an SEI message.
[0226] In the same embodiment, according to the illustration of Fig. 21, the sub-picture identifier, sub_pic_id[i], for the i-th sub-picture may be signaled when the number of sub-pictures is greater than 1 (max_subpics_minus1>0).
[0227] In the same or an alternative embodiment, one or more syntax elements indicating the identifier of a sub-picture belonging to each layer of each set of output layers may be signaled in the VPS. According to the illustration of Fig. 22, the value sub_pic_id_layer[i][j][k] indicates that the k-th sub-picture is present in the j-th layer of the i-th set of output layers. Using this information, the decoder can recognize which of the sub-pictures can be decoded and output for each layer of a given set of output layers.
[0228] In one embodiment, a picture header (PH) is a syntactic structure that contains syntactic elements applicable to all slices of a coded picture. A picture unit (PU) is a set of NAL units that are linked according to a specific classification rule, follow each other in decoding order, and contain exactly one coded picture. A picture unit may comprise a picture header (PH) and one or more coded slice NAL units or VCL NAL units that make up the coded picture.
[0229] In one embodiment, the SPS set (RBSP) may be available in the decoding procedure before it is referenced, it may be included in at least one access unit with a TemporalId equal to 0, or provided by external means.
[0230] In one embodiment, the SPS set (RBSP) may be available in the decoding procedure before it is referenced, it may be included in at least one access unit with a TemporalId equal to 0 in a CVS that contains one or more PPS sets referencing this SPS, or provided by external means.
[0231] In one embodiment, the SPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more PPS sets, it may be included in at least one image block with a Nuh_layer_id equal to the smallest Nuh_layer_id value among the PPS NAL units referencing the SPS NAL unit in the CVS that contains one or more PPS sets referencing this SPS, or provided by external means.
[0232] In one embodiment, the SPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more PPS sets, it may be included in at least one image block with a TemporalId equal to 0 and a Nuh_layer_id equal to the smallest Nuh_layer_id value among the PPS NAL units referencing the SPS NAL unit in the CVS that contains one or more PPS sets referencing this SPS, or provided by external means.
[0233] In one embodiment, the SPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more PPS sets, it may be included in at least one image block with a TemporalId equal to 0 and with a Nuh_layer_id equal to the smallest Nuh_layer_id value among the PPS NAL units referencing the SPS NAL unit in the CVS that contains one or more PPS sets referencing this SPS, or provided by external means.
[0234] In this or an alternative embodiment, the value of pps_seq_parameter_set_id determines the value of sps_seq_parameter_set_id for the referenced SPS set. The value of pps_seq_parameter set_id may be the same in all SPS sets referenced by encoded images in CLVS.
[0235] In this or an alternative embodiment, all SPS NAL units that have the same sps_seq_parameter_set_id value in CVS may have the same content.
[0236] In this or an alternative embodiment, regardless of the nuh_layer_id values, NALSPS blocks may share a common sps_seq_parameter_set_id value space.
[0237] In the same or an alternative embodiment, the nuh layer id value in the SPS NAL unit may be equal to the smallest nuh_layer_id value among the PPS NAL units that reference the given SPS NAL unit.
[0238] In one embodiment, when an SPS set with nuh_layer_id equal to m is referenced in one or more PPS sets with nuh_layer_id equal to n, a layer with nuh_layer_id equal to m may be the same as a layer with nuh_layer_id equal to n, or a (direct or indirect) reference layer for a layer with nuh_laye_id equal to m.
[0239] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before it is referenced, it may be included in at least one access unit with a TemporalId equal to the TemporalId of the PPS NAL unit, or provided by external means.
[0240] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before referencing it, it may be included in at least one access unit with a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS that contains one or more sets of picture headers, RNs (or coded slice NAL units) referencing this PPS, or provided by external means.
[0241] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more picture headers, PH (or coded slice NAL units), it may be included in at least one picture block with a nuh_layer_id equal to the smallest nuh_layer_id value among the coded slice NAL units referencing the PPS NAL unit in the CVS that contains one or more sets of picture headers, PH (or coded slice NAL units) referencing this PPS, or provided by external means.
[0242] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more picture headers, PHs (or coded slice NAL units), it may be included in at least one picture block with a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the smallest nuh_layer_id value among the coded slice NAL units referencing the PPS NAL unit in the CVS that contains one or more sets of picture headers, PHs (or coded slice NAL units) referencing this PPS, or provided by external means.
[0243] In this or an alternative embodiment, the ph_pic_parameter_set_id value in the image header, PH, determines the pps_pic_parameter_set_id value for the referenced PPS used. The pps_seq_parameter_set_id value may be the same in all PPS sets referenced by encoded images in CLVS.
[0244] In this or an alternative embodiment, all PPS NAL units with the same pps_pic_parameter_set_id value within a picture unit must have the same content.
[0245] In this or an alternative embodiment, regardless of the nuh_layer_id values, PPS NAL units may share a common pps_seq_parameter_set_id value space.
[0246] In the same or an alternative embodiment, the nuh_layer_id value in the SPS NAL unit may be equal to the smallest nuh_layer_id value among the coded slice NAL units referencing the given PPS NAL unit.
[0247] In one embodiment, when a PPS set with nuh_layer_id equal to m is referenced in one or more NAL units of coded slices with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n, or a (direct or indirect) reference layer for the layer with nuh_layer_id equal to m.
[0248] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before it is referenced, it may be included in at least one access unit with a TemporalId equal to the TemporalId of the PPS NAL unit, or provided by external means.
[0249] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before referencing it, it may be included in at least one access unit with a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS that contains one or more sets of picture headers, RNs (or coded slice NAL units) referencing this PPS, or provided by external means.
[0250] In one embodiment, the PPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more picture headers, PH (or coded slice NAL units), it may be included in at least one picture block with a nuh_layer_id equal to the smallest nuh_layer_id value among the coded slice NAL units referencing the PPS NAL unit in the CVS that contains one or more sets of picture headers, PH (or coded slice NAL units) referencing this PPS, or provided by external means.
[0251] In one embodiment, a PPS set (RBSP) may be available in the decoding procedure before it is referenced in one or more picture headers, PHs (or coded slice NAL units), it may be included in at least one picture block with a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the smallest nuh_layer_id value among the coded slice NAL units referencing the PPS NAL unit in a CVS that contains one or more sets of picture headers, PHs (or coded slice NAL units) referencing this PPS, or provided by external means.
[0252] In this or an alternative embodiment, the ph_pic_parameter_set_id value in the image header, PH, determines the pps_pic_parameter_set_id value for the referenced PPS used. The pps_seq_parameter_set_id value may be the same in all PPS sets referenced by encoded images in CLVS.
[0253] In this or an alternative embodiment, all PPS NAL units with the same pps_pic_parameter_set_id value within a picture unit must have the same content.
[0254] In this or an alternative embodiment, regardless of the nuh_layer_id values, PPS NAL units may share a common pps_seq_parameter_set_id value space.
[0255] In the same or an alternative embodiment, the nuh_layer_id value in the SPS NAL unit may be equal to the smallest nuh_layer_id value among the coded slice NAL units referencing the given PPS NAL unit.
[0256] In one embodiment, when a PPS set with nuh_layer_id equal to m is referenced in one or more NAL units of coded slices with nuh_layer_id equal to n, the layer with nuh_layer_id equal to m may be the same as the layer with nuh_layer_id equal to n, or a (direct or indirect) reference layer for the layer with nuh_layer_id equal to m.
[0257] An output level is a level from an output level set that is to be output. An output level set (OLS) is a set of levels consisting of a specified set of levels, in which one or more levels are specified as output levels. The index of a level in an output level set (OLS) is the index of the level in the OLS level list.
[0258] A sublayer is a temporal scale layer or a temporal scale bitstream consisting of VCL NAL units with a specific TemporalId variable value, as well as corresponding non-VLC NAL units. A sublayer representation is a subset of the bitstream consisting of the NAL units of a particular sublayer and the sublayers below it.
[0259] The VPS RBSP sequence may be accessible in the decoding procedure before being referenced; it may be included in at least one access unit with a TemporalId of 0 or provided externally. All VPS NAL units with a given vps_video_parameter_set_id value in the CVS may have the same content.
[0260] The vps_video_parameter_set_id value contains the VPS ID for reference from other syntax elements. The vps_video_parameter_set_id value may be greater than 0.
[0261] The value of vps_max_layers_minus 1 plus 1 defines the maximum number of layers allowed in each CVS referencing this VPS.
[0262] The vps_max_sublayers_minus1 value plus 1 defines the maximum number of temporary sublayers that can be contained in a layer in each CVS referencing this VPS. The vps_max_sublayers_minus 1 value can range from 0 to 6, inclusive.
[0263] The vps_all_layers_same_num_sublayers_flag flag, equal to 1, indicates that the number of temporary sublayers is the same for all layers in every CVS referencing this VPS.
[0264] The vps_all_layers_same_num_sublayers_flag flag, set to 0, indicates that layers in each CVS referencing this VPS do not necessarily have the same number of temporary sublayers. If this flag is absent, the value of vps_all_layers_same_num_sublayers_flag is set to 1.
[0265] The vps_all_independent_layers_flag flag, set to 1, specifies that all layers in CVS are encoded independently, without using inter-layer prediction.
[0266] The vps_all_independent_layers_flag flag, when set to 0, specifies that inter-layer prediction can be used for one or more layers in CVS. When this flag is absent, the value of vps_all_independent_layers_flag is set to 1.
[0267] The value of vps_layer_id[i] defines the value of nuh_layer_id of the i-th layer. For any two non-negative values m and n, when m is less than n, the value of vps_layer_id[m] may be less than vps_layer_id[n].
[0268] The flag, vps_independent_layer_flag[i], equal to 1, specifies that inter-layer prediction is not used for the layer with index i. The flag, vps_independent_layer_flag[i], equal to 0, specifies that inter-layer prediction and syntax elements can be used for the layer with index i.
[0269] The VPS contains the vps_direct_ref_layer_flag[i][j] flags for j in the range from i-1 inclusive. If this flag is absent, the value of vps_independent_layer_flag[i] is set to 1. A vps_direct_ref_layer_flag[i][j] value of 0 indicates that the layer with ordinal number j is not a direct reference layer for the layer with ordinal number i. A vps_direct_ref_layer_flag[i][j] value of 1 indicates that the layer with ordinal number j is a direct reference layer for the layer with ordinal number i.
[0270] When the value of vps_direct_ref_layer_flag[i][j] is absent for i and j in the range from 0 to vps_max_layers_minus1 inclusive, its value is taken to be 0. When the flag vps_independent_layer_flag[i] is 0, at least one value of j in the range from 0 to i 1 inclusive may be present, that is, the value of vps_direct_ref_layer_flag[i][j] is 1. The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are calculated as follows:
[0271] The GeneralLayerIdx[i] variable, which defines the ordinal number of the layer with nuh_layer_id equal to vps_layer_id[i], is calculated as follows:
[0272] For any two distinct values of i and j, both in the range from 0 to vps_max_layers_minus1 inclusive, when DependencyFlag[i][j] is 1, it is a bitstream compatibility requirement that the values of chroma_format_idc and bit_depth_minus8 applicable to the i-th layer may be equal to the values of chroma_format_idc and bit_depth_minus8, respectively, applicable to the j-th layer.
[0273] A max_tid_ref_present_flag[i] value of 1 indicates that the max_tid_il_ref_pics_plus1[i] syntax element is present. A max_tid_ref_present_flag[i] value of 0 indicates that the max_tid_il_ref_pics_plus1[i] syntax element is not present.
[0274] A max_tid_il_ref_pics_plus1[i] value of 0 specifies that inter-layer prediction is not used for non-IRAP images of the i-th level. A max_tid_il_refpics_plus 1[i] value greater than 0 specifies that images with a Temporalid greater than max_tid_il_ref_pics_plus1[i] - 1 cannot be used as ILRP for decoding images of the i-th level. If this value is absent, the value of max_tid_il_ref_pics_plus1[i] is taken to be 7.
[0275] The each_layer_is_an_ols_flag flag, equal to 1, specifies that each OLS set contains only one layer, and each level in the CVS referencing this VPS is an OLS set with a single contained layer, which is the only output layer. The each_layer_is_an_ols_flag flag, equal to 0, specifies that an OLS set can contain more than one layer. If vps_max_layers_minus1 is 0, the value of each_layer_is_an_ols_flag is set to 1. Otherwise, when vps_all_independent_layers_flag is 0, the value of each_layer_is_an_ols_flag is set to 0.
[0276] The ols mode idc value of 0 specifies that the total number of OLS sets specified by the VPS set is vps_max_layers_minus1+1, where the i-th OLS set contains layers with ordinal numbers from 0 to i inclusive, and for each OLS, only the highest layer in it is the output layer.
[0277] The ols mode idc value of 1 specifies that the total number of OLS sets specified by the VPS set is vps_max_layers_minus1+1, where the i-th OLS set contains layers with ordinal numbers from 0 to i inclusive, and for each OLS, all layers in it are output layers.
[0278] The ols_mode_idc value of 2 specifies that the total number of OLS sets specified by the VPS set are signaled explicitly, and for each OLS, the output levels are signaled explicitly, with the remaining levels being direct or indirect reference levels for the OLS output levels.
[0279] The ols_mode_idc value can range from 0 to 2 inclusive. The ols_mode_idc value of 3 is reserved by ITU-T / ISO / IEC for future use.
[0280] When the vps_all_independent_layers_flag flag is 1 and the each_layer_is_an_ols_flag flag is 0, the value of ols_mode_idc is set to 2.
[0281] The value of num_output_layer_sets_minus 1 plus 1 defines the total number of OLSs specified by the VPS when ols_mode_idc is 2.
[0282] The TotalNumOlss variable, which defines the total number of OLS specified by the VPS, is calculated as follows:
[0283] The vps_all_layers_same_num_sublayers_flag flag, set to 0, indicates that layers in each CVS referencing this VPS do not necessarily have the same number of temporary sublayers. If this flag is absent, the value of vps_all_layers_same_num_sublayers_flag is set to 1.
[0284] The vps_all_independent_layers_flag flag, set to 1, specifies that all layers in CVS are encoded independently, without using inter-layer prediction.
[0285] The ols_output_layer_flag[i][j] flag, when set to 1, specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the ith OLS set when ols_mode_idc is 2. The ols-output-layer-flag[i][j] flag, when set to 0, specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the ith OLS set when ols_mode_idc is 2.
[0286] The variable NumOutputLayersInOls[i] defining the number of output layers in the i-th OLS set, the variable NumSubLayersInLayerInOLS[i][j] defining the number of sublayers in the j-th layer of the i-th OLS set, the variable OutputLayerIdInOls[i][j] defining the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] defining whether the k-th layer is used as an output layer in at least one OLS set are calculated as follows:
[0287] For each value of i in the range from 0 to vps_max_layers_minus1, inclusive, both LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] may be 0 simultaneously. In other words, there cannot be a layer that is neither an output layer of at least one OLS set nor a reference layer of some other layer.
[0288] Each OLS may contain at least one output layer. In other words, for any value of i in the range from 0 to TotalNumOlss - 1, inclusive, the value of NumOutputLayersInOls[i] may be greater than or equal to 1.
[0289] The variable NumLayersInOls[i], which defines the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j], which defines the nuh_layer_id value of the j-th layer in the i-th OLS set, are calculated as follows:
[0290] The variable OlsLayerIdx[i][j], which defines the ordinal number of the OLS layer with nuh_layer_id equal to LayerIdInOls[i][j], is calculated as follows:
[0291] The lowest layer in each OLS set can be an independent layer. That is, for each i in the range from 0 to TotalNumOlss - 1 inclusive, the value of vps_independent_layer_flagf GeneralLayerIdx[LayerIdInOls[i][0]]] must be equal to 1.
[0292] Each layer must be a part of at least one OLS defined by the VPS set. In other words, for each layer with a particular nuh_layer_id, nuhLayerId value equal to one of vps_layer_id[k] for k in the range from 0 to vps_max_layers_minus1 inclusive, there can be at least one pair of values i and j, where i lies in the range from 0 to TotalNumOlss - 1 inclusive, and aj lies in the range NumLayersInOls[i] - 1 inclusive, such that the value LayerIdInOls[i][j] is equal to nuhLayerId.
[0293] In one embodiment, the decoding procedure of the current CurrPic picture is performed as follows: - PictureOutputFlag is set as follows: - if one of the following conditions is met ("true"), the PictureOutputFlag is set to 0: - the current picture is a RASL picture, and the NoOutputBeforeRecoveryFlag of the corresponding IRAP picture is 1. - gdr_enabled_flag is 1 and the current picture is a GDR picture with the NoOutputBeforeRecoveryFlag equal to 1. - gdr_enabled_flag is 1, the current picture is a GDR picture with the NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the corresponding GDR picture. - sps_video_parameter_set_id is 0, ols_mode_idc is 0 and the current access unit, AU, contains image picA that satisfies all of the following conditions: - PicA has PictureOutputFlag equal to 1.
[0000] is equal to nuhLid).- sps_video_parameter_set_id is greater than 0, ols_mode_idc is 2 and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is 0.- Otherwise, PictureOutputFlag is set equal to pic_output_flag.
[0294] After all slices of the current picture are decoded, it is marked as "used as short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "used as short-term reference".
[0295] In this or an alternative embodiment, when each layer is a set of output layers, PictureOutputFlag is set equal to picoutputflag, regardless of the value of ols_mode_idc.
[0296] In this or an alternative embodiment, PictureOutputFlag is set to 0 when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is 0, ols_mode_idc is 0, and the current access unit, AU, contains a picture picA that satisfies all of the following conditions: PicA has PictureOutputFlag equal to 1, PicA has a nuh_layer_id, nuhLid greater than that of the current picture, and PicA belongs to an output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0297] In this or an alternative embodiment, PictureOutputFlag is set to 0 when sps_video_parameter_set_id is greater than 0, each_layer_is_an_ols_flag is 0, ols_mode_idc is 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is 0.
[0298] Fig. 23 shows an example of a syntax table for specifying an output level set using the output level set mode.
[0299] Fig. 24 shows a flow chart of a bitstream decoding procedure according to one aspect of the present invention. Namely, Fig. 24 shows a flow chart of an algorithm on the decoder side for specifying an output level set using an output level set mode, according to one embodiment.
[0300] According to one aspect of the present invention, a decoding method may include: receiving a bitstream including compressed video / image data (operation 1001 in Fig. 24). The bitstream may have multiple layers.
[0301] The decoding method may also include an operation 1002 of analyzing or deriving, from the bitstream, an output level set mode indicator (e.g., ols_mode_idc) in a video parameter set (VPS).
[0302] The decoding method may also include an operation 1003 of identifying the output level set signaling based on the output level set mode indicator.
[0303] The decoding method may also include an operation 1004 (for example, operations 1004A, 1004B or 1004C) of identifying one or more image output levels based on the identified signaling of the set of output levels.
[0304] The decoding method may also include an operation 1005 of decoding one or more identified image output levels. The decoded one or more image output levels can be displayed.
[0305] Identifying the signaling of a set of output levels based on the output level set mode indicator may include: in the case where the output level set mode indicator in the VPS has a first value, determining that the uppermost level in the bit stream is said one or more image output levels (see, for example, Fig. 25A); in the case where the output level set mode indicator in the VPS has a second value, determining that all levels in the bit stream are said one or more image output levels (see, for example, Fig. 25B); and in the case where the output level set mode indicator in the VPS has a third value, identifying one or more image output levels based on explicit signaling in the VPS (see, for example, Fig. 25C).
[0306] In Fig. 25A-25C, the output image is underlined. According to Fig. 25A-25C, there may be five levels in the bitstream, and only some of them are output for display. For example, level 3 may be the output (see, for example, Fig. 25C).
[0307] According to one embodiment of the present invention, the ols_mode_idc indicator in the VPS can indicate the method (mechanism) for signaling a set of output levels. For example, if it is 0, the topmost level in the bitstream can be the only output level; if it is 1, all levels in the bitstream can be output levels; and if it is 2, one or more output levels can be explicitly signaled in the VPS. That is, output images can be determined by signaling a set of output levels.
[0308] For example, the output of images of each level can be determined by signaling a set of output levels, and the signaling method of the output levels can be determined by ols_mode_idc.
[0309] According to one embodiment of the present invention, an output level mode may be signaled in each bit stream, and the output level mode may change over time.
[0310] According to one embodiment of the present invention shown in Fig. 25A, 5 levels (level 4 level 0) may be present in the bitstream. According to the illustration of Fig. 25A, at the time K, when the output level set mode indicator (ols_mode_idc)=0, the uppermost level may be output. That is, according to the illustration of Fig. 25A, at the time K, the output image of level 4 is output. At the time K+1, level 3, which is the uppermost level, may be output, and so on for the times K+2 and K+3.
[0311] According to the illustration of Fig. 25B, where the output level set mode indicator (ols_mode_idc) is 1, all levels can be output, for example, at all times from K to K+3.
[0312] According to the illustration of Fig. 25C, explicit signaling can be performed. For example, according to the illustration of Fig. 26, image level 3 (which is shown as the output underlined) is the output level, for example, at times from K to K+3.
[0313] The first value may be different from the second value and may be different from the third value, and the second value may be different from the third value.
[0314] The first value may be 0, the second value may be 1, and the third value may be 2. However, other values may be used, without limiting the present invention to the use of the values 0, 1, and 2 described above.
[0315] Identifying one or more image output levels by explicit signaling in a VPS may include: (i) obtaining from the VPS, by analysis or deduction, an output level flag, and (ii) assigning levels whose output level flag is equal to 1 to said one or more image output levels.
[0316] The output level set signaling identification based on the output level set mode indicator may include: in the case where the output level set mode indicator in the VPS has a predetermined value, the output level set signaling includes identifying one or more image output levels based on explicit signaling in the VPS.
[0317] Identifying one or more image output levels by explicit signaling in a VPS includes: (i) obtaining from the VPS, by analyzing or deducing, an output level flag, and (ii) assigning levels whose output level flag is equal to 1 to said one or more image output levels, wherein the number of levels in said plurality of levels is greater than 2.
[0318] The output level set mode signaling may include identifying one or more image output levels based on explicit signaling in the VPS when the output level set mode indicator is 2 and the number of levels in said plurality of levels is greater than 2.
[0319] The output level set signaling may include determining that the topmost level in the bit stream or all levels in the bit stream are said one or more image output levels, by logically inferring one or more image output levels when the output level set mode indicator is less than 2, and the number of levels in said plurality of levels is 2, and the output level set mode indicator is actually less than 2, and the number of levels in said plurality of levels is actually 2.
[0320] According to one embodiment of the present invention, the indicator of the number of sets of output levels minus 1 in the VPS indicates the number of output levels.
[0321] According to one embodiment of the present invention, the maximum level indicator in the VPS minus 1 in the VPS indicates the number of levels in the bit stream.
[0322] According to one embodiment of the present invention, the output level set flag [i][j] indicates whether the j-th level of the i-th set of output levels is an output level or not.
[0323] According to one embodiment of the present invention, if all of said plurality of levels are independent, and the "all levels in the VPS are independent" flag in the VPS is 1, the output level set mode indicator is not signaled, and the value of the output level set mode indicator is taken to be equal to said second value.
[0324] According to one embodiment of the present invention, when each level is an output level set, the image output flag in the VPS is set equal to the image output flag signaled in the image header, regardless of the value of the output level set mode indicator.
[0325] Note: Images in the output layer may or may not have PictureOutputFlag set to 1. Images in the non-output layer have PictureOutputFlag set to 0. Images with PictureOutputFlag set to 1 are output for display. Images with PictureOutputFlag set to 0 are not output for display.
[0326] According to one embodiment of the present invention, the image output flag is assigned to 0 when the VPS identifier in the sequence parameter set (SPS) is greater than 0, which indicates: there is more than one level in the bitstream, the "each level is an output level set" flag in the VPS is 0, which indicates: not all levels among the plurality of levels in the bitstream are independent, the output level set mode indicator is 0, and the current access block contains an image that satisfies all of the following conditions, including: the presence of an image output flag equal to 1, the presence of a level identifier nuh greater than that of the current image, and belonging to an output level from the output level set.
[0327] According to one embodiment of the present invention, the image output flag in the VPS is set to 0 when the sequence parameter set (SPS) identifier in the VPS is greater than 0, the "each layer is an output layer set" flag is 0, the output layer set mode indicator is 2, and the "output layer of the output layer set" flag [target OLS indicator] [common layer indicator [NUH layer identifier]] is 0.
[0328] According to one embodiment of the present invention, the method may further include controlling the display to display the decoded one or more image output layers.
[0329] According to one aspect of the present invention, there is provided a computer-readable storage medium that stores instructions that, when executed, cause a system or device that includes one or more processors to perform the following: receiving a bit stream that includes compressed video / image data, wherein said bit stream has a plurality of levels; obtaining from the bit stream, by analyzing or deriving, an indicator of a mode of a set of output levels in a video parameter set (VPS); identifying a signaling of a set of output levels based on the mode indicator of a set of output levels; identifying one or more image output levels based on the identified signaling of a set of output levels; and decoding one or more identified image output levels.
[0330] According to one embodiment of the present invention, said instructions are also configured to ensure that a system or device including one or more processors performs the following: controlling a display to display decoded one or more image output levels.
[0331] According to one aspect of the present invention, an apparatus may include: at least one memory configured to store a computer program code; and at least one processor configured to access said at least one memory and to perform operations in accordance with the computer program code. According to one embodiment of the present invention, the computer program code may include: a receiving code configured to cause said at least one processor to receive a bitstream including compressed video / image data, wherein said bitstream has a plurality of layers; an obtaining code configured to cause said at least one processor to obtain, by analyzing or deriving, a mode indicator of a set of output layers in a video parameter set (VPS);an output level signaling identification code configured to cause said at least one processor to identify a signaling of a set of output levels based on an output level set mode indicator; an image output level identification code configured to cause said at least one processor to identify one or more image output levels based on the identified signaling of the set of output levels; and a decoding code configured to cause said at least one processor to perform the following: decoding the one or more identified image output levels.
[0332] According to one embodiment of the present invention, the computer program code may further include: a display control code configured to control, by said at least one processor, the display for displaying one or more image output levels.
[0333] The above-described methods for decoding, displaying, and signaling adaptive resolution parameters may be implemented in the form of computer software that uses machine-readable instructions and is physically stored on one or more machine-readable media. For example, Fig. 7 shows a computer system 700 suitable for implementing some of the embodiments of the present invention.
[0334] Computer software may be encoded using any suitable machine code or programming language that can be processed by assembly, compilation, linking, or similar mechanisms, resulting in code that includes instructions that are executed directly or through interpretation, microcode execution, etc., by computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0335] The instructions can be executed on computers, or computer components, of various types, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0336] The components of the computer system 700 shown in Fig. 7 are provided for exemplary purposes only and do not imply any limitations on the scope of application or functionality of the computer software implementing embodiments of the present invention. Likewise, the illustrated configuration of components should not be interpreted as having any dependency or requirement associated with any component illustrated in the example embodiment of the computer system 700, or a combination of such components.
[0337] The computer system 700 may include input devices as part of a user interface. The input devices in the user interface may respond to input from one or more users using, for example, tactile input (e.g., keystrokes, finger swipes, cyberglove movements), audio input (e.g., voice, hand claps), visual input (e.g., gestures), olfactory input (not shown in the drawing). The user interface devices may also be used to capture various types of media data not necessarily associated with conscious input from a person, such as audio (e.g., voice, music, environmental sounds), images (e.g., scanned images, photographic images obtained from a camera), video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0338] The user interface devices may include one or more of the following (only one device of each type is shown in the drawing): keyboard 701, mouse 702, trackpad 703, touch screen 710, cyber glove 704, joystick 705, microphone 706, scanner 707, camera 708.
[0339] The computer system 700 may also include user interface output devices. The user interface output devices may act on the senses of one or more users, for example, via tactile output, sound, light, and / or smell / taste. The user interface output devices may include haptic output devices (e.g., haptic feedback from a touchscreen 710, a cyber glove 704, or a joystick 705, but haptic output devices that are not input devices may also be present).For example, such devices may be audio output devices (e.g., speakers 709, headphones (not shown in the drawing)), visual output devices (e.g., screens 710, including CRT screens, LCD screens, plasma screens, OLED screens, both touch and without touch input functions, both with and without haptic feedback functions, and some of the screens may be capable of outputting two-dimensional visual information or more than three-dimensional visual information, using such means as, for example, stereographic output; virtual reality glasses (not shown in the drawing), holographic displays, smoke machines, as well as printers (not shown in the drawing).
[0340] The computer system 700 may also include storage devices accessible to users and data carriers associated therewith, such as optical media, including CD / DVD ROM / RW 720, with CD / DVD media 721 or the like, a flash drive 722, a removable hard disk or solid state disk 723, previously used magnetic media, such as tapes or floppy disks (not shown in the drawing), specialized ROM / ASIC / PLD-based devices, such as hardware keys (not shown in the drawing), etc.
[0341] Those skilled in the art will understand that the expression "machine-readable storage medium" used in connection with the present invention does not include data transmission media, carrier waves or other transient signals.
[0342] The computer system 700 may also have an interface with one or more communication networks. The networks may be, for example, wireless, wired, or optical. The networks may also be local, wide area, metropolitan, located on vehicles or industrial facilities, real-time networks, delay-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area networks of digital television broadcasting, including cable television, satellite television and terrestrial broadcasting networks, networks of vehicles and industrial facilities, including CANBus, etc. Some networks require external network interface adapters connected to data ports or general-purpose peripheral buses (749) (for example, USB ports of the computer system 700).Other networks may be integrated into the internal structure of the computer system 700 by connecting to the system bus, as described below (e.g., an Ethernet interface in a personal computer system or a cellular network interface in a smartphone-based computer system). The use of any of these networks allows the computer system 700 to communicate with other objects. Such communication may be unidirectional, receiving only (e.g., television broadcasting), unidirectional, transmitting only (e.g., a CANbus network to certain CANbus devices), or bidirectional, for example, to other computer systems using local or wide area networks. Any of the networks and network interfaces described above may employ appropriate protocols and protocol stacks.
[0343] The user interface devices, user-accessible storage devices, and network interfaces described above can be connected to the basic internal structure 740 of the computer system 700.
[0344] The basic internal structure 740 may include one or more central processing units (Central Processing Units, CPU) 741, graphic processing units (Graphics Processing Units, GPU) 742, specialized programmable data processing units in the form of electrically programmable gate matrices (Field Programmable Gate Areas, FPGA) 743, hardware accelerators 744 for specific tasks, etc. These devices, together with the memory 745 in the read-only memory (ROM) mode 945, the memory 746 with random access, the internal memory of large capacity, such as internal, not accessible to the user hard drives, SSD-drives and similar memory 747, can be combined by the system bus 748. In some computer systems, the system bus 748 can be provided with access in the form of one or more physical connectors, allowing the expansion of the system with additional CPUs, GPUs, etc.Peripheral devices can be connected either directly to the base system bus 748 or to the peripheral bus 749. Examples of peripheral bus architectures include PCI, USB, and the like.
[0345] The CPU 741, the GPU 742, the FPGA 743, and the accelerators 744 can execute instructions that, taken together, can constitute the computer code described above. The computer code can be stored in the ROM 745 or the RAM 746. Temporary data can be stored in the RAM 746, while persistent data can be stored, for example, in the internal high-capacity memory 747. High speed of storing data in memory devices and retrieving data from them can be ensured by using a cache memory, which can be closely connected with one or more of the CPU 741, the GPU 742, the high-capacity memory 747, the ROM 745, the RAM 746, etc.
[0346] Machine-readable storage media may store computer code for performing various computer-executable operations. Such media and computer code may be specially designed and manufactured for the purposes of the present invention or may be widely available and known to those skilled in the art of computer software.
[0347] As a non-limiting example, a computer system with architecture 700, and in particular, base structure 740, can provide the required functionality as a result of execution by a processor (or processors) (including CPUs, GPUs, FPGAs, accelerators, etc.), of software implemented on one or more tangible computer-readable storage media. Such computer-readable storage media can be media associated with the above-described mass storage devices that are accessible to users, or non-volatile storage devices in the base structure 740, for example, a built-in mass storage device 747 or ROM 745. Software that implements various embodiments of the present invention can be stored in such devices and executed by the internal structure 740 of the computer system.The computer-readable storage medium, depending on the specific requirements, may include one or more memory devices or memory chips. The software may ensure that the base structure 740, and in particular the processors included therein (including CPUs, GPUs, FPGAs, etc.), execute the necessary procedures or portions of the necessary procedures described in this document, including the creation of data structures stored in RAM 746 and the modification of these data structures in accordance with the procedures defined by the software. In addition or alternatively, the computer system may provide the required functionality as a result of the operation of logic hard-coded or otherwise embodied in the electrical circuit (e.g., accelerator 744), which may operate in conjunction with the software, or instead of it, to perform the required procedures or portions of the required procedures described in this document.References to software, where appropriate, may include such logic, and vice versa. References to a machine-readable medium, where appropriate, may include an electrical circuit (e.g., an integrated circuit) storing executable software, an electrical circuit implementing executable logic, or both. The scope of the present invention includes all appropriate combinations of hardware and software.
[0348] This document describes several embodiments of the invention; however, modifications, variations, and equivalent substitutions are possible and fall within the scope of the invention. Accordingly, it should be understood that those skilled in the art are capable of creating numerous systems and methods that, although not explicitly described herein, embody the intent of the invention and, accordingly, are within its scope and spirit.
Claims
1. A decoding method comprising: processing an encoded video bitstream containing multiple levels; defining, from the encoded video bitstream, the ols_mode_idc syntax element in the video parameter set (VPS); decoding one or more levels of the current image, wherein the decoding comprises: setting the PictureOutputFlag syntax element of the current image to 0 when the sps_video_parameter_set_id syntax element is greater than 0, the each_layer_is_an_ols_flag syntax element is 0, and the ols_mode_idc syntax element is 2, otherwise, setting the PictureOutputFlag syntax element of the current image to the pic_output_flag syntax element signaled in the image header; and controlling the display of one or more image output levels from the decoded one or more levels of the current image, wherein the current image level that is not the image output level has PictureOutputFlag equal to 0 and is not displayed.
2. The method according to paragraph 1, also comprising identifying the signaling of a set of output levels based on the syntax element ols_mode_idc, which includes: if the ols_mode_idc syntax element in the VPS has the first value, an identification of the highest level in the encoded video bitstream as said one or more image output levels; if the ols_mode_idc syntax element in the VPS has a second value, identifying all levels in the encoded video bitstream as said one or more image output levels; and if the ols_mode_idc syntax element in the VPS has a third value, the identification of one or more image output levels based on explicit signaling in the VPS, where the first value is different from the second value and is different from the third value, and the second value is different from the third value.
3. The method of claim 2, wherein the first value is 0, the second value is 1, and the third value is 2.
4. The method of claim 2, wherein identifying one or more image output layers by explicit signaling in the VPS comprises: (i) obtaining from the VPS, by parsing or inferring, a syntax element output_layer_flag, and (ii) designating layers that have the syntax element output_layer_flag equal to 1 as said one or more image output layers.
5. The method according to paragraph 1, also comprising identifying the signaling of the set of output levels based on the syntax element ols_mode_idc, wherein: If the ols_mode_idc syntax element in the VPS has a predetermined value, identifying the signaling of the set of output levels includes identifying one or more image output levels based on explicit signaling in the VPS.
6. The method of claim 5, wherein identifying one or more image output layers by explicit signaling in the VPS comprises: (i) obtaining from the VPS, by parsing or inferring, a syntax element output_layer_flag, and (ii) assigning layers that have the syntax element output_layer_flag equal to 1 to said one or more image output layers, wherein the number of layers in said plurality of layers is greater than 2.
7. The method according to claim 5, wherein identifying the signaling of the set of output levels includes identifying one or more image output levels based on explicit signaling in the VPS, when the syntax element ols_mode_idc is equal to 2 and the number of levels in said set of levels is greater than 2.
8. The method according to paragraph 1, also comprising: identifying a signaling of a set of output levels, which includes identifying the topmost level in the encoded video bitstream or all levels in the encoded video bitstream as said one or more image output levels by logically inferring the one or more image output levels when the syntax element ols_mode_idc is less than 2 and the number of levels in the set of levels is 2.
9. The method according to claim 8, wherein the syntax element num_output_layer_sets_minus1 in the VPS indicates the total number of output layer sets.
10. The method of claim 9, wherein the vps_max_layers_minus1 syntax element in the VPS indicates the maximum allowed number of layers in each coded video sequence (CVS) referencing the VPS.
11. The method of claim 10, wherein the syntax element ols_output_layer_flag[i][j] in the VPS indicates whether the j-th layer of the i-th set of output layers is an output layer.
12. The method according to claim 2, wherein if all of the plurality of layers in the encoded video bitstream are independent layers that do not have analysis or decoding dependencies on other layers, and the syntax element vps_all_independent_layers_flag in the VPS is equal to 1, then the syntax element ols_mode_idc is not signaled, and the value of the syntax element ols_mode_idc is taken to be equal to the mentioned second value.
13. The method according to claim 1, wherein the decoding comprises setting the syntax element PictureOutputFlag in the VPS equal to the syntax element pic_output_flag signaled in the image header, regardless of the value of the syntax element ols_mode_idc.
14. The method according to claim 1, wherein the decoding comprises setting the syntax element PictureOutputFlag in the VPS to 0 when the syntax element sps_video_parameter_set_id is greater than 0, the syntax element each_layer_is_an_ols_flag is 0, and the syntax element ols_mode_idc is 2.
15. A coding method comprising: definition of the ols_mode_idc syntax element that must be signaled in the video parameter set (VPS); encoding one or more levels of the current image based on the ols_mode_idc syntax element to form an encoded video bitstream containing a plurality of levels; wherein the encoding comprises encoding one or more image output levels from one or more levels of the current image, where the current image level, which is not the image output level, has a PictureOutputFlag syntax element equal to 0 and should not be displayed; wherein the sps_video_parameter_set_id syntax element is greater than 0, the each_layer_is_an_ols_flag syntax element is equal to 0, and the ols_mode_idc syntax element is equal to 2, indicating that the PictureOutputFlag syntax element is equal to 0; and The sps_video_parameter_set_id syntax element that is not greater than 0, the each_layer_is_an_ols_flag syntax element that is not equal to 0, and the ols_mode_idc syntax element that is not equal to 2 indicate that the PictureOutputFlag syntax element of the current image is equal to the pic_output_flag syntax element signaled in the image header.
16. A method for processing video data, comprising: forming a bit stream using the method according to paragraph 15 and transmitting the generated bit stream to the decoder or storing the generated bit stream on a machine-readable medium.