Method, computer system and computer program for alignment across layers in encoded video stream
Patent Information
- Application Number
- JP2025071975
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-10-05
- Filing Date
- 2025-04-24
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2040-10-19
AI Technical Summary
Existing video encoding and decoding technologies struggle with efficiently handling multiple semantically independent source pictures that require different adaptive resolution settings, particularly in applications like 360 coding or surveillance, where separate resolution adjustments are necessary for each scene component, leading to inefficiencies in encoding and decoding processes.
Implementing a method and system for alignment between layers in encoded video data by identifying sub-picture regions, including background and foreground regions, and selectively decoding and displaying enhanced sub-pictures based on user selection, allowing for adaptive resolution change (ARC) signaling across multiple parts of the picture.
Enhances encoding and decoding efficiency by allowing separate resolution adjustments for different scene components, improving video quality and adaptability in applications requiring diverse resolution settings, such as 360-degree video and surveillance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] This application claims priority from U.S. Provisional Patent Application No. 62 / 954,844, filed on December 30, 2019, and U.S. Patent Application No. 17 / 063,025, filed on October 5, 2020, the entireties of which are hereby incorporated by reference.
[0002] This disclosure generally relates to the field of video encoding and decoding, and more specifically, to the reference and scope of parameter sets in an encoded video stream.
Background Art
[0003] Video encoding and decoding using inter-picture prediction with motion compensation has been known for decades. Uncompressed digital video is composed of a series of pictures, each picture having spatial dimensions, for example, of 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate (also known informally as the frame rate), for example, 60 pictures per second, i.e., a picture rate of 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires storage space exceeding 600 GB.
[0004] One purpose of video encoding and decoding can be considered to be the reduction of redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth requirements or storage space requirements, sometimes by two or more orders of magnitude. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to techniques that can reconstruct an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to be useful for the intended application of the reconstructed signal. In the case of video, lossy compression is widely used. The amount of acceptable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television contribution applications. The achievable compression ratio reflects this, and higher acceptable / tolerable distortion can result in a higher compression ratio.
[0005] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy encoding, some of which are introduced below.
[0006] Historically, video encoders and decoders have tended to operate on a given picture size that was typically defined and remained constant for a coded video sequence (CVS), a Group of Pictures (GOP), or a similar multi-picture temporal frame. For example, in MPEG-2, a system design is known that varies the horizontal resolution (and thus the picture size) depending on factors such as the activity of a scene, but that is only in I pictures and thus typically only for a GOP. Resampling of reference pictures for use within a CVS at different resolutions is known, for example, from ITU-T Recommendation H.263 Annex P. However, there, only the reference pictures are resampled while the picture size remains the same, potentially only the portion of the picture canvas used (in the case of downsampling) or only the portion of the scene captured (in the case of upsampling). Also, H.263 Annex Q enables resampling of individual macroblocks by a factor of 2 in the up or down direction (in each dimension). Again, the picture size remains the same. The size of the macroblocks is fixed in H.263 and thus does not need to be signaled.
[0007] In modern video coding, changing the picture size in predicted pictures has become more mainstream. For example, VP9 enables resampling of reference pictures and changing the resolution for the entire picture. Similarly, certain proposals made towards VVC (e.g., Hendry et al., “On adaptive resolution change (ARC) for VVC”, Joint Video Team document JVET-M0135-v1, from January 9 to 19, 2019, which is hereby incorporated by reference in its entirety) enable resampling of the entire reference picture to a different resolution, either higher or lower. In that document, it is proposed that multiple different candidate resolutions be coded within the sequence parameter set and referenced by per-picture syntax elements within the picture parameter set.
Summary of the Invention
[0008] Embodiments relate to a method, system, and computer-readable medium for alignment between layers in encoded video data. According to one aspect, a method for alignment between layers in encoded video data is provided. The method may include decoding a video bitstream having a plurality of layers. One or more sub-picture regions are identified from among the plurality of layers of the decoded video bitstream, and the sub-picture region includes a background region and one or more foreground sub-picture regions. Based on a determination that a foreground sub-picture region is selected, an enhanced sub-picture is decoded and displayed. Based on a determination that no foreground sub-picture region is selected, the background region is decoded and displayed.
[0009] According to another aspect, a computer system for alignment between layers in encoded video data is provided. The computer system can include one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored in at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, whereby the computer system can execute the method. The method may include decoding a video bitstream having a plurality of layers. One or more sub-picture regions are identified from among the plurality of layers of the decoded video bitstream, and the sub-picture region includes a background region and one or more foreground sub-picture regions. Based on a determination that a foreground sub-picture region is selected, an enhanced sub-picture is decoded and displayed. Based on a determination that no foreground sub-picture region is selected, the background region is decoded and displayed.
[0010] According to yet another aspect, there is provided a computer-readable medium for alignment between layers in encoded video data. The computer-readable medium may include one or more computer-readable storage devices and program instructions executable by a processor stored in at least one of the one or more storage devices. The program instructions are executable by a processor to perform a method, which may include decoding a video bitstream having a plurality of layers accordingly. One or more sub-picture regions are identified from among the plurality of layers of the decoded video bitstream, and the sub-picture region includes a background region and one or more foreground sub-picture regions. Based on a determination that a foreground sub-picture region is selected, an enhanced sub-picture is decoded and displayed. Based on a determination that a foreground sub-picture region is not selected, the background region is decoded and displayed.
Brief Description of the Drawings
[0011] These and other objects, features, and advantages will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. The drawings are for the purpose of facilitating understanding by those skilled in the art in connection with the detailed description and are not drawn to scale. In the drawings:
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Best Mode for Carrying Out the Invention
[0012] Detailed embodiments of the structure and method according to the claims are disclosed herein, but it should be understood that the disclosed embodiments are merely illustrative of the structure and method according to the claims that can be implemented in various forms. Those structures and methods, however, can be embodied in many different forms and should not be construed as limited to the exemplary embodiments described herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope to those skilled in the art. In the description, details of well-known mechanisms and techniques may be omitted so as not to unnecessarily obscure the presented embodiments.
[0013] Embodiments generally relate to the field of data processing, and more specifically to media processing. The exemplary embodiments described below provide, among other things, a system, method, and computer program that enable alignment between multiple layers of encoded video data. Accordingly, some embodiments have the ability to improve the field of computing through improved video encoding and decoding.
[0014] As described above, video encoders and decoders typically operate on a given picture size that is defined and remains constant for a coded video sequence (CVS), a Group of Pictures (GOP), or a similar multi-picture temporal frame. For example, in MPEG-2, a system design is known that varies the horizontal resolution (and thus the picture size) depending on factors such as the activity of the scene, but that is only in I pictures and thus typically only for a GOP. Resampling of reference pictures for use at different resolutions within a CVS is known, for example, from ITU-T Recommendation H.263 Annex P. However, there, only the reference pictures are resampled while the picture size remains the same, and potentially only the portion of the picture canvas used (in the case of downsampling) or only the portion of the scene captured (in the case of upsampling). Also, H.263 Annex Q enables resampling of individual macroblocks by a factor of two (in each dimension) up or down. Again, the picture size remains the same. The size of the macroblocks is fixed in H.263 and thus does not need to be signaled.
[0015] However, for example, in the context of 360 coding or certain surveillance applications, multiple semantically independent source pictures (e.g., the six cube surfaces of a cube-projected 360 scene, or the individual camera inputs in the case of a multi-camera surveillance setup) may require separate adaptive resolution settings to handle different activities for each scene at a given point in time. In other words, the encoder may choose to use different resampling factors for multiple different semantically independent pictures that make up the 360 scene or surveillance scene as a whole at a given point in time. When combined into a single picture, it requires that resampling of the reference picture be performed and that adaptive resolution coding signaling be available for multiple parts of the picture being encoded. Thus, it may be advantageous to use the available adaptive resolution coding signaling data for better alignment, encoding, decoding, and display of video layers.
[0016] FIG. 1 illustrates a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The system (100) may include at least two terminals (110 - 120) interconnected via a network (150). In one-way data transmission, a first terminal (110) may encode video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the encoded video data of the other terminal from the network (150), decode the encoded data, and display the restored video data. One-way data transmission may be common in media service providing applications and the like.
[0017] FIG. 1 illustrates a second pair of terminals (130, 140) provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. In bidirectional transmission of data, each terminal (130, 140) may encode video data captured at a local location for transmission to the other terminal via a network (150). Each terminal (130, 140) may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the restored video data on a local display device.
[0018] In FIG. 1, the terminals (110-140) may be illustrated as a server, a personal computer, and a smartphone, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that transfer encoded video data between terminals (110-140), including, for example, a wired communication network and / or a wireless communication network. The communication network (150) may exchange data over a circuit-switched channel and / or a packet-switched channel. Representative networks include long-distance communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (150) may not be important for the operation of the present disclosure, unless otherwise described below.
[0019] FIG. 2 illustrates the arrangement of a video encoder and a decoder in a streaming environment as an example of an application related to the subject matter of the present disclosure. The subject matter of the present disclosure may be equally applicable to other uses where video can be used, including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.
[0020] The streaming system can include a capture subsystem (213), which can include a video source (201), such as a digital camera, that produces, for example, an uncompressed video sample stream (202). The sample stream (202) is drawn as a thick line to emphasize that it has a high data volume compared to the encoded video bitstream and can be processed by an encoder (203) coupled to the camera 201. The encoder (203) can include hardware, software, or a combination thereof to enable or implement aspects of the present disclosure described in more detail hereinafter. The encoded video bitstream (204) is drawn as a thin line to emphasize that it has a low data volume compared to the sample stream and can be stored in a streaming server (205) for later use. One or more streaming clients (206, 208) can access the streaming server (205) to retrieve copies (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210) that decodes an incoming copy (207) of the encoded video bitstream and produces an outgoing video sample stream (211), and the outgoing video sample stream (211) can be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to a particular video encoding / compression standard. Examples of those standards include ITU-T Recommendation H.265. A video encoding standard known informally as Versatile Video Coding or VVC is under development. Aspects of the present disclosure can be used in the context of VVC.
[0021] FIG. 3 can be a functional block diagram of a video decoder (210) according to one or more embodiments.
[0022] The receiver (310) can receive one or more encoded video sequences to be decoded by the decoder (210). In the same or other embodiments, it can receive one encoded video sequence at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence can be received from a channel (312) that can be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data together with other data, such as, for example, encoded audio data and / or auxiliary data streams, and those data can be transferred to their respective using entities (not shown). The receiver (310) can separate the encoded video sequence from other data. To counter network jitter, a buffer memory (315) can be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter, “parser”). When the receiver (310) is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from a synchronous network, the buffer (315) may not be needed or can be made smaller. For use on a best-effort packet network such as the Internet, the buffer (315) may be needed, made relatively large, and advantageously, of an adaptable size.
[0023] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-encoded video sequence. The categories of those symbols include information used to manage the operation of the decoder (210) and may also include information for controlling a rendering device such as, for example, a display (212). A rendering device such as a display (212) is not an integral part of the decoder but can be coupled to the decoder as shown in FIG. 2. Control information for the (one or more) rendering devices may be in the form of Supplementary Enhancement Information (SEI) messages or video user capability information (VUI) parameter set fragments (not shown). The parser (320) may syntax analyze / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be according to a video encoding technique or standard and can follow principles well-known to those skilled in the art, including variable length encoding, Huffman encoding, arithmetic encoding with or without context dependence, etc. The parser (320) can extract a set of subgroup parameters regarding at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroups can include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The entropy decoder / parser can also extract information such as, for example, transform coefficients, quantization parameter values, motion vectors, etc. from the encoded video sequence information.
[0024] The parser (320) may perform entropy decoding / syntax analysis processing on the video sequence received from the buffer (315) to produce symbols (321).
[0025] For the reconstruction of symbol (321), multiple different units may be involved according to the type of the encoded video picture or its part and other factors (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.). How each unit is involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (320). Such a flow of subgroup control information between the parser (320) and the following multiple units is not shown for clarity.
[0026] Beyond the aforementioned functional blocks, the decoder 210 can conceptually be subdivided into a number of functional units as described later. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the matters related to the present disclosure, the conceptual subdivision into the following functional units is appropriate.
[0027] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives the quantized transform coefficients together with control information including which transform to use, block size, quantization coefficient, quantization scaling matrix, etc. as the symbol(s) (321) from the parser (320). This can output a block with sample values that can be input to the aggregator (355).
[0028] In some cases, the output samples of the scaler / inverse transform (351) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra picture prediction unit (352). In some cases, the intra picture prediction unit (352) generates blocks of the same size and shape as the block being reconstructed, using surrounding already reconstructed information fetched from the current (partially reconstructed) picture (356). The aggregator (355) may, in some cases, add, for each sample, the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0029] In other cases, the output samples of the scaler / inverse transform unit (351) may relate to inter-coded, potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols (321) related to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (in this case, called the residual samples or residual signal) to generate the output sample information. The address in the reference picture memory from which the motion compensation unit fetches the prediction samples can be controlled by, for example, the motion vectors available to the motion compensation unit in the form of symbols (321) having X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory when exact sub-sample motion vectors are used, and motion vector prediction mechanisms, etc.
[0030] The output samples of the aggregator (355) can be subjected to various loop filtering techniques in the loop filter unit (356). The video compression technique can include in-loop filtering techniques, which are controlled by parameters that can be included in the encoded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also respond to meta information obtained during decoding of the preceding part (in decoding order) of the encoded picture or encoded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.
[0031] The output of the loop filter unit (356) can be made into a sample stream that can be output to the rendering device (212), and this can also be stored in the reference picture memory (357) for use in future inter-picture prediction.
[0032] When a particular encoded picture is fully reconstructed, it can be used as a reference picture for future prediction. When an encoded picture is fully reconstructed and that encoded picture is specified as a reference picture (e.g., by the parser (320)), the current reference picture (356) can become part of the reference picture buffer (357), and a new current picture memory can be reallocated before starting the reconstruction of the next encoded picture.
[0033] The video decoder 210 may perform decoding processing according to a predetermined video compression technique that can be documented in a standard such as ITU-T Recommendation H.235. The encoded video sequence may comply with the syntax defined by the video compression technique or standard used, in the sense of faithfully observing the syntax of the video compression technique or standard as defined in the video compression technique document or standard, particularly the profile document therein. Also, for compliance, it is also necessary that the complexity of the encoded video sequence is within the limit determined by the level of the video compression technique or standard. Optionally, the level restricts, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured in megasamples per second, for example), the maximum reference picture size, etc. The limitations set by the level may optionally be further restricted through the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0034] In one embodiment, the receiver (310) may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the (one or more) encoded video sequences. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0035] FIG. 4 may be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0036] The encoder (203) may receive video samples from a video source (201) (not part of the encoder) that may capture the (one or more) video images to be encoded by the encoder (203).
[0037] The video source (201) can provide a source video sequence to be encoded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCB, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service providing system, the video source (201) can be a storage device storing pre-prepared videos. In a video conferencing system, the video source (201) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures that convey motion when viewed in sequence. Those pictures themselves can be organized as a spatial array of pixels, and each pixel can have one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can immediately understand the relationship between pixels and samples. The following description focuses on samples.
[0038] According to one embodiment, the encoder (203) can encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under other time constraints required by the application. Enforcing an appropriate encoding speed is one function of the controller (450). The controller controls other functional units as described later and is functionally coupled to those units. That coupling is not shown for clarity. Parameters set by the controller can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can immediately identify other functions of the controller (450) as being related to a video encoder (203) optimized for a specific system design.
[0039] Some video encoders operate in what those skilled in the art would immediately recognize as an "encoding loop." As an overly simplified explanation, the encoding loop can consist of an encoder's encoding portion (430) (hereinafter, the "source coder," which is responsible for creating symbols based on the input picture to be encoded and the (one or more) reference pictures) and a (local) decoder (433) embedded in the encoder (203). The (local) decoder (433) can reconstruct the symbols to generate sample data that the (remote) decoder could also create (in the video compression techniques under consideration in this disclosure, any compression between the symbols and the encoded video bitstream is reversible). The reconstructed sample stream is input into the reference picture memory (434). Since the decoding of the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the contents of the reference picture buffer are also bit-exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the same sample values as reference picture samples that the decoder would "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift in the case where synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.
[0040] The operation of the "local" decoder (433) can be considered the same as that of the "remote" decoder (210), which has already been described in detail above in relation to FIG. 3. However, referring briefly to FIG. 3 as well, since symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy coder (445) and the parser (320) can be assumed to be reversible, the entropy decoding portion of the decoder (210), including the channel (312), the receiver (310), the buffer (315), and the parser (320), need not be fully implemented in the local decoder (433).
[0041] At this point, it can be noticed that any decoder technology, except for syntax analysis / entropy decoding existing in the decoder, must necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the matters related to the present disclosure focus on decoder operations. Since the description of encoder technology is the reverse of the thoroughly described decoder technology, it can be omitted. Only in specific fields is a more detailed description required, which is provided below.
[0042] As part of its operation, the source coder (430) may perform motion-compensated predictive coding that predictively encodes an input frame with respect to one or more previously encoded frames designated as "reference frames" from a video sequence. Thus, the encoding engine (432) encodes the difference between a pixel block of the input frame and a pixel block of the (one or more) reference frames that can be selected as the (one or more) prediction references for the input frame.
[0043] The local video decoder (433) may decode the encoded video data of the frame that can be designated as a reference frame based on the symbols created by the source coder (430). The operation of the encoding engine (432) may advantageously be an irreversible process. When the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence with some error. The local video decoder (433) may replicate the decoding process that can be performed by the video decoder on the reference frame and cause the reconstructed reference frame to be stored in the reference picture cache (434). Thus, the encoder (203) may locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame that would be obtained by the far-end video decoder.
[0044] Predictor (435) may perform prediction search for the encoding engine (432). That is, for a new frame to be encoded, predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) that can serve as an appropriate prediction reference for the new picture or for specific metadata such as, for example, reference picture motion vectors and block shapes. Predictor (435) may operate on a per-pixel block basis to find an appropriate prediction reference. Optionally, the input picture may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search results obtained by predictor (435).
[0045] Controller (450) may manage the encoding process of video coder (430), including, for example, setting parameters and subgroup parameters used to encode video data.
[0046] The outputs of all the aforementioned functional units may be subjected to entropy encoding in entropy coder (445). The entropy coder converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, etc.
[0047] Transmitter (440) may buffer the encoded video sequence generated by entropy coder (445) and prepare it for transmission via communication channel (460). Communication channel (460) may be a hardware / software link to a storage device that stores the encoded video data. Transmitter (440) may merge the encoded video data from video coder (430) with other data to be transmitted, such as, for example, encoded audio data and / or auxiliary data streams (sources not shown).
[0048] The controller (450) may manage the operation of the encoder (203). In encoding, the controller (450) may assign to each encoded picture a specific encoded picture type that may affect the encoding technique applicable to that picture. For example, a picture may often be assigned as one of the following frame types.
[0049] An intra picture (I picture) may be encoded and decoded without using other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example including independent decoder refresh pictures. Those skilled in the art know those variants of I pictures, as well as their respective uses and characteristics.
[0050] A predicted picture (P picture) may be encoded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index to predict the sample values of each block.
[0051] A bi - directionally predicted picture (B picture) may be encoded and decoded using intra prediction or inter prediction, using at most two motion vectors and a reference index to predict the sample values of each block. Similarly, multiple - prediction pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0052] Source pictures are generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be encoded block by block. The blocks can be predictively encoded with reference to other (already encoded) blocks determined by the encoding assignment applied to each of those blocks in their respective pictures. For example, blocks of an I picture can be encoded non-predictively, or they can be encoded predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be encoded non-predictively or via spatial or temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be encoded non-predictively or via spatial or temporal prediction with reference to one or two previously encoded reference pictures.
[0053] The video coder (203) can perform an encoding process according to a predetermined video encoding technique or standard such as ITU-T Recommendation H.265. In its operation, the video coder (203) can perform various compression processes including a predictive encoding process that exploits the temporal and spatial redundancies in the input video sequence. The encoded video data can thus conform to the syntax defined by the video encoding technique or standard being used.
[0054] In one embodiment, the transmitter (440) can transmit additional data along with the encoded video. The video coder (430) can include such data as part of the encoded video sequence. The additional data can have temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, video user usability information (VUI) parameter set fragments, and the like.
[0055] Before describing specific aspects of the disclosed subject matter in more detail, it is necessary to introduce some terms that will be referred to in the remainder of this description.
[0056] Hereinafter, in some cases, a sub-picture refers to a sample, block, macroblock, coding unit, or similar entity in a rectangular configuration that can be semantically grouped and independently coded at a changed resolution. A picture can be formed by one or more sub-pictures. One or more coded sub-pictures can form a coded picture. One or more sub-pictures can be assembled into one picture, and one or more sub-pictures can be extracted from a picture. In a specific environment, one or more coded sub-pictures can be assembled into a coded picture in the compression domain without transcoding to the sample level, and in the same or certain other cases, one or more coded sub-pictures can be extracted from a coded picture in the compression domain.
[0057] Adaptive Resolution Change (ARC) hereinafter refers to a mechanism that enables the change of the resolution of a picture or sub-picture in a coded video sequence, for example, by reference picture resampling. Hereinafter, ARC parameters refer to the control information required to perform adaptive resolution change, which can include, for example, filter parameters, scaling factors, the resolution of the output and / or reference pictures, various control flags, etc.
[0058] The above description focuses on encoding and decoding a single semantically independent coded video picture. Before explaining the meaning of encoding / decoding of multiple sub-pictures with independent ARC parameters and the additional complexity it brings, the options for signaling ARC parameters will be explained.
[0059] Referring to FIG. 5, several new options for signaling ARC parameters are shown. As mentioned for each of these options, they have certain advantages and certain disadvantages from the viewpoints of coding efficiency, complexity, and architecture. A video coding standard or technology may select one or more of these options or options known from the prior art for signaling ARC parameters. These options are not mutually exclusive and may be interchanged with each other as far as possible, based on the application needs, the standard technologies involved, or the encoder selection.
[0060] The class of ARC parameters may include the following: - Upsampling / downsampling factors, separate or combined, in the X and Y dimensions - Upsampling / downsampling factors with an additional time dimension, indicating a constant speed of zooming in / out for a given number of pictures Either of the above two may involve encoding one or more probably short syntax elements that may point within a table containing (one or more) factors. - Resolution in the X or Y dimension, in units of samples, blocks, macroblocks, CUs, or other suitable granularity, of the input picture, output picture, reference picture, coded picture, either in combination or separately. If there are two or more resolutions (for example, one for the input picture and one for the reference picture), in certain cases, one set of values may be inferred from another set of values. This may be gated, for example, by the use of a flag. For more detailed examples, see below. - "Warping" coordinates similar to those used in Annex P of H.263, at a suitable granularity as described above. Annex P of H.263 defines one efficient method for encoding such warping coordinates, but other potentially more efficient methods can be devised wherever possible. For example, the variable-length reversible "Huffman" encoding of the warping coordinates in Annex P may be replaced by a binary encoding of a suitable length, in which case the length of the binary codeword is derived, for example, from the maximum picture size, and may be multiplied by a certain coefficient and offset by a certain value to enable "warping" outside the boundaries of the maximum picture size. - Up or downsampling filter parameters. In the simplest case, there may be only a single filter for upsampling and / or downsampling. However, in certain cases, it can be advantageous to allow more flexibility in filter design, which may require signaling of filter parameters. Such parameters can be selected via an index within a list of possible filter designs, the filter can be fully specified (e.g., via a list of filter coefficients, using appropriate entropy coding techniques), or the filter can alternatively be implicitly selected through the up / downsampling ratio that is signaled according to one of the mechanisms described above.
[0061] Henceforth, this description assumes the encoding of a finite set of up / downsampling coefficients (the same coefficients are used in both the X and Y dimensions) indicated through codewords. The codewords can advantageously be variable-length encoded using, for example, the Ext-Golomb code, which is popular for certain syntax elements in video coding specifications such as H.264 and H.265.
[0062] Many similar mappings can be devised according to the needs of the application and the capabilities of the upscaling and downscaling mechanisms available in the video compression technology or standard. This table may be extended to more values. The values may also be represented by an entropy coding mechanism other than the Ext-Golomb code, for example using binary coding. This may have certain advantages, for example in the case where the resampling factor is of interest outside of the video processing engine (encoder and decoder first) itself, such as by MANE. In the most common case where (presumably) no resolution change is required, a short Ext-Golomb code of only 1 bit can be selected in the above table. This may have an advantage in coding efficiency over using a binary code for the most common cases.
[0063] The number of entries in the table and their semantics can be made fully or partially configurable. For example, the basic framework of the table can be conveyed in a "high" parameter set such as a sequence or decoder parameter set. Instead, or in addition, one or more such tables can be defined in the video coding technology or standard and can be selected, for example, through the decoder or sequence parameter set.
[0064] Hereinafter, how the upsampling / downsampling factors (ARC information) encoded as described above are included in the syntax of the video coding technology or standard will be described. Similar considerations can apply to one or a few codewords that control the up / downsampling filter. See below for the discussion in the case where a relatively large amount of data is required for the filter or other data structure.
[0065] Annex P of H.263 includes ARC information (502) in the form of four warping coordinates in the picture header (501), specifically in the H.263 PLUS PTYPE (503) header extension. This can be a reasonable design choice when a) there is an available picture header and b) frequent changes to the ARC information are expected. However, the overhead when using H.263-style signaling can be very high, and since the picture header can be of a temporary nature, the scaling factors may not be appropriate between picture boundaries.
[0066] JVCET-M135-v1 cited above includes ARC reference information (505) (index) located within the picture parameter set (504), and instead indexes a table (506) that includes the target resolution located within the sequence parameter set (507). The placement of possible resolutions in the table (506) within the sequence parameter set (507) can be justified by using the SPS as an interoperability trade-off point during capability interchange, according to the author's verbal description. The resolution can vary from picture to picture, within the limits set by the values in the table (506), by referring to the appropriate picture parameter set (504).
[0067] Still referring to FIG. 5, the following further options may exist for transmitting ARC information within the video bitstream. Each of these options has certain advantages over the existing technologies described above. These options may coexist simultaneously within the same video coding technology or standard.
[0068] In one embodiment, ARC information (509), such as, for example, a resampling (zoom) factor, may be present within a slice header, GOB header, tile header, or tile group header (hereinafter, tile group header) (508). This may be appropriate when the ARC information is small, for example, as shown above, such as a single variable length ue(v) or a fixed length codeword of several bits. Having the ARC information directly within the tile group header has the additional advantage of making the ARC information applicable to a sub-picture represented by, for example, that tile group rather than the entire picture. See also below. Further, even if the video compression technology or standard assumes only an adaptive resolution change for the entire picture (as opposed to, for example, tile group-based adaptive resolution change), putting the ARC information in the tile group header has certain advantages from the perspective of error resilience compared to putting it in an H.263-like picture header.
[0069] In the same or another embodiment, the ARC information (512) itself may be present within an appropriate parameter set (511), such as, for example, a picture parameter set, header parameter set, tile parameter set, adaptive parameter set (the adaptive parameter set is shown). The scope of this parameter set may advantageously not be larger than the picture, for example, such as a tile group. The use of the ARC information is implicit by the activation of the relevant parameter set. For example, if the video coding technology or standard contemplates only picture-based ARC, a picture parameter set or equivalent may be appropriate.
[0070] In the same or another embodiment, the ARC reference information (513) may be present within a tile group header (514) or a similar data structure. This reference information (513) can refer to a subset (515) of the ARC information available within a parameter set (516) having a scope that exceeds a single picture, such as, for example, a sequence parameter set or decoder parameter set.
[0071] For this additional level, the indirect implicit activation of the PPS from the tile group header, PPS, and SPS, as used in JVET-M0135-v1, does not seem necessary. This is because the picture parameter set can be used for capability negotiation or announcement, similar to the sequence parameter set (and is available in certain standards such as RFC3984 for example). However, if the ARC information is to be applicable to sub-pictures represented by, for example, tile groups, a parameter set with a limited activation scope for the tile group, such as an adaptive parameter set or a header parameter set, may be a better choice. Also, if the ARC information is larger than negligible size and includes filter control information such as a large number of filter coefficients, for example, the parameters may be a better choice from the perspective of coding efficiency than directly using the header (508). This is because their settings can be made reusable by future pictures or sub-pictures by referring to the same parameter set.
[0072] When using a sequence parameter set or another higher-level parameter set with a scope spanning multiple pictures, the following considerations may apply.
[0073] The parameter set storing the ARC information table (516) can be a sequence parameter set in some cases, but advantageously can be a decoder parameter set in other cases. The decoder parameter set can have a validity range of a plurality of CVSs, specifically encoded video streams, that is, all the encoded video bits from session start to session end. Such a range can be even more appropriate. This is because the possible ARC coefficients can probably be a decoder function implemented in hardware, and the hardware function does not tend to change with the CVS (which is a group of pictures with a length of less than 1 second in at least some entertainment systems). That being said, putting the table into the sequence parameter set is clearly included in the placement options described herein.
[0074] The ARC reference information (513) can advantageously be placed directly within the picture / slice style / GOB / tile group header (hereinafter the tile group header) (514), rather than within the picture parameters as in JVCET-M0135-v1. The reasons are as follows. When the encoder wants to change a single value within the picture parameter set, such as the ARC reference information, it has to create a new PPS and refer to that new PPS. Assume that only the ARC reference information changes and other information, such as the quantization matrix information within the PPS, remains the same. Such information can be of a fairly large size and has to be resent to complete the new PPS. The ARC reference information can be a single codeword, such as an index to a table (513), and since it is the only value that changes, resending all of the quantization matrix information, for example, can be cumbersome and wasteful. In that case, avoiding the roundabout way through the PPS, as proposed in JVET-M0135-v1, can be quite beneficial from the perspective of coding efficiency. Similarly, placing the ARC reference information within the PPS has the further drawback that since the scope of picture parameter set activation is the picture, the ARC information referred to by the ARC reference information (513) has to be applied to the entire picture rather than being applied to sub-pictures.
[0075] In the same or another embodiment, the signaling of the ARC parameters can follow the detailed example outlined in FIGS. 6A-6B. FIG. 6 shows a syntax diagram in a notation such as that used in video coding standards since at least 1993. The notation of such a syntax diagram roughly follows C-style programming. The thick lines indicate the syntax elements present in the bitstream, and the non-thick lines often indicate control flow and variable settings.
[0076] As an exemplary syntax structure of a header applicable to a (possibly rectangular) part of a picture, a tile group header (601) can conditionally include a variable-length Exp-Golomb coded syntax element dec_pic_size_idx (602) (shown in bold). The presence of this syntax element within the tile group header can be gated by the use of an adaptive resolution (603), which is the value of a flag not shown in bold here, meaning that at the location where it occurs in the syntax diagram, the flag is present in the bitstream. Whether adaptive resolution is used for this picture or a part thereof can be signaled by any high-level syntax structure inside or outside the bitstream. In the illustrated example, it is signaled within the sequence parameter set as outlined below.
[0077] Still referring to FIG. 6, an excerpt of the sequence parameter set (610) is also shown. The first syntax element shown is the adplicative_pic_resolution_change_flag (611). When true, this flag can indicate the use of adaptive resolution, which in turn may require specific control information. In this example, such control information is conditionally present based on the value of a flag based on the parameter set (612) and the if() statement within the tile group header (601).
[0078] When adaptive resolution is used, in this example, the output resolution is coded in sample units (613). The reference sign 613 refers to both output_pic_width_in_luma_samples and output_pic_height_in_luma_samples, and together they can determine the resolution of the output picture. In some video coding technologies or standards, specific restrictions for any value may be defined. For example, the total number of output samples, which can be the product of the values of these two syntax elements, may be restricted by the level specification. Also, a specific video coding technology or standard, or an external technology or standard such as a system standard, may restrict the numbering range (for example, one or both dimensions must be divisible by a power of 2) or the aspect ratio (for example, the width and height must be in a relationship such as 4:3 or 16:9). Such restrictions may be introduced to facilitate hardware implementation or for other reasons and are well-known technically.
[0079] In certain applications, it may be desirable for the encoder to instruct the decoder to use a predetermined reference picture size instead of implicitly assuming its size to be the output picture size. In this example, the syntax element reference_pic_size_present_flag (614) gates the conditional presence of the reference picture dimensions (615) (here again, this reference sign refers to both the width and height).
[0080] Finally, a table of possible decoded picture widths and heights is shown. Such a table can be represented, for example, by table indication (num_dec_pic_size_in_luma_samples_minus1) (616). "minus1" can refer to the interpretation of the value of this syntax element. For example, if the coded value is zero, there is one table entry, and if the value is 5, there are six table entries. In each "line" of the table, the width and height of the decoded picture are included in the syntax (617).
[0081] The indicated table entry (617) can be indexed using the syntax element dec_pic_size_idx (602) within the tile group header, thereby enabling different decoding sizes (in effect, zoom factors) for each tile group.
[0082] Certain video coding techniques or standards, such as VP9, support spatial scalability by implementing a particular form of reference picture resampling, along with temporal scalability, to enable spatial scalability. In particular, a particular reference picture can be upsampled to a higher resolution using an ARC-style technique to form the basis of a spatial enhancement layer. These upsampled pictures can be refined using normal prediction mechanisms at their higher resolution to add detail.
[0083] The subject matter of the disclosure can be used in such an environment. In certain cases, in the same or another embodiment, values such as, for example, the Temporal ID field within the NAL unit header can be used to indicate not only the temporal layer but also the spatial layer. Doing so has certain advantages with respect to a particular system design, for example, enabling an existing selected forwarding unit (SFU) created and optimized for a temporal layer selected based on the temporal ID value of the NAL unit header to be used without modification for a scalable environment. To enable this, it can be assumed that a mapping between the coded picture size and the temporal layer needs to be indicated by the temporal ID field within the NAL unit header.
[0084] In some video encoding techniques, an access unit (AU) can refer to one or more encoded pictures, slices, tiles, NAL units, etc. that are captured at a given instance in time and then combined into a picture / slice / tile / NAL unit bitstream. This instance in time can be the composition time.
[0085] In HEVC and other specific video encoding techniques, a picture order count (POC) value can be used to indicate a reference picture selected from among a plurality of reference pictures stored in a decoded picture buffer (DPB). If an access unit (AU) has one or more pictures, slices, or tiles, each picture, slice, or tile belonging to the same AU can carry the same POC value, from which it can be deduced that they are created from content of the same composition time. In other words, in a scenario where two pictures / slices / titles carry the same POC value, it can be assumed that these two pictures / slices / titles belong to the same AU and have the same composition time. Conversely, two pictures / titles / slices with different POC values can indicate that these pictures / slices / titles belong to different AUs and have different composition times.
[0086] In one embodiment of the disclosed subject matter, the strict relationship described above can be relaxed in that an access unit can have a plurality of pictures, slices, or tiles with different POC values. By allowing multiple different POC values within one AU, it becomes possible to use the POC value to identify potentially independently decodable pictures / slices / titles with equal presentation times. This, in turn, can enable support for multiple scalable layers without changing the reference picture selection signaling (e.g., reference picture set signaling or reference picture list signaling), as will be described in more detail later.
[0087] However, it is still desirable that the access unit (AU) to which a picture / slice / tile belongs can be identified from the POC value alone with respect to other pictures / slices / titles having different POC values. This can be achieved as described below.
[0088] In the same or other embodiments, the access unit count (AUC) can be signaled in a high-level syntax structure such as, for example, a NAL unit header, a slice header, a tile group header, a SEI message, a parameter set, or an AU delimiter. The value of the AUC can be used to identify which NAL unit, picture, slice, or tile belongs to a given AU. The value of the AUC can correspond to distinguishable composition time instances. The AUC value can be equal to a multiple of the POC value. The AUC value can be calculated by dividing the POC value by an integer value. In certain cases, the division operation can impose a certain burden on decoder implementation. In such cases, a small constraint in the numbering space of the AUC value can make it possible to replace the division operation with a shift operation. For example, the AUC value can be equal to the most significant bit (MSB) value of the POC value range.
[0089] In the same embodiment, the value of the POC cycle per AU (poc_cycle_au) can be signaled in a high-level syntax structure such as, for example, a NAL unit header, a slice header, a tile group header, a SEI message, a parameter set, or an AU delimiter. The poc_cycle_au can indicate how many consecutive different POC values can be associated with the same AU. For example, when the value of the poc_cycle_au is equal to 4, pictures, slices, or tiles having POC values equal to 0-3 including both ends are associated with an AU having an AUC value equal to 0, and pictures, slices, or tiles having POC values equal to 4-7 including both ends are associated with an AU having an AUC value equal to 1. Therefore, the value of the AUC can be estimated by dividing the POC value by the value of the poc_cycle_au.
[0090] In the same or another embodiment, the value of poc_cyle_au may be derived from information identifying the number of spatial or SNR layers in an encoded video sequence, for example, located within a video parameter set (VPS). Such possible relationships are briefly described below. The above derivation may save a few bits in the VPS and thus improve the encoding efficiency, but it may be advantageous to explicitly encode poc_cyle_au in an appropriate high-level syntax structure under the video parameter set hierarchically so that poc_cycle_au can be minimized for a given small portion of the bitstream, such as a picture. This optimization may save more bits than can be saved through the above derivation process because the POC value (and / or the value of a syntax element that indirectly references the POC) may be encoded in a low-level syntax structure.
[0091] In the same or another embodiment, FIG. 9 shows a syntax table for signaling the syntax element of vps_poc_cycle_au in a VPS (or SPS) that indicates poc_cycle_au used for all pictures / slices in an encoded video sequence, and the syntax element of slice_poc_cycle_au that indicates the poc_cycle_au of the current slice in a slice header. When the POC value increases uniformly for each AU, vps_contant_poc_cycle_per_au in the VPS is set to 1, and vps_poc_cycle_au is signaled in the VPS. In this case, slice_poc_cycle_au is not explicitly signaled, and the value of AUC for each AU is calculated by dividing the POC value by vps_poc_cycle_au. When the POC value does not increase uniformly for each AU, vps_contant_poc_cycle_per_au in the VPS is set equal to 0. In this case, vps_access_unit_cnt is not signaled, and slice_access_unit_cnt is signaled in the slice header of each slice or picture. Each slice or picture may have a different value of slice_access_unit_cnt. The value of AUC for each AU is calculated by dividing the POC value by slice_poc_cycle_au. FIG. 10 shows a block diagram illustrating the related workflow.
[0092] In the same or another embodiment, even though the POC values of pictures, slices, or tiles may be different, pictures, slices, or tiles corresponding to AUs having the same AUC value may be associated with the same decoding or output time instance. Thus, all or a subset of the pictures, slices, or tiles associated with the same AU can be decoded in parallel and output at the same time instance without inter-syntax analysis / decoding dependencies across the pictures, slices, or tiles within the same AU.
[0093] In the same or another embodiment, pictures, slices, or tiles corresponding to AUs having the same AUC value may be associated with the same composition / display time instance, even if the POC values of the pictures, slices, or tiles may differ. If the composition time is included in a container format, even if the pictures correspond to different AUs, those pictures can be displayed at the same time instance if they have the same composition time.
[0094] In the same or another embodiment, each picture, slice, or tile may have the same temporal identifier (temporal_id) within the same AU. All or a subset of the pictures, slices, or tiles corresponding to a time instance may be associated with the same temporal sublayer. In the same or another embodiment, each picture, slice, or tile may have the same or different spatial layer IDs (layer_id) within the same AU. All or a subset of the pictures, slices, or tiles corresponding to a time instance may be associated with the same or different spatial layers.
[0095] FIG. 8 shows an example of a video sequence structure having a combination of temporal_id, layer_id, POC, and AUC values with adaptive resolution change. In this example, pictures, slices, tiles within the first AU with AUC = 0 can have temporal_id = 0 and layer_id = 0 or 1, and pictures, slices, tiles of the second AU with AUC = 1 can have temporal_id = 1 and layer_id = 0 or 1. The value of POC is incremented by 1 for each picture, regardless of the values of temporal_id and layer_id. In this example, the value of poc_cycle_au can be assumed to be equal to 2. Preferably, the value of poc_cycle_au can be set equal to the number of (spatial scalability) layers. In this example, therefore, the value of POC is incremented by 2, and the value of AUC is incremented by 1.
[0096] In the above embodiment, by using the existing reference picture set (RPS) signaling or reference picture list (RPL) signaling in HEVC, all or a subset of the inter-picture or inter-layer prediction structure and reference picture indication can be supported. In RPS or RPL, the selected reference picture is indicated by signaling the value of the picture order count (POC) or the POC delta value between the current picture and the selected reference picture. In the matters disclosed, there are the following constraints when using RPS and RPL, but the inter-picture or inter-layer prediction structure can be indicated without changing the signaling. If the value of the temporal_id of the reference picture is greater than the value of the temporal_id of the current picture, the current picture may not use that reference picture for motion compensation or other prediction. If the value of the layer_id of the reference picture is greater than the value of the layer_id of the current picture, the current picture may not use that reference picture for motion compensation or other prediction.
[0097] In the same or other embodiments, motion vector scaling based on the POC difference for temporal motion vector prediction may be disabled across multiple pictures within an access unit. Thus, each picture within an access unit may have a different POC value, but the motion vectors are not scaled and are not used for temporal motion vector prediction within the access unit. This is because reference pictures with different POCs within the same AU are considered to have the same temporal instance. Thus, in an embodiment, if the reference picture belongs to the AU associated with the current picture, the motion vector scaling function may return 1.
[0098] In the same or other embodiments, motion vector scaling based on the POC difference for temporal motion vector prediction may optionally be disabled across multiple pictures if the spatial resolution of the reference picture is different from the spatial resolution of the current picture. When motion vector scaling is enabled, the motion vectors are scaled based on both the POC difference and the spatial resolution ratio between the current picture and the reference picture.
[0099] In the same embodiment or another embodiment, the motion vectors may be scaled based on the AUC difference instead of the POC difference for temporal motion vector prediction, especially when poc_cycle_au has a non-uniform value (when vps_contant_poc_cycle_per_au == 0). Otherwise (when vps_contant_poc_cycle_per_au == 1), the motion vector scaling based on the AUC difference may be the same as the motion vector scaling based on the POC difference.
[0100] In the same or other embodiments, when the motion vectors are scaled based on the AUC difference, the reference motion vectors in the same AU (having the same AUC value) as the current picture are not scaled based on the AUC difference and are used for motion vector prediction without scaling or using scaling based on the spatial resolution ratio between the current picture and the reference picture.
[0101] In the same or other embodiments, the AUC value may be used to identify the boundaries of the AU and may be used in the operation of a hypothetical reference decoder (HRD) that requires input and output timings at the AU granularity. In most cases, the decoded picture using the highest layer within the AU may be output for display. The AUC value and the layer_id value can be used to identify the output picture.
[0102] In one embodiment, a picture may be composed of one or more sub - pictures. Each sub - picture may cover a local area or the entire area of the picture. The area supported by a sub - picture may or may not overlap with the area supported by another sub - picture. The area composed of one or more sub - pictures may or may not cover the entire area of the picture. When the picture is composed of one sub - picture, the area supported by that sub - picture is the same as the area supported by the picture.
[0103] In the same embodiment, a sub - picture may be encoded by an encoding method similar to the encoding method used for the picture to be encoded. The sub - picture may be encoded independently or dependently on another sub - picture or the encoded picture. The sub - picture may or may not have some syntactic analysis dependencies from another sub - picture or the encoded picture.
[0104] In the same embodiment, the encoded sub - pictures may be included in one or more layers. The encoded sub - pictures within a layer may have different spatial resolutions. The original sub - picture may be spatially resampled (upsampled or downsampled), encoded with different spatial resolution parameters, and included in the bit - stream corresponding to the layer.
[0105] In the same embodiment or another embodiment, assuming that W represents the width of the sub - picture and H represents the height of the sub - picture, a sub - picture with (W, H) is encoded and included in the encoded bit - stream corresponding to layer 0, and S w,k , S h,k represents the resampling ratios in the horizontal and vertical directions. A sub - picture upsampled (or downsampled) from the original - resolution sub - picture of (W * S w,k , H * S h,k ) may be encoded and included in the encoded bit - stream corresponding to layer k. S w,k, S h,k If the value of S is greater than 1, the resampling is equal to upsampling. S w,k , S h,k If the value of S is less than 1, the resampling is equal to downsampling.
[0106] In the same embodiment or another embodiment, the encoded sub - pictures within a layer can have a different visual quality from those of the encoded sub - pictures of another layer within the same sub - picture or a different sub - picture. For example, sub - picture i within layer n is encoded with quantization parameter Q i,n and sub - picture j within layer m is encoded with quantization parameter Q j,m .
[0107] In the same embodiment or another embodiment, the encoded sub - pictures within a layer can be independently decodable without syntactic analysis or decoding dependence from the encoded sub - pictures of another layer within the same local region. A sub - picture layer that can be considered independently decodable without referring to another sub - picture layer of the same local region is an independent sub - picture layer. The encoded sub - pictures within an independent sub - picture layer may or may not have decoding or syntactic analysis dependence from previously encoded sub - pictures within the same sub - picture layer, but the encoded sub - pictures can be considered to have no dependence from the encoded pictures of another sub - picture layer.
[0108] In the same embodiment or another embodiment, the encoded sub - pictures within a layer can be dependently decodable with syntactic analysis or decoding dependence from the encoded sub - pictures of another layer within the same local region. A sub - picture layer that can be considered dependently decodable by referring to another sub - picture layer of the same local region is a dependent sub - picture layer. The encoded sub - pictures within a dependent sub - picture layer can refer to encoded sub - pictures belonging to the same sub - picture, previously encoded sub - pictures within the same sub - picture layer, or both reference sub - pictures.
[0109] In the same or another embodiment, an encoded subpicture is composed of one or more independent subpicture layers and one or more dependent subpicture layers. However, at least one independent subpicture layer may exist in the encoded subpicture. An independent subpicture layer may have a layer identifier (layer_id) value equal to 0 and may exist within an NAL unit header or other high-level syntax structure. A subpicture layer having a layer_id equal to 0 may be regarded as a base subpicture layer.
[0110] In the same or another embodiment, a picture is composed of one or more foreground subpictures and one background subpicture. The area supported by the background subpicture may be equal to the area of the picture. The area supported by the foreground subpicture may overlap with the area supported by the background subpicture. The background subpicture can be a base subpicture layer, and the foreground subpicture can be a non-base (enhancement) subpicture layer. One or more non-base subpicture layers may refer to the same base layer for decoding. Assuming a is greater than b, each non-base subpicture layer having a layer_id equal to a may refer to a non-base subpicture layer having a layer_id equal to b.
[0111] In the same or another embodiment, a picture may be composed of one or more foreground subpictures with or without a background subpicture. Each subpicture may have its own base subpicture layer and one or more non-base (enhancement) layers. Each base subpicture layer may be referred to by one or more non-base subpicture layers. Assuming a is greater than b, each non-base subpicture layer having a layer_id equal to a may refer to a non-base subpicture layer having a layer_id equal to b.
[0112] In the same or another embodiment, a picture may be composed of one or more foreground sub - pictures, with or without a background sub - picture. Each encoded sub - picture within a (base or non - base) sub - picture layer may be referenced by one or more non - base layer sub - pictures that belong to the same sub - picture and one or more non - base layer sub - pictures that do not belong to the same sub - picture.
[0113] In the same or another embodiment, a picture may be composed of one or more foreground sub - pictures, with or without a background sub - picture. The sub - pictures within layer a may be further divided into a plurality of sub - pictures within the same layer. One or more encoded sub - pictures within layer b may reference the divided sub - pictures within layer a.
[0114] In the same or another embodiment, an encoded video sequence (CVS) may be a group of encoded pictures. The CVS can be composed of one or more encoded sub - picture sequences (CSPS), where a CSPS may be a group of encoded sub - pictures that cover the same local region of a picture. The CSPS may have the same or a different temporal resolution than the encoded video sequence.
[0115] In the same or another embodiment, a CSPS may be encoded and included in one or more layers. A CSPS may be composed of one or more CSPS layers. By decoding one or more CSPS layers corresponding to a CSPS, a sequence of sub - pictures corresponding to the same local region can be reconstructed.
[0116] In the same or another embodiment, the number of CSPS layers corresponding to a CSPS may be the same as or different from the number of CSPS layers corresponding to another CSPS.
[0117] In the same or another embodiment, the CSPS layer may have a different temporal resolution (e.g., frame rate) from another CSPS layer. The original (uncompressed) sub-picture sequence may be temporally resampled (upsampled or downsampled), encoded with different temporal resolution parameters, and included in the bitstream corresponding to the layer.
[0118] In the same or another embodiment, a sub-picture sequence having a frame rate F is encoded and included in the encoded bitstream corresponding to layer 0, and S t,k is assumed to represent the temporal sampling ratio for layer k, a sub-picture sequence temporally upsampled (or downsampled) from the original sub-picture sequence having F*S t,k may be encoded and included in the encoded bitstream corresponding to layer k. When the value of S t,k is greater than 1, the temporal resampling process is equivalent to frame rate up-conversion. When the value of S t,k is less than 1, the temporal resampling process is equivalent to frame rate down-conversion.
[0119] In the same or another embodiment, when a sub-picture having CSPS layer a is referenced by a sub-picture having CSPS layer b for motion compensation or some inter-layer prediction, if the spatial resolution of CSPS layer a is different from the spatial resolution of CSPS layer b, the decoded pixels in CSPS layer a are resampled and used for reference. This resampling process may require upsampling filtering or downsampling filtering.
[0120] FIG. 11 shows an example of a video stream including a background video CSPS having a layer_id equal to 0 and a plurality of foreground CSPS layers. Encoded sub-pictures can be composed of one or more CSPS layers, and a background area that does not belong to any of the foreground CSPS layers can be composed of a base layer. The base layer can include a background area and a foreground area, and the enhancement CSPS layer includes a foreground area. The enhancement CSPS layer can have better visual quality than the base layer in the same area. The enhancement CSPS layer can refer to the reconstructed pixels and motion vectors of the base layer corresponding to the same area.
[0121] In the same embodiment or another embodiment, the video bitstream corresponding to the base layer is included in a track, and the CSPS layer corresponding to each sub-picture is included in a separate track within the video file.
[0122] In the same embodiment or another embodiment, the video bitstream corresponding to the base layer is included in a track, and the CSPS layer having the same layer ID is included in a separate track. In this example, the track corresponding to layer k includes only the CSPS layer corresponding to layer k.
[0123] In the same embodiment or another embodiment, each CSPS layer of each sub-picture is stored in a separate track. Each track may or may not have a syntax analysis or decoding dependency from one or more other tracks.
[0124] In the same embodiment or another embodiment, assuming 0 < i <= j <= k and k is the top layer of CSPS, each track can include the bitstream corresponding to the CSPS layers from layer i to layer j of all or a subset of the sub-pictures.
[0125] In the same or another embodiment, the picture is composed of one or more associated media data including a depth map, an alpha map, 3D geometry data, an occupancy map, etc. Such associated time-stamped media data can each be split into one or more data sub-streams, each corresponding to one sub-picture.
[0126] In the same or another embodiment, FIG. 12 shows an example of a video conference based on the multi-layered sub-picture method. The video stream includes one base layer video bitstream corresponding to the background picture and one or more enhancement layer video bitstreams corresponding to the foreground sub-pictures. Each enhancement layer video bitstream corresponds to a CSPS layer. By default, the picture corresponding to the base layer is displayed on the display. This includes the picture-in-picture (PIP) of one or more users. When a specific user is selected by the control of the client, the enhancement CSPS layer corresponding to the selected user is decoded and displayed with enhanced quality or spatial resolution. FIG. 13 shows a diagram of the operation.
[0127] In the same or another embodiment, a network intermediate box (e.g., a router, etc.) can select a subset of the plurality of layers to be sent to the user according to its bandwidth. Picture / sub-picture composition can be used for bandwidth adaptation. For example, when the user has no bandwidth, the router can strip layers or select some sub-pictures according to importance or based on the usage setup, which can be done dynamically to adapt to the bandwidth.
[0128] Figure 14 shows a use case of 360 video. When a spherical 360 picture is projected onto a planar picture, the projected 360 picture can be divided into a plurality of sub - pictures as a base layer. An enhancement layer of a specific sub - picture can be encoded and sent to a client. A decoder may be able to decode both the base layer including all sub - pictures and the enhancement layer of the selected sub - pictures. When the current viewport is the same as the selected sub - picture, the displayed picture can have higher quality using the decoded sub - picture with the enhancement layer. Otherwise, the decoded picture with the base layer can be displayed with low quality.
[0129] In the same embodiment or another embodiment, some layout information for display may exist in the file as supplementary information (such as SEI messages or metadata, etc.). One or more decoded sub - pictures can be rearranged and displayed according to the signaled layout information. The layout information may be signaled by a streaming server or a broadcaster, or may be regenerated by a network entity or a cloud server, or may be determined by the user's customization settings.
[0130] In one embodiment, when an input picture is divided into one or more (rectangular) sub-regions, each sub-region can be encoded as an independent layer. Each independent layer corresponding to a local region can have a unique layer_id value. For each independent layer, sub-picture size and position information can be signaled. For example, picture size (width, height), offset information (x_offset, y_offset) of the upper left corner. FIG. 15 shows an example of the layout of the divided sub-pictures, their sub-picture sizes and position information, and the corresponding picture prediction structure. This layout information including (one or more) sub-picture sizes and (one or more) sub-picture positions can be signaled in a high-level syntax structure such as, for example, (one or more) parameter sets, slice or tile group headers, or SEI messages.
[0131] In the same embodiment, each sub-picture corresponding to an independent layer can have its own POC value within the AU. When a certain reference picture among the plurality of pictures stored in the DPB is indicated by using (one or more) syntax elements within the RPS or RPL structure, the (one or more) POC values of each sub-picture corresponding to the layer can be used.
[0132] In the same embodiment or another embodiment, POC (delta) values may be used instead of layer_id to indicate the (inter-layer) prediction structure.
[0133] In the same embodiment, a sub-picture having a POC value equal to N corresponding to a layer (or local region) can be used or not used as a reference picture of a sub-picture having a POC value equal to N+K corresponding to the same layer (or the same local region) for motion compensation prediction. In most cases, the value of the number K can be equal to the number of sub-regions and can be equal to the maximum number of (independent) layers.
[0134] In the same or another embodiment, FIG. 16 shows an extended case of FIG. 15. When the input picture is divided into a plurality of (e.g., four) sub-regions, each local region can be encoded with one or more layers. In this case, the number of independent layers can be equal to the number of sub-regions, and one or more layers can correspond to the sub-regions. Thus, each sub-region can be encoded with one or more independent layers and zero or more dependent layers.
[0135] In the same embodiment, in FIG. 16, the input picture can be divided into four sub-regions. The upper-right sub-region can be encoded as two layers, layer 1 and layer 4, and the lower-right sub-region can be encoded as two layers, layer 3 and layer 5. In this case, layer 4 can refer to layer 1 for motion compensation prediction, and layer 5 can refer to layer 3 for motion compensation.
[0136] In the same or another embodiment, in-loop filtering (e.g., deblocking filtering, adaptive in-loop filtering, reshaper, bilateral filtering, or any deep learning-based filtering) across layer boundaries can be (optionally) disabled.
[0137] In the same or another embodiment, motion compensation prediction or intra-block copy across layer boundaries can be (optionally) disabled.
[0138] In the same or another embodiment, boundary padding for motion compensation prediction or in-loop filtering at the boundaries of sub-pictures can be optionally processed. A flag indicating whether the boundary padding is processed can be signaled in a high-level syntax structure such as (one or more) parameter sets (VPS, SPS, PPS, or APS), slice or tile group headers, or SEI messages.
[0139] In the same embodiment or another embodiment, the layout information of the (one or more) sub-regions (or (one or more) sub-pictures) may be signaled within the VPS or SPS. FIG. 17 shows an example of syntax elements in the VPS and SPS. In this example, the vps_sub_picturing_dividing_flag is signaled within the VPS. This flag may indicate whether the (one or more) input pictures are divided into a plurality of sub-regions. If the value of the vps_sub_picture_dividing_flag is equal to 0, the (one or more) input pictures within the (one or more) coded video sequences corresponding to the current VPS may not be divided into a plurality of sub-regions. In this case, the input picture size may be equal to the coded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) that is signaled within the SPS. If the value of the vps_sub_picture_dividing_flag is equal to 1, the (one or more) input pictures may be divided into a plurality of sub-regions. In this case, the syntax elements vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples are signaled within the VPS. The values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be equal to the width and height of the (one or more) input pictures, respectively.
[0140] In the same embodiment, the values of vps_full_pic_width_in_luma_samples and vps_full_pic_height_in_luma_samples may be used for synthesis and display without being used for decoding.
[0141] In the same embodiment, when the value of vps_sub_picture_dividing_flag is equal to 1, the syntax elements pic_offset_x and pic_offset_y can be signaled in the SPS that corresponds to a specific (one or more) layer. In this case, the coded picture size (pic_width_in_luma_samples, pic_height_in_luma_samples) signaled in the SPS can be equal to the width and height of the sub-region corresponding to the specific layer. Also, the position (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region can also be signaled in the SPS.
[0142] In the same embodiment, the position information (pic_offset_x, pic_offset_y) of the upper left corner of the sub-region may be used for synthesis and display without being used for decoding.
[0143] In the same embodiment or another embodiment, the layout information (size and position) of all or a subset of (one or more) sub-regions of (one or more) input pictures, and the dependency information between (one or more) layers may be signaled within a parameter set or an SEI message. FIG. 18 shows an example of a syntax element indicating the layout information of sub-regions, the dependencies between layers, and the relationship between sub-regions and one or more layers. In this example, the syntax element num_sub_region indicates the number of (rectangular) sub-regions within the currently encoded video sequence. The syntax element num_layers indicates the number of layers within the currently encoded video sequence. The value of num_layers may be equal to or greater than the value of num_sub_region. When any sub-region is encoded as a single layer, the value of num_layers may be equal to the value of num_sub_region. When one or more sub-regions are encoded as multiple layers, the value of num_layers may be greater than the value of num_sub_region. The syntax element direct_dipendency_flag[i][j] indicates the dependency from the j-th layer to the i-th layer. num_layers_for_region[i] indicates the number of layers associated with the i-th sub-region. sub_region_layer_id[i][j] indicates the layer_id of the j-th layer associated with the i-th sub-region. sub_region_offset_x[i] and sub_region_offset_y[i] indicate the horizontal and vertical positions of the upper left corner of the i-th sub-region, respectively. sub_region_width[i] and sub_region_height[i] indicate the width and height of the i-th sub-region, respectively.
[0144] In one embodiment, one or more syntax elements that specify an output layer set to indicate one or more layers output with or without profile tier level information may be signaled in a high-level syntax structure such as, for example, a VPS, DPS, SPS, PPS, APS, or SEI message. Referring to FIG. 19, a syntax element num_output_layer_sets indicating the number of output layer sets (OLSs) in an encoded video sequence that references a VPS may be signaled within the VPS. For each output layer set, an output_layer_flag may be signaled as many times as the number of output layers.
[0145] In the same embodiment, an output_layer_flag[i] equal to 1 specifies that the i-th layer is output. A vps_output_layer_flag[i] equal to 0 specifies that the i-th layer is not output.
[0146] In the same or another embodiment, one or more syntax elements that define profile tier level information for each output layer set may be signaled in a high-level syntax structure such as, for example, a VPS, DPS, SPS, PPS, APS, or SEI message. Still referring to FIG. 19, a syntax element num_profile_tile_level indicating the number of profile tier level information for each OLS in an encoded video sequence that references a VPS may be signaled within the VPS. For each output layer set, a set of syntax elements for profile tier level information, or an index indicating a particular profile tier level information among the entries in the profile tier level information, may be signaled as many times as the number of output layers.
[0147] In the same embodiment, profile_tier_level_idx[i][j] defines the index of the profile_tier_level() syntax structure applied to the j-th layer of the i-th OLS in the list of profile_tier_level() syntax structures within the VPS.
[0148] In the same embodiment or another embodiment, referring to FIG. 20, when the maximum number of layers is greater than 1 (vps_max_layers_minus1>0), the syntax elements num_profile_tile_level and / or num_output_layer_sets may be signaled.
[0149] In the same embodiment or another embodiment, referring to FIG. 20, the syntax element vps_output_layers_mode[i] indicating the mode of output layer signaling for the i-th output layer set may be present within the VPS.
[0150] In the same embodiment, vps_output_layers_mode[i] equal to 0 specifies that only the top layer is output in the i-th output layer set. vps_output_layer_mode[i] equal to 1 specifies that all layers are output in the i-th output layer set. vps_output_layer_mode[i] equal to 2 specifies that the layers to be output are the layers having vps_output_layer_flag[i][j] equal to 1 in the i-th output layer set. More values may be reserved.
[0151] In the same embodiment, output_layer_flag[i][j] may or may not be signaled depending on the value of vps_output_layers_mode[i] for the i-th output layer set.
[0152] In the same embodiment or another embodiment, referring to FIG. 20, the flag vps_ptl_signal_flag[i] may be present for the i-th output layer set. Depending on the value of vps_ptl_signal_flag[i], the profile tile level information for the i-th output layer set may or may not be signaled.
[0153] In the same embodiment or another embodiment, referring to FIG. 21, the number of sub-pictures max_subpics_minus1 currently in the CVS may be signaled in a high-level syntax structure such as, for example, a VPS, DPS, SPS, PPS, APS, or SEI message.
[0154] In the same embodiment, referring to FIG. 21, if the number of sub-pictures is greater than 1 (max_subpics_minus1>0), the sub-picture identifier sub_pic_id[i] for the i-th sub-picture may be signaled.
[0155] In the same embodiment or another embodiment, one or more syntax elements indicating the sub-picture identifiers belonging to each layer of each output layer set may be signaled within the VPS. Referring to FIG. 21, sub_pic_id_layer[i][j][k] indicates the k-th sub-picture present in the j-th layer of the i-th output layer set. Using this information, the decoder can recognize which sub-pictures can be decoded and output for each layer of a specific output layer set.
[0156] In one embodiment, a picture header (PH) is a syntax structure that includes syntax elements applicable to all slices of an encoded picture. A picture unit (PU) is a set of NAL units that are associated with each other according to defined classification rules, are consecutive in decoding order, and contain exactly one encoded picture. A PU may include a picture header (PH) and one or more VCL NAL units having the encoded picture.
[0157] In one embodiment, the SPS (RBSP) may be made available for the decoding process before it is referenced, included in at least one AU having a TemporalId equal to 0, or provided via external means.
[0158] In one embodiment, the SPS (RBSP) is included in at least one AU having a TemporalId equal to 0 within a CVS that can be made available for the decoding process before it is referenced and may include one or more PPSs that reference the SPS, or may be provided via external means.
[0159] In one embodiment, the SPS (RBSP) is included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit within a CVS that can be made available for the decoding process before it is referenced by one or more PPSs and may include one or more PPSs that reference the SPS, or may be provided via external means.
[0160] In one embodiment, the SPS (RBSP) is included in at least one PU having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit, which can be made available for the decoding process before it is referenced by one or more PPSs, or may be provided via external means.
[0161] In one embodiment, the SPS (RBSP) is included in at least one PU having a TemporalId equal to 0 and a nuh_layer_id equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit within a CVS that can be made available for the decoding process before it is referenced by one or more PPSs and may include one or more PPSs that reference the SPS, or may be provided via external means.
[0162] In the same or another embodiment, pps_seq_parameter_set_id defines the value of sps_seq_parameter_set_id for the referenced SPS. The value of pps_seq_parameter_set_id can be the same in all PPSs referenced by coded pictures within the CLVS.
[0163] In the same or another embodiment, all SPS NAL units having a particular value of sps_seq_parameter_set_id within the CVS can have the same content.
[0164] In the same or another embodiment, regardless of the nuh_layer_id value, SPS NAL units can share the same value space of sps_seq_parameter_set_id.
[0165] In the same or another embodiment, the nuh_layer_id value of an SPS NAL unit can be equal to the lowest nuh_layer_id value of the PPS NAL unit that references the SPS NAL unit.
[0166] In one embodiment, when an SPS having an nuh_layer_id equal to m is referenced by one or more PPSs having an nuh_layer_id equal to n, the layer having an nuh_layer_id equal to m can be the same as the layer having an nuh_layer_id equal to n or the (direct or indirect) reference layer of the layer having an nuh_layer_id equal to m.
[0167] In one embodiment, the PPS (RBSP) can be made available for the decoding process before it is referenced, included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit, or provided via external means.
[0168] In one embodiment, the PPS (RBSP) is made available for use in the decoding process before it is referenced, and is included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit in the CVS that contains one or more PHs (or coded slice NAL units) that reference the PPS, or can be provided via external means.
[0169] In one embodiment, the PPS (RBSP) is made available for use in the decoding process before it is referenced by one or more PHs (or coded slice NAL units), and is included in at least one PU having an nuh_layer_id equal to the lowest nuh_layer_id of the coded slice NAL units that reference the PPS NAL unit in the CVS that contains one or more PHs (or coded slice NAL units) that reference the PPS, or can be provided via external means.
[0170] In one embodiment, the PPS (RBSP) is made available for use in the decoding process before it is referenced by one or more PHs (or coded slice NAL units), and is included in at least one PU having a TemporalId equal to the TemporalId of the PPS NAL unit and an nuh_layer_id equal to the lowest nuh_layer_id of the coded slice NAL units that reference the PPS NAL unit in the CVS that contains one or more PHs (or coded slice NAL units) that reference the PPS, or can be provided via external means.
[0171] In the same or another embodiment, the ph_pic_parameter_set_id within the PH defines the value of the pps_pic_parameter_set_id for the reference PPS used. The value of the pps_seq_parameter_set_id can be the same for all PPSs referenced by the coded pictures in the CLVS.
[0172] In the same embodiment or in another embodiment, all PPS NAL units having a specific value of pps_pic_parameter_set_id within the PU have the same content.
[0173] In the same embodiment or in another embodiment, regardless of the nuh_layer_id value, PPS NAL units may share the same value space of pps_pic_parameter_set_id.
[0174] In the same embodiment or in another embodiment, the nuh_layer_id value of a PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL units that reference the NAL units that reference the PPS NAL unit.
[0175] In one embodiment, when a PPS having a nuh_layer_id equal to m is referenced by one or more coded slice NAL units having a nuh_layer_id equal to n, the layer having a nuh_layer_id equal to m may be the same as the (direct or indirect) reference layer of the layer having a nuh_layer_id equal to n or the layer having a nuh_layer_id equal to m.
[0176] In one embodiment, the PPS (RBSP) may be made available to the decoding process before it is referenced, included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit, or provided via external means.
[0177] In one embodiment, the PPS (RBSP) may be made available to the decoding process before it is referenced, included in at least one AU having a TemporalId equal to the TemporalId of the PPS NAL unit within the CVS that includes one or more PHs (or coded slice NAL units) that reference the PPS, or provided via external means.
[0178] In one embodiment, the PPS (RBSP) is made available for the decoding process before it is referenced by one or more PHs (or coded slice NAL units), and is included in at least one PU having a nuh_layer_id equal to the lowest nuh_layer_id of the coded slice NAL units that reference the PPS NAL unit within the CVS that includes one or more PHs (or coded slice NAL units) that reference the PPS, or can be provided via external means.
[0179] In one embodiment, the PPS (RBSP) is made available for the decoding process before it is referenced by one or more PHs (or coded slice NAL units), and is included in at least one PU having a TemporalId equal to the TemporalId of the PPS NAL unit and a nuh_layer_id equal to the lowest nuh_layer_id of the coded slice NAL units that reference the PPS NAL unit within the CVS that includes one or more PHs (or coded slice NAL units) that reference the PPS, or can be provided via external means.
[0180] In the same or another embodiment, the ph_pic_parameter_set_id within the PH specifies the value of the pps_pic_parameter_set_id for the reference PPS used. The value of the pps_seq_parameter_set_id can be the same for all PPSs referenced by the coded pictures within the CLVS.
[0181] In the same or another embodiment, all PPS NAL units having a specific value of pps_pic_parameter_set_id within the PU have the same content.
[0182] In the same or another embodiment, regardless of the nuh_layer_id value, the PPS NAL units can share the same value space of pps_pic_parameter_set_id.
[0183] In the same embodiment or another embodiment, the nuh_layer_id value of the PPS NAL unit may be equal to the lowest nuh_layer_id value of the coded slice NAL unit that references the NAL unit that references the PPS NAL unit.
[0184] In one embodiment, when a PPS having an nuh_layer_id equal to m is referenced by one or more coded slice NAL units having an nuh_layer_id equal to n, the layer having an nuh_layer_id equal to m may be the same as the (direct or indirect) reference layer of the layer having an nuh_layer_id equal to n or the layer having an nuh_layer_id equal to m.
[0185] In one embodiment, when the flag no_temporal_sublayer_switching_flag is signaled within the DPS, VPS, or SPS, the TemporalId value of the PPS that references the parameter set including the flag equal to 1 may be equal to 0, and the TemporalId value of the PPS that references the parameter set including the flag equal to 1 may be equal to or greater than the TemporalId value of the parameter set.
[0186] In one embodiment, each PPS (RBSP) is made available for use in the decoding process before it is referenced and is included in at least one AU having a TemporalId less than or equal to the TemporalId of the encoded slice NAL unit (or PH NAL unit) that references it, or can be provided via external means. When the PPS NAL unit is included in an AU that is earlier than the AU containing the encoded slice NAL unit that references the PPS, a VCL NAL unit that enables temporal up-layer switching, or a VCL NAL unit having a nal_unit_type equal to STSA_NUT (which indicates that the picture in the VCL NAL unit can be a stepwise temporal sub-layer access (STSA) picture) may not exist after the PPS NAL unit and before the encoded slice NAL unit that references the PPS.
[0187] In the same embodiment or another embodiment, the PPS NAL unit and the encoded slice NAL unit (and its PH NAL unit) that reference the PPS can be included in the same AU.
[0188] In the same embodiment or another embodiment, the PPS NAL unit and the STSA NAL unit that reference the PPS can be included in the same AU that precedes the encoded slice NAL unit (and its PH NAL unit).
[0189] In the same embodiment or another embodiment, the STSA NAL unit, the PPS NAL unit, and the encoded slice NAL unit (and its PH NAL unit) that reference the PPS can exist within the same AU.
[0190] In the same embodiment or another embodiment, the TemporalId value of the VCL NAL unit containing the PPS can be equal to the TemporalId value of the preceding STSA NAL unit.
[0191] In the same embodiment, the picture order count (POC) value of the PPS NAL unit may be equal to or greater than the POC value of the STSA NAL unit.
[0192] In the same embodiment, the picture order count (POC) value of an encoded slice or PH NAL unit that references a PPS NAL unit may be equal to or greater than the POC value of the referenced PPS NAL unit.
[0193] In one embodiment, since all VCL NAL units within an AU have the same TemporalId value, the value of sps_max_sublayers_minus1 is the same across all layers in the encoded video sequence. The value of sps_max_sublayers_minus1 is the same in all SPSs referenced by the encoded pictures within the CVS.
[0194] In one embodiment, assuming layer A is the direct reference layer of layer B, the chroma_format_idc value of the SPS referenced by one or more encoded pictures within layer A is equal to the chroma_format_idc value in the SPS referenced by one or more encoded pictures within layer B. This is because any encoded picture has the same chroma_format_idc value as its reference picture. The chroma_format_idc value of the SPS referenced by one or more encoded pictures within layer A is equal to the chroma_format_idc value in the SPS referenced by one or more encoded pictures within the direct reference layer of layer A within the CVS.
[0195] In one embodiment, assuming that layer A is the direct reference layer of layer B, the values of subpics_present_flag and sps_subpic_id_present_flag in the SPS referred to by one or more coded pictures in layer A are equal to the values of subpics_present_flag and sps_subpic_id_present_flag in the SPS referred to by one or more coded pictures in layer B. This is because the subpicture layout needs to be aligned or associated between layers. Otherwise, subpictures with multiple layers may not be accurately extracted. The values of subpics_present_flag and sps_subpic_id_present_flag in the SPS referred to by one or more coded pictures in layer A are equal to the values of subpics_present_flag and sps_subpic_id_present_flag in the SPS referred to by one or more coded pictures in the direct reference layer of layer A within the CVS.
[0196] In one embodiment, when a STSA picture in layer A is referred to by a picture in the direct reference layer of layer A within the same AU, the picture referring to the STSA picture is a STSA picture. Otherwise, time sublayer switch-up cannot be synchronized between layers. When a STSA NAL unit in layer A is referred to by a VCL NAL unit in the direct reference layer of layer A within the same AU, the nal_unit_type value of the VCL NAL unit referring to the STSA NAL unit is equal to STSA_NUT.
[0197] In one embodiment, when a RASL picture in layer A is referenced by a picture in the direct reference layer of layer A within the same AU, the picture referencing the RASL picture is a RASL picture. Otherwise, the picture cannot be correctly decoded. When a RASL NAL unit in layer A is referenced by a VCL NAL unit in the direct reference layer of layer A within the same AU, the nal_unit_type value of the VCL NAL unit referencing the RASL NAL unit is equal to RASL_NUT.
[0198] The technology for signaling the above adaptive resolution parameters can be implemented as computer software using computer-readable instructions and can also be physically stored on one or more computer-readable media. For example, FIG. 7 shows a computer system 700 suitable for implementing a particular embodiment of the matters disclosed herein.
[0199] The computer software can be coded using any suitable machine code or computer language such that, when subjected to assembly, compilation, linking, or similar mechanisms, it can create code having instructions executable directly or via interpretation, microcode execution, and the like by a computer central processing unit (CPU), a graphics processing unit (GPU), and the like.
[0200] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0201] The components shown in FIG. 7 with respect to computer system 700 are exemplary in nature and are not intended to suggest any limitation as to the use or functionality of the computer software implementing embodiments of the present disclosure. Also, the component configuration should not be construed as having any dependency or requirement with respect to any one or combination of the components shown in this exemplary embodiment of computer system 700.
[0202] Computer system 700 may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, moving a data glove, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), olfactory input (not shown). The human interface device may also be used to capture certain media that is not necessarily directly related to conscious input by a human, such as, for example, audio (e.g., conversation, music, ambient sound, etc.), images (e.g., scanned images, photographic images obtained from a still camera, etc.), video (e.g., 2D video, 3D video including stereoscopic video, etc.).
[0203] The input human interface device may include one or more of keyboard 701, mouse 702, trackpad 703, touch screen 710, data glove 704, joystick 705, microphone 706, scanner 707, camera 708 (each shown only once).
[0204] Computer system 700 may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by touch screen 710, data glove 704, or joystick 705, although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speaker 709, headphones (not shown), etc.), visual output devices (e.g., screens 710 including CRT screens, LCD screens, plasma screens, OLED screens (each may or may not have a touch screen input function, each may or may not have a tactile feedback function. Some of these may be capable of outputting two-dimensional visual output or outputting output of four dimensions or more through means such as stereoscopic output, etc.), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and printers (not shown).
[0205] Computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 720 having a CD / DVD or similar medium 721, thumb drive 722, removable hard drive or solid state drive 723, legacy magnetic media such as tapes and floppy disks (registered trademark, not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0206] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the matters disclosed herein does not include transmission media, carrier waves, or other transient signals.
[0207] The computer system 700 may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet (registered trademark), wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus, and the like. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (749) (e.g., a USB port of the computer system 700), while others are generally integrated into the core of the computer system 700 by attachment to the system bus described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 700 can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANbus to a specific CANbus device), or bidirectional, for example, to other computer systems using a local or wide area digital network. Specific protocols and protocol stacks can be used on each of the networks and network interfaces as described above.
[0208] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core 740 of the computer system 700.
[0209] The core 740 may include one or more central processing units (CPUs) 741, a graphics processing unit (GPU) 742, a special programmable processing unit in the form of a field programmable gate array (FPGA) 743, a hardware accelerator 744 for specific tasks, and the like. These devices may be connected via a system bus 748 together with internal mass storage 747 such as read-only memory (ROM) 745, random access memory 746, for example, internal hard drives, SSDs, and the like that are not accessible to internal users 747. In some computer systems, the system bus 748 may be made accessible in the form of one or more physical plugs to allow for expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached either directly to the core's system bus 748 or via a peripheral bus 749. Peripheral bus architectures include PCI, USB, and the like.
[0210] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may execute specific instructions that can be combined to form the aforementioned computer code. The computer code may be stored in the ROM 745 or the RAM 746. Transient data can also be stored in the RAM 746, and permanent data can be stored, for example, in the internal mass storage 747. Fast storage and retrieval to any of the memory devices may be enabled by the use of cache memory that may be associated near one or more CPUs 741, GPUs 742, mass storage 747, ROM 745, RAM 746, and the like.
[0211] A computer-readable medium can have computer code thereon for performing various computer-implemented processes. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure or, alternatively, they may be of the kind well known and available to those skilled in the computer software arts.
[0212] As an example, and not by way of limitation, a computer system having an architecture 700, particularly a core 740, can provide functionality as a result of software embodied in one or more tangible computer-readable media being executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, and the like). Such computer-readable media can be specific storage of the core 740 that is non-transitory in nature, such as the mass storage 747 inside the core or the ROM 745, and media associated with user-accessible mass storage as introduced above. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 740. The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core 740 and particularly the processors therein (including CPUs, GPUs, FPGAs, and the like) to define data structures stored in the RAM 746 and modify such data structures according to processes defined by the software, thereby executing the specific processes described herein or specific parts of the specific processes. Additionally, or alternatively, the computer system can provide functionality as a result of logic wired or otherwise embodied in a circuit (e.g., an accelerator 744) that operates instead of or in conjunction with software to execute the specific processes described herein or specific parts of the specific processes. References to software include logic, and vice versa where appropriate. References to computer-readable media can include circuits (e.g., integrated circuits (ICs), etc.) storing software for execution, circuits embodying logic for execution, or both where appropriate. The present disclosure includes suitable combinations of hardware and software.
[0213] While this disclosure describes several exemplary embodiments, there are changes, substitutions, and various equivalent alternatives that fall within the scope of the disclosure. Thus, it is understood that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly illustrated or described herein, embody the principles of the disclosure and are therefore within its spirit and scope.
Claims
1. 1. A method for video coding executed by at least one processor, comprising: Parsing a bitstream associated with the coded video data; determining whether a picture order count value increases uniformly for each access unit based on a value of a first syntax element in a video parameter set (VPS); determining an access unit count for each access unit based on dividing the picture order count value by a picture order count cycle access unit value based on determining that the picture order count value increases uniformly for each access unit; determining the access unit count for each of the access units based on dividing the picture order count value by a slice picture order count cycle access unit value based on determining that the picture order count value does not increase uniformly from access unit to access unit; decoding the coded video data based on the access unit count; A method having the following.
2. The method of claim 1 , wherein the picture order count cycle access unit value is signaled within the VPS.
3. The method of claim 1 , wherein the slice picture order count cycle access unit value is signaled in a slice header.
4. 2. The method of claim 1, wherein the slice picture order count cycle access unit value is not explicitly signaled within the coded video data when the picture order count value increases uniformly per access unit.
5. 2. The method of claim 1, wherein a VPS access unit count value is not explicitly signaled within the coded video data when the picture order count value does not increase uniformly per access unit.
6. The method of claim 1 , wherein each picture in the same access unit has the same temporal identifier.
7. The method of claim 6 , wherein each picture within the same access unit is associated with the same temporal sub-layer.
8. The method of claim 1 , wherein one or more pictures within the same access unit have different spatial layer identifiers.
9. The method of claim 1 , wherein at least a subset of pictures associated with the same access unit are decoded in parallel and output at the same time instance.
10. 1. A computer system for video coding, comprising: one or more memories storing a computer program; one or more processors; and The computer program causes the one or more processors to perform the method of any one of claims 1 to 9. Computer system.
11. A computer program causing a computer to carry out the method according to any one of claims 1 to 9.