Signaling subpicture id in subpicture-based video coding
By setting flags in the SPS or PPS to manage sub-picture ID changes, the method enhances coding efficiency and reduces redundancy in video coding, especially in scenarios involving sub-bitstream extraction and merging, resulting in improved user experience.
Patent Information
- Application Number
- JP2025227104
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-17
- Filing Date
- 2025-12-03
- Publication Date
- 2026-02-27
AI Technical Summary
Existing video coding technologies face challenges in efficiently signaling sub-picture identifiers (IDs) when they change within a coded video sequence, particularly in scenarios involving both sub-bitstream extraction and sub-bitstream merging, leading to redundancy and reduced coding efficiency.
Setting a flag in the sequence parameter set (SPS) or picture parameter set (PPS) to indicate whether sub-picture IDs in the CVS may change and where they are located, allowing for efficient signaling of sub-picture IDs.
This approach reduces redundancy and increases coding efficiency, providing a better user experience by improving video coding processes for transmission, reception, and viewing.
Smart Images

Figure 2026034483000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 901,552, entitled "Signaling Subpicture IDs in Subpicture Based Video Coding," filed September 17, 2019, by Ye-Kui Wang, which is incorporated herein by reference.
[0002] FIELD This disclosure relates generally to video coding, and more particularly to signaling sub-picture identifiers (IDs) in sub-picture-based video coding. [Background technology]
[0003] The amount of video data required to render even a relatively short video can be considerable, which can pose challenges when the data is to be streamed or otherwise transmitted over communication networks with limited bandwidth capacity. Therefore, video data is typically compressed before being transmitted over modern telecommunications networks. Because memory resources can be limited, video size can also be an issue when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data needed to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and an ever-increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little or no sacrifice in image quality are desirable. Summary of the Invention [Means for solving the problem]
[0004] A first aspect relates to a method performed by a decoder, comprising: receiving, by the decoder, a bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with sub-picture identifier (ID) mappings, wherein the SPS includes an SPS flag; determining, by the decoder, whether the SPS flag has a first value or a second value, wherein an SPS flag having the first value specifies that the sub-picture ID mappings are signaled in the SPS and an SPS flag having the second value specifies that the sub-picture ID mappings are signaled in the PPS; obtaining, by the decoder, the sub-picture ID mappings from the SPS when the SPS flag has the first value, or from the PPS when the SPS flag has the second value; and decoding, by the decoder, the plurality of sub-pictures using the sub-picture ID mappings.
[0005] The method provides a technique to ensure efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, coders / decoders (also known as "codecs") in video coding are improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0006] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the PPS includes a PPS flag, and the method further comprises: determining, by a decoder, whether the PPS flag has a first value or a second value, wherein a PPS flag having the first value specifies that sub-picture ID mapping is signaled in the PPS, and a PPS flag having the second value specifies that sub-picture ID mapping is not signaled in the PPS; and obtaining, by the decoder, the sub-picture ID mapping from the PPS when the PPS flag has the first value.
[0007] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value.
[0008] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the first value is 1 and the second value is 0.
[0009] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the SPS includes a second SPS flag, and the second SPS flag specifies whether sub-picture ID mapping is explicitly signaled within the SPS or the PPS.
[0010] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping is allowed to change within the coded video sequence (CVS) of the bitstream.
[0011] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream comprises a merged bitstream, and the sub-picture ID mapping has changed within the CVS of the bitstream.
[0012] A second aspect relates to a method performed by an encoder, comprising: encoding, by a decoder, a bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with sub-picture identifier (ID) mappings, where the SPS includes an SPS flag; setting, by the decoder, the SPS flag to a first value when the sub-picture ID mappings are signaled in the SPS, and to a second value when the sub-picture ID mappings are signaled in the PPS; and storing, by the decoder, the bitstream for communication to the decoder.
[0013] The method provides a technique for ensuring efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, the coder / decoder (also known as a "codec") in video coding is improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0014] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides the decoder with a PPS flag in the PPS that is set to a first value when sub-picture ID mapping is signaled in the PPS, and to a second value when sub-picture ID mapping is not signaled in the PPS.
[0015] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value.
[0016] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the first value is 1 and the second value is 0.
[0017] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the SPS includes a second SPS flag, and the second SPS flag specifies whether sub-picture mapping is explicitly signaled within the SPS or the PPS.
[0018] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping may change within the coded video sequence (CVS) of the bitstream.
[0019] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream comprises a merged bitstream, and the sub-picture ID mapping has changed within the CVS of the bitstream.
[0020] A third aspect relates to a decoding device comprising: a receiver configured to receive a video bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with sub-picture identifier (ID) mappings, wherein the SPS includes an SPS flag; a memory coupled to the receiver, the memory storing instructions; and a processor coupled to the memory, the processor executing the instructions to cause a decoding device to determine whether the SPS flag has a first value or a second value, wherein an SPS flag having the first value specifies that a sub-picture ID mapping is signaled in the SPS and an SPS flag having the second value specifies that the sub-picture ID mapping is signaled in the PPS; obtain the sub-picture ID mapping from the SPS when the SPS flag has the first value and from the PPS when the SPS flag has the second value; and decode the plurality of sub-pictures using the sub-picture ID mappings.
[0021] A decoding device provides a technique for ensuring efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, coders / decoders (also known as "codecs") in video coding are improved over current codecs. As a practical matter, an improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0022] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that when the PPS flag has a second value, the SPS flag has a first value, and when the PPS flag has the first value, the SPS flag has a second value, the first value being 1 and the second value being 0.
[0023] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the SPS includes a second SPS flag, and the second SPS flag specifies whether sub-picture mapping is explicitly signaled within the SPS or the PPS.
[0024] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping may change within the coded video sequence (CVS) of the bitstream.
[0025] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream comprises a merged bitstream, and the sub-picture ID mapping has changed within the CVS of the bitstream.
[0026] A fourth aspect relates to an encoding device comprising: a memory including instructions; a processor coupled to the memory; the processor configured to execute the instructions to cause the encoding device to encode a bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with sub-picture identifier (ID) mappings, where the SPS includes an SPS flag; setting the SPS flag to a first value when the sub-picture ID mapping is signaled in the SPS and to a second value when the sub-picture ID mapping is signaled in the PPS; and setting the PPS flag in the PPS to the first value when the sub-picture ID mapping is signaled in the PPS and to the second value when the sub-picture ID mapping is not signaled in the PPS; and a transmitter coupled to the processor; the transmitter configured to transmit the bitstream to a video decoder.
[0027] The encoding device provides a technique for ensuring efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or a picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, the coder / decoder (also known as a "codec") in video coding is improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0028] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value.
[0029] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the first value is 1 and the second value is 0.
[0030] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping may change within the coded video sequence (CVS) of the bitstream.
[0031] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the PPS ID specifies a value of a second PPS ID for the PPS in use, and the second PPS ID identifies the PPS for reference by the syntax element.
[0032] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the bitstream comprises a merged bitstream, and the sub-picture ID mapping has changed within the CVS of the bitstream.
[0033] A fifth aspect relates to a coding apparatus including: a receiver configured to receive pictures to be encoded or to receive a bitstream to be decoded, a transmitter coupled to the receiver and configured to transmit the bitstream to a decoder or to transmit decoded images to a display, a memory coupled to at least one of the receiver or the transmitter and configured to store instructions, and a processor coupled to the memory and configured to execute the instructions stored in the memory to perform any of the methods disclosed herein.
[0034] The coding apparatus provides a technique for ensuring efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, the coder / decoder (also known as a "codec") in video coding is improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0035] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides a display configured to display the decoded picture.
[0036] A sixth aspect relates to a system, the system including an encoder and a decoder in communication with the encoder, the encoder or decoder including a decoding device, encoding device, or coding apparatus disclosed herein.
[0037] The system provides techniques to ensure efficient signaling of sub-picture identifiers (IDs) even when they change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where they are located. This reduces redundancy and increases coding efficiency. Therefore, coders / decoders (also known as "codecs") in video coding are improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0038] A seventh aspect relates to a means for coding, the means for coding including receiving means configured to receive a picture to be encoded or to receive a bitstream to be decoded, transmitting means coupled to the receiving means, the transmitting means configured to transmit the bitstream to the decoding means or to transmit a decoded image to the display means, storage means coupled to at least one of the receiving means or the transmitting means, the storage means configured to store instructions, and processing means coupled to the storage means, the processing means configured to execute the instructions stored in the storage means to perform any of the methods disclosed herein.
[0039] The coding method provides a technique for ensuring efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or a picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, the coder / decoder (also known as a "codec") in video coding is improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0040] For purposes of clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.
[0041] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.
[0042] For a more complete understanding of this disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts. [Brief explanation of the drawings]
[0043] [Figure 1] 1 is a flowchart of an example method for coding a video signal. [Figure 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding. [Figure 3]FIG. 1 is a schematic diagram illustrating an example video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an example video decoder. [Figure 5] FIG. 1 is a schematic diagram illustrating an example bitstream and sub-bitstreams extracted from the bitstream. [Figure 6A] 1 illustrates an example mechanism for creating an extractor track for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6B] 1 illustrates an example mechanism for creating an extractor track for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6C] 1 illustrates an example mechanism for creating an extractor track for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6D] 1 illustrates an example mechanism for creating an extractor track for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6E] 1 illustrates an example mechanism for creating an extractor track for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 7A] 1 illustrates an example mechanism for updating extractor tracks based on changes in viewing orientation in a VR application. [Figure 7B] 1 illustrates an example mechanism for updating extractor tracks based on changes in viewing orientation in a VR application. [Figure 7C] 1 illustrates an example mechanism for updating extractor tracks based on changes in viewing orientation in a VR application. [Figure 7D] 1 illustrates an example mechanism for updating extractor tracks based on changes in viewing orientation in a VR application. [Figure 7E] 1 illustrates an example mechanism for updating extractor tracks based on changes in viewing orientation in a VR application. [Figure 8] FIG. 1 illustrates an embodiment of a method for decoding a coded video bitstream. [Figure 9] FIG. 1 illustrates an embodiment of a method for encoding a coded video bitstream. [Figure 10] 1 is a schematic diagram of a video coding device. [Figure 11] FIG. 1 is a schematic diagram of an embodiment of a means for coding; DETAILED DESCRIPTION OF THE INVENTION
[0044] While example implementations of one or more embodiments are provided below, it should be understood at the outset that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The disclosure should not be limited in any way to the example implementations, diagrams, and techniques illustrated below, including the example designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.
[0045] The following terms are defined as follows, unless used in a contrary context herein. Specifically, the following definitions are intended to provide additional clarity to this disclosure. However, terms may be explained differently in different contexts. Therefore, the following definitions should be considered supplemental, and not limiting, of any other definitions provided for such terms herein.
[0046] A bitstream is a sequence of bits containing video data compressed for transmission between an encoder and a decoder. An encoder is a device configured to compress video data into a bitstream using an encoding process. A decoder is a device configured to reconstruct video data from a bitstream for display using a decoding process. A picture is an array of luma samples and / or chroma samples that make up a frame or a field thereof. The picture being encoded or decoded may be referred to as the current picture for clarity of description. A reference picture is a picture containing reference samples that may be used when coding other pictures by reference according to inter-prediction and / or inter-layer prediction. A reference picture list is a list of reference pictures used for inter-prediction and / or inter-layer prediction. Some video coding systems utilize two reference picture lists, which may be denoted as Reference Picture List 1 and Reference Picture List 0. A reference picture list structure is an addressable syntax structure that contains multiple reference picture lists. Inter-prediction is a mechanism for coding samples of a current picture by reference to indicated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer. A reference picture list structure entry is an addressable location within the reference picture list structure that indicates a reference picture associated with the reference picture list. A slice header is a part of a coded slice that contains data elements pertaining to all video data within tiles represented in the slice. A sequence parameter set (SPS) is a parameter set that contains data related to a sequence of pictures. A picture parameter set (PPS) is a syntax structure that contains syntax elements that apply to zero or more entire coded pictures, as determined by the syntax elements found in each picture header.
[0047] A flag is a variable or single-bit syntax element that can take one of two possible values: 0 and 1. A subpicture is a rectangular region of one or more slices within a picture. A subpicture identifier (ID) is a number, letter, or other indicator that uniquely identifies a subpicture. A subpicture ID (also known as a tile identifier) is used to identify a specific subpicture using a subpicture index, which may be referred to herein as a subpicture ID mapping. In other words, a subpicture ID mapping is a table that provides a one-to-one mapping between a list of subpicture indexes and subpicture IDs. That is, a subpicture ID mapping provides a separate subpicture ID for each subpicture.
[0048] An access unit (AU) is a set of one or more coded pictures associated with the same display time (e.g., the same picture order count) for output from a decoded picture buffer (DPB) (e.g., for display to a user). An access unit delimiter (AUD) is an indicator or data structure used to mark the start of an AU or the boundary between AUs. A decoded video sequence is a sequence of pictures reconstructed by a decoder in preparation for display to a user.
[0049] A CVS is a sequence of access units (AUs) that includes, in decoding order, a coded video sequence start (CVSS) AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AUs that are CVSS AUs. A CVSS AU is an AU where there is a prediction unit (PU) for each layer specified by a video parameter set (VPS), and the coded pictures within each PU are CVS pictures. In one embodiment, each picture is within an AU. A PU is a set of Network Abstraction Layer (NAL) units associated with each other according to specified classification rules, are contiguous in decoding order, and contain exactly one coded picture.
[0050] The following acronyms are used: Adaptive Loop Filter (ALF), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Instantaneous Decoding Refresh (IDR), Intra-Random Access Point (IRAP), Joint Video Experts Team (JVET), Least Significant Bit (LSB), Most Significant Bit (MSB), Motion-Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Picture Parameter Set (PPS), Raw Byte Sequence Payload (RBS), and Raw Byte Sequence Payload (RMS). The following are used herein: Frame Reconstruction Payload (RBSP), Sample Adaptive Offset (SAO), Sequence Parameter Set (SPS), Temporal Motion Vector Prediction (TMVP), Versatile Video Coding (VVC), and Working Draft (WD).
[0051] Many video compression techniques can be utilized to reduce the size of video files with minimal loss of data. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) can be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a picture can be coded by utilizing spatial prediction with respect to reference samples in neighboring blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or image, and a reference picture may be referred to as a reference frame and / or image. Spatial or temporal prediction results in a predictive block that represents an image block. Residual data represents pixel differences between the original image block and the predictive block. Thus, inter-coded blocks are encoded according to a motion vector that points to a block of reference samples that form the predictive block, and residual data that indicates the difference between the coded block and the predictive block. Intra-coded blocks are encoded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain. These result in residual transform coefficients that may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to create a one-dimensional vector of transform coefficients.Entropy coding may be applied to achieve even greater compression. Such video compression techniques are discussed in more detail below.
[0052] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to a corresponding video coding standard, including International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).
[0053] There is also an emerging video coding standard named Versatile Video Coding (VVC) being developed by the ITU-T and ISO / IEC Joint Video Experts Team (JVET). The VVC standard has several working drafts, but one working draft (WD) of VVC in particular is referenced here: B. Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 5)," JVET-N1001-v3, 13th JVET Meeting, March 27, 2019 (VVC Draft 5).
[0054] HEVC includes four different picture partitioning schemes, namely, normal slice, dependent slice, tile, and Wavefront Parallel Processing (WPP), which can be applied for maximum transmission unit (MTU) size alignment, parallel processing, and reduced end-to-end delay.
[0055] Normal slices are similar to those in H.264 / AVC. Each normal slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, normal slices can be reconstructed independently from other normal slices in the same picture (although there may still be interdependencies due to loop filtering operations).
[0056] Regular slices are the only tool in H.264 / AVC that can be used for parallelization that is also available in a substantially identical format. Regular slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is typically much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, the use of regular slices can incur significant coding overhead due to the bit cost of slice headers and the lack of prediction across slice boundaries. Furthermore, due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated within its own NAL unit, regular slices also serve as the primary mechanism for bitstream partitioning to align with MTU size requirements (as opposed to the other tools described below). In many cases, the goals of parallelization and MTU size alignment impose conflicting demands on slice layout within a picture. The realization of this situation led to the development of the parallelization tools described below.
[0057] Dependent slices have short slice headers and allow for partitioning of the bitstream at treeblock boundaries without breaking any intra-picture prediction. Essentially, dependent slices provide fragmentation of a normal slice into multiple NAL units to provide reduced end-to-end delay by allowing parts of the normal slice to be sent out before the encoding of the entire normal slice is finished.
[0058] In WPP, a picture is partitioned into a single row of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows; the start of decoding a CTB row is delayed by two CTBs to ensure that data associated with the CTBs above and to the right of the current CTB is available before the current CTB is decoded. Using this staggered start (which looks like a wavefront when represented diagrammatically), parallelization is possible using up to the same number of processors / cores as the picture contains CTB rows. Because intra-picture prediction between adjacent tree block rows within a picture is allowed, the inter-processor / inter-core communication required to enable intra-picture prediction can be significant. WPP partitioning does not result in the generation of additional NAL units compared to when it is not applied, and therefore WPP is not a tool for MTU size alignment. However, if MTU size alignment is required, regular slicing can be used with WPP, with some coding overhead.
[0059] Tiles define horizontal and vertical boundaries that partition a picture into tile columns and rows. The scan order of the CTBs is changed to be local within the tile (in the tile's CTB raster scan order) before decoding the top-left CTB of the next tile in the picture's tile raster scan order. Similar to regular slices, tiles break intra-picture prediction dependencies as well as entropy decoding dependencies. However, tiles are not required to be contained in individual NAL units (as with WPP), and therefore tiles cannot be used for MTU size alignment. Each tile can be processed by a single processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent tiles is limited to conveying a shared slice header if the slice spans more than one tile and loop filtering-related sharing of reconstructed samples and metadata. When a slice contains more than one tile or WPP segment, the entry point byte offset for each tile or WPP segment, other than the initial entry point byte offset within the slice, is signaled in the slice header.
[0060] For simplicity, constraints on the application of four different picture partitioning schemes are specified in HEVC. A given coded video sequence cannot contain both tiles and wavefronts for most of the profiles specified in HEVC. For each slice and tile, one or both of the following conditions must be met: 1) all coded treeblocks in a slice belong to the same tile; 2) all coded treeblocks in a tile belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is in use, if a slice starts within a CTB row, it must end within the same CTB row.
[0061] Recent amendments to HEVC are specified in JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, GJ Sullivan, A. Tourapis, and Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)", October 24, 2017, publicly available at http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. With this amendment, HEVC specifies three MCTS-related supplemental enhancement information (SEI) messages: the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nesting SEI message.
[0062] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to point to full sample positions within the MCTS and fractional sample positions that require only full sample positions within the MCTS for interpolation, and the use of motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction is not permitted. In this way, each MCTS can be decoded independently without the presence of tiles not included in the MCTS.
[0063] The MCTS Extraction Information Set SEI message provides supplemental information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for the MCTS set. The information consists of several extraction information sets, each of which defines several MCTS sets and includes the RBSP bytes of the replacement video parameter set (VPS), SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) will typically need to have different values, so the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated.
[0064] In the latest VVC draft specification, a picture may be partitioned into multiple sub-pictures, each covering a rectangular area and containing an integer number of complete slices. The sub-picture partitioning persists across all pictures in a coded video sequence (CVS), and the partitioning information (i.e., sub-picture position and size) is signaled in the SPS. A sub-picture may be indicated as being coded without using sample values from any other sub-picture for motion compensation.
[0065] In JVET contribution JVET-O0141, publicly available at http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 13_Marrakech / wg11 / JVET-M0261-v1.zip, the subpicture design is similar to that in the latest VVC draft specification with some differences, one of which is explicit signaling of the subpicture ID for each subpicture in the SPS, along with signaling of the subpicture ID in each slice header, to enable subpicture-based sub-bitstream extraction without having to modify the coded slice NAL units. In this approach, the subpicture ID for each subpicture remains constant across all pictures in the CVS.
[0066] During the previous discussion, it was mentioned that a different approach for signaling sub-picture IDs should be used to enable sub-picture-based sub-bitstream merging without the need to modify the coded slice NAL units.
[0067] For example, as described more fully below, in viewport-dependent 360-degree video streaming, the client selects which subpictures are received and merged into the bitstream to be decoded. When the viewing orientation changes, some subpicture ID values may remain the same, while other subpicture ID values are changed, compared to the subpicture ID values received before the viewing orientation changed. Thus, when the viewing orientation changes, the subpicture IDs that the decoder receives in the merged bitstream change.
[0068] One approach that was considered is as follows.
[0069] Subpicture locations and sizes are indicated in the SPS, for example, using the syntax in the latest VVC draft specification. Subpicture indices are assigned (0 to N-1, where N is the number of subpictures). A mapping of subpicture IDs to subpicture indices exists in the PPS because the mapping may need to be changed in the middle of the CVS. This can be a loop i=0 to N-1, providing subpicture_id[i] that maps to subpicture index i. This subpicture ID mapping needs to be rewritten by the client when it selects a new set of subpictures to be decoded. The slice header contains the encoder-selected subpicture ID (e.g., subpicture_id) (and this does not need to be rewritten during sub-bitstream merging).
[0070] Unfortunately, there are problems with existing sub-picture ID signaling approaches. Many application scenarios using sub-picture-based video coding involve sub-bitstream extraction but not sub-bitstream merging. For example, each extracted sub-bitstream may be decoded by its own decoder instance. Therefore, there is no need to merge the extracted sub-bitstreams into one bitstream to be decoded by only one decoder instance. In these scenarios, the sub-picture ID for each sub-picture does not change in the CVS. Therefore, the sub-picture ID can be signaled in the SPS. Signaling such sub-picture IDs in the SPS instead of the PPS is beneficial both from the perspective of bit saving and session negotiation.
[0071] In application scenarios involving both sub-bitstream extraction and sub-bitstream merging, the sub-picture IDs do not change in the CVS in the original bitstream. Therefore, signaling the sub-picture IDs in the SPS allows the use of the sub-picture IDs both in the SPS and in the slice headers, which is useful for the sub-bitstream extraction part of the process. However, this approach is only feasible when the sub-picture IDs do not change in the CVS in the original bitstream.
[0072] Disclosed herein is a technique for ensuring efficient signaling of sub-picture identifiers (IDs) even when the IDs change within a coded video sequence (CVS) in application scenarios involving both sub-bitstream extraction and sub-bitstream merging. Efficient signaling is achieved by setting a flag in a sequence parameter set (SPS) or picture parameter set (SPS) to indicate whether the sub-picture IDs in the CVS may change and, if so, where the sub-picture IDs are located. This reduces redundancy and increases coding efficiency. Therefore, coders / decoders (also known as "codecs") in video coding are improved over current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0073] 1 is a flowchart of an example operational method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size allows the compressed video file to be transmitted to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to coherently reconstruct the video signal.
[0074] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed sequentially, create the visual impression of movement. The frames include pixels represented in terms of light, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0075] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels by 64 pixels). CTUs contain both luma and chroma samples. A coding tree can be used to divide the CTUs into blocks and then recursively subdivide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of a frame can be subdivided until each block contains relatively homogeneous illumination values. Furthermore, the chroma component of a frame can be subdivided until each block contains relatively homogeneous color values. Thus, the partitioning scheme varies depending on the content of the video frame.
[0076] In step 105, various compression mechanisms are utilized to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be utilized. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Thus, a block depicting an object in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a constant position over multiple frames. Thus, the table may be described once, and adjacent frames can reference back to the reference frame. A pattern matching mechanism may be utilized to match objects across multiple frames. Furthermore, a moving object may be depicted across multiple frames, for example, due to object motion or camera motion. As a specific example, a video may depict a car moving across the screen over multiple frames. Motion vectors may be utilized to describe such motion. A motion vector is a two-dimensional vector that provides an offset from the coordinates of the object in a frame to the coordinates of the object in a reference frame. Therefore, inter-prediction may encode an image block in a current frame as a set of motion vectors indicating an offset from a corresponding block in a reference frame.
[0077] Intra prediction encodes blocks within a common frame. Intra prediction takes advantage of the fact that luma and chroma components tend to cluster within a frame. For example, a green spot in a tree tends to be located adjacent to similar green spots. Intra prediction utilizes multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar / the same as samples of neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. Planar mode effectively indicates a smooth transition of light / color across a row / column by utilizing a relatively constant gradient in changing values. DC mode is utilized for boundary smoothing, indicating that the block is similar / the same as the average value associated with samples of all neighboring blocks associated with the angular direction of the directional prediction mode. Thus, intra-predicted blocks can represent image blocks as values of various associated prediction modes instead of actual values. Furthermore, inter-predicted blocks can represent image blocks as values of motion vectors instead of actual values. In either case, the prediction block may in some cases not exactly represent the image block. Any differences are stored in a residual block. To further compress the file, a transform may be applied to the residual block.
[0078] In step 107, various filtering techniques may be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction discussed above may result in the generation of blocky images in a decoder. Furthermore, the block-based prediction scheme may encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to a block / frame. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference block, so that the artifacts are less likely to create additional artifacts in subsequent blocks that are encoded based on the reconstructed reference block.
[0079] Once the video signal has been segmented, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data discussed above along with any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder on demand. The bitstream may also be broadcast and / or multicast to multiple decoders. Generating the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously on multiple frames and blocks. The order depicted in FIG. 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to any particular order.
[0080] In step 111, a decoder receives the bitstream and begins the decoding process. Specifically, the decoder converts the bitstream into corresponding syntax and video data using an entropy decoding scheme. In step 111, the decoder uses syntax data from the bitstream to determine a partition for the frame. The partition should match the result of the block partitioning in step 103. Entropy encoding / decoding as used in step 111 is now described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible choices based on the spatial arrangement of values in the input image. Signaling the exact choice may utilize multiple bins. As used herein, a bin is a binary value (e.g., a bit value that can change depending on the situation) treated as a variable. Entropy coding allows the encoder to discard any options that are clearly not feasible for a particular case, leaving a set of allowable options. Each allowable option is then assigned a codeword. The length of the codeword is based on the number of allowable choices (e.g., one bin for two choices, two bins for three to four choices, etc.). The encoder then encodes a codeword for the selected choice. This scheme reduces the size of the codeword because the codeword is only as large as desired to uniquely indicate a choice from a small subset of allowable choices, as opposed to uniquely indicating a choice from a potentially large set of all possible choices. The decoder then decodes the choices by determining the set of allowable choices in a manner similar to the encoder. By determining the set of allowable choices, the decoder can read the codeword and determine the choices made by the encoder.
[0081] In step 113, the decoder performs block decoding. Specifically, the decoder generates a residual block using an inverse transform. Then, the decoder reconstructs an image block according to the partition using the residual block and a corresponding prediction block. The prediction block may include both intra-prediction blocks and inter-prediction blocks, such as those generated in the encoder in step 105. The reconstructed image block is then arranged into a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding, as discussed above.
[0082] In step 115, filtering is performed on the frames of the reconstructed video signal at the encoder in a manner similar to step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal can be output to a display in step 117 for viewing by an end user.
[0083] 2 is a schematic diagram of an example coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support the implementation of operational method 100. Codec system 200 is generalized to depict components utilized in both encoders and decoders. Codec system 200 receives and segments a video signal as discussed with respect to steps 101 and 103 in operational method 100, which results in a segmented video signal 201. When operating as an encoder as discussed with respect to steps 105, 107, and 109 in method 100, codec system 200 then compresses the segmented video signal 201 into a coded bitstream. When operating as a decoder, codec system 200 generates an output video signal from the bitstream as discussed with respect to steps 111, 113, 115, and 117 in operational method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, while dashed lines indicate the movement of control data that controls the operation of other components. The components of codec system 200 may all reside within an encoder. A decoder may include a subset of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are now described.
[0084] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree utilizes various partitioning modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. Blocks can be referred to as nodes in the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. In some cases, the partitioned blocks can be included in a coding unit (CU). For example, a CU can be a subpart of a coding unit (CTU) that includes a luma block, a red differential chroma (Cr) block, and a blue differential chroma (Cb) block, along with corresponding syntax instructions for the CU. Partitioning modes can include a binary tree (BT), a ternary tree (TT), and a quad tree (QT), which are used to partition a node into two, three, or four child nodes, respectively, of varying shapes depending on the partitioning mode used. The segmented video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0085] The generic coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the generic coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be made based on storage space / bandwidth availability and image resolution requirements. The generic coder control component 211 also manages buffer utilization, taking transmission rate into account, to mitigate buffer underrun and overrun issues. To manage these issues, the generic coder control component 211 manages segmentation, prediction, and filtering by other components. For example, the generic coder control component 211 may dynamically increase compression complexity to increase resolution and increase bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the generic coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality with bitrate concerns. The generic coder control component 211 produces control data that controls the operation of the other components. The control data is also forwarded to the header formatting and CABAC component 231 to be encoded into the bitstream for signaling parameters for decoding at the decoder.
[0086] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes to, for example, select an appropriate coding mode for each block of video data.
[0087] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are illustrated separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors, which estimate motion for a video block. A motion vector may indicate, for example, the displacement of a coded object relative to a predictive block. A predictive block is a block that is found to closely match a block to be coded in terms of pixel differences. A predictive block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference measures. HEVC utilizes several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into CTBs, which may then be divided into CBs for inclusion within a CU. A CU may be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and may select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance both coding efficiency (e.g., size of the final encoding) and quality of video reconstruction (e.g., amount of data loss due to compression).
[0088] In some examples, the codec system 200 may calculate values for sub-integer pixel locations of a reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values for quarter-pixel locations, eighth-pixel locations, or other fractional pixel locations of a reference picture. Accordingly, the motion estimation component 221 may perform motion searches for full-pixel and fractional pixel locations to output motion vectors with fractional-pixel accuracy. The motion estimation component 221 calculates motion vectors for PUs of video blocks in inter-coded slices by comparing the positions of the PUs with the positions of predictive blocks of the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting and CABAC component 231 for encoding and outputs motion to the motion compensation component 219.
[0089] The motion compensation performed by the motion compensation component 219 may involve fetching or generating a predictive block based on a motion vector determined by the motion estimation component 221. Again, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector for the PU of the current video block, the motion compensation component 219 may locate the predictive block to which the motion vector points. A residual video block is then formed by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predictive block and the residual block are forwarded to the transform scaling and quantization component 213.
[0090] The partitioned video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are illustrated separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block relative to blocks in the current frame as an alternative to the inter-prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames, as described above. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode to encode the current block from multiple tested intra-prediction modes. The selected intra-prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.
[0091] For example, the intra picture estimation component 215 may calculate rate-distortion values for various tested intra prediction modes using a rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to create the encoded block, along with the bit rate (e.g., number of bits) used to create the encoded block. The intra picture estimation component 215 may calculate a ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. In addition, the intra picture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0092] The intra-picture prediction component 217, when implemented in an encoder, may generate a residual block from the prediction block based on a selected intra-prediction mode determined by the intra-picture estimation component 215, or, when implemented in a decoder, may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both luma and chroma components.
[0093] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to create a video block comprising residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized with different granularity, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transform coefficients, which are forwarded to the header formatting and CABAC component 231 to be encoded into the bitstream.
[0094] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct a residual block in the pixel domain for later use, for example, as a reference block that may become a predictive block for another current block. The motion estimation component 221 and / or motion compensation component 219 may calculate a reference block by adding the residual block back to the corresponding predictive block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to mitigate artifacts created during scaling, quantization, and transform. Such artifacts may otherwise cause inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.
[0095] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 may be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are depicted separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes multiple parameters for adjusting how such filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such filter should be applied and sets the corresponding parameters. Such data is forwarded to the header formatting and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., on reconstructed pixel blocks) or in the frequency domain, depending on the example.
[0096] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores and forwards the reconstructed and filtered blocks as part of the output video signal toward the display. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0097] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to a decoder. Specifically, the header formatting and CABAC component 231 generates various headers to encode control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, along with residual data in the form of quantized transform coefficient data, are all encoded within the bitstream. The final bitstream contains all information desired by a decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra-prediction mode, indications of partition information, and so on. Such data may be encoded by utilizing entropy coding. For example, the information may be encoded by utilizing context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream may be transmitted to another device (eg, a video decoder) or archived for later transmission or retrieval.
[0098] 3 is a block diagram illustrating an example video encoder 300. Video encoder 300 may be utilized to implement the encoding functionality of codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of operating method 100. Encoder 300 segments an input video signal, resulting in a segmented video signal 301 that is substantially similar to segmented video signal 201. Segmented video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.
[0099] Specifically, the partitioned video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on a reference block in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual block. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with associated control data) are forwarded to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231 .
[0100] The transformed and quantized residual block and / or the corresponding prediction block are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter within the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as discussed with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0101] 4 is a block diagram illustrating an example video decoder 400. Video decoder 400 may be utilized to implement the decoding functionality of codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of operating method 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.
[0102] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may utilize header information to provide context for interpreting additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion information, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into a residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0103] The reconstructed residual block and / or predictive block are forwarded to the intra-picture prediction component 417 for reconstruction into an image block based on an intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 utilizes a prediction mode to locate a reference block within a frame and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and corresponding inter-prediction data are forwarded to the decoded picture buffer component 423 via an in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are forwarded to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and applies a residual block to the result to reconstruct an image block. The resulting reconstructed block may also be forwarded to the decoded picture buffer component 423 via an in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks, which can be reconstructed into frames via partition information. Such frames may also be arranged in a sequence. The sequence is output to a display as a reconstructed output video signal.
[0104] 5 is a schematic diagram illustrating an example bitstream 500 including an encoded video sequence. For example, bitstream 500 may be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400. As another example, bitstream 500 may be generated by an encoder in step 109 of method 100 for use by a decoder in step 111.
[0105] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPSs) 512, a tile group header 514, and image data 520. The SPS 510 includes sequence data common to all pictures in a video sequence included in the bitstream 500. Such data may include picture sizing, bit depth, coding tool parameters, bit rate constraints, etc. In one embodiment, the SPS 510 includes an SPS flag 565. In one embodiment, the SPS flag 565 has a first value (e.g., 1) or a second value (e.g., 0). An SPS flag 565 with a first value specifies that a sub-picture ID mapping 575 is signaled in the SPS 510, and an SPS flag 565 with a second value specifies that a sub-picture ID mapping 575 is signaled in the PPS 512.
[0106] The PPS 512 includes parameters that are specific to one or more corresponding pictures. Thus, each picture in a video sequence may reference one PPS 512. The PPS 512 may indicate available coding tools, quantization parameters, offsets, picture-specific coding tool parameters (e.g., filter control), etc. for tiles in the corresponding picture. In one embodiment, the PPS 512 includes a PPS flag 567. In one embodiment, the PPS flag 567 has a first value (e.g., 1) or a second value (e.g., 0). A PPS flag 567 with the first value specifies that a sub-picture ID mapping 575 is signaled in the PPS 512, and a PPS flag 567 with the second value specifies that a sub-picture ID mapping 575 is not signaled in the PPS 512.
[0107] The tile group header 514 contains parameters that are specific to each tile group in a picture. Thus, there may be one tile group header 514 for each tile group in a video sequence. The tile group header 514 may include tile group information, a picture order count (POC), a reference picture list, prediction weights, tile entry points, deblocking parameters, etc. It should be noted that some systems refer to the tile group header 514 as a slice header and use such information to support slices instead of tile groups.
[0108] The image data 520 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. Such image data 520 is sorted according to the partition used to partition the image before encoding. For example, the images in the image data 520 include one or more pictures 521. A sequence or series of pictures 527 may be referred to as a CVS 527. As used herein, a CVS 527 is a sequence of access units (AUs) including, in decoding order, a coded video sequence start (CVSS) AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AUs that are CVSS AUs. A CVSS AU is an AU where there is a prediction unit (PU) for each layer specified by a video parameter set (VPS), and the coded pictures within each PU are CVS pictures. In one embodiment, each picture 521 is within an AU. A PU is a set of NAL units that are associated with each other according to specified classification rules, are consecutive in decoding order, and contain exactly one coded picture.
[0109] A CVS is a coded video sequence for every coded layer video sequence (CLVS) in a video bitstream. In particular, a CVS and a CLVS are the same when a video bitstream contains a single layer. A CVS and a CLVS are different only when a video bitstream contains multiple layers.
[0110] Each picture 521 may be divided into tiles 523. A tile 523 is a partitioned portion of a picture created by horizontal and vertical boundaries. The tiles 523 may be rectangular and / or square. Specifically, a tile 523 includes four sides connected at right angles. The four sides include two pairs of parallel sides. Furthermore, the sides in a pair of parallel sides are of equal length. Therefore, the tiles 523 may be any rectangular shape, and a square is a special case of a rectangle where all four sides are of equal length. An image / picture may include one or more tiles 523.
[0111] A picture (e.g., picture 521) may be partitioned into rows and columns of tiles 523. A tile row is a set of tiles 523 arranged in a horizontally adjacent manner to create a continuous line from the left boundary to the right boundary of the picture (or vice versa). A tile column is a set of tiles 523 arranged in a vertically adjacent manner to create a continuous line from the top boundary to the bottom boundary of the picture (or vice versa).
[0112] A tile 523 may or may not allow prediction based on other tiles 523, depending on the example. Each tile 523 may have a unique tile index within a picture. The tile index is a procedurally selected numeric identifier that can be used to distinguish one tile 523 from another tile 523. For example, the tile index may increase numerically in raster scan order. The raster scan order is from left to right and from top to bottom.
[0113] It should be noted that in some examples, tiles 523 may also be assigned a tile identifier (ID). A tile ID is an assigned identifier that can be used to distinguish one tile 523 from another tile 523. In some examples, calculations may utilize the tile ID instead of the tile index. Further, in some examples, the tile ID can be assigned to have the same value as the tile index. The tile index and / or ID may be signaled to indicate the tile group that includes the tile 523. For example, the tile index and / or ID may be utilized to map picture data associated with the tile 523 to an appropriate location for display. A tile group is a related set of tiles 523 that can be extracted and coded separately, e.g., to support display of a region of interest and / or to support parallel processing. Tiles 523 within a tile group can be coded without reference to tiles 523 outside the tile group. Each tile 523 may be assigned to a corresponding tile group; thus, a picture can include multiple tile groups.
[0114] The tiles 523 are further divided into coding tree units (CTUs), which are further divided into coding blocks based on the coding tree, which can then be encoded / decoded according to a prediction mechanism.
[0115] 6A through 6E illustrate an example mechanism 600 for creating an extractor track 610 (also known as a merged bitstream) for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. The mechanism 600 may be utilized to support an example use case of the method 100. For example, the mechanism 600 may be utilized to generate the bitstream 500 for transmission from the codec system 200 and / or the encoder 300 to the codec system 200 and / or the decoder 400. As a particular example, the mechanism 600 may be utilized for use with VR, Omnidirectional Media Format (OMAF), 360-degree video, etc.
[0116] In VR, only a portion of a video is displayed to the user. For example, a VR video may be shot to include a sphere surrounding the user. A user may utilize a head-mounted display (HMD) to view the VR video. The user may point the HMD toward a region of interest. The region of interest is displayed to the user, and other video data is discarded. In this way, at any given moment, the user sees only a portion of the VR video selected by the user. This approach mimics the user's perception and thus allows the user to experience a virtual environment in a manner that mimics a real environment. One problem with this approach is that although the entire VR video may be transmitted to the user, only the current viewport of the video is actually used, and the rest is discarded. To increase signaling efficiency for streaming applications, the user's current viewport can be transmitted at a higher first resolution, and other viewports can be transmitted at a lower second resolution. In this way, viewports that are likely to be discarded occupy less bandwidth than viewports that are likely to be seen by the user. If the user selects a new viewport, lower resolution content may be displayed until the decoder may request that a different current viewport be transmitted at the higher first resolution. To support this functionality, a mechanism 600 may be utilized to create an extractor track 610 as depicted in Figure 6E. The extractor track 610 is a track of image data that encapsulates a picture (e.g., picture 521) at multiple resolutions for use as described above.
[0117] The mechanism 600 encodes the same video content in a first resolution 611 and a second resolution 612, as shown in FIGS. 6A and 6B, respectively. As a specific example, the first resolution 611 may be 5120×2560 luma samples, and the second resolution 612 may be 2560×1280 luma samples. Pictures of the video may be partitioned into tiles 601 in the first resolution 611 and tiles 603 in the second resolution 612, respectively. As used herein, tiles may be referred to as subpictures. In the illustrated example, tiles 601 and 603 are each partitioned into a 4×2 grid. Furthermore, an MCTS may be coded for the location of each tile 601 and 603. The pictures in the first resolution 611 and the second resolution 612 each result in an MCTS sequence that describes the video over time at the corresponding resolution. Each coded MCTS sequence is stored as a subpicture track or a tile track. The mechanism 600 can then create segments using the picture to support viewport-adaptive MCTS selection. For example, a range of viewing orientations is considered, which causes different selections of high-resolution and low-resolution MCTSs. In the illustrated example, four tiles 601 containing MCTSs at a first resolution 611 and four tiles 603 containing MCTSs at a second resolution 612 are obtained.
[0118] The mechanism 600 can then create an extractor track 610 for each possible viewport-adaptive MCTS selection. Figures 6C and 6D illustrate an example viewport-adaptive MCTS selection. Specifically, a set of selected tiles 605 and 607 are selected in a first resolution 611 and a second resolution 612, respectively. The selected tiles 605 and 607 are illustrated in shades of gray. In the illustrated example, the selected tile 605 is a tile 601 in the first resolution 611 that is to be displayed to the user, and the selected tile 607 is a tile 603 in the second resolution 612 that would likely be discarded but is retained to support display if the user selects a new viewport. The selected tiles 605 and 607 are then combined into a single picture that includes image data in both the first resolution 611 and the second resolution 612. Such a picture is combined to create the extractor track 610. For illustrative purposes, Figure 6E illustrates a single picture from a corresponding extractor track 610. As shown, the picture in the extractor track 610 includes selected tiles 605 and 607 from a first resolution 611 and a second resolution 612. As noted above, Figures 6C through 6E illustrate single-viewport adaptive MCTS selection. To enable user selection of any viewport, an extractor track 610 should be created for each possible combination of selected tiles 605 and 607.
[0119] In the depicted example, each selection of tiles 603 encapsulating content from a bitstream at a second resolution 612 includes two slices. A RegionWisePackingBox may be included in the extractor track 610 to create a mapping between the packed picture and a projected picture in equirectangular projection (ERP) format. In the presented example, the bitstream parsed from the extractor track 610 has a resolution of 3200x2560. Thus, a 4000-sample (4K)-capable decoder may decode content whose viewport is extracted from a coded bitstream having a 5000-sample 5K (5120x2560) resolution.
[0120] As shown in FIG. 6C , selected tiles 605 (represented in gray) at a first resolution 611 have the following tile identifiers: 10, 11, 14, and 15. As used herein, tile identifiers may also be referred to as subpicture identifiers. Selected tiles 607 at a second resolution 612 have the following identifiers: 1, 5, 4, and 8. Thus, extractor track 610 includes the following tile identifiers: 10, 11, 14, 15, 1, 5, 4, and 8. The tile identifiers are used to identify specific subpictures using subpicture indices, which may be referred to herein as subpicture ID mappings.
[0121] Figures 7A through 7E illustrate an example mechanism 700 for creating extractor tracks 710 for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in VR applications when the user changes viewport relative to the viewport chosen for Figures 6A through 6E. That is, Figures 7A through 7E illustrate how a new extractor track 710 is created when there is a change in viewing orientation in the CVS, which includes extractor track 610 and extractor track 710.
[0122] As depicted in Figures 7A-7B, the video pictures were partitioned into tiles 701 at a first resolution 711 and tiles 703 at a second resolution 712, respectively. However, there was a change in viewing orientation for mechanism 700 relative to mechanism 600. Thus, as depicted in Figures 7C-7D, selected tiles 705 (represented in gray) at first resolution 711 now have the following tile identifiers: 9, 10, 13, and 14, and selected tiles 707 at second resolution 712 now have the following identifiers: 3, 4, 7, and 8. Thus, extractor track 710 includes the following tile identifiers due to the change in viewing orientation: 3, 4, 7, 8, 9, 10, 13, and 14.
[0123] When the CVS being transmitted contains a certain set of sub-pictures, the associated sub-picture IDs are left in the SPS (all others are removed). When merging occurs (e.g., to form one of the extractor tracks 610 or 710), the sub-picture IDs are moved to the PPS. In either case, a flag is set in the bitstream to indicate where the sub-picture IDs are now located.
[0124] Typically, a new IRAP picture must be sent when a change in viewing orientation occurs. An IRAP picture is a coded picture for which all VCL NAL units have the same value of NAL unit type. An IRAP picture provides two important functions / benefits: First, the presence of an IRAP picture indicates that the decoding process can start from that picture. This function enables a random access function in which the decoding process starts at a position in the bitstream, not necessarily at the beginning of the bitstream, as long as an IRAP picture exists at that position. Second, the presence of an IRAP picture refreshes the decoding process so that coded pictures starting at the IRAP picture are coded without any reference to previous pictures, except for random access skipped leading (RASL) pictures. Therefore, the presence of an IRAP picture in the bitstream stops any errors that may occur during the decoding of coded pictures before the IRAP picture from propagating to the IRAP picture and those pictures that follow the IRAP picture in decoding order.
[0125] Although IRAP pictures provide important functionality, they come with a penalty to compression efficiency. The presence of IRAP pictures causes a sudden increase in bitrate. This penalty to compression efficiency is due to two reasons. First, because IRAP pictures are intra-predicted pictures, the picture itself requires relatively more bits to represent when compared to other pictures (e.g., preceding and following pictures) that are inter-predicted. Second, because the presence of an IRAP picture truncates temporal prediction (this is because the decoder refreshes the decoding process, and one of the actions of the decoding process for this is to remove previous reference pictures in the decoded picture buffer (DPB)), IRAP pictures have fewer reference pictures for their inter-predictive coding, making the coding of pictures that follow the IRAP picture in decoding order less efficient (i.e., requiring more bits to represent).
[0126] In one embodiment, IRAP pictures, along with random access decodable (RADL) pictures, are referred to as clean random access (CRA) pictures or instantaneous decoder refresh (IDR) pictures. In HEVC, IDR pictures, CRA pictures, and Broken Link Access (BLA) pictures are all considered IRAP pictures. For VVC, it was agreed during the 12th JVET meeting in October 2018 to have both IDR and CRA pictures as IRAP pictures. In one embodiment, Broken Link Access (BLA) and Gradual Decoder Refresh (GDR) pictures may also be considered IRAP pictures. The decoding process for a coded video sequence always starts at an IRAP picture.
[0127] In contrast to sending a new IRAP picture as described above, a better approach is to continue sending any tiles (also known as subpictures) that are shared between extractor track 610 and extractor track 710. That is, continue sending tiles with the following tile IDs: 4, 8, 10, and 14, because they are in both extractor track 610 and extractor track 710. In doing so, new IRAP pictures need to be sent only for those tiles in extractor track 710 that were not also in extractor track 610. That is, when the viewing orientation changes, new IRAP pictures need to be sent only for tiles with the following tile IDs: 1, 5, 9, and 13. However, changing subpicture IDs within a CVS can cause problems in signaling.
[0128] To solve at least the signaling problem, an indication (e.g., a flag) is signaled in the bitstream to indicate whether the tile ID (also known as subpicture ID) for each tile (also known as subpicture) can change within the CVS. Such a flag may be signaled within the SPS or the PPS. In addition to signaling whether the tile ID can change within the CVS, the flag may also serve other functions.
[0129] In one approach, subpicture IDs are signaled either in the SPS syntax (when the subpicture ID for each subpicture is indicated to not change in the CVS) or in the PPS syntax (when the subpicture ID for each subpicture is indicated to be variable in the CVS). In one embodiment, subpicture IDs are not signaled at all in both the SPS and PPS syntax.
[0130] In another approach, the subpicture ID is always signaled in the SPS syntax, and when a flag indicates that the subpicture ID for each subpicture may change in the CVS, the subpicture ID value signaled in the SPS may be overwritten by the subpicture ID signaled in the PPS syntax.
[0131] In yet another approach, subpicture IDs are signaled only in the PPS syntax and not in the SPS syntax, and an indication of whether the subpicture ID for each subpicture can change in the CVS is also signaled only in the PPS syntax.
[0132] In this approach, other subpicture information such as the position and size of each subpicture, the length of the subpicture ID in bits, along with the flags subpic_treated_as_pic_flag[i] and loop_filter_across_subpic_enabled_flag[i] as in the latest VVC draft specification may also be signaled in the PPS instead of in the SPS, but everything is the same for all PPSs referenced by coded pictures in the CVS.
[0133] When sub-picture positions and sizes are signaled within a PPS, instead of being signaled as in the latest VVC draft specification, they can be signaled based on the slices included in each sub-picture. For example, for each sub-picture, the slice index or ID of the slice located in the upper-left corner of the sub-picture and the slice index or ID or slice located in the lower-right corner of the sub-picture can be signaled for deriving the sub-picture position and size, and the signaling can be delta-based; in some specific cases, signaling of the slice index or ID or its delta can be avoided, and its value can be inferred, for example, similar to how the upper-left and lower-right brick indices are signaled for a rectangular slice in the latest VVC draft specification.
[0134] In this embodiment, the sub-picture ID is signaled as follows:
[0135] When the SPS syntax indicates that the number of sub-pictures in each picture in the CVS is greater than one, the following applies:
[0136] A flag (e.g., a specified subpicture_ids_signalled_in_sps_flag or sps_subpic_id_mapping_present_flag) is signaled in the SPS. In one embodiment, the flag has the following semantics: subpicture_ids_signalled_in_sps_flag equal to 1 specifies that subpicture IDs are signaled one per subpicture in the SPS, and that the subpicture ID value for a particular subpicture does not change in the CVS. subpicture_ids_signalled_in_sps_flag equal to 0 specifies that subpicture IDs are not signaled in the SPS, but are instead signaled in the PPS, and that the subpicture ID value for a particular subpicture may change in the CVS. When subpicture_ids_signalled_in_sps_flag is equal to 1, subpicture IDs are signaled per subpicture in the SPS.
[0137] Within the PPS syntax, a flag (e.g., named subpicture_ids_signalled_in_pps_flag or pps_subpic_id_mapping_present_flag) is signaled within the PPS. The flag has the following semantics: subpicture_ids_signalled_in_pps_flag equal to 1 specifies that subpicture IDs are signaled one per subpicture within the PPS, and that the subpicture ID value for each particular subpicture may vary within the CVS. subpicture_ids_signalled_in_pps_flag equal to 0 specifies that subpicture IDs are not signaled within the PPS, but are instead signaled within the SPS, and that the subpicture ID value for each particular subpicture does not vary within the CVS.
[0138] In one embodiment, the value of subpicture_ids_signalled_in_pps_flag shall be equal to 1-subpicture_ids_signalled_in_sps_flag. When subpicture_ids_signalled_in_pps_flag is equal to 1, subpicture IDs are signaled for each subpicture within the PPS.
[0139] Within the slice header syntax, the sub-picture IDs are signaled regardless of the number of sub-pictures specified by the referenced SPS.
[0140] Alternatively, when the number of subpictures in each picture is greater than one, the SPS syntax always includes one subpicture ID per subpicture, and also includes flags that specify whether the subpicture IDs can be overridden by subpicture IDs signaled in the PPS syntax and whether the subpicture ID value for a particular subpicture can change in the CVS. Subpicture ID overriding can be performed either always for all of the subpicture IDs, or only for a selected subset of all subpicture IDs.
[0141] FIG. 8 illustrates one embodiment of a decoding method 800 performed by a video decoder (e.g., video decoder 400). Method 800 may be performed after a decoded bitstream is received directly or indirectly from a video encoder (e.g., video encoder 300). Method 800 improves the decoding process in application scenarios involving both sub-bitstream extraction and sub-bitstream merging by ensuring efficient signaling of sub-picture IDs even when the sub-picture IDs change within a CVS. This reduces redundancy and increases coding efficiency. Thus, the coder / decoder (also known as a "codec") in video coding is improved relative to current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0142] In block 802, a video decoder receives a video bitstream (e.g., bitstream 500) that includes an SPS (e.g., SPS 510), a PPS (e.g., PPS 512), and multiple sub-pictures (e.g., tiles 605 and 607) associated with a sub-picture identifier (ID) mapping 575. As noted above, sub-picture ID mapping 575 is a mapping of sub-picture IDs to specific sub-pictures via sub-picture indices (e.g., sub-picture ID 8 corresponds to sub-picture index 8, which identifies a specific sub-picture from multiple sub-pictures). In one embodiment, SPS 510 includes an SPS flag 565. In one embodiment, PPS 512 includes a PPS flag 567.
[0143] In block 804, the video decoder determines whether the SPS flag 565 has a first value or a second value. The SPS flag 565 having the first value specifies that the sub-picture ID mapping 575 is included in the SPS 510, and the SPS flag 565 having the second value specifies that the sub-picture ID mapping 575 is signaled within the PPS 512.
[0144] When SPS flag 565 has the second value, in block 806, the video decoder determines whether PPS flag 567 has the first value or the second value. PPS flag 567 having the first value specifies that sub-picture ID mapping 575 is included in PPS 512, and PPS flag 567 having the second value specifies that sub-picture ID mapping 575 is not signaled within PPS 512.
[0145] In one embodiment, when PPS flag 567 has a second value, SPS flag 656 has a first value. In one embodiment, when PPS flag 567 has a first value, SPS flag 565 has a second value. In one embodiment, the first value is 1 and the second value is 0.
[0146] In one embodiment, SPS 510 includes second SPS flag 569. Second SPS flag 569 specifies whether sub-picture mapping 575 is explicitly signaled within SPS 510 or PPS 512. In one embodiment, bitstream 500 further comprises CVS change flag 571. CVS change flag 571 indicates whether sub-picture ID mapping 575 can change within CVS 590 of bitstream 500. In one embodiment, CVS change flag 571 is included in SPS 510, PPS 512, or another parameter set or header within bitstream 500.
[0147] In block 808, the video decoder obtains sub-picture ID mapping 575 from SPS 510 when SPS flag 565 has a first value, and from PPS 512 when SPS has a second value and / or PPS flag 567 has a first value. In one embodiment, the bitstream comprises a merged bitstream. In one embodiment, sub-picture ID mapping 575 has changed within CVS 590 of bitstream 500.
[0148] In block 810, the video decoder decodes the multiple sub-pictures using the sub-picture ID mapping 575. Once decoded, the multiple sub-pictures may be used to generate or create an image or video sequence for presentation to a user on a display or screen of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.).
[0149] 9 is an embodiment of a method 900 for encoding a video bitstream, performed by a video encoder (e.g., video encoder 300). Method 900 may be performed when pictures (e.g., from a video) are to be encoded into a video bitstream and then transmitted toward a video decoder (e.g., video decoder 400). Method 900 improves the encoding process in application scenarios involving both sub-bitstream extraction and sub-bitstream merging by ensuring efficient signaling of sub-picture IDs even when the sub-picture IDs change within a CVS. This reduces redundancy and increases coding efficiency. Thus, coders / decoders (also known as "codecs") in video coding are improved relative to current codecs. As a practical matter, the improved video coding process provides users with a better user experience when videos are transmitted, received, and / or viewed.
[0150] In block 902, the video encoder encodes a bitstream that includes an SPS (e.g., SPS 510), a PPS (e.g., PPS 512), and multiple sub-pictures (e.g., tiles 605 and 607) associated with a sub-picture identifier (ID) mapping. As noted above, the sub-picture ID mapping 575 is a mapping of sub-picture IDs to specific sub-pictures via sub-picture indices (e.g., sub-picture ID 8 corresponds to sub-picture index 8, which identifies a specific sub-picture from multiple sub-pictures). In one embodiment, the SPS 510 includes an SPS flag 565, and the PPS 512 includes a PPS flag 567.
[0151] In block 904, the video encoder sets the SPS flag 565 to a first value when the sub-picture ID mapping 575 is included in the SPS 510, and to a second value when the sub-picture ID mapping 575 is signaled within the PPS 512.
[0152] In block 906, the video encoder sets the PPS flag 567 to a first value when the sub-picture ID mapping 575 is included in the PPS 512, and to a second value when the sub-picture ID mapping 575 is not signaled within the PPS 512.
[0153] In one embodiment, when PPS flag 567 has a second value, SPS flag 565 has a first value. In one embodiment, when PPS flag 567 has a first value, SPS flag 565 has a second value. In one embodiment, the first value is 1 and the second value is 0.
[0154] In one embodiment, SPS 510 includes second SPS flag 569. Second SPS flag 569 specifies whether sub-picture ID mapping 575 is explicitly signaled within SPS 510 or PPS 512. In one embodiment, bitstream 500 further comprises CVS change flag 571. CVS change flag 571 indicates whether sub-picture ID mapping 575 can change within CVS 590 of bitstream 500. In one embodiment, CVS change flag 571 is included in SPS 510, PPS 512, or another parameter set or header within bitstream 500.
[0155] At block 908, the video encoder stores the bitstream for communication to the decoder. The bitstream may be stored in memory until the video bitstream is transmitted to the video decoder. Once received by the video decoder, the encoded video bitstream may be decoded (e.g., as described above) to generate or create images or video sequences for display to a user on a display or screen of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.).
[0156] In one embodiment, the subpicture ID is signaled as follows:
[0157] The SPS syntax does not include signaling of subpicture IDs.
[0158] Within the PPS syntax, a value indicating the number of sub-pictures is signaled and is required to be the same for all PPSs referenced by coded pictures within the CVS; when the indicated number of sub-pictures is greater than 1, the following applies:
[0159] A flag, for example named subpicture_id_unchanging_in_cvs_flag, is signaled in the PPS with the following semantics: subpicture_id_unchanging_in_cvs_flag equal to 1 specifies that the subpicture ID for each particular subpicture signaled in the PPS does not change in the CVS; subpicture_id_unchanging_in_cvs_flag equal to 0 specifies that the subpicture ID for each particular subpicture signaled in the PPS may change in the CVS.
[0160] A value indicating the number of sub-pictures is signaled within the PPS. The indicated number of sub-pictures shall be the same for all PPSs referenced by coded pictures within the CVS.
[0161] A subpicture ID is signaled for each subpicture within a PPS. When subpicture_id_unchanging_in_cvs_flag is equal to 1, the subpicture ID for a particular subpicture shall be the same for all PPSs referenced by coded pictures within the CVS.
[0162] In the slice header syntax, the sub-picture ID is signaled regardless of the number of sub-pictures specified by the referenced SPS.
[0163] It should also be understood that the steps of the example methods described herein are not necessarily required to be performed in the order described, and the order of steps in such methods should be understood to be merely exemplary. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined, in methods consistent with various embodiments of the present disclosure.
[0164] 10 is a schematic diagram of a video coding device 1000 (e.g., video encoder 300 or video decoder 400) according to one embodiment of the disclosure. The video coding device 1000 is suitable for implementing the disclosed embodiments as described herein. The video coding device 1000 comprises an ingress port 1010 and a receiver unit (Rx) 1020 for receiving data, a processor, logic unit, or central processing unit (CPU) 1030 for processing data, a transmitter unit (Tx) 1040 and an egress port 1050 for transmitting data, and a memory 1060 for storing data. The video coding device 1000 may also comprise optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to the ingress port 1010, the receiver unit 1020, the transmitter unit 1040, and the egress port 1050 for the egress or ingress of optical or electrical signals.
[0165] The processor 1030 is implemented by hardware and software. The processor 1030 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1030 is in communication with the ingress port 1010, the receiver unit 1020, the transmitter unit 1040, the egress port 1050, and the memory 1060. The processor 1030 includes a coding module 1070. The coding module 1070 implements the disclosed embodiments described above. For example, the coding module 1070 implements, processes, prepares, or provides various codec functions. Thus, the inclusion of the coding module 1070 provides significant improvements to the functionality of the video coding device 1000 and results in the transformation of the video coding device 1000 into different states. Alternatively, the coding module 1070 is implemented as instructions stored in the memory 1060 and executed by the processor 1030.
[0166] Video coding device 1000 may also include input and / or output (I / O) devices 1080 for communicating data to and from a user. I / O devices 1080 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. I / O devices 1080 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0167] The memory 1060 may comprise one or more disks, tape drives, and solid state drives, and may be used as overflow data storage devices, for storing programs when such programs are selected for execution, for storing instructions and data read during program execution, etc. The memory 1060 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).
[0168] 11 is a schematic diagram of one embodiment of a means for coding 1100. In one embodiment, the means for coding 1100 is implemented within a video coding device 1102 (e.g., video encoder 300 or video decoder 400). The video coding device 1102 includes a means for receiving 1101. The means for receiving 1101 is configured to receive a picture to be encoded or to receive a bitstream to be decoded. The video coding device 1102 includes a means for transmitting 1107 coupled to the means for receiving 1101. The means for transmitting 1107 is configured to transmit the bitstream to a decoder or to transmit a decoded image to a display means (e.g., one of I / O devices 1080).
[0169] The video coding device 1102 includes a storage means 1103. The storage means 1103 is coupled to at least one of the receiving means 1101 or the transmitting means 1107. The storage means 1103 is configured to store instructions. The video coding device 1102 also includes a processing means 1105. The processing means 1105 is coupled to the storage means 1103. The processing means 1105 is configured to execute the instructions stored in the storage means 1103 to perform the methods disclosed herein.
[0170] Although several embodiments have been provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples should be considered as illustrative and not restrictive, and the intention is not to be limited to the details provided herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0171] Additionally, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of modifications, substitutions, and alterations are ascertainable by those skilled in the art and could be made without departing from the spirit and scope of what is disclosed herein. [Explanation of symbols]
[0172] 100 How it works 200 Coding and Decoding (Codec) Systems 201 segmented video signal 211 General-Purpose Coder Control Component 213 Transform Scaling and Quantization Components 215 Intra-picture estimation components 217 Intra-picture prediction components 219 Motion Compensation Components 221 Motion Estimation Components 223 Decoded Picture Buffer Components 225 In-loop filter components 227 Filter Control Analysis Components 229 Scaling and Inverse Transformation Components 231 Header Formatting and Context-Adaptive Binary Arithmetic Coding (CABAC) Components 300 Video Encoder 301 Segmented Video Signal 313 Transformation and Quantization Components 317 Intra-picture prediction components 321 Motion Compensation Components 323 Decoded Picture Buffer Components 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Components 400 Video Decoder 417 Intra-picture prediction components 421 Motion Compensation Components 423 Decoded Picture Buffer Components 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Components 500 bitstream 510 Sequence Parameter Set (SPS) 512 Picture Parameter Set (PPS) 514 Tile Group Header 520 Image Data 521 Pictures 523, 601, 603, 605, 607 tiles 527 A sequence or series of pictures 565 SPS Flag 567 PPS Flag 569 Second SPS Flag 571 CVS change flags 575 Subpicture Identifier (ID) Mapping 590 Coded Video Sequence (CVS) 600 Mechanism 605, 607 tiles 610 Extractor Truck 611 First Resolution 612 Second Resolution 700 Mechanism 701, 703, 705, 707 tiles 710 Extractor Truck 711 First Resolution 712 Second Resolution 800 ways 900 methods 1000 Video Coding Devices 1010 Inlet port 1020 Receiver Unit (Rx) 1030 Processor, Logical Unit, or Central Processing Unit (CPU) 1040 Transmitter Unit (Tx) 1050 Exit Port 1060 memory 1070 Coding Module 1080 input and / or output (I / O) devices 1100 Coding Tools 1101 Receiving means 1102 Video Coding Device 1103 Memory means 1105 Processing means 1107 Transmission means
Claims
1. A method performed by a decoder, comprising: receiving, by the decoder, a bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with a sub-picture identifier (ID) mapping, wherein the SPS includes an SPS flag; determining by the decoder whether the SPS flag has a first value or a second value, wherein the SPS flag having the first value specifies that the sub-picture ID mapping is signaled in the SPS, and the SPS flag having the second value specifies that the sub-picture ID mapping is signaled in the PPS; obtaining, by the decoder, the sub-picture ID mapping from the SPS when the SPS flag has the first value, and from the PPS when the SPS flag has the second value; decoding, by the decoder, the plurality of sub-pictures using the sub-picture ID mapping; A method for providing the above.
2. the PPS includes a PPS flag, and the method comprises: determining by the decoder whether the PPS flag has the first value or the second value, wherein the PPS flag having the first value specifies that the sub-picture ID mapping is signaled within the PPS, and the PPS flag having the second value specifies that the sub-picture ID mapping is not signaled within the PPS; obtaining, by the decoder, the sub-picture ID mapping from the PPS when the PPS flag has the first value; The method of claim 1 further comprising:
3. 3. The method according to claim 1, wherein when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value.
4. 4. The method of claim 1, wherein the first value is 1 and the second value is 0.
5. 5. The method of claim 1, wherein the SPS includes a second SPS flag, the second SPS flag specifying whether the sub-picture ID mapping is explicitly signaled within the SPS or the PPS.
6. 6. The method of claim 1, wherein the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping is allowed to change within the coded video sequence (CVS) of the bitstream.
7. The method of claim 1 , wherein the bitstream comprises a merged bitstream and the sub-picture ID mapping has changed within the CVS of the bitstream.
8. A method performed by an encoder, comprising: encoding, by the decoder, a bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with a sub-picture identifier (ID) mapping, wherein the SPS includes an SPS flag; setting, by the decoder, the SPS flag to a first value when the sub-picture ID mapping is signaled in the SPS and to a second value when the sub-picture ID mapping is signaled in the PPS; storing the bitstream by the decoder for communication thereto; A method for providing the above.
9. 9. The method of claim 8, further comprising setting, by the decoder, a PPS flag in the PPS to the first value when the sub-picture ID mapping is signaled in the PPS, and to the second value when the sub-picture ID mapping is not signaled in the PPS.
10. 10. The method according to claim 8, wherein when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value.
11. 11. The method of claim 8, wherein the first value is 1 and the second value is 0.
12. 12. The method of claim 8, wherein the SPS includes a second SPS flag, the second SPS flag specifying whether the sub-picture mapping is explicitly signaled within the SPS or the PPS.
13. 13. The method of claim 8, wherein the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping may change within a coded video sequence (CVS) of the bitstream.
14. The method of claim 8 , wherein the bitstream comprises a merged bitstream and the sub-picture ID mapping has changed within the CVS of the bitstream.
15. 1. A decoding device, comprising: a receiver configured to receive a video bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with a sub-picture identifier (ID) mapping, the SPS including an SPS flag; a memory coupled to the receiver, the memory storing instructions; and a processor coupled to the memory; wherein the processor executes the instructions to cause the decoding device to: determining whether the SPS flag has a first value or a second value, wherein the SPS flag having the first value specifies that the sub-picture ID mapping is signaled in the SPS and the SPS flag having the second value specifies that the sub-picture ID mapping is signaled in the PPS; obtaining the sub-picture ID mapping from the SPS when the SPS flag has the first value and from the PPS when the SPS flag has the second value; decoding the plurality of sub-pictures using the sub-picture ID mapping; and a decoding device configured to cause
16. the PPS includes a PPS flag, and the processor: determining whether the PPS flag has the first value or the second value, wherein the PPS flag having the first value specifies that the sub-picture ID mapping is signaled within the PPS, and the PPS flag having the second value specifies that the sub-picture ID mapping is not signaled within the PPS; obtaining the sub-picture ID mapping from the PPS when the PPS flag has the first value; 16. The decoding device of claim 15, further configured to:
17. 17. A decoding device according to claim 15, wherein when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value, the first value being 1 and the second value being 0.
18. 18. A decoding device according to claim 15, wherein the SPS includes a second SPS flag, the second SPS flag specifying whether the sub-picture mapping is explicitly signaled within the SPS or the PPS.
19. 19. The decoding device of claim 15, wherein the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping may change within the coded video sequence (CVS) of the bitstream.
20. 20. A decoding device according to claim 15, wherein the bitstream comprises a merged bitstream and the sub-picture ID mapping has changed within the CVS of the bitstream.
21. 1. An encoding device, comprising: a memory containing instructions; a processor coupled to the memory that executes the instructions to cause the encoding device to: encoding a bitstream including a sequence parameter set (SPS), a picture parameter set (PPS), and a plurality of sub-pictures associated with a sub-picture identifier (ID) mapping, wherein the SPS includes an SPS flag; setting the SPS flag to a first value when the sub-picture ID mapping is signaled in the SPS and to a second value when the sub-picture ID mapping is signaled in the PPS; setting a PPS flag in the PPS to the first value when the sub-picture ID mapping is signaled in the PPS and to the second value when the sub-picture ID mapping is not signaled in the PPS; a processor configured to cause a transmitter coupled to the processor, the transmitter configured to transmit the bitstream to a video decoder; An encoding device comprising:
22. 22. The encoding device of claim 21, wherein when the PPS flag has the second value, the SPS flag has the first value, and when the PPS flag has the first value, the SPS flag has the second value.
23. 23. The encoding device according to claim 21, wherein the first value is 1 and the second value is 0.
24. 24. The encoding device of claim 21, wherein the bitstream further comprises a coded video sequence (CVS) change flag, the CVS change flag indicating whether the sub-picture ID mapping may change within the coded video sequence (CVS) of the bitstream.
25. 25. An encoding device as described in any of claims 21 to 24, wherein the PPS ID specifies a value of a second PPS ID for a PPS in use, the second PPS ID identifying the PPS for reference by the syntax element.
26. 26. The encoding device of claim 21, wherein the bitstream comprises a merged bitstream, and the sub-picture ID mapping has changed within the CVS of the bitstream.
27. 1. A coding device, comprising: a receiver configured to receive pictures to be encoded or to receive a bitstream to be decoded; a transmitter coupled to the receiver, the transmitter configured to transmit the bitstream to a decoder or to transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter, the memory configured to store instructions; a processor coupled to the memory, the processor configured to execute the instructions stored in the memory to perform the method of any of claims 1 to 7 and any of claims 8 to 14; A coding device comprising:
28. 28. The coding apparatus of claim 27, further comprising a display configured to display the decoded picture.
29. 1. A system comprising: An encoder; 28. A system comprising a decoder in communication with the encoder, the encoder or decoder comprising a decoding device, encoding device or coding apparatus according to any one of claims 15 to 27.
30. receiving means configured to receive a picture to be encoded or to receive a bitstream to be decoded; transmitting means coupled to said receiving means, said transmitting means configured to transmit said bitstream to decoding means or to transmit a decoded image to display means; a storage means coupled to at least one of the receiving means or the transmitting means, the storage means configured to store instructions; processing means coupled to said storage means, said processing means being configured to execute the instructions stored in said storage means to perform the method of any one of claims 1 to 7 and any one of claims 8 to 14; A means for coding comprising: