VIDEO ENCODER, VIDEO DECODER, AND CORRESPONDING METHODS - Patent application
By signaling sub-picture layout information in the SPS rather than the PPS, the inefficiencies in managing sub-pictures are addressed, enhancing coding efficiency and reducing resource usage and errors in video coding systems.
Patent Information
- Application Number
- JP2024201024
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-09
- Filing Date
- 2024-11-18
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2040-01-09
AI Technical Summary
Existing video coding systems face inefficiencies in managing sub-pictures, leading to redundant signaling of layout information, increased resource usage, and potential errors due to packet loss, especially in region-of-interest applications and sub-picture-based access schemes.
Incorporating sub-picture layout information into the Sequence Parameter Set (SPS) instead of the Picture Parameter Set (PPS) to signal sub-picture positions and sizes only once per sequence, reducing redundant signaling and improving coding efficiency, error resilience, and resource utilization.
This approach enhances coding efficiency, reduces resource usage, and minimizes errors by ensuring sub-picture layout is signaled only once per sequence, thereby optimizing network, memory, and processing resources in both encoders and decoders.
Smart Images

Figure 0007778895000015 
Figure 0007778895000016 
Figure 0007778895000017
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to video coding, and more particularly to sub-picture management in video coding. [Background technology]
[0002] The amount of video data required to render even a relatively short video can be considerable, which can pose challenges when the data is streamed or otherwise transmitted over communication networks with limited bandwidth capacity. Therefore, video data is typically compressed before being transmitted over modern telecommunications networks. Because memory resources can be limited, video size can also be an issue when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data needed to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and an ever-increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little or no sacrifice in image quality are desirable. Summary of the Invention [Means for solving the problem]
[0003] In an embodiment, the disclosure includes a method implemented in a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream comprising subpictures partitioned from a picture and a sequence parameter set (SPS) comprising subpicture sizes and subpicture positions of the subpictures; parsing, by a processor of the decoder, the SPS to obtain the subpicture sizes and positions; decoding, by the processor, the subpictures based on the subpicture sizes and positions to produce a video sequence; and transmitting, by the processor, the video sequence for display. Because tiles and subpictures are smaller than pictures, some systems include tiling and subpicture information in the picture parameter set (PPS). However, subpictures may be used to support region-of-interest (ROI) applications and subpicture-based access schemes, which do not change from picture to picture. The disclosed example includes layout information for subpictures in the SPS instead of the PPS. The subpicture layout information includes subpicture positions and subpicture sizes. The subpicture position is the offset between the top-left sample of the subpicture and the top-left sample of the picture. The subpicture size is the height and width of the subpicture, measured in luma samples. A video sequence may contain a single SPS (or one per video segment) or as many as one PPS per picture. Placing the layout information for subpictures within the SPS ensures that the layout is signaled only once per sequence / segment, rather than being signaled redundantly for each PPS. Signaling the subpicture layout within the SPS therefore increases coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Some systems also have subpicture information derived by the decoder.Signaling subpicture information reduces the probability of errors in case of lost packets and supports additional functionality with respect to extracting subpictures. Thus, signaling the subpicture layout within the SPS improves the functionality of the encoder and / or decoder.
[0004] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the subpicture is a temporal motion constrained subpicture, and that the subpicture size and the subpicture position indicate a layout of the temporal motion constrained subpicture.
[0005] Optionally, in any of the preceding aspects, another implementation of the aspect provides further comprising determining, by the processor, a size of the sub-picture relative to a size of the display based on the sub-picture size.
[0006] Optionally, in any of the preceding aspects, another implementation of the aspect provides further comprising determining, by the processor, a position of the sub-picture relative to the display based on the sub-picture position.
[0007] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the sub-picture position includes an offset distance between a top-left sample of the sub-picture and a top-left sample of the picture.
[0008] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the sub-picture size includes a sub-picture height in luma samples and a sub-picture width in luma samples.
[0009] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the SPS further comprises a sub-picture identifier (ID) for each sub-picture partitioned from the picture.
[0010] In an embodiment, the disclosure includes a method implemented in an encoder, the method comprising: encoding, by a processor of the encoder, a subpicture segmented from a picture into a bitstream; encoding, by the processor, a subpicture size and a subpicture position of the subpicture into an SPS in the bitstream; and storing the bitstream in the encoder's memory for communication to a decoder. Because tiles and subpictures are smaller than pictures, some systems include tiling and subpicture information in a picture parameter set (PPS). However, subpictures may be used to support region of interest (ROI) applications and subpicture-based access schemes, which do not change from picture to picture. A disclosed example includes layout information for subpictures in an SPS instead of a PPS. The subpicture layout information includes a subpicture position and a subpicture size. The subpicture position is the offset between the top-left sample of the subpicture and the top-left sample of the picture. The subpicture size is the height and width of the subpicture as measured in luma samples. A video sequence may contain a single SPS (or one per video segment) or as many as one PPS per picture. Placing layout information for subpictures within an SPS ensures that the layout is signaled only once for the sequence / segment, rather than being signaled redundantly for each PPS. Signaling subpicture layout within an SPS therefore increases coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Some systems also have subpicture information derived by the decoder. Signaling subpicture information reduces the probability of errors in the event of lost packets and supports additional functionality with respect to extracting subpictures. Thus, signaling subpicture layout within an SPS improves encoder and / or decoder functionality.
[0011] Optionally, in any of the preceding aspects, another implementation of the aspect provides further comprising encoding, by the processor, a flag in the SPS to indicate that the subpicture is a temporal motion constrained subpicture.
[0012] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the sub-picture size and the sub-picture position indicate a layout of the temporal motion constrained sub-pictures.
[0013] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the sub-picture position includes an offset distance between a top-left sample of the sub-picture and a top-left sample of the picture.
[0014] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the sub-picture size includes a sub-picture height in luma samples and a sub-picture width in luma samples.
[0015] Optionally, in any of the preceding aspects, another implementation of the aspect provides further comprising encoding, by the processor, a sub-picture ID for each of the sub-pictures partitioned from the picture into the SPS.
[0016] Optionally, in any of the preceding aspects, another implementation of the aspect provides further comprising encoding, by the processor, into the SPS the number of sub-pictures partitioned from the picture.
[0017] In an embodiment, the disclosure includes a video coding device comprising a processor, a memory, a receiver coupled to the processor, and a transmitter coupled to the processor, wherein the processor, memory, receiver, and transmitter are configured to perform the method of any of the preceding aspects.
[0018] In an embodiment, the disclosure includes a non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to perform the method of any of the preceding aspects.
[0019] In an embodiment, the disclosure includes a decoder comprising: receiving means for receiving a bitstream comprising sub-pictures partitioned from a picture and an SPS comprising sub-picture sizes and sub-picture positions of the sub-pictures; analyzing means for analyzing the SPS to obtain the sub-picture sizes and sub-picture positions; decoding means for decoding the sub-pictures based on the sub-picture sizes and sub-picture positions to produce a video sequence; and transferring means for transferring the video sequence for display.
[0020] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the decoder is further configured to perform the method of any of the preceding aspects.
[0021] In an embodiment, the disclosure includes an encoder comprising encoding means for encoding sub-pictures partitioned from a picture into a bitstream and encoding the sub-picture size and sub-picture position of the sub-picture into an SPS within the bitstream, and storage means for storing the bitstream for communication to a decoder.
[0022] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the encoder is further configured to perform the method of any of the preceding aspects.
[0023] For purposes of clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.
[0024] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.
[0025] For a more complete understanding of this disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts. [Brief explanation of the drawings]
[0026] [Figure 1] 1 is a flowchart of an example method for coding a video signal. [Figure 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding. [Figure 3] FIG. 1 is a schematic diagram illustrating an example video encoder. [Figure 4] FIG. 1 is a schematic diagram illustrating an example video decoder. [Figure 5] FIG. 1 is a schematic diagram illustrating an example bitstream and sub-bitstreams extracted from the bitstream. [Figure 6] FIG. 1 is a schematic diagram illustrating an example picture partitioned into sub-pictures. [Figure 7] FIG. 10 is a schematic diagram illustrating an example mechanism for associating slices with subpicture layouts. [Figure 8] FIG. 10 is a schematic diagram illustrating another example picture partitioned into sub-pictures. [Figure 9] 1 is a schematic diagram of an example video coding device. [Figure 10] 10 is a flowchart of an example method for encoding a subpicture layout in a picture bitstream to support subpicture extraction. [Figure 11]10 is a flowchart of an example method for decoding a subpicture bitstream based on a signaled subpicture layout. [Figure 12] FIG. 1 is a schematic diagram of an example system for signaling subpicture layouts via a bitstream. DETAILED DESCRIPTION OF THE INVENTION
[0027] While example implementations of one or more embodiments are provided below, it should be understood at the outset that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The disclosure should not be limited in any way to the example implementations, diagrams, and techniques illustrated below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims along with their full scope of equivalents.
[0028] Various acronyms are utilized herein, such as coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded video sequence (CVS), Joint Video Experts Team (JVET), motion constrained tile set (MCTS), maximum transmission unit (MTU), network abstraction layer (NAL), picture order count (POC), raw byte sequence payload (RBSP), sequence parameter set (SPS), versatile video coding (VVC), and working draft (WD).
[0029] Many video compression techniques can be utilized to reduce the size of video files with minimal loss of data. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a picture may be coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or image, and a reference picture may be referred to as a reference frame and / or image. Spatial or temporal prediction results in a predictive block that represents an image block. Residual data represents pixel differences between the original image block and the predictive block. Thus, inter-coded blocks are encoded according to a motion vector that points to a block of reference samples that form the predictive block, and residual data that indicates the difference between the coded block and the predictive block. Intra-coded blocks are encoded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain. These result in residual transform coefficients that may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to create a one-dimensional vector of transform coefficients.Entropy coding may be applied to achieve even greater compression. Such video compression techniques are discussed in more detail below.
[0030] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to corresponding video coding standards, including International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has begun developing a video coding standard called Versatile Video Coding (VVC). VVC is included in working drafts (WD), including JVET-L1001-v9.
[0031] To code a video image, the image is first partitioned, and the partitions are coded into a bitstream. Various picture partitioning schemes are available. For example, an image can be partitioned into normal slices, dependent slices, tiles, and / or according to wavefront parallelism (WPP). For simplicity, HEVC restricts the encoder so that only normal slices, dependent slices, tiles, WPP, and combinations thereof can be used when partitioning slices into groups of CTBs for video coding. Such partitioning can be applied to support maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay. The MTU represents the maximum amount of data that can be transmitted in a single packet. If a packet payload exceeds the MTU, the payload is split into two packets through a process called fragmentation.
[0032] A regular slice, also simply referred to as a slice, is a partitioned portion of an image that can be reconstructed independently of other regular slices within the same picture, despite some interdependence due to loop filtering operations. Each regular slice is encapsulated in its own Network Abstraction Layer (NAL) unit for transmission. Furthermore, intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries can be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, parallelization based on regular slices utilizes minimal inter-processor or inter-core communication. However, because each regular slice is independent, each slice is associated with a separate slice header. The use of regular slices can incur significant coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Furthermore, regular slices can be utilized to support matching MTU size requirements. Specifically, because regular slices can be encapsulated in separate NAL units and coded independently, each regular slice should be smaller than the MTU in the MTU scheme to avoid splitting the slice into multiple packets. Therefore, the goals of parallelization and MTU size matching may impose conflicting requirements on the slice layout in a picture.
[0033] Dependent slices are similar to regular slices but have shortened slice headers, allowing for partitioning of picture treeblock boundaries without breaking intra-picture prediction. Dependent slices therefore allow regular slices to be fragmented into multiple NAL units, which provides reduced end-to-end delay by allowing parts of a regular slice to be sent out before the encoding of the entire regular slice is complete.
[0034] A tile is a partitioned portion of an image created by horizontal and vertical boundaries that create tile columns and rows. Tiles may be coded in raster scan order (right-to-left and top-to-bottom). The scan order of CTBs is local within a tile. Thus, the CTB in the first tile is coded in raster scan order before proceeding to the CTB in the next tile. Similar to regular slices, tiles break the intra-picture prediction dependency as well as the entropy decoding dependency. However, tiles may not be contained in individual NAL units, and therefore tiles may not be used for MTU size matching. Each tile may be processed by a single processor / core, and inter-processor / inter-core communication utilized for intra-picture prediction between processing units decoding neighboring tiles may be limited to carrying a shared slice header (when adjacent tiles are in the same slice) and performing loop filtering-related sharing of reconstructed samples and metadata. When more than one tile is included in a slice, the entry point byte offset for each tile, other than the initial entry point offset within the slice, may be signaled in the slice header. For each slice and tile, at least one of the following conditions should be met: 1) all coded treeblocks in a slice belong to the same tile, and 2) all coded treeblocks in a tile belong to the same slice.
[0035] In WPP, an image is partitioned into single rows of CTBs. The entropy decoding and prediction mechanisms may use data from CTBs in other rows. Parallel processing is enabled through parallel decoding of CTB rows. For example, the current row can be decoded in parallel with the previous row. However, decoding of the current row is delayed by the decoding process of the previous row using two CTBs. This delay ensures that data related to the CTBs above and to the right of the current CTB in the current row is available before the current CTB is coded. This approach appears as a wavefront when represented graphically. This staggered start allows parallelization using up to the same number of processors / cores as the number of CTB rows the image contains. Because intra-picture prediction between neighboring treeblock rows within a picture is allowed, inter-processor / inter-core communication to enable intra-picture prediction can be significant. WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support MTU size matching. However, regular slicing can be used in conjunction with WPP with some coding overhead to implement MTU size matching on demand.
[0036] A tile may also include a motion constrained tile set. A motion constrained tile set (MCTS) is a tile set designed such that associated motion vectors are restricted to point to full sample locations within the MCTS and fractional sample locations that require only full sample locations within the MCTS for interpolation. Furthermore, the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is not permitted. In this way, each MCTS can be independently decoded without the presence of tiles not included in the MCTS. A temporal MCTS Supplemental Enhancement Information (SEI) message indicates the presence of an MCTS in a bitstream and can be used to signal the MCTS. The MCTS SEI message provides supplemental information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a compliant bitstream for the MCTS set. The information includes several extraction information sets, each defining a number of MCTS sets and including raw byte sequence payload (RBSP) bytes of replacement video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS) to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values in the extracted sub-bitstream.
[0037] A picture may also be partitioned into one or more subpictures. A subpicture is a rectangular set of tile groups / slices, starting with the tile group with tile_group_address equal to 0. Each subpicture may reference a different PPS and therefore have a different tile partition. Subpictures may be treated like pictures in the decoding process. A reference subpicture for decoding the current subpicture is generated by extracting an area in the decoded picture buffer from a reference picture that is co-located with the current subpicture. The extracted area is treated as the decoded subpicture. Inter-prediction may be performed between subpictures of the same size and the same location within a picture. A tile group, also known as a slice, is a sequence of related tiles within a picture or subpicture. To determine the location of a subpicture within a picture, several items may be derived. For example, each current subpicture may be placed at the next unoccupied position in CTU raster scan order within a picture large enough to contain the current subpicture within the picture boundary.
[0038] Furthermore, picture partitioning may be based on picture-level tiles and sequence-level tiles. Sequence-level tiles may include MCTS functionality and be implemented as subpictures. For example, a picture-level tile may be defined as a rectangular region of coding tree blocks within a particular tile column and a particular tile row within a picture. A sequence-level tile may also be defined as a set of rectangular regions of coding tree blocks contained in different frames, each of which further comprises one or more picture-level tiles, such that a set of rectangular regions of coding tree blocks is independently decodable from any other set of similar rectangular regions. A sequence-level tile group set (STGPS) is a group of such sequence-level tiles. STGPS may be signaled within a non-video coding layer (VCL) NAL unit with an associated identifier (ID) within the NAL unit header.
[0039] The preceding sub-picture-based partitioning scheme may be associated with certain challenges. For example, when sub-pictures are enabled, tiling within the sub-picture (partitioning of the sub-picture into tiles) can be used to support parallel processing. The tile partitioning of the sub-picture for the purpose of parallel processing can vary from picture to picture (e.g., for the purpose of balancing parallel processing load) and therefore may be managed at the picture level (e.g., in the PPS). However, sub-picture partitioning (partitioning of a picture into sub-pictures) may be utilized to support region of interest (ROI) and sub-picture-based picture access. In such cases, signaling of the sub-picture or MCTS in the PPS is inefficient.
[0040] In another example, when any subpicture in a picture is coded as a temporal motion constrained subpicture, all subpictures in the picture may be coded as temporal motion constrained subpictures. Such picture partitioning may be restrictive. For example, coding a subpicture as a temporal motion constrained subpicture may reduce coding efficiency in exchange for additional functionality. However, in region-of-interest-based applications, typically only one or a few of the subpictures use the functionality based on temporal motion constrained subpictures. Therefore, the remaining subpictures suffer from reduced coding efficiency without providing any practical benefit.
[0041] In another example, syntax elements for specifying the size of a subpicture may be specified in units of luma CTU size. Thus, both the width and height of the subpicture should be integer multiples of CtbSizeY. This mechanism for specifying the width and height of a subpicture may result in various problems. For example, subpicture partitioning is only applicable to pictures with picture widths and / or picture heights that are integer multiples of CtbSizeY. This renders subpicture partitioning unavailable for pictures with dimensions that are not integer multiples of CtbSizeY. If subpicture partitioning is applied to the picture width and / or height when the picture dimensions are not integer multiples of CtbSizeY, the derivation of the subpicture width and / or subpicture height in luma sample units for the rightmost and bottommost subpictures will be incorrect. Such an incorrect derivation causes incorrect results in some coding tools.
[0042] In another example, the position of a sub-picture within a picture may not be signaled. Instead, the position is derived using the following rule: The current sub-picture is placed at the next unoccupied position in CTU raster scan order within a picture large enough to contain the sub-picture within its picture boundaries. Deriving sub-picture positions in such a manner may cause errors in some cases. For example, if a sub-picture is lost in transmission, the positions of other sub-pictures will be derived incorrectly and decoded samples will be placed in the wrong positions. The same issue applies when sub-pictures arrive in the wrong order.
[0043] In another example, decoding a sub-picture may require extraction of a co-located sub-picture within a reference picture, which may impose additional complexity and a resulting burden in terms of processor and memory resource usage.
[0044] In another example, when a subpicture is designed as a temporal motion constrained subpicture, the loop filter across the subpicture boundary is disabled. This occurs regardless of whether the loop filter across the tile boundary is enabled. Such constraints may be too restrictive and may result in visual artifacts for video pictures that utilize multiple subpictures.
[0045] In another example, the relationship between the SPS, STGPS, PPS, and tile group header is as follows: STGPS references the SPS, PPS references the STGPS, and tile group header / slice header references the PPS. However, STGPS and PPS should be orthogonal, rather than PPS referencing STGPS. The preceding configuration may also not allow all tile groups of the same picture to reference the same PPS.
[0046] In another example, each STGPS may include IDs for the four sides of a subpicture. Such IDs can be used to identify subpictures that share the same boundary, thereby defining their relative spatial relationships. However, such information may in some cases not be sufficient to derive position and size information for a sequence-level tile group set. In other cases, signaling position and size information may be redundant.
[0047] In another example, the STGPS ID may be signaled in the NAL unit header of a VCL NAL unit using 8 bits. This may aid in subpicture extraction. Such signaling may unnecessarily increase the length of the NAL unit header. Another issue is that one tile group may be associated with multiple sequence-level tile group sets unless the sequence-level tile group sets are constrained to prevent overlaps.
[0048] To address one or more of the above-mentioned challenges, various mechanisms are disclosed herein. In a first example, layout information for subpictures is included in the SPS instead of the PPS. The subpicture layout information includes the subpicture position and the subpicture size. The subpicture position is the offset between the top-left sample of the subpicture and the top-left sample of the picture. The subpicture size is the height and width of the subpicture as measured in luma samples. As noted above, some systems include tiling information in the PPS because tiles can change from picture to picture. However, subpictures can be used to support ROI application and subpicture-based access. These features do not change from picture to picture. Furthermore, a video sequence may contain a single SPS (or one per video segment) or as many as one PPS per picture. Placing the layout information for subpictures in the SPS ensures that the layout is signaled only once per sequence / segment, rather than being redundantly signaled for each PPS. Signaling the subpicture layout within the SPS therefore increases coding efficiency and therefore reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Some systems also have subpicture information derived by the decoder. Signaling the subpicture information reduces the probability of errors in the event of lost packets and supports additional functionality with respect to extracting subpictures. Signaling the subpicture layout within the SPS therefore improves the functionality of the encoder and / or decoder.
[0049] In the second example, the subpicture width and subpicture height are constrained to be multiples of the CTU size. However, these constraints are removed when the subpicture is placed on the right border of a picture or the bottom border of a picture, respectively. As noted above, some video systems may restrict subpictures to include heights and widths that are multiples of the CTU size. This prevents subpictures from working correctly with many picture layouts. By allowing the bottom and right subpictures to include heights and widths that are not multiples of the CTU size, respectively, subpictures can be used with any picture without causing decoding errors. This results in increased encoder and decoder capabilities. Furthermore, the increased capabilities allow the encoder to code pictures more efficiently, which reduces the use of network, memory, and / or processing resources in the encoder and decoder.
[0050] In a third example, a subpicture is constrained to encompass a picture without gaps or overlaps. As noted above, some video coding systems allow subpictures to include gaps and overlaps. This creates the possibility that a tile group / slice may be associated with multiple subpictures. If this is allowed in an encoder, the decoder must be built to support such a coding scheme, even when the decoding scheme is rarely used. By not allowing subpicture gaps and overlaps, decoder complexity can be reduced because the decoder is not required to consider potential gaps and overlaps when determining the size and position of the subpicture. Furthermore, not allowing subpicture gaps and overlaps reduces the complexity of the rate-distortion optimization (RDO) process in the encoder because the encoder can omit considering gap and overlap cases when selecting an encoding for a video sequence. Therefore, avoiding gaps and overlaps may reduce the use of memory and / or processing resources in the encoder and decoder.
[0051] In a fourth example, a flag can be signaled within the SPS to indicate when a subpicture is a temporal motion constrained subpicture. As noted above, some systems may collectively set all subpictures to be temporal motion constrained subpictures or may not allow the use of temporal motion constrained subpictures at all. Such temporal motion constrained subpictures provide independent extraction capability at the expense of reduced coding efficiency. However, in region-of-interest-based applications, the region of interest should be coded for independent extraction, while regions outside the region of interest do not require such capability. Thus, the remaining subpictures suffer reduced coding efficiency without providing any practical benefit. Thus, the flag allows for a mix of temporal motion constrained subpictures providing independent extraction capability and non-motion constrained subpictures for increased coding efficiency when independent extraction is not desired. Thus, the flag allows for increased capability and / or increased coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0052] In a fifth example, a complete set of subpicture IDs is signaled within the SPS, and slice headers include subpicture IDs that indicate the subpictures that contain the corresponding slices. As noted above, some systems signal subpicture positions relative to other subpictures. This creates problems if subpictures are lost or extracted separately. By specifying each subpicture by ID, the subpictures can be positioned and sized without referencing other subpictures. This supports error correction, along with applications that only extract some subpictures and avoid transmitting others. A complete list of all subpicture IDs can be transmitted within the SPS, along with associated size information. Each slice header may include a subpicture ID that indicates the subpicture that contains the corresponding slice. In this way, subpictures and corresponding slices can be extracted and positioned without referencing other subpictures. Thus, subpicture IDs support increased functionality and / or increased coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0053] In a sixth example, a level is signaled for each subpicture. In some video coding systems, a level is signaled for a picture. The level indicates the hardware resources required to decode the picture. As noted above, different subpictures may have different capabilities in some cases and may therefore be treated differently during the coding process. Therefore, a picture-based level may not be useful for decoding some subpictures. Therefore, this disclosure includes a level for each subpicture. In this way, each subpicture can be coded independently of other subpictures without unnecessarily burdening the decoder by setting too high decoding requirements for subpictures coded according to less complex mechanisms. Signaled subpicture level information supports increased functionality and / or increased coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0054] 1 is a flowchart of an example operational method 100 for coding a video signal. Specifically, a video signal is encoded in an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size allows the compressed video file to be transmitted to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to an end user. The decoding process generally mirrors the encoding process to enable the decoder to coherently reconstruct the video signal.
[0055] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed sequentially, create the visual impression of movement. The frames include pixels represented in terms of light, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0056] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels by 64 pixels). CTUs contain both luma and chroma samples. A coding tree can be used to partition the CTUs into blocks and then recursively subdivide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of a frame can be subdivided until each block contains relatively homogeneous illumination values. Furthermore, the chroma component of a frame can be subdivided until each block contains relatively homogeneous color values. Thus, the partitioning scheme varies depending on the content of the video frame.
[0057] In step 105, various compression mechanisms are utilized to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be utilized. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Thus, a block depicting an object in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a constant position over multiple frames. Thus, the table may be described once, and adjacent frames may reference back to the reference frame. A pattern matching mechanism may be utilized to match objects across multiple frames. Furthermore, a moving object may be depicted across multiple frames, for example, due to object motion or camera motion. As a specific example, a video may depict a car moving across the screen over multiple frames. Motion vectors may be utilized to describe such motion. A motion vector is a two-dimensional vector that provides an offset from the coordinates of the object in a frame to the coordinates of the object in a reference frame. Therefore, inter-prediction may encode an image block in a current frame as a set of motion vectors indicating an offset from a corresponding block in a reference frame.
[0058] Intra prediction encodes blocks within a common frame. Intra prediction takes advantage of the fact that luma and chroma components tend to cluster within a frame. For example, a green spot in a tree tends to be located adjacent to similar green spots. Intra prediction utilizes multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar / the same as samples of neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. Planar mode effectively indicates a smooth transition of light / color across a row / column by utilizing a relatively constant gradient in changing values. DC mode is utilized for boundary smoothing, indicating that the block is similar / the same as the average value associated with samples of all neighboring blocks associated with the angular direction of the directional prediction mode. Thus, intra-predicted blocks can represent image blocks as values of various associated prediction modes instead of actual values. Furthermore, inter-predicted blocks can represent image blocks as values of motion vectors instead of actual values. In either case, the prediction block may in some cases not exactly represent the image block. Any differences are stored in a residual block. To further compress the file, a transform may be applied to the residual block.
[0059] Various filtering techniques may be applied in step 107. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction discussed above may result in the generation of blocky images in a decoder. Furthermore, block-based prediction schemes may encode blocks and then reconstruct the encoded blocks for later use as reference blocks. The in-loop filtering scheme iteratively applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference blocks, so that the artifacts are less likely to create additional artifacts in subsequent blocks that are encoded based on the reconstructed reference blocks.
[0060] Once the video signal has been segmented, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data discussed above along with any signaling data desired to support reconstruction of the proper video signal at the decoder. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder on demand. The bitstream may also be broadcast and / or multicast to multiple decoders. Generating the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously on multiple frames and blocks. The order depicted in FIG. 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to any particular order.
[0061] In step 111, a decoder receives the bitstream and begins the decoding process. Specifically, the decoder converts the bitstream into corresponding syntax and video data using an entropy decoding scheme. In step 111, the decoder uses syntax data from the bitstream to determine a partition for the frame. The partition should match the result of the block partitioning in step 103. Entropy encoding / decoding as used in step 111 is now described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible choices based on the spatial arrangement of values in the input image. Signaling the exact choice may utilize multiple bins. As used herein, a bin is a binary value (e.g., a bit value that can change depending on the situation) treated as a variable. Entropy coding allows the encoder to discard any options that are clearly not feasible for a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of allowable choices (e.g., one bin for two choices, two bins for three to four choices, etc.). The encoder then encodes a codeword for the selected choice. This scheme reduces the size of the codeword because the codeword is as large as desired to uniquely indicate a choice from a small subset of allowable choices, as opposed to uniquely indicating a choice from a potentially large set of all possible choices. The decoder then decodes the choices by determining the set of allowable choices in a similar manner as the encoder. By determining the set of allowable choices, the decoder can read the codeword and determine the choices made by the encoder.
[0062] In step 113, the decoder performs block decoding. Specifically, the decoder generates a residual block using an inverse transform. Then, the decoder reconstructs an image block according to the partition using the residual block and a corresponding prediction block. The prediction block may include both intra-prediction blocks and inter-prediction blocks, such as those generated in the encoder in step 105. The reconstructed image block is then arranged into a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding, as discussed above.
[0063] In step 115, filtering is performed on the frames of the reconstructed video signal at the encoder in a manner similar to step 107. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal can be output to a display in step 117 for viewing by an end user.
[0064] 2 is a schematic diagram of an example coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support the implementation of operational method 100. Codec system 200 is generalized to depict components utilized in both encoders and decoders. Codec system 200 receives and segments a video signal as discussed with respect to steps 101 and 103 in operational method 100, which results in a segmented video signal 201. When operating as an encoder as discussed with respect to steps 105, 107, and 109 in method 100, codec system 200 then compresses the segmented video signal 201 into a coded bitstream. When operating as a decoder, codec system 200 generates an output video signal from the bitstream as discussed with respect to steps 111, 113, 115, and 117 in operational method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, while dashed lines indicate the movement of control data that controls the operation of other components. The components of codec system 200 may all reside within an encoder. A decoder may include a subset of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are now described.
[0065] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree utilizes various partitioning modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. Blocks can be referred to as nodes in the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. In some cases, the partitioned blocks can be included in a coding unit (CU). For example, a CU can be a subpart of a coding unit (CTU) that includes a luma block, a red differential chroma (Cr) block, and a blue differential chroma (Cb) block, along with corresponding syntax instructions for the CU. Partitioning modes can include a binary tree (BT), a ternary tree (TT), and a quad tree (QT), which are used to partition a node into two, three, or four child nodes, respectively, of varying shapes depending on the partitioning mode used. The segmented video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0066] The generic coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the generic coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be made based on storage space / bandwidth availability and image resolution requirements. The generic coder control component 211 also manages buffer utilization, taking transmission rate into account, to mitigate buffer underrun and overrun issues. To manage these issues, the generic coder control component 211 manages segmentation, prediction, and filtering by other components. For example, the generic coder control component 211 may dynamically increase compression complexity to increase resolution and increase bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the generic coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality with bitrate concerns. The generic coder control component 211 produces control data that controls the operation of the other components. The control data is also forwarded to the header formatting and CABAC component 231 to be encoded into the bitstream for signaling parameters for decoding at the decoder.
[0067] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes to, for example, select an appropriate coding mode for each block of video data.
[0068] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are illustrated separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors, which estimate motion for a video block. A motion vector may indicate, for example, the displacement of a coded object relative to a predictive block. A predictive block is a block that is found to closely match a block to be coded in terms of pixel differences. A predictive block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference measures. HEVC utilizes several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into CTBs, which may then be divided into CBs for inclusion within a CU. A CU may be encoded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and may select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance both coding efficiency (e.g., size of the final encoding) and quality of the video reconstruction (e.g., amount of data loss due to compression).
[0069] In some examples, the codec system 200 may calculate values for sub-integer pixel locations of a reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values for quarter-pixel locations, eighth-pixel locations, or other fractional pixel locations of a reference picture. Accordingly, the motion estimation component 221 may perform motion searches for full-pixel and fractional pixel locations to output motion vectors with fractional-pixel accuracy. The motion estimation component 221 calculates motion vectors for PUs of video blocks in inter-coded slices by comparing the positions of the PUs with the positions of predictive blocks of the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting and CABAC component 231 for encoding and outputs motion to the motion compensation component 219.
[0070] The motion compensation performed by the motion compensation component 219 may involve fetching or generating a predictive block based on a motion vector determined by the motion estimation component 221. Again, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector for the PU of the current video block, the motion compensation component 219 may locate the predictive block to which the motion vector points. A residual video block is then formed by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predictive block and the residual block are forwarded to the transform scaling and quantization component 213.
[0071] The partitioned video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are illustrated separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block relative to blocks in the current frame as an alternative to the inter-prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames, as described above. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode to encode the current block from multiple tested intra-prediction modes. The selected intra-prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.
[0072] For example, the intra picture estimation component 215 may calculate rate-distortion values for various tested intra prediction modes using a rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to create the encoded block, along with the bit rate (e.g., number of bits) used to create the encoded block. The intra picture estimation component 215 may calculate a ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. In addition, the intra picture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0073] The intra-picture prediction component 217, when implemented in an encoder, may generate a residual block from the prediction block based on a selected intra-prediction mode determined by the intra-picture estimation component 215, or, when implemented in a decoder, may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both luma and chroma components.
[0074] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to create a video block comprising residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized with different granularity, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transform coefficients, which are forwarded to the header formatting and CABAC component 231 to be encoded into the bitstream.
[0075] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct a residual block in the pixel domain for later use, for example, as a reference block that may become a predictive block for another current block. The motion estimation component 221 and / or motion compensation component 219 may calculate a reference block by adding the residual block back to the corresponding predictive block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to mitigate artifacts created during scaling, quantization, and transform. Such artifacts may otherwise cause inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.
[0076] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 may be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are depicted separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes multiple parameters for adjusting how such filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such filter should be applied and sets the corresponding parameters. Such data is forwarded to the header formatting and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., on reconstructed pixel blocks) or in the frequency domain, depending on the example.
[0077] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as discussed above. When operating as a decoder, the decoded picture buffer component 223 stores and forwards the reconstructed and filtered blocks as part of the output video signal toward the display. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0078] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to a decoder. Specifically, the header formatting and CABAC component 231 generates various headers to encode control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, along with residual data in the form of quantized transform coefficient data, are all encoded within the bitstream. The final bitstream contains all information desired by a decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra-prediction mode, indications of partition information, and so on. Such data may be encoded by utilizing entropy coding. For example, the information may be encoded by utilizing context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream may be transmitted to another device (eg, a video decoder) or archived for later transmission or retrieval.
[0079] 3 is a block diagram illustrating an example video encoder 300. Video encoder 300 may be utilized to implement the encoding functionality of codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of operating method 100. Encoder 300 segments an input video signal, resulting in a segmented video signal 301 that is substantially similar to segmented video signal 201. Segmented video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.
[0080] Specifically, the partitioned video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on a reference block in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual block. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with associated control data) are forwarded to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231 .
[0081] The transformed and quantized residual block and / or the corresponding prediction block are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter within the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as discussed with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0082] 4 is a block diagram illustrating an example video decoder 400. Video decoder 400 may be utilized to implement the decoding functionality of codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of operating method 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.
[0083] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may utilize header information to provide context for interpreting additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion information, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into a residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0084] The reconstructed residual block and / or predictive block are forwarded to the intra-picture prediction component 417 for reconstruction into an image block based on an intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 utilizes a prediction mode to locate a reference block within a frame and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and corresponding inter-prediction data are forwarded to the decoded picture buffer component 423 via an in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are forwarded to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and applies a residual block to the result to reconstruct an image block. The resulting reconstructed block may also be forwarded to the decoded picture buffer component 423 via an in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks, which can be reconstructed into frames via partition information. Such frames may also be arranged in a sequence. The sequence is output to a display as a reconstructed output video signal.
[0085] 5 is a schematic diagram illustrating an example bitstream 500 and a sub-bitstream 501 extracted from bitstream 500. For example, bitstream 500 may be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400. As another example, bitstream 500 may be generated by an encoder in step 109 of method 100 for use by a decoder in step 111.
[0086] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPSs) 512, multiple slice headers 514, image data 520, and one or more SEI messages 515. The SPS 510 includes sequence data common to all pictures in the video sequence included in the bitstream 500. Such data may include picture size, bit depth, coding tool parameters, bit rate limits, etc. The PPS 512 includes parameters specific to one or more corresponding pictures. Thus, each picture in the video sequence may point to one PPS 512. The PPS 512 may indicate available coding tools, quantization parameters, offsets, picture-specific coding tool parameters (e.g., filter control), etc. for tiles in the corresponding picture. The slice header 514 includes parameters specific to one or more corresponding slices 524 in the picture. Thus, each slice 524 in the video sequence may reference a slice header 514. The slice header 514 may include slice type information, a picture order count (POC), a reference picture list, prediction weights, tile entry points, deblocking parameters, etc. In some examples, the slice 524 may be referred to as a tile group. In such cases, the slice header 514 may be referred to as a tile group header. The SEI message 515 is an optional message that includes metadata that is not required for block decoding, but can be utilized for related purposes such as indicating picture output timing, display settings, loss detection, loss concealment, etc.
[0087] The image data 520 includes video data to be encoded according to inter-prediction and / or intra-prediction along with corresponding transformed and quantized residual data. Such image data 520 is classified according to a partition used to partition the image before encoding. For example, a video sequence is partitioned into pictures 521. The pictures 521 may be further partitioned into sub-pictures 522, which are further partitioned into slices 524. The slices 524 may be further partitioned into tiles and / or CTUs. The CTUs are further partitioned into coding blocks based on a coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism. For example, a picture 521 can include one or more sub-pictures 522. A sub-picture 522 can include one or more slices 524. The picture 521 references the PPS 512, and the slices 524 reference the slice headers 514. The sub-pictures 522 can be partitioned coherently across the entire video sequence (also known as segments) and therefore can reference the SPS 510. Each slice 524 may include one or more tiles. Each slice 524, and therefore the picture 521 and subpicture 522, may also include multiple CTUs.
[0088] Each picture 521 may include the entire set of visual data associated with a video sequence for a corresponding moment in time. However, certain applications may desire to display only a portion of picture 521 in some cases. For example, a virtual reality (VR) system may display a user-selected region of picture 521, which creates the feeling of being present in the scene depicted in picture 521. The regions the user may wish to view are not known when bitstream 500 is encoded. Thus, picture 521 may include each possible region the user could potentially view as a sub-picture 522, which can be decoded and displayed separately based on user input. Other applications may display regions of interest separately. For example, a television with picture-in-picture may desire to display a particular region, and thus a sub-picture 522, from one video sequence over a picture 521 of an unrelated video sequence. In yet another example, a teleconferencing system may display an entire picture 521 of a user who is currently speaking and a sub-picture 522 of a user who is not currently speaking. Thus, sub-picture 522 may include a defined region of picture 521. Temporal motion constrained sub-picture 522 may be separately decodable from the rest of picture 521. Specifically, a temporal motion constrained sub-picture is encoded without reference to samples outside of the temporal motion constrained sub-picture and therefore contains sufficient information for complete decoding without reference to the rest of picture 521.
[0089] Each slice 524 may be a rectangle defined by a CTU in the upper-left corner and a CTU in the lower-right corner. In some examples, slices 524 include a series of tiles and / or CTUs in a raster scan order proceeding from left to right and top to bottom. In other examples, slices 524 are rectangular slices. Rectangular slices may not traverse the entire width of the picture in raster scan order. Instead, rectangular slices may include rectangular and / or square regions of picture 521 and / or subpicture 522 defined in terms of CTU and / or tile rows and CTU and / or tile columns. Slices 524 are the smallest units capable of being separately displayed by a decoder. Thus, slices 524 from picture 521 may be assigned to different subpictures 522 to separately depict desired regions of picture 521.
[0090] A decoder may display one or more subpictures 523 of picture 521. A subpicture 523 is a user-selected or pre-defined subgroup of subpictures 522. For example, picture 521 may be divided into nine subpictures 522, but the decoder may display only a single subpicture 523 from the group of subpictures 522. A subpicture 523 includes slices 525, which are selected or pre-defined subgroups of slices 524. To enable separate display of the subpictures 523, a sub-bitstream 501 may be extracted (529) from the bitstream 500. Extraction 529 may occur on the encoder side, such that the decoder receives only the sub-bitstream 501. In other cases, the entire bitstream 500 is transmitted to the decoder, and the decoder extracts (529) the sub-bitstream 501 for separate decoding. It should be noted that the sub-bitstream 501 may also be referred to generally as a bitstream in some cases. The sub-bitstream 501 includes an SPS 510 , a PPS 512 , selected sub-pictures 523 , along with slice headers 514 and SEI messages 515 related to the sub-pictures 523 and / or slices 525 .
[0091] This disclosure signals various data to support efficient coding of subpictures 522 for selection and display of subpictures 523 at a decoder. SPS 510 includes subpicture size 531, subpicture position 532, and subpicture ID 533 for the complete set of subpictures 522. Subpicture size 531 includes the subpicture height in luma samples and the subpicture width in luma samples for the corresponding subpicture 522. Subpicture position 532 includes the offset distance between the top-left sample of the corresponding subpicture 522 and the top-left sample of picture 521. Subpicture position 532 and subpicture size 531 define the layout of the corresponding subpicture 522. Subpicture ID 533 includes data that uniquely identifies the corresponding subpicture 522. Subpicture ID 533 may be the raster scan index of the subpicture 522 or another defined value. Thus, a decoder can read the SPS 510 and determine the size, position, and ID of each subpicture 522. In some video coding systems, subpictures 522 are partitioned from pictures 521, so data related to subpictures 522 may be included in the PPS 512. However, the partitioning used to create subpictures 522 may be used by applications such as ROI-based applications, VR applications, etc., that rely on subpicture 522 partitioning being consistent across a video sequence / segment. Therefore, the partitioning of subpictures 522 generally does not change from picture to picture. Placing the layout information for subpictures 522 within the SPS 510 ensures that the layout is signaled only once per sequence / segment, rather than redundantly signaled for each PPS 512 (which in some cases may be signaled for each picture 521). Also, signaling sub-picture 522 information instead of relying on the decoder to derive such information reduces the probability of error in the event of a lost packet and supports additional functionality with respect to extracting sub-picture 523.Thus, signaling the layout of the sub-pictures 522 within the SPS 510 improves the functionality of the encoder and / or decoder.
[0092] The SPS 510 also includes motion constrained subpicture flags 534 associated with the complete set of subpictures 522. The motion constrained subpicture flags 534 indicate whether each subpicture 522 is a temporal motion constrained subpicture. Thus, a decoder can read the motion constrained subpicture flags 534 to determine which of the subpictures 522 can be extracted and displayed separately without decoding the other subpictures 522. This allows selected subpictures 522 to be coded as temporal motion constrained subpictures, while allowing other subpictures 522 to be coded without such constraints for increased coding efficiency.
[0093] The sub-picture IDs 533 are also included in the slice headers 514. Each slice header 514 contains data related to a corresponding set of slices 524. Thus, a slice header 514 contains only the sub-picture IDs 533 corresponding to the slices 524 associated with the slice header 514. Therefore, a decoder can receive a slice 524, obtain the sub-picture IDs 533 from the slice header 514, and determine which sub-picture 522 contains the slice 524. The decoder can also use the sub-picture IDs 533 from the slice header 514 to correlate with the associated data in the SPS 510. Therefore, the decoder can determine how to arrange the sub-pictures 522 / 523 and slices 524 / 525 by reading the SPS 510 and the associated slice header 514. This allows the sub-pictures 523 and slices 525 to be decoded even if some sub-pictures 522 are lost in transmission or intentionally omitted to increase coding efficiency.
[0094] The SEI message 515 may also include a sub-picture level 535. The sub-picture level 535 indicates the hardware resources required to decode the corresponding sub-picture 522. In this way, each sub-picture 522 can be coded independently of the other sub-pictures 522. This ensures that each sub-picture 522 can be allocated the correct amount of hardware resources at the decoder. Without such a sub-picture level 535, each sub-picture 522 would be allocated sufficient resources to decode the most complex sub-picture 522. Thus, the sub-picture level 535 prevents the decoder from over-allocating hardware resources if the sub-pictures 522 are associated with varying hardware resource requirements.
[0095] 6 is a schematic diagram illustrating an example picture 600 partitioned into sub-pictures 622. For example, picture 600 may be encoded in and decoded from bitstream 500, e.g., by codec system 200, encoder 300, and / or decoder 400. Furthermore, picture 600 may be partitioned and / or included in sub-bitstreams 501 to support encoding and decoding according to method 100.
[0096] Picture 600 may be substantially similar to picture 521. Furthermore, picture 600 may be partitioned into sub-pictures 622, which are substantially similar to sub-pictures 522. The sub-pictures 622 each include a sub-picture size 631, which may be included in bitstream 500 as sub-picture size 531. The sub-picture size 631 includes a sub-picture width 631a and a sub-picture height 631b. The sub-picture width 631a is the width of the corresponding sub-picture 622 in units of luma samples. The sub-picture height 631b is the height of the corresponding sub-picture 622 in units of luma samples. The sub-pictures 622 each include a sub-picture ID 633, which may be included in bitstream 500 as sub-picture ID 633. The sub-picture ID 633 may be any value that uniquely identifies each sub-picture 622. In the depicted example, subpicture ID 633 is an index of subpicture 622. Subpictures 622 each include a position 632, which may be included in bitstream 500 as subpicture position 532. Position 632 is expressed as an offset between the top-left sample of the corresponding subpicture 622 and the top-left sample 642 of picture 600.
[0097] As also shown, some subpictures 622 may be temporal motion constrained subpictures 634, while others may not. In the shown example, the subpicture 622 with a subpicture ID 633 of 5 is a temporal motion constrained subpicture 634. This indicates that the subpicture 622 identified as 5 is coded without reference to any other subpictures 622, and therefore can be extracted and decoded separately without considering data from the other subpictures 622. An indication of which subpictures 622 are temporal motion constrained subpictures 634 may be signaled in the bitstream 500 in a motion constrained subpicture flag 534.
[0098] As shown, subpictures 622 can be constrained to encompass picture 600 without gaps or overlaps. A gap is a region of picture 600 that is not included in any subpicture 622. An overlap is a region of picture 600 that is included in more than one subpicture 622. In the example shown in FIG. 6, subpicture 622 is partitioned from picture 600 to prevent both gaps and overlaps. A gap causes a sample of picture 600 to be left outside of subpicture 622. An overlap causes an associated slice to be included in multiple subpictures 622. Thus, gaps and overlaps can cause samples to be affected by different treatment when subpictures 622 are coded differently. If this is allowed in an encoder, a decoder must support such a coding scheme, even if the decoding scheme is rarely used. By not allowing gaps and overlaps of sub-pictures 622, decoder complexity can be reduced because the decoder is not required to consider potential gaps and overlaps when determining sub-picture size 631 and position 632. Furthermore, not allowing gaps and overlaps of sub-pictures 622 reduces the complexity of the RDO process in the encoder because the encoder can omit considering cases of gaps and overlaps when selecting encodings for a video sequence. Thus, avoiding gaps and overlaps may reduce the use of memory and / or processing resources in the encoder and decoder.
[0099] 7 is a schematic diagram illustrating an example mechanism 700 for associating slices 724 with the layout of subpictures 722. For example, mechanism 700 may be applied to picture 600. Furthermore, mechanism 700 can be applied based on data in bitstream 500, for example, by codec system 200, encoder 300, and / or decoder 400. Furthermore, mechanism 700 can be utilized to support encoding and decoding according to method 100.
[0100] Mechanism 700 may be applied to slices 724 within subpicture 722, such as slices 524 / 525 and subpictures 522 / 523, respectively. In the illustrated example, subpicture 722 includes a first slice 724a, a second slice 724b, and a third slice 724c. The slice header for each of slices 724 includes a subpicture ID 733 for subpicture 722. The decoder may match the subpicture ID 733 from the slice header with a subpicture ID 733 in the SPS. The decoder may then determine the position 732 and size of subpicture 722 from the SPS based on the subpicture ID 733. Subpicture 722 may be positioned relative to an upper-left sample at the upper-left corner 742 of the picture using the position 732. The size may be used to set the height and width of subpicture 722 relative to the position 732. Slice 724 may then be included in subpicture 722. Thus, slice 724 can be placed in the correct subpicture 722 based on subpicture ID 733 without reference to other subpictures. This supports error correction, as other lost subpictures do not change the decoding of subpicture 722. This also supports applications that extract only subpicture 722 and avoid transmitting other subpictures. Thus, subpicture ID 733 supports increased capacity and / or increased coding efficiency, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0101] 8 is a schematic diagram illustrating another example picture 800 partitioned into sub-pictures 822. Picture 800 may be substantially similar to picture 600. Additionally, picture 800 may be encoded into and decoded from bitstream 500, for example, by codec system 200, encoder 300, and / or decoder 400. Furthermore, picture 800 may be partitioned and / or included in sub-bitstream 501 to support encoding and decoding according to method 100 and / or mechanism 700.
[0102] Picture 800 includes a subpicture 822, which may be substantially similar to subpictures 522, 523, 622, and / or 722. The subpicture 822 is divided into multiple CTUs 825. The CTUs 825 are the basic coding units in standardized video coding systems. The CTUs 825 are further subdivided by a coding tree into coding blocks, which are coded according to inter-prediction or intra-prediction. As shown, some subpictures 822a are constrained to include subpicture widths and subpicture heights that are multiples of the size of the CTUs 825. In the example shown, subpicture 822a has a height of six CTUs 825 and a width of five CTUs 825. This constraint is removed for subpicture 822b, located on the right border 801 of the picture, and for subpicture 822c, located on the bottom border 802 of the picture. In the illustrated example, subpicture 822b has a width between five and six CTUs 825. However, subpictures 822b that are not placed on the bottom boundary 802 of the picture are still constrained to maintain a subpicture height that is a multiple of the size of the CTUs 825. In the illustrated example, subpicture 822c has a height between six and seven CTUs 825. However, subpictures 822c that are not placed on the right boundary 801 of the picture are still constrained to maintain a subpicture width that is a multiple of the size of the CTUs 825.
[0103] As noted above, some video systems may restrict subpictures 822 to include heights and widths that are multiples of the size of CTU 825. This may prevent subpictures 822 from working correctly with many picture layouts, for example, with pictures 800 that include a total width or height that is not a multiple of the size of CTU 825. By allowing bottom subpicture 822c and right subpicture 822b to include heights and widths, respectively, that are not multiples of the size of CTU 825, subpictures 822 can be used with any picture 800 without causing decoding errors. This results in increased encoder and decoder capabilities. Furthermore, the increased capabilities allow the encoder to code pictures more efficiently, which reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0104] As described herein, this disclosure describes the design of sub-picture-based picture partitioning in video coding. A sub-picture is a rectangular area within a picture that can be independently decoded using a decoding process similar to that used for a picture. This disclosure relates to the signaling of sub-pictures in a coded video sequence and / or bitstream, as well as a process for sub-picture extraction. The description of the techniques is based on VVC by JVET of ITU-T and ISO / IEC. However, the techniques also apply to other video codec specifications. Below are example embodiments described herein. Such embodiments may be applied individually or in combination.
[0105] Information related to sub-pictures that may exist in a coded video sequence (CVS) may be signaled in a sequence-level parameter set, such as an SPS. Such signaling may include the following information: The number of sub-pictures present in each picture of the CVS may be signaled in the SPS. In the context of an SPS or a CVS, sub-pictures that are positioned at the same location for all access units (AUs) may be collectively referred to as a sub-picture sequence. A loop for further specifying information describing the attributes of each sub-picture may also be included in the SPS. This information may comprise sub-picture identification information, the sub-picture's location (e.g., the offset distance between the sub-picture's top-left corner luma sample and the picture's top-left corner luma sample), and the sub-picture's size. In addition, the SPS may signal whether each sub-picture is a motion-constrained sub-picture (including the functionality of MCTS). Profile, tier, and level information for each sub-picture may also be signaled or derivable at the decoder. Such information may be utilized to determine profile, tier, and level information for a bitstream created by extracting subpictures from the original bitstream. The profile and tier of each subpicture may be derived to be the same as the profile and tier of the entire bitstream. The level for each subpicture may be explicitly signaled. Such signaling may be present in a loop contained in the SPS. Sequence-level hypothetical reference decoder (HRD) parameters may be signaled in the video availability information (VUI) section of the SPS for each subpicture (or equivalently, each subpicture sequence).
[0106] When a picture is not partitioned into two or more subpictures, the attributes of the subpictures (e.g., position, size, etc.) may not be present / signaled in the bitstream, except for the subpicture ID. When subpictures of a picture in the CVS are extracted, each access unit in the new bitstream may not contain a subpicture. In this case, pictures in each AU in the new bitstream are not partitioned into multiple subpictures. Therefore, there is no need to signal subpicture attributes such as position and size in the SPS, because such information can be derived from the picture attributes. However, subpicture identification information may still be signaled, since the ID can be referenced by the VCL NAL unit / tile group contained in the extracted subpicture. This may allow the subpicture ID to remain the same when extracting subpictures.
[0107] The position (x offset and y offset) of a sub-picture within a picture may be signaled in units of luma samples. The position represents the distance between the top-left corner luma sample of the sub-picture and the top-left corner luma sample of the picture. Alternatively, the position of a sub-picture within a picture may be signaled in units of the minimum coding luma block size (MinCbSizeY). Alternatively, the unit of the sub-picture position offset may be explicitly indicated by a syntax element in the parameter set. The unit may be CtbSizeY, MinCbSizeY, luma samples, or other values.
[0108] The subpicture size (subpicture width and subpicture height) may be signaled in units of luma samples. Alternatively, the subpicture size may be signaled in units of minimum coding luma block size (MinCbSizeY). Alternatively, the unit of the subpicture size value may be explicitly indicated by a syntax element in the parameter set. The unit may be CtbSizeY, MinCbSizeY, luma samples, or other values. When the right boundary of the subpicture does not coincide with the right boundary of the picture, the subpicture width may be required to be an integer multiple of the luma CTU size (CtbSizeY). Similarly, when the bottom boundary of the subpicture does not coincide with the bottom boundary of the picture, the subpicture height may be required to be an integer multiple of the luma CTU size (CtbSizeY). If the subpicture width is not an integer multiple of the luma CTU size, the subpicture may be required to be positioned at the rightmost position within the picture. Similarly, if the height of the sub-picture is not an integer multiple of the luma CTU size, the sub-picture may be required to be positioned at the bottom-most position within the picture. In some cases, the width of the sub-picture may be signaled in units of the luma CTU size, but the width of the sub-picture is not an integer multiple of the luma CTU size. In this case, the actual width in luma samples may be derived based on the offset position of the sub-picture. The width of the sub-picture may be derived based on the luma CTU size, and the height of the picture may be derived based on the luma samples. Similarly, the height of the sub-picture may be signaled in units of the luma CTU size, but the height of the sub-picture is not an integer multiple of the luma CTU size. In such cases, the actual height in luma samples may be derived based on the offset position of the sub-picture. The height of the sub-picture may be derived based on the luma CTU size, and the height of the picture may be derived based on the luma samples.
[0109] For any subpicture, the subpicture ID may be different from the subpicture index. The subpicture index may be the index of the subpicture as signaled in the subpicture loop within the SPS. The subpicture ID may be the index of the subpicture in the subpicture raster scan order within the picture. When the value of the subpicture ID for each subpicture is the same as the subpicture index, the subpicture ID may be signaled or derived. When the subpicture ID for each subpicture is different from the subpicture index, the subpicture ID is explicitly signaled. The number of bits for signaling the subpicture ID may be signaled within the same parameter set (e.g., within the SPS) that contains the subpicture attributes. Some values for the subpicture ID may be reserved for certain purposes. For example, when a tile group header includes a sub-picture ID to specify which sub-pictures comprise the tile group, the value 0 may be reserved and unused for the sub-picture to ensure that the first few bits of the tile group header are not all 0 to prevent the accidental inclusion of emulation prevention code. In the optional case where a sub-picture of a picture does not encompass the entire area of the picture without gaps and without overlaps, a value (e.g., value 1) may be reserved for the tile group that is not part of any sub-picture. Alternatively, the sub-picture IDs of the remaining area are explicitly signaled. The number of bits for signaling the sub-picture ID may be constrained as follows: The range of values should be sufficient to uniquely identify all sub-pictures in the picture, including the reserved value of the sub-picture ID. For example, the minimum number of bits for the sub-picture ID may be the value of Ceil(Log2(number of sub-pictures in the picture + number of reserved sub-picture IDs).
[0110] It may be constrained that the union of sub-pictures must encompass the entire picture, without gaps and without overlaps. When this constraint is applied, for each sub-picture, there may be a flag to specify whether the sub-picture is a motion-constrained sub-picture, which indicates that the sub-picture can be extracted. Alternatively, the union of sub-pictures may not encompass the entire picture, but overlaps may not be allowed.
[0111] To aid the sub-picture extraction process without requiring the extractor to parse the rest of the NAL unit bits, a sub-picture ID may be present immediately after the NAL unit header. For VCL NAL units, the sub-picture ID may be present in the first bits of the tile group header. For non-VCL NAL units, the following may apply: For an SPS, the sub-picture ID does not need to be present immediately after the NAL unit header. For a PPS, if all tile groups of the same picture are constrained to reference the same PPS, the sub-picture ID does not need to be present immediately after its NAL unit header. If tile groups of the same picture are allowed to reference different PPSs, the sub-picture ID may be present in the first bits of the PPS (e.g., immediately after the NAL unit header). In this case, any tile groups of a picture may be allowed to share the same PPS. Alternatively, when tile groups of the same picture are allowed to reference different PPSs and different tile groups of the same picture are also allowed to share the same PPS, the sub-picture ID may not be present in the PPS syntax. Alternatively, when tile groups of the same picture are allowed to reference different PPSs and different tile groups of the same picture are allowed to share the same PPS, a list of sub-picture IDs may be present in the PPS syntax. This list indicates the sub-pictures to which the PPS applies. For other non-VCL NAL units, if a non-VCL unit (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) applies at or above the picture level, the sub-picture ID may not be present immediately after the NAL unit header. Otherwise, the sub-picture ID may be present immediately after the NAL unit header.
[0112] Using the above SPS signaling, tile partitioning within individual subpictures can be signaled within the PPS. Tile groups within the same picture may be allowed to reference different PPSs. In this case, tile grouping can only occur within each subpicture. The concept of tile grouping is the partitioning of a subpicture into tiles.
[0113] Alternatively, a parameter set is defined to describe tile partitioning within individual subpictures. Such a parameter set may be called a subpicture parameter set (SPPS). The SPPS references the SPS. Syntax elements referencing the SPS ID are present within the SPPS. The SPPS may include a subpicture ID. For the purpose of subpicture extraction, the syntax element referencing the subpicture ID is the first syntax element within the SPPS. The SPPS includes the tile structure (e.g., number of columns, number of rows, uniform tile spacing, etc.). The SPPS may include a flag to indicate whether the loop filter is enabled across the associated subpicture boundary. Alternatively, subpicture attributes for each subpicture may be signaled within the SPPS instead of within the SPS. Tile partitioning within individual subpictures may still be signaled within the PPS. Tile groups within the same picture are allowed to reference different PPSs. Once the SPPS is activated, it continues for a sequence of consecutive AUs in decode order. However, an SPPS may be deactivated / activated in an AU that is not the start of a CVS. At any moment during the decoding process of a single-layer bitstream with multiple subpictures in some AUs, multiple SPPSs may be active. An SPPS may be shared by different subpictures of an AU. Alternatively, an SPPS and a PPS may be merged into one parameter set. In such cases, it may not be required that all tile groups of the same picture refer to the same PPS. A constraint may be applied such that all tile groups in the same subpicture may refer to the same parameter set resulting from the merger between an SPPS and a PPS.
[0114] The number of bits used to signal the sub-picture ID may be signaled in the NAL unit header. When present in the NAL unit header, such information may aid the sub-picture extraction process in parsing the sub-picture ID value at the beginning of the NAL unit payload (e.g., the first few bits immediately after the NAL unit header). For such signaling, some of the reserved bits in the NAL unit header (e.g., 7 reserved bits) may be used to avoid increasing the length of the NAL unit header. The number of bits for such signaling may comprise the value of sub-picture-ID-bit-len. For example, 4 bits of the 7 reserved bits of the VVC NAL unit header may be used for this purpose.
[0115] When decoding a subpicture, the position of each coding tree block (e.g., xCtb and yCtb) may be adjusted to the actual luma sample position in the picture instead of the luma sample position in the subpicture. In this way, because the coding tree blocks are decoded with reference to the picture instead of the subpicture, extraction of the same positioned subpicture from each reference picture may be avoided. To adjust the position of the coding tree block, variables SubpictureXOffset and SubpictureYOffset may be derived based on the subpicture position (subpic_x_offset and subpic_y_offset). The values of the variables may be added to the values of the x and y coordinates of the luma sample position of each coding tree block in the subpicture, respectively.
[0116] The subpicture extraction process may be defined as follows: The input to the process is the target subpicture to be extracted. This may be in the form of a subpicture ID or a subpicture location. When the input is a subpicture location, the associated subpicture ID may be resolved by parsing the subpicture information in the SPS. For non-VCL NAL units, the following applies: Syntax elements in the SPS related to picture size and level may be updated with the subpicture size and level information. The following non-VCL NAL units are retained unchanged: PPS, Access Unit Delimiter (AUD), End of Sequence (EOS), End of Bitstream (EOB), and any other non-VCL NAL units applicable at or above the picture level. Remaining non-VCL NAL units with subpicture IDs not equal to the target subpicture ID may be removed. VCL NAL units with subpicture IDs not equal to the target subpicture ID may also be removed.
[0117] A sequence-level subpicture nesting SEI message may be used to nest AU-level or subpicture-level SEI messages for a set of subpictures. This may include buffering period, picture timing, and non-HRD SEI messages. The syntax and semantics of this subpicture nesting SEI message may be as follows: For system operation, such as in an Omni-Directional Media Format (OMAF) environment, a set of subpicture sequences that encompass a viewport may be requested and decoded by an OMAF player. Thus, a sequence-level SEI message is used to convey information about a set of subpicture sequences that collectively encompass a rectangular picture area. This information may be used by the system to indicate the required decoding capability along with the bitrate of the set of subpicture sequences. This information indicates the level of a bitstream that includes only the set of subpicture sequences. This information also indicates the bitrate of a bitstream that includes only the set of subpicture sequences. Optionally, a sub-bitstream extraction process may be specified for a set of subpicture sequences. The benefit of doing this is that bitstreams containing only a set of sub-picture sequences may also become compliant. The drawback is that when considering the possibility of different viewport sizes, there may be many such sets, in addition to the already large possible number of individual sub-picture sequences.
[0118] In one example embodiment, one or more of the disclosed examples may be implemented as follows: A subpicture may be defined as a rectangular region of one or more tile groups within a picture. An allowed bisection process may be defined as follows: Inputs to this process are a bisection mode btSplit, a coding block width cbWidth, a coding block height cbHeight, a position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture, a multi-type tree depth mttDepth, a maximum multi-type tree depth with offset maxMttDepth, a maximum bisection tree size maxBtSize, and a partition index partIdx. The output of this process is the variable allowBtSplit.
[0119] [Table 1]
[0120] The variables parallelTtSplit and cbSize are derived as specified above. The variable allowBtSpit is derived as follows: if one or more of the following conditions are true: cbSize is less than or equal to MinBtSizeY, cbWidth is greater than maxBtSize, cbHeight is greater than maxBtSize, and mttDepth is greater than or equal to maxMttDepth, then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions are true: btSplit is equal to SPLIT_BT_VER, and y0 + cbHeight is greater than SupPicBottomBorderInPic, then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions are true: btSplit equals SPLIT_BT_HOR, x0+cbWidth is greater than SupPicRightBorderInPic, and y0+cbHeight is less than or equal to SubPicBottomBorderInPic, then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions are true: mttDepth is greater than 0, partIdx is equal to 1, and MttSplitMode[x0][y0][mttDepth-1] is equal to parallelTtSplit, then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions are true: btSplit equals SPLIT_BT_VER, cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY, then allowBtSplit is set equal to FALSE. Otherwise, if all of the following conditions are true: btSplit equals SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY, then allowBtSplit is set equal to FALSE. Otherwise, allowBtSplit is set equal to TRUE.
[0121] The allowed three-way split process may be defined as follows: The inputs to this process are the three-way split mode ttSplit, the coding block width cbWidth, the coding block height cbHeight, the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture, the multi-type tree depth mttDepth, the maximum multi-type tree depth with offset maxMttDepth, and the maximum binary tree size maxTtSize. The output of this process is the variable allowTtSplit.
[0122] [Table 2]
[0123] The variable cbSize is derived as specified above. The variable allowTtSplit is derived as follows: If one or more of the following conditions are true: cbSize is less than or equal to 2*MinTtSizeY, cbWidth is greater than Min(MaxTbSizeY,maxTtSize), cbHeight is greater than Min(MaxTbSizeY,maxTtSize), mttDepth is greater than or equal to maxMttDepth, x0+cbWidth is greater than SupPicRightBoderInPic, and y0+cbHeight is greater than SubPicBottomBorderInPic, then allowTtSplit is set equal to FALSE. Otherwise, allowTtSplit is set equal to TRUE.
[0124] The syntax and semantics of the sequence parameter set RBSP are as follows:
[0125] [Table 3]
[0126] pic_width_in_luma_samples specifies the width of each decoded picture in units of luma samples. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. pic_height_in_luma_samples specifies the height of each decoded picture in units of luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. num_subpicture_minus1 plus 1 specifies the number of subpictures partitioned in a coded picture belonging to the coded video sequence. subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax elements subpic_id[i] in an SPS, spps_subpic_id in an SPPS that references an SPS, and tile_group_subpic_id in a tile group header that references an SPS. The value of subpic_id_len_minus1 shall be in the range of Ceil(Log2(num_subpic_minus1+2)) to 8, inclusive. subpic_id[i] specifies the subpicture ID of the ith subpicture of the picture referencing the SPS. The length of subpic_id[i] is subpic_id_len_minus1+1 bits. The value of subpic_id[i] shall be greater than 0. subpic_level_idc[i] indicates the level at which the CVS resulting from the extraction of the ith subpicture complies with the specified resource requirements. The bitstream shall not contain values of subpic_level_idc[i] other than those specified. Other values of subpic_level_idc[i] are reserved. When not present, the value of subpic_level_idc[i] is inferred to be equal to the value of general_level_idc.
[0127] subpic_x_offset[i] specifies the horizontal offset of the top-left corner of the ith subpicture relative to the top-left corner of the picture. When not present, the value of subpic_x_offset[i] is inferred to be equal to 0. The subpicture x offset value is derived as follows: SubpictureXOffset[i] = subpic_x_offset[i]. subpic_y_offset[i] specifies the vertical offset of the top-left corner of the ith subpicture relative to the top-left corner of the picture. When not present, the value of subpic_y_offset[i] is inferred to be equal to 0. The subpicture y offset value is derived as follows: SubpictureYOffset[i] = subpic_y_offset[i]. subpic_width_in_luma_samples[i] specifies the width of the ith decoded subpicture for which this SPS is the active SPS. When the sum of SubpictureXOffset[i] and subpic_width_in_luma_samples[i] is less than pic_width_in_luma_samples, the value of subpic_width_in_luma_samples[i] shall be an integer multiple of CtbSizeY. When not present, the value of subpic_width_in_luma_samples[i] is inferred to be equal to the value of pic_width_in_luma_samples. subpic_height_in_luma_samples[i] specifies the height of the ith decoded subpicture for which this SPS is the active SPS. When the sum of SubpictureYOffset[i] and subpic_height_in_luma_samples[i] is less than pic_height_in_luma_samples, the value of subpic_height_in_luma_samples[i] shall be an integer multiple of CtbSizeY. When not present, the value of subpic_height_in_luma_samples[i] is inferred to be equal to the value of pic_height_in_luma_samples.
[0128] It is a bitstream compliance requirement that the union of subpictures encompass the entire area of the picture with no overlaps or gaps. subpic_motion_constrained_flag[i] equal to 1 specifies that the i-th subpicture is a temporal motion constrained subpicture. subpic_motion_constrained_flag[i] equal to 0 specifies that the i-th subpicture may or may not be a temporal motion constrained subpicture. When not present, the value of subpic_motion_constrained_flag is inferred to be equal to 0.
[0129] The variables SubpicWidthInCtbsY, SubpicHeightInCtbsY, SubpicSizeInCtbsY, SubpicWidthInMinCbsY, SubpicHeightInMinCbsY, SubpicSizeInMinCbsY, SubpicSizeInSamplesY, SubpicWidthInSamplesC, and SubpicHeightInSamplesC are derived as follows: SubpicWidthInLumaSamples[i]=subpic_width_in_luma_samples[i] SubpicHeightInLumaSamples[i]=subpic_height_in_luma_samples[i] SubPicRightBorderInPic[i]=SubpictureXOffset[i]+PicWidthInLumaSamples[i] SubPicBottomBorderInPic[i]=SubpictureYOffset[i]+PicHeightInLumaSamples[i] SubpicWidthInCtbsY[i]=Ceil(SubpicWidthInLumaSamples[i]÷CtbSizeY) SubpicHeightInCtbsY[i]=Ceil(SubpicHeightInLumaSamples[i]÷CtbSizeY) SubpicSizeInCtbsY[i]=SubpicWidthInCtbsY[i]*SubpicHeightInCtbsY[i] SubpicWidthInMinCbsY[i]=SubpicWidthInLumaSamples[i] / MinCbSizeY SubpicHeightInMinCbsY[i]=SubpicHeightInLumaSamples[i] / MinCbSizeY SubpicSizeInMinCbsY[i]=SubpicWidthInMinCbsY[i]*SubpicHeightInMinCbsY[i] SubpicSizeInSamplesY[i]=SubpicWidthInLumaSamples[i]*SubpicHeightInLumaSamples[i] SubpicWidthInSamplesC[i]=SubpicWidthInLumaSamples[i] / SubWidthC SubpicHeightInSamplesC[i]=SubpicHeightInLumaSamples[i] / SubHeightC
[0130] The syntax and semantics of the subpicture parameter set RBSP are as follows:
[0131] [Table 4]
[0132] spps_subpic_id identifies the subpicture to which the SPPS belongs. The length of spps_subpic_id is subpic_id_len_minus1 + 1 bits. spps_subpic_parameter_set_id identifies the SPPS for reference by other syntax elements. The value of spps_subpic_parameter_set_id shall be in the range of 0 to 63, inclusive. spps_seq_parameter_set_id specifies the value of spps_seq_parameter_set_id for the active SPS. The value of spps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive. single_tile_in_subpic_flag equal to 1 specifies that there is only one tile in each subpicture that references the SPPS. single_tile_in_subpic_flag equal to 0 specifies that there is more than one tile in each subpicture that references the SPPS. num_tile_columns_minus1 plus one specifies the number of tile columns that partition the subpicture. num_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY[spps_subpic_id]-1, inclusive. When not present, the value of num_tile_columns_minus1 is inferred to be equal to 0. num_tile_rows_minus1 plus one specifies the number of tile rows that partition the subpicture. num_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY[spps_subpic_id]-1, inclusive. When not present, the value of num_tile_rows_minus1 is inferred to be equal to 0. The variable NumTilesInPic is set equal to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1).
[0133] When single_tile_in_subpic_flag is equal to 0, NumTilesInPic shall be greater than 0. uniform_tile_spacing_flag equal to 1 specifies that the tile column borders, and similarly the tile row borders, are uniformly distributed across the subpicture. uniform_tile_spacing_flag equal to 0 specifies that the tile column borders, and similarly the tile row borders, are not uniformly distributed across the subpicture, but are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When not present, the value of uniform_tile_spacing_flag is inferred to be equal to 1. tile_column_width_minus1[i] plus 1 specifies the width of the ith tile column in units of CTBs. tile_row_height_minus1[i] plus 1 specifies the height of the ith tile row in units of CTBs.
[0134] The following variables: a list ColWidth[i] for i ranging from 0 to num_tile_columns_minus1, inclusive, specifying the width of the ith tile column in CTBs; a list RowHeight[j] for j ranging from 0 to num_tile_rows_minus1, inclusive, specifying the height of the jth tile row in CTBs; a list ColBd[i] for i ranging from 0 to num_tile_columns_minus1+1, inclusive, specifying the location of the ith tile column boundary in CTBs; a list RowBd[j] for j ranging from 0 to num_tile_rows_minus1+1, inclusive, specifying the location of the jth tile row boundary in CTBs; a list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY-1, inclusive, specifying the conversion from CTB addresses in the CTB raster scan of the picture to CTB addresses in the tile scan. a list CtbAddrTsToRs[ctbAddrTs] of ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 inclusive that specifies the conversion from CTB addresses in tile scan to CTB addresses in the CTB raster scan of the picture; a list TileId[ctbAddrTs] of ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 inclusive that specifies the conversion from CTB addresses to tile IDs in tile scan; a list NumCtusInTile[tileIdx] of tileIdx ranging from 0 to PicSizeInCtbsY-1 inclusive that specifies the conversion from tile index to the number of CTUs in the tile; a list FirstCtbAddrTs[tileIdx] of tileIdx ranging from 0 to NumTilesInPic-1 inclusive that specifies the conversion from tile ID to CTB address in tile scan of the first CTB in the tile;The list ColumnWidthInLumaSamples[i] for i, ranging from 0 to num_tile_columns_minus1 inclusive, and the list RowHeightInLumaSamples[j] for j, ranging from 0 to num_tile_rows_minus1 inclusive, which specifies the height of the jth tile row in units of luma samples, are derived by invoking the CTB raster and tile scan conversion process. The values of ColumnWidthInLumaSamples[i] for i, ranging from 0 to num_tile_columns_minus1 inclusive, and RowHeightInLumaSamples[j] for j, ranging from 0 to num_tile_rows_minus1 inclusive, shall all be greater than 0.
[0135] loop_filter_across_tiles_enabled_flag equal to 1 specifies that in-loop filtering operations may be performed across tile boundaries in subpictures that reference an SPPS. loop_filter_across_tiles_enabled_flag equal to 0 specifies that in-loop filtering operations are not performed across tile boundaries in subpictures that reference an SPPS. In-loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of loop_filter_across_tiles_enabled_flag is inferred to be equal to 1. loop_filter_across_subpic_enabled_flag equal to 1 specifies that in-loop filtering operations may be performed across subpicture boundaries in subpictures that reference an SPPS. loop_filter_across_subpic_enabled_flag equal to 0 specifies that in-loop filtering operations are not performed across subpicture boundaries in subpictures that reference an SPPS. In-loop filtering operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of loop_filter_across_subpic_enabled_flag is inferred to be equal to the value of loop_filter_across_tiles_enabed_flag.
[0136] The syntax and semantics of a general tile group header are as follows:
[0137] [Table 5]
[0138] The values of the tile group header syntax elements tile_group_pic_parameter_set_id and tile_group_pic_order_cnt_lsb shall be the same in all tile group headers of a coded picture. The value of the tile group header syntax element tile_group_subpic_id shall be the same in all tile group headers of a coded subpicture. tile_group_subpic_id identifies the subpicture to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1+1 bits. tile_group_subpic_parameter_set_id specifies the value of spps_subpic_parameter_set_id for the SPPS in use. The value of tile_group_spps_parameter_set_id shall be in the range 0 to 63, inclusive.
[0139] The following variables are derived and override the respective variables derived from the active SPS: PicWidthInLumaSamples=SubpicWidthInLumaSamples[tile_group_subpic_id] PicHeightInLumaSamples=PicHeightInLumaSamples[tile_group_subpic_id] SubPicRightBorderInPic=SubPicRightBorderInPic[tile_group_subpic_id] SubPicBottomBorderInPic=SubPicBottomBorderInPic[tile_group_subpic_id] PicWidthInCtbsY=SubPicWidthInCtbsY[tile_group_subpic_id] PicHeightInCtbsY=SubPicHeightInCtbsY[tile_group_subpic_id] PicSizeInCtbsY=SubPicSizeInCtbsY[tile_group_subpic_id] PicWidthInMinCbsY=SubPicWidthInMinCbsY[tile_group_subpic_id] PicHeightInMinCbsY=SubPicHeightInMinCbsY[tile_group_subpic_id] PicSizeInMinCbsY=SubPicSizeInMinCbsY[tile_group_subpic_id] PicSizeInSamplesY=SubPicSizeInSamplesY[tile_group_subpic_id] PicWidthInSamplesC=SubPicWidthInSamplesC[tile_group_subpic_id] PicHeightInSamplesC=SubPicHeightInSamplesC[tile_group_subpic_id]
[0140] The coding tree unit syntax is as follows:
[0141] [Table 6]
[0142] [Table 7]
[0143] The syntax and semantics of the coding quadtree are as follows:
[0144] [Table 8]
[0145] qt_split_cu_flag[x0][y0] specifies whether the coding unit is split into coding units with half the horizontal and vertical sizes. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When qt_split_cu_flag[x0][y0] does not exist, the following applies. If one or more of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. If treeType is equal to DUAL_TREE_CHROMA or, otherwise, greater than MaxBtSizeY, then x0+(1<<log2CbSize) is greater than SubPicRightBorderInPic and (1<<log2CbSize) is greater than MaxBtSizeC. If treeType is equal to DUAL_TREE_CHROMA or, otherwise, greater than MaxBtSizeY, then y0+(1<<log2CbSize) is greater than SubPicBottomBorderInPic and (1<<log2CbSize) is greater than MaxBtSizeC.
[0146] Otherwise, if all of the following conditions are true, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 1. If treeType is equal to DUAL_TREE_CHROMA or, otherwise, greater than MinQtSizeY, then x0+(1<<log2CbSize) is greater than SubPicRightBorderInPic, y0+(1<<log2CbSize) is greater than SubPicBottomBorderInPic, and (1<<log2CbSize) is greater than MinQtSizeC. Otherwise, the value of qt_split_cu_flag[x0][y0] is assumed to be equal to 0.
[0147] The syntax and semantics of the multi-type tree are as follows.
[0148] [Table 9A] [Table 9B] [Table 9C]
[0149] mtt_split_cu_flag equal to 0 specifies that the coding unit is not split. mtt_split_cu_flag equal to 1 specifies that the coding unit is split into two coding units using bisection, or into three coding units using trisection, as indicated by the syntax element mtt_split_cu_binary_flag. The bisection or trisection can be either vertical or horizontal, as indicated by the syntax element mtt_split_cu_vertical_flag. When mtt_split_cu_flag is not present, the value of mtt_split_cu_flag is inferred as follows: The value of mtt_split_cu_flag is inferred to be equal to 1 if one or more of the following conditions are true: x0 + cbWidth is greater than SubPicRightBorderInPic, and y0 + cbHeight is greater than SubPicBottomBorderInPic. Otherwise, the value of mtt_split_cu_flag is inferred to be equal to 0.
[0150] The derivation process for temporal luma motion vector prediction is as follows: The output of this process is a motion vector prediction with 1 / 16 fractional sample precision, mvLXCol, and an availability flag, availableFlagLXCol. The variable currCb specifies the current luma coding block at luma position (xCb, yCb). The variables mvLXCol and availableFlagLXCol are derived as follows: If tile_group_temporal_mvp_enabled_flag is equal to 0, or if the reference picture is the current picture, then both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1 and the reference picture is not the current picture), the following ordered steps are applied: The bottom-right co-located motion vector is derived as follows: xColBr=xCb+cbWidth (8-355) yColBr=yCb+cbHeight (8-356)
[0151] If yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than SubPicBottomBorderInPic, and xColBr is less than SubPicRightBorderInPic, the following applies: The variable colCb specifies the luma coding block that contains the modified position given by ((xColBr>>3)<<3,(yColBr>>3)<<3) within the co-located picture specified by ColPic. The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the co-located picture specified by ColPic. The derivation process for co-located motion vectors is invoked with inputs currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag set equal to 0, and the output is assigned to mvLXCol and availableFlagLXCol. Otherwise, both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0.
[0152] The derivation process for temporal triangle merge candidates is as follows: The variables mvLXColC0, mvLXColC1, availableFlagLXColC0, and availableFlagLXColC1 are derived as follows: If tile_group_temporal_mvp_enabled_flag is equal to 0, then both components mvLXColC0 and mvLXColC1 are set equal to 0, and availableFlagLXColC0 and availableFlagLXColC1 are set equal to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the following ordered steps are applied: The bottom right co-located motion vector mvLXColC0 is derived as follows: xColBr=xCb+cbWidth (8-392) yColBr=yCb+cbHeight (8-393)
[0153] If yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than SubPicBottomBorderInPic, and xColBr is less than SubPicRightBorderInPic, the following applies: The variable colCb specifies the luma coding block that contains the modified position given by ((xColBr>>3)<<3,(yColBr>>3)<<3) within the co-located picture specified by ColPic. The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the co-located picture specified by ColPic. The derivation process for co-located motion vectors is called with inputs currCb, colCb, (xColCb, yColCb), refIdxLXColC0, and sbFlag set equal to 0, and the output is assigned to mvLXColC0 and availableFlagLXColC0. Otherwise, both components of mvLXColC0 are set equal to 0, and availableFlagLXColC0 is set equal to 0.
[0154] The derivation process for the constructed affine control point motion vector merge candidate is as follows: For X=0 and 1, the fourth (co-located bottom-right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list usage flag predFlagLXCorner[3], and availability flag availableFlagCorner[3] are derived as follows: For X=0 or 1, the reference index refIdxLXCorner[3] for the temporal merge candidate is set equal to 0. For X=0 or 1, the variables mvLXCol and availableFlagLXCol are derived as follows: If tile_group_temporal_mvp_enabled_flag is equal to 0, then both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the following applies: xColBr=xCb+cbWidth (8-566) yColBr=yCb+cbHeight (8-567)
[0155] If yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than SubPicBottomBorderInPic, and xColBr is less than SubPicRightBorderInPic, the following applies: The variable colCb specifies the luma coding block that contains the modified position given by ((xColBr>>3)<<3,(yColBr>>3)<<3) within the co-located picture specified by ColPic. The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the co-located picture specified by ColPic. The derivation process for co-located motion vectors is called with inputs currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag set equal to 0, and the output is assigned to mvLXCol and availableFlagLXCol. Otherwise, both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0. Replace all occurrences of pic_width_in_luma_samples with PicWidthInLumaSamples. Replace all occurrences of pic_height_in_luma_samples with PicHeightInLumaSamples.
[0156] In the second example embodiment, the syntax and semantics of the sequence parameter set RBSP are as follows:
[0157] [Table 10]
[0158] subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element subpic_id[i] in an SPS, spps_subpic_id in an SPPS that references an SPS, and tile_group_subpic_id in a tile group header that references an SPS. The value of subpic_id_len_minus1 shall be in the range Ceil(Log2(num_subpic_minus1+3)) to 8, inclusive. It is a bitstream compliance requirement that there be no overlap among subpicture[i] for i between 0 and num_subpic_minus1, inclusive. Each subpicture may be a temporal motion constrained subpicture.
[0159] The general semantics of a tile group header are as follows: tile_group_subpic_id identifies the subpicture to which the tile group belongs. The length of tile_group_subpic_id is subpic_id_len_minus1+1 bits. tile_group_subpic_id equal to 1 indicates that the tile group does not belong to any subpicture.
[0160] In the third example embodiment, the syntax and semantics of the NAL unit header are as follows:
[0161] [Table 11]
[0162] nuh_subpicture_id_len specifies the number of bits used to represent the syntax element that specifies the subpicture ID. When the value of nuh_subpicture_id_len is greater than 0, the first nuh_subpicture_id_len-th bits after nuh_reserved_zero_4bits specify the ID of the subpicture to which the payload of the NAL unit belongs. When nuh_subpicture_id_len is greater than 0, the value of nuh_subpicture_id_len shall be equal to the value of subpic_id_len_minus1 in the active SPS. The value of nuh_subpicture_id_len for non-VCL NAL units is constrained as follows: If nal_unit_type is equal to SPS_NUT or PPS_NUT, nuh_subpicture_id_len shall be equal to 0. nuh_reserved_zero_3bits shall be equal to '000'. A decoder SHALL ignore (eg, remove from the bitstream and discard) any NAL unit whose value of nuh_reserved_zero_3bits is not equal to '000'.
[0163] In a fourth example embodiment, the subpicture nesting syntax is as follows:
[0164] [Table 12]
[0165] all_sub_pictures_flag equal to 1 specifies that the nested SEI message applies to all subpictures. all_sub_pictures_flag equal to 1 specifies that the subpictures to which the nested SEI message applies are explicitly signaled by subsequent syntax elements. nesting_num_sub_pictures_minus1 plus 1 specifies the number of subpictures to which the nested SEI message applies. nesting_sub_picture_id[i] indicates the subpicture ID of the ith subpicture to which the nested SEI message applies. The nesting_sub_picture_id[i] syntax element is represented by Ceil(Log2(nesting_num_sub_pictures_minus1 + 1)) bits. sub_picture_nesting_zero_bit shall be equal to 0.
[0166] FIG. 9 is a schematic diagram of an example video coding device 900. The video coding device 900 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 900 includes a downstream port 920, an upstream port 950, and / or a transceiver unit (Tx / Rx) 910 including a transmitter and / or receiver for communicating data upstream and / or downstream over a network. The video coding device 900 also includes a processor 930 including a logic unit and / or central processing unit (CPU) for processing data, and a memory 932 for storing data. The video coding device 900 may also include electrical, optical-electrical (OE) components, electrical-optical (EO) components, and / or wireless communication components coupled to the upstream port 950 and / or downstream port 920 for communication of data over an electrical, optical, or wireless communication network. The video coding device 900 may also include input and / or output (I / O) devices 960 for communicating data to and from a user. The I / O devices 960 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. The I / O devices 960 may also include input devices such as a keyboard, mouse, trackball, etc. and / or corresponding interfaces for interacting with such output devices.
[0167] The processor 930 is implemented in hardware and software. The processor 930 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 930 communicates with the downstream port 920, the Tx / Rx 910, the upstream port 950, and the memory 932. The processor 930 includes a coding module 914. The coding module 914 implements the disclosed embodiments described above, such as the methods 100, 1000, 1100, and / or the mechanism 700, which may utilize the bitstream 500, the picture 600, and / or the picture 800. The coding module 914 may also implement any other method / mechanism described herein. Additionally, the coding module 914 may implement the codec system 200, the encoder 300, and / or the decoder 400. For example, the coding module 914 may be utilized to signal and / or obtain the location and size of subpictures within an SPS. In another example, the coding module 914 may constrain subpicture widths and subpicture heights to be multiples of the CTU size, unless such subpictures are located at the right border of a picture or the bottom border of a picture, respectively. In another example, the coding module 914 may constrain subpictures to encompass a picture without gaps or overlaps. In another example, the coding module 914 may be utilized to signal and / or obtain data indicating that some subpictures are temporal motion constrained subpictures and others are not. In another example, the coding module 914 may signal the complete set of subpicture IDs within an SPS and include the subpicture ID in each slice header to indicate the subpicture that contains the corresponding slice. In another example, the coding module 914 may signal the level for each subpicture.Thus, coding module 914 allows video coding device 900 to provide additional functionality and avoid certain processing to reduce processing overhead and / or increase coding efficiency when partitioning and coding video data. Thus, coding module 914 addresses challenges inherent in the art of video coding and improves the functionality of video coding device 900. Furthermore, coding module 914 effects transformation of video coding device 900 into a different state. Alternatively, coding module 914 can be implemented as instructions stored in memory 932 and executed by processor 930 (e.g., as a computer program product stored on a non-transitory medium).
[0168] Memory 932 comprises one or more memory types such as a disk, a tape drive, a solid state drive, a read-only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM), a static random access memory (SRAM), etc. Memory 932 may be used as an overflow data storage device for storing programs when such programs are selected for execution and for storing instructions and data read during program execution.
[0169] 10 is a flowchart of an example method 1000 of encoding a subpicture layout within a bitstream, such as bitstream 500, of a picture to support extraction of subpictures such as subpictures 522, 523, 622, 722, and / or 822. Method 1000 may be utilized by an encoder, such as codec system 200, encoder 300, and / or video coding device 900, when performing method 100.
[0170] Method 1000 may begin when an encoder receives a video sequence including multiple pictures and decides to encode the video sequence into a bitstream, for example, based on user input. The video sequence is partitioned into pictures / images / frames for further partitioning before encoding. In step 1001, a picture is partitioned into multiple sub-pictures, including a current sub-picture, hereafter referred to as a sub-picture. In step 1003, the sub-pictures are encoded into a bitstream.
[0171] In step 1005, the subpicture size and subpicture position of the subpicture are encoded into the SPS in the bitstream. The subpicture position includes the offset distance between the top-left sample of the subpicture and the top-left sample of the picture. The subpicture size includes the subpicture height in luma samples and the subpicture width in luma samples. A flag may also be encoded in the SPS to indicate that the subpicture is a motion-constrained subpicture. In such a case, the subpicture size and subpicture position indicate the layout of the motion-constrained subpicture.
[0172] In step 1007, a sub-picture ID is encoded into the SPS for each sub-picture partitioned from the picture. The number of sub-pictures partitioned from the picture may also be encoded into the SPS. In step 1009, the bitstream is stored for communication to the decoder. The bitstream may then be transmitted to the decoder as desired. In some examples, the sub-bitstream may be extracted from the encoded bitstream. In such cases, the transmitted bitstream is the sub-bitstream. In other examples, the encoded bitstream may be transmitted for extraction of the sub-bitstream at the decoder. In yet other examples, the encoded bitstream may be decoded and displayed without extraction of the sub-bitstream. In any of these examples, the size, position, ID, number of sub-pictures, and / or motion constraint sub-picture flags may be used to efficiently signal the sub-picture layout to the decoder.
[0173] 11 is a flowchart of an example method 1100 of decoding a bitstream, such as bitstream 500 and / or sub-bitstream 501, of a subpicture, such as subpictures 522, 523, 622, 722, and / or 822, based on a signaled subpicture layout. Method 1100 may be utilized by a decoder, such as codec system 200, decoder 400, and / or video coding device 900, when performing method 100. For example, method 1100 may be applied to decode a bitstream produced as a result of method 1000.
[0174] Method 1100 may begin when a decoder begins receiving a bitstream including subpictures. The bitstream may include a complete video sequence, or the bitstream may be a sub-bitstream including a reduced set of subpictures for separate extraction. In step 1101, a bitstream is received. The bitstream comprises subpictures partitioned from a picture. The bitstream also comprises an SPS. The SPS comprises subpicture sizes and subpicture positions. In some examples, the subpictures are temporal motion constrained subpictures. In such cases, the subpicture sizes and subpicture positions indicate the layout of the motion constrained subpictures. In some examples, the SPS may further comprise a subpicture ID for each subpicture partitioned from the picture.
[0175] In step 1103, the SPS is parsed to obtain a subpicture size and a subpicture position. The subpicture size may include a subpicture height in luma samples and a subpicture width in luma samples. The subpicture position may include an offset distance between the top-left sample of the subpicture and the top-left sample of the picture. The subpicture may also be parsed to obtain other subpicture-related data, such as a temporal motion constraint subpicture flag and / or a subpicture ID.
[0176] In step 1105, a size of the sub-picture may be determined relative to the size of the display based on the sub-picture size. Further, a position of the sub-picture may be determined relative to the display based on the sub-picture position. The decoder may also determine whether the sub-picture can be independently decoded based on the temporal motion constraint sub-picture flag. Thus, the decoder may determine the layout of the sub-picture based on parsed data from the SPS and / or corresponding data from slice headers associated with slices included in the sub-picture.
[0177] In step 1107, the sub-pictures are decoded based on the sub-picture size, sub-picture position, and / or other information obtained from the SPS, PPS, slice header, SEI message, etc. The sub-pictures are decoded to produce a video sequence. In step 1109, the video sequence can then be transmitted for display.
[0178] 12 is a schematic diagram of an example system 1200 for signaling subpicture layouts, such as layouts for subpictures 522, 523, 622, 722, and / or 822, via bitstreams, such as bitstream 500 and / or sub-bitstream 501. System 1200 may be implemented by an encoder and decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 900. Additionally, system 1200 may be utilized when implementing methods 100, 1000, and / or 1100.
[0179] System 1200 includes a video encoder 1202. Video encoder 1202 comprises a partition module 1201 for partitioning a picture into multiple sub-pictures, including a current sub-picture. Video encoder 1202 further comprises an encoding module 1203 for encoding the sub-pictures partitioned from the picture into a bitstream and encoding the sub-picture sizes and sub-picture positions of the sub-pictures into SPSs within the bitstream. Video encoder 1202 further comprises a storage module 1205 for storing the bitstream for communication to a decoder. Video encoder 1202 further comprises a transmission module 1207 for transmitting the bitstream including the sub-pictures, the sub-picture sizes, and the sub-picture positions to the decoder. Video encoder 1202 may be further configured to perform any of the steps of method 1000.
[0180] System 1200 also includes a video decoder 1210. The video decoder 1210 comprises a receiving module 1211 for receiving a bitstream comprising sub-pictures partitioned from a picture and an SPS comprising sub-picture sizes and sub-picture positions of the sub-pictures. The video decoder 1210 further comprises a parsing module 1213 for parsing the SPS to obtain the sub-picture sizes and sub-picture positions. The video decoder 1210 further comprises a decoding module 1215 for decoding the sub-pictures based on the sub-picture sizes and sub-picture positions to produce a video sequence. The video decoder 1110 further comprises a transport module 1217 for transporting the video sequence for display. The video decoder 1210 may be further configured to perform any of the steps of method 1100.
[0181] A first component is directly coupled to a second component when there are no intervening components other than wires, traces, or another medium between the first and second components. A first component is indirectly coupled to a second component when there are intervening components other than wires, traces, or another medium between the first and second components. The term "coupled" and variations thereof include both directly coupled and indirectly coupled. The use of the term "about" means a range that includes ±10% of the subsequent number, unless otherwise stated.
[0182] It should also be understood that the steps of the exemplary methods described herein are not necessarily required to be performed in the order described, and the order of steps in such methods should be understood to be exemplary only. Likewise, additional steps may be included in such methods, and certain steps may be omitted or combined, in methods consistent with various embodiments of the present disclosure.
[0183] Although several embodiments have been provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples should be considered as illustrative and not restrictive, and the intention is not to be limited to the details provided herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0184] Additionally, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of modifications, substitutions, and alterations are ascertainable by those skilled in the art and could be made without departing from the spirit and scope of what is disclosed herein. [Explanation of symbols]
[0185] 200 Codec System 201 segmented video signal 211 General-purpose coder control component 213 Transform Scaling and Quantization Components 215 Intra-picture Estimation Component 217 Intra-picture Prediction Component 219 Motion Compensation Component 221 Motion Estimation Component 223 Decoded Picture Buffer Component 225 In-Loop Filter Components 227 Filter Control Analysis Component 229 Scaling and Inverse Transformation Components 231 Header Formatting and CABAC Components 301 Segmented Video Signal 313 Transform and Quantize Components 317 Intra-picture Prediction Component 321 Motion Compensation Component 323 Decoded Picture Buffer Component 325 In-Loop Filter Components 329 Inverse Transform and Quantization Components 331 Entropy Coding Component 417 Intra-picture Prediction Component 421 Motion Compensation Component 423 Decoded Picture Buffer Component 425 In-Loop Filter Components 429 Inverse Transform and Quantization Components 433 Entropy Decoding Component 500 bitstream 501 Sub-Bitstream 510 SPS 512 PPS 514 slice header 515 SEI Message 520 Image Data 521 Pictures 522 Subpictures 523 Subpictures 524 slices 525 slices 531 Subpicture Size 532 Subpicture Position 533 Subpicture ID 534 Motion Constraint Subpicture Flag 535 Subpicture Level 600 pictures 622 Subpicture 631 Subpicture Size 631a Subpicture Width 631b Subpicture Height 632 position 633 Subpicture ID 634 Time-Motion Constrained Subpictures 642 Upper left sample 700 mechanism 722 Subpicture 724 slices 733 Subpicture ID 742 top left corner 801 Picture right border 802 Picture bottom border 822 Subpicture 825 CTU 900 Video Coding Device 910 Transmitter / Receiver 914 Coding Module 920 downstream ports 930 processor 932 memory 950 upstream ports 960 I / O devices 1200 System 1201 Division Module 1202 Video Encoder 1203 Encoding Module 1205 Memory Module 1207 Transmitting Module 1210 Video Decoder 1211 Receiver Module 1213 Analysis Module 1215 Decoding Module 1217 Transfer Module
Claims
1. 1. A method implemented in a decoder, comprising: receiving a bitstream comprising a current subpicture identified by a subpicture identifier (ID) and a sequence parameter set (SPS) comprising a subpicture size and a subpicture position of the current subpicture, wherein the current subpicture is partitioned from a picture, and the SPS further comprises a subpicture ID for each subpicture partitioned from the picture, the subpicture ID comprising the subpicture ID of the current subpicture; Parsing the SPS to obtain the subpicture size, the subpicture position, and the subpicture ID of the current subpicture; decoding the current subpicture based on the subpicture size, the subpicture position, and the subpicture ID of the current subpicture; A method for providing the above.
2. The method of claim 1 , wherein the subpicture position comprises an offset distance between a top-left sample of the current subpicture and a top-left sample of the picture.
3. The method of claim 1 or 2, wherein the subpicture size includes a subpicture height and a subpicture width.
4. The method of claim 3 , wherein the sub-picture height and the sub-picture width are specified in units of luma coding tree unit size (CtbSizeY).
5. 5. The method of claim 1, wherein when the right boundary of the current subpicture does not coincide with the right boundary of the picture, the subpicture width is an integer multiple of a luma coding tree unit size (CtbSizeY).
6. 6. The method of claim 1, wherein when the bottom boundary of the current subpicture does not coincide with the bottom boundary of the picture, the subpicture height is an integer multiple of a luma coding tree unit size (CtbSizeY).
7. 1. A method implemented in an encoder, comprising: encoding a current sub-picture identified by a sub-picture identifier (ID) into a bitstream, wherein the current sub-picture is partitioned from a picture; encoding a sub-picture size and a sub-picture position of the current sub-picture into a Sequence Parameter Set (SPS) of the bitstream; encoding a sub-picture ID for each sub-picture partitioned from the picture into the SPS, the sub-picture ID comprising the sub-picture ID of the current sub-picture; A method for providing the above.
8. The method of claim 7 , wherein the subpicture position comprises an offset distance between a top-left sample of the current subpicture and a top-left sample of the picture.
9. 9. The method of claim 7, wherein the subpicture size comprises a subpicture height and a subpicture width, the subpicture height and the subpicture width being specified in units of a luma coding tree unit size (CtbSizeY).
10. 10. The method of claim 7, wherein when the right boundary of the current subpicture does not coincide with the right boundary of the picture, the subpicture width is an integer multiple of a luma coding tree unit size (CtbSizeY).
11. 11. The method of claim 7, wherein when the bottom boundary of the current subpicture does not coincide with the bottom boundary of the picture, the subpicture height is an integer multiple of a luma coding tree unit size (CtbSizeY).
12. a receiving means for receiving a bitstream comprising a current subpicture identified by a subpicture identifier (ID) and a sequence parameter set (SPS) comprising a subpicture size and a subpicture position of the current subpicture, wherein the current subpicture is partitioned from a picture, and the SPS further comprises a subpicture ID for each subpicture partitioned from the picture, and the subpicture ID comprises the subpicture ID of the current subpicture; parsing means for parsing the SPS to obtain the subpicture size, the subpicture position, and the subpicture ID of the current subpicture; decoding means for decoding the current subpicture based on the subpicture size, the subpicture position, and the subpicture ID of the current subpicture; A decoder comprising:
13. Encoding a current subpicture identified by a subpicture identifier (ID) into a bitstream; encoding a sub-picture size and a sub-picture position of the current sub-picture into a sequence parameter set (SPS) of the bitstream; encoding means for encoding a sub-picture ID for each sub-picture separated from a picture into the SPS, the current sub-picture being separated from the picture, the sub-picture ID comprising the sub-picture ID of the current sub-picture; An encoder comprising:
14. 12. A computer program comprising a program code for performing the method according to any one of claims 1 to 11, when the computer program is run on a computer or processor.
15. 1. An electronic device comprising: a memory configured to store data including a bitstream, the bitstream comprising a current subpicture identified by a subpicture identifier (ID) and a sequence parameter set (SPS) comprising a subpicture size and a subpicture position of the current subpicture, the current subpicture being partitioned from a picture, the SPS further comprising a subpicture ID for each subpicture partitioned from the picture, the subpicture ID comprising the subpicture ID of the current subpicture; a transmitter configured to transmit the bitstream towards another device; An electronic device comprising:
16. 1. A method for storing a bitstream, comprising: receiving the bitstream, the bitstream comprising a current subpicture identified by a subpicture identifier (ID) and a sequence parameter set (SPS) comprising a subpicture size and a subpicture position of the current subpicture, the current subpicture being partitioned from a picture, the SPS further comprising a subpicture ID for each subpicture partitioned from the picture, the subpicture ID comprising the subpicture ID of the current subpicture; storing the bitstream on a computer-readable storage medium; A method for providing the above.
17. The method of claim 16 , further comprising transmitting the bitstream toward another device.
18. A decoder comprising processing circuitry for carrying out the method of any one of claims 1 to 6.
19. An encoder comprising processing circuitry for carrying out the method according to any one of claims 7 to 11.
20. 18. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by a video coding device, cause the video coding device to perform a method according to any one of claims 1 to 11 and 16 to 17.
Citation Information
Patent Citations
Coding schemes for virtual reality (VR) sequences
US20180101967A1
Concept for picture / video data streams allowing efficient reducibility or efficient random access
WO2017137444A1
Advanced video data stream extraction and multi-resolution video transmission
WO2018172234A2