Video coding bitstream extraction using identifier signaling - Patents.com
By including sub-picture information and a flag in the sub-bitstream, the mechanism addresses coding errors and resource inefficiencies in sub-bitstream extraction, improving efficiency and reducing resource usage.
Patent Information
- Application Number
- JP2022500125
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-05
- Filing Date
- 2020-06-15
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2040-06-15
AI Technical Summary
Existing video coding technologies face challenges in efficiently extracting and decoding sub-bitstreams when only a subset of sub-pictures is present, leading to coding errors and increased resource usage due to the need to transmit the entire bitstream.
Incorporating a mechanism that includes sub-picture information and a flag in the sub-bitstream to indicate the presence of this information, allowing decoders to correctly decode the subset of sub-pictures, thereby reducing resource usage and preventing errors.
This mechanism enhances coding efficiency by enabling sub-bitstream extraction without transmitting the entire bitstream, reducing processor, memory, and network resource usage in encoders and decoders.
Smart Images

Figure 0007810638000004 
Figure 0007810638000005 
Figure 0007810638000006
Abstract
Description
[Technical Field]
[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 870,892, filed July 5, 2019 by Ye-Kui Wang, entitled “Handling Signaled Slice Id for Bitstream Extraction,” which is incorporated herein by reference.
[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to video coding, and more particularly to bitstream extraction in video coding. [Background technology]
[0003] The amount of video data required to render even a relatively short video is substantial, which can pose challenges when the data is to be streamed or otherwise communicated over communication networks with limited bandwidth capacity. Therefore, video data is typically compressed before being communicated over modern telecommunications networks. Because memory resources can be limited, video size can also be an issue when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to code video data before transmission or storage, thereby reducing the amount of data needed to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and ever-increasing demands for higher video quality, improved compression and decompression techniques that increase compression ratios with little sacrifice in image quality are desirable. Summary of the Invention
[0004] In one embodiment, the present disclosure includes a method implemented in a decoder, the method including: receiving, by a receiver of the decoder, an extracted bitstream that is the result of a sub-bitstream extraction process from an input bitstream that includes a set of sub-pictures, wherein the extracted bitstream includes only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process; determining, by a processor of the decoder, a flag from the extracted bitstream that is set to indicate that sub-picture information related to the subset of sub-pictures is present in the extracted bitstream; obtaining, by the processor, one or more sub-picture identifiers (IDs) for the subset of sub-pictures based on the flag; decoding, by the processor, the subset of sub-pictures based on the sub-picture IDs, and decoding, by the processor, the subset of sub-pictures of the sub-pictures.
[0005] Some video coding sequences may include pictures that are encoded as a set of sub-pictures. A sub-picture may be associated with a sub-picture ID that can be used to indicate the sub-picture's location relative to the picture. In some cases, such sub-picture information can be inferred. In such cases, this sub-picture information can be excluded from the bitstream to increase coding efficiency. Certain processes may extract sub-bitstreams from the bitstream for independent transmission to end users. In such cases, the sub-bitstream contains only a subset of the sub-pictures included in the original bitstream. While sub-picture information can be inferred when all sub-pictures are present, such inference may be impossible at a decoder when only a subset of sub-pictures is present. This example includes a mechanism to prevent coding errors during sub-bitstream extraction. Specifically, when a sub-bitstream is extracted from a bitstream, the encoder and / or splicer includes sub-picture information for at least a subset of the sub-pictures in the sub-bitstream. Furthermore, the encoder / splice r includes a flag indicating that sub-picture information is included in the sub-bitstream. A decoder can read this flag, obtain the correct sub-picture information, and decode the sub-bitstream. Thus, the disclosed mechanism creates additional functionality in the encoder and / or decoder by avoiding errors. Furthermore, the disclosed mechanism may improve coding efficiency by enabling sub-bitstream extraction rather than transmitting the entire bitstream. This may reduce processor, memory, and / or network resource usage in the encoder and / or decoder.
[0006] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides further comprising obtaining, by the processor, a bit length of a syntax element that includes one or more sub-picture IDs.
[0007] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the flag, sub-picture ID, and length are obtained from a sequence parameter set (SPS) in the extracted bitstream.
[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the flag is a subpicture information present flag (subpic_info_present_flag).
[0009] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the subpicture ID is included within an SPS subpicture identifier (sps_subpic_id[i]) syntax structure.
[0010] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the length is included in the SPS subpic_id_len_minus1 plus 1 syntax structure.
[0011] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that a flag is required to be set to 1 to specify that sub-picture information is present for a coded layer video sequence (CLVS) and that each picture in the CLVS contains more than one sub-picture when the extracted bitstream is the result of a sub-bitstream extraction process from the input bitstream.
[0012] In one embodiment, the present disclosure includes a method implemented in an encoder, the method including: encoding, by a processor of the encoder, an input bitstream including a set of subpictures; performing, by the processor, a sub-bitstream extraction process on the input bitstream to generate an extracted bitstream including only a subset of the subpictures of the input bitstream; encoding, by the processor, into the extracted bitstream one or more subpicture IDs for the subset of subpictures in the extracted bitstream; setting, by the processor, a flag in the extracted bitstream to indicate that subpicture information related to the subset of subpictures is present in the extracted bitstream; and storing, by a memory coupled to the processor, the bitstream for communication to a decoder.
[0013] Some video coding sequences may include pictures that are encoded as a set of sub-pictures. A sub-picture may be associated with a sub-picture ID that can be used to indicate the sub-picture's location relative to the picture. In some cases, such sub-picture information can be inferred. In such cases, this sub-picture information can be excluded from the bitstream to increase coding efficiency. Certain processes may extract sub-bitstreams from the bitstream for independent transmission to end users. In such cases, the sub-bitstream contains only a subset of the sub-pictures included in the original bitstream. While sub-picture information can be inferred when all sub-pictures are present, such inference may be impossible at a decoder when only a subset of sub-pictures is present. This example includes a mechanism to prevent coding errors during sub-bitstream extraction. Specifically, when a sub-bitstream is extracted from the bitstream, the encoder and / or splicer includes sub-picture information for at least a subset of the sub-pictures in the sub-bitstream. Furthermore, the encoder / splicer includes a flag indicating that sub-picture information is included in the sub-bitstream. A decoder can read this flag, obtain the correct sub-picture information, and decode the sub-bitstream. Thus, the disclosed mechanism creates additional functionality in the encoder and / or decoder by avoiding errors. Furthermore, the disclosed mechanism may improve coding efficiency by enabling sub-bitstream extraction rather than transmitting the entire bitstream. This may reduce processor, memory, and / or network resource usage in the encoder and / or decoder.
[0014] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides further comprising encoding, by the processor, a bit length of a syntax element including one or more sub-picture IDs into the extracted bitstream.
[0015] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the flag, sub-picture ID, and length are encoded into the SPS in the extracted bitstream.
[0016] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the flag is subpic_info_present_flag.
[0017] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the subpicture ID is included in the sps_subpic_id[i] syntax structure.
[0018] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the length is included in the sps_subpic_id_len_minus1 plus 1 syntax structure.
[0019] Optionally, in any of the above aspects, another implementation of the aspect provides that a flag is required to be set to 1 to specify that sub-picture information is present for the CLVS and that each picture in the CLVS contains more than one sub-picture when the extracted bitstream is the result of a sub-bitstream extraction process from the input bitstream.
[0020] In one embodiment, the present disclosure includes a video coding device having: a processor; a receiver coupled to the processor; a memory coupled to the processor; and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to perform the method of any of the aforementioned aspects.
[0021] In one embodiment, the present disclosure is a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to perform a method of any of the foregoing aspects.
[0022] In one embodiment, the present disclosure includes a decoder having: receiving means for receiving an extracted bitstream that is the result of a sub-bitstream extraction process from an input bitstream that includes a set of sub-pictures, where the extracted bitstream includes only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process; determining means for determining that a flag from the extracted bitstream is set to indicate that sub-picture information related to the subset of sub-pictures is present in the extracted bitstream; obtaining means for obtaining one or more sub-picture IDs for the subset of sub-pictures based on the flag; decoding means for decoding the subset of sub-pictures based on the sub-picture IDs; and forwarding means for forwarding the subset of sub-pictures for display as part of a decoded video sequence.
[0023] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the decoder is further configured to perform the method of any of the aforementioned aspects.
[0024] In one embodiment, the present disclosure includes an encoder having: first encoding means for encoding an input bitstream including a set of sub-pictures; bitstream extraction means for performing a sub-bitstream extraction process on the input bitstream to generate an extracted bitstream including only a subset of the sub-pictures of the input bitstream; second encoding means for encoding one or more sub-picture IDs for the subset of sub-pictures in the extracted bitstream into the extracted bitstream; setting means for setting a flag in the extracted bitstream to indicate that sub-picture information associated with the subset of sub-pictures is present in the extracted bitstream; and storage means for storing the bitstream for communication to a decoder.
[0025] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the encoder is further configured to perform the method of any of the aforementioned aspects.
[0026] For clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.
[0027] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. [Brief explanation of the drawings]
[0028] For a more complete understanding of this disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0029] [Figure 1] 1 is a flowchart of an exemplary method for coding a video signal.
[0030] [Figure 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding.
[0031] [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder.
[0032] [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder.
[0033] [Figure 5] FIG. 1 is a schematic diagram illustrating multiple sub-picture video streams extracted from a picture video stream.
[0034] [Figure 6] FIG. 2 is a schematic diagram illustrating an exemplary bitstream split into sub-bitstreams.
[0035] [Figure 7] 1 is a schematic diagram of an exemplary video coding device.
[0036] [Figure 8] 1 is a flowchart of an example method for encoding a video sequence into a bitstream and extracting sub-bitstreams while mitigating ID errors.
[0037] [Figure 9] 1 is a flowchart of an exemplary method for decoding a video sequence from a sub-bitstream extracted from a bitstream.
[0038] [Figure 10] 1 is a schematic diagram of an example system for coding a video sequence of images in a bitstream and extracting sub-bitstreams while mitigating ID errors; DETAILED DESCRIPTION OF THE INVENTION
[0039] First, while exemplary implementations of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or in existence. The present disclosure is in no way limited to the exemplary implementations, drawings, and technologies shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the scope of the appended claims, along with their full range of equivalents.
[0040] The following terms are defined as follows, unless used herein in a contrary context. Specifically, the following definitions are intended to provide further clarity to the present disclosure. However, terms may be described differently in different contexts. Therefore, the following definitions should be considered supplemental and not limiting of other definitions provided for such terms herein.
[0041] A bitstream is a sequence of bits containing video data that is compressed for transmission between an encoder and a decoder. An encoder is a device configured to use an encoding process to compress video data into a bitstream. A decoder is a device configured to use a decoding process to reconstruct video data from a bitstream for display. A picture is an array of luma samples and / or chroma samples that generate a frame or its fields. The picture being encoded or decoded is sometimes referred to as the current picture for clarity of discussion. A subpicture is a rectangular region of one or more slices within a picture. The subbitstream extraction process is a specified mechanism that removes Network Abstraction Layer (NAL) units from a bitstream that are not part of a target set, resulting in an output subbitstream that contains NAL units included in the target set. A NAL unit is a syntax structure that contains bytes of data and an indication of the type of data contained therein. NAL units include Video Coding Layer (VCL) NAL units, which contain video data, and non-VCL NAL units, which contain supporting syntax data. An input bitstream is a bitstream containing the complete set of NAL units before the application of the sub-bitstream extraction process. An extracted bitstream, also known as a sub-bitstream, is a bitstream output from the bitstream extraction process that contains a subset of NAL units from the input bitstream. A set is a collection of individual items. A subset is a collection of items such that each item in the subset is included in the set and at least one item from the set is excluded from the subset. Subpicture information is any data that describes a subpicture. A flag is a data structure containing a sequence of bits that can be set to indicate corresponding data.A subpicture identifier (ID) is a data item that uniquely identifies the corresponding subpicture. The length of the data structure is the number of bits contained in the data structure. A coded layered video sequence (CLVS) is a sequence of encoded video data that contains one or more layers of pictures. A CLVS is sometimes called a coded video sequence (CVS) when it contains a single layer or when the CLVS is discussed outside of a layer-specific context. A sequence parameter set (SPS) is a parameter set that contains data related to a sequence of pictures. A decoded video sequence is a sequence of pictures reconstructed by a decoder in preparation for presentation to a user.
[0042] The following acronyms are used herein: Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Sub-Picture Unit (SPU), Versatile Video Coding (VVC), and Working Draft (WD).
[0043] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples of neighboring blocks in the same picture. Video blocks in an intra-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a picture may be coded by using spatial prediction with respect to reference samples in neighboring blocks in the same picture or by using temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a prediction block that represents an image block. Residual data represents pixel differences between the original image block and the prediction block. Thus, an intra-coded block is encoded according to a motion vector that points to a block of reference samples that form the prediction block and the residual data that indicates the difference between the coded block and the prediction block. An intra-coded block is encoded according to an intra-coding mode and the residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain.These result in residual transform coefficients, which may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding may be applied to achieve more compression. Such video compression techniques are described in more detail below.
[0044] To ensure that the encoded video is accurately decoded, the video is encoded and decoded according to a corresponding video coding standard, including Advanced Video Coding (AVC), also known as International Telecommunication Union (ITU-T) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC) and Multiview Video Coding plus Depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has initiated the development of a video coding standard called Versatile Video Coding (VVC). VVC is contained in working drafts (WDs) including JVET-N1001-v8.
[0045] A video coding sequence includes a sequence of pictures. In some cases, such a picture can be further partitioned into a set of sub-pictures, with each sub-picture comprising a separate region of the picture. Sub-pictures allow different spatial portions of a picture to be treated differently at the decoder. For example, in a virtual reality (VR) context, only a portion of the entire picture is displayed to the user. Thus, sub-pictures can be used to transmit different portions of a picture to the decoder at different resolutions and / or even omit certain portions of the picture. This can increase coding efficiency. In another example, a videoconferencing application may dynamically increase the size and / or resolution of an image of an actively speaking participant and decrease the size and resolution of a participant's image when the participant stops speaking. Including each participant in a different sub-picture allows such dynamic changes for one participant without affecting the images for other participants. Sub-pictures can be associated with a sub-picture ID, which uniquely identifies the corresponding sub-picture. Thus, the sub-picture ID can be used to indicate the position of the sub-picture relative to the picture and / or to modify the sub-picture level coding process. In some cases, sub-picture information such as the sub-picture ID can be inferred. For example, if a picture contains nine sub-pictures, the sub-picture ID can be inferred by the decoder to be an index ranging from zero to eight. In such cases, this sub-picture information can be omitted from the bitstream to increase coding efficiency.
[0046] However, some processes may extract sub-bitstreams from a bitstream for independent transmission to an end user. In such cases, the sub-bitstream contains only a subset of the sub-pictures included in the original bitstream. While sub-picture information can be inferred when all sub-pictures are present, such inference may not be possible at a decoder when only a subset of sub-pictures is present. As an example, an encoder may send only sub-pictures 3 of 9 and 4 of 9 to a decoder. If sub-picture information is omitted, the decoder may not be able to determine which sub-pictures have been received and how such sub-pictures should be displayed. In such cases, the bitstream is considered to be a conforming bitstream because the missing data associated with the bitstream can be inferred. However, the extracted sub-bitstream is not conforming because some missing data associated with the sub-bitstream cannot be inferred.
[0047] Disclosed herein is a mechanism for preventing coding errors during sub-bitstream extraction. Specifically, when a sub-bitstream is extracted from a bitstream, an encoder and / or splicer encodes sub-picture information for at least a subset of the sub-bitstream's sub-pictures into the sub-bitstream's parameters. Furthermore, the encoder / splice includes a flag indicating that the sub-picture information is included in the sub-bitstream. A decoder can read the flag, obtain the correct sub-picture information, and decode the sub-bitstream. Such sub-picture information may include a sub-picture ID in a syntax element and a length data element indicating the bit length of the sub-picture ID syntax element. Thus, the disclosed mechanism creates additional functionality for the encoder and / or decoder by avoiding coding errors related to sub-pictures. Furthermore, the disclosed mechanism may improve coding efficiency by enabling sub-bitstream extraction rather than transmitting the entire bitstream. This may reduce processor, memory, and / or network resource usage in the encoder and / or decoder.
[0048] 1 is a flowchart of an exemplary operational method 100 for coding a video signal. Specifically, a video signal is encoded by an encoder. The encoding process compresses the video signal by using various mechanisms to reduce the video file size. The smaller file size allows the compressed video file to be transmitted to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process, allowing the decoder to consistently reconstruct the video signal.
[0049] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device, such as a video camera, and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, create the visual impression of movement. The frames include pixels represented in terms of light, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional displays.
[0050] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels by 64 pixels). CTUs contain both luma samples and chroma samples. The coding tree can be used to divide the CTUs into blocks and then recursively subdivide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of a frame can be subdivided until each block contains relatively uniform illumination values. Furthermore, the chroma component of a frame can be subdivided until each block contains relatively uniform color values. Thus, the partitioning mechanism varies depending on the content of the video frame.
[0051] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that common scene objects tend to appear in consecutive frames. Thus, a block depicting an object in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a constant position across multiple frames. Thus, the table may be described once, and adjacent frames may reference the reference frame. A pattern matching mechanism may be used to match objects across multiple frames. Furthermore, moving objects may be represented across multiple frames, for example, due to object motion or camera motion. As a specific example, a video may show a car moving across the screen over multiple frames. To describe such motion, a motion vector may be used. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. In this manner, inter prediction may encode an image block in a current frame as a set of motion vectors that indicate an offset from a corresponding block in a reference frame.
[0052] Intra prediction encodes blocks within a common frame. Intra prediction takes advantage of the fact that luma and chroma components tend to cluster within a frame. For example, a green patch in a tree tends to be located adjacent to similar green patches. Intra prediction uses multi-directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar / identical to samples in neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. Planar mode effectively indicates a smooth transition of light / color across a row / column by using a relatively constant slope in changing values. DC mode is used for boundary smoothing and indicates that the block is similar / identical to the average value associated with samples in all neighboring blocks associated with the angular direction of the directional prediction mode. Therefore, intra-predicted blocks can represent image blocks as various related prediction mode values instead of their actual values. Furthermore, inter-predicted blocks can represent image blocks as motion vector values instead of actual values. In either case, the predicted block may not exactly represent the image block in some cases. Any differences are stored in a residual block. Transforms may be applied to the residual block to further compress the file.
[0053] Various filtering techniques may be applied in step 107. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may result in the generation of blocky images at the decoder. Furthermore, the block-based prediction scheme may encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to a block / frame. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference block, such that the artifacts are less likely to generate additional artifacts in subsequent blocks that are encoded based on the reconstructed reference block.
[0054] Once the video signal has been partitioned, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data described above, as well as any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Generating the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across many frames and blocks. The order depicted in FIG. 1 is for clarity and ease of discussion and is not intended to limit the video coding process to any particular order.
[0055] The decoder receives the bitstream and begins the decoding process in step 111. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. The decoder uses syntax data from the bitstream to determine the frame partitions in step 111. The partitioning should be consistent with the results of the block partitioning in step 103. The entropy encoding / decoding used in step 111 is now described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible choices based on the spatial positioning of values in the input image(s). Signaling the exact selection may use multiple bins. As used herein, a bin is a binary value (e.g., a bit value that can change depending on the context) that is treated as a variable. Entropy coding allows the encoder to discard any options that are clearly infeasible in a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of allowable options (e.g., one bin for two options, two bins for three to four options, etc.). The encoder then encodes the codeword for the selected option. This scheme reduces the size of the codeword because it is desirable for the codeword to be large enough to uniquely indicate a selection from a small subset of allowable options, as opposed to uniquely indicating a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of allowable options in a similar manner as the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the selection made by the encoder.
[0056] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate a residual block. The decoder then uses the residual block and a corresponding prediction block to reconstruct an image block according to the partitioning. The prediction block may include both intra-predicted blocks and inter-predicted blocks, as generated in the encoder in step 105. The reconstructed image block is then positioned within a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding, as described above.
[0057] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters may be applied to the frames to remove blocking artifacts. Once the frames have been filtered, the video signal can be output to a display in step 117 for viewing by an end user.
[0058] 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support the implementation of operational method 100. Codec system 200 is generalized to illustrate components used in both encoders and decoders. Codec system 200 receives and partitions a video signal, as described with reference to steps 101 and 103 of operational method 100, resulting in partitioned video signal 201. When operating as an encoder, codec system 200 then compresses partitioned video signal 201 into a coded bitstream, as described with reference to steps 105, 107, and 109 of method 100. When operating as a decoder, codec system 200 generates an output video signal from the bitstream, as described with reference to steps 111, 113, 115, and 117 of operational method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All of the components of codec system 200 may reside within an encoder. A decoder may include a subset of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described next.
[0059] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to subdivide the blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. The blocks can be referred to as nodes on the coding tree. Large parent nodes are divided into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. In some cases, the partitioned blocks can be included in a coding unit (CU). For example, a CU can be a subpart of a CTU that includes a luma block, red-difference chroma (Cr) block(s), and blue-difference chroma (Cb) block(s), along with the corresponding syntax instructions for the CU. Partitioning modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are used to divide a node into two, three, or four child nodes of different shapes, respectively, depending on the partitioning mode used. The partitioned video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0060] The general coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be made based on storage / bandwidth availability and image resolution requirements. The general coder control component 211 also manages buffer utilization in relation to transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general coder control component 211 manages partitioning, prediction, and filtering by other components. For example, the general coder control component 211 may dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the general coder control component 211 controls other components of the codec system 200 to balance bitrate concerns with video signal reconstruction quality. The general coder control component 211 generates control data that controls the operation of other components. Control data is also forwarded to the header format and CABAC component 231 to be encoded in the bitstream to signal parameters for decoding at the decoder.
[0061] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0062] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are illustrated separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors that estimate the motion of video blocks. A motion vector may indicate, for example, the displacement of a coded object relative to a predictive block. A predictive block is a block that is found to closely match a block to be coded in terms of pixel differences. A predictive block is also referred to as a reference block. Such pixel differences may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, CTUs are divided into CTBs, which are then divided into CBs for inclusion in CUs. CUs can be encoded as prediction units (PUs), which contain prediction data, and / or transform units (TUs), which contain the transform residual data of the CUs. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for a current block / frame and select the reference block, motion vector, etc. with the best rate-distortion performance, which balances both the quality of the video reconstruction (e.g., the amount of data lost due to compression) and the coding efficiency (e.g., the size of the final encoding).
[0063] In some examples, codec system 200 may calculate values for sub-integer pixel positions of reference pictures stored in decoded picture buffer component 223. For example, video codec system 200 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of reference pictures. Accordingly, motion estimation component 221 may perform a motion search for whole pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision. Motion estimation component 221 calculates motion vectors for PUs of video blocks in inter-coded slices by comparing the positions of the PUs with the positions of predictive blocks of reference pictures. Motion estimation component 221 outputs the calculated motion vectors as motion data to header formatting and CABAC component 231 for encoding and outputs motion to motion compensation component 219.
[0064] The motion compensation performed by motion compensation component 219 may include fetching or generating a predictive block based on a motion vector determined by motion estimation component 221. Again, in some examples, motion estimation component 221 and motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the PU of the current video block, motion compensation component 219 may identify the location of the predictive block to which the motion vector points. A residual video block is then formed by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. Generally, motion estimation component 221 performs motion estimation on the luma component, and motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predictive block and residual block are forwarded to transform scaling and quantization component 213.
[0065] The partitioned video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are illustrated separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block relative to blocks within the current frame, instead of the inter-prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219, as described above. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode for encoding the current block from multiple tested intra-picture prediction modes. The selected intra-prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.
[0066] For example, the intra picture estimation component 215 calculates rate-distortion values for various tested intra prediction modes using rate-distortion analysis and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block encoded to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. In addition, the intra picture estimation component 215 may be configured to code depth blocks of the depth map using a rate-distortion optimization (RDO)-based depth modeling mode (DMM).
[0067] The intra-picture prediction component 217 may generate a residual block from the prediction block based on a selected intra-prediction mode determined by the intra-picture estimation component 215 when implemented in an encoder, or may read the residual block from the bitstream when implemented in a decoder. The residual block contains value differences between the prediction block and the original block represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both the luma and chroma components.
[0068] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block, generating a video block including residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling may involve applying a scale factor to the residual information so that different frequency information is quantized with different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transform coefficients, which are forwarded to the header formatting and CABAC component 231 for encoding in the bitstream.
[0069] The scaling and inverse transform component 229 applies the inverse operation of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct a residual block in the pixel domain, for example, for use as a reference block that may later become a predictive block for another current block. The motion estimation component 221 and / or motion compensation component 219 may calculate a reference block by adding the residual block back to the corresponding predictive block for use in motion estimation for a later block / frame. A filter is applied to the reconstructed reference block to mitigate artifacts generated during scaling, quantization, and transform. Such artifacts may otherwise cause inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.
[0070] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 may be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes multiple parameters for adjusting how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is forwarded to the header format and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., of reconstructed pixel blocks) or the frequency domain, depending on the example.
[0071] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded picture buffer component 223 stores and forwards the reconstructed and filtered blocks toward a display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0072] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to a decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all information desired by a decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, an indication of the most likely intra-prediction mode, an indication of partition information, etc. Such data may be encoded using entropy coding. For example, the information may be encoded using context-adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0073] 3 is a block diagram illustrating an example video encoder 300. Video encoder 300 may be used to implement the encoding functionality of codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of method of operation 100. Encoder 300 partitions an input video signal, resulting in a partitioned video signal 301, which is substantially similar to partitioned video signal 201. Partitioned video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.
[0074] Specifically, the partitioned video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and corresponding prediction blocks (along with associated control data) are forwarded to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially similar to the header format and CABAC component 231 .
[0075] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into a reference block used by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter of the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters, as discussed with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0076] 4 is a block diagram illustrating an exemplary video decoder 400. Video decoder 400 may be used to implement the decoding functionality of codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of operating method 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.
[0077] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into the residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0078] The reconstructed residual block and / or predictive block are forwarded to the intra-picture prediction component 417 for reconstruction into an image block based on an intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to locate a reference block within a frame and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and corresponding inter-prediction data are forwarded to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses motion vectors from reference blocks to generate prediction blocks and applies a residual block to the result to reconstruct an image block. The resulting reconstructed blocks may also be forwarded to the decoded picture buffer component 423 via an in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that can be reconstructed into frames via partition information. Such frames may also be arranged in a sequence. This sequence is output to a display as a reconstructed output video signal.
[0079] 5 is a schematic diagram illustrating multiple sub-picture video streams 501, 502, and 503 extracted from a picture video stream 500. For example, each of the sub-picture video streams 501-503 and / or the picture video stream 500 may be encoded by an encoder, such as codec system 200 and / or encoder 300, in accordance with method 100. Furthermore, the sub-picture video streams 501-503 and / or the picture video stream 500 may be decoded by a decoder, such as codec system 200 and / or decoder 400.
[0080] Picture video stream 500 includes multiple pictures presented over time. As shown in FIG. 5, picture video stream 500 is configured for use in virtual reality (VR) applications. VR works by coding a sphere of video content that can be displayed as if the user were at the center of the sphere. Each picture contains the entire sphere, while only a portion of the picture, known as a viewport, is displayed to the user. For example, a user may use a head-mounted display (HMD) that selects and displays a viewport of the sphere based on the user's head movement. This creates the impression of being physically present in the virtual space depicted by the video. To achieve this result, each picture in a video sequence contains the entire sphere of video data at the corresponding instant in time. However, only a small portion of the image (e.g., a single viewport) is displayed to the user. The remainder of the picture is discarded at the decoder without being rendered. The entire picture can be transmitted so that different viewports can be dynamically selected and displayed depending on the user's head movement.
[0081] In the illustrated example, each picture in picture video stream 500 can be subdivided into sub-pictures based on available viewports. Thus, each picture and corresponding sub-picture includes a temporal position (e.g., picture order) as part of the temporal presentation. Sub-picture video streams 501-503 are created when sub-division is applied consistently over time. Such consistent sub-division produces sub-picture video streams 501-503, each including a set of sub-pictures of a predetermined size, shape, and spatial position relative to the corresponding picture in picture video stream 500. Furthermore, the sets of sub-pictures in sub-picture video streams 501-503 vary in temporal position over presentation time. Thus, the sub-pictures in sub-picture video streams 501-503 can be aligned in the temporal domain based on their temporal position. The sub-pictures from the sub-picture video streams 501-503 at each temporal position can then be merged in the spatial domain based on their predetermined spatial positions to reconstruct the picture video stream 500 for display. Specifically, the sub-picture video streams 501-503 can each be encoded into a separate sub-bitstream. When such sub-bitstreams are merged together, they result in a bitstream that includes the entire set of pictures spanning time. The resulting bitstream can be sent to a decoder for decoding and displayed based on the user's currently selected viewport.
[0082] One issue with VR video is that all of the sub-picture video streams 501-503 can be transmitted to the user at high quality (e.g., high resolution). This allows the decoder to dynamically select the user's current viewport and display the sub-picture(s) from the corresponding sub-picture video streams 501-503 in real time. However, the user may only see a single viewport, for example, from the sub-picture video stream 501, while the sub-picture video streams 502-503 are discarded. Transmitting the sub-picture video streams 502-503 at high quality in this way can use a significant amount of bandwidth without providing a corresponding benefit to the user. To improve coding efficiency, VR video can be encoded into multiple video streams 500, with each video stream 500 encoded at a different quality / resolution. In this way, the decoder can transmit a request for the current sub-picture video stream 501. In response, the encoder (or intermediate slicer or other content server) can select high-quality sub-picture video stream 501 from high-quality video stream 500 and select low-quality sub-picture video streams 502-503 from low-quality video stream 500. The encoder can then merge such sub-bitstreams together into a fully encoded bitstream for transmission to the decoder. In this way, the decoder receives a series of images where the current viewport is at high quality and the other viewports are at lower quality. Furthermore, the highest-quality sub-pictures are generally displayed to the user (without head movement), and the lower-quality sub-pictures are generally discarded, which balances functionality and coding efficiency.
[0083] If the user switches from viewing sub-picture video stream 501 to sub-picture video stream 502, the decoder requests that the new sub-picture video stream 502 be transmitted at high quality. The encoder can then change the merging mechanism accordingly.
[0084] Picture video stream 500 is included to illustrate a practical use of sub-pictures. Note that sub-pictures have many uses, and this disclosure is not limited to VR technology. For example, sub-pictures may also be used in videoconferencing systems. In such cases, each user's video feed is included in a sub-picture bitstream, such as sub-picture video streams 501, 502, and / or 503. The system can receive such sub-picture video streams 501, 502, and / or 503 and combine them at different positions, resolutions, etc. to generate a complete picture video stream 500 for transmission back to the user. This allows the videoconferencing system to dynamically modify picture video stream 500 based on changing user input, for example, by increasing or decreasing the size of sub-picture video streams 501, 502, and / or 503 to emphasize a user who is currently speaking or de-emphasize a user who is no longer speaking. Sub-pictures therefore have many applications that allow for dynamic runtime modification of picture video stream 500 based on changes in user behavior. This functionality can be achieved by extracting and / or combining sub-picture video streams 501, 502, and / or 503 from and / or into picture video stream 500.
[0085] 6 is a schematic diagram illustrating an exemplary bitstream 600 split into sub-bitstreams 601. Bitstream 600 may include a picture video stream such as picture video stream 500, and sub-bitstream 601 may include sub-picture video streams such as sub-picture video streams 501, 502, and / or 503. For example, bitstream 600 and sub-bitstream 601 may be generated by codec system 200 and / or encoder 300 and / or decoder 400 for decoding by codec system 200. As another example, bitstream 600 and sub-bitstream 601 may be generated by an encoder in step 109 of method 100 for use by a decoder in step 111.
[0086] The bitstream 600 includes a sequence parameter set (SPS) 610, multiple picture parameter sets (PPSs) 611, multiple slice headers 615, and image data 620. The SPS 610 includes sequence data common to all pictures in the video sequence included in the bitstream 600. Such data may include picture size, bit depth, coding tool parameters, bit rate limits, etc. The PPS 611 includes parameters that apply to the entire picture. Thus, each picture in the video sequence may reference the PPS 611. Note that while each picture references the PPS 611, a single PPS 611 can include data for multiple pictures in some instances. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS 611 can include data for such similar pictures. The PPS 611 can indicate the coding tools available for slices in the corresponding picture, quantization parameters, offsets, etc. The slice header 615 includes parameters specific to each slice in the picture. Thus, there may be one slice header 615 per slice in a video sequence. The slice header 615 may include slice type information, a picture order count (POC), a reference picture list, prediction weights, tile entry points, deblocking parameters, etc. Note that the slice header 615 may also be referred to as a tile group header in some contexts. Note that in some examples, the bitstream 600 may also include a picture header, which is a syntax structure that includes parameters that apply to all slices in a single picture. For this reason, the picture header and the slice header 615 may be used interchangeably in some contexts. For example, certain parameters may be moved between the slice header 615 and the picture header depending on whether such parameters are common to all slices in the picture.
[0087] The image data 620 includes video data encoded according to inter-prediction, intra-prediction, and / or inter-layer prediction, as well as corresponding transformed and quantized residual data. For example, a video sequence includes multiple pictures 621. A picture 621 is an array of luma samples and / or chroma samples that generate a frame or a field thereof. A frame is a complete image intended to be displayed, fully or partially, to a user at a corresponding instant in time in a video sequence. A picture 621 includes one or more slices. A slice may be defined as an integer number of complete tiles of the picture 621 or an integer number of consecutive complete CTU rows (e.g., within a tile) contained exclusively in a single NAL unit. A slice is further divided into CTUs and / or CTBs. A CTU is a group of samples of a predefined size that can be partitioned by a coding tree. A CTB is a subset of a CTU and contains the luma or chroma components of the CTU. CTUs / CTBs are further divided into coding blocks based on the coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.
[0088] Picture 621 can be divided into multiple sub-pictures 623 and 624. Sub-pictures 623 and / or 624 are rectangular regions of one or more slices within picture 621. Thus, each of the slices, and its sub-divisions, can be assigned to a sub-picture 623 and / or 624. This allows different regions of picture 621 to be treated differently from a coding perspective, depending on which sub-pictures 623 and / or 624 are included in such region.
[0089] Sub-bitstream 601 can be extracted from bitstream 600 according to sub-bitstream extraction process 605. Sub-bitstream extraction process 605 is a designated mechanism that removes NAL units from the bitstream that are not part of a target set, resulting in an output sub-bitstream that includes NAL units included in the target set. NAL units include slices. Thus, sub-bitstream extraction process 605 retains a target set of slices and removes other slices. The target set can be selected based on sub-picture boundaries. In the illustrated example, slices included in sub-picture 623 are included in the target set, and slices included in sub-picture 624 are not included in the target set. As such, sub-bitstream extraction process 605 generates sub-bitstream 601, which is substantially similar to bitstream 600 but includes sub-picture 623, while excluding sub-picture 624. Sub-bitstream extraction process 605 can be performed by an encoder and / or associated slicer configured to dynamically modify bitstream 600 based on user behavior / requests.
[0090] Thus, sub-bitstream 601 is an extracted bitstream that is the result of sub-bitstream extraction process 605 applied to input bitstream 600. Input bitstream 600 includes a set of subpictures. However, the extracted bitstream (e.g., sub-bitstream 601) includes only a subset of the subpictures of input bitstream 600 to sub-bitstream extraction process 605. In the illustrated example, the set of subpictures included in input bitstream 600 includes subpictures 623 and 624, while the subset of subpictures of sub-bitstream 601 includes subpicture 623 but not subpicture 624. Note that any number of subpictures 623-624 can be used. For example, bitstream 600 may include N subpictures 623-624, and sub-bitstream 601 may include N-1 or fewer subpictures 623, where N is any integer value.
[0091] The sub-bitstream extraction process 605 may, in some cases, generate coding errors. For example, subpictures 623-624 may be associated with subpicture information, such as a subpicture ID. The subpicture ID uniquely identifies the corresponding subpicture, such as subpicture 623 or 624. Thus, the subpicture ID can be used to indicate the location of subpictures 623-624 relative to picture 621 and / or to modify the subpicture-level coding process. In some cases, subpicture information can be inferred based on the location of subpictures 623-624. Thus, the bitstream 600 may omit such subpicture information associated with subpictures 623 and 624 to reduce the amount of data in the bitstream 600 for increased coding efficiency. However, if subpicture 623 or subpicture 624 is not present, a decoder may not be able to infer such subpicture information. Thus, a simple sub-bitstream extraction process 605 can be applied to a conforming bitstream 600 to produce a non-conforming sub-bitstream 601. A bitstream 600 / sub-bitstream 601 is conforming when the bitstream 600 / sub-bitstream 601 conforms to a standard such as VVC, and therefore can be correctly decoded by any decoder that also conforms to the standard. As such, a simple sub-bitstream extraction process 605 can convert a decodable bitstream 600 into a non-decodable sub-bitstream 601.
[0092] To address this issue, this disclosure includes an improved sub-bitstream extraction process 605. Specifically, the sub-bitstream extraction process 605 encodes sub-picture IDs for the sub-picture(s) 623 in the sub-bitstream 601, even if such sub-picture IDs were omitted from the bitstream 600. For example, the sub-picture IDs may be included in an SPS sub-picture identifier (sps_subpic_id[i]) syntax structure 635. The sps_subpic_id[i] syntax structure 635 is included in the SPS 610 and includes i sub-picture IDs, where i is the number of sub-picture(s) 623 included in the sub-bitstream 601. Furthermore, the sub-bitstream extraction process 605 may also encode the length in bits of syntax elements (e.g., sps_subpic_id[i] syntax structure 635) that include one or more sub-picture IDs in the extracted bitstream. For example, the length may be included in the SPS Subpicture ID Length Minus 1 (sps_subpic_id_len_minus1) syntax structure 633. The sps_subpic_id_len_minus1 syntax structure 633 may contain the length in bits of the sps_subpic_id[i] syntax structure 635 minus 1 (minus 1). The minus 1 coding approach encodes the value as 1 less than the actual value to save bits. A decoder can derive the actual value by adding 1. Thus, the sps_subpic_id_len_minus1 syntax structure 633 may also be referred to as sps_subpic_id_len_minus1+1. Thus, a decoder can use the sps_subpic_id_len_minus1 syntax structure 633 to determine the number of bits associated with the sps_subpic_id[i] syntax structure 635, and can therefore use the sps_subpic_id_len_minus1 syntax structure 633 to interpret the sps_subpic_id[i] syntax structure 635.The decoder can then decode the subpicture 623 based on the sps_subpic_id_len_minus1 syntax structure 633 and the sps_subpic_id[i] syntax structure 635.
[0093] Additionally, the sub-bitstream extraction process 605 may encode / set a flag in the sub-bitstream 601 to indicate that sub-picture information related to the sub-picture 623 is present in the sub-bitstream 601. As a specific example, the flag may be encoded as a sub-picture information present flag (subpic_info_present_flag) 631. Accordingly, the subpic_info_present_flag 631 may be set to indicate that sub-picture information related to a subset of the sub-pictures is present in the extracted bitstream (sub-bitstream 601), such as the sps_subpic_id_len_minus1 syntax structure 633 and the sps_subpic_id[i] syntax structure 635. Additionally, a decoder can read subpic_info_present_flag 631 to determine whether subpicture information related to a subset of subpictures, such as sps_subpic_id_len_minus1 syntax structure 633 and sps_subpic_id[i] syntax structure 635, is present in the extracted bitstream (sub-bitstream 601). As a particular example, the encoder / slicer can require that a flag be set to 1 to specify that subpicture information is present for a coded layered video sequence (CLVS) and that each picture 621 of the CLVS contains more than one subpicture 623 and 624 when the extracted bitstream (sub-bitstream 601) is the result of a sub-bitstream extraction process 605 from the input bitstream 600. A CLVS is a sequence of encoded video data that includes one or more layers of pictures. A layer is a set of NAL units that all have a particular layer ID value. The pictures 621 may or may not be organized into multiple layers, where all the pictures 621 of a corresponding layer have similar characteristics, such as size, resolution, signal-to-noise ratio (SNR), etc.
[0094] The foregoing information is described in more detail herein below. HEVC may use regular slices, dependent slices, tiles, and wavefront parallel processing (WPP) as partitioning schemes. These partitioning schemes may be applied to maximum transfer unit (MTU) size matching, parallel processing, and end-to-end delay reduction. Each regular slice may be encapsulated in a separate NAL unit. Entropy coding dependency and in-picture prediction, including intra-sample prediction, motion information prediction, and coding mode prediction, may be disabled across slice boundaries. Therefore, regular slices can be reconstructed independently from other regular slices within the same picture. However, slices may still have some interdependencies due to loop filtering operations.
[0095] Regular slice-based parallelization may not require significant inter-processor or inter-core communication. One exception is that inter-processor and / or inter-core data sharing may be important for motion compensation when decoding predictively coded pictures. Such processes may involve more processing resources than inter-processor or inter-core data sharing for in-picture prediction. However, for the same reasons, the use of regular slices may result in substantial coding overhead due to the bit cost of slice headers and the lack of prediction across slice boundaries. Furthermore, regular slices also serve as a mechanism for bitstream partitioning to meet MTU size requirements due to the in-picture independence of regular slices and the fact that each regular slice is encapsulated in a separate NAL unit. In many cases, the goals of parallelization and MTU size matching conflict with requirements on the slice layout of a picture.
[0096] Dependent slices have short slice headers and allow for bitstream partitioning at treeblock boundaries without breaking in-picture prediction. Dependent slices provide for the division of regular slices into multiple NAL units, which reduces end-to-end delay by allowing parts of a regular slice to be transmitted before the encoding of the entire regular slice is finished.
[0097] In WPP, a picture is partitioned into a single row of CTBs. Entropy decoding and prediction may use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows. The start of decoding of a CTB row may be delayed by one or two CTBs, depending on the example, to ensure that data associated with the CTB above and to the right of the subject CTB is available before the subject CTB is decoded. This staggered start creates the appearance of a wavefront. This process supports parallelization with up to as many processors / cores as there are CTB rows in the picture. Because in-picture prediction is possible between adjacent treeblock rows within a picture, inter-processor / inter-core communication to enable in-picture prediction can be important. WPP partitioning does not result in the generation of additional NAL units. Therefore, WPP may not be used for MTU size matching. However, if MTU size matching is required, regular slices can be used with WPP with constant coding overhead.
[0098] Tiles define horizontal and vertical boundaries that partition a picture into tile columns and rows. The scan order of the CTBs may be local within a tile, in the order of the tile's CTB raster scan. Thus, a tile may be completely decoded before decoding the top-left CTB of the next tile in the picture's tile raster scan order. Similar to regular slices, tiles break not only entropy decoding dependencies but also in-picture prediction dependencies. However, tiles may not be included in individual NAL units. Thus, tiles may not be used for MTU size matching. Each tile can be processed by one processor / core. Inter-processor / inter-core communication used for in-picture prediction between processing units decoding adjacent tiles may be limited to conveying a shared slice header when a slice contains more than one tile and loop filtering related to the sharing of reconstructed samples and metadata. When more than one tile or WPP segment is included in a slice, the entry point byte offset for each tile or WPP segment other than the first tile or WPP segment of the slice may be signaled in the slice header.
[0099] For simplicity, HEVC uses certain restrictions on the application of the four different picture partitioning schemes. A coded video sequence may not contain both tiles and wavefronts in most profiles specified by HEVC. Furthermore, for each slice and / or tile, one or both of the following conditions must be met: All coded treeblocks in a slice are contained in the same tile. Furthermore, all coded treeblocks in a tile are contained in the same slice. Additionally, a wavefront segment contains exactly one CTB row. When WPP is used, a slice that starts within a CTB row must end in the same CTB row.
[0100] In VVC, tiles define horizontal and vertical boundaries that partition a picture into columns and rows of tiles. VVC allows tiles to be further divided horizontally to form bricks. Tiles that are not further divided can also be considered bricks. The scan order of the CTBs is changed to be local within a brick (e.g., the order of the CTB raster scan of the brick). The current brick is completely decoded before decoding the top-left CTB of the next brick in the brick raster scan order of the picture.
[0101] A slice in VVC may contain one or more bricks. Each slice is encapsulated in a separate NAL unit. Entropy coding dependencies and in-picture prediction, including intra-sample prediction, motion information prediction, and coding mode prediction, may be disabled across slice boundaries. Therefore, regular slices can be reconstructed independently from other regular slices within the same picture. VVC includes rectangular slices and raster scan slices. A rectangular slice may contain one or more bricks that occupy a rectangular area within a picture. A raster scan slice may contain one or more bricks that are in the raster scan order of the bricks within the picture.
[0102] VVC-based WPP is similar to HEVC WPP, except that HEVC WPP has two CTU delays, while VVC WPP has one CTU delay. With HEVC WPP, a new decoding thread can start decoding the first CTU of an assigned CTU row after the first two CTUs of the previous CTU row have already been decoded. With VVC WPP, a new decoding thread can start decoding the first CTU of an assigned CTU row after the first CTU of the previous CTU row has been decoded.
[0103] An example signaling of tiles, bricks, and slices in a PPS is as follows: [Table 1-1] [Table 1-2] [Table 1-3]
[0104] Prior systems have several problems. For example, when a bitstream is initially encoded, slices in a picture in the bitstream may be partitioned into rectangular slices. In this case, slice IDs may be omitted from the PPS. In this case, the value of signalled_slice_id_flag may be set equal to zero in the PPS of the bitstream. However, when one or more rectangular slices from a bitstream are extracted to form another bitstream, slice IDs should be present in the PPS of the bitstream generated from such an extraction process.
[0105] Generally, this disclosure describes handling of signaled slice IDs to assist the bitstream extraction process. The description of this technique is based on VVC, but may also be applied to other video codec specifications.
[0106] An exemplary mechanism for addressing the above-listed problems is as follows: A method is disclosed for extracting one or more slices from a picture of a bitstream, denoted as bitstream A, and generating a new bitstream B from the extraction process. Bitstream A includes at least one picture. The picture includes multiple slices. The method includes parsing a parameter set from bitstream A and rewriting the parameters to bitstream B. The value of signalled_slice_id_flag is set to 1 in the rewritten parameter set. When a signalled_slice_id_length_minus1 syntax element is present in the parameter set of bitstream A, the value of signalled_slice_id_flag is copied to the rewritten parameter set. When the signalled_slice_id_length_minus1 syntax element is not present in the parameter set of bitstream A, the value of signalled_slice_id_flag is set in the rewritten parameter set. For example, signalled_slice_id_flag may be set to Ceil(Log2(num_slices_in_pic_minus1 + 1))-1, where num_slices_in_pic_minus1 is equal to the number of slices in the picture of bitstream A minus one. One or more slices are extracted from bitstream A. An extracted bitstream B is then generated.
[0107] Example PPS semantics are as follows: signalled_slice_id_flag set to 1 may specify that the slice ID of each slice is signaled. signalled_slice_id_flag set to zero may specify that the slice ID is not signaled. When rect_slice_flag is equal to zero, the value of signalled_slice_id_flag may be inferred to be equal to zero. In the case of a bitstream that is the result of a sub-bitstream extraction and that results in a bitstream containing a subset of the slices originally contained in the picture, the value of signalled_slice_id_flag should be set equal to 1 for the PPS. signalled_slice_id_length_minus1 plus 1 may specify the number of bits used to represent the syntax element slice_id[i], when present, and the syntax element slice_address in the slice header. The value of signalled_slice_id_length_minus1 may range from 0 to 15, inclusive. If not present, the value of signalled_slice_id_length_minus1 may be inferred to be equal to Ceil(Log2(num_slices_in_pic_minus1 + 1))-1. If this is the result of a sub-bitstream extraction, and the result is a bitstream containing a subset of the slices originally contained in the picture, the value of signalled_slice_id_length_minus1 for the PPS should remain unchanged.
[0108] 7 is a schematic diagram of an exemplary video coding device 700. The video coding device 700 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 700 includes a downstream port 720, an upstream port 750, and / or a transceiver unit (Tx / Rx) 710 including a transmitter and / or receiver for communicating data upstream and / or downstream over a network. The video coding device 700 also includes a processor 730 including a logic unit and / or central processing unit (CPU) for processing data and a memory 732 for storing data. The video coding device 700 may also include electrical, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the upstream port 750 and / or downstream port 720 for communicating data over an electrical, optical, or wireless communication network. Video coding device 700 may also include input and / or output (I / O) devices 760 for communicating data to and from a user. I / O devices 760 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. I / O devices 760 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0109] The processor 730 is implemented in hardware and software. The processor 730 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 730 is in communication with the downstream port 720, the Tx / Rx 710, the upstream port 750, and the memory 732. The processor 730 includes a coding module 714. The coding module 714 may implement the disclosed embodiments described herein, such as methods 100, 800, and / or 900, using the bitstream 600 and / or sub-bitstream 601 including the picture video stream 500 and / or the sub-picture video streams 501-503. The coding module 714 may also implement any other method / mechanism described herein. Furthermore, coding module 714 may implement codec system 200, encoder 300, and / or decoder 400. For example, coding module 714 can be used to extract sub-bitstreams from a bitstream, include sub-picture information in the sub-bitstream during the extraction process, and / or include a flag in the sub-bitstream to indicate that the sub-bitstream contains sub-picture information. Thus, coding module 714 allows video coding device 700 to provide additional functionality and / or coding efficiency when coding video data. Thus, coding module 714 improves the functionality of video coding device 700 and addresses problems specific to video coding techniques. Furthermore, coding module 714 affects the transformation of video coding device 700 into different states. Alternatively, coding module 714 is implemented as instructions stored in memory 732 and executed by processor 730 (e.g., as a computer program product stored on a non-transitory medium).
[0110] Memory 732 may include one or more memory types such as a disk, tape drive, solid state drive, read-only memory (ROM), random access memory (RAM), flash memory, ternary content addressable memory (TCMA), static random access memory (SRAM), etc. Memory 732 may be used as an overflow data storage device to store programs when they are selected for execution and to store instructions and data read during program execution.
[0111] 8 is a flowchart of an example method 800 for encoding a video sequence into a bitstream, such as bitstream 600, and extracting a sub-bitstream, such as sub-bitstream 601, while mitigating ID errors. Method 800 may be used by an encoder, such as codec system 200, encoder 300, and / or video coding device 700, when performing method 100 to encode picture video stream 500 and / or sub-picture video streams 501-503.
[0112] Method 800 may begin when an encoder receives a video sequence including multiple pictures and determines, for example, based on user input, to encode the video sequence into a bitstream. In step 801, the encoder encodes an input bitstream, such as picture video stream 500 and / or bitstream 600, that includes a set of subpictures. For example, the bitstream may include VR video data and / or videoconferencing video data. The set of subpictures may include multiple subpictures. Furthermore, the subpictures may be associated with a subpicture ID.
[0113] In step 803, the encoder and / or associated slicer performs a sub-bitstream extraction process on the input bitstream to generate an extracted bitstream, such as sub-picture video streams 501-503 and / or sub-bitstream 601. The extracted bitstream includes only a subset of the subpictures of the input bitstream. Specifically, the extracted bitstream includes only subpictures included in the set of subpictures of the input bitstream. Furthermore, the extracted bitstream excludes one or more of the subpictures from the set of subpictures of the input bitstream. Thus, the input bitstream may include the CLVS of a picture, and the extracted bitstream may include the CLVS of the subpictures of the picture.
[0114] In step 805, the encoder encodes into the extracted bitstream one or more subpicture IDs for a subset of the subpictures of the extracted bitstream. For example, such subpicture IDs may be excluded from the input bitstream. Thus, the encoder may encode such subpicture IDs into the extracted bitstream to support decoding of the subpictures included in the extracted bitstream. For example, the subpicture IDs may be included / encoded in an sps_subpic_id[i] syntax structure in the extracted bitstream.
[0115] In step 807, the encoder encodes the bit length of the syntax element containing one or more subpicture IDs into the extracted bitstream. For example, the length of the subpicture IDs may be excluded from the input bitstream. Thus, the encoder may encode the length of the subpicture IDs into the extracted bitstream to support decoding of the subpictures included in the extracted bitstream. For example, the length may be included / encoded in an sps_subpic_id_len_minus1 plus 1 syntax structure in the extracted bitstream.
[0116] In step 809, the encoder may set a flag in the extracted bitstream to indicate that subpicture information related to a subset of subpictures is present in the extracted bitstream. The flag may indicate to the decoder that subpicture IDs and / or subpicture ID lengths are present in the extracted bitstream. For example, the flag may be subpic_info_present_flag. In certain examples, the flag is required to be set to 1 to specify that subpicture information is present in the CLVS (e.g., included in the input bitstream and / or the extracted bitstream) and that each picture in the CLVS contains more than one subpicture when the extracted bitstream is the result of a subbitstream extraction process from the input bitstream. In some examples, the flag, subpicture IDs, and lengths are encoded into the SPS in the extracted bitstream.
[0117] In step 811, the encoder stores the bitstream for communication to the decoder. In some examples, the bitstream can then be transmitted to the decoder. For example, the bitstream can be transmitted to the decoder upon request by the decoder, for example, based on a user request.
[0118] 9 is a flowchart of an example method 900 for decoding a video sequence from a sub-bitstream, such as sub-bitstream 601, extracted from a bitstream, such as bitstream 600. Method 900 may be used by decoders, such as codec system 200, decoder 400, and / or video coding device 700, when performing method 100 to decode picture video stream 500 and / or sub-picture video streams 501-503.
[0119] Method 900 may begin when a decoder begins receiving a sub-bitstream extracted from a bitstream, for example, as a result of method 800. In step 901, the decoder receives the extracted bitstream. The extracted bitstream is the result of a sub-bitstream extraction process from an input bitstream that includes a set of sub-pictures. The extracted bitstream includes only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process. Specifically, the extracted bitstream includes only the sub-pictures included in the set of sub-pictures of the input bitstream. Furthermore, the extracted bitstream excludes one or more sub-pictures from the set of sub-pictures of the input bitstream. Thus, the input bitstream may include a CLVS of a picture, and the extracted bitstream may include a CLVS of the sub-pictures of the picture. The received extracted bitstream may also be referred to as a sub-bitstream. For example, the extracted bitstream may include sub-picture(s) including VR video data and / or videoconferencing video data.
[0120] In step 903, the decoder determines that a flag from the extracted bitstream is set to indicate that subpicture information related to a subset of subpictures is present in the extracted bitstream. The flag may indicate that a subpicture ID and / or a subpicture ID length is present in the extracted bitstream. For example, the flag may be subpic_info_present_flag. In a particular example, the flag is required to be set to 1 to specify that subpicture information is present in the CLVS (e.g., included in the input bitstream and / or the extracted bitstream) and that each picture in the CLVS contains more than one subpicture when the extracted bitstream is the result of a subbitstream extraction process from the input bitstream.
[0121] In step 905, the decoder obtains the length in bits of the syntax element containing one or more subpicture IDs. For example, the length of the subpicture IDs may be excluded from the input bitstream but included in the extracted bitstream. For example, the length may be included in an sps_subpic_id_len_minus1 plus 1 syntax structure in the received extracted bitstream.
[0122] In step 907, the decoder obtains one or more subpicture IDs for the subset of subpictures based on the flags and / or based on the lengths. For example, the decoder can use the flags to determine that subpicture IDs are present. The decoder can then use the lengths to determine the boundaries of the subpicture ID data within the bitstream. For example, subpicture IDs may be excluded from the input bitstream but included in the extracted bitstream. For example, subpicture IDs may be included in an sps_subpic_id[i] syntax structure within the extracted bitstream. In some examples, the flags, subpicture IDs, and lengths are obtained from the SPS in the extracted bitstream.
[0123] In step 909, the decoder may decode a subset of the sub-pictures in the extracted bitstream based on the sub-picture IDs obtained in step 907. The decoder may then, in step 911, forward the subset of sub-pictures for display as part of the decoded video sequence.
[0124] 10 is a schematic diagram of an example system 1000 for coding a video sequence of images in a bitstream, such as bitstream 600, and extracting a sub-bitstream, such as sub-bitstream 601, while mitigating ID errors. Accordingly, system 1000 may be used to code picture video stream 500 and / or sub-picture video streams 501-503. System 1000 may be implemented by an encoder and decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 700. Furthermore, system 1000 may be used when implementing methods 100, 800, and / or 900.
[0125] The system 1000 includes a video encoder 1002. The video encoder 1002 comprises a first encoding module 1001 for encoding an input bitstream including a set of subpictures. The video encoder 1002 further comprises a bitstream extraction module 1004 for performing a sub-bitstream extraction process on the input bitstream to generate an extracted bitstream including only a subset of the subpictures of the input bitstream. The video encoder 1002 further comprises a second encoding module 1003 for encoding one or more subpicture IDs for the subset of subpictures in the extracted bitstream into the extracted bitstream. The video encoder 1002 further comprises a setting module 1005 for setting a flag in the extracted bitstream to indicate that subpicture information associated with the subset of subpictures is present in the extracted bitstream. The video encoder 1002 further comprises a storage module 1007 for storing the bitstream for communication to a decoder. The video encoder 1002 further comprises a transmission module 1009 for transmitting the bitstream toward a video decoder 1010. The video encoder 1002 may be further configured to perform any of the steps of the method 800.
[0126] The system 1000 also includes a video decoder 1010. The video decoder 1010 has a receiving module 1011 for receiving an extracted bitstream that is the result of a sub-bitstream extraction process from an input bitstream that includes a set of sub-pictures, where the extracted bitstream includes only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process. The video decoder 1010 further has a determining module 1013 for determining that a flag from the extracted bitstream is set to indicate that sub-picture information related to the subset of sub-pictures is present in the extracted bitstream. The video decoder 1010 further has an obtaining module 1015 for obtaining one or more sub-picture IDs for the subset of sub-pictures based on the flag. The video decoder 1010 further has a decoding module 1017 for decoding the subset of sub-pictures based on the sub-picture IDs. The video decoder 1010 further has a transferring module 1019 for transferring the subset of sub-pictures for display as part of a decoded video sequence. The video decoder 1010 may be further configured to perform any of the steps of the method 900.
[0127] A first component is directly coupled to a second component when there are no intervening components, other than lines, traces, or another medium, between the first and second components. A first component is indirectly coupled to a second component when there are intervening components, other than lines, traces, or other medium, between the first and second components. The term "coupled" and variations thereof include both directly coupled and indirectly coupled. The use of the term "about," unless otherwise specified, means a range that includes ±10% of the subsequent number.
[0128] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of steps of such methods should be understood to be merely exemplary. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in methods consistent with various embodiments of the present disclosure.
[0129] While several embodiments are provided in this disclosure, it will be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples should be considered illustrative and not limiting, and the intention is not to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or features may be omitted, or not implemented.
[0130] Additionally, techniques, systems, subsystems, and methods described and illustrated as separate or distinct in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alterations will be ascertainable by those skilled in the art and could be made without departing from the spirit and scope disclosed herein.
Claims
1. 1. A method implemented in an encoder, the method comprising: encoding, by a processor of the encoder, an input bitstream; performing, by the processor, a sub-bitstream extraction process on the input bitstream to generate an extracted bitstream; encoding, by the processor, into the extracted bitstream, a picture parameter set (PPS) and a subset of the rectangular slices originally included in the picture, wherein for the extracted bitstream, a value of a flag in the PPS will be set equal to 1; setting, by the processor, the flag in the extracted bitstream, wherein the flag equal to 1 specifies that a slice identifier (ID) for each of the subset of the plurality of rectangular slices is signaled in the PPS, and the flag equal to 0 specifies that a slice ID for the subset of the plurality of rectangular slices is not signaled in the PPS; storing, by a memory coupled to the processor, the bitstream for communication to a decoder; method.
2. and encoding, by the processor, a bit length of a syntax element including the slice ID into the extracted bitstream. The method of claim 1.
3. 1. A video coding device comprising:
3. A system comprising: a processor; a receiver coupled to the processor; a memory coupled to the processor; and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method of claim 1 or 2. Video coding device.
4. 3. A non-transitory computer-readable medium containing a computer program for use by a video coding device, the computer program comprising computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, causes the video coding device to perform the method of claim 1 or 2. Non-transitory computer-readable medium.
5. An encoder comprising: first encoding means for encoding an input bitstream; bitstream extraction means for performing a sub-bitstream extraction process on the input bitstream to generate an extracted bitstream; second encoding means for encoding a picture parameter set (PPS) and a subset of the plurality of rectangular slices originally included in the picture into the extracted bitstream, wherein for the extracted bitstream, a value of a flag in the PPS will be set equal to 1; setting means for setting the flag in the extracted bitstream, wherein the flag equal to 1 specifies that a slice identifier (ID) for each of the subset of the plurality of rectangular slices is signaled in the PPS, and the flag equal to 0 specifies that a slice ID for the subset of the plurality of rectangular slices is not signaled in the PPS; storage means for storing said bitstream for communication to a decoder; Encoder.
6. The encoder is further configured to perform the method of claim 1 or 2. The encoder of claim 5 .