Explicit address indication in video coding
By adopting a new addressing scheme in video decoding, using identifiers and striped address lengths to determine the address of the sub-image, the problem of difficult sub-image address management in the prior art is solved, and efficient sub-code stream extraction and display are achieved.
Patent Information
- Application Number
- CN202510252073.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-06
- Filing Date
- 2019-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-12-31
AI Technical Summary
In video decoding, it is difficult for the prior art to effectively manage the address of the sub-image, especially in applications such as virtual reality, which causes the upper left corner index of the sub-image received by the decoder to be no 0, affecting the normal operation of the index-based addressing scheme.
Using a new addressing scheme, by encoding the identifiers and stripe address lengths of the sub-image-related characters and stripe address lengths in the code stream, the decoder can determine the stripe address from the stripe header based on this information, realizing the correct decoding and display of the sub-image without rewriting the stripe header.
This solution can extract and display subcode streams without rewriting the stripe header, reducing the use of network resources, memory resources and processing resources, and improving coding efficiency and system performance.
Smart Images

Figure CN120050434A_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 201980087375.7, the original application date is December 31, 2019, and the entire content of the original application is incorporated herein by reference.
[0002] Cross - reference to related applications
[0003] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 787,110, filed on December 31, 2018, entitled "Explicit Tile Identifier (ID) Signaling" by FNU Hendry et al., and U.S. Provisional Patent Application No. 62 / 883,537, filed on August 6, 2019, entitled "Explicit Tile Identifier (ID) Signaling" by FNU Hendry et al., the content of which is incorporated herein by reference. Technical field
[0004] The present invention generally relates to video coding, and more particularly to address management when extracting sub - images from an image in video coding. Background art
[0005] Even relatively short videos require a large amount of video data to describe, which can cause difficulties when the data is to be streamed or otherwise transmitted over a communication network with limited bandwidth capacity. Therefore, video data is usually compressed first and then transmitted in modern telecommunication networks. Since memory resources may be limited, the size of the video can also be a problem when storing the video on a storage device. Video compression devices typically use software and / or hardware on the source side to encode the video data and then transmit or store it, thereby reducing the amount of data required to represent a digital video image. Then, the compressed data is received on the destination side by a video decompression device that decodes the video data. In the context of limited network resources and the growing demand for higher video quality, improved compression and decompression techniques are needed, which can increase the compression ratio with little impact on image quality. Summary of the invention
[0006] In one embodiment, the present invention includes a method implemented in a decoder. The method includes: a receiver of the decoder receives a sub-bitstream, wherein the sub-bitstream includes a sub-image in an image, wherein the image is segmented into a plurality of strips including a first strip, a parameter set related to the image and the sub-image, and a strip header related to the first strip; a processor of the decoder parses the parameter set to obtain an identifier and a strip address length of the first strip; the processor determines the strip address of the first strip from the strip header according to the identifier and the strip address length; the processor decodes the sub-bitstream to generate a video sequence including the sub-image of the first strip; the processor sends the video sequence including the sub-image for display. In some video coding systems, a slice (also referred to as a tile group) can be addressed according to a set of indices. These indices can start from index 0 at the upper left corner of the image and increase in raster scan order, ending at index N at the lower right corner of the image, where N is the number of indices minus 1. Such a system is suitable for most applications. However, certain applications, such as virtual reality (VR), only display sub-images in an image. Some systems only send the sub-bitstream in the bitstream to the decoder when streaming VR content, which can improve coding efficiency, wherein the sub-bitstream includes the sub-image to be displayed. In this case, since the upper left corner of the sub-image received by the decoder is usually an index other than 0, the index-based addressing scheme may stop functioning properly. To solve these problems, it may be necessary for the encoder (or related stripper) to rewrite each strip header to change the index of the sub-image so that the upper left corner index starts from 0 and the remaining sub-image strips are adjusted accordingly. Dynamically rewriting the strip header (e.g., according to each user request) may require a lot of calculations by the processor. The disclosed system employs an addressing scheme that can extract the sub-bitstream including the sub-image without the need to rewrite the strip header. Each strip is addressed according to an identifier (ID) other than the index (e.g., sub-image ID). In this way, regardless of which sub-image is received and regardless of the position of the received sub-image relative to the upper left corner of the complete image, the decoder can consistently determine all relevant addresses. Since the ID is arbitrarily defined (e.g., selected by the encoder), the ID is encoded in a variable length field. Accordingly, the strip address length is also indicated. The ID related to the sub-image is also indicated. The length is used to parse the strip address, and the sub-image ID is used to map the strip address from the image-based position to the sub-image-based position. By adopting these mechanisms, the encoder, decoder, and / or related stripper can be improved. For example, the sub-bitstream can be extracted and sent instead of the entire bitstream, which can reduce the use of network resources, memory resources, and / or processing resources.In addition, extracting the sub-bitstream in this way can avoid rewriting each strip header according to each user request, which further reduces the use of network resources, memory resources, and / or processing resources.
[0007] Optionally, according to any of the above aspects, in another implementation of the aspect, the identifier is related to the sub-image.
[0008] Optionally, according to any of the above aspects, in another implementation of the aspect, the strip address length represents the number of bits included in the strip address.
[0009] Optionally, according to any of the above aspects, in another implementation of the aspect, determining the strip address of the first strip includes: the processor uses the length in the parameter set to determine the bit boundary to parse the strip address from the strip header; the processor uses the identifier and the strip address to map the strip address from the image-based position to the sub-image-based position.
[0010] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes: the processor parses the parameter set to obtain an identifier (ID) flag, where the ID flag indicates that the mapping relationship can be used to map the strip address from the image-based position to the sub-image-based position.
[0011] Optionally, according to any of the above aspects, in another implementation of the aspect, the mapping relationship between the image-based position and the sub-image-based position aligns the strip header with the sub-image without rewriting the strip header.
[0012] Optionally, in any of the above aspects, in another implementation of the aspect, the strip address includes a defined value but does not include an index.
[0013] In one embodiment, the present invention includes a method implemented in an encoder. The method includes: a processor of the encoder encoding an image in a bitstream, wherein the image includes a plurality of slices, and the plurality of slices includes a first slice; the processor encoding a slice header in the bitstream, wherein the slice header includes a slice address of the first slice; the processor encoding a parameter set in the bitstream, wherein the parameter set includes an identifier and a slice address length of the first slice; the processor extracting a sub-bitstream from the bitstream by: extracting the first slice without rewriting the slice header according to the slice address of the first slice, the slice address length, and the identifier; storing the sub-bitstream in a memory of the encoder for transmission to a decoder. In some video coding systems, a slice (also referred to as a tile group) can be addressed according to a set of indices. These indices can start from index 0 at the upper left corner of the image and increase in raster scan order, ending at index N at the lower right corner of the image, where N is the number of indices minus 1. Such a system is suitable for most applications. However, certain applications, such as virtual reality (VR), only display a sub-image in the image. Some systems only send the sub-bitstream in the bitstream to the decoder when streaming VR content, which can improve coding efficiency, wherein the sub-bitstream includes the sub-image to be displayed. In this case, since the upper left corner of the sub-image received by the decoder is usually an index other than 0, the index-based addressing scheme may stop functioning properly. To solve these problems, it may be necessary for the encoder (or related slicer) to rewrite each slice header to change the index of the sub-image so that the upper left corner index starts from 0 and the remaining sub-image slices are adjusted accordingly. Dynamically rewriting the slice header (e.g., according to each user request) may require a lot of calculations by the processor. The disclosed system employs an addressing scheme that can extract a sub-bitstream including a sub-image without rewriting the slice header. Each slice is addressed according to an identifier (ID) other than an index (e.g., a sub-image ID). In this way, the decoder can consistently determine all relevant addresses regardless of which sub-image is received and regardless of the position of the received sub-image relative to the upper left corner of the complete image. Since the ID is arbitrarily defined (e.g., selected by the encoder), the ID is encoded in a variable-length field. Accordingly, the slice address length is also indicated. The sub-image ID related to the sub-image is also indicated. The length is used to parse the slice address, and the sub-image ID is used to map the slice address from an image-based position to a sub-image-based position. By adopting these mechanisms, the encoder, decoder, and / or related slicer can be improved.For example, instead of extracting and sending the entire bitstream, a sub-bitstream can be extracted and sent, which can reduce the use of network resources, memory resources, and / or processing resources. In addition, extracting the sub-bitstream in this way does not require rewriting each strip header according to each user request, which further reduces the use of network resources, memory resources, and / or processing resources.
[0014] Optionally, according to any of the above aspects, in another implementation of the aspect, the identifier is related to a sub-image.
[0015] Optionally, according to any of the above aspects, in another implementation of the aspect, the strip address length represents the number of bits included in the strip address.
[0016] Optionally, according to any of the above aspects, in another implementation of the aspect, the length in the parameter set includes data sufficient to parse the strip address from the strip header, and the identifier includes data sufficient to map the strip address from an image-based location to a sub-image-based location.
[0017] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further includes: the processor encodes an identifier (ID) flag in the parameter set, where the ID flag indicates that the mapping relationship can be used to map the strip address from the image-based location to the sub-image-based location.
[0018] Optionally, according to any of the above aspects, in another implementation of the aspect, the strip address includes a defined value but does not include an index.
[0019] Optionally, according to any of the above aspects, in another implementation of the aspect, extracting the sub-bitstream from the bitstream includes: extracting a sub-image from the image, where the sub-image includes the first strip, and the sub-bitstream includes the sub-image, the strip header, and the parameter set.
[0020] In one embodiment, the present invention includes a video decoding device. The video decoding device includes: a processor, a memory, a receiver coupled to the processor, and a transmitter coupled to the processor, where the processor, the memory, the receiver, and the transmitter are configured to execute the method according to any of the above aspects.
[0021] In one embodiment, the present invention includes a non-transitory computer-readable medium. The non-transitory computer-readable medium includes a computer program product for use by a video decoding device; the computer program product includes computer-executable instructions stored in the non-transitory computer-readable medium; when the processor executes the computer-executable instructions, the video decoding device is caused to perform the method according to any of the above aspects.
[0022] In one embodiment, the present invention includes a decoder. The decoder includes: a receiving module for receiving a sub-bitstream, where the sub-bitstream includes a sub-image in an image, where the image is segmented into a plurality of strips including a first strip, a parameter set related to the image and the sub-image, and a strip header related to the first strip; a parsing module for parsing the parameter set to obtain an identifier and a strip address length of the first strip; a determining module for determining the strip address of the first strip from the strip header according to the identifier and the strip address length; a decoding module for decoding the sub-bitstream to generate a video sequence including the sub-image of the first strip; and a sending module for sending the video sequence including the sub-image of the first strip for display.
[0023] Optionally, according to any of the above aspects, in another implementation of the aspect, the decoder is further configured to perform the method according to any of the above aspects.
[0024] In one embodiment, the present invention includes an encoder. The encoder includes: an encoding module for: encoding an image in a bitstream, where the image includes a plurality of strips, the plurality of strips including a first strip; encoding a strip header in the bitstream, where the strip header includes the strip address of the first strip; encoding a parameter set in the bitstream, where the parameter set includes an identifier and a strip address length of the first strip; an extraction module for extracting a sub-bitstream from the bitstream by: extracting the first strip without rewriting the strip header according to the strip address of the first strip, the strip address length, and the identifier; and a storage module for storing the sub-bitstream for sending to a decoder.
[0025] Optionally, according to any of the above aspects, in another implementation of the aspect, the encoder is further configured to perform the method according to any of the above aspects.
[0026] For clarity, any of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present invention.
[0027] These and other features can be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] To understand the present invention more thoroughly, reference is now made to the following description of the drawings in conjunction with the specific embodiments, in which like reference numerals represent like components.
[0029] Figure 1 A flowchart of an exemplary method for encoding a video signal.
[0030] Figure 2 A schematic diagram of an exemplary encoding and decoding (codec) system for video decoding.
[0031] Figure 3 A schematic diagram of an exemplary video encoder.
[0032] Figure 4 A schematic diagram of an exemplary video decoder.
[0033] Figure 5 A schematic diagram of an example sub-bitstream extracted from a bitstream.
[0034] Figure 6 A schematic diagram of an exemplary image segmented for encoding.
[0035] Figure 7 A schematic diagram of an example sub-image extracted from an image.
[0036] Figure 8 A schematic diagram of an exemplary video decoding device.
[0037] Figure 9 A flowchart of an exemplary method for encoding a bitstream of an image, which extracts a sub-bitstream including a sub-image without rewriting the slice header through explicit address indication.
[0038] Figure 10 A flowchart of an exemplary method for decoding a sub-bitstream including a sub-image through explicit address indication.
[0039] Figure 11 A schematic diagram of an exemplary system for transmitting a sub-bitstream including a sub-image through explicit address indication. DETAILED DESCRIPTION
[0040] First, it should be understood that although the following provides illustrative implementations of one or more embodiments, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0041] This document uses various abbreviations, such as Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Draft (WD).
[0042] Many video compression techniques can be used to reduce video files while minimizing data loss. For example, video compression techniques can include performing spatial (e.g., intra-frame) prediction and / or temporal (e.g., inter-frame) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video strip (e.g., a video image or a portion of a video image) can be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Intra-coded (I) strips of video blocks within an image are coded using spatial prediction of reference samples in adjacent blocks within the same image. Inter-coded unidirectional prediction (P) or bidirectional prediction (B) strips of video blocks within an image can be coded using spatial prediction of reference samples in adjacent blocks within the same image or using temporal prediction of reference samples in other reference images. A picture / image may be referred to as a frame, and a reference image may be referred to as a reference frame. Spatial or temporal prediction produces a predicted block representing an image block. Residual data represents the pixel difference between the original image block and the predicted block. Accordingly, inter-coded blocks are coded based on a motion vector and residual data, where the motion vector points to the block of reference samples that make up the predicted block, and the residual data indicates the difference between the coded block and the predicted block. Intra-coded blocks are coded based on an intra-coding mode and residual data. For further compression, the residual data can be transformed from the pixel domain to a transform domain, thereby producing residual transform coefficients that can then be quantized. The quantized transform coefficients are initially arranged in a two-dimensional array. The quantized transform coefficients can be scanned to produce a one-dimensional vector of transform coefficients. Entropy coding can be applied to achieve further compression. Such video compression techniques are discussed in more detail below.
[0043] To ensure that the encoded video can be correctly decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extended versions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three dimensional (3D) AVC (3D-AVC). HEVC includes extended versions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC has started to develop a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD), and the WD includes JVET-L1001-v7.
[0044] To encode a video image, the image is first segmented, and then the resulting segmented parts are encoded in a bitstream. There are various image segmentation schemes. For example, an image can be segmented into regular slices, non-independent slices, tiles, and / or segmented according to Wavefront Parallel Processing (WPP). For simplicity, HEVC places restrictions on the encoder such that when slicing the image into Coding Tree Block (CTB) groups for video decoding, only regular slices, non-independent slices, tiles, WPP, and their combinations can be used. This segmentation enables maximum transfer unit (MTU) size matching, parallel processing, and reduced end-to-end latency. The MTU represents the maximum amount of data that can be sent in a single data packet. If the data packet payload exceeds the MTU, the payload is divided into two data packets through a process called fragmentation.
[0045] A regular slice, also simply referred to as a slice, is a part obtained after segmenting an image and can be reconstructed independently of other regular slices within the same image. However, due to the existence of loop filtering operations, there are still dependencies. Each regular slice is encapsulated in its own network abstraction layer (NAL) unit for transmission. Additionally, intra-frame prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries can be disabled to support independent reconstruction. This independent reconstruction enables parallelization operations. For example, parallelization based on regular slices reduces inter-processor or inter-core communication. However, since regular slices are all independent, each slice is associated with a separate slice header. Since each slice has a slice header bit cost and lacks prediction across slice boundaries, using regular slices incurs a large amount of coding overhead. Additionally, regular slices can be used to support MTU size matching requirements. Specifically, since regular slices are encapsulated in separate NAL units and can be encoded independently, each regular slice needs to be less than the MTU in the MTU size matching scheme to avoid splitting the slice into multiple data packets. Therefore, to achieve parallelization and MTU size matching, the slice layout within the image is conflicting.
[0046] A non-independent slice is similar to a regular slice but has a shortened slice header and can segment the image tree block boundary without breaking intra-frame prediction. Correspondingly, a non-independent slice can divide a regular slice into multiple NAL units, such that the entire regular slice is encoded first and then a part of the regular slice is sent out, thereby reducing end-to-end latency.
[0047] A tile is a segmented part within an image formed by horizontal and vertical boundaries that form tile columns and tile rows. Tiles can be encoded in raster scan order (from right to left, top to bottom). The scan order of CTBs is the order in which scanning is performed within a tile. Accordingly, the CTBs within the first tile are encoded in raster scan order first, and then the CTBs within the next tile are processed. Similar to a regular strip, a tile breaks the dependency on intra-frame prediction and entropy decoding. However, a tile may not be included in a single NAL unit, and thus, tiles cannot be used to achieve MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / core communication used for intra-frame prediction between processing units decoding adjacent tiles can be limited to sending the shared strip header (when adjacent tiles are within the same strip) and sharing the reconstructed samples and metadata related to loop filtering. When a strip includes multiple tiles, the byte offset of the entry point for each tile can be indicated in the strip header in addition to the entry point offset of the first entry point in the strip. For each strip and tile, at least one of the following conditions needs to be satisfied: (1) all coding tree blocks in the strip belong to the same tile; (2) all coding tree blocks in the tile belong to the same strip.
[0048] In WPP, the image is segmented into single-row CTBs. The entropy decoding and prediction mechanisms can use the data of CTBs in other rows. Parallel processing can be achieved through parallel decoding of CTB rows. For example, the current row can be decoded in parallel with the previous row. However, the decoding of the current row is delayed by two CTBs compared to the decoding processes of the previous rows. This delay ensures that the data related to the CTBs above and to the right of the current CTB located in the current row is available before decoding the current CTB. When represented graphically, this method looks like a wavefront. This staggered start of decoding can be parallelized using as many processors / cores as there are CTB rows included in the image. Since intra-frame prediction between adjacent tree block rows within the image is supported, a large amount of inter-processor / core communication may be required to achieve intra-frame prediction. The WPP segmentation does not consider the NAL unit size. Therefore, WPP does not support MTU size matching. However, regular strips can be used in combination with WPP, incurring a certain coding overhead to achieve MTU size matching as needed.
[0049] The tile may also include a motion constrained tile set. A motion constrained tile set (MCTS) is a tile set such that the associated motion vectors are restricted to point to integer pixel positions inside the MCTS and to fractional pixel positions that only require integer pixel positions inside the MCTS for interpolation. In addition, motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS are not allowed. Thus, each MCTS can be decoded independently without the tiles included in the MCTS. A temporal MCTS supplemental enhancement information (SEI) message may be used to indicate the presence of the MCTS in the bitstream and to signal the MCTS. The MCTS SEI message provides supplemental information (detailed as part of the SEI message semantics) that can be used to extract the MCTS sub-bitstream to generate a compliant bitstream for the MCTS set. This information includes the number of extraction information sets, each extraction information set defining multiple MCTS sets and including the raw bytes sequence payload (RBSP) bytes of the video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) to be used in the MCTS sub-bitstream extraction process. Since one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) may take different values in the extracted sub-bitstream, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced and the slice header updated when extracting the sub-bitstream according to the MCTS sub-bitstream extraction process.
[0050] The above - mentioned solution may have some problems. In some systems, when there are multiple tiles / strips in an image, a syntax element (such as tile_group_address) can be used to indicate the address of a tile group as an index in the tile group header. tile_group_address represents the tile address of the first tile in the tile group. The length of tile_group_address can be determined to be Ceil(Log2(NumTilesInPic)) bits, where NumTilesInPic includes the number of tiles in the image. The value range of tile_group_address can be from 0 to NumTilesInPic–1 (including the end values), and the value of tile_group_address can be not equal to the value of tile_group_address of any other encoded tile group NAL unit in the same encoded image. When tile_group_address does not exist in the bitstream, it can be inferred that tile_group_address is 0. The tile address described above includes the tile index. However, using the tile index as the address of each tile group may reduce the coding efficiency to some extent.
[0051] For example, directly on the client side or in some network - based media processing entity before sending a sub - bitstream to a decoder, some situations may require modifying the AVC or HEVC slice segment header between encoding and decoding. An example of such a situation is tile - based streaming. In tile - based streaming, panoramic videos are encoded using HEVC tiles, but the decoder only decodes a subset of these tiles. By rewriting the HEVC slice segment header (SSH) and SPS / PPS, the bitstream can be controlled to change the subset of tiles being decoded and their spatial arrangement in the decoded video frame. One reason for the CPU processing overhead is that the AVC and HEVC slice segment headers use variable - length fields and have a byte - alignment field at the end. That is, whenever a field in the SSH is changed, it affects the byte - alignment field at the end of the SSH, and then the byte - alignment field has to be rewritten. Moreover, since all fields are variable - length encoded, the only way to know the position of the byte - alignment field is to parse all the previous fields. This results in a large amount of processing overhead, especially when using tiles, as the video per second may include hundreds of NALs. Some systems support explicit indication of tile identifiers (IDs). However, some syntax elements may not be optimized and may include unnecessary bits and / or redundant bits when indicating. In addition, some constraints related to the explicit tile ID indication are not specified.
[0052] For example, the above mechanism allows for the segmentation and compression of images. For example, an image can be segmented into strips, blocks, and / or groups of blocks. In some examples, groups of blocks can be used interchangeably with strips. Such strips and / or groups of blocks can be addressed according to a set of indices. These indices can start from index 0 at the upper left corner of the image and increment in raster scan order, ending at index N at the lower right corner of the image. In this case, N is the number of indices minus 1. Such a system is suitable for most applications. However, certain applications, such as virtual reality (VR), only display sub-images within an image. Such sub-images can be referred to as regions of interest in some contexts. Some systems only send a sub-bitstream in the bitstream to the decoder when streaming VR content, which can improve the encoding efficiency, where the sub-bitstream includes the sub-image to be displayed. In this case, since the upper left corner of the sub-image received by the decoder is usually an index other than 0, the index-based addressing scheme may stop working properly. To solve these problems, it may be necessary for the encoder (or related stripper) to rewrite each strip header to change the index of the sub-image so that the upper left corner index starts from 0 and the remaining sub-image strips are adjusted accordingly. Dynamically rewriting the strip header (e.g., according to each user request) may require a lot of calculations by the processor.
[0053] This disclosure presents various mechanisms to improve encoding efficiency and reduce processing overhead when extracting a sub-bitstream including a sub-image from an encoded bitstream including an image. The disclosed system employs an addressing scheme that can extract a sub-bitstream including a sub-image without rewriting the slice header. Each slice / tile group is addressed based on an ID other than an index. For example, a slice can be addressed by a value that can be mapped to an index and stored in the slice header. In this way, the decoder reads the slice address from the slice header and maps the address from an image-based location to a sub-image-based location. Since the slice address is not a predefined index, the slice address is encoded in a variable-length field. Accordingly, the slice address length is also indicated. The ID associated with the sub-image is also indicated. The sub-image ID and length can be indicated in the PPS. A flag can also be indicated in the PPS to indicate that an explicit addressing scheme is employed. After reading the flag, the decoder can obtain the length and the sub-image ID. The length is used to parse the slice address from the slice header. The sub-image ID is used to map the slice address from an image-based location to a sub-image-based location. In this way, the decoder can consistently determine all relevant addresses regardless of which sub-image is received and regardless of the position of the received sub-image relative to the upper left corner of the complete image. Additionally, this mechanism can make these decisions without rewriting the slice header to change the slice address value and / or without changing the byte alignment field associated with the slice address. By adopting the above mechanism, the encoder, decoder, and / or related slicer can be improved. For example, instead of extracting and sending the entire bitstream, a sub-bitstream can be extracted and sent, which can reduce the use of network resources, memory resources, and / or processing resources. Moreover, extracting the sub-bitstream in this way can avoid rewriting each slice header according to each user request, which further reduces the use of network resources, memory resources, and / or processing resources.
[0054] Figure 1 Flowchart of an exemplary operation method 100 for encoding a video signal. Specifically, the encoder encodes the video signal. During the encoding process, various mechanisms are employed to compress the video signal to reduce the video file. The file being smaller, the compressed video file can be sent to the user while reducing the associated bandwidth overhead. Then, the decoder decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is typically the inverse of the encoding process such that the video signal reconstructed by the decoder is consistent with the video signal on the encoder side.
[0055] In step 101, a video signal is input into an encoder. For example, the video signal can be an uncompressed video file stored in a memory. As another example, the video file can be captured by a video capture device (such as a camera) and encoded to support live streaming of the video. The video file can include an audio component and a video component. The video component includes a series of image frames. When these image frames are viewed in sequence, they give a visual effect of motion. These frames include pixels represented by light, referred to herein as the luminance component (or luminance samples), and also include pixels represented by color, referred to as the chrominance component (or color samples). In some examples, these frames can also include depth values to support three-dimensional viewing.
[0056] In step 103, the video is segmented into blocks. The segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frame can first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). These CTUs include luminance samples and chrominance samples. The CTUs can be divided into blocks using a coding tree, and then these blocks can be repeatedly subdivided until a configuration is obtained that supports further encoding. For example, the luminance component of the frame can be subdivided until each block includes relatively uniform luminance values. Additionally, the chrominance component of the frame can be subdivided until each block includes relatively uniform color values. Thus, the segmentation mechanism varies according to the content of the video frame.
[0057] In step 105, various compression mechanisms are employed to compress the image blocks segmented in step 103. For example, inter-frame prediction and / or intra-frame prediction can be employed. Inter-frame prediction takes advantage of the fact that objects in common scenarios often appear in consecutive frames. Therefore, blocks representing objects in a reference frame do not need to be repeatedly represented in adjacent frames. Specifically, an object (such as a table) may remain in a fixed position in multiple frames. Therefore, the table is represented once, and adjacent frames can refer back to this reference frame. A pattern matching mechanism can be employed to match objects in multiple frames. Additionally, due to reasons such as object movement or camera movement, moving objects can be represented in multiple frames. In a particular example, the video can show a car moving across the screen in multiple frames. Motion vectors can be used to represent such movement. A motion vector is a two-dimensional vector that provides the offset between the coordinates of an object in one frame and the coordinates of the same object in a reference frame. Thus, inter-frame prediction can encode an image block in the current frame as a set of motion vectors that indicate the offset between the image block in the current frame and the corresponding block in the reference frame.
[0058] Intra prediction is used to encode blocks in a common frame. Intra prediction takes advantage of the fact that luminance and chrominance components tend to cluster in a frame. For example, a patch of green in a part of a tree often has several adjacent patches of similar green. Intra prediction employs multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. These directional modes indicate that the samples of the current block are similar / identical to the samples of the adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks in a row / column (e.g., a plane) can be interpolated based on the adjacent blocks at the edge of that row. The planar mode actually represents a smooth transition of light / color between rows / columns with a relatively constant slope with numerical variations. The DC mode is used for boundary smoothing and indicates that the block is similar / identical to the average of the samples of all adjacent blocks that are related to the angular directions of the directional prediction modes. Thus, an intra prediction block can represent an image block as various relationship prediction mode values rather than actual values. Additionally, an inter prediction block can represent an image block as motion vector values rather than actual values. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in a residual block. The residual block can be transformed to further compress the file.
[0059] In step 107, various filtering techniques can be used. In HEVC, filters are used according to an in-loop filtering scheme. The block-based prediction described above may produce a blocky image at the decoder side. Additionally, a block-based prediction scheme can encode a block and then reconstruct the encoded block for subsequent use as a reference block. The in-loop filtering scheme iteratively applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce block artifacts so that the encoded file can be accurately reconstructed. Additionally, these filters reduce the artifacts in the reconstructed reference blocks, making it less likely that other artifacts will be generated in subsequent blocks encoded based on the reconstructed reference blocks.
[0060] Once the video signal is segmented, compressed, and filtered, in step 109, the resulting data is encoded in a bitstream. The bitstream includes the data described above and any indication data needed to support proper video signal reconstruction at the decoder side. For example, this data can include segmentation part data, prediction data, residual blocks, and various flags that provide encoding instructions to the decoder. The bitstream can be stored in a memory for transmission to the decoder upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. The generation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 can be performed continuously and / or simultaneously over multiple frames and blocks.Figure 1 The order shown is for purposes of clarity and ease of discussion and is not intended to limit the video decoding process to a particular order.
[0061] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data in the bitstream to determine the segmented parts of the frame. The segmentation should match the result of the block segmentation in step 103. The entropy encoding / decoding employed in step 111 is described below. The encoder makes many choices during the compression process, such as selecting a block segmentation scheme from several possible choices based on the spatial location of values in one or more input images. Indicating the exact choices may use a large number of binary symbols (bins). As used herein, a "binary symbol" is a binary value that serves as a variable (e.g., a bit value that may vary depending on the context). Entropy encoding enables the encoder to discard any options that are clearly not suitable for a particular situation, leaving a set of available options. Then, a codeword is assigned to each available option. The length of the codeword depends on the number of available options (e.g., one binary symbol corresponds to two options, two binary symbols correspond to three to four options, and so on). Then, the encoder encodes the codewords of the selected options. This scheme reduces the codewords because the codewords are as large as necessary to uniquely indicate a selection from a small subset of available options rather than uniquely indicating a selection from a potentially large set of all possible options. Then, the decoder decodes the options by determining the set of available options in a manner similar to the encoder. By determining the set of available options, the decoder can read the codewords and determine the choices made by the encoder.
[0062] In step 113, the decoder performs block decoding. Specifically, the decoder performs an inverse transform to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct image blocks according to the segmentation. The prediction blocks may include intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. Then, the reconstructed image blocks are placed in the frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax for step 113 may also be indicated in the bitstream by the entropy encoding described above.
[0063] In step 115, the frame of the reconstructed video signal is filtered in a manner similar to step 107 on the encoder side. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove block artifacts. Once the frame has been filtered, then in step 117, the video signal may be output to a display for viewing by the end user.
[0064] Figure 2Schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the functions of the codec system 200 can implement the operation method 100. The codec system 200 is generally applicable to describe components used in both encoders and decoders. The codec system 200 receives a video signal and segments the video signal to obtain a segmented video signal 201, as described in conjunction with steps 101 and 103 in the operation method 100. When acting as an encoder, the codec system 200 compresses the segmented video signal 201 into an encoded bitstream, as described in conjunction with steps 105, 107, and 109 in the method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in conjunction with steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general decoder control component 211, a transform scaling and quantization component 213, an intra-frame estimation component 215, an intra-frame prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In Figure 2 it, the black lines represent the motion of the data to be encoded / decoded, and the dashed lines represent the motion of the control data that controls the operation of other components. All components in the codec system 200 can be present in the encoder. The decoder can include a subset of the components in the codec system 200. For example, the decoder can include an intra-frame prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components are described below.
[0065] The split video signal 201 is split into a captured video sequence, which has been split into pixel blocks by an encoding tree. The encoding tree uses various partitioning modes to subdivide the pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into smaller blocks. These blocks can be referred to as nodes on the encoding tree. Larger parent nodes are partitioned into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / encoding tree. In some cases, the partitioned blocks can be included in a coding unit (CU). For example, a CU can be a sub-part of a CTU, including a luminance block, one or more chrominance red difference (Cr) blocks, and one or more chrominance blue difference (Cb) blocks, as well as syntax instructions corresponding to the CU. The partitioning modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT) for splitting a node into two, three, or four child nodes of different shapes (depending on the partitioning mode used). The split video signal 201 is sent to an overall encoder control component 211, a transform scaling and quantization component 213, an intra prediction component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0066] The overall decoder control component 211 is used to determine, according to application constraints, which images in the video sequence are encoded in the bitstream. For example, the overall decoder control component 211 manages the optimization of the bitrate / bitstream size and the reconstruction quality. These decisions can be made based on storage space / bandwidth availability and image resolution requests. The overall decoder control component 211 also manages the buffer utilization according to the transmission speed to alleviate buffer underflow and overflow. To address these issues, the overall decoder control component 211 manages the splitting, prediction, and filtering performed by other components. For example, the overall decoder control component 211 can dynamically increase the compression complexity to improve the resolution and bandwidth utilization, or decrease the compression complexity to reduce the resolution and bandwidth utilization. Therefore, the overall decoder control component 211 controls other components in the codec system 200, balancing the video signal reconstruction quality and the bitrate. The overall decoder control component 211 generates control data, which is used to control the operations of other components. The control data is also sent to a header format and CABAC component 231 for encoding in the bitstream to indicate the parameters used by the decoder for decoding.
[0067] The split video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. The frames or strips in the split video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame predictive coding on the received video blocks with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple decoding rounds to select a suitable coding mode for each block in the video data, and so on.
[0068] The motion estimation component 221 and the motion compensation component 219 can be highly integrated, but are described separately for conceptual purposes. Motion estimation performed by the motion estimation component 221 is a process of generating motion vectors, where these motion vectors are used to estimate the motion of video blocks. For example, a motion vector can represent the displacement of an encoded object relative to a predicted block. A predicted block is a block that highly matches the block to be encoded in terms of pixel difference. A predicted block can also be referred to as a reference block. Such pixel difference can be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. HEVC employs several encoded objects, including CTU, coding tree block (CTB), and CU. For example, a CTU can be divided into CTBs, and then the CTBs can be divided into CUs to be included in CUs. A CU can be encoded as a prediction unit (PU) including prediction data and / or a transform unit (TU) including transform residual data of the CU. The motion estimation component 221 performs rate-distortion analysis as part of a rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and can select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics can maintain a balance between the quality of video reconstruction (e.g., the amount of data loss caused by compression) and the coding efficiency (e.g., the final encoded size).
[0069] In some examples, the codec system 200 may calculate values of sub-integer pixel positions of a reference image stored in the decoded image buffer component 223. For example, the video codec system 200 may interpolate at a quarter-pixel position, an eighth-pixel position, or other fractional pixel positions of the reference image. Thus, the motion estimation component 221 may perform a motion search relative to integer pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy. The motion estimation component 221 compares the position of the PU with the position of a predicted block in the reference image to calculate a motion vector for the PU of a video block in an inter-coded strip. The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and CABAC component 231 for encoding and as motion data to the motion compensation component 219.
[0070] The motion compensation performed by the motion compensation component 219 may include obtaining or generating a predicted block according to the motion vector determined by the motion estimation component 221. Similarly, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vector of the PU of the current video block, the motion compensation component 219 may locate the predicted block pointed to by the motion vector. Then, the pixel values of the predicted block are subtracted from the pixel values of the current video block being encoded to obtain a pixel difference, thereby forming a residual video block. Generally, the motion estimation component 221 performs motion estimation with respect to the luminance component, and the motion compensation component 219 uses the motion vector calculated according to the luminance component for both the chrominance component and the luminance component. The predicted block and the residual block are sent to the transform scaling and quantization component 213.
[0071] The segmented video signal 201 is also sent to the intra estimation component 215 and the intra prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra estimation component 215 and the intra prediction component 217 may be highly integrated, but are described separately for conceptual purposes. The intra estimation component 215 and the intra prediction component 217 perform intra prediction on the current block with respect to each block in the current frame to replace the inter prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra estimation component 215 determines an intra prediction mode to encode the current block. In some examples, the intra estimation component 215 selects a suitable intra prediction mode from multiple tested intra prediction modes to encode the current block. Then, the selected intra prediction mode is sent to the header format and CABAC component 231 for encoding.
[0072] For example, the intra prediction component 215 performs rate-distortion analysis on various tested intra prediction modes, calculates rate-distortion values, and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis is generally used to determine the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to generate the encoded block, and to determine the bit rate (e.g., the number of bits) used to generate the encoded block. The intra prediction component 215 calculates a ratio based on the distortion and rate of various encoded blocks to determine the intra prediction mode that yields the best rate-distortion value for the block. Additionally, the intra prediction component 215 can be used to encode depth blocks in a depth image using a depth modeling mode (DMM) according to rate-distortion optimization (RDO).
[0073] When implemented on the encoder, the intra prediction component 217 can generate a residual block based on the predicted block according to the intra prediction mode determined by the intra prediction component 215, or when implemented on the decoder, read the residual block from the bitstream. The residual block includes the difference between the predicted block and the original block, represented as a matrix. Then, the residual block is sent to the transform scaling and quantization component 213. The intra prediction component 215 and the intra prediction component 217 can operate on the luminance component and the chrominance component.
[0074] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 performs a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform on the residual block, thereby generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be performed. The transform can convert the residual information from the pixel value domain to the transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also used to scale the transform residual information according to frequency and the like. This scaling involves applying a scaling factor to the residual information in order to quantize different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 can then perform a scan on the matrix including the quantized transform coefficients. The quantized transform coefficients are sent to the header format and CABAC component 231 for encoding in the bitstream.
[0075] The scaling and inverse transform component 229 performs operations opposite to those of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 performs inverse scaling, inverse transform, and / or inverse quantization to reconstruct the residual block in the pixel domain, for example, to be subsequently used as a reference block. This reference block may become the predicted block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may add the residual block back to the corresponding predicted block to calculate the reference block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce the artifacts generated during scaling, quantization, and transformation. These artifacts may cause inaccurate prediction (and generate additional artifacts) when predicting subsequent blocks.
[0076] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding predicted block from the intra prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, a filter may be applied to the reconstructed image block. In some examples, a filter may be applied to the residual block. Like Figure 2 the other components in, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are described separately for conceptual purposes. The filters applied to the reconstructed reference block are applied to specific spatial regions, and these filters include multiple parameters to adjust the way these filters are used. The filter control analysis component 227 analyzes the reconstructed reference block to determine the locations where these filters need to be used and sets the corresponding parameters. This data is sent to the header format and CABAC component 231 for encoding as filter control data. The in-loop filter component 225 uses these filters according to the filter control data. These filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. These filters may be applied in the spatial domain / pixel domain (e.g., to reconstruct pixel blocks) or in the frequency domain according to examples.
[0077] When operating as an encoder, the filtered reconstructed image block, residual block, and / or predicted block are stored in the decoded image buffer component 223 for subsequent use in motion estimation as described above. When operating as a decoder, the decoded image buffer component 223 stores the filtered reconstructed block and sends it to the display as part of the output video signal. The decoded image buffer component 223 may be any storage device capable of storing predicted blocks, residual blocks, and / or reconstructed image blocks.
[0078] The header format and CABAC component 231 receive data from various components in the codec system 200 and encode this data in an encoded bitstream for transmission to a decoder. Specifically, the header format and CABAC component 231 generate various headers to encode control data (such as overall control data and filter control data). Additionally, predictive data (including intra prediction and motion data) as well as residual data in the form of quantized transform coefficient data are both encoded in the bitstream. The final bitstream includes all the information needed by the decoder to reconstruct the original segmented video signal 201. This information may also include an intra prediction mode index table (also referred to as a codeword mapping table), definitions of the coding contexts for various blocks, indications of the most probable intra prediction modes, indications of segmentation information, etc. This data can be encoded using entropy coding. For example, context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques can be used to encode the above information. After entropy coding, the encoded bitstream can be sent to another device (such as a video decoder) or archived for subsequent transmission or retrieval.
[0079] Figure 3 FIG. is a block diagram of an exemplary video encoder 300. The video encoder 300 can be used to implement the encoding function of the codec system 200 and / or perform steps 101, 103, 105, 107, and / or 109 in the operation method 100. The encoder 300 segments the input video signal to obtain a segmented video signal 301 that is substantially similar to the segmented video signal 201. Then, the components in the encoder 300 compress and encode the segmented video signal 301 in a bitstream.
[0080] Specifically, the segmented video signal 301 is sent to the intra prediction component 317 for intra prediction. The intra prediction component 317 can be substantially similar to the intra estimation component 215 and the intra prediction component 217. The segmented video signal 301 is also sent to the motion compensation component 321 for inter prediction based on the reference blocks in the decoded picture buffer component 323. The motion compensation component 321 can be substantially similar to the motion estimation component 221 and the motion compensation component 219. The predicted blocks and residual blocks from the intra prediction component 317 and the motion compensation component 321 are sent to the transform and quantization component 313 to transform and quantize the residual blocks. The transform and quantization component 313 can be substantially similar to the transform scaling and quantization component 213. The transform quantized residual blocks and the corresponding predicted blocks (along with the relevant control data) are sent to the entropy coding component 331 to be encoded in the bitstream. The entropy coding component 331 can be substantially similar to the header format and CABAC component 231.
[0081] The transform quantized residual blocks and / or the corresponding predicted blocks are also sent from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed as reference blocks for the motion compensation component 321 to use. The inverse transform and quantization component 329 can be substantially similar to the scaling and inverse transform component 229. According to an example, the in-loop filter in the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 can include multiple filters as described in connection with the in-loop filter component 225. Then, the filtered blocks are stored in the decoded picture buffer component 323 to be used as reference blocks for the motion compensation component 321. The decoded picture buffer component 323 can be substantially similar to the decoded picture buffer component 223.
[0082] Figure 4 It is a block diagram of an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding function of the codec system 200 and / or perform the steps 111, 113, 115 and / or 117 in the operation method 100. The decoder 400 receives the bitstream from the encoder 300 etc. and generates a reconstructed output video signal according to the bitstream to display to the end user.
[0083] The bitstream is received by an entropy decoding component 433. The entropy decoding component 433 is used to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context for parsing other data encoded as codewords in the bitstream. The above decoding information includes any information required to decode the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients in residual blocks. The quantized transform coefficients are sent to an inverse transform and quantization component 429 to be reconstructed into a residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0084] The reconstructed residual block and / or prediction block is sent to an intra prediction component 417 to be reconstructed into an image block according to an intra prediction operation. The intra prediction component 417 may be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses a prediction mode to locate a reference block in the frame and applies the residual block to the above result to reconstruct an intra prediction image block. The above reconstructed intra prediction image block and / or residual block and the corresponding inter prediction data are sent to a decoded image buffer component 423 through an in-loop filter component 425. The decoded image buffer component 423 and the in-loop filter component 425 may be substantially similar to the decoded image buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the above reconstructed image block, residual block, and / or prediction block. This information is stored in the decoded image buffer component 423. The reconstructed image block from the decoded image buffer component 423 is sent to a motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector of the reference block to generate a prediction block and applies the residual block to the above result to reconstruct an image block. The resulting reconstructed block may also be sent to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks. These reconstructed image blocks may be reconstructed into frames through segmentation information. These frames may also be placed in a sequence. The sequence is output as a reconstructed output video signal to a display.
[0085] Figure 5 A schematic diagram of an exemplary bitstream 500 including an encoded video sequence. For example, the bitstream 500 may be generated by the codec system 200 and / or the encoder 300 and decoded by the codec system 200 and / or the decoder 400. Again, for example, the bitstream 500 may be generated by the encoder in step 109 of the method 100 and used by the decoder in step 111.
[0086] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPS) 512, multiple slice headers 514, and picture data 520. The SPS 510 includes sequence data common to all pictures in the video sequence included in the bitstream 500. This data may include picture size, bit depth, coding tool parameters, bitrate limit conditions, etc. The PPS 512 includes one or more picture-specific parameters. Thus, each picture in the video sequence can refer to a PPS 512. The PPS 512 may represent coding tools available for block partitioning in the corresponding picture, quantization parameters, offsets, picture-specific coding tool parameters (such as filter control parameters), etc. The slice header 514 includes parameters specific to one or more corresponding slices in the picture. Thus, each slice in the video sequence can refer to the slice header 514. The slice header 514 may include slice type information, picture order count (POC), reference picture list, prediction weights, block entry points, deblocking filter parameters, etc. In some examples, a slice may be referred to as a block group. In this case, the slice header 514 may be referred to as a block group header.
[0087] The picture data 520 includes video data encoded according to inter-frame prediction and / or intra-frame prediction, as well as corresponding transform and quantization residual data. These picture data 520 are sorted according to the segmentation mode of the picture before encoding. For example, the video sequence is divided into pictures 521, the pictures 521 are divided into slices 523, the slices 523 can be further divided into blocks and / or CTUs, the CTUs are further divided into coding blocks according to the coding tree, and then, the coding blocks can be encoded / decoded according to the prediction mechanism. For example, a picture 521 may include one or more slices 523. The picture 521 refers to the PPS 512, and the slice 523 refers to the slice header 514. Each slice 523 may include one or more blocks. Each slice 523 and / or picture 521 may then include multiple CTUs.
[0088] Each picture 521 may include the entire visual data set related to the video sequence at the corresponding moment. The VR system can display the area selected by the user in the picture 521. In this way, a feeling of being in the scene depicted in the picture 521 is generated. When the bitstream 500 is encoded, the area that the user may wish to view is unknown. Therefore, the picture 521 may include every possible area that the user is likely to view. However, in the VR scenario, the corresponding codec can be designed based on the assumption that the user only views the area selected in the picture 521 and discards the rest of the picture 521.
[0089] Each strip 523 can be a rectangle defined by the top-left CTU and the bottom-right CTU. In some examples, strip 523 includes a series of blocks and / or CTUs that are raster scanned in a left-to-right, top-to-bottom order. In other examples, strip 523 is a rectangular strip. The rectangular strip may not traverse the entire width of the image in raster scan order. Instead, the rectangular strip can include rectangular and / or square regions in image 521 defined according to CTU and / or block rows and CTU and / or block columns. Strip 523 is the smallest unit that can be separately displayed by the decoder. Thus, strips 523 in image 521 can be assigned to different sub-images 522 to depict desired regions of image 521 respectively. For example, in a VR scenario, image 521 can include the entire visible data range, but the user can only view sub-image 522 including one or more strips 523 on a head-mounted display.
[0090] As described above, the video codec can assume that non-selected regions of image 521 are to be discarded on the decoder side. Thus, sub-bitstream 501 can be extracted from bitstream 500. The extracted sub-bitstream 501 can include the selected sub-image 522 and related syntax. Non-selected regions of image 521 can be sent at a lower resolution or ignored to improve encoding efficiency. Sub-image 522 is the selected region of image 521 and can include one or more related strips 524. Strip 524 is a subset of strip 523 and depicts the selected region of image 521 related to sub-image 522. Sub-bitstream 501 also includes SPS 510, PPS 512, strip header 514, and / or sub-parts thereof related to sub-image 522 and strip 524.
[0091] Sub-bitstream 501 can be extracted from bitstream 500. For example, a user using the decoder can watch a video. The user can select the corresponding region of image 521. The decoder can request subsequent sub-images 522 related to the region the user is currently viewing. Then, the encoder can send the sub-image 522 related to the selected region at a higher resolution and the remaining region of image 521 at a lower resolution. To implement this functionality, the decoder can extract 529 one or more sub-bitstreams 501 from bitstream 500. Extraction 529 includes storing sub-image 522 (including strips 524 in sub-image 522) in sub-bitstream 501. Extraction 529 also includes storing related SPS 510, PPS 512, and strip header 514 in the sub-bitstream as needed to support decoding of sub-image 522 and strip 524.
[0092] One problem with sub - bitstream 501 extraction 529 is that the addressing relative to image 521 may be different from the addressing relative to sub - image 522. The addressing problem is described in detail below. In some systems, the strip header 514 may be rewritten to adjust for these addressing differences. However, the sub - bitstream 501 may include many strip headers 514 (e.g., on the order of one or two per image 521), and these strip headers 514 are rewritten dynamically for each user. Thus, rewriting the strip headers 514 in this way may require a lot of computation by the processor. The present invention includes a mechanism that can extract 529 the strip headers 514 into the sub - bitstream 501 without rewriting the strip headers 514.
[0093] In systems that rewrite the strip header 514, strips 523 and 524 are addressed according to index values such as strip index, block index, CTU index, etc. The values of these indexes increase in raster scan order. To solve the addressing mismatch problem, the disclosed embodiments use ID values defined for each strip, block, and / or CTU. These defined IDs can be default values and / or can be selected by the encoder. The defined IDs can increase uniformly in raster scan order, but these defined IDs may not be monotonically increasing. Thus, there may be gaps in the values of the defined IDs to enable address management. For example, the index may increase monotonically (e.g., 0, 1, 2, 3, etc.), while the defined ID may increase by a defined multiple (e.g., 0, 10, 20, 30, etc.). The encoder can carry the mapping relationship 535 in the bitstream 500 and the sub - bitstream 501 so that the decoder can map the defined ID to an index that the decoder can parse.
[0094] Parameter sets such as SPS 510 and / or PPS 512 may include an ID flag 531. The ID flag 531 can be set to indicate that the mapping relationship 535 is available to map the strip address from the position based on the image 521 to the position based on the sub - image 522. Accordingly, the ID flag 531 can be set to indicate to the decoder the disclosed mechanism used in the bitstream 500 and the sub - bitstream 501. For example, the ID flag 531 can be encoded as an explicit block ID flag, sps_subpic_id_present_flag, or other syntax element. The ID flag 531 can be encoded in the bitstream 500 and extracted 529 into the sub - bitstream 501.
[0095] Parameter sets such as SPS 510 and / or PPS 512 may also include a syntax element ID 532. The ID 532 may indicate a sub-image 522 within an image 521. For example, some ID 532s may be included in the PPS 512 of the bitstream 500. When extracting 529 the sub-bitstream 501, one or more ID 532s associated with one or more sub-images 522 to be sent to the decoder may be included in the PPS 512 of the sub-bitstream 501. In other examples, a flag (point) pointing to the relevant ID 532 may be inserted into the PPS 512 of the sub-bitstream 501 so that the decoder can determine the correct ID 532. For example, the ID 532 may be encoded as SubPicIdx, Tile_id_val[i], or other syntax elements indicating the boundaries of the sub-image 522.
[0096] Parameter sets such as SPS 510 and / or PPS 512 may also include a syntax element strip address length 533. Additionally, the strip header 514 may include a strip address 534 for the strip 523. The strip address 534 is included as a defined ID value. The strip address 534 may be directly extracted 529 into the strip header 514 of the sub-bitstream 501 without modification to avoid overwriting the strip header 514. For example, the strip address 534 may be encoded as slice_address, tile_group_address, or other syntax elements indicating the boundaries of strip 523 and strip 524. The strip address length 533 may then be used to parse the strip address 534. For example, the strip address 534 includes an encoder-defined value and is thus encoded as a variable-length value before a byte-aligned field. The strip address length 533 may indicate the number of bits included in the corresponding strip address 534 and can thus indicate the boundaries of the strip address 534 to the decoder. Thus, the decoder may use the strip address length 533 (e.g., from the PPS 512) to parse the strip address 534. Thus, the strip header 514 does not need to be overwritten to adjust the byte-aligned field after the strip address 534. For example, the strip address length 533 may be encoded as subpic_id_len_minus1, tile_id_len_minus1, or other syntax elements indicating the strip address length 533. The strip address length 533 may be included in the PPS 512 of the bitstream 500 and then extracted 529 into the PPS 512 of the sub-bitstream 501.
[0097] The mapping relationship 535 can also be sent via parameter sets such as the SPS 510, PPS 512, etc. and / or the strip header 514. The mapping relationship 535 represents a mechanism for mapping the strip address from the position based on the image 521 to the position based on the sub-image 522. The mapping relationship 535 can be encoded in the bitstream 500 and extracted 529 into the corresponding parameter set of the sub-bitstream 501. For example, the mapping relationship 535 can be encoded as the syntax element SliceSubpicToPicIdx[SubPicIdx][slice_address], the syntax element tileIdToIdx[Tile_group_address], or other syntax elements representing a mechanism for mapping the strip address from the position based on the image 521 to the position based on the sub-image 522.
[0098] Accordingly, the decoder can read the sub-bitstream 501 and obtain the ID flag 531 to determine that the strip 524 is addressed by the defined address instead of the index. The decoder can obtain the ID 532 to determine the sub-image 522 included in the sub-bitstream 501. The decoder can also obtain one or more strip addresses 534 and the strip address length 533 to parse one or more strip addresses 534. Then, the decoder can obtain the mapping relationship 53 to map one or more strip addresses 534 into a format that the decoder can parse. Finally, the decoder can use one or more strip addresses 534 when decoding and displaying the sub-image 522 and the corresponding strip 524.
[0099] Figure 6 Schematic diagram for splitting the exemplary image 600 for encoding. For example, the image 600 can be encoded in the bitstream 500 by the codec system 200, the encoder 300, and / or the decoder 400, etc. and decoded from the bitstream 500. In addition, the image 600 can be split into and / or included in the sub-images of the sub-bitstream 501 to implement the encoding and decoding in the method 100.
[0100] The image 600 can be split into strips 623, and the strips 623 can be substantially similar to the strips 523. The strips 623 can be further split into blocks 625 and CTUs 627. In Figure 6 it, the strips 623 are represented by thick lines, and each strip 623 is graphically distinguished by an alternating white background and hash. The blocks 625 are represented by dashed lines. The block 625 boundaries located on the strip 623 boundaries are shown as thick dashed lines, while the block 625 boundaries not located on the strip 623 boundaries are shown as thin dashed lines. The CTU 627 boundaries are shown as thin solid lines, except for the CTU 627 boundaries covered by the block 625 boundaries or the strip 623 boundaries. In this example, the image 600 includes 9 strips 623, 24 blocks 625, and 216 CTUs 627.
[0101] As shown in the figure, stripe 623 is a rectangle with boundaries that can be defined by the included tiles 625 and / or CTUs 627. Stripe 623 may not span the entire width of image 600. Tiles 625 may be generated in stripe 623 according to rows and columns. CTUs 627 may be split out from tiles 625 and / or stripe 623 to produce segmented portions of image 600, and these segmented portions may be subdivided into coding blocks for encoding according to inter-frame prediction and / or intra-frame prediction. Image 600 may be encoded in a bitstream such as bitstream 500. Regions of image 600 may be included in a sub-image and extracted into sub-bitstreams such as sub-image 522 and sub-bitstream 501.
[0102] Figure 7 It is a schematic diagram of an exemplary sub-image 722 extracted from image 700. For example, image 700 may be substantially similar to image 600. In addition, image 700 may be encoded in bitstream 500 by an encoding and decoding system 200 and / or an encoder 300, etc. Sub-image 722 may be extracted into sub-bitstream 501 and decoded from sub-bitstream 501 by an encoding and decoding system 200, an encoder 300, and / or a decoder 400, etc. In addition, image 700 may be used to implement encoding and decoding in method 100.
[0103] As shown in the figure, image 700 includes an upper left corner 702 and a lower right corner 704. Sub-image 722 includes one or more stripes 723 in image 700. When using indices, the upper left corner 702 and the lower right corner 704 are respectively associated with the first index and the last index. However, the decoder can only display sub-image 722 and not the entire image 700. In addition, the stripe address 734 of the first stripe 723a may not be aligned with the upper left corner 702, and the stripe address 734 of the third stripe 723c may not be aligned with the lower right corner 704. Therefore, the stripe address 734 relative to sub-image 722 is not aligned with the stripe address 734 relative to image 700. The present invention uses the ID defined for the stripe address 734 instead of an index. The decoder can use a mapping relationship to map the stripe address 734 from the position based on image 700 to the position based on sub-image 722. Then, the decoder can place the first stripe 723a at the upper left corner 702 of the decoder display, place the third stripe 723c at the lower left corner 704 of the decoder display, and place the second stripe 723b between the first stripe 723a and the third stripe 723c using the mapped stripe address 734.
[0104] As described herein, the present invention describes improvements to explicit tile ID indication in video coding, where tiles are used for image segmentation. The above technical description is based on VVC developed by JVET of ITU-T and ISO / IEC. However, these techniques are also applicable to other video coding specifications. The following are exemplary embodiments described herein.
[0105] The concepts of tile index and tile ID can be distinguished. The tile ID of a tile can be equal to or different from the tile index of that tile. When the tile ID is different from the tile index, the mapping relationship between the tile ID and the tile index can be indicated in the PPS. The tile ID can be used to indicate the tile group address in the tile group header instead of using the tile index. In this way, the value of the tile ID can remain the same when extracting the tile group from the original bitstream. This can be achieved by updating the mapping relationship between the tile ID and the tile index in the PPS referenced by the tile group. This method solves the problem that the value of the tile index may change according to the sub-picture to be extracted. It should be noted that when performing MCTS-based sub-bitstream extraction, it may still be necessary to rewrite other parameter sets (such as parameter sets other than the slice header).
[0106] The above can be achieved by using a flag in the parameter set indicating tile information. For example, the PPS can be used as the parameter set. For example, explicit_tile_id_flag can be used for this purpose. Regardless of the number of tiles in the image, explicit_tile_id_flag can be indicated, and explicit_tile_id_flag can indicate that explicit tile indication is used. A syntax element can also be used to represent the number of bits used for tile ID value indication (such as the mapping relationship between the tile index and the tile ID). This syntax element can also be used to indicate the tile ID / address in the tile group header. For example, the syntax element tile_id_len_minus1 can be used for this purpose. When explicit_tile_id_flag is equal to 0 (for example, when the tile ID is set to the tile index), tile_id_len_minus1 may not exist. When tile_id_len_minus1 does not exist, it can be inferred that the value of tile_id_len_minus1 is equal to the value of Ceil(Log2(NumTilesInPic)). Another constraint requires that the bitstream as a result of MCTS sub-bitstream extraction can include explicit_tile_id_flag set to 1 for the active PPS, unless the sub-bitstream includes the top-left tile in the original bitstream.
[0107] In an exemplary embodiment, the video coding syntax may be modified as described below to implement the functions described herein. The exemplary CTB raster and block scanning process may be described as follows. The list TileId[ctbAddrTs] (where ctbAddrTs ranges from 0 to PicSizeInCtbsY–1, inclusive) represents the conversion from the CTB address to the tile ID under block scanning, and the list NumCtusInTile[tileIdx] (where tileIdx ranges from 0 to PicSizeInCtbsY–1, inclusive) represents the conversion from the tile index to the number of CTUs in the tile. The derivations of both are as follows:
[0108]
[0109] The list NumCtusInTile[tileIdx] (where tileIdx ranges from 0 to PicSizeInCtbsY - 1, inclusive) represents the conversion from the tile index to the number of CTUs in the tile and can be derived as follows:
[0110]
[0111] The set TileIdToIdx[tileId] for a set of NumTilesInPic tileId values represents the conversion from the tile ID to the tile index and can be derived as follows:
[0112]
[0113] The exemplary picture parameter set RBSP syntax may be described as follows.
[0114]
[0115]
[0116] The exemplary tile group header syntax may be described as follows.
[0117]
[0118] The exemplary tile group data syntax may be described as follows.
[0119]
[0120]
[0121] The semantics of an exemplary image parameter set RBSP can be described as follows. When the explicit_tile_id_flag is set to 1, it indicates that the tile ID of each tile is explicitly indicated. When the explicit_tile_id_flag is set to 0, it indicates that the tile ID is not explicitly indicated. For a bitstream that is the result of MCTS sub-bitstream extraction, the value of the explicit_tile_id_flag can be set to 1 for activating the PPS, unless the resulting bitstream includes the top-left tile in the original bitstream. Tile_id_len_minus1 + 1 represents the number of bits used to represent the syntax elements tile_id_val[i] and tile_group_address in the tile group header that references the PPS. The value range of tile_id_len_minus1 can be Ceil(Log2(NumTilesInPic)) to 15 (including the end values). When tile_id_len_minus1 does not exist, it can be inferred that the value of tile_id_len_minus1 is equal to Ceil(Log2(NumTilesInPic)). It should be noted that in some cases, the value of tile_id_len_minus1 can be greater than Ceil(Log2(NumTilesInPic)). This is because the current bitstream may be the result of MCTS sub-bitstream extraction. At this time, the tile ID can be the tile index in the original bitstream, and Ceil(Log2(OrgNumTilesInPic)) bits can be used to represent it, where OrgNumTilesInPic is the NumTilesInPic of the original bitstream, which is greater than the NumTilesInPic of the current bitstream. tile_id_val[i] represents the tile ID of the i-th tile in the image that references the PPS. The length of tile_id_val[i] is (tile_id_len_minus1 + 1) bits. For any integers m and n within the range 0 to NumTilesInPic - 1 (including the end values), when m is not equal to n, tile_id_val[m] can be not equal to tile_id_val[n], and when m is less than n, tile_id_val[m] can be less than tile_id_val[n].
[0122] The following variables can be derived by calling the CTB raster and tile scan conversion: The list ColWidth[i] (where i ranges from 0 to num_tile_columns_minus1, inclusive) represents the width of the i-th tile column in CTBs; the list RowHeight[j] (where j ranges from 0 to num_tile_rows_minus1, inclusive) represents the height of the j-th tile row in CTBs; the list ColBd[i] (where i ranges from 0 to num_tile_columns_minus1+1, inclusive) represents the position of the boundary of the i-th tile column in CTBs; the list RowBd[j] (where j ranges from 0 to num_tile_rows_minus1+1, inclusive) represents the position of the boundary of the j-th tile row in CTBs; the list CtbAddrRsToTs[ctbAddrRs] (where ctbAddrRs ranges from 0 to PicSizeInCtbsY-1, inclusive) represents the conversion from the CTB address in the CTB raster scan of the image to the CTB address in the tile scan; the list CtbAddrTsToRs[ctbAddrTs] (where ctbAddrTs ranges from 0 to PicSizeInCtbsY-1, inclusive) represents the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the image; the list TileId[ctbAddrTs] (where ctbAddrTs ranges from 0 to PicSizeInCtbsY-1, inclusive) represents the conversion from the CTB address in the tile scan to the tile ID; the list NumCtusInTile[tileIdx] (where tileIdx ranges from 0 to PicSizeInCtbsY-1, inclusive) represents the conversion from the tile index to the number of CTUs in the tile; the list FirstCtbAddrTs[tileIdx] (where tileIdx ranges from 0 to NumTilesInPic-1, inclusive) represents the conversion from the tile ID to the CTB address of the first CTB in the tile in the tile scan; the set TileIdToIdx[tileId] for a set of NumTilesInPic tileId values represents the conversion from the tile ID to the tile index, and the list FirstCtbAddrTs[tileIdx] (where tileIdx ranges from 0 to NumTilesInPic-1, inclusive) represents the conversion from the tile ID to the CTB address of the first CTB in the tile in the tile scan;The list ColumnWidthInLumaSamples[i] (where i ranges from 0 to num_tile_columns_minus1, inclusive) represents the width of the i-th tile column, in luma samples; the list RowHeightInLumaSamples[j] (where j ranges from 0 to num_tile_rows_minus1, inclusive) represents the height of the j-th tile row, in luma samples.
[0123] tile_group_address represents the tile ID of the first tile in a tile group. tile_group_address is (tile_id_len_minus1 + 1) bits in length. The value range of tile_group_address can be from 0 to 2tile_id_len_minus1+1 - 1 (inclusive), and the value of tile_group_address can be different from the tile_group_address values of any other coded tile group NAL unit within the same coded picture.
[0124] Figure 8 FIG. is a schematic diagram of an exemplary video decoding device 800. The video decoding device 800 is suitable for implementing the disclosed examples / embodiments described herein. The video decoding device 800 includes a downlink port 820, an uplink port 850, and / or a transceiver (Tx / Rx) 810. The transceiver 810 includes a transmitter and / or a receiver for data communication over a network in the uplink and / or downlink. The video decoding device 800 further includes a processor 830 and a memory 832. The processor 830 includes a logic unit and / or a central processing unit (CPU) for processing data. The memory 832 is used to store the data. The video decoding device 800 may further include electrical components, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the uplink port 850 and / or the downlink port 820 for data communication over an electrical communication network, an optical communication network, or a wireless communication network. The video decoding device 800 may further include an input and / or output (I / O) device 860 for data communication with a user. The I / O device 860 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 860 may further include input devices such as a keyboard, a mouse, a trackball, etc. and / or corresponding interfaces for interacting with the above output devices.
[0125] The processor 830 is implemented by hardware and software. The processor 830 can be implemented as one or more CPU chips, one or more cores (e.g., implemented as a multi-core processor), one or more field-programmable gate arrays (FPGAs), one or more application specific integrated circuits (ASICs), and one or more digital signal processors (DSPs). The processor 830 communicates with the downlink port 820, the transceiver unit 810, the uplink port 850, and the memory 832. The processor 830 includes a decoding module 814. The decoding module 814 implements the disclosed embodiments described herein, such as method 100, method 900, and / or method 1000, which may employ the bitstream 500, the image 600, and / or the image 700. The decoding module 814 can also implement any other method / mechanism described herein. In addition, the decoding module 814 can implement the codec system 200, the encoder 300, and / or the decoder 400. For example, when acting as an encoder, the decoding module 814 can indicate flags, sub-image IDs, and lengths in the PPS. The decoding module 814 can also encode the strip address in the strip header. Then, the decoding module 814 can extract a sub-bitstream including the sub-image from the bitstream of the image without rewriting the strip header. When acting as a decoder, the decoding module 814 can read the flags to determine whether to use an explicit strip address instead of an index. The decoding module 814 can also read the length and sub-image ID from the PPS and the strip address from the strip header. Then, the decoding module 814 can use the length to resolve the strip address and use the sub-image ID to map the strip address from an image-based address to a sub-image-based address. Thus, the decoding module 814 can determine the desired position of the strip, regardless of the selected sub-image, and without rewriting the strip header to accommodate changes in the sub-image-based address. Therefore, when segmenting and encoding video data, the decoding module 814 enables the video decoding device 800 to provide other functions, avoid certain processing to reduce processing overhead, and / or improve encoding efficiency. Accordingly, the decoding module 814 improves the functionality of the video decoding device 800 while solving problems specific to the field of video coding. In addition, the decoding module 814 affects the transition of the video decoding device 800 to different states. Alternatively, the encoding module 814 can be implemented as instructions stored in the memory 832 and executed by the processor 830 (e.g., implemented as a computer program product stored in a non-transitory medium).
[0126] The memory 832 includes one or more types of memories, such as magnetic disks, tape drives, solid state drives, read only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), static random-access memory (SRAM), etc. The memory 832 can be used as an overflow data storage device to store these programs when a program is selected for execution and to store the instructions and data read during the execution of the program.
[0127] Figure 9 FIG. 900 is a flowchart of an exemplary method for encoding a bitstream (e.g., bitstream 500) of an image (e.g., image 600) to support extraction of a sub-bitstream (e.g., sub-bitstream 501) including a sub-image (e.g., sub-image 522) without rewriting a slice header by explicit address indication. The method 900 can be performed by an encoder such as the codec system 200, the encoder 300, and / or the video decoding device 800 when executing the method 100.
[0128] The method 900 can begin with: the encoder receiving a video sequence including a plurality of images and determining to encode the video sequence in a bitstream according to a user input or the like. The video sequence is segmented into pictures / frames before encoding for further segmentation. In step 901, the images in the video sequence are encoded in the bitstream. The image can include a plurality of slices, and the plurality of slices includes a first slice. The first slice can be any slice in the image, but for the sake of clear discussion, it is described as the first slice. For example, the upper left corner of the first slice may not be aligned with the upper left corner of the image.
[0129] In step 903, the slice header associated with the slice is encoded in the bitstream. The slice header includes the slice address of the first slice. The slice address can include a defined value, such as a numerical value selected by the encoder. Such a value can be any value, but it can be incremented in raster scan order (e.g., from left to right and top to bottom) to support consistent encoding functions. The slice address may not include an index. In some examples, the slice address can be the syntax element slice_address.
[0130] In step 905, the PPS is encoded in the bitstream. The identifier and the strip address length of the first strip can be encoded in the PPS of the bitstream. The identifier can be a sub-image identifier. The strip address length can represent the number of bits included in the strip address. For example, the strip address length in the PPS can include data sufficient to parse the strip address from the strip header (encoded in step 903). In some examples, the length can be the syntax element subpic_id_len_minus1. Additionally, the identifier can include data sufficient to map the strip address from an image-based location to a sub-image-based location. In some examples, the identifier can be the syntax element subPicIdx. For example, multiple sub-image-based identifiers can be included in the PPS. When extracting a sub-image, the corresponding sub-image ID can be indicated in the PPS by using a flag / pointer and / or removing unused sub-image IDs, etc. In some examples, an explicit ID flag can also be encoded in the parameter set. The flag can indicate to the decoder that a mapping relationship is available to map the strip address from the image-based location to the sub-image-based location. In some examples, the mapping relationship can be the syntax element SliceSubpicToPicIdx[SubPicIdx][slice_address]. Accordingly, the flag can indicate that the strip address is not an index. In some examples, the flag can be the sps_subpic_id_present_flag.
[0131] In step 907, the sub-bitstream in the bitstream is extracted. For example, extracting the sub-bitstream in the bitstream can include: extracting the first strip without rewriting the strip header according to the strip address of the first strip, the strip address length, and the identifier. In a specific example, such extraction can also include extracting the sub-image in the image. At this time, the sub-image includes the first strip. The parameter set can also be included in the sub-bitstream. For example, the sub-bitstream can include the sub-image, the strip header, the PPS, SPS, etc.
[0132] In step 909, the sub-bitstream is stored for sending to the decoder. The sub-bitstream can then be sent to the decoder as needed.
[0133] Figure 10Flowchart of an exemplary method 1000 for decoding a sub-bitstream (e.g., sub-bitstream 501) including a sub-image (e.g., sub-image 522), where the sub-bitstream including the sub-image is extracted from a bitstream (e.g., bitstream 500) of an image (e.g., image 600) by explicit address indication. Method 1000 may be executed by a decoder such as codec system 200, decoder 400, and / or video coding device 800 when performing method 100.
[0134] Method 1000 may begin with the decoder starting to receive a sub-bitstream extracted from the bitstream, e.g., the result of method 900. In step 1001, the sub-bitstream is received. The sub-bitstream includes a sub-image in the image. For example, the bitstream encoded on the encoder side may include an image, and the encoder and / or stripper extracts the sub-bitstream from the bitstream, the sub-bitstream including a sub-image that includes one or more regions of the image in the bitstream. The received sub-image may be segmented into a plurality of strips. The plurality of strips may include a strip designated as the first strip. The first strip may be any strip in the image, but for clarity of discussion, it is described as the first strip. For example, the upper left corner of the first strip may not be aligned with the upper left corner of the image. The sub-bitstream also includes a PPS that describes the syntax related to the image and thus also describes the syntax related to the sub-image. The sub-bitstream also includes a strip header that describes the syntax related to the first strip.
[0135] In step 1003, parameter sets such as PPS and / or SPS are parsed to obtain an explicit ID flag. The ID flag can indicate that the mapping relationship can be used to map the strip address from an image-based location to a sub-image-based location. Correspondingly, the flag can indicate that the corresponding strip address includes a defined value and does not include an index. In some examples, the flag can be sps_subpic_id_present_flag. According to the value of the ID flag, the PPS can be parsed to obtain an identifier and the strip address length of the first strip. The identifier can be a sub-image identifier. The strip address length can indicate the number of bits included in the corresponding strip address. For example, the strip address length in the PPS can include data sufficient to parse the strip address from the strip header. In some examples, the length can be the syntax element subpic_id_len_minus1. Additionally, the identifier can include data sufficient to map the strip address from an image-based location to a sub-image-based location. In some examples, the identifier can be the syntax element subPicIdx. For example, multiple sub-image-based identifiers can be included in the PPS. When extracting sub-images, the corresponding sub-image ID can be indicated in the PPS by using a flag / pointer and / or removing unused sub-image IDs, etc.
[0136] In step 1005, according to the identifier and the strip address length, the strip address of the first strip is determined from the strip header. For example, the length in the PPS can be used to determine the bit boundary to parse the strip address from the strip header. The identifier and the strip address can then be used to map the strip address from an image-based location to a sub-image-based location. For example, the mapping relationship between the image-based location and the sub-image-based location can be used to align the strip header with the sub-image. In this way, the decoder can solve the problem of address mismatch between the strip header and the image addressing scheme caused by the encoder and / or stripper extracting the sub-bitstream without needing to rewrite the strip header. In some examples, the mapping relationship can be the syntax element SliceSubpicToPicIdx[SubPicIdx][slice_address].
[0137] In step 1007, the sub-bitstream is decoded to generate a video sequence of sub-images. The sub-image can include the first strip. Correspondingly, the first strip is also decoded. Then, a video sequence including the sub-image (including the decoded first strip) can be sent for display through a head-mounted display or other display device.
[0138] Figure 11Schematic diagram of an exemplary system 1100 for transmitting a sub-bitstream (e.g., sub-bitstream 501) including a sub-image (e.g., sub-image 522). The sub-bitstream including the sub-image is extracted from a bitstream (e.g., bitstream 500) of an image (e.g., image 600) through explicit address indication. System 1100 can be implemented by encoders and decoders such as codec system 200, encoder 300, decoder 400, and / or video coding device 800. In addition, system 1100 can be used to implement method 100, method 900, and / or method 1000.
[0139] System 1100 includes video encoder 1102. Video encoder 1102 includes an encoding module 1101 for: encoding an image in a bitstream, where the image includes a plurality of slices, and the plurality of slices includes a first slice; encoding a slice header in the bitstream, where the slice header includes the slice address of the first slice; encoding PPS in the bitstream, where the PPS includes an identifier and the slice address length of the first slice. Video encoder 1102 further includes an extraction module 1103 for extracting a sub-bitstream from the bitstream without rewriting the slice header according to the slice address of the first slice, the slice address length, and the identifier. Video encoder 1102 further includes a storage module 1105 for storing the bitstream for transmission to a decoder. Video encoder 1102 further includes a transmission module 1107 for transmitting the sub-bitstream to the decoder, where the sub-bitstream includes the slice header, the PPS, the first slice, and / or the corresponding sub-image. Video encoder 1102 can also be used to perform any step of method 900.
[0140] System 1100 further includes video decoder 1110. Video decoder 1110 includes a receiving module 1111 for receiving a sub-bitstream, where the sub-bitstream includes a sub-image in an image, where the image is segmented into a plurality of slices including a first slice, PPS related to the image and the sub-image, and a slice header related to the first slice. The video decoder 1110 further includes a parsing module 1113 for parsing the PPS to obtain an identifier and the slice address length of the first slice. Video decoder 1110 further includes a determination module 1115 for determining the slice address of the first slice from the slice header according to the identifier and the slice address length. Video decoder 1110 further includes a decoding module 1117 for decoding the sub-bitstream to generate a video sequence including the sub-image of the first slice. Video decoder 1110 further includes a transmission module 1119 for transmitting the video sequence including the sub-image of the first slice for display. Video decoder 1110 can also be used to perform any step of method 1000.
[0141] When there is no intermediate component between the first component and the second component other than a wire, trace, or other medium, the first component and the second component are directly coupled. When there is an intermediate component between the first component and the second component in addition to a wire, trace, or other medium, the first component and the second component are indirectly coupled. The term "coupled" and its variants include direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" means a range including ±10% of the subsequent number.
[0142] It should also be understood that the steps of the exemplary methods described herein do not necessarily need to be performed in the order described, and the order of these method steps should be understood to be merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, these methods may include other steps, and certain steps may be omitted or combined.
[0143] Although the present invention provides several embodiments, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be considered illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0144] In addition, without departing from the scope of the present invention, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques, or methods. Other instances of variations, substitutions, and alterations can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. A method implemented in a decoder, characterized in that, the method comprises: receiving a bitstream, wherein the bitstream includes sub-images in an image, wherein the image is segmented into a plurality of strips including a first strip, a set of parameters related to the image and the sub-images, and a strip header related to the first strip; parsing the bitstream to obtain a sequence parameter set SPS and an image parameter set PPS, wherein the PPS includes a first parameter for deriving the strip address length of the first strip, and the strip address length is inferred to be equal to Ceil(Log2(NumTilesInPic)) according to the presence or absence of the first parameter, where NumTilesInPic represents the number of tiles in the image, and the SPS includes a second parameter for indicating the identifier (ID) of the sub-image in the image; determining the strip address of the first strip from the strip header according to the strip address length; decoding the sub-image including the first strip according to the strip address and the second parameter.
2. The method according to claim 1, characterized in that, the method further comprises: sending a video sequence including the sub-image for display.
3. The method according to claim 1 or 2, characterized in that, the strip address length represents the number of bits included in the strip address.
4. The method according to claim 1 or 2, characterized in that, the determining the strip address of the first strip includes: using the length in the parameter set to determine a bit boundary to parse the strip address from the strip header; mapping the strip address from an image-based position to a sub-image-based position using the identifier and the strip address.
5. The method according to claim 1 or 2, characterized in that, the strip address includes a defined value but does not include an index.
6. A method implemented in an encoder, characterized in that, the method comprises: encoding an image into a bitstream, wherein the image is segmented into a plurality of strips including a first strip; encoding a parameter set into the bitstream, the parameter set including a sequence parameter set SPS and an image parameter set PPS, wherein the PPS includes a first parameter for deriving the strip address length of the first strip, and the strip address length is inferred to be equal to Ceil(Log2(NumTilesInPic)) according to the presence or absence of the first parameter, where NumTilesInPic represents the number of tiles in the image, and the SPS includes a second parameter for indicating the identifier (ID) of the sub-image in the image; encoding the strip address of the first strip into the strip header of the bitstream according to the length of the strip address and the second parameter.
7. The method according to claim 6, characterized in that, the method further comprises: storing the bitstream in at least one storage medium.
8. The method according to claim 6 or 7, characterized in that, The strip address length represents the number of bits included in the strip address.
9. The method according to claim 6 or 7, wherein, the strip address includes a defined value but does not include an index.
10. A decoder, wherein, the decoder includes: a receiver for receiving a bitstream, wherein the bitstream includes a sub-image in an image, wherein the image is segmented into a plurality of strips including a first strip, a set of parameters related to the image and the sub-image, and a strip header related to the first strip; a processor for: parsing the bitstream to obtain a sequence parameter set SPS and an image parameter set PPS, wherein the PPS includes a first parameter for deriving the strip address length of the first strip, and the strip address length is inferred to be equal to Ceil(Log2(NumTilesInPic)) according to whether the first parameter exists, where NumTilesInPic represents the number of tiles in the image, and the SPS includes a second parameter for indicating the identifier (ID) of the sub-image in the image; determining the strip address of the first strip from the strip header according to the strip address length; decoding the sub-image including the first strip according to the strip address and the second parameter.
11. The decoder according to claim 10, wherein, the processor is further configured to: send a video sequence including the sub-image for display.
12. The decoder according to claim 10 or 11, wherein, the strip address length represents the number of bits included in the strip address.
13. The decoder according to claim 10 or 11, wherein, the processor is specifically configured to: use the length in the parameter set to determine a bit boundary to parse the strip address from the strip header; map the strip address from an image-based position to a sub-image-based position using the identifier and the strip address.
14. The decoder according to claim 10 or 11, wherein, the strip address includes a defined value but does not include an index.
15. An encoder, wherein, the encoder includes: a processor for: encoding an image into a bitstream, wherein the image is segmented into a plurality of strips including a first strip; encoding a parameter set into the bitstream, the parameter set including a sequence parameter set SPS and an image parameter set PPS, wherein the PPS includes a first parameter for deriving the strip address length of the first strip, and the strip address length is inferred to be equal to Ceil(Log2(NumTilesInPic)) according to whether the first parameter exists, where NumTilesInPic represents the number of tiles in the image, and the SPS includes a second parameter for indicating the identifier (ID) of the sub-image in the image; Encode the stripe address of the first stripe into the stripe header of the bitstream according to the length of the stripe address and the second parameter; A transmitter for transmitting the bitstream to a decoder.
16. The encoder according to claim 15, wherein, the method further comprises: storing the bitstream in at least one storage medium.
17. The encoder according to claim 15 or 16, wherein, the stripe address length represents the number of bits included in the stripe address.
18. The encoder according to claim 15 or 16, wherein, the stripe address includes a defined value but does not include an index.
19. A video decoding device, wherein, the video decoding device comprises: a processor, a memory, a receiver coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the memory, the receiver, and / or the transmitter are configured to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable medium, wherein, the non-transitory computer-readable medium includes a computer program product for use by a video decoding device; the computer program product includes computer-executable instructions stored in the non-transitory computer-readable medium; when the processor executes the computer-executable instructions, the video decoding device is caused to perform the method according to any one of claims 1 to 9.
21. A non-transitory computer-readable medium, wherein, the non-transitory computer-readable medium stores a bitstream, the bitstream comprising: data representing an image, the image including sub-images, the image being segmented into a plurality of stripes including a first stripe; a sequence parameter set SPS and an image parameter set PPS, wherein the PPS includes a first parameter for deriving the stripe address length of the first stripe, the stripe address length being inferred to be equal to Ceil(Log2(NumTilesInPic)) according to the presence or absence of the first parameter, where NumTilesInPic represents the number of tiles in the image, and the SPS includes a second parameter for indicating the identifier (ID) of the sub-image in the image.
22. The non-transitory computer-readable medium according to claim 21, wherein, the stripe address length represents the number of bits included in the stripe address.
23. The non-transitory computer-readable medium according to claim 21 or 22, wherein, the stripe address includes a defined value but does not include an index.
Citation Information
Patent Citations
Image processing apparatus, image processing method, and image processing system
CN103297806A
Image encoding method, image decoding method, image encoding device, image decoding device, and image encoding / decoding device
CN104737541A
Image decoding apparatus, image decoding method
US20160381393A1
Image coding method, image decoding method, image coding apparatus, image decoding apparatus, and image coding and decoding apparatus
US20180084282A1
Image encoding method, image decoding method, image encoding device, image decoding device, and image encoding / decoding device
WO2014050038A1