Video coding method and equipment
By signaling a flag to distinguish raster scan and rectangular tile groups, the mechanism improves coding efficiency and reduces resource usage in video coding systems, addressing inefficiencies in existing systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-06-05
- Publication Date
- 2026-05-26
AI Technical Summary
Existing video coding systems face inefficiencies in handling both raster scan and rectangular tile groups, leading to increased memory, processing, and network resource usage, particularly in applications like virtual reality where user-selected sub-pictures are typically rectangular.
Implementing a mechanism that signals a flag indicating whether the tile group is raster scan or rectangular, allowing the encoder and decoder to efficiently determine tile group membership, thereby reducing resource usage by omitting the need for a complete list of tiles in the bitstream.
This approach enhances coding efficiency by minimizing memory, processing, and network resource usage, while supporting multiple tile group schemes for various applications.
Smart Images

Figure 0007864982000006 
Figure 0007864982000007 
Figure 0007864982000008
Abstract
Description
Technical Field
[0001] [Technical Field] The present disclosure generally relates to video coding, and more particularly, to a mechanism for partitioning an image into tile groups to support increased compression in video coding.
Background Art
[0002] The amount of video data required to depict even relatively short videos can be substantial. This can pose difficulties when the data is streamed across a communication network having limited bandwidth capabilities or otherwise communicated. Thus, video data is typically compressed before being communicated across today's telecommunications networks. The size of the video can also be an issue when the video is stored on a storage device, as memory resources may be limited. Video compression devices often use software and / or hardware to encode video data at the source prior to transmission or storage, thereby reducing the amount of data necessary to represent digital video images. The compressed data is then received at the destination by a video decompression device that decodes the video data. With limited network resources and the ever increasing demands for higher video quality, improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desirable.
Summary of the Invention
[0003] In an embodiment, the present disclosure includes a method implemented in an encoder, the method comprising: partitioning a picture into a plurality of tiles by a processor of the encoder; assigning the number of the tiles to tile groups by the processor; The process involves the processor encoding a flag, which is set to a first value when the tile group is a raster scan tile group and to a second value when the tile group is a rectangular tile group, wherein the flag is encoded into a parameter set of the bitstream. The processor performs the steps of encoding the tiles into the bitstream based on the tile group, The method includes the step of storing the bitstream in the encoder's memory for communication to the decoder. Some video coding systems utilize tile groups containing tiles specified in raster scan order. Other systems utilize rectangular tile groups to support subpicture extraction in virtual reality (VR), teleconferencing, and other areas of interest based on coding schemes. Still other systems allow the encoder to select which type of tile group to use depending on the type of video coding application. Aspects of the present invention include a flag indicating whether the corresponding tile group is raster scan or rectangular. This approach alerts the decoder to the appropriate tile group coding scheme to support proper decoding. Thus, the flags of the disclosure enable the encoder / decoder (codec) to support multiple tile group schemes for different use cases, thereby improving the functionality of both the encoder and the decoder. Furthermore, signaling of the flags of the disclosure improves coding efficiency and thus reduces memory resource usage, processing resource usage, and / or network resource usage in the encoder and / or decoder.
[0004] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the flag is a rectangular tile group flag.
[0005] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that the parameter set on which the flag is encoded is a sequence parameter set.
[0006] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the parameter set on which the flag is encoded is a picture parameter set.
[0007] Optionally, in any of the embodiments described above, another implementation of the embodiment further includes the step of the processor encoding an identifier for the first tile in the tile group and an identifier for the last tile in the tile group into the bitstream in order to indicate the tiles included in the tile group.
[0008] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the identifier of the first tile of the tile group and the identifier of the last tile of the tile group are encoded in the tile group header in the bitstream.
[0009] Optionally, in any of the embodiments described above, another implementation of this embodiment is such that when the tile group is the raster scan tile group, the inclusion of tiles into the tile group is The steps include determining the number of tiles between the first tile and the last tile of the tile group as the total number of tiles in the tile group, A step of determining the inclusion of tiles based on the number of tiles in the tile group, It provides that this is determined by [the specified method / method].
[0010] Optionally, in any of the embodiments described above, another implementation of this embodiment is such that when the tile group is the rectangular tile group, the inclusion of tiles into the tile group is A step of determining the delta value between the first tile and the last tile of the tile group, The steps include determining the number of tile group rows based on the delta value and the number of tile columns in the picture, The steps include determining the number of tile group columns based on the delta value and the number of tile columns in the picture, A step of determining the inclusion of the tile based on the number of tile group rows and the number of tile group columns, It provides that this is determined by [the specified method / method].
[0011] In embodiments, the present disclosure includes a method implemented in a decoder, the method being: The decoder's processor receives a bitstream containing a picture partitioned into multiple tiles via a receiver, wherein the number of tiles is included in a tile group. The processor performs the steps of obtaining a flag from the parameter set of the bitstream, The processor determines that the tile group is a raster scan tile group when the flag is set to a first value, The processor determines that the tile group is a rectangular tile group when the flag is set to the second value, The processor performs the steps of determining whether the tile group is a raster scan tile group or a rectangular tile group, and determining whether the tile group includes tiles. The processor performs the steps of decoding the tiles based on the tile group to generate decoded tiles, The processor provides the steps of generating a reconstructed video sequence for display based on the decoded tiles, This includes: Some video coding systems utilize tile groups containing tiles specified in raster scan order. Other systems utilize rectangular tile groups to support subpicture extraction in VR, teleconferencing, and other areas of interest based on coding schemes. Still other systems allow the encoder to select which type of tile group to use depending on the type of video coding application. Aspects of the present invention include a flag indicating whether the corresponding tile group is raster scan or rectangular. This approach alerts the decoder to the appropriate tile group coding scheme to support proper decoding. Thus, the flags of the disclosure enable the codec to support multiple tile group schemes for different use cases, thereby improving the functionality of both the encoder and the decoder. Furthermore, signaling of the flags of the disclosure improves coding efficiency and thus reduces memory resource usage, processing resource usage, and / or network resource usage in the encoder and / or decoder.
[0012] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the flag is a rectangular tile group flag.
[0013] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the parameter set including the flag is a sequence parameter set.
[0014] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the parameter set including the flag is a picture parameter set.
[0015] Optionally, in any of the embodiments described above, another implementation of the embodiment provides that the processor further includes the step of obtaining an identifier for the first tile in the tile group and an identifier for the last tile in the tile group in order to determine which tiles are included in the tile group.
[0016] Optionally, in any of the above aspects, another implementation of this aspect provides that the identifier of the first tile of the tile group and the identifier of the last tile of the tile group are obtained from a tile group header in the bitstream.
[0017] Optionally, in any of the above aspects, another implementation of this aspect provides that when the tile group is a raster scan tile group, the inclusion of tiles in the tile group determining the number of tiles between the first tile of the tile group and the last tile of the tile group as the number of tiles in the tile group; determining tile inclusion based on the number of tiles in the tile group; is determined by.
[0018] Optionally, in any of the above aspects, another implementation of this aspect provides that when the tile group is a rectangular tile group, the inclusion of tiles in the tile group determining a delta value between the first tile of the tile group and the last tile of the tile group; determining the number of tile group rows based on the delta value and the number of tile columns in the picture; determining the number of tile group columns based on the delta value and the number of tile columns in the picture; determining tile inclusion based on the number of tile group rows and the number of tile group columns; is determined by.
[0019] In an embodiment, the present disclosure is a video coding device, A video coding device includes a processor, a receiver connected to the processor, and a transmitter connected to the processor, and the processor, the receiver, and the transmitter are configured to execute any of the methods in the above-described manner.
[0020] In one embodiment, the present disclosure includes a non-transitory computer-readable medium for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute any of the methods in the above-described manner.
[0021] In an embodiment, the present disclosure is an encoder, partitioning means for partitioning a picture into a plurality of tiles; inclusion means for including the number of the tiles in a tile group; encoding means for encoding a flag set to a first value when the tile group is a raster scan tile group and to a second value when the tile group is a rectangular tile group, the flag being encoded in a parameter set of a bitstream; encoding means for encoding the tiles into the bitstream based on inclusion of the tiles; storage means for storing the bitstream for communication to a decoder; including an encoder.
[0022] Optionally, in any of the above-described manners, another implementation of this aspect provides that the encoder is further configured to execute any of the methods in the above-described manner.
[0023] In an embodiment, the present disclosure is a decoder, A receiving means for receiving a bitstream containing a picture partitioned into multiple tiles, wherein the number of tiles is included in a tile group, and An acquisition means for obtaining a flag from the parameter set of the bitstream, A means of decision-making, When the flag is set to the first value, it is determined that the tile group is a raster scan tile group. When the aforementioned flag is set to the second value, it is determined that the tile group is a rectangular tile group. A determination means for determining whether the tile group is a raster scan tile group or a rectangular tile group, and for determining whether the tile group is a raster scan tile group or a rectangular tile group. A decoding means that decodes the tiles based on the tile group and generates decoded tiles, A generation means for generating a reconstructed video sequence for display based on the decoded tiles, Includes a decoder.
[0024] Optionally, in any of the embodiments described above, another implementation of this embodiment provides that the decoder is further configured to perform any of the methods described above.
[0025] For clarity, any one of the embodiments described above may be combined with any one or more of the other embodiments described above to generate a new embodiment within the scope of this disclosure.
[0026] The features described above and other features will be understood more clearly from the following detailed description, which is incorporated in conjunction with the attached drawings and claims. [Brief explanation of the drawing]
[0027] For a more complete understanding of this disclosure, the following brief description is made in relation to the accompanying drawings and detailed description. Herein, similar reference numerals represent similar parts.
[0028] [Figure 1] This is a flowchart illustrating an exemplary method for coding a video signal.
[0029] [Figure 2] This is a schematic diagram illustrating an exemplary coding and decoding (codec) system for video coding.
[0030] [Figure 3] This is a schematic diagram illustrating an example video encoder.
[0031] [Figure 4] This is a schematic diagram illustrating an exemplary video decoder.
[0032] [Figure 5] This is a schematic diagram showing an exemplary bitstream containing an encoded video sequence.
[0033] [Figure 6] This is a schematic diagram showing exemplary pictures partitioned into raster scantile groups.
[0034] [Figure 7] This is a schematic diagram showing an exemplary picture partitioned into rectangular tile groups.
[0035] [Figure 8] This is a schematic diagram of an exemplary video coding device.
[0036] [Figure 9] This is a flowchart illustrating an exemplary method for encoding a picture into a bitstream.
[0037] [Figure 10] This is a flowchart illustrating an exemplary method for decoding a picture from a bitstream.
[0038] [Figure 11]This is a schematic diagram of an exemplary system for coding a video-picture sequence within a bitstream. [Modes for carrying out the invention]
[0039] It should be understood from the outset that while one or more explanatory implementations of embodiments are provided below, the systems and / or methods of the disclosure may be implemented using any number of techniques, whether currently known or existing. This disclosure is by no means limited to the explanatory implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, and may be modified within the scope of the appended claims, along with the entire scope of their equivalents.
[0040] Various abbreviations are used here, such as coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded video sequence (CVS), Joint Video Experts Team (JVET), motion-constrained tile set (MCTS), maximum transfer unit (MTU), network abstraction layer (NAL), picture order count (POC), raw byte sequence payload (RBSP), sequence parameter set (SPS), versatile video coding (VVC), and working draft (WD).
[0041] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or eliminate data redundancy in a video sequence. In block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be called tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU), and / or coding nodes. Video blocks in an intra-coding (I) slice of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks in an intercoding unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may be called a frame and / or image, and a reference picture may be called a reference frame and / or reference image. Spatial or temporal predictions produce predicted blocks representing image blocks. Residual data represents the pixel difference between the original image blocks and the predicted blocks. Thus, an intercoding block is encoded according to a motion vector pointing to a block of reference samples forming the predicted block, and residual data indicating the difference between the coding block and the predicted block. An intracoding block is encoded according to the intracoding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to the transformation domain. These produce residual transformation coefficients that may be quantized. The quantized transformation coefficients may first be organized into a two-dimensional array. The quantized transformation coefficients may be scanned to generate transformation coefficients in a one-dimensional vector. Entropy coding may be applied to achieve even greater compression.Such video compression techniques will be discussed in more detail below.
[0042] To ensure that encoded video is accurately decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include Advanced Video Coding (AVC), also known as ITU-T H.261, Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262, or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264, or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The ITU-T and ISO / IEC Joint Video Experts Team (JVET) has begun developing a video coding standard called Versatile Video Coding (VVC).VVC is included in the Working Draft (WD) which includes JVET-L1001-v5.
[0043] To code a video image, the image is first partitioned, and the partitions are coded into bitstreams. Various picture partitioning schemes are available. For example, an image can be partitioned according to normal slices, dependent slices, tiles, and / or Wavefront Parallel Processing (WPP). For simplicity, HEVC restricts the encoder so that only normal slices, dependent slices, tiles, WPP, and combinations thereof can be used when partitioning slices into groups of CTBs for video coding. Such partitioning can be applied to support Maximum Transfer Unit (MTU) size fitting, parallel processing, and reduced end-to-end latency. The MTU indicates the maximum amount of data that can be transmitted in a single packet. If the packet payload exceeds the MTU, the payload is split into two packets through a process called fragmentation.
[0044] A normal slice, also simply called a slice, is a partitioned portion of an image that can be reconstructed independently of other normal slices within the same picture, despite any interdependencies due to loop filtering operations. Each normal slice is encapsulated within its own Network Abstraction Layer (NAL) unit for transmission. Furthermore, in-picture predictions (intra-sample predictions, motion information predictions, coding mode predictions) and entropy coding dependencies across slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, parallelization based on normal slices utilizes minimal interprocessor and intercore communication. However, since each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can result in considerable coding overhead due to the bit cost of the slice header per slice and the lack of predictions across slice boundaries. Furthermore, normal slices may be used to support compliance with MTU size requirements. Specifically, since a typical slice is usually encapsulated in a separate NAL unit and can be coded independently, each typical slice should be smaller than the MTU in the MTU scheme to prevent the slice from being broken down into multiple packets. Thus, the objectives of parallelization and MTU size conformance can impose conflicting requirements on the slice layout within a picture.
[0045] Dependent slices are similar to normal slices but have a shortened slice header and allow partitioning at image tree block boundaries without breaking in-picture predictions. Thus, dependent slices allow normal slices to be fragmented into multiple NAL units, which results in reduced end-to-end delay by making portions of the normal slice available for transmission before the encoding of the entire normal slice is complete.
[0046] A tile is a partitioned portion of an image generated by horizontal and vertical boundaries that create columns and rows of tiles. Tiles may be coded in raster scan order (right to left and top to bottom). The scan order of CTBs is local within a tile. Thus, CTBs in the first tile are coded in raster scan order before proceeding to CTBs in the next tile. As with regular slices, tiles break in-picture prediction dependencies, as well as entropy decoding dependencies. However, tiles do not have to be contained within individual NAL units, and therefore tiles do not have to be used for MTU size fitting. Each tile can be processed by one processor / core, and interprocessor / intercore communication used for in-picture prediction between processing units decoding neighboring tiles may be limited to carrying a shared slice header (when adjacent tiles are in the same slice) and performing the sharing of reconstructed samples and metadata related to loop filtering. When a slice contains more than one tile, the entry point byte offsets of each tile, other than the first entry point offset in the slice, may be signaled in the slice header. For each slice and tile, at least one of the following conditions must be met: 1) All coding tree blocks within a slice belong to the same tile, and 2) All coding tree blocks within a tile belong to the same slice.
[0047] In WPP, an image is partitioned into a single row of CTBs. The entropy decoding and prediction mechanism may use data from CTBs in other rows. Parallel processing is enabled through parallel decoding of CTB rows. For example, the current row may be decoded in parallel with the preceding row. However, the decoding of the current row lags behind the decoding of the preceding row by only 2CTBs. This delay ensures that data related to the CTB above and to the upper right of the current CTB in the current row is available before the current CTB is coded. This approach appears as a wavefront graphically. This time-delayed start allows for parallelization by up to the same number of processors / cores as the number of CTB rows the image contains. Since intra-picture prediction between neighboring tree block rows within a picture is allowed, interprocessor / intercore communication that enables intra-picture prediction can be important. WPP partitions consider NAL unit size. Therefore, WPP does not support MTU size fitting. However, typical slicing involves certain coding overhead and can be used in conjunction with WPP to achieve the desired MTU size conformance.
[0048] Tiles may include motion-constrained tile sets. A motion-constrained tile set (MCTS) is a tile set designed so that the associated motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that require only full sample positions within the MCTS for interpolation. Furthermore, the use of motion vector candidates for time motion vector prediction derived from blocks outside the MCTS is not permitted. Thus, each MCTS may be decoded independently, without the presence of tiles not included in the MCTS. Time MCTS supplemental enhancement information (SEI) messages may be used to indicate the presence of MCTS in the bitstream and to signal MCTS. MCTS SEI messages provide supplemental information that can be used in MCTS sub-bitstream extraction (specified as the semantic part of the SEI message) to generate an MCTS confirmation bitstream. The information includes the number of extracted information sets, each defining the number of MCTS and including raw bytes sequence payload (RBSP) bytes for the replacement video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) to be used during the MCTS sub-bitstream extraction process. When extracting sub-bitstreams according to the MCTS sub-bitstream extraction process, one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) may have different values in the extracted sub-bitstream, so the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced and the slice header may be updated.
[0049] This disclosure relates to various tiling schemes. Specifically, when an image is partitioned into tiles, such tiles can be assigned to tile groups. A tile group is a set of related tiles that can be individually extracted and coded, for example, to support the display of a region of interest and / or to support parallel processing. Tiles can be assigned to tile groups, enabling the application of corresponding parameters, functions, coding tools, etc., on a group-by-group basis. For example, a tile group may include MCTS. As another example, tile groups may be processed and / or extracted individually. Some systems utilize a raster scan mechanism to generate corresponding tile groups. When used here, a raster scan tile group is a tile group generated by assigning tiles in a raster scan order. The raster scan order proceeds sequentially from right to left and top to bottom between the first and last tiles. Raster scan tile groups can be useful for some applications, for example, to support parallel processing.
[0050] However, raster scan tile groups can be inefficient in some cases. For example, in virtual reality (VR) applications, the environment is recorded as a sphere encoded in a picture. The user can then experience the environment by browsing user-selected sub-pictures of the picture. The user-selected sub-pictures may be called regions of interest. Allowing the user to selectively perceive parts of the environment generates the sense that the user is present within that environment. Thus, the unselected parts of the picture may not be visible and are therefore discarded. Consequently, user-selected sub-pictures may be treated differently from unselected sub-pictures (for example, unselected sub-pictures may be signaled at a lower resolution and processed using a simpler mechanism during rendering, etc.). Tile groups allow for such different treatment among sub-pictures. However, user-selected sub-pictures are typically rectangular and / or square regions. Therefore, raster scan tile groups may not be useful in such use cases.
[0051] To overcome these problems, some systems utilize rectangular tile groups. A rectangular tile group is a group of tiles containing a set of tiles that, when viewed as a whole, produce a rectangular shape. A rectangular shape, when used here, is a shape in which all four sides are connected precisely such that each side is connected to two other sides at an angle of 90°. Both tile group approaches (e.g., raster scan tile groups and rectangular tile groups) may have advantages and disadvantages. Therefore, a video coding system may want to support both approaches. However, a video coding system may not be able to efficiently signal the use of tile groups when both approaches are available. For example, a simple merge of the signaling of these approaches may result in an inefficient and / or processor-intensive complex syntax structure in the encoder and / or decoder. This disclosure presents a mechanism for solving these and other problems in video coding technology.
[0052] Disclosed herein are various mechanisms for harmonizing the use of raster scan tile groups and rectangular tile groups by utilizing simple and compact signaling. Such signaling improves coding efficiency and therefore reduces memory resource usage, processing resource usage, and / or network resource usage in encoders and / or decoders. To harmonize these approaches, the encoder may signal a flag indicating which type of tile group is being used. For example, the flag may be a rectangular tile group flag that may be signaled in a parameter set such as SPS and / or PPS. The flag can indicate whether the encoder is using a raster scan tile group or a rectangular tile group. The encoder can therefore indicate tile group membership by simply signaling the first and last tiles in the tile group. Based on the first tile, the last tile, and the indication of the tile group type, the decoder can determine which tiles are included in the tile group. Thus, a complete list of all tiles in each tile group can be omitted from the bitstream, which improves coding efficiency. For example, if a tile group is a raster scan tile group, the tiles assigned to the tile group can be determined by determining the number of tiles between the first and last tiles of the tile group, and then adding that number of tiles, each having an identifier between the first and last tiles, to the tile group. If the tile group is a rectangular tile group, a different approach can be used. For example, the delta value between the first and last tiles of the tile group can be determined. Then, the number of rows and columns of the tile group can be determined based on the delta value and the number of tile columns in the picture. The tiles within the tile group can then be determined based on the number of rows and columns of the tile group. These and other examples are described in detail below.
[0053] Figure 1 is a flowchart of an exemplary operation method 100 for coding a video signal. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. Smaller file sizes allow for the transmission of a compressed video file to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video signal to reconstruct the original video signal for display to the end user. The decoding process is typically a mirror of the encoding process, allowing the decoder to reconstruct the video signal without inconsistencies.
[0054] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may contain both an audio component and a video component. The video component contains a series of image frames that, when viewed in a sequence, give a visual impression of motion. Each frame contains pixels, which are represented here in terms of light, called the lumina component (or lumina sample), and color, called the chroma component (or chroma sample). In some examples, the frame may also include depth values to support three-dimensional display.
[0055] In step 103, the video is partitioned into blocks. Partitioning involves subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU contains both lumens and chroma samples. The coding tree may be used to divide the CTUs into blocks, and then to repeatedly subdivide the blocks until a configuration supporting further coding is achieved. For example, the lumens component of a frame may be subdivided until the individual blocks contain relatively similar light values. Furthermore, the chroma component of a frame may be subdivided until the individual blocks contain relatively similar color values. Thus, the partitioning mechanism varies depending on the content of the video frame.
[0056] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, interpretation and / or intrapretation may be used. Interpretation is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Therefore, a block depicting an object in a reference frame does not need to be shown repeatedly in adjacent frames. Specifically, an object such as a table may remain in the same position across multiple frames. Thus, a table is shown once, and adjacent frames can refer back to the reference frame. Pattern matching mechanisms may be used to match objects across multiple frames. Furthermore, moving objects may be shown across multiple frames, for example, due to the movement of an object or the movement of the camera. As a particular example, a video may show a car moving horizontally across the screen across multiple frames. Motion vectors can be used to show such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of that object in a reference frame. Therefore, interpretation can encode the image blocks in the current frame as a set of motion vectors indicating the offset from the corresponding block in the reference frame.
[0057] Intra-prediction encodes blocks within a common frame. Intra-prediction takes advantage of the fact that luma and chroma components tend to be densely packed within a frame. For example, some green patches in a tree tend to be located adjacent to similar green patches. Intra-prediction utilizes multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional modes indicate that a current block is similar to / identical to samples of neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the ends of the row. Planar mode effectively shows smooth light / color transitions across rows / columns by utilizing a relatively constant gradient of changing values. DC mode is used for boundary smoothing and indicates that a block is similar to / identical to the mean value related to samples of all neighboring blocks related to the angular direction of the directional prediction mode. Thus, intra-prediction blocks can represent image blocks as various relational prediction modes instead of actual values. Furthermore, intra-prediction blocks can represent image blocks as motion vector values instead of actual values. In any case, the prediction block may not accurately represent the image in some instances. Any differences are stored in the residual block. To further compress the file, transformations may be applied to the residual block.
[0058] In step 107, various filtering techniques may be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may result in the generation of an image with uneven grayscale in the decoder. Furthermore, the block-based prediction scheme may encode blocks and then reconstruct the encoded blocks for later use as reference blocks. The in-loop filtering scheme repeatedly applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to the blocks / filters. These filters mitigate such uneven grayscale artifacts, and as a result, the encoded file can be accurately reconstructed. Moreover, these filters mitigate artifacts in the reconstructed reference blocks, and as a result, there is a low probability of additional artifacts occurring in subsequent blocks encoded based on the reconstructed reference blocks.
[0059] Once the video signal is partitioned, compressed, and filtered, the resulting data is encoded in a bitstream in step 109. The bitstream contains the data described above, and any desired signaling data to support proper video signal reconstruction in the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may be broadcast and / or multicast to multiple decoders. The generation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across a large number of frames and blocks. The order shown in Figure 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to a specific order.
[0060] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax from the bitstream to determine the partitions of the frame. The partitions should match the result of the block partitions in step 103. The entropy coding / decoding used in step 111 is described below. During the compression process, the encoder generates many choices, such as selecting a block partition scheme from several possible options based on the spatial location of values in the input image. Signaling the exact choices can utilize a vast number of bins. As used here, bins are binary values treated as variables (e.g., bit values that can change depending on the context). Entropy coding allows the encoder to discard any choices that are obviously not feasible in particular, leaving a set of acceptable choices. Each acceptable choice is then assigned a codeword. The length of the codeword is based on the number of acceptable choices (e.g., one bin for two choices, two bins for three to four choices, etc.). The encoder then encodes a codeword for the selected choice. This method reduces the size of the codeword by being large enough to uniquely represent a choice from a small subset of possible choices, rather than uniquely representing a choice from a potentially large set of all possible choices. The decoder then decodes the choice by determining the set of acceptable choices in a similar manner to the encoder. By determining the set of acceptable choices, the decoder can read the codeword and determine the choice made by the encoder.
[0061] In step 113, the decoder performs block decoding. Specifically, the decoder generates residual blocks using the inverse transform. Next, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the partitions. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks generated in step 105 in the encoder. The reconstructed image blocks are then positioned into frames of the reconstructed video signal according to the partition data determined in step 111. The syntax of step 113 may also be signaled in the bitstream by entropy coding as described above.
[0062] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to that in step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters may be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal can be output to a display in step 117 for viewing by the end user.
[0063] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions to support the implementation of operation method 100. The codec system 200 is generalized to show the components used in both the encoder and the decoder. The codec system 200 receives and partitions the video signal described above with respect to steps 101 and 103 of operation method 100, which produces a partitioned video signal 201. The codec system 200 then compresses the partitioned video signal 201 into a coding bitstream when operating as an encoder described above with respect to steps 105, 107, and 109 of method 100. When operating as a decoder, the codec system 200 generates an output video signal from the bitstream as described above with respect to steps 111, 113, 115, and 117 of operation method 100. The codec system 200 includes a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter-controlled analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and ontext adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown in the figure. In Figure 2, the black lines show the movement of the data to be encoded / decoded, while the dashed lines show the movement of the control data that controls the operation of other components. All components of the codec system 200 may reside within the encoder. The decoder may include some of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described here.
[0064] The partitioned video signal 201 is a captured video sequence partitioned into blocks of pixels by a coding tree. The coding tree subdivides blocks of pixels into smaller blocks of pixels using various partitioning modes. These blocks can then be further subdivided into even smaller blocks. Blocks may be called nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the node / coding tree depth. In some cases, the partitioned blocks may be contained within a coding unit (CU). For example, a CU may be a sub-part of a CTU containing a luma block, a red difference chroma (Cr) block, and a blue difference chroma (Cb) block, as well as the corresponding syntax instructions of the CU. Partitioning modes may include binary trees (BT), triple trees (TT), and quad trees (QT), which are used to partition a node into 2, 3, or 4 child nodes, respectively, with shapes that vary depending on the partitioning mode used. The partitioned video signal 201 is transferred for compression to a general-purpose coder control component 211, a transformation scaling and quantization component 213, an intrapicture estimation component 215, a filter control analysis component 227, and a motion estimation component 221.
[0065] The general-purpose coder control component 211 is configured to make decisions related to coding images of a video sequence into a bitstream, in accordance with application constraints. For example, the general-purpose coder control component 211 manages the optimization of bitrate / bitstream size for reconstruction quality. Such decisions may be based on memory space / bandwidth availability and image resolution requirements. The general-purpose coder control component 211 also manages buffer utilization in terms of conversion speed to mitigate buffer underrun and overrun problems. To address these issues, the general-purpose coder control component 211 manages partitioning, prediction, and filtering by other components. For example, the general-purpose coder control component 211 may dynamically increase compression complexity to increase resolution, and increase bandwidth usage or decrease compression complexity to reduce resolution and bandwidth usage. Thus, the general-purpose coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality and bitrate concerns. The general-purpose coder control component 211 generates control data that controls the operation of other components. The control data is also transferred to the header format and CABAC component 231 so that it is encoded within the bitstream to signal the parameters for decoding in the decoder.
[0066] The partitioned video signal 201 is also transmitted to the motion estimation component 221 and the motion compensation component 219 for interpretation. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform interpretation coding of the received video blocks in relation to one or more blocks in one or more reference frames, and provide a temporary prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0067] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is the process of generating motion vectors that estimate the motion for a video block. The motion vectors may, for example, indicate the placement of coding objects in relation to a predicted block. A predicted block is a block that has been found to exactly match the block to be coded in terms of pixel difference. A predicted block may also be called a reference block. Such pixel differences may be determined by the sum of absolute differences (SAD), the sum of square differences (SSD), or other difference metrics. HEVC utilizes several coding objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be split into CTBs, and a CTB can then be split into CBs to be included in a CU. A CU can be encoded as a prediction unit (PU) containing the predicted data and / or a transform unit (TU) containing the transformed residual data of the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and select the reference blocks, motion vectors, etc. that have the optimal rate distortion characteristics. The optimal rate distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and coding efficiency (e.g., the size of the final encoding).
[0068] In some examples, the codec system 200 may calculate the value of a sub-integer picture position of the reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate the value of a quarter-pixel position, an eighth-pixel position, or other fractional pixel position of the reference picture. Thus, the motion estimation component 221 may perform motion search in relation to the complete pixel position and fractional pixel position and output a motion vector with fractional pixel precision. The motion estimation component 221 calculates a motion vector for the PU of a video block in the intercoding slice by comparing the PU position with the predicted block position of the reference picture. The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and CABAC component 231 for encoding, and the motion to the motion compensation component 219.
[0069] Motion compensation performed by the motion compensation component 219 may include fetching or generating a predicted block based on a motion vector determined by the motion estimation component 221. Here again, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector of the current video block's PU, the motion compensation component 219 may identify the location of the predicted block pointed to by the motion vector. Next, a residual video block is formed by subtracting the pixel values of the predicted block from the coded pixel values of the current video block to form a pixel difference value. Generally, the motion estimation component 221 performs motion estimation in relation to the lumens component, and the motion compensation component 219 uses a motion vector calculated based on the lumens component for both the chromens and lumens components. The predicted block and residual block are transferred to the transformation scaling and quantization component 213.
[0070] The partitioned video signal 201 is also transmitted to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are shown separately for conceptual purposes. Instead of inter-prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 as described above, the intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block in relation to the block in the current frame. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to be used to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode for encoding the current block from several tested intra-prediction modes. The selected intra-prediction mode is then transmitted to the header format and CABAC component 231 for encoding.
[0071] For example, the intrapicture estimation component 215 calculates rate distortion values for various tested intraprediction modes using rate distortion analysis and selects the intraprediction mode with the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that coded and generated the coded block, as well as the bit rate (e.g., number of bits) used to generate the coded block. The intrapicture estimation component 215 calculates a ratio from the distortion and rate for various coded blocks to determine which intraprediction mode exhibits the best rate distortion value for a block. Furthermore, the intrapicture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0072] When implemented in an encoder, the intra-picture prediction component 217 may generate residual blocks from prediction blocks based on a selected intra-prediction mode determined by the intra-picture estimation component 215, or when implemented in a decoder, it may read residual blocks from the bitstream. The residual blocks contain the difference in values between the prediction blocks and the original blocks, expressed as a matrix. The residual blocks are then transferred to the transformation scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate for both luma and chroma components.
[0073] The transformation scaling and quantization component 213 is configured to further compress the residual block. The transformation scaling and quantization component 213 applies a transformation such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transformation to the residual block to generate a video block containing residual transformation coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transformations may also be used. The transformation may convert the residual information from a pixel value domain to a transformation domain such as a frequency domain. The transformation scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a magnification factor to the residual information. As a result, different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transformation scaling and quantization component 213 is also configured to quantize the transformation coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be changed by adjusting the quantization parameters. In some examples, the transformation scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transformation coefficients. The quantized transformation coefficients are then transferred to the header format and CABAC component 231 for encoding within the bitstream.
[0074] The scaling and inverse transform component 229 applies the inverse of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel domain for later use as a reference block that may become a predicted block for another current block, for example. The motion estimation component 221 and / or motion compensation component 219 may compute the reference block by adding the residual block back to the corresponding predicted block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transform. Such artifacts could otherwise lead to inaccurate predictions (and generate additional artifacts) when subsequent blocks are predicted.
[0075] The filter-controlled analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding predictive block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may be applied to the residual block instead. As with the other components in Figure 2, the filter-controlled analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes several parameters to adjust how such a filter is applied. The filter-controlled analysis component 227 analyzes the reconstructed reference block to determine when such a filter should be applied and sets the corresponding parameters. Such data is transferred to the header format and CABAC component 231 as filter-controlled data for encoding. The in-loop filter component 225 applies such filters based on filter control data. These filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Depending on the example, these filters may be applied in the spatial / pixel domain (e.g., on a reconstructed pixel block) or in the frequency domain.
[0076] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them to the display. The decoded picture buffer component 223 may be any memory device capable of storing the prediction blocks, residual blocks, and / or reconstructed image blocks.
[0077] The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coding bitstream for transmission to the decoder. Specifically, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Furthermore, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transformation coefficient data are all encoded within the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of coding contexts for various blocks, indications for the most promising intra-prediction modes, indications for partition information, etc. Such data may be encoded using entropy coding. For example, information may be encoded using context-adaptive variable length coding (CAVLC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. The bitstream coded according to entropy coding may be transmitted to another device (e.g., a video decoder) or stored for later transmission or retrieval.
[0078] Figure 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107 and / or 109 of the operation method 100. The encoder 300 partitions the input video signal to produce a partitioned video signal 301 which is substantially similar to the partitioned video signal 201. The partitioned video signal 301 is then compressed and encoded into a bitstream by the components of the encoder 300.
[0079] Specifically, the partitioned video signal 301 is transferred to the intra-picture prediction component 317 for intra-prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also transferred to the motion compensation component 321 for inter-prediction based on the reference block in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transformation and quantization component 313 for transformation and quantization of the residual block. The transformation and quantization component 313 may be substantially the same as the transformation scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (along with the associated control data) are transferred to the entropy coding component 313 for coding into a bitstream. The entropy coding component 331 may be substantially the same as the header format and the CABAC component 231.
[0080] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantize component 313 to the inverse transform and quantize component 329 for reconstruction into reference blocks for use by the motion compensation component 321. The inverse transform and quantize component 329 may be substantially the same as the scaling and inverse transform component 229. The in-loop filters in the in-loop filter component 325 are also applied to the residual blocks and / or reconstructed reference blocks, depending on the example. The in-loop filter component 325 may be substantially the same as the filter-controlled analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may contain multiple filters, as discussed with respect to the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223.
[0081] Figure 4 is a block diagram illustrating an exemplary video decoder 400. A video encoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115 and / or 117 of the operation method 100. The decoder 400 receives a bitstream from, for example, the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to the end user.
[0082] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context for interpreting additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transformation coefficients from residual blocks. The quantized transformation coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0083] The reconstructed residual blocks and / or predicted blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on the intra-prediction operation. The intra-picture prediction component 417 may be the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses the prediction mode to determine the position of the reference block in the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block, and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the decoded picture buffer component 223 and the in-loop filter component 425, which may be substantially the same as the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predicted block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially the same as the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses motion vectors from a reference block to generate a predicted block and provides a residual block to reconstruct the image block. The resulting reconstructed block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 may continue to store additional reconstructed image blocks that can be reconstructed into frames by partition information. Such frames may be placed in a sequence. The sequence is output to a display as a reconstructed output video signal.
[0084] Figure 5 is a schematic diagram showing an exemplary bitstream 500 containing an encoded video sequence. For example, bitstream 500 can be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400. As another example, bitstream 500 may be generated by encoder in step 109 of method 100 for use by decoder in step 111.
[0085] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPS) 512, a tile group header 514, and image data 520. The SPS 510 contains sequence data common to all pictures in the video sequence contained in the bitstream 500. Such data may include picture sizing, bit depth, coding tool parameters, bitrate constraints, etc. The PPS 512 contains parameters specific to one or more corresponding pictures. Thus, each picture in the video sequence may refer to one PPS 512. The PPS 512 may indicate the coding tools available for the tiles in the corresponding picture, quantization parameters, offsets, picture-specific coding tool parameters (e.g., filter controls), etc. The tile group header 514 contains parameters specific to each tile group in the picture. Thus, there may be one tile group header 514 for each tile group in the video sequence. The tile group header 514 may include tile group information, picture order count (POC), reference picture list, prediction weights, tile entry point, deblocking parameters, etc. It should be noted that some systems refer to the tile group header 514 as a slice header and use such information to support slices instead of tile groups.
[0086] Image data 520 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed quantized residual data. Such image data 520 is sorted according to the partitions used to partition the image before encoding. For example, the image in image data 520 is divided into one or more tile groups 521. Each tile group 521 contains one or more tiles 523. The tiles 523 are further divided into coding tree units (CTUs). The CTUs are further divided into coding blocks based on the coding tree. The coding blocks can then be encoded / decoded according to the prediction mechanism. An image / picture may contain one or more tile groups 521 and one or more tiles 523.
[0087] A tile group 521 is a set of related tiles 523 that can be extracted and coded separately, for example, to support the display of a region of interest and / or to support parallel processing. A picture may contain one or more tile groups 521. Each tile group 521 refers to a coding tool in its corresponding tile group header 514. Thus, a tile group 521 can now be coded using a different coding tool than other tile groups 521 by modifying the data in its corresponding tile group header 514. A tile group 521 can be described in terms of the mechanism used to assign tiles 523 to the tile group 521. A tile group 521 containing tiles 523 assigned in raster scan order may be called a raster scan tile group. A tile group 521 containing tiles 523 assigned to generate rectangles (or squares) may be called a rectangle tile group. Figures 6-7 include examples of raster scan tile groups and rectangle tile groups, respectively, which will be discussed in more detail below.
[0088] A tile 523 is a partitioned portion of a picture generated by horizontal and vertical boundaries. A tile 523 may be rectangular and / or square. The picture may be partitioned into rows and columns of tiles 523. A row of tiles 523 is a set of tiles 523 positioned horizontally adjacent to generate a continuous line from the left boundary to the right boundary of the picture (and vice versa). A column of tiles 523 is a set of tiles 523 positioned vertically adjacent to generate a continuous line from the top boundary to the bottom boundary of the picture (and vice versa). Depending on the example, a tile 523 may or may not allow predictions based on other tiles 523. For example, a tile group 521 may contain a set of tiles 523 designated as an MCTS. Tiles 523 within an MCTS can be coded by predictions from other tiles 523 within the MCTS, rather than by tiles 523 outside the MCTS. Tiles 523 can be further partitioned into CTUs. A coding tree can be used to partition a CTU into coding blocks, which can then be coded according to intra-prediction or inter-prediction.
[0089] Each tile 523 may have a unique tile index 524 within the picture. The tile index 524 is a procedurally selected numerical identifier that can be used to distinguish one tile 523 from another. For example, the tile index 524 may be numerically incremented in the raster scan order, which is left to right and top to bottom. It should be noted that in some examples, a tile 523 may also be assigned a tile identifier (ID). The tile ID is an assigned identifier that can be used to distinguish one tile 523 from another. In some examples, calculations may utilize the tile ID instead of the tile index 524. In some examples, the tile ID may also be assigned to have the same value as the tile index 524. In some examples, the tile index 524 and / or ID may be signaled to indicate the boundary of a tile group 521 containing the tile 523. Furthermore, the tile index 524 and / or ID may be used to map the image data 520 associated with the tile 523 to the correct position for display.
[0090] As described above, tile group 521 may be a raster scan tile group or a rectangular tile group. This disclosure includes a signaling mechanism to enable the codec to support both types of tile group 521 in a manner that supports improved coding efficiency and reduced complexity. The tile group flag 531 is a data unit that can be used to signal whether the corresponding tile group 521 is raster scan or rectangular. The tile group flag 531 can be signaled in SPS 510 or PPS 512, depending on the example. Tiles 523 assigned to tile group 521 can be signaled by indicating the first tile 532 and the last tile 533 in the bitstream 500. For example, the first tile 532 may contain the tile index 524 or ID of the first tile 523 in the tile group 521. The first position is the top-left corner for a rectangular tile group and the minimum index / ID for a raster scan tile group. Furthermore, the last tile 533 may contain the tile index 524 or ID of the last tile 523 in the tile group 521. The last position is the bottom right corner in a rectangular tile group and the maximum index / ID in a raster scan tile group.
[0091] The tile group flag 531, the first tile 532, and the last tile 533 provide sufficient information for the decoder to determine which tile 523 is in tile group 521. For example, a raster scan mechanism can determine which tile 523 is in a raster scan tile group based on the first tile 532 and the last tile 533. Furthermore, a rectangle mechanism can determine which tile 523 is in a rectangle tile group based on the first tile 532 and the last tile 533. This makes it possible to omit the tile index 524 of other tiles 523 in the corresponding tile group 521 from the bitstream 500. This reduces the size of the bitstream 500 and therefore improves coding efficiency. Thus, the tile group flag 531 provides sufficient information for the decoder to determine which mechanism should be used to determine which tile 523 is assigned to tile group 521.
[0092] Therefore, the encoder can determine whether to use a raster scan or a rectangular tile group for the bitstream 500 or a sub-part thereof. The encoder can then set the tile group flag 531. Furthermore, the encoder can assign tile 523 to tile group 521 and include the first tile 532 and the last tile 533 in the bitstream 500. The hypothetical reference decoder (HRD) in the encoder can then determine the assignment of tile 523 to tile group 521 based on the tile group flag 531, the first tile 532, and the last tile 533. The HRD is a set of encoder-side modules that predict the decoding result in the decoder as part of selecting the optimal coding approach during RDO. Furthermore, the decoder can receive the bitstream 500 and determine the assignment to tile group 521 based on the tile group flag 531, the first tile 532, and the last tile 533. Specifically, both the HRD in the encoder and the HRD in the decoder can select either a raster scan mechanism or a rectangular mechanism based on the tile group flag 531. The HRD and decoder can then use the selected mechanism to determine the assignment of tile 523 to tile group 521 based on the first tile 523 and the last tile 533.
[0093] The following is a specific example of the mechanism described above.
number
[0094] In this example, the tile group flag 531 is shown as rectangular_tile_group_flag and can be used to select either the rectangular mechanism (e.g., an if statement) or the raster scan mechanism (e.g., an else statement). The rectangular mechanism determines the delta value between the first tile and the last tile in the tile group. The number of rows in the tile group is determined by dividing the delta value by the number of columns of tiles in the picture plus 1. The number of columns in the tile group is determined by the delta value modulo the number of columns of tiles in the picture plus 1. Tile assignment can then be determined based on the number of rows and columns in the tile group (e.g., a for loop in an if statement). On the other hand, the raster scan mechanism determines the number of tiles between the first tile and the last tile in the tile group. Since tiles are indexed in raster scan order, the raster scan mechanism can then add the determined number of tiles to the tile group in raster scan order (e.g., a for loop in an else statement).
[0095] Figure 6 is a schematic diagram showing an exemplary picture 600 partitioned into raster scantile group 621. For example, picture 600 can be encoded into a bitstream 500 by, for example, a codec system 200, an encoder 300, and / or a decoder 400, and then decoded. Furthermore, picture 600 can be partitioned to support encoding and decoding according to method 100.
[0096] Picture 600 includes tiles 623 assigned to raster scan tile groups 621, 624, and 625, which may be substantially the same as tile groups 521 and tile 523, respectively. Each tile 623 is assigned to raster scan tile groups 621, 624, and 625 in raster scan order. To clearly indicate the boundaries between raster scan tile groups 621, 624, and 625, each tile group is enclosed in a bold outline. Furthermore, tile group 621 is indicated by a shadow to further distinguish it from the tile group boundaries. It should also be noted that picture 600 may be partitioned into any number of raster scan tile groups 621, 624, and 625. For the sake of clarity, the following description relates to raster scan tile group 621. However, tile 623 is assigned to raster scan tile groups 624 and 625 in a similar manner to raster scan tile group 621.
[0097] As shown, the first tile 623a, the last tile 623b, and all the shaded tiles between the first tile 623a and the last tile 623b are assigned to tile group 621 in the raster scan order. As shown, the mechanism that proceeds according to the raster scan order (e.g., how it operates on the processor) assigns the first tile 623a to tile group 621, and then proceeds to assign each tile 623 to tile group 621 (from left to right) until it reaches the right picture 600 boundary (unless it reaches the last tile 623b). The raster scan order then proceeds to the next row of tile 623 (e.g., from top row to bottom row). In this case, the first tile 623a is in the first row, and therefore the next row is the second row. Specifically, the raster scan order proceeds to the first tile of the second row, which is at the left picture 600 boundary, and then proceeds horizontally from left to right through the second row until it reaches the right picture 600 boundary. The raster scan then moves to the next row, which in this case is the third row, and proceeds with the assignment starting from the first tile of the third row, which is on the left boundary of picture 600. The raster scan then moves horizontally to the right of the third row. This sequence continues until the last tile 623b is reached. At this point, tile group 621 is complete. Any additional tiles 623 below and / or to the right of tile group 621 can be assigned to tile group 625 in a similar manner using the raster scan sequence. Any tiles 623 above and / or to the left of tile group 621 are assigned to tile group 624 in a similar manner.
[0098] Figure 7 is a schematic diagram showing an exemplary picture 700 partitioned into rectangular tile groups 721. For example, the picture 700 can be encoded into a bitstream 500 by, for example, a codec system 200, an encoder 300, and / or a decoder 400, and then decoded. Furthermore, the picture 700 can be partitioned to support encoding and decoding according to method 100.
[0099] Picture 700 includes tiles 723 assigned to rectangular tile group 721, which may be substantially the same as tile group 521 and tile 523, respectively. Tiles 723 assigned to rectangular tile group 721 are shown in Figure 7 as being surrounded by a bold outline. Furthermore, the selected rectangular tile group 721 is shaded to clearly depict the spaces between rectangular tile groups 721. As shown, rectangular tile group 721 includes a set of tiles 723 that form a rectangular shape. It should be noted that a square is a particular case of a rectangle, so rectangular tile group 721 may be a square. As shown, a rectangle has four sides, each side connected to two other sides by a right angle (e.g., a 90° angle). Rectangular tile group 721a includes the first tile 723a and the last tile 723b. The first tile 723a is the upper left corner of rectangular tile group 721a, and the last tile is the lower right corner of rectangular tile group 721a. Tiles 723 contained in or between the rows and columns containing the first tile 723a and the last tile 723b are also assigned to a rectangular tile group 721a, tile by tile. As shown, this method differs from raster scanning. For example, tile 723c, which is between the first tile 723a and the last tile 723b in the raster scan order, is not included in the same rectangular tile group 721a. Rectangular tile groups 721a can be computationally more complex than raster scan tile groups 621 due to the associated geometry. However, rectangular tile groups 721 are more flexible. For example, rectangular tile group 721a may include tiles 723 from different rows, but not all tiles between the first tile 723 and the right boundary of picture 700 (e.g., tile 723c). Rectangular tile group 721a may exclude selected tiles between the left picture boundary and the last tile 723b. For example, tile 723d is excluded from tile group 721a.
[0100] Therefore, the rectangular tile group 721 and the raster scan tile group 621 each have different advantages and may be better suited to different use cases. For example, the raster scan tile group 621 may be more advantageous when the entire picture 600 is displayed, while the rectangular tile group 721 may be more advantageous when only a sub-picture is displayed. However, as described above, when only the first and last tile indices are signaled in the bitstream, different mechanisms may be used to determine which tiles are assigned to a tile group. Thus, a flag indicating which tile group type is used can be used by the decoder or HRD to select the appropriate raster scan or rectangular mechanism. The assignment of tiles to a tile group can then be determined by utilizing the first and last tiles within the tile group.
[0101] By utilizing the above, video coding systems can be improved. Accordingly, this disclosure describes various improvements to tile grouping in video coding. More specifically, this disclosure describes signaling and derivation processes for supporting two different tile group concepts: raster scan-based tile groups and rectangular tile groups. In one example, a flag is used in a parameter set directly or indirectly referenced by the corresponding tile group. The flag specifies which tile group approach is used. The flag can be signaled in a parameter set such as a sequence parameter set, a picture parameter set, or another type of parameter set directly or indirectly referenced by the tile group. In a specific example, the flag may be rectangular_tile_group_flag. In some examples, an instruction having two or more bits may be defined and signaled in a parameter set directly or indirectly referenced by the corresponding tile group. The instruction may specify which tile group approach is used. Using such an instruction, two or more tile group approaches can be supported. The number of bits for signaling the instruction depends on the number of tile group approaches to be supported. In some cases, flags or instructions can be signaled within the tile group header.
[0102] Signaling information indicating the first and last tiles in a tile group may be sufficient to indicate which tiles belong to a raster scan tile group or a rectangular tile group. The derivation of the tiles in a tile group may depend on the tile group approach used (which may be indicated by a flag or instruction), the information of the first tile in the tile group, and the information of the last tile in the tile group. Information for identifying a particular tile may be one of the following: tile index, tile ID (if different from the tile index), CTUs contained in the tile (e.g., the first CTU contained in the tile), or luma samples contained in the tile (e.g., the first luma sample contained in the tile).
[0103] The following is a specific embodiment of the mechanism described above. The picture parameter set RBSP syntax may be as follows: [Table 1]
[0104] The value of tile_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element tile_id_val[i][j] in the PPS, if present, and the syntax elements first_tile_id and last_tile_id in the tile group header referencing the PPS. The value of tile_id_len_minus1 may be within the range of Ceil(Log2(NumTilesInPic)~15, including both ends. When rectangular_tile_group_flag is set to equal 1, it may specify that the tile group referencing the PPS contains one or more tiles that form a rectangular region of the picture. When rectangular_tile_group_flag is set to equal 0, it may specify that the tile group referencing the PPS contains one or more tiles that are consecutive in the raster scan order of the picture.
[0105] The tile group header syntax may be as follows: [Table 2]
[0106] `single_tile_in_tile_group_flag`, when set to 1, may specify that there is only one tile in the tile group. `single_tile_in_tile_group_flag`, when set to 0, may specify that there is more than one tile in the tile group. `first_tile_id` may specify the tile ID of the first tile in the tile group. The length of `first_tile_id` may be `tile_id_len_minus1+1` bits. The value of `first_tile_id` must not be equal to the value of `first_tile_id` of any other coding tile group in the same coding picture. When there is more than one tile group in the picture, the decoding order of the tile groups in the picture may be in ascending order of the values of `first_tile_id`. `last_tile_id` may specify the tile ID of the last tile in the tile group. The length of `last_tile_id` may be `tile_id_len_minus1+1` bits. When it does not exist, the value of `last_tile_id` may be assumed to be equal to `first_tile_id`.
[0107] The variables NumTilesInTileGroup, which specify the number of tiles in a tile group, and TgTileIdx[i], which specifies the tile index of the i-th tile in the tile group, can be derived as follows.
number
[0108] The general tile group data syntax may be as follows: [Table 3]
[0109] Figure 8 is a schematic diagram of an exemplary video coding apparatus 800. The video coding apparatus 800 is suitable for carrying out the examples / embodiments of the disclosure described herein. The video coding apparatus 800 includes a downstream port 820, an upstream port 850, and / or a transceiver unit (Tx / Rx) 810 including a transmitter and / or receiver for communicating upstream and / or downstream data over a network. The video coding apparatus 800 also includes a processor 830 including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 832 for storing data. The video coding apparatus 800 may also include electrical, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components connected to the upstream port 850 and / or downstream port 820 for communication of data over an electrical, optical, or wireless communication network. The video coding device 800 may also include an input and / or output (I / O) device 860 for communicating data to and from the user. The I / O device 860 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 860 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interfacing with such output devices.
[0110] The processor 830 is implemented in hardware and software. The processor 830 may be implemented as one or more CPU chips, cores (e.g., a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 830 communicates with downstream port 820, Tx / Rx 810, upstream port 850, and memory 832. The processor 830 includes a coding module 814. The coding module 814 implements embodiments of the disclosure described herein, such as methods 100, 900, and 1000, which may utilize bitstream 500, picture 600, and / or picture 700. The coding module 814 may also implement any other methods / mechanisms described herein. Furthermore, the coding module 814 may implement a codec system 200, an encoder 300, and / or a decoder 400. For example, the coding module 814 can partition an image into tile groups and / or tiles, the tiles into CTUs, the CTUs into blocks, and encode the blocks when operating as an encoder. Furthermore, the coding module 814 can select a raster scan or a rectangular tile group and signal such selection in the bitstream. The coding module 814 may also signal the first and last tiles to support the decision of assigning tiles to a tile group. When operating as a decoder or HRD, the coding module 814 can determine the type of tile group to be used and determine the tiles assigned to the tile group based on the first and last tiles. Thus, the coding module 814 causes the video coding device 800 to provide additional functionality and / or coding efficiency when partitioning and coding video data.Therefore, the coding module 814 improves the functionality of the video coding device 800 and solves problems specific to video coding technology. Furthermore, the coding module 814 performs conversions of the video coding device 800 to different states. Alternatively, the coding module 814 can be implemented as instructions stored in memory 832 and executed by processor 830 (for example, as a computer program product stored in a non-temporary medium).
[0111] Memory 832 includes one or more memory types such as disks, tape drives, solid drives, read-only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), static random-access memory (SRAM), etc. Memory 832 may be used as an overflow data storage device for storing a program when it is selected for execution, and for storing instructions and data read during program execution.
[0112] Figure 9 is a flowchart of an exemplary method 900 for encoding pictures such as picture 600 and / or 700 into a bitstream such as bitstream 500. Method 900 may be used with an encoder such as a codec system 200, encoder 300, and / or video coding device 800 when performing method 100.
[0113] Method 900 may begin when an encoder receives a video sequence containing multiple pictures and decides to encode the video sequence into a bitstream, for example, based on user input. The video sequence is partitioned into pictures / images / frames for further partitioning before encoding. In step 901, the pictures are partitioned into multiple tiles. Furthermore, the tiles are assigned to multiple tile groups, so that a subset of tiles is assigned to a tile group. In some examples, the tile groups are raster scan tile groups. In other examples, the tile groups are rectangular tile groups.
[0114] In step 903, the flag is encoded into the bitstream. The flag can be set to a first value when the tile group is a raster scan tile group, and to a second value when the tile group is a rectangular tile group. The flag may be encoded into a parameter set of the bitstream. For example, the parameter set into which the flag is encoded may be a sequence parameter set or a picture parameter set. In some examples, the flag is a rectangular tile group flag.
[0115] In step 905, the identifiers of the first tile and the last tile in the tile group are encoded into the bitstream. The identifiers of the first and last tile in the tile group may be used to indicate the tiles assigned to the tile group. In some examples, the identifiers of the first and last tile in the tile group are encoded into the tile group header in the bitstream.
[0116] The flags, the first tile of the tile group, and the last tile of the tile group can be used by the decoder and / or the HRD in the encoder to determine the assignment of tiles to the tile group. When the tile group is a raster scan tile group, as indicated by the flags, the assignment of tiles to the tile group can be determined as follows: The number of tiles between the first tile and the last tile of the tile group can be determined as the number of tiles in the tile group. The assignment of tiles can then be determined based on the number of tiles in the tile group. When the tile group is a rectangular tile group, as indicated by the flags, the assignment of tiles to the tile group can be determined as follows: The delta value between the first tile and the last tile of the tile group can be determined. The number of rows in the tile group can be determined based on the delta value and the number of tile columns in the picture. The number of columns in the tile group can also be determined based on the delta value and the number of tile columns in the picture. The assignment of tiles can then be determined based on the number of rows and the number of columns in the tile group.
[0117] In step 907, the tiles are encoded into a bitstream based on their tile assignments. In step 909, the bitstream may be stored for communication to the decoder.
[0118] Figure 10 is a flowchart of an exemplary method 1000 for decoding pictures such as picture 600 and / or 700 from a bitstream such as bitstream 500. Method 1000 may be used by a decoder such as a codec system 200, a decoder 400, and / or a video coding device 800 when performing method 100. For example, method 1000 may be used in response to method 900.
[0119] Method 1000 may begin, for example, when the decoder begins receiving a bitstream of coding data representing a video sequence as a result of Method 900. In step 1001, the bitstream is received by the decoder. The bitstream contains pictures partitioned into multiple tiles. The tiles are assigned to multiple tile groups, and thus subsets of tiles are assigned to tile groups. In some examples, the tile groups are raster scan tile groups. In other examples, the tile groups are rectangular tile groups.
[0120] In step 1003, a flag is obtained from the bitstream parameter set. When the flag is set to the first value, the tile group is determined to be a raster scan tile group. When the flag is set to the second value, the tile group is determined to be a rectangle tile group. For example, the parameter containing the flag may be a sequence parameter set or a picture parameter set. In some examples, the flag is a rectangle tile group flag.
[0121] In step 1005, the identifiers of the first tile and the last tile in the tile group are obtained to support the determination of which tile is assigned to the tile group. In some examples, the identifiers of the first and last tile in the tile group are obtained from the tile group header in the bitstream.
[0122] In step 1007, the assignment of tiles to a tile group is determined based on whether the tile group is a raster scan tile group or a rectangular tile group. For example, the flag, the first tile of the tile group, and the last tile of the tile group can be used to determine the assignment of tiles for the tile group. When the tile group is a raster scan tile group, as indicated by the flag, the assignment of tiles to the tile group can be determined as follows: The number of tiles between the first tile of the tile group and the last tile of the tile group can be determined as the number of tiles in the tile group. The assignment of tiles can then be determined based on the number of tiles in the tile group. When the tile group is a rectangular tile group, as indicated by the flag, the assignment of tiles to the tile group can be determined as follows: The delta value between the first tile of the tile group and the last tile of the tile group can be determined. The number of rows in the tile group can be determined based on the delta value and the number of tile columns in the picture. The number of columns in the tile group can also be determined based on the delta value and the number of tile columns in the picture. The assignment of tiles can then be determined based on the number of rows in the tile group and the number of columns in the tile group.
[0123] In step 1009, the tiles are decoded to generate decoded tiles based on the tile assignment to the tile group. A reconstructed video sequence may also be generated for display based on the decoded tiles.
[0124] Figure 11 is a schematic diagram of an exemplary system 1100 for coding a video sequence of pictures, such as picture 600 and / or 700, within a bitstream such as bitstream 500. System 1100 may be implemented by an encoder and decoder such as codec system 200, encoder 300, decoder 400, and / or video coding device 800. Furthermore, system 1100 may be used when implementing methods 100, 900, and / or 1000.
[0125] System 1100 includes a video encoder 1102. The video encoder 1102 includes a partition module 1101 for partitioning a picture into multiple tiles. The video encoder 1102 further includes an inclusion module 1103 for including the number of tiles in a tile group. The video encoder 1102 further includes an encoding module 1105 for encoding a flag set to a first value when the tile group is a raster scan tile group and to a second value when the tile group is a rectangular tile group, the flag being encoded into a bitstream parameter set, and encoding the tiles into a bitstream based on the tile group. The video encoder 1102 further includes a storage module 1107 for storing the bitstream for communication toward a decoder. The video encoder 1102 further includes a transmission module 1109 for transmitting the bitstream to support determining the type of tile group to be used and the tiles to be included in the tile group. The video encoder 1102 may be further configured to perform any of the steps of method 900.
[0126] System 1100 also includes a video decoder 1110. The video decoder 1110 includes a receive module 1111 that receives a bitstream containing a picture partitioned into multiple tiles, where the number of tiles is included in the tile group. The video decoder 1110 further includes an acquire module 1113 that obtains a flag from a parameter set of the bitstream. The video decoder 1110 further includes a determination module 1115 that determines that the tile group is a raster scan tile group when the flag is set to a first value, determines that the tile group is a rectangular tile group when the flag is set to a second value, and determines tile inclusion for the tile group based on whether the tile group is a raster scan tile group or a rectangular tile group. The video decoder 1110 further includes a decode module 1117 that decodes the tiles and generates decoded tiles based on the tile group. The video decoder 1110 further includes a generate module 1119 that generates a reconstructed video sequence for display based on the decoded tiles. The video decoder 1110 may be further configured to perform any of the steps of method 1000.
[0127] When there are no intermediary components between the first and second components, except for lines, traces, or other media, the first component is directly connected to the second component. When there are intermediary components other than lines, traces, or other media between the first and second components, the first component is indirectly connected to the second component. The term "connected" and its variations include both direct and indirect connections. The use of the term "about" means a range including ±10% of the following number, unless otherwise specified.
[0128] Furthermore, it should be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of the steps in such methods should be understood as merely an example. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in methods according to various embodiments of this disclosure.
[0129] While several embodiments are provided in this disclosure, it should be understood that the systems and methods of the disclosure can be implemented in many other specific forms without departing from the spirit or scope of this disclosure. The examples of the invention should be considered for illustrative purposes only and not limiting, and are not intended to limit the details given herein. For example, various elements or components may be combined or integrated into another system, or certain functions may be omitted or not implemented.
[0130] Furthermore, the technologies, systems, subsystems, and methods described and shown in various embodiments may be combined with or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. Other examples of modifications, substitutions, and alterations will be seen by those skilled in the art and may be made without departing from the spirit and scope disclosed herein. [Explanation of Symbols]
[0131] 101 Input Video Signals 103 Block Partition 105-block compression 107 Filtering 109 bitstream 111 Determine the partition 113 Block Decryption 115 Filtering 117 Output video signal
Claims
1. A device for storing video encoded data, wherein the device is The picture is partitioned into multiple tiles, and some of the multiple tiles are included in at least one tile group. A flag in the parameter set is set, the flag being rectangular_tile_group_flag, the flag indicating that when the flag is set to a first value, the at least one tile group is a raster scan tile group, and when the flag is set to a second value, the at least one tile group is a rectangular tile group, the raster scan tile group includes tiles that are consecutive in the raster scan order of the picture, and the rectangular tile group includes tiles that collectively form a rectangular region of the picture. The system generates encoded video data, the encoded data having a bitstream structure including data representing the picture partitioned into the plurality of tiles, the parameter set, and the flags in the parameter set. It is configured in such a way, The aforementioned device is A storage module configured to store the encoded data of the video, A transmission module configured to transmit the encoded data to another device, A device that includes this.
2. The apparatus according to claim 1, further comprising a receiving module configured to receive information from another device.
3. The apparatus according to claim 1 or 2, wherein the parameter set is a picture parameter set.
4. The apparatus according to claim 1 or 2, wherein the first value is 0 and the second value is 1.
5. A method for storing encoded video data, wherein the method is performed by an apparatus, and the apparatus The picture is partitioned into multiple tiles, and some of these tiles are included in at least one tile group. A flag in the parameter set is set, the flag being rectangular_tile_group_flag, the flag indicating that when the flag is set to a first value, the at least one tile group is a raster scan tile group, and when the flag is set to a second value, the at least one tile group is a rectangular tile group, the raster scan tile group includes tiles that are consecutive in the raster scan order of the picture, and the rectangular tile group includes tiles that collectively form a rectangular region of the picture. The encoded data of the video is stored, and the encoded data has a bitstream structure including data representing the picture partitioned into a plurality of tiles, the parameter set, and the flags in the parameter set. The encoded data is transmitted to another device. A method that includes the act of doing so.
6. The method according to claim 5, further comprising the device receiving information from another device.
7. The method according to claim 5 or 6, wherein the parameter set is a picture parameter set.
8. The method according to claim 5 or 6, wherein the first value is 0 and the second value is 1.