Subpicture size in video coding

By allowing sub-image width and height to be less of a multiple of CTU at image boundaries and indicating sub-image layout information in SPS, the decoding errors and resource waste caused by sub-image segmentation are resolved, achieving more efficient video decoding.

CN122120439APending Publication Date: 2026-05-29HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2020-01-09
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing video decoding technologies suffer from sub-image size constraints when processing image segmentation, leading to incorrect sub-image operation in many image layouts, resulting in decoding errors and high resource utilization, especially when image boundaries are not integer multiples of the CTU size.

Method used

By allowing the width and height of sub-images to no longer be forced to be multiples of the CTU size when they are at the right or bottom edge of the image, combined with indicating sub-image layout information in SPS, including position and size, flexible sub-image segmentation and independent extraction are supported, reducing resource utilization.

Benefits of technology

It improves the decoding efficiency of encoders and decoders, reduces the utilization of network, memory and processing resources, reduces the possibility of decoding errors, and supports flexible image layout and region of interest applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120439A_ABST
    Figure CN122120439A_ABST
Patent Text Reader

Abstract

A video coding mechanism is disclosed. The mechanism includes receiving a bitstream including one or more sub-pictures from a picture partitioned such that a sub-picture width included by each sub-picture is an integer multiple of a coding tree unit (CTU) size when a right boundary included by each sub-picture does not coincide with a right boundary of the picture. The bitstream is parsed to obtain the one or more sub-pictures. The one or more sub-pictures are decoded to create a video sequence. The video sequence is transmitted for display.
Need to check novelty before this filing date? Find Prior Art

Description

Related applications cross-application

[0001] This application is a divisional application. The original application has the application number 202080008756.4 and the original application date is January 9, 2020. The entire contents of the original application are incorporated herein by reference.

[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 790,207, filed January 9, 2019, entitled "Sub-Pictures in Video Coding," which is incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to video decoding, and specifically to the management of sub-images in video decoding. Background Technology

[0004] Even with shorter videos, a large amount of video data needs to be described, which can be challenging when the data needs to be transmitted over bandwidth-constrained communication networks or otherwise. Therefore, video data is typically compressed before transmission over modern telecommunications networks. Video size can also be an issue when storing video on storage devices due to potentially limited memory resources. Video compression devices typically use software and / or hardware at the source side to decode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device used to decode the video data. Given limited network resources and the growing demand for higher video quality, there is a need to improve compression and decompression techniques that can increase compression ratios with minimal impact on image quality. Summary of the Invention

[0005] In one embodiment, the present invention includes a method implemented in a decoder, the method comprising: a receiver of the decoder receiving a bitstream comprising one or more sub-images obtained by segmenting an image, such that when the right boundary of a first sub-image coincides with the right boundary of the image, the width of the first sub-image comprises an incomplete coding tree unit (CTU); a processor of the decoder parsing the bitstream to obtain the one or more sub-images; the processor decoding the one or more sub-images to create a video sequence; and the processor sending the video sequence for display. Some video systems may restrict the height and width of a sub-image to multiples of the CTU size. However, the height and width of an image may not be multiples of the CTU size. Therefore, sub-image size constraints cause sub-images to operate incorrectly in many image layouts. In the disclosed example, the sub-image width and sub-image height are constrained to multiples of the CTU size. However, these constraints need to be removed when the sub-image is located at the right boundary or the bottom boundary of the image, respectively. By allowing the height and width of the lower and right sub-images to be non-multiples of the CTU size, respectively, the sub-images can be used with any image without causing decoding errors. This enhances the capabilities of both the encoder and decoder. Furthermore, these enhancements allow the encoder to decode images more efficiently, reducing network resource utilization, memory resource utilization, and / or processing resource utilization for both the encoder and decoder.

[0006] Optionally, according to any of the above aspects, in another implementation of said aspect, when the lower boundary of the second sub-image does not coincide with the lower boundary of the image, the height of the second sub-image includes an integer number of complete CTUs.

[0007] Optionally, according to any of the above aspects, in another implementation of said aspect, when the right boundary of the third sub-image does not coincide with the right boundary of the image, the width of the sub-image included in the third sub-image includes an integer number of complete CTUs.

[0008] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, when the lower boundary of the fourth sub-image coincides with the lower boundary of the image, the height of the sub-image included in the fourth sub-image includes an incomplete CTU.

[0009] In one embodiment, the invention includes a method implemented in a decoder, the method comprising: a receiver of the decoder receiving a bitstream comprising one or more sub-images obtained by segmenting an image, such that the width of each sub-image is an integer multiple of the coding tree unit (CTU) size when the right boundary of each sub-image does not coincide with the right boundary of the image; a processor of the decoder parsing the bitstream to obtain the one or more sub-images; the processor decoding the one or more sub-images to create a video sequence; and the processor sending the video sequence for display. Some video systems may restrict the height and width of sub-images to multiples of the CTU size. However, the height and width of an image may not be multiples of the CTU size. Therefore, sub-image size constraints cause sub-images to operate incorrectly in many image layouts. In the disclosed example, the sub-image width and sub-image height are constrained to multiples of the CTU size. However, these constraints need to be removed when the sub-image is located at the right boundary or the bottom boundary of the image, respectively. By allowing the height and width of the lower and right sub-images to be non-multiples of the CTU size, respectively, the sub-images can be used with any image without causing decoding errors. This enhances the capabilities of both the encoder and decoder. Furthermore, these enhancements allow the encoder to decode images more efficiently, reducing network resource utilization, memory resource utilization, and / or processing resource utilization for both the encoder and decoder.

[0010] Alternatively, according to any of the above aspects, in another implementation of said aspect, when the lower boundary of each sub-image does not coincide with the lower boundary of the image, the height of each sub-image is an integer multiple of the CTU size.

[0011] Alternatively, according to any of the above aspects, in another implementation of said aspect, when the right boundary of each sub-image coincides with the right boundary of the image, the width of at least one sub-image in the sub-image is not an integer multiple of the CTU size.

[0012] Optionally, according to any of the foregoing aspects, in another implementation of said aspect, when the lower boundary of each sub-image coincides with the lower boundary of the image, the height of at least one sub-image in the sub-images is not an integer multiple of the CTU size.

[0013] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the image includes an image width that is not an integer multiple of the CTU size.

[0014] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the image includes an image height that is not an integer multiple of the CTU size.

[0015] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the CTU size is measured in units of luminance pixels.

[0016] In one embodiment, the invention includes a method implemented in an encoder, the method comprising: a processor of the encoder segmenting an image into a plurality of sub-images such that the width of each sub-image is an integer multiple of the CTU size when the right boundary of each sub-image does not coincide with the right boundary of the image; the processor encoding one or more sub-images into a bitstream; and storing the bitstream in the memory of the encoder for transmission to a decoder. Some video systems may restrict the height and width of sub-images to multiples of the CTU size. However, the height and width of an image may not be multiples of the CTU size. Therefore, sub-image size constraints cause sub-images to operate incorrectly in many image layouts. In the disclosed example, the sub-image width and sub-image height are constrained to multiples of the CTU size. However, these constraints need to be removed when the sub-images are located at the right boundary or the bottom boundary of the image, respectively. By allowing the height and width of the lower and right sub-images to be non-multiples of the CTU size, respectively, the sub-images can be used with any image without causing decoding errors. This enhances the functionality of the encoder and decoder. In addition, the enhanced functionality enables the encoder to decode images more efficiently, which reduces the utilization of network resources, memory resources, and / or processing resources of the encoder and decoder.

[0017] Alternatively, according to any of the above aspects, in another implementation of said aspect, when the lower boundary of each sub-image does not coincide with the lower boundary of the image, the height of each sub-image is an integer multiple of the CTU size.

[0018] Alternatively, according to any of the above aspects, in another implementation of said aspect, when the right boundary of each sub-image coincides with the right boundary of the image, the width of at least one sub-image in the sub-image is not an integer multiple of the CTU size.

[0019] Optionally, according to any of the foregoing aspects, in another implementation of said aspect, when the lower boundary of each sub-image coincides with the lower boundary of the image, the height of at least one sub-image in the sub-images is not an integer multiple of the CTU size.

[0020] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the image includes an image width that is not an integer multiple of the CTU size.

[0021] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the image includes an image height that is not an integer multiple of the CTU size.

[0022] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the CTU size is measured in units of luminance pixels.

[0023] In one embodiment, the present invention includes a video decoding device comprising: a processor, a memory, a receiver coupled to the processor, and a transmitter, the processor, memory, receiver, and transmitter being configured to perform the method according to any one of the foregoing aspects.

[0024] In one embodiment, the present invention includes a non-transitory computer-readable medium comprising a computer program product for a video decoding apparatus, the computer program product including computer-executable instructions stored in the non-transitory computer-readable medium, which, when executed by a processor, cause the video decoding apparatus to perform the method according to any of the foregoing aspects.

[0025] In one embodiment, the present invention includes a decoder comprising: a receiving module for receiving a bitstream comprising one or more sub-images obtained by segmenting an image, such that the width of each sub-image is an integer multiple of the CTU size when the right boundary of each sub-image does not coincide with the right boundary of the image; a parsing module for parsing the bitstream to obtain the one or more sub-images; a decoding module for decoding the one or more sub-images to create a video sequence; and a sending module for sending the video sequence for display.

[0026] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the decoder is also used to perform the method described in any of the foregoing aspects.

[0027] In one embodiment, the present invention includes an encoder comprising: a segmentation module for segmenting an image into multiple sub-images, such that when the right boundary of each sub-image does not coincide with the right boundary of the image, the width of each sub-image is an integer multiple of the CTU size; an encoding module for encoding one or more of the sub-images into a bitstream; and a storage module for storing the bitstream, the bitstream being sent to a decoder.

[0028] Alternatively, according to any of the foregoing aspects, in another implementation of said aspect, the encoder is also used to perform the method described in any of the foregoing aspects.

[0029] For clarity, any of the above embodiments may be combined with any of the other above embodiments to create new embodiments within the scope of the present invention.

[0030] These and other features will become clearer from the following detailed description taken in conjunction with the accompanying drawings and claims. Attached Figure Description

[0031] To gain a more complete understanding of the invention, reference is made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein the same reference numerals denote the same parts.

[0032] Figure 1 A flowchart illustrating an exemplary method for decoding video signals.

[0033] Figure 2 This is a schematic diagram of an example encoding and decoding (codec) system used for video decoding.

[0034] Figure 3 This is a schematic diagram of an exemplary video encoder.

[0035] Figure 4 This is a schematic diagram of an exemplary video decoder.

[0036] Figure 5 This is a schematic diagram of an exemplary bitstream and sub-bitstream extracted from a bitstream.

[0037] Figure 6 This is a schematic diagram of an example image segmented into sub-images.

[0038] Figure 7 This is a schematic diagram of an exemplary mechanism for associating stripes with sub-image layouts.

[0039] Figure 8 This is a schematic diagram of another exemplary image segmented into sub-images.

[0040] Figure 9 This is a schematic diagram of an exemplary video decoding device.

[0041] Figure 10 A flowchart illustrating an exemplary method for encoding the bitstream of a sub-image using adaptive size constraints.

[0042] Figure 11 A flowchart illustrating an exemplary method for decoding the bitstream of a sub-image using adaptive size constraints.

[0043] Figure 12A schematic diagram of an exemplary system for indicating the bitstream of a sub-image using adaptive size constraints. Detailed Implementation

[0044] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The invention should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.

[0045] This paper uses various abbreviations, such as: coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded videosequence (CVS), Joint Video Experts Team (JVET), motion constrained tile set (MCTS), maximum transfer unit (MTU), network abstraction layer (NAL), picture order count (POC), raw byte sequence payload (RBSP), sequence parameter set (SPS), versatile video coding (VVC), and working draft (WD).

[0046] Many video compression techniques can be used to reduce the size of video files while minimizing data loss. For example, video compression techniques may include performing spatial (e.g., intra-frame) prediction and / or temporal (e.g., inter-frame) prediction to reduce or remove data redundancy in a video sequence. For block-based video decoding, a video slice (e.g., a video image or a portion of a video image) can be divided into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or decoding nodes. Intra-frame decoding (I) of an image: Video blocks in a slice are decoded using spatial prediction for reference pixels in adjacent slices within the same image. Inter-frame decoding (P or B) of an image: Video blocks in a slice may use spatial prediction for reference pixels in adjacent slices within the same image, or temporal prediction for reference pixels in other reference images. An image may be referred to as a frame, and a reference image may be referred to as a reference frame. Spatial or temporal prediction produces prediction blocks representing image slices. The residual data represents the pixel difference between the original image block and the predicted block. Therefore, inter-frame decoded blocks are encoded based on the motion vector pointing to the reference pixel block that constitutes the predicted block, and the residual data representing the difference between the decoded block and the predicted block. Intra-frame decoded blocks are encoded based on the intra-frame decoding mode and the residual data. For further compression, the residual data can be transformed from the pixel domain to the transform domain to generate residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, can be scanned to produce a one-dimensional transform coefficient vector. Entropy decoding can be applied to achieve greater compression. This video compression technique is discussed in detail below.

[0047] To ensure accurate decoding of encoded video, video is encoded and decoded according to the corresponding video decoding standards. These standards include ITU-T H.261, ISO / IEC Motion Picture Experts Group (MPEG) Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has begun developing a video coding standard called Versatile Video Coding (VVC). VVC is included in a Working Draft (WD), which includes JVET-L1001-v9.

[0048] To decode video images, the image is first segmented, and the resulting parts are decoded into the bitstream. There are various image segmentation schemes. For example, the image can be segmented into regular stripes, correlated stripes, tiles, and / or segmented according to Wavefront Parallel Processing (WPP). For simplicity, HEVC restricts the encoder so that only regular stripes, correlated stripes, tiles, WPP, and combinations thereof can be used when segmenting stripes into CTB groups for video decoding. This segmentation can be used to support Maximum Transfer Unit (MTU) size matching, parallel processing, and reduce end-to-end latency. MTU represents the maximum amount of data that can be sent in a single message. If the message payload exceeds the MTU, the payload is divided into two messages through a process called segmentation.

[0049] Regular stripes, also simply stripes, are segmented portions of an image that can be reconstructed independently of other regular stripes within the same image, despite some interdependencies due to cyclic filtering operations. Each regular stripe is encapsulated within its own network abstraction layer (NAL) unit and transmitted. Furthermore, intra-frame prediction (intra-pixel prediction, motion information prediction, decoding pattern prediction) and entropy decoding dependencies across stripe boundaries can be disabled to support independent reconstruction. This independent reconstruction supports parallelization. For example, parallelization based on regular stripes uses minimal inter-processor or inter-core communication. However, since each regular stripe is independent, each stripe is associated with a separate stripe header. Using regular stripes incurs significant decoding overhead due to the bit cost of each stripe header and the lack of prediction across stripe boundaries. Additionally, regular stripes can be used to support MTU size matching requirements. Specifically, since regular stripes are encapsulated in separate NAL units and can be decoded independently, each regular stripe should be smaller than the MTU in the MTU scheme to avoid splitting the stripe into multiple packets. Therefore, in order to achieve parallelization and MTU size matching, the stripe layout in the image may be contradictory.

[0050] Correlated stripes are similar to regular stripes, but they have a shortened strip header and allow segmentation of image tree block boundaries without affecting intra-frame prediction. Therefore, correlated stripes divide regular stripes into multiple NAL units, which reduces end-to-end latency by sending a portion of the regular stripe before the entire regular stripe is encoded.

[0051] A block is a segmented portion of an image created by horizontal and vertical boundaries, which define the blocks' columns and rows. Blocks can be decoded in raster scan order (right-to-left, top-to-bottom). The CTB scan order is performed within a block. Therefore, the CTB in the first block is decoded in raster scan order before proceeding to the CTB in the next block. Similar to regular striping, blocks eliminate intra-frame prediction dependencies and entropy decoding dependencies. However, individual NAL units may not include blocks; therefore, blocks may not be used for MTU size matching. Each block can be processed by a single processor / core. Inter-processor / inter-core communication (used for intra-frame prediction between processing units, where processing units decode adjacent blocks) can be limited to sending a shared stripe header (when adjacent blocks are in the same stripe) and sharing cyclic filtering-related information for reconstructed pixels and metadata. When a stripe includes multiple blocks, the entry point byte offset for each block can be signaled in the stripe header, in addition to the first entry point offset within the stripe. For each stripe and block, at least one of the following conditions must be met: (1) all decoder blocks in the stripe belong to the same block; (2) all decoder blocks in the block belong to the same stripe.

[0052] In WPP, the image is segmented into single-line CTBs. Entropy decoding and prediction mechanisms can utilize data from CTBs in other lines. Parallel processing is supported by decoding CTB lines in parallel. For example, the current line can be decoded in parallel with the previous line. However, decoding the current line is delayed by two CTBs compared to decoding the previous few lines. This delay ensures that data related to the CTBs above and to the right of the current CTB in the current line is available before decoding the current CTB. When represented graphically, this approach is represented as wavefronts. If the image includes CTB lines, this staggered start allows for parallel processing by as many processors / cores as possible. Intra-frame prediction is possible through inter-processor / inter-core communication, as intra-frame prediction is supported between adjacent tree block lines within the image. WPP segmentation does not consider NAL unit size. Therefore, WPP does not support MTU size matching. However, regular striping can be used in conjunction with WPP to achieve MTU size matching as needed, which requires some decoding overhead.

[0053] Tiles can also include motion-constrained tilesets. A motion-constrained tileset (MCTS) is a set of tiles used to constrain relevant motion vectors to point to full-pixel positions within the MCTS and to fractional-pixel positions, which only require interpolation of full-pixel positions within the MCTS. Furthermore, motion vector candidates are not allowed for temporal motion vector prediction derived from blocks outside the MCTS. This allows each MCTS to be decoded independently without including tiles within the MCTS. Supplemental enhancement information (SEI) messages for temporal MCTS can be used to indicate the presence of MCTSs in the bitstream and to indicate MCTSs. MCTS SEI messages provide supplemental information that can be used for MCTS sub-bitstream extraction (represented as part of the SEI message semantics) to generate a consistent bitstream for the MCTS. This information includes multiple extraction information sets, each defining multiple MCTSs, and includes raw byte sequence payload (RBSP) bytes of the replacement video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) used in the MCTS sub-stream extraction process. When extracting sub-streams according to the MCTS sub-stream extraction process, the parameter sets (VPS, SPS, PPS) can be rewritten or replaced, and the slice header can be updated because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) can use different values ​​in the extracted sub-streams.

[0054] An image can also be segmented into one or more sub-images. A sub-image is a set of rectangular tile groups / stripes, starting with a tile group whose `tile_group_address` is equal to 0. Each sub-image can reference a separate PPS and therefore can have its own tile segmentation. During decoding, sub-images can be processed like the main image. The reference sub-image used to decode the current sub-image is generated by extracting a region juxtaposed with the current sub-image from the reference image in the decoded image buffer. The extracted region is considered the decoded sub-image. Inter-frame prediction can occur between sub-images of the same size and location within the image. A tile group, also called a stripe, is a sequence of related tiles in an image or sub-image. Several terms can be derived to determine the position of a sub-image within the image. For example, each current sub-image can be located in the next unoccupied position in the image according to the CTU raster scan order, a position large enough to include the current sub-image within the image boundary.

[0055] Furthermore, image segmentation can be based on image-level tiles and sequence-level tiles. Sequence-level tiles can include the functionality of MCTS and can be implemented as sub-images. For example, an image-level tile can be defined as a rectangular region of decoded tree blocks within a specific tile column and row in an image. A sequence-level tile can be defined as a set of rectangular regions comprising decoded tree blocks in different frames, where each rectangular region also includes one or more image-level tiles, and the set of rectangular regions of decoded tree blocks can be decoded independently from any other similar set of rectangular regions. A sequence-level tile groupset (STGPS) is a set of such sequence-level tiles. STGPS can be indicated in the non-video coding layer (VCL) NAL unit using the associated identifier (ID) in the NAL unit header.

[0056] The sub-image-based segmentation schemes described above may be relevant to certain issues. For example, when using sub-images, they can be segmented into tiles to support parallel processing. The tile segmentation of sub-images for parallel processing can vary depending on the image (e.g., for load balancing in parallel processing) and can therefore be managed at the image level (e.g., in PPS). However, sub-image segmentation (dividing an image into sub-images) can be used to support regions of interest (ROIs) and sub-image-based image access. In this case, the indicative efficiency of sub-images or MCTS in PPS is not high.

[0057] In another example, when any sub-image in an image is decoded into a time-motion-constrained sub-image, all sub-images in the image can be decoded into time-motion-constrained sub-images. This image segmentation can be limiting. For example, decoding sub-images into time-motion-constrained sub-images may reduce decoding efficiency but benefit other functionalities. However, in region-of-interest (ROI) applications, typically only one or a few sub-images are time-motion-constrained. Therefore, the decoding efficiency of the remaining sub-images is reduced without providing any real benefit.

[0058] In another example, the syntax element used to represent the size of a subimage can be represented in units of luminance CTU. Therefore, the width and height of the subimage should both be integer multiples of CtbSizeY. This mechanism of representing subimage width and height can lead to various problems. For example, subimage segmentation only applies to images where the image width and / or image height is an integer multiple of CtbSizeY. This makes subimage segmentation unusable for images whose included dimensions are not integer multiples of CtbSizeY. If subimage segmentation is performed on the image width and / or height when the image size is not an integer multiple of CtbSizeY, the derivation of the subimage width and / or subimage height in luminance pixels for the rightmost and bottommost subimages will be incorrect. This incorrect derivation will cause some decoding tools to produce incorrect results.

[0059] In another example, the position of a sub-image within the image may not be indicated. Instead, the position is derived using the following rule: The current sub-image is located in the next unoccupied position in the image according to the CTU raster scan order, a position large enough to include the sub-image within the image boundary. In some cases, deriving the sub-image position in this way can lead to errors. For example, if a sub-image is lost during transmission, the positions of other sub-images will be incorrectly derived, and decoded pixels will be placed in the wrong locations. The same problem occurs when sub-images arrive in the wrong order.

[0060] In another example, decoding a sub-image can extract a juxtaposed sub-image from a reference image. This can increase processor complexity and load, as well as memory resource utilization.

[0061] In another example, when a subimage is specified as a subimage with a temporal motion constraint, the loop filter that traverses the subimage boundaries is disabled. This occurs regardless of whether the loop filter that traverses the block boundaries is enabled. This constraint can be overly restrictive and may cause visual artifacts in video images that use multiple subimages.

[0062] In another example, the relationship between SPS, STGPS, PPS, and the tile group header is as follows: STGPS references SPS, PPS references STGPS, and the tile group header / strip header references PPS. However, STGPS and PPS should be data-independent, not PPS referencing STGPS. This arrangement also allows all tile groups of the same image to not reference the same PPS.

[0063] In another example, each STGPS may include the IDs of the four sides of a sub-image. These IDs are used to identify sub-images sharing the same boundary so that their relative spatial relationships can be defined. However, in some cases, this information may be insufficient to derive the location and size information of sequence-level block sets. In other cases, the indication of location and size information may be redundant.

[0064] In another example, the STGPS ID can be indicated using 8 bits in the NAL cell header of the VCL NAL cell. This facilitates sub-image extraction. However, this indication may unnecessarily increase the length of the NAL cell header. Another issue is that unless the sequence-level block sets are constrained to prevent overlap, one block set can be associated with multiple sequence-level block sets.

[0065] This paper discloses various mechanisms for addressing one or more of the aforementioned problems. In a first example, the sub-image layout information is included in the SPS instead of the PPS. The sub-image layout information includes the sub-image position and sub-image size. The sub-image position is the offset between the top-left pixel of the sub-image and the top-left pixel of the image. The sub-image size is the height and width of the sub-image measured in luminance pixels. As mentioned above, some systems include chunking information in the PPS because chunking may vary based on the image. However, sub-images can be used to support ROI applications and sub-image-based access. These functionalities do not change on a per-image basis. Furthermore, a video sequence may include a single SPS (or one SPS per video segment) and may include at most one PPS per image. Placing the sub-image layout information in the SPS ensures that the layout is indicated only once for the sequence / segment, rather than for each PPS. Therefore, indicating the sub-image layout in the SPS improves decoding efficiency, thereby reducing the network resource utilization, memory resource utilization, and / or processing resource utilization of the encoder and decoder. In addition, some systems have sub-image information derived by the decoder. Indicating sub-image information reduces the likelihood of errors when messages are lost and supports other functions for extracting sub-images. Therefore, indicating the sub-image layout in SPS enhances the functionality of the encoder and / or decoder.

[0066] In the second example, the sub-image width and height are constrained to multiples of the CTU size. However, these constraints need to be removed when the sub-images are located at the right or bottom edge of the image, respectively. As mentioned above, some video systems may restrict the height and width of sub-images to multiples of the CTU size. This causes sub-images to operate incorrectly in multiple image layouts. By allowing the height and width of the lower and right sub-images to be non-multiples of the CTU size, respectively, the sub-images can be used with any image without causing decoding errors. This enhances the capabilities of both the encoder and decoder. Furthermore, this enhancement allows the encoder to decode images more efficiently, reducing network resource utilization, memory resource utilization, and / or processing resource utilization for both the encoder and decoder.

[0067] In the third example, the sub-images are constrained to cover images without gaps or overlaps. As mentioned above, some video decoding systems allow sub-images to include gaps and overlaps. This allows chunks / strips to be associated with multiple sub-images. If the encoder allows this, the decoder must be built to support this decoding scheme, even if the decoding scheme is rarely used. By disallowing sub-image gaps and overlaps, the complexity of the decoder can be reduced because the decoder does not need to consider potential gaps and overlaps when determining the size and position of the sub-images. Furthermore, disallowing sub-image gaps and overlaps reduces the complexity of the rate distortion optimization (RDO) process in the encoder, because the encoder can ignore gaps and overlaps when selecting encodings for the video sequence. Therefore, avoiding gaps and overlaps can reduce the memory resource utilization and / or processing resource utilization of both the encoder and decoder.

[0068] In the fourth example, a flag can be indicated in the SPS to specify when a subimage is a temporally motion-constrained subimage. As mentioned above, some systems may set all subimages as temporally motion-constrained subimages or completely disallow the use of temporally motion-constrained subimages. Such temporally motion-constrained subimages provide independent extraction capabilities but reduce decoding efficiency. However, in region-of-interest (ROI) based applications, the ROI should be decoded for independent extraction, while regions outside the ROI do not require this capability. Therefore, the decoding efficiency of the remaining subimages is reduced without providing any practical benefit. Therefore, this flag allows the existence of temporally motion-constrained subimages that provide independent extraction capabilities, as well as non-motion-constrained subimages, to improve decoding efficiency when independent extraction is not required. Thus, this flag enhances functionality and / or improves decoding efficiency, which reduces the network resource utilization, memory resource utilization, and / or processing resource utilization of the encoder and decoder.

[0069] In the fifth example, the SPS indicates the complete set of sub-image IDs, and the strip header includes the sub-image ID, which represents the sub-image that includes the corresponding strip. As mentioned above, some systems indicate the position of a sub-image relative to other sub-images. This can cause problems if a sub-image is missing or extracted in isolation. By specifying each sub-image with an ID, the sub-image can be located and resized without referencing other sub-images. This, in turn, supports error correction and applications that extract only a portion of the sub-images and avoid sending other sub-images. The SPS can send a complete list of all sub-image IDs along with the corresponding size information. Each strip header can include a sub-image ID, which represents the sub-image that includes the corresponding strip. In this way, sub-images and corresponding strips can be extracted and located without referencing other sub-images. Therefore, the sub-image ID enhances functionality and / or improves decoding efficiency, which reduces the network resource utilization, memory resource utilization, and / or processing resource utilization of the encoder and decoder.

[0070] In the sixth example, the level of each sub-image is indicated. In some video decoding systems, the level of an image is indicated. The level represents the hardware resources required to decode the image. As mentioned above, in some cases, different sub-images may have different functions and therefore can be processed differently during the decoding process. Therefore, image-based levels may be disadvantageous for decoding some sub-images. Therefore, the present invention includes a level for each sub-image. In this way, each sub-image can be decoded independently of other sub-images, without unnecessarily overloading the decoder by setting excessively high decoding requirements for sub-images based on uncomplicated mechanisms. The indicated sub-image level information enhances functionality and / or improves decoding efficiency, which reduces the network resource utilization, memory resource utilization, and / or processing resource utilization of the encoder and decoder.

[0071] Figure 1 This is a flowchart of an exemplary method 100 for decoding a video signal. Specifically, the video signal is encoded at the encoder side. The encoding process compresses the video signal using various mechanisms to reduce the video file size. A smaller file size helps to send the compressed video file to the user while reducing associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process typically corresponds to the encoding process to ensure consistent reconstruction of the video signal by the decoder.

[0072] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. In another example, the video file may be captured by a video capture device (e.g., a video camera) and encoded to support live video streaming. The video file may include audio and video components. The video components consist of a series of image frames that, when viewed sequentially, produce a visual effect of motion. These frames include pixels represented according to light (referred herein to as luminance components (or luminance pixels)) and color (referred to as chrominance components (or color pixels)). In some examples, the frames may also include depth values ​​to support three-dimensional viewing.

[0073] In step 103, the video is segmented into blocks. Segmentation involves subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), frames can first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64×64 pixels). A CTU includes luma pixels and chroma pixels. CTUs can be divided into blocks using the coding tree, and these blocks can then be subdivided recursively until a configuration supporting further encoding is obtained. For example, the luma component of a frame can be subdivided until the blocks contain relatively uniform luma values. Similarly, the chroma component of a frame can be subdivided until the blocks contain relatively uniform chroma values. Therefore, the segmentation mechanism varies depending on the content of the video frame.

[0074] In step 105, various compression mechanisms are used to compress the image patches segmented in step 103. For example, inter-frame prediction and / or intra-frame prediction can be used. Inter-frame prediction aims to take advantage of the fact that objects in common scenes tend to appear in consecutive frames. Therefore, there is no need to repeatedly describe the blocks depicting objects in the reference frame in adjacent frames. Specifically, an object (such as a table) may maintain a constant position in multiple frames. Therefore, the table is described only once, and adjacent frames can re-reference the reference frame. Pattern matching mechanisms can be used to match objects in multiple frames. Furthermore, moving objects can be represented over multiple frames due to object movement or camera movement, etc. In a particular example, video can be displayed over multiple frames as a moving car on the screen. This movement can be described using motion vectors. A motion vector is a two-dimensional vector that provides the offset from the coordinates of an object in a frame to the coordinates of that object in a reference frame. Therefore, inter-frame prediction can encode image patches in the current frame as sets of motion vectors representing offsets relative to the corresponding blocks in the reference frame.

[0075] Intra-frame prediction encodes blocks within a common frame. Intra-frame prediction leverages the fact that luma and chroma components tend to cluster within a frame. For example, a patch of green in a section of a tree is often adjacent to several similar patches of green. Intra-frame prediction employs various directional prediction modes (e.g., the 33 modes in HEVC), planar modes, and direct current (DC) modes. Directional modes indicate that the current block is similar / identical to the pixels of its neighboring blocks in the corresponding direction. Planar modes represent interpolation of a series of blocks in a row / column (e.g., a plane) based on neighboring blocks at row edges. In practice, planar modes represent smooth transitions of light / color along rows / columns by employing a relatively constant slope for the changing values. DC modes are used for boundary smoothing, indicating that the block is similar / identical to the average value associated with the pixels of all its neighboring blocks, which are angularly related to the directional prediction mode. Therefore, intra-frame predicted blocks can represent image blocks as values ​​from various relational prediction modes rather than actual values. Furthermore, inter-frame predicted blocks can represent image blocks as motion vector values ​​rather than actual values. In either case, the predicted block may not accurately represent the image block in some situations. All differences are stored in residual blocks. Transformations can be applied to the residual blocks to further compress the file.

[0076] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above can create a blocky image on the decoder side. Furthermore, the block-based prediction scheme can encode blocks and then reconstruct the encoded blocks for later use as reference blocks. The in-loop filtering scheme iteratively applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / frames. These filters reduce such block artifacts, allowing for accurate reconstruction of the encoded file. Furthermore, these filters reduce artifacts in the reconstructed reference block, making it less likely that artifacts will generate other artifacts in subsequent blocks encoded based on the reconstructed reference block.

[0077] In step 109, once the video signal has been segmented, compressed, and filtered, the resulting data is encoded into a bitstream. The bitstream includes the aforementioned data as well as any indicative data necessary to support proper video signal reconstruction at the decoder side. For example, such data may include segmentation data, prediction data, residual blocks, and various flags providing decoding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Creating the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may occur consecutively and / or simultaneously across multiple frames and blocks. Figure 1The order shown is presented for clarity and ease of description and is not intended to restrict the video decoding process to a specific order.

[0078] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data in the bitstream to determine frame segmentation. The segmentation should match the result of block segmentation in step 103. The entropy encoding / decoding used in step 111 is described here. The encoder makes many choices during compression, such as selecting a block segmentation scheme from multiple possible choices based on the spatial location of values ​​in one or more input images. Indicating the exact choice can occupy a large number of bits. Bits used herein are binary values ​​treated as variables (e.g., bit values ​​that may vary depending on the context). Entropy decoding causes the encoder to discard any options that are obviously unsuitable for a particular situation, leaving a set of usable options. A codeword is then assigned to each usable option. The length of the codeword depends on the number of usable options (e.g., one bit for two options, two bits for three or four options, etc.). The encoder then encodes the codewords for the selected options. This scheme reduces the codeword size because the codeword size is as large as the codeword required to uniquely represent a choice from a small subset of available options, rather than uniquely representing a choice from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of available options in a similar manner to the encoder. By determining the set of available options, the decoder can read the codeword and determine the selection made by the encoder.

[0079] In step 113, the decoder performs block decoding. Specifically, the decoder performs an inverse transform to generate residual blocks. Then, the decoder uses the residual blocks and corresponding prediction blocks to reconstruct image blocks based on the segmentation. The prediction blocks may include intra-frame prediction blocks and inter-frame prediction blocks generated on the encoder side in step 105. The reconstructed image blocks are then placed into frames of the reconstructed video signal based on the segmentation data determined in step 111. The syntax of step 113 can also be indicated in the bitstream by the entropy decoding described above.

[0080] In step 115, the frames of the reconstructed video signal are filtered on the encoder side in a manner similar to that in step 107. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters can be applied to the frames to remove block artifacts. In step 117, once the frames have been filtered, the video signal is output to a display for viewing by the end user.

[0081] Figure 2This is a schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 is capable of implementing operation method 100. The codec system 200 broadly describes the components used in the encoder and decoder. The codec system 200 receives a video signal and segments the video signal, as described in steps 101 and 103 of operation method 100, thereby generating a segmented video signal 201. Then, when acting as an encoder, the codec system 200 compresses the segmented video signal 201 into an encoded bitstream, as described in steps 105, 107, and 109 of method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in steps 111, 113, 115, and 117 of operation method 100. The encoding / decoding system 200 includes a general decoder control component 211, a transform scaling and quantization component 213, an intra-frame estimation component 215, an intra-frame prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control and analysis component 227, an intra-loop filter component 225, a decoded image buffer component 223, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown in the figure. Figure 2 In the diagram, black lines represent the movement of data to be encoded / decoded, while dashed lines represent the movement of control data controlling the operation of other components. All components of the encoding / decoding system 200 can be used in the encoder. The decoder may include a subset of the components of the encoding / decoding system 200. For example, the decoder may include an intra-frame prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components are described herein.

[0082] The segmented video signal 201 is an acquired video sequence that has been segmented into pixel blocks by a decoding tree. The decoding tree uses various partitioning patterns to subdivide pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into smaller blocks. These blocks can be referred to as nodes on the decoding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the node / decoding tree depth. In some cases, the segmented blocks can be included in a coding unit (CU). For example, a CU can be a sub-part of a CTU, including a luma block, one or more red chromatic aberration (Cr) blocks, and one or more blue chromatic aberration (Cb) blocks, as well as the corresponding syntax instructions of the CU. The partitioning patterns can include binary trees (BT), triple trees (TT), and quad trees (QT), used to segment nodes into two, three, or four child nodes of different shapes, depending on the partitioning pattern used. The segmented video signal 201 is sent to the general decoder control component 211, the transform scaling and quantization component 213, the intra-frame estimation component 215, the filter control and analysis component 227, and the motion estimation component 221 for compression.

[0083] The general-purpose decoder control component 211 makes decisions related to decoding images of a video sequence into a bitstream based on application constraints. For example, the general-purpose decoder control component 211 manages the optimization of bitrate / bitstream size relative to reconstruction quality. Such decisions can be made based on storage space / bandwidth availability and image resolution requests. The general-purpose decoder control component 211 also manages buffer utilization based on transmission speed to mitigate buffer underloading and overloading issues. To address these issues, the general-purpose decoder control component 211 manages segmentation, prediction, and filtering performed by other components. For example, the general-purpose decoder control component 211 can dynamically increase compression complexity to improve resolution and bandwidth utilization, or decrease compression complexity to reduce resolution and bandwidth utilization. Therefore, the general-purpose decoder control component 211 controls other components of the encoding / decoding system 200 to balance video signal reconstruction quality with bitrate. The general-purpose decoder control component 211 creates control data that controls the operation of other components. The control data is also sent to the header formatting and CABAC component 231 to be encoded into the bitstream, thereby indicating the parameters used for decoding on the decoder side.

[0084] The segmented video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. Frames or stripes of the segmented video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame prediction decoding on the received video blocks relative to one or more blocks in one or more reference frames to provide timing prediction. The encoding / decoding system 200 can perform multiple decoding processes to select an appropriate decoding mode for each video data block, etc.

[0085] Motion estimation component 221 and motion compensation component 219 can be highly integrated, but are described separately for conceptual purposes. Motion estimation performed by motion estimation component 221 is the process of generating motion vectors, which are used to estimate the motion of video blocks. For example, motion vectors can represent the displacement of a decoded object relative to a prediction block. A prediction block is a block found to be highly matched to the block to be decoded in terms of pixel differences. A prediction block can also be referred to as a reference block. Such pixel differences can be determined using the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference measures. HEVC employs several decoded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into multiple CTBs, and then a CTB can be divided into multiple CBs included in a CU. A CU can be encoded as a prediction unit (PU) including prediction data and / or one or more transform units (TUs) including the variational residual data of the CU. Motion estimation component 221 uses rate-distortion analysis as part of a rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc., of the current block / frame, and can select reference blocks, motion vectors, etc., with optimal rate-distortion characteristics. Optimal rate-distortion characteristics balance the quality of video reconstruction (e.g., the amount of data loss due to compression) and decoding efficiency (e.g., the size of the final encoded value).

[0086] In some examples, the codec system 200 can compute values ​​for sub-integer pixel positions of a reference image stored in the decoded image buffer component 223. For example, the video codec system 200 can interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Therefore, the motion estimation component 221 can perform motion search relative to full-pixel and fractional-pixel positions and output motion vectors with fractional-pixel precision. The motion estimation component 221 computes the motion vector of the PU (PU) of a video block in the inter-frame decoding strip by comparing the position of the PU with the position of the predicted block in the reference image. The motion estimation component 221 outputs the computed motion vectors as motion data to the header formatting and CABAC component 231 for encoding and outputs the motion to the motion compensation component 219.

[0087] The motion compensation performed by motion compensation component 219 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation component 221. Additionally, in some examples, motion estimation component 221 and motion compensation component 219 may be functionally integrated. After receiving the motion vector of the PU for the current video block, motion compensation component 219 can locate the prediction block to which the motion vector points. Then, the pixel values ​​of the prediction block are subtracted from the pixel values ​​of the decoded current video block to form a pixel difference, thus forming a residual video block. Typically, motion estimation component 221 performs motion estimation relative to the luma component, and motion compensation component 219 uses the motion vectors calculated based on the luma component for both the chroma and luma components. The prediction block and residual block are then sent to transform scaling and quantization component 213.

[0088] The segmented video signal 201 is also sent to the intra-frame estimation component 215 and the intra-frame prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-frame estimation component 215 and the intra-frame prediction component 217 can be highly integrated, but are described separately for conceptual purposes. The intra-frame estimation component 215 and the intra-frame prediction component 217 perform intra-frame prediction relative to blocks in the current frame, instead of the inter-frame prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra-frame estimation component 215 determines an intra-frame prediction mode for encoding the current block. In some examples, the intra-frame estimation component 215 selects an appropriate intra-frame prediction mode from multiple tested intra-frame prediction modes to encode the current block. The selected intra-frame prediction mode is then sent to the header formatting and CABAC component 231 for encoding.

[0089] For example, intra-frame estimation component 215 performs rate-distortion analysis on various tested intra-prediction modes to calculate rate-distortion values ​​and selects the intra-prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between a coded block and the original uncoded block encoded to produce the coded block, as well as the bit rate (e.g., number of bits) used to generate the coded block. Intra-frame estimation component 215 calculates the ratio based on the distortion and rate of various coded blocks to determine the intra-prediction mode that exhibits the best rate-distortion value for the block. Furthermore, intra-frame estimation component 215 can be used to decode depth blocks of a depth map using a depth modeling mode (DMM) according to rate-distortion optimization (RDO).

[0090] When implemented on the encoder, intra-prediction component 217 can generate residual blocks from the prediction blocks according to the selected intra-prediction mode determined by intra-estimation component 215, or, when implemented on the decoder, read residual blocks from the bitstream. The residual blocks comprise the value difference between the prediction blocks and the original blocks, represented as a matrix. The residual blocks are then sent to transform-scaling and quantization component 213. Intra-estimation component 215 and intra-prediction component 217 can operate on both the luma and chroma components.

[0091] Transform scaling and quantization component 213 is used to further compress the residual block. Transform scaling and quantization component 213 applies a transform to the residual block, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to produce a video block that includes the values ​​of the residual transform coefficients. Wavelet transform, integer transform, subband transform, or other types of transforms can also be used. The transform can convert the residual information from the pixel value domain to the transform domain, such as the frequency domain. For example, transform scaling and quantization component 213 is also used to scale the residual information of the transform according to frequencies, etc. This scaling involves applying a scaling factor to the residual information to quantize different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. Transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, transform scaling and quantization component 213 can then scan a matrix that includes the quantized transform coefficients. The quantization transform coefficients are sent to the header formatting and CABAC component 231 for encoding into the bitstream.

[0092] The scaling and inverse transform component 229 applies the inverse operation of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 uses inverse scaling, inverse transform, and / or inverse quantization to reconstruct residual blocks in the pixel domain, for example, to be used later as reference blocks, which can become prediction blocks for another current block. The motion estimation component 221 and / or the motion compensation component 219 can compute the reference blocks by adding the residual blocks back to the corresponding prediction blocks for motion estimation in subsequent blocks / frames. Filters are applied to the reconstructed reference blocks to reduce artifacts generated during scaling, quantization, and transform. Such artifacts can further lead to inaccurate predictions (and generate other artifacts) when predicting subsequent blocks.

[0093] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, the transform residual block of the scaling and inverse transform component 229 can be merged with the corresponding prediction block of the intra-prediction component 217 and / or motion compensation component 219 to reconstruct the original image block. The filter can then be applied to the reconstructed image block. In some examples, the filter can instead be applied to the residual block. As... Figure 2 The other components, filter control analysis component 227 and in-loop filter component 225, are highly integrated and can be implemented together, but are described separately for conceptual purposes. Filters applied to the reconstructed reference block are applied to a specific spatial region and include multiple parameters to adjust how such filters are applied. Filter control analysis component 227 analyzes the reconstructed reference block to determine where such filters should be applied and sets the corresponding parameters. This data is sent as filter control data to header formatting and CABAC component 231 for encoding. In-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied, for example, in the spatial / pixel domain (e.g., for reconstructed pixel blocks) or the frequency domain.

[0094] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded image buffer assembly 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded image buffer assembly 223 stores the reconstructed and filtered blocks and sends them to the display as part of the output video signal. The decoded image buffer assembly 223 can be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0095] The header formatting and CABAC component 231 receives data from various components of the encoding / decoding system 200 and encodes this data into the decoded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers to encode control data (such as overall control data and filter control data). Furthermore, prediction data (including intra-frame prediction) and motion data, as well as residual data in the form of quantization transform coefficient data, are encoded into the bitstream. The final bitstream contains all the information required by the decoder to reconstruct the original segmented video signal 201. This information may also include an intra-frame prediction mode index table (also called a codeword map), definitions of the coding contexts of various blocks, representations of the most likely intra-frame prediction modes, representations of segmentation information, etc. This data may be encoded using entropy decoding. For example, context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or other entropy decoding techniques can be used to encode the information. After entropy decoding, the decoded bitstream can be sent to another device (e.g., a video decoder) or archived for subsequent transmission or retrieval.

[0096] Figure 3 This is a block diagram of an exemplary video encoder 300. The video encoder 300 can be used to implement the encoding function of the encoding / decoding system 200 and / or implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 segments the input video signal to generate a segmented video signal 301, wherein the segmented video signal 301 is substantially similar to the segmented video signal 201. The segmented video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.

[0097] Specifically, the segmented video signal 301 is sent to the intra-prediction component 317 for intra-frame prediction. The intra-prediction component 317 can be substantially similar to the intra-estimation component 215 and the intra-prediction component 217. The segmented video signal 301 is also sent to the motion compensation component 321 for inter-frame prediction based on the reference block in the decoded image buffer component 323. The motion compensation component 321 can be substantially similar to the motion estimation component 221 and the motion compensation component 219. The predicted blocks and residual blocks from the intra-prediction component 317 and the motion compensation component 321 are sent to the transform and quantization component 313 for transforming and quantizing the residual blocks. The transform and quantization component 313 can be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding predicted blocks (along with associated control data) are sent to the entropy decoding component 331 for decoding into the bitstream. The entropy decoding component 331 can be substantially similar to the header formatting and CABAC component 231.

[0098] The transformed and quantized residual blocks and / or corresponding prediction blocks are also sent from the transform and quantization component 313 to the inverse transform and quantization component 329 to reconstruct a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. According to the example, intra-loop filters in the intra-loop filter component 325 are also applied to the residual blocks and / or the reconstructed reference blocks. The intra-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the intra-loop filter component 225. The intra-loop filter component 325 may include multiple filters as described in the intra-loop filter component 225. The filtered blocks are then stored in the decoded image buffer component 323 for use as a reference block by the motion compensation component 321. The decoded image buffer component 323 may be substantially similar to the decoded image buffer component 223.

[0099] Figure 4 This is a block diagram of an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding function of the encoding / decoding system 200 and / or implement steps 111, 113, 115, and / or 117 of the operation method 100. For example, the decoder 400 receives a bitstream from the encoder 300 and generates a reconstructed output video signal based on the bitstream for display to the end user.

[0100] The bitstream is received by entropy decoding component 433. Entropy decoding component 433 implements entropy decoding schemes such as CAVLC, CABAC, SBAC, PIPE decoding, or other entropy decoding techniques. For example, entropy decoding component 433 can use header information to provide context for interpreting other data encoded as codewords in the bitstream. Decoding information includes any information required to decode the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantization transform coefficients in the residual block. The quantization transform coefficients are sent to inverse transform and quantization component 429 to reconstruct the residual block. Inverse transform and quantization component 429 can be similar to inverse transform and quantization component 329.

[0101] The reconstructed residual block and / or predicted block are sent to the intra-prediction component 417 to reconstruct image blocks according to the intra-prediction operation. The intra-prediction component 417 may be similar to the intra-estimation component 215 and the intra-prediction component 217. Specifically, the intra-prediction component 417 uses a prediction mode to locate a reference block in the frame and applies the residual block to the result to reconstruct the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block, along with the corresponding inter-frame prediction data, are sent to the decoding image buffer component 423 via the intra-loop filter component 425. The decoding image buffer component 423 and the intra-loop filter component 425 may be substantially similar to the decoding image buffer component 223 and the intra-loop filter component 225, respectively. The intra-loop filter component 425 filters the reconstructed image block, the residual block, and / or the predicted block and stores this information in the decoding image buffer component 423. The reconstructed image block from the decoding image buffer component 423 is sent to the motion compensation component 421 for inter-frame prediction. Motion compensation component 421 may be substantially similar to motion estimation component 221 and / or motion compensation component 219. Specifically, motion compensation component 421 uses the motion vectors of a reference block to generate a prediction block and applies the residual block to the result to reconstruct an image block. The resulting reconstructed block can also be sent to decoding image buffer component 423 via in-loop filter component 425. Decoding image buffer component 423 continues to store other reconstructed image blocks, which can be reconstructed into frames using segmentation information. Such frames can also be placed in a sequence. The sequence is output to a display as the reconstructed output video signal.

[0102] Figure 5 This is a schematic diagram of an exemplary bitstream 500 and sub-bitstream 501 extracted from bitstream 500. For example, bitstream 500 may be generated by encoding / decoding system 200 and / or encoder 300, and decoded by encoding / decoding system 200 and / or decoder 400. In another example, bitstream 500 may be generated by encoder in step 109 of method 100 and used by decoder in step 111.

[0103] Bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPS) 512, multiple stripe headers 514, picture data 520, and one or more SEI messages 515. SPS 510 includes sequence data shared by all pictures in the video sequence contained in bitstream 500. This data may include picture size, bit depth, decoding tool parameters, bit rate limits, etc. PPS 512 includes one or more picture-specific parameters. Therefore, each picture in the video sequence can reference a PPS 512. PPS 512 may represent decoding tools, quantization parameters, offsets, picture-specific decoding tool parameters (e.g., filter controls), etc., that can be used for blocks in the corresponding picture. Stripe headers 514 include one or more picture-specific parameters for a corresponding strip 524. Therefore, each strip 524 in the video sequence can reference a stripe header 514. The strip header 514 may include strip type information, image order count (POC), a list of reference images, prediction weights, block entry points, deblocking parameters, etc. In some examples, strip 524 may be referred to as a block group. In this case, strip header 514 may be referred to as a block group header. SEI message 515 is an optional message that includes metadata not needed for block decoding, but can be used for related purposes, such as indicating image output timing, display settings, loss detection, loss hiding, etc.

[0104] Image data 520 includes video data encoded according to inter-frame prediction and / or intra-frame prediction, as well as corresponding transform and quantization residual data. This image data 520 is classified according to the segmentation used to segment the images before encoding. For example, the video sequence is divided into images 521. Images 521 can be further divided into sub-images 522, and sub-images 522 into stripes 524. Stripes 524 can be further divided into blocks and / or CTUs. CTUs are further divided into decoding blocks according to the decoding tree. The decoding blocks can then be encoded / decoded according to the prediction mechanism. For example, image 521 may include one or more sub-images 522. Sub-images 522 may include one or more stripes 524. Image 521 references PPS 512, and stripe 524 references strip header 514. Sub-images 522 can be consistently segmented over the entire video sequence (also called segments), and therefore may reference SPS 510. Each stripe 524 may include one or more blocks. Each strip 524, as well as image 521 and sub-image 522, may also include multiple CTUs.

[0105] Each image 521 may include the entire visual dataset associated with the video sequence at the corresponding moment. However, in some cases, some applications may display only a portion of image 521. For example, a virtual reality (VR) system may display a user-selected region of image 521, creating the sensation of appearing in the scene depicted in image 521. When encoding bitstream 500, the region the user may wish to view is unknown. Therefore, image 521 may include each possible region the user might view as a sub-image 522, which can be individually decoded and displayed based on user input. Other applications may display regions of interest separately. For example, to display one image within another, a television may display a specific region of a video sequence on image 521 of an unrelated video sequence, and thus display sub-image 522. In yet another example, a teleconference system may display the entire image 521 of the currently speaking user and sub-images 522 of users who are not currently speaking. Therefore, sub-image 522 may include the defined region of image 521. Sub-image 522, temporarily subject to motion constraints, can be decoded separately from the rest of image 521. Specifically, the temporal motion constraint subimage is encoded without reference to pixels outside the temporal motion constraint subimage, and thus includes enough information to complete decoding without reference to the rest of image 521.

[0106] Each strip 524 can be a rectangle defined by a CTU in the upper left corner and a CTU in the lower right corner. In some examples, strip 524 comprises a series of blocks and / or CTUs arranged in a left-to-right and top-to-bottom order. In other examples, strip 524 is a rectangular strip. A rectangular strip may not traverse the entire width of the image in raster scan order. Instead, a rectangular strip may comprise rectangular and / or square regions of image 521 and / or sub-image 522 defined by CTUs and / or block rows and CTUs and / or block columns. Strip 524 is the smallest unit that the decoder can display individually. Therefore, strip 524 of image 521 can be assigned to different sub-images 522 to depict desired areas of image 521 respectively.

[0107] The decoder can display one or more sub-images 523 of image 521. Sub-images 523 are subgroups of user-selected or predefined sub-images 522. For example, image 521 can be divided into nine sub-images 522, but the decoder can display only a single sub-image 523 from that group of sub-images 522. Sub-images 523 include stripes 525, which are subgroups of selected or predefined stripes 524. To display sub-images 523 separately, sub-streams 501 can be extracted (529) from bitstream 500. Extraction (529) can occur on the encoder side, such that the decoder receives only sub-streams 501. In other cases, the entire bitstream 500 is sent to the decoder, and the decoder extracts (529) sub-streams 501 for separate decoding. It should be noted that sub-streams 501 can also be collectively referred to as bitstreams in some cases. Substream 501 includes SPS 510, PPS 512, selected subimage 523, strip header 514, and SEI message 515 associated with subimage 523 and / or strip 525.

[0108] This invention instructs various data to support efficient decoding of sub-images 522 in the decoder, so as to select and display sub-images 523. SPS 510 includes a sub-image size 531, a sub-image position 532, and a sub-image ID 533 associated with the complete set of sub-images 522. The sub-image size 531 includes the sub-image height and width in luminance pixels of the corresponding sub-image 522. The sub-image position 532 includes the offset distance between the top-left pixel of the corresponding sub-image 522 and the top-left pixel of image 521. The sub-image position 532 and sub-image size 531 define the layout of the corresponding sub-image 522. The sub-image ID 533 includes data that uniquely identifies the corresponding sub-image 522. The sub-image ID 533 can be the raster scan index of the sub-image 522 or other defined values. Therefore, the decoder can read SPS 510 and determine the size, position, and ID of each sub-image 522. In some video decoding systems, data associated with sub-image 522 may be included in PPS 512 because sub-image 522 is segmented from image 521. However, the portion used to create sub-image 522 can be used by applications, such as ROI-based applications, VR applications, etc., which rely on a consistent portion of sub-image 522 across the video sequence / segment. Therefore, the sub-image 522 portion typically does not change on a per-image basis. Placing the layout information of sub-image 522 in SPS 510 ensures that the layout is indicated only once for the sequence / segment, rather than for each PPS 512 (which in some cases could be indicated for each image 521). Furthermore, indicating the sub-image 522 information instead of relying on the decoder to derive such information reduces the likelihood of errors in the event of lost messages and supports other functions for extracting sub-image 523. Therefore, indicating the sub-image 522 layout in SPS 510 enhances the functionality of the encoder and / or decoder.

[0109] SPS 510 also includes a motion-constrained sub-image flag 534 associated with the complete sub-image set 522. The motion-constrained sub-image flag 534 indicates whether each sub-image 522 is a temporally motion-constrained sub-image. Therefore, the decoder can read the motion-constrained sub-image flag 534 and determine which sub-images 522 can be extracted and displayed individually without decoding other sub-images 522. The selected sub-images 522 are decoded as temporally motion-constrained sub-images, while other sub-images 522 are decoded without such constraints, thus improving decoding efficiency.

[0110] Sub-image ID 533 is also included in the stripe header 514. Each stripe header 514 includes data associated with the corresponding stripe set 524. Therefore, the stripe header 514 only includes the sub-image ID 533 corresponding to the stripe 524 associated with the stripe header 514. In this way, the decoder can receive the stripe 524, obtain the sub-image ID 533 from the stripe header 514, and determine which sub-image 522 includes the stripe 524. The decoder can also use the sub-image ID 533 from the stripe header 514 to correlate with the relevant data in the SPS 510. In this way, the decoder can determine how to locate the sub-images 522 / 523 and the stripes 524 / 525 by reading the SPS 510 and the relevant stripe header 514. This makes it possible to decode the sub-images 523 and the stripes 525 even if some sub-images 522 are lost or deliberately omitted during transmission to improve decoding efficiency.

[0111] SEI message 515 may also include a sub-image level 535. The sub-image level 535 represents the hardware resources required to decode the corresponding sub-image 522. This allows each sub-image 522 to be decoded independently of other sub-images 522. This ensures that the correct amount of hardware resources is allocated to each sub-image 522 in the decoder. Without such a sub-image level 535, sufficient resources would be allocated to decode the most complex sub-image 522. Therefore, if a sub-image 522 is associated with varying hardware resource requirements, the sub-image level 535 prevents the decoder from over-allocating hardware resources.

[0112] Figure 6 This is a schematic diagram of an exemplary image 600 segmented into sub-images 622. For example, image 600 can be encoded in and decoded from bitstream 500 by encoding / decoding system 200, encoder 300, and / or decoder 400. Furthermore, image 600 can be segmented and / or included in sub-bitstream 501 to support encoding and decoding according to method 100.

[0113] Image 600 can be substantially similar to image 521. Furthermore, image 600 can be divided into sub-images 622, which are substantially similar to sub-image 522. Each sub-image 622 includes a sub-image size 631, which can be included in bitstream 500 as sub-image size 531. Sub-image size 631 includes a sub-image width 631a and a sub-image height 631b. Sub-image width 631a is the width of the corresponding sub-image 622 in luminance pixels. Sub-image height 631b is the height of the corresponding sub-image 622 in luminance pixels. Each sub-image 622 includes a sub-image ID 633, which can be included in bitstream 500 as sub-image ID 633. Sub-image ID 633 can be any value that uniquely identifies each sub-image 622. In the example shown, sub-image ID 633 is the index of sub-image 622. Each sub-image 622 includes a position 632, which can be included in the bitstream 500 as sub-image position 532. Position 632 is represented as the offset between the top-left pixel of the corresponding sub-image 622 and the top-left pixel 642 of image 600.

[0114] Furthermore, as shown in the figure, some sub-images 622 may be time-motion-constrained sub-images 634, while others may not be. In the example shown, sub-image 622 with sub-image ID 633 of 5 is a time-motion-constrained sub-image 634. This means that sub-image 622 identified as 5 can be decoded without referencing any other sub-images 622, thus allowing extraction and separate decoding without considering data from other sub-images 622. The motion-constrained sub-image tag 534 can be used in the bitstream 500 to indicate which sub-images 622 are time-motion-constrained sub-images 634.

[0115] As shown in the figure, sub-image 622 can be constrained to cover image 600 without gaps or overlaps. A gap is a region in image 600 that is not included in any sub-image 622. An overlap is a region of image 600 that is included in multiple sub-images 622. Figure 6In the example shown, sub-images 622 are segmented from image 600 to prevent gaps and overlaps. Gaps cause pixels of image 600 to remain outside sub-image 622. Overlaps cause related stripes to be included in multiple sub-images 622 simultaneously. Therefore, gaps and overlaps can lead to different pixel processing effects when sub-images 622 are decoded in different ways. If the encoder allows this, the decoder must support this decoding scheme, even if the decoding scheme is rarely used. By disallowing gaps and overlaps in sub-images 622, the complexity of the decoder can be reduced because the decoder does not need to consider potential gaps and overlaps when determining the sub-image size 631 and position 632. Furthermore, disallowing gaps and overlaps in sub-images 622 reduces the complexity of the RDO process in the encoder. This is because the encoder can ignore gaps and overlaps when selecting encodings for the video sequence. Therefore, avoiding gaps and overlaps can reduce the memory resource utilization and / or processing resource utilization of both the encoder and decoder.

[0116] Figure 7 This is a schematic diagram of an exemplary mechanism 700 for relating stripe 724 to the layout of sub-image 722. For example, mechanism 700 can be applied to image 600. Furthermore, mechanism 700 can be applied based on data in bitstream 500 by encoding / decoding system 200, encoder 300, and / or decoder 400, etc. Additionally, mechanism 700 can be used for encoding and decoding according to method 100.

[0117] Mechanism 700 can be applied to stripes 724 in sub-image 722, such as stripes 524 / 525 and sub-images 522 / 523. In the example shown, sub-image 722 includes a first stripe 724a, a second stripe 724b, and a third stripe 724c. The strip header of each stripe 724 includes a sub-image ID 733 of sub-image 722. The decoder can match the sub-image ID 733 in the strip header with the sub-image ID 733 in the SPS. The decoder can then determine the position 732 and size of sub-image 722 from the SPS based on the sub-image ID 733. Using position 732, sub-image 722 can be positioned relative to the top-left pixel of the top-left corner 742 of the image. This size can be used to set the height and width of sub-image 722 relative to position 732. Stripe 724 can then be included in sub-image 722. Therefore, without referencing other sub-images, stripe 724 can be placed in the correct sub-image 722 based on sub-image ID 733. This supports error correction because other missing sub-images do not affect the decoding of sub-image 722. This also supports applications that extract only sub-image 722 and avoid sending other sub-images. Therefore, sub-image ID 733 enhances functionality and / or improves decoding efficiency, which reduces the network resource utilization, memory resource utilization, and / or processing resource utilization of the encoder and decoder.

[0118] Figure 8 This is a schematic diagram of another exemplary image 800 segmented into sub-images 822. Image 800 may be substantially similar to image 600. Furthermore, image 800 can be encoded in and decoded from bitstream 500 by encoding / decoding system 200, encoder 300, and / or decoder 400. Additionally, image 800 can be segmented and / or included in sub-bitstream 501 to support encoding and decoding according to method 100 and / or mechanism 700.

[0119] Image 800 includes sub-image 822, which may be substantially similar to sub-images 522, 523, 622, and / or 722. Sub-image 822 is divided into multiple CTUs 825. A CTU 825 is a basic decoding unit in a standardized video decoding system. The CTU 825 is subdivided into decoding blocks using a decoding tree, and the decoding blocks are decoded based on inter-frame prediction or intra-frame prediction. As shown, the width and height of some sub-images 822a are constrained to multiples of the CTU 825 size. In the example shown, sub-image 822a has a height of 6 CTUs 825 and a width of 5 CTUs 825. This constraint needs to be removed for sub-image 822b located on the right boundary 801 of the image and sub-image 822c located on the lower boundary 802 of the image. In the example shown, sub-image 822b has a width between 5 and 6 CTUs 825. Therefore, subimage 822b has a width of five full CTUs 825 and one incomplete CTU 825. However, the subimage height of subimage 822b, which is not located on the lower boundary 802 of the image, is still constrained to a multiple of the CTU 825 size. In the example shown, subimage 822c has a height between 6 and 7 CTUs 825. Therefore, subimage 822c has a height of six full CTUs 825 and one incomplete CTU 825. However, the subimage width of subimage 822c, which is not located on the right boundary 801 of the image, is still constrained to a multiple of the CTU 825 size. It should be noted that the right boundary 801 and the lower boundary 802 of the image can also be referred to as the right boundary and the lower boundary of the image, respectively. It should also be noted that the CTU 825 size is a user-defined value. The CTU 825 size can be any value between the minimum and maximum CTU 825 size. For example, the minimum CTU 825 size can be 16 luminance pixels in height and 16 luminance pixels in width. Conversely, the maximum CTU 825 size can be 128 luminance pixels in height and 128 luminance pixels in width.

[0120] As described above, some video systems may restrict the height and width of sub-image 822 to multiples of the CTU 825 size. This causes sub-image 822 to operate incorrectly in many image layouts, such as in image 800 where the total width or height is not a multiple of the CTU 825 size. By ensuring that the height and width of the lower sub-image 822c and the right sub-image 822b are not multiples of the CTU 825 size, respectively, sub-image 822 can be used with any image 800 without causing decoding errors. This enhances the capabilities of both the encoder and decoder. Furthermore, this enhanced capability allows the encoder to decode images more efficiently, which reduces the network resource utilization, memory resource utilization, and / or processing resource utilization of both the encoder and decoder.

[0121] As described herein, this invention describes a design for sub-image-based image segmentation in video decoding. A sub-image is a rectangular region in an image that can be decoded independently using the same decoding process as the image. This invention relates to the indication of sub-images in decoded video sequences and / or bitstreams, and the process for extracting sub-images. The description of these techniques is based on ITU-T and ISO / IEC JVET's VVC. However, these techniques are also applicable to other video codec specifications. Exemplary embodiments described herein are given below. Such embodiments can be applied individually or in combination.

[0122] Information related to sub-images that may exist in a coded video sequence (CVS) can be indicated in a sequence-level parameter set (SPS). Such indications may include the following: The number of sub-images present in each image of the CVS can be indicated in the SPS. In the context of the SPS or CVS, juxtaposed sub-images of all access units (AUs) can be collectively referred to as a sub-image sequence. The SPS may also include loops to further represent information describing the attributes of each sub-image. This information may include the sub-image identifier, the sub-image's position (e.g., the offset distance between the top-left luminance pixel of the sub-image and the top-left luminance pixel of the image), and the sub-image's size. Furthermore, the SPS may indicate whether each sub-image is a motion-constrained sub-image (including the functionality of MCTS). Profile, layer, and level information for each sub-image can also be indicated or derived in the decoder. Such information can be used to determine the profile, layer, and level information of the bitstream created by extracting sub-images from the original bitstream. The profile and layer of each sub-image can be derived to be the same as the profile and layer of the entire bitstream. The level of each sub-image can be explicitly indicated. Such indications can appear in loops included in SPS. Sequence-level hypothetical reference decoder (HRD) parameters can be indicated for each sub-image (or equivalently, each sub-image sequence) in the video usability information (VUI) section of SPS.

[0123] When an image is not segmented into two or more sub-images, the attributes of the sub-images (e.g., position, size, etc.) other than the sub-image ID may not exist or be indicated in the bitstream. When extracting sub-images from an image in the CVS, each access unit in the new bitstream may not include the sub-image. In this case, the image in each AU of the new bitstream is not segmented into multiple sub-images. Therefore, it is not necessary to indicate sub-image attributes such as position and size in the SPS, as such information can be derived from the image attributes. However, the sub-image identifier can still be indicated, as the ID can be referenced by the VCL NAL unit / block group included in the extracted sub-image. This maintains the same sub-image ID when extracting sub-images.

[0124] The position (x-offset and y-offset) of a subimage within an image can be indicated in luminance pixels. This position represents the distance between the top-left luminance pixel of the subimage and the top-left luminance pixel of the image. Alternatively, the position can be indicated in the minimum decoded luminance block size (MinCbSizeY). Alternatively, the unit of subimage position offset can be explicitly indicated by a syntax element in the parameter set. The unit can be CtbSizeY, MinCbSizeY, luminance pixels, or other values.

[0125] The size of a subimage (subimage width and height) can be indicated in luminance pixels. Alternatively, the size can be indicated in the minimum decoded luminance block size (MinCbSizeY). The unit of the subimage size value can also be explicitly indicated by a syntax element in the parameter set. The unit can be CtbSizeY, MinCbSizeY, luminance pixels, or other values. When the right boundary of the subimage does not coincide with the right boundary of the image, the width of the subimage can be required to be an integer multiple of the luminance CTU size (CtbSizeY). Similarly, when the bottom boundary of the subimage does not coincide with the bottom boundary of the image, the height of the subimage can be required to be an integer multiple of the luminance CTU size (CtbSizeY). If the width of the subimage is not an integer multiple of the luminance CTU size, the subimage can be required to be located at the rightmost position in the image. Similarly, if the height of the subimage is not an integer multiple of the luminance CTU size, the subimage can be required to be located at the bottommost position in the image. In some cases, the width of a subimage can be indicated in units of luminance CTU size, but the subimage width is not an integer multiple of the luminance CTU size. In this case, the actual width in luminance pixels can be derived based on the subimage's offset position. The subimage width can be derived based on the luminance CTU size, and the image height can be derived based on luminance pixels. Similarly, the height of a subimage can be indicated in units of luminance CTU size, but the subimage height is not an integer multiple of the luminance CTU size. In this case, the actual height based on luminance pixels can be derived based on the subimage's offset position. The subimage height can be derived based on the luminance CTU size, and the image height can be derived based on luminance pixels.

[0126] For any given subimage, the subimage ID and subimage index can be different. The subimage index can be the index of the subimage indicated in the subimage cycle in SPS. The subimage ID can be the index of the subimage in the subimage raster scan sequence within the image. When the value of the subimage ID for each subimage is the same as the subimage index, the subimage ID can be indicated or derived. When the subimage ID for each subimage is different from the subimage index, the subimage ID is explicitly indicated. The number of bits used to indicate the subimage ID can be indicated in the same parameter set that includes the subimage attributes (e.g., in SPS). Some values ​​of the subimage ID can be reserved for certain purposes. For example, when the block group header includes a subimage ID to specify which subimage includes the block group, a value of 0 can be reserved for the subimage and not used for the subimage to ensure that the first few bits of the block group header are not all 0 to avoid accidentally including anti-counterfeiting codes. In the optional case where the subimages of the image do not cover the entire area of ​​the image without gaps and overlaps, a value (e.g., value 1) can be reserved for block groups that do not belong to any subimage. Alternatively, the subimage IDs for the remaining areas can be explicitly indicated. The number of bits used to indicate a sub-image ID can be constrained as follows. The range of this value should be sufficient to uniquely identify all sub-images in an image, including reserved sub-image ID values. For example, the minimum number of bits for a sub-image ID could be the value of Ceil(Log2(number of sub-images in an image + number of reserved sub-image IDs)).

[0127] The combination of subimages must cover the entire image without gaps or overlaps, which may be subject to constraints. When imposing such constraints, for each subimage, there can be a flag indicating whether the subimage is a motion-constrained subimage, meaning that the subimage can be extracted. Alternatively, the combination of subimages may not cover the entire image, but overlap is not allowed.

[0128] Sub-image IDs can immediately follow the NAL unit header. This facilitates the sub-image extraction process, eliminating the need for the extractor to parse the remaining NAL unit bits. For VCL NAL units, the sub-image ID can be among the first few bits of the chunk group header. For non-VCL NAL units, the following applies: For SPS, the sub-image ID does not need to immediately follow the NAL unit header. For PPS, if all chunk groups of the same image are constrained to refer to the same PPS, the sub-image ID does not need to immediately follow the NAL unit header. If chunk groups of the same image are allowed to refer to different PPSs, the sub-image ID can be among the first few bits of the PPS (e.g., immediately following the NAL unit header). In this case, any chunk groups of an image can share the same PPS. Alternatively, when chunk groups of the same image are allowed to refer to different PPSs, and different chunk groups of the same image are allowed to share the same PPS, the PPS syntax may not contain a sub-image ID. Alternatively, when chunk groups of the same image are allowed to refer to different PPSs, and different chunk groups of the same image are allowed to share the same PPS, the PPS syntax may contain a list of sub-image IDs. This list represents the sub-images to which PPS is applied. For other non-VCL NAL units, if the non-VCL unit applies to the image level or higher (such as access unit separators, sequence ends, stream ends, etc.), the sub-image ID may not immediately follow the NAL unit header. Otherwise, the sub-image ID may immediately follow the NAL unit header.

[0129] Using the SPS indicator described above, block segmentation within each sub-image can be indicated in the PPS. It is permissible for block groups within the same image to reference different PPSs. In this case, block grouping may only occur within each sub-image. The concept of block grouping is to divide a sub-image into blocks.

[0130] Alternatively, a parameter set can be defined to describe the block segmentation within each sub-image. This parameter set can be called a sub-picture parameter set (SPPS). SPPS references SPS. A syntax element referencing the SPSID exists within the SPPS. SPPS can include sub-image IDs. For sub-image extraction, the syntax element referencing the sub-image ID is the first syntax element in the SPPS. SPPS includes block structure (e.g., number of columns, number of rows, uniform block spacing, etc.). SPPS can include flags indicating whether a cyclic filter is enabled across relevant sub-image boundaries. Alternatively, sub-image attributes for each sub-image can be indicated in the SPPS instead of the SPS. Block segmentation within each sub-image can still be indicated in the PPS. It is permissible for block groups within the same image to reference different PPSs. Once an SPPS is activated, it persists for a series of consecutive AUs in the decoding order. However, an SPPS can be deactivated / activated in an AU that is not the beginning of a CVS. In some AUs, multiple SPPSs can be active at any point during the decoding process of a single-layer bitstream with multiple sub-images. Different sub-images of an AU can share a single SPPS. Alternatively, SPPS and PPS can be merged into a single parameter set. In this case, all patch groups within the same image may not need to reference the same PPS. Constraints can be imposed such that all patch groups within the same sub-image can reference the same parameter set generated by the fusion of SPPS and PPS.

[0131] The number of bits used to indicate the sub-picture ID can be specified in the NAL unit header. When this information is present in the NAL unit header, it helps the sub-picture extraction process parse the sub-picture ID value at the beginning of the NAL unit's payload (e.g., the first few bits immediately following the NAL unit header). For such a specification, some reserved bits in the NAL unit header (e.g., seven reserved bits) can be used to avoid increasing the length of the NAL unit header. The number of bits for this specification can override the value of `sub-picture-ID-bit-len`. For example, four of the seven reserved bits in the VVC NAL unit header can be used for this purpose.

[0132] When decoding sub-images, the positions of each decoder block (e.g., xCtb and yCtb) can be adjusted to the actual luminance pixel positions in the image, rather than the luminance pixel positions in the sub-image. This avoids extracting and placing sub-images from each reference image, since the decoder blocks are decoded from the reference image, not the reference sub-image. When adjusting the decoder block positions, the variables SubpictureXOffset and SubpictureYOffset can be derived from the sub-image positions (subpic_x_offset and subpic_y_offset). The values ​​of these variables can then be added to the x and y coordinates of the luminance pixel positions for each decoder block in the sub-image.

[0133] The sub-image extraction process can be defined as follows. The input to this process is the target sub-image to be extracted. This can be either the sub-image ID or the sub-image location. When the input is the sub-image location, the relevant sub-image ID can be parsed from the sub-image information in the SPS. For non-VCL NAL units, the following applies: Syntax elements in the SPS related to image size and level can be updated to the sub-image size and level information. The following non-VCL NAL units remain unchanged: PPS, access unit delimiter (AUD), end of sequence (EOS), end of bitstream (EOB), and non-VCL NAL units applicable to image level or higher. Other non-VCL NAL units where the sub-image ID is not equal to the target sub-image ID can be removed. VCL NAL units where the sub-image ID is not equal to the target sub-image ID can also be removed.

[0134] Sequence-level sub-image nested SEI messages can be used to nest AU-level or sub-image-level SEI messages for sub-image sets. This can include buffering periods, image timing, and non-HRD SEI messages. The syntax and semantics of this sub-image nested SEI message can be as follows. For system operation, such as in an omnidirectional media format (OMAF) environment, an OMAF player can request and decode a set of sub-image sequences covering a viewpoint. Therefore, a sequence-level SEI message is used to carry information about a set of sub-image sequences collectively covering a rectangular image region. The system can use this information, which indicates the required decoding capability and the bitrate of the sub-image sequence set. This information indicates the level of the bitstream that includes only the sub-image sequence set. This information also indicates the bitrate of the bitstream that includes only the sub-image sequence set. Optionally, a sub-bitstream extraction process can be specified for the sub-image sequence set. The advantage of doing so is that the bitstream that includes only the sub-image sequence set can also be consistent. The disadvantage is that, when considering the possibility of different viewpoint sizes, many such sets can exist in addition to the already large number of individual sub-image sequences.

[0135] In one exemplary embodiment, one or more of the disclosed examples can be implemented as follows. A sub-image can be defined as a rectangular region of one or more blocks in an image. The allowed binary partitioning process can be defined as follows. The inputs to the process are: binary partitioning mode btSplit, decoded block width cbWidth, decoded block height cbHeight, the position (x0, y0) of the top-left luminance pixel of the considered decoded block relative to the top-left luminance pixel of the image, multi-type tree depth mttDepth, maximum multi-type tree depth maxMttDepth with offset, maximum binary tree size maxBtSize, and partition index partIdx. The output of the process is the variable allowBtSplit.

[0136] Based on btSplit, parallelTtSplit and cbSize specifications

[0137] The variables parallelTtSplit and cbSize are derived as described above. The variable allowBtSplit is derived as follows: allowBtSplit is set to false if one or more of the following conditions are true: cbSize is less than or equal to MinBtSizeY, cbWidth is greater than maxBtSize, cbHeight is greater than maxBtSize, and mttDepth is greater than or equal to maxMttDepth. Otherwise, allowBtSplit is set to false if all of the following conditions are true: btSplit is equal to SPLIT_BT_VER, and y0 + cbHeight is greater than SubPicBottomBorderInPic. Otherwise, allowBtSplit is set to false if all of the following conditions are true: btSplit is equal to SPLIT_BT_HOR, x0 + cbWidth is greater than SubPicRightBorderInPic, and y0 + cbHeight is less than or equal to SubPicBottomBorderInPic. Otherwise, allowBtSplit is set to FALSE if all of the following conditions are true: mttDepth is greater than 0, partIdx is equal to 1, and MttSplitMode[x0][y0][mttDepth - 1] is equal to parallelTtSplit. Otherwise, allowBtSplit is set to false if all of the following conditions are true: btSplit is equal to SPLIT_BT_VER, cbWidth is less than or equal to MaxTbSizeY, and cbHeight is greater than MaxTbSizeY. Otherwise, allowBtSplit is set to false if all of the following conditions are true: btSplit is equal to SPLIT_BT_HOR, cbWidth is greater than MaxTbSizeY, and cbHeight is less than or equal to MaxTbSizeY. Otherwise, allowBtSplit is set to true.

[0138] The allowed ternary partitioning process can be defined as follows. The inputs to this process are: the ternary partitioning pattern `ttSplit`, the decoded block width `cbWidth`, the decoded block height `cbHeight`, the position (x0, y0) of the top-left luminance pixel of the considered decoded block relative to the top-left luminance pixel of the image, the multi-type tree depth `mttDepth`, the maximum multi-type tree depth with offset `maxMttDepth`, and the maximum binary tree size `maxTtSize`. The output of this process is the variable `allowTtSplit`.

[0139] cbSize specification based on ttSplit.

[0140]

[0141] The variable `cbSize` is derived as described above. The variable `allowTtSplit` is derived as follows: `allowTtSplit` is set to false if one or more of the following conditions are true: `cbSize` is less than or equal to 2. MinTtSizeY, cbWidth is greater than Min(MaxTbSizeY, maxTtSize), cbHeight is greater than Min(MaxTbSizeY, maxTtSize), mttDepth is greater than or equal to maxMttDepth, x0 + cbWidth is greater than SubPicRightBorderInPic, and y0 + cbHeight is greater than SubPicBottomBorderInPic. Otherwise, set allowTtSplit to true.

[0142] The syntax and semantics of the Sequence Parameter Set (RBSP) are as follows.

[0143]

[0144] `pic_width_in_luma_samples` represents the width of each decoded image in luminance pixels. `pic_width_in_luma_samples` should not be equal to 0 and should be an integer multiple of `MinCbSizeY`. `pic_height_in_luma_samples` represents the height of each decoded image in luminance pixels. `pic_height_in_luma_samples` should not be equal to 0 and should be an integer multiple of `MinCbSizeY`. `Num_subpicture_minus1` plus 1 indicates the number of subpictures segmented within the decoded images belonging to the decoded video sequence. `subpic_id_len_minus1` plus 1 indicates the number of bits used to represent the syntax element `subpic_id[i]` in the SPS. `spps_subpic_id` in the SPS references the SPS, and `tile_group_subpic_id` in the tile group header references the SPS. The value of `subpic_id_len_minus1` should range from Ceil(Log2(num_subpic_minus1 + 2)) to 8 (inclusive). `subpic_id[i]` represents the subpick ID of the i-th subpick in the image referencing the SPS. The length of `subpic_id[i]` is `subpic_id_len_minus1 + 1` bits. The value of `subpic_id[i]` should be greater than 0. `subpic_level_idc[i]` represents the level of CVS extracted from the i-th subpick that meets the specified resource requirements. The bitstream should not include the value of `subpic_level_idc[i]`, except for the specified value. Other values ​​of `subpic_level_idc[i]` are reserved. If it does not exist, the value of `subpic_level_idc[i]` is assumed to be equal to the value of `general_level_idc`.

[0145] `subpic_x_offset[i]` represents the horizontal offset of the top-left corner of the i-th sub-image relative to the top-left corner of the image. If it does not exist, the value of `subpic_x_offset[i]` is inferred to be 0. The derivation of the sub-image x-offset value is as follows: `SubpictureXOffset[i] = subpic_x_offset[i]`. `subpic_y_offset[i]` represents the vertical offset of the top-left corner of the i-th sub-image relative to the top-left corner of the image. If it does not exist, the value of `subpic_y_offset[i]` is inferred to be 0. The derivation of the sub-image y-offset value is as follows: `SubpictureYOffset[i] = subpic_y_offset[i]`. `subpic_width_in_luma_samples[i]` represents the width of the i-th decoded sub-image for which the SPS is active. When the sum of SubpictureXOffset[i] and subpic_width_in_luma_samples[i] is less than pic_width_in_luma_samples, the value of subpic_width_in_luma_samples[i] should be an integer multiple of CtbSizeY. When it does not exist, it is inferred that the value of subpic_width_in_luma_samples[i] is equal to the value of pic_width_in_luma_samples. subpic_height_in_luma_samples[i] represents the height of the i-th decoded sub-image of the active SPS. When the sum of SubpictureYOffset[i] and subpic_height_in_luma_samples[i] is less than pic_height_in_luma_samples, the value of subpic_height_in_luma_samples[i] should be an integer multiple of CtbSizeY. If it does not exist, then it is inferred that the value of subpic_height_in_luma_samples[i] is equal to the value of pic_height_in_luma_samples.

[0146] The combination of sub-images should cover the entire region of the image without overlap or gaps, which is a requirement for bitstream consistency. `subpic_motion_constrained_flag[i]` equals 1, indicating that the i-th sub-image is a time-motion-constrained sub-image. `subpic_motion_constrained_flag[i]` equals 0, indicating that the i-th sub-image may or may not be a time-motion-constrained sub-image. When it does not exist, it is inferred that the value of `subpic_motion_constrained_flag` is equal to 0.

[0147] The variables SubpicWidthInCtbsY, SubpicHeightInCtbsY, SubpicSizeInCtbsY, SubpicWidthInMinCbsY, SubpicHeightInMinCbsY, SubpicSizeInMinCbsY, SubpicSizeInSamplesY, SubpicWidthInSamplesC, and SubpicHeightInSamplesC are derived as follows: SubpicWidthInLumaSamples[ i ] = subpic_width_in_luma_samples[ i ] SubpicHeightInLumaSamples[ i ] = subpic_height_in_luma_samples[ i ] SubPicRightBorderInPic[ i ] = SubpictureXOffset[ i ]+PicWidthInLumaSamples[ i ] SubPicBottomBorderInPic[ i ] = SubpictureYOffset[ i ]+PicHeightInLumaSamples[ i ] SubpicWidthInCtbsY[ i ] = Ceil( SubpicWidthInLumaSamples[ i ]÷CtbSizeY ) SubpicHeightInCtbsY[ i ] = Ceil( SubpicHeightInLumaSamples[ i ]÷CtbSizeY ) SubpicSizeInCtbsY[ i ] = SubpicWidthInCtbsY[ i ] SubpicHeightInCtbsY[ i ] SubpicWidthInMinCbsY[ i ] = SubpicWidthInLumaSamples[ i ] / MinCbSizeY SubpicHeightInMinCbsY[ i ] = SubpicHeightInLumaSamples[ i ] / MinCbSizeY SubpicSizeInMinCbsY[ i ] = SubpicWidthInMinCbsY[ i ] SubpicHeightInMinCbsY[ i ] SubpicSizeInSamplesY[ i ] = SubpicWidthInLumaSamples[ i ] SubpicHeightInLumaSamples[ i ] SubpicWidthInSamplesC[ i ] = SubpicWidthInLumaSamples[ i ] / SubWidthC SubpicHeightInSamplesC[ i ] = SubpicHeightInLumaSamples[ i ] / SubHeightC The syntax and semantics of the sub-picture parameter set RBSP are as follows.

[0148]

[0149] `spps_subpic_id` identifies the subpic of which the SPPS belongs. The length of `spps_subpic_id` is `subpic_id_len_minus1 + 1` bits. `spps_subpic_parameter_set_id` identifies the SPPS for reference by other syntax elements. The value of `spps_subpic_parameter_set_id` should range from 0 to 63 (inclusive). The syntax element `spps_seq_parameter_set_id` represents the value of `sps_seq_parameter_set_id` that activates the SPS. The value of `spps_seq_parameter_set_id` should range from 0 to 15 (inclusive). `single_tile_in_subpic_flag` equals 1, indicating that there is only one tile in each subpic of the referenced SPPS. `single_tile_in_subpic_flag` equals 0, indicating that there are multiple tiles in each subpic of the referenced SPPS. `num_tile_columns_minus1 + 1` represents the number of tile columns that divide the subpic of the SPPS. `num_tile_columns_minus1` should be in the range of 0 to `PicWidthInCtbsY[spps_subpic_id] - 1` (inclusive). If it does not exist, the value of `num_tile_columns_minus1` is assumed to be 0. `num_tile_rows_minus1` plus 1 indicates the number of tile rows in the segmented subimage. `num_tile_rows_minus1` should be in the range of 0 to `PicHeightInCtbsY[spps_subpic_id] - 1` (inclusive). If it does not exist, the value of `num_tile_rows_minus1` is assumed to be 0. Set the variable `NumTilesInPic` to `(num_tile_columns_minus1 + 1)`. (num_tile_rows_minus1 + 1).

[0150] When `single_tile_in_subpic_flag` equals 0, `NumTilesInPic` should be greater than 0. `uniform_tile_spacing_flag` equal to 1 indicates that the tile column and row boundaries are evenly distributed across the subimage. `uniform_tile_spacing_flag` equal to 0 indicates that the tile column and row boundaries are not evenly distributed across the subimage, but are explicitly indicated using the syntax elements `tile_column_width_minus1[i]` and `tile_row_height_minus1[i]`. When these elements are not present, the value of `uniform_tile_spacing_flag` is inferred to be 1. `tile_column_width_minus1[i]` plus 1 represents the width (in CTB) of the i-th tile column. `tile_row_height_minus1[i]` plus 1 represents the height (in CTB) of the i-th tile row.

[0151] By calling the CTB raster and block scan conversion process, the following variables are derived: List ColWidth[i], where i ranges from 0 to num_tile_columns_minus1 (inclusive), and ColWidth[i] represents the width of the i-th block column (in CTB); List RowHeight[j], where j ranges from 0 to num_tile_rows_minus1 (inclusive), and RowHeight[j] represents the height of the j-th block row (in CTB); List ColBd[i], where i ranges from 0 to num_tile_columns_minus1 + 1 (inclusive), and ColBd[i] represents the position of the boundary of the i-th block column (in CTB); List RowBd[j], where j ranges from 0 to num_tile_rows_minus1 + 1 (inclusive), and RowBd[j] represents the width of the i-th block column (in CTB); List RowBd[j], where j ranges from 0 to num_tile_rows_minus1 + 1 (inclusive), and RowBd[j] represents the width of the i-th block column (in CTB); The list `CtbAddrRsToTs[ctbAddrRs]` represents the position of the j-th block row boundary (in CTB units); `CtbAddrRs` ranges from 0 to `PicSizeInCtbsY - 1` (inclusive); `CtbAddrRsToTs[ctbAddrRs]` represents the conversion of CTB addresses in the CTB raster scan to CTB addresses in the block scan; `CtbAddrTsToRs[ctbAddrTs]` ranges from 0 to `PicSizeInCtbsY - 1` (inclusive); `CtbAddrTsToRs[ctbAddrTs]` represents the conversion of CTB addresses in the block scan to CTB addresses in the CTB raster scan; `TileId[ctbAddrTs]` ranges from 0 to `PicSizeInCtbsY - 1`. 1 (including end values), the list TileId[ ctbAddrTs ] represents the conversion from CTB address to block ID in block scan; the list NumCtusInTile[ tileIdx ], the range of tileIdx is 0 to PicSizeInCtbsY - 1 (including end values), the list NumCtusInTile[ tileIdx ] represents the conversion from block index to the number of CTUs in the block; the list FirstCtbAddrTs[ tileIdx ], the range of tileIdx is 0 to NumTilesInPic - 1 (including end values), the list FirstCtbAddrTs[ tileIdx ] represents the conversion from block ID to the CTB address of the first CTB in the block in block scan;The list `ColumnWidthInLumaSamples[i]`, where `i` ranges from 0 to `num_tile_columns_minus1` (inclusive), represents the width of the `i`-th column (in pixels of brightness). The list `RowHeightInLumaSamples[j]`, where `j` ranges from 0 to `num_tile_rows_minus1` (inclusive), represents the height of the `j`-th row (in pixels of brightness). Both `ColumnWidthInLumaSamples[i]` and `RowHeightInLumaSamples[j]` should be greater than 0, with `i` ranging from 0 to `num_tile_columns_minus1` (inclusive) and `j` ranging from 0 to `num_tile_rows_minus1` (inclusive).

[0152] `loop_filter_across_tiles_enabled_flag equal to 1` indicates that in-loop filtering can be performed across tile boundaries in sub-images referencing SPPS. `loop_filter_across_tiles_enabled_flag equal to 0` indicates that in-loop filtering is not performed across tile boundaries in sub-images referencing SPPS. In-loop filtering operations include deblocking filtering, adaptive sampling offset filtering, and adaptive loop filtering. When it does not exist, the value of `loop_filter_across_tiles_enabled_flag` is inferred to be 1. `loop_filter_across_subpic_enabled_flag equal to 1` indicates that in-loop filtering can be performed across sub-image boundaries in sub-images referencing SPPS. `loop_filter_across_subpic_enabled_flag equal to 0` indicates that in-loop filtering is not performed across sub-image boundaries in sub-images referencing SPPS. In-loop filtering operations include deblocking filtering, adaptive sampling offset filtering, and adaptive loop filtering. If it does not exist, then it is inferred that the value of loop_filter_across_subpic_enabled_flag is equal to the value of loop_filter_across_tiles_enabled_flag.

[0153] The general block group header syntax and semantics are as follows:

[0154] In all block headers of the decoded image, the values ​​of the block header syntax elements `tile_group_pic_parameter_set_id` and `tile_group_pic_order_cnt_lsb` should be the same. In all block headers of the decoded sub-image, the value of the block header syntax element `tile_group_subpic_id` should be the same. `tile_group_subpic_id` identifies the sub-image to which the block group belongs. The length of `tile_group_subpic_id` is `subpic_id_len_minus1 + 1` bits. `tile_group_subpic_parameter_set_id` represents the value of `spps_subpic_parameter_set_id` for the currently used SPPS. The value range of `tile_group_spps_parameter_set_id` should be from 0 to 63 (inclusive).

[0155] Derive the following variables and use them to override the corresponding variables derived from the activation SPS: PicWidthInLumaSamples = SubpicWidthInLumaSamples[tile_group_subpic_id] PicHeightInLumaSamples = PicHeightInLumaSamples[tile_group_subpic_id] SubPicRightBorderInPic = SubPicRightBorderInPic[tile_group_subpic_id] SubPicBottomBorderInPic = SubPicBottomBorderInPic[tile_group_subpic_id] PicWidthInCtbsY = SubPicWidthInCtbsY[tile_group_subpic_id] PicHeightInCtbsY = SubPicHeightInCtbsY[tile_group_subpic_id] PicSizeInCtbsY = SubPicSizeInCtbsY[tile_group_subpic_id] PicWidthInMinCbsY = SubPicWidthInMinCbsY[tile_group_subpic_id] PicHeightInMinCbsY = SubPicHeightInMinCbsY[tile_group_subpic_id] PicSizeInMinCbsY = SubPicSizeInMinCbsY[tile_group_subpic_id] PicSizeInSamplesY = SubPicSizeInSamplesY[tile_group_subpic_id] PicWidthInSamplesC = SubPicWidthInSamplesC[tile_group_subpic_id] PicHeightInSamplesC = SubPicHeightInSamplesC[tile_group_subpic_id] The syntax of the decoding tree unit is as follows:

[0156]

[0157] The syntax and semantics of the decoded quadtree are as follows:

[0158] qt_split_cu_flag[ x0 ][ y0 ] indicates whether the decoding unit is divided into decoding units with half the horizontal size and half the vertical size. The array indices x0, y0 represent the position (x0, y0) of the upper-left luminance pixel point of the considered decoding block relative to the upper-left luminance pixel point of the image. When qt_split_cu_flag[ x0 ][ y0 ] does not exist, the following applies: If one or more of the following conditions are true, then it is inferred that the value of qt_split_cu_flag[ x0 ][ y0 ] is equal to 1. Otherwise, if treeType is equal to DUAL_TREE_CHROMA or greater than MaxBtSizeY, then x0 + ( 1<<log2CbSize ) is greater than SubPicRightBorderInPic, and ( 1<<log2CbSize ) is greater than MaxBtSizeC. Otherwise, if treeType is equal to DUAL_TREE_CHROMA or greater than MaxBtSizeY, then y0 + ( 1<<log2CbSize ) is greater than SubPicBottomBorderInPic, and ( 1<<log2CbSize ) is greater than MaxBtSizeC.

[0159] Otherwise, if treeType is equal to DUAL_TREE_CHROMA or greater than MinQtSizeY, if all of the following conditions are true, then it is inferred that the value of qt_split_cu_flag[ x0 ][ y0 ] is equal to 1: x0 + ( 1<<log2CbSize ) is greater than SubPicRightBorderInPic, y0 + ( 1<<log2CbSize ) is greater than SubPicBottomBorderInPic, ( 1<<log2CbSize ) is greater than MinQtSizeC. Otherwise, it is inferred that the value of qt_split_cu_flag[ x0 ][ y0 ] is equal to 0.

[0160] The multi-type tree syntax and semantics are as follows:

[0161] A value of 0 for `mtt_split_cu_flag` indicates that no decoding unit is divided. A value of 1 for `mtt_split_cu_flag` indicates that a binary partition is used to divide the decoding unit into two decoding units, or a ternary partition is used to divide the decoding unit into three decoding units, as indicated by the syntax element `mtt_split_cu_binary_flag`. Binary or ternary partitions can be vertical or horizontal, as indicated by the syntax element `mtt_split_cu_vertical_flag`. When `mtt_split_cu_flag` does not exist, its value is inferred as follows: If one or more of the following conditions are true, then `mtt_split_cu_flag` is inferred to be equal to 1: `x0 + cbWidth` is greater than `SubPicRightBorderInPic`, and `y0 + cbHeight` is greater than `SubPicBottomBorderInPic`. Otherwise, the value of `mtt_split_cu_flag` is inferred to be equal to 0.

[0162] The derivation of temporal brightness motion vector prediction is as follows. The output of this process is: a motion vector prediction `mvLXCol` with a 1 / 16 fractional sampling precision, and an availability flag `availableFlagLXCol`. The variable `currCb` represents the current brightness decoded block at the brightness position (xCb, yCb). The derivation of variables `mvLXCol` and `availableFlagLXCol` is as follows: If `tile_group_temporal_mvp_enabled_flag` equals 0, or if the reference image is the current image, then both components of `mvLXCol` are set to 0, and `availableFlagLXCol` is set to 0. Otherwise (`tile_group_temporal_mvp_enabled_flag` equals 1, and the reference image is not the current image), the following steps are executed sequentially. The derivation of the bottom-right juxtaposed motion vector is as follows: xColBr = xCb + cbWidth (8-355) yColBr = yCb + cbHeight (8-356) If yCb >> CtbLog2SizeY equals yColBr >> CtbLog2SizeY, yColBr is less than SubPicBottomBorderInPic, and xColBr is less than SubPicRightBorderInPic, then the following applies. The variable colCb represents a luminance decoding block that covers the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the juxtaposed image represented by ColPic. The luminance position (xColCb, yColCb) is set to be equal to the top-left pixel of the juxtaposed luminance decoding block represented by colCb relative to the top-left luminance pixel of the juxtaposed image represented by ColPic. The juxtaposition motion vector derivation process is invoked, taking currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag as inputs and assigning the outputs to mvLXCol and availableFlagLXCol. Otherwise, both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.

[0163] The derivation process for the temporal triangulation candidates is as follows. The variables mvLXColC0, mvLXColC1, availableFlagLXColC0, and availableFlagLXColC1 are derived as follows: If tile_group_temporal_mvp_enabled_flag equals 0, then the components of mvLXColC0 and mvLXColC1 are all set to 0, and the components of availableFlagLXColC0 and availableFlagLXColC1 are also set to 0. Otherwise (tile_group_temporal_mvp_enabled_flag equals 1), the following steps are executed sequentially. The derivation of the lower right juxtaposed motion vector mvLXColC0 is as follows: xColBr = xCb + cbWidth (8-392) yColBr = yCb + cbHeight (8-393) If yCb >> CtbLog2SizeY equals yColBr >> CtbLog2SizeY, yColBr is less than SubPicBottomBorderInPic, and xColBr is less than SubPicRightBorderInPic, then the following applies. The variable colCb represents a luminance decoding block that covers the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the juxtaposed image represented by ColPic. The luminance position (xColCb, yColCb) is set to be equal to the top-left pixel of the juxtaposed luminance decoding block represented by colCb relative to the top-left luminance pixel of the juxtaposed image represented by ColPic. The juxtaposition motion vector derivation process is invoked, taking currCb, colCb, (xColCb, yColCb), refIdxLXC0, and sbFlag as inputs and assigning the outputs to mvLXColC0 and availableFlagLXColC0. Otherwise, both components of mvLXColC0 are set to 0, and availableFlagLXColC0 is set to 0.

[0164] The derivation process of the constructed affine control point motion vector fusion candidate is as follows. The fourth (lower right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list use flag predFlagLXCorner[3] and availability flag availableFlagCorner[3] (where X is 0 and 1) are derived as follows. The reference index refIdxLXCorner[3] (where X is 0 or 1) of the time fusion candidate is set to 0. The variables mvLXCol and availableFlagLXCol (X is 0 or 1) are derived as follows. If tile_group_temporal_mvp_enabled_flag is equal to 0, then both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the following applies: xColBr = xCb + cbWidth (8-566) yColBr = yCb + cbHeight (8-567) If yCb >> CtbLog2SizeY equals yColBr >> CtbLog2SizeY, yColBr is less than SubPicBottomBorderInPic, and xColBr is less than SubPicRightBorderInPic, then the following applies. The variable colCb represents a luminance decoding block that covers the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the juxtaposed image represented by ColPic. The luminance position (xColCb, yColCb) is set to be equal to the top-left pixel of the juxtaposed luminance decoding block represented by colCb relative to the top-left luminance pixel of the juxtaposed image represented by ColPic. The process of juxtaposing motion vectors is invoked, taking currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag as inputs and assigning the outputs to mvLXCol and availableFlagLXCol. Otherwise, both components of mvLXCol are set to 0, and availableFlagLXCol is also set to 0. All occurrences of pic_width_in_luma_samples are replaced with PicWidthInLumaSamples. All occurrences of pic_height_in_luma_samples are replaced with PicHeightInLumaSamples.

[0165] In the second exemplary embodiment, the sequence parameter set RBSP syntax and semantics are as follows.

[0166]

[0167] The increment of 1 in `subpic_id_len_minus1` indicates the number of bits used to represent the syntax element `subpic_id[i]` in the SPS. `spps_subpic_id` in the SPS references the SPS, and `tile_group_subpic_id` in the chunk group header also references the SPS. The value of `subpic_id_len_minus1` should range from Ceil(Log2(num_subpic_minus1 + 3)) to 8 (inclusive). Subpics [i] from 0 to num_subpic_minus1 (inclusive) should not overlap, which is a requirement for bitstream consistency. Each subpic can be a time-motion-constrained subpic.

[0168] The semantics of the general tile group header are as follows. `tile_group_subpic_id` identifies the sub-image to which the tile group belongs. The length of `tile_group_subpic_id` is `subpic_id_len_minus1 + 1` bits. A `tile_group_subpic_id` equal to 1 indicates that the tile group does not belong to any sub-image.

[0169] In the third exemplary embodiment, the NAL unit header syntax and semantics are as follows.

[0170]

[0171] `nuh_subpicture_id_len` represents the number of bits used in the syntax element to represent the subpicture ID. When the value of `nuh_subpicture_id_len` is greater than 0, the first `nuh_subpicture_id_len-th` bits in the last `nuh_reserved_zero_4 bits` represent the subpicture ID to which the payload of the NAL unit belongs. When `nuh_subpicture_id_len` is greater than 0, the value of `nuh_subpicture_id_len` should be equal to the value of `subpic_id_len_minus1` for activating SPS. The value of `nuh_subpicture_id_len` for non-VCL NAL units is constrained as follows: if `nal_unit_type` is equal to `SPS_NUT` or `PPS_NUT`, then `nuh_subpicture_id_len` should be equal to 0. `nuh_reserved_zero_3 bits` should be equal to "000". The decoder should ignore (e.g., remove and discard from the bitstream) NAL units whose `nuh_reserved_zero_3 bits` value is not equal to "000".

[0172] In the fourth exemplary embodiment, the sub-image nesting syntax is as follows.

[0173]

[0174] `all_sub_pictures_flag` equal to 1 indicates that the nested SEI message applies to all subpictures. The subpicture to which the nested SEI message is applied is explicitly indicated by subsequent syntax elements. `nesting_num_sub_pictures_minus1` plus 1 indicates the number of subpictures to which the nested SEI message is applied. `nesting_sub_picture_id[i]` represents the subpicture ID of the i-th subpicture to which the nested SEI message is applied. The `nesting_sub_picture_id[i]` syntax element is represented by the Ceil(Log2(nesting_num_sub_pictures_minus1 + 1)) bits. `sub_picture_nesting_zero_bit` should be equal to 0.

[0175] Figure 9 This is a schematic diagram of an exemplary video decoding device 900. As described herein, the video decoding device 900 is suitable for implementing the disclosed examples / implementations. The video decoding device 900 includes a downlink port 920, an uplink port 950, and / or a transceiver unit (Tx / Rx) 910, including a transmitter and / or receiver for transmitting data upstream and / or downstream over a network. The video decoding device 900 also includes a processor 930, which includes a logic unit and / or a central processing unit (CPU) for processing data and a memory 932 for storing data. The video decoding device 900 may also include electrical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the uplink port 950 and / or the downlink port 920 for transmitting data over an electrical, optical, or wireless communication network. The video decoding device 900 may also include input and / or output (I / O) devices 960 for transmitting data with a user. I / O device 960 may include output devices, such as a display for showing video data, a speaker for outputting audio data, etc. I / O device 960 may also include input devices, such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with these output devices.

[0176] Processor 930 is implemented in both hardware and software. Processor 930 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 930 communicates with downlink port 920, Tx / Rx 910, uplink port 950, and memory 932. Processor 930 includes a decoding module 914. Decoding module 914 implements the embodiments disclosed above, such as methods 100, 1000, 1100, and / or mechanism 700, and can use bitstream 500, image 600, and / or image 800. Decoding module 914 can also implement any other methods / mechanisms described herein. Furthermore, decoding module 914 can implement codec system 200, encoder 300, and / or decoder 400. For example, decoding module 914 can be used to indicate and / or acquire sub-image positions and sizes in SPS. In another example, the decoding module 914 can constrain the width and height of sub-images to multiples of the CTU size, unless these sub-images are located on the right or bottom boundary of the image, respectively. In another example, the decoding module 914 can constrain sub-images to cover the image without gaps or overlaps. In another example, the decoding module 914 can be used to indicate and / or acquire data indicating that some sub-images are time-motion-constrained sub-images while others are not. In another example, the decoding module 914 can indicate the complete set of sub-image IDs in the SPS and include the sub-image ID in each strip header to indicate the sub-image including the corresponding strip. In another example, the decoding module 914 can indicate the level of each sub-image. Therefore, the decoding module 914 enables the video decoding device 900 to have additional functions, avoiding certain processing during video data segmentation and decoding to reduce processing overhead and / or improve decoding efficiency. Thus, the decoding module 914 improves the functionality of the video decoding device 900 and solves problems specific to the field of video decoding. Furthermore, the decoding module 914 can transform the video decoding device 900 into different states. Alternatively, the decoding module 914 can be implemented as instructions (e.g., a computer program product stored in a non-transient medium) stored in the memory 932 and executed by the processor 930.

[0177] Memory 932 includes one or more memory types, such as disks, tape drives, solid-state drives, read-only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), static random-access memory (SRAM), etc. Memory 932 can be used as an overflow data storage device to store programs when a program is selected for execution, and to store instructions and data read during program execution.

[0178] Figure 10 A flowchart of an exemplary method 1000 for encoding a bitstream (e.g., bitstream 500 and / or sub-bitstream 501) of sub-images (e.g., sub-images 522, 523, 622, 722 and / or 822) using adaptive size constraints. Method 1000 can be used by an encoder (e.g., encoding / decoding system 200, encoder 300 and / or video decoding device 900) when performing method 100.

[0179] Method 1000 can begin when the encoder receives a video sequence comprising multiple images and determines, for example, to encode the video sequence into a bitstream based on user input. The video sequence is segmented into images / frames for further segmentation before encoding. In step 1001, the images are segmented into multiple sub-images. Adaptive size constraints are applied when segmenting the sub-images. Each sub-image includes a sub-image width and a sub-image height. When the right boundary of the current sub-image does not coincide with the right boundary of the image, the sub-image width of the current sub-image is constrained to an integer multiple of the CTU (e.g., sub-image 822 except 822b). Therefore, when the right boundary of each sub-image coincides with the right boundary of the image, the width of at least one sub-image in the sub-image (e.g., sub-image 822b) may not be an integer multiple of the CTU size. Furthermore, when the lower boundary of the current sub-image does not coincide with the lower boundary of the image, the sub-image height of the current sub-image is constrained to an integer multiple of the CTU (e.g., sub-image 822 except 822c). Therefore, when the lower boundary of each sub-image does not coincide with the lower boundary of the image, the height of at least one sub-image included in the sub-image (e.g., sub-image 822c) may not be an integer multiple of the CTU size. This adaptive size constraint allows for the segmentation of sub-images from the image, where the image width and / or image height are not integer multiples of the CTU size. The CTU size can be measured in luminance pixels.

[0180] In step 1003, one or more sub-images are encoded into a bitstream. In step 1005, the bitstream is stored for transmission to the decoder. The bitstream can then be sent to the decoder as needed. In some examples, sub-bitstreams can be extracted from the encoded bitstream. In this case, the transmitted bitstream is the sub-bitstream. In other examples, the encoded bitstream can be transmitted for sub-bitstream extraction in the decoder. In yet another example, no sub-bitstream extraction is required; the encoded bitstream can be decoded and displayed. In any of these examples, adaptive size constraints allow for sub-image segmentation from images whose height or width is not a multiple of the CTU size, thus enhancing the encoder's capabilities.

[0181] Figure 11 This is a flowchart of an exemplary method 1100 for decoding a bitstream (e.g., bitstream 500 and / or sub-bitstream 501) of sub-images (e.g., sub-images 522, 523, 622, 722, and / or 822) using adaptive size constraints. Method 1100 can be used by a decoder (e.g., codec system 200, decoder 400, and / or video decoding device 900) when performing method 100. For example, method 1100 can be used to decode a bitstream created by method 1000.

[0182] Method 1100 may begin with the decoder starting to receive a bitstream including sub-images. The bitstream may include a complete video sequence, or the bitstream may be a sub-bitstream comprising a reduced set of sub-images for individual extraction. In step 1101, the bitstream is received. The bitstream comprises one or more sub-images segmented from an image according to adaptive size constraints. Each sub-image includes a sub-image width and a sub-image height. When the right boundary of the current sub-image does not coincide with the right boundary of the image, the sub-image width of the current sub-image is constrained to an integer multiple of the CTU (e.g., sub-image 822 other than 822b). Therefore, when the right boundary of each sub-image coincides with the right boundary of the image, the width of at least one sub-image included in the sub-image (e.g., sub-image 822b) may not be an integer multiple of the CTU size. Furthermore, when the lower boundary of the current sub-image does not coincide with the lower boundary of the image, the sub-image height of the current sub-image is constrained to an integer multiple of the CTU (e.g., sub-image 822 other than 822c). Therefore, when the lower boundary of each sub-image does not coincide with the lower boundary of the image, the height of at least one sub-image included in the sub-image (e.g., sub-image 822c) may not be an integer multiple of the CTU size. This adaptive size constraint allows for the segmentation of sub-images from the image, where the image width and / or image height are not integer multiples of the CTU size. The CTU size can be measured in luminance pixels.

[0183] In step 1103, the bitstream is parsed to obtain one or more sub-images. In step 1105, the one or more sub-images are decoded to create a video sequence. The video sequence can then be sent for display. Therefore, the adaptive size constraint allows for sub-image segmentation from images whose height or width is not a multiple of the CTU size. Thus, the decoder can use sub-image-based functions, such as individual sub-image extraction and / or display, on images whose height or width is not a multiple of the CTU size. Therefore, the application of the adaptive size constraint enhances the functionality of the decoder.

[0184] Figure 12 This is a schematic diagram of an exemplary system 1200 that indicates the bitstream (e.g., bitstream 500 and / or sub-bitstream 501) of sub-images (e.g., sub-images 522, 523, 622, 722, and / or 822) using adaptive size constraints. System 1200 can be implemented by an encoder and a decoder (e.g., encoding / decoding system 200, encoder 300, decoder 400, and / or video decoding device 900). Furthermore, system 1200 can be used to implement methods 100, 1000, and / or 1100.

[0185] System 1200 includes a video encoder 1202. The video encoder 1202 includes a segmentation module 1201 for segmenting an image into multiple sub-images, such that the width of each sub-image is an integer multiple of the CTU size, provided that the right boundary of each sub-image does not coincide with the right boundary of the image. The video encoder 1202 also includes an encoding module 1203 for encoding one or more sub-images into a bitstream. The video encoder 1202 further includes a storage module 1205 for storing the bitstream for transmission to a decoder. The video encoder 1202 also includes a transmission module 1207 for transmitting the bitstream, including the sub-images, to the decoder. The video encoder 1202 can also be used to perform any step of method 1000.

[0186] System 1200 also includes a video decoder 1210. The video decoder 1210 includes a receiving module 1211 for receiving a bitstream comprising one or more sub-images segmented from an image, such that the width of each sub-image is an integer multiple of the coding tree unit (CTU) size, provided that the right boundary of each sub-image does not coincide with the right boundary of the image. The video decoder 1210 also includes a parsing module 1213 for parsing the bitstream to obtain one or more sub-images. The video decoder 1210 also includes a decoding module 1215 for decoding one or more sub-images to create a video sequence. The video decoder 1210 also includes a sending module 1217 for sending the video sequence for display. The video decoder 1210 can also be used to perform any step of method 1100.

[0187] When there is no intermediate component between the first component and the second component other than a line, trace, or other medium, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component other than a line, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupling" and its synonyms include direct coupling and indirect coupling. Unless otherwise stated, the term "about" means a range including ±10% of the following quantity.

[0188] It should also be understood that the steps of the exemplary methods described herein do not necessarily need to be performed in the order described, and the order of the steps of these methods should be understood as merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, these methods may include other steps, and some steps may be omitted or combined.

[0189] While several embodiments have been provided in this invention, it is to be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of the invention. These examples are to be regarded as illustrative rather than limiting and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0190] Furthermore, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments, without departing from the scope of the invention, may be combined or integrated with other systems, components, techniques, or methods. Other examples of changes, substitutions, and modifications can be determined by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. A decoding method, characterized in that, The method includes: The received bitstream includes encoded data of one or more sub-images obtained by segmenting an image. When the width of the image is not an integer multiple of the CTU size, the right boundary of the image includes an incomplete CTU; or when the height of the image is not an integer multiple of the CTU size, the bottom boundary of the image includes an incomplete CTU. The bitstream also includes a sequence-level parameter set (SPS), which includes the sub-image identifier ID, sub-image position, and sub-image size of the current image. The current sub-image is one of the one or more sub-images. Parse the bitstream to obtain the sub-image ID, sub-image position, and sub-image size of the current sub-image; The current sub-image is decoded based on its sub-image ID, sub-image position, and sub-image size.

2. The method according to claim 1, characterized in that, When the lower boundary of the sub-image does not coincide with the lower boundary of the image, the sub-image height is an integer multiple of the CTU size.

3. The method according to claim 1 or 2, characterized in that, When the right boundary of a sub-image does not coincide with the right boundary of the image, the width of the sub-image is an integer multiple of the CTU size.

4. The method according to any one of claims 1 to 3, characterized in that, When the lower boundary of a sub-image coincides with the lower boundary of the image, the height of the sub-image is not an integer multiple of the CTU size.

5. The method according to any one of claims 1 to 4, characterized in that, When the right boundary of a sub-image coincides with the right boundary of the image, the width of the sub-image is not an integer multiple of the CTU size.

6. An encoding method, characterized in that, The method includes: The image is divided into one or more sub-images; when the width of the image is not an integer multiple of the CTU size, the right boundary of the image includes an incomplete CTU; or when the height of the image is not an integer multiple of the CTU size, the bottom boundary of the image includes an incomplete CTU. Encode the one or more sub-images into a bitstream; The sequence-level parameter set (SPS) is encoded into the bitstream, wherein the SPS includes the sub-image identifier ID, sub-image position, and sub-image size of the current image; the current sub-image is one of the one or more sub-images.

7. The method according to claim 6, characterized in that, When the lower boundary of the sub-image does not coincide with the lower boundary of the image, the sub-image height is an integer multiple of the CTU size.

8. The method according to claim 6 or 7, characterized in that, When the right boundary of a sub-image does not coincide with the right boundary of the image, the width of the sub-image is an integer multiple of the CTU size.

9. The method according to any one of claims 6 to 8, characterized in that, When the lower boundary of a sub-image coincides with the lower boundary of the image, the height of the sub-image is not an integer multiple of the CTU size.

10. The method according to any one of claims 6 to 9, characterized in that, When the right boundary of a sub-image coincides with the right boundary of the image, the width of the sub-image is not an integer multiple of the CTU size.

11. A video decoding device, characterized in that, include: A processor and a memory, the processor being configured to perform the method according to any one of claims 1 to 5 or 6 to 10.

12. A non-transitory computer-readable medium, characterized in that, It includes computer-executable instructions that, when executed by a processor, cause a video decoding device to perform the method according to any one of claims 1 to 5 or 6 to 10.

13. A decoder, characterized in that, include: A receiving module is used to receive a bitstream, which includes encoded data of one or more sub-images obtained by segmenting an image. When the width of the image is not an integer multiple of the CTU size, the right boundary of the image includes an incomplete CTU; or when the height of the image is not an integer multiple of the CTU size, the bottom boundary of the image includes an incomplete CTU. The bitstream also includes a sequence-level parameter set (SPS), which includes the sub-image identifier ID, sub-image position, and sub-image size of the current image. The current sub-image is one of the one or more sub-images. The parsing module is used to parse the bitstream to obtain the sub-image ID, sub-image position, and sub-image size of the current sub-image; The decoding module is used to decode the current sub-image based on its sub-image ID, sub-image position, and sub-image size.

14. The decoder according to claim 13, characterized in that, The decoder is also used to perform the method according to any one of claims 2 to 5.

15. An encoder, characterized in that, include: A segmentation module is used to segment an image into one or more sub-images; when the width of the image is not an integer multiple of the CTU size, the right boundary of the image includes an incomplete CTU; or when the height of the image is not an integer multiple of the CTU size, the lower boundary of the image includes an incomplete CTU. An encoding module is used to encode the one or more sub-images into a bitstream; The encoding module is used to encode the sequence-level parameter set (SPS) into the bitstream. The SPS includes the sub-image identifier ID, sub-image position, and sub-image size of the current image. The current sub-image is one of the one or more sub-images.

16. The encoder according to claim 15, characterized in that, The encoder is also used to perform the method according to any one of claims 7 to 10.

17. A method for storing a bitstream, characterized in that, The method includes: The bitstream is received, comprising encoded data of one or more sub-images obtained by segmenting an image. When the width of the image is not an integer multiple of the CTU size, the right boundary of the image includes an incomplete CTU; or when the height of the image is not an integer multiple of the CTU size, the bottom boundary of the image includes an incomplete CTU. The bitstream also includes a sequence-level parameter set (SPS), which includes a sub-image identifier ID, sub-image position, and sub-image size for the current image. The current sub-image is one of the one or more sub-images. The bitstream is stored in one or more memories.

18. A device for storing bitstreams, characterized in that, The system includes a storage medium and a communication interface for receiving a bitstream. The bitstream includes encoded data of one or more sub-images obtained by segmenting an image. When the width of the image is not an integer multiple of the CTU size, the right boundary of the image includes an incomplete CTU; or when the height of the image is not an integer multiple of the CTU size, the bottom boundary of the image includes an incomplete CTU. The bitstream also includes a sequence-level parameter set (SPS), which includes a sub-image identifier ID, sub-image position, and sub-image size for the current image. The current sub-image is one of the one or more sub-images. The storage medium is used to store the bitstream.

19. A computer program product, characterized in that, The computer program product includes a computer executable program, which, when executed by a computer or processor, performs the method as described in any one of claims 1 to 5 or 6 to 10.