Method and device for signaling sub-picture partition information - Patents.com

By using the sub-image existence flag in video encoding to signal sub-image segmentation information, the problem of difficulty in efficient signal sub-image segmentation information in the prior art is solved, and the encoding efficiency and decoding quality are improved.

JP7675718B2Active Publication Date: 2025-05-13ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022535057
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2020-12-18
Publication Date
2025-05-13
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively segment information of signal sub-images in video encoding, resulting in limited encoding efficiency and decoding quality.

Method used

By determining whether sub-image information is included in the bitstream, sub-picture information present flag is used to determine whether to signal the segmentation information of the sub-image, including the number, width, height, position, identification map, etc. of the sub-image.

Benefits of technology

The encoding efficiency and decoding quality in video encoding are improved, and information transmission and processing in the encoding process are optimized by segmenting information by effective signal sub-images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675718000001
    Figure 0007675718000001
  • Figure 0007675718000002
    Figure 0007675718000002
  • Figure 0007675718000003
    Figure 0007675718000003
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for signaling sub-picture split information. An example method includes determining whether a bitstream includes sub-picture information according to a sub-picture information present flag signaled in the bitstream, and signaling in the bitstream at least one of the number of sub-pictures in a picture, the width, height, position and identifier (ID) mapping of a target sub-picture, a subpic_treated_as_pic_flag, and a loop_filter_across_subpic_enabled_flag in response to the bitstream including the sub-picture information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 954,014, filed December 27, 2019, the entirety of which is incorporated by reference herein.

[0002] Technical Field

[0002] The present disclosure relates generally to video processing, and more particularly, to a method and apparatus for signaling sub-picture split information. [Background technology]

[0003] background

[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, that specify certain video coding formats. As more and more advanced video coding techniques are adopted into video standards, the coding efficiency of the new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]

[0004]

[0004] In some embodiments, an exemplary method for signaling subpicture split information includes determining whether the bitstream includes subpicture information according to a subpicture information present flag signaled in the bitstream, and in response to the bitstream including subpicture information, signaling in the bitstream at least one of the number of subpictures in the picture, the width, height, position and identifier (ID) mapping of the target subpicture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag.

[0005] In some embodiments, an exemplary video processing device includes at least one memory for storing instructions and at least one processor, the at least one processor being configured to execute instructions to cause the device to determine whether the bitstream includes sub-picture information according to a sub-picture information present flag signaled in the bitstream, and in response to the bitstream including the sub-picture information, signal in the bitstream at least one of a number of sub-pictures in a picture, a width, height, position and an identifier (ID) mapping of a target sub-picture, a subpic_treated_as_pic_flag, and a loop_filter_across_subpic_enabled_flag.

[0006] In some embodiments, an exemplary non-transitory computer-readable storage medium stores a set of instructions executable by one or more processing devices to cause a video processing device to determine whether the bitstream includes sub-picture information according to a sub-picture information present flag signaled in the bitstream, and in response to the bitstream including sub-picture information, signal in the bitstream at least one of a number of sub-pictures in a picture, a width, height, position and an identifier (ID) mapping of a target sub-picture, a subpic_treated_as_pic_flag, and a loop_filter_across_subpic_enabled_flag.

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which the various features illustrated are not drawn to scale. [Brief description of the drawings]

[0008] [Figure 1]

[0008] FIG. 1 is a schematic diagram illustrating the structure of an exemplary video sequence according to some embodiments of the present disclosure. [Figure 2A]

[0009] 1 shows a schematic diagram of an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 2B]

[0010] 4 shows a schematic diagram of another example encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 3A]

[0011] 2 shows a schematic diagram of an exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 3B]

[0012] 4 shows a schematic diagram of another example decoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 4]

[0013] 1 shows a block diagram of an exemplary device for encoding or decoding video in accordance with some embodiments of the present disclosure. [Diagram 5]

[0014] 1 is a schematic diagram illustrating an example of a picture divided into coding tree units (CTUs) according to some embodiments of the present disclosure. [Figure 6]

[0015] FIG. 2 is a schematic diagram illustrating an example of a picture divided into tiles and raster scan slices according to some embodiments of the present disclosure. [Figure 7]

[0016] FIG. 2 is a schematic diagram illustrating an example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure. [Figure 8]

[0017] FIG. 2 is a schematic diagram illustrating another example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure. [Figure 9]

[0018] FIG. 2 is a schematic diagram illustrating an example of a picture being divided into sub-pictures according to some embodiments of the present disclosure. [Figure 10]

[0019] 1 shows an example Table 1 illustrating an example sequence parameter set (SPS) syntax for sub-picture partitioning according to some embodiments of the present disclosure. [Figure 11]

[0020] 1 shows an example Table 2 illustrating an example SPS syntax for a sub-picture identifier according to some embodiments of the present disclosure. [Figure 12]

[0021] 1 shows an example Table 3 illustrating an example Picture Parameter Set (PPS) syntax for a sub-picture identifier according to some embodiments of the present disclosure. [Figure 13]

[0022] 1 shows an example Table 4 illustrating an example picture header (PH) syntax for a sub-picture identifier according to some embodiments of the present disclosure. [Figure 14]

[0023] 1 is a schematic diagram illustrating example bitstream adaptation constraints according to some embodiments of the present disclosure. [Figure 15]

[0024] 13 shows an example Table 5 illustrating another example PH syntax for a sub-picture identifier according to some embodiments of the present disclosure. [Figure 16]

[0025] 13 shows an example Table 6 illustrating another example PH syntax for a sub-picture identifier according to some embodiments of the present disclosure. [Figure 17A]

[0026] 7 shows an example Table 7A illustrating an example SPS syntax according to some embodiments of the present disclosure. [Figure 17B]

[0027] 7B illustrates an example Table 7B showing another example SPS syntax according to some embodiments of the present disclosure. [Figure 18]

[0028] 1 illustrates an example Table 8 showing another example SPS syntax according to some embodiments of the present disclosure. [Figure 19]

[0029] 1 illustrates an example Table 9 showing another example SPS syntax according to some embodiments of the present disclosure. [Figure 20]

[0030] 1 illustrates an example Table 10 showing another example SPS syntax according to some embodiments of the present disclosure. [Figure 21]

[0031] 1 illustrates a flowchart of an exemplary video processing method according to some embodiments of the present disclosure. [Figure 22]

[0032] 1 illustrates a flowchart of another exemplary video processing method according to some embodiments of the present disclosure. [Diagram 23]

[0033] 1 illustrates a flowchart of another exemplary video processing method according to some embodiments of the present disclosure. [Figure 24]

[0034] 1 illustrates a flowchart of another exemplary video processing method according to some embodiments of the present disclosure. [Diagram 25]

[0035] 1 illustrates a flowchart of another exemplary video processing method according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Detailed Description

[0036] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations illustrated in the following description of the exemplary embodiments do not represent all implementations in accordance with the present invention. Rather, they are merely examples of apparatus and methods in accordance with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In case of conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0010]

[0037] The ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.

[0011]

[0038] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing techniques that go beyond HEVC using the Joint Search Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0012]

[0039] The VVC standard is a recent development and continues to incorporate more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0013]

[0040] A video is a set of static pictures (or "frames") arranged in a time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in a time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display capability) can be used to display such pictures in a time sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for supervision, conferencing, live broadcast, etc.

[0014]

[0041] To reduce the storage space and transmission bandwidth required by such applications, video may be compressed before storage and transmission, and decompressed before display. Compression and decompression may be performed by software executed by a processor (e.g., a processor of a general-purpose computer) or by specialized hardware. A module for compression is generally referred to as an "encoder," and a module for decompression is generally referred to as a "decoder." The encoder and decoder may be collectively referred to as a "codec." The encoder and decoder may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder may include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithms or processes fixed in a computer-readable medium. Video compression and decompression may be performed by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, or the like. In some applications, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."

[0015]

[0042] A video coding process can identify and keep useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process may be called "lossy". Otherwise, it may be called "lossless". Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.

[0016]

[0043] Useful information of the picture being coded (called the "current picture") includes changes with respect to a reference picture (e.g., a previously coded and reconstructed picture). Such changes may include changes in pixel position, brightness, or color, among which position changes are the most important. Changes in position of a group of pixels representing an object may reflect the motion of the object between the reference picture and the current picture.

[0017]

[0044] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture". A picture that is coded using a previous picture as a reference picture is called a "P-picture". A picture that is coded using both previous and future pictures as reference pictures (i.e., the references are "bidirectional") is called a "B-picture".

[0018]

[0045] 1 illustrates the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. Video sequence 100 may be live video or captured and archived video. Video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). Video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface for receiving video from a video content provider (e.g., a video broadcast transceiver).

[0019]

[0046] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture whose reference picture is picture 102 itself. Picture 104 is a P-picture whose reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture whose reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately preceding or following that picture. For example, the reference picture of picture 104 may be a picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of reference pictures to the examples shown in FIG.

[0020]

[0047] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, they may divide a picture into elementary segments and encode or decode a picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, a picture is divided into 4x4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units can have variable sizes in pictures or any shape and size of pixels, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, etc. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for a picture.

[0021]

[0048] A basic processing unit may be a logical unit that may include a group of different kinds of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing colorless luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements, where the luma and chroma components may have the same size of a basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be performed repeatedly on each of its luma and chroma components.

[0022]

[0049] Video coding has multiple operation stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large for processing, and therefore may be further divided into segments, referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing subunits are also logical units that may include groups of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit may be repeatedly performed on each of its luma and chroma components. Note that such division may be performed to further levels as required for processing. Also note that different stages may use different schemes to divide the basic processing units.

[0023]

[0050] For example, in the mode decision stage (an example of which is shown in FIG. 2B), the encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.

[0024]

[0051] As another example, in the prediction stage (examples of which are shown in FIG. 2A-2B), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which the prediction operations can be performed.

[0025]

[0052] As another example, in the transform stage (examples of which are shown in FIG. 2A-2B), the encoder can perform transform operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level the transform operations can be performed. Note that the division scheme of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0026]

[0053] 1, the basic processing units 112 are further divided into 3×3 basic processing sub-units, the boundaries of which are shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different ways.

[0027]

[0054] In some implementations, to provide parallel processing and error resilience capabilities to video encoding and decoding, a picture can be divided into regions for processing, such that the encoding or decoding process does not depend on information about a region of the picture from any other region of the picture. In other words, each region of the picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, thus increasing the coding efficiency. Also, when data of a region is corrupted during processing or lost during network transmission, the codec can also correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience capabilities. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different pictures of the video sequence 100 can have different partitioning schemes for dividing the picture into regions.

[0028]

[0055] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines inside structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0029]

[0056] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 in FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to the structure 110 in FIG. 1, each original picture of the video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform the process 200A at the level of the basic processing units for each original picture of the video sequence 202. For example, the encoder may perform the process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions of each original picture of video sequence 202 (eg, regions 114-118).

[0030]

[0057] 2A , the encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224, which is used in a prediction stage 204 for the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A may be referred to as a “reconstruction path.” The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0031]

[0058] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate (in the reconstruction path) a prediction reference 224 for encoding the next original BPU of the original picture. After encoding all the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0032]

[0059] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, acquiring, getting, obtaining, reading, accessing, or any act in any manner for inputting data.

[0033]

[0060] In the prediction step 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, perform a prediction operation, and generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting the prediction data 206, which can be used to reconstruct the original BPU as a predicted BPU 208 from the prediction data 206 and the prediction reference 224.

[0034]

[0061] Ideally, the predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate the residual BPU 210. For example, the encoder may subtract the values ​​(e.g., grayscale or RGB values) of pixels of the predicted BPU 208 from the values ​​of corresponding pixels of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of such a subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, which may be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0035]

[0062] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transform coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a change frequency (e.g., frequency of luminance change) component of the residual BPU 210. No basis pattern can be reconstructed from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the changes in the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0036]

[0063] Different transform algorithms may use different basis patterns. For example, various transform algorithms may be used in transform stage 212, such as discrete cosine transform, discrete sine transform, or the like. The transform in transform stage 212 is invertible. That is, the encoder may recover residual BPU 210 by inverting the transform (referred to as an "inverse transform"). For example, to recover pixels of residual BPU 210, the inverse transform may multiply the values ​​of corresponding pixels of the basis pattern by their associated coefficients and add up the products to generate a weighted sum. For video coding standards, both the encoder and the decoder may use the same transform algorithm (and thus the same basis pattern). Therefore, the encoder may record only the transform coefficients, and the decoder may reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to residual BPU 210, the transform coefficients may have fewer bits, but they may be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0037]

[0064] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore high-frequency change information without significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate the quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and the transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-value quantized transform coefficients 216, which further compresses the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").

[0038]

[0065] The quantization stage 214 may be lossy because the encoder ignores such division remainders in rounding operations. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits are needed for the quantized transform coefficients 216. To obtain different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0039]

[0066] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, besides the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, a type of transformation in the transformation stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may generate a video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0040]

[0067] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0041]

[0068] It should be noted that other variations of the process 200A may also be used to encode the video sequence 202. In some embodiments, the stages of the process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of the process 200A may be combined into a single stage. In some embodiments, a single stage of the process 200A may be split into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, the process 200A may include additional stages. In some embodiments, the process 200A may omit one or more stages in FIG. 2A.

[0042]

[0069] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. The process 200B may be modified from the process 200A. For example, the process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared to the process 200A, the forward path of the process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of the process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0043]

[0070] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels from one or more already coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions from one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0044]

[0071] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs in the same picture that are coded (in the forward path) and reconstructed (in the reconstruction path). The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. The extrapolation technique may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform the extrapolation at a pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 can include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, or the like.

[0045]

[0072] As another example, in the temporal prediction stage 2044, the encoder may perform inter prediction. For an original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference picture (referred to as a "search window"). The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may be centered at a location in the reference picture that has the same coordinates as the original BPU in the current picture, and may extend outward for a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., by using a pixel-recursive algorithm, a block matching algorithm, or the like), the encoder may determine such a region as a matching region. The matching region may have different dimensions (e.g., smaller than, equal to, larger than, or of a different shape) than the original BPU. Because the reference picture and the current picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region may be considered to "move" toward the location of the original BPU as time progresses. The encoder may record the direction and distance of such movement as a "motion vector." When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the encoder may search for the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder may weight the pixel values ​​of the matching region in each matching reference picture.

[0046]

[0073] Motion estimation may be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. For inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, a number of reference pictures, weights associated with the reference pictures, or the like.

[0047]

[0074] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move the matching region of the reference picture according to the motion vector, in which case the encoder may predict the original BPU of the current picture. When multiple reference pictures are used (e.g., like picture 106 in FIG. 1), the encoder may move the matching region of the reference picture according to the respective motion vectors and average the pixel values ​​of the matching region. In some embodiments, if the encoder weights the pixel values ​​of the matching region of each matching reference picture, the encoder may add a weighted sum of pixel values ​​to the moved matching region.

[0048]

[0075] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures of the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter predicted picture where the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter predicted picture where the reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0049]

[0076] Still referring to the forward path of the process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in a mode decision stage 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of the process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction modes and the distortion of the reconstructed reference picture under such candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding prediction BPU 208 and prediction data 206.

[0050]

[0077] Within the reconstruction path of the process 200B, if an intra prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current BPU encoded and reconstructed in the current picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by the inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as, for example, deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop filtered reference picture may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode parameters of the loop filter (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.

[0051]

[0078] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be similar to video sequence 202. However, due to information loss in the compression and decompression process (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, in which case the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0052]

[0079] In FIG. 3A, the decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as an “encoded BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may provide the prediction reference 224 to a prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0053]

[0080] The decoder may iteratively perform the process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0054]

[0081] In the binary decoding stage 302, the decoder may perform an inverse operation of the binary encoding technique used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder may depacketize the video bitstream 228 before providing it to the binary decoding stage 302.

[0055]

[0082] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0056]

[0083] In the process 300B, for a coding basic processing unit (referred to as a "current BPU") of a coding picture being decoded (referred to as a "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various kinds of data depending on what prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, or the like. The parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, or the like. As another example, if inter prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, or the like. Parameters for the inter prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors respectively associated with the matching regions, or the like.

[0057]

[0084] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or perform temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224 as described in FIG. 3A.

[0058]

[0085] In the process 300B, the decoder can provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform the prediction operation in the next iteration of the process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture decoded by all the BPUs), the encoder can provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply the loop filter to the prediction reference 224 in the manner as described in FIG. 2B. The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., to be used as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include parameters of a loop filter (e.g., loop filter strength).

[0059]

[0086] FIG. 4 is a block diagram of an exemplary device 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 may become a specialized machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any number and combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system on a chip (SoC), an application specific integrated circuit (ASIC), or the like. In some embodiments, the processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in FIG. 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0060]

[0087] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for performing steps in a process 200A, 200B, 300A, or 300B) as well as data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access (e.g., via bus 410) the program instructions and data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number and combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. The memory 404 may also be a group of memories (not shown in FIG. 4) grouped as a single logical entity.

[0061]

[0088] Bus 410 may be a communication device that transfers data between components internal to device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0062]

[0089] For ease of explanation and without creating ambiguity, the processor 402 and other data processing circuitry are collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single independent module, or may be fully or partially combined with any other components of the device 400.

[0063]

[0090] The device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, the network interface 406 may include any number and combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0064]

[0091] In some embodiments, optionally, the apparatus 400 may further include a peripheral interface 408 for providing a connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., an input interface coupled to a camera or a video archive), or the like.

[0065]

[0092] It should be noted that a video codec (e.g., a codec performing processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within device 400. For example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as specialized data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).

[0066]

[0093] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIG. 2A or 2B, inverse quantization 218 in FIG. 3A or 3B), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for coding of a picture or slice may be signaled at a high level, for example, using the init_qp_minus26 syntax element in the picture parameter set (PPS) and using the slice_qp_delta syntax element in the slice header. Furthermore, the QP value can be adapted at a local level per CU using delta QP values ​​sent with the granularity of the quantization group.

[0067]

[0094] In the disclosed embodiments, to code a frame, a picture is divided into a sequence of coding tree units (CTUs). Multiple CTUs may form a tile, slice, or sub-picture. A picture is divided into a sequence of CTUs. For a picture with three sample arrays, a CTU consists of an N×N block of luma samples with corresponding two blocks of chroma samples. Figure 5 shows an example of a picture divided into multiple CTUs according to some embodiments of the present disclosure.

[0068]

[0095] According to some embodiments, the maximum allowed size of a luma block in a CTU is specified to be 128x128 (although the maximum size of a luma transform block may be 64x64), and the minimum allowed size of a luma block in a CTU is specified to be 32x32.

[0069]

[0096] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that cover a rectangular area of ​​the picture. A slice contains an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. Two slice modes may be supported: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a sequence of complete tiles within a raster scan of tiles of a picture. In rectangular slice mode, a slice contains several complete tiles that collectively form a rectangular area of ​​the picture, or several consecutive complete CTU rows of one tile that collectively form a rectangular area of ​​the picture. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular area corresponding to the slice.

[0070]

[0097] A subpicture contains one or more slices that collectively cover a rectangular area of ​​the picture.

[0071]

[0098] 6 illustrates an example of a picture divided into tiles and raster scan slices according to some embodiments of the present disclosure. As shown in FIG. 6, the picture is divided into 12 tiles (4 tile rows and 3 tile columns) and 3 raster scan slices.

[0072]

[0099] 7 illustrates an example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure. As shown in FIG 7, the picture is divided into 20 tiles (5 tile rows and 4 tile columns) and 9 rectangular slices.

[0073]

[0100] 8 illustrates another example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure. As shown in FIG. 8, the picture is divided into four tiles (two tile rows and two tile columns) and four rectangular slices.

[0074]

[0101] Figure 9 illustrates an example of a picture divided into sub-pictures according to some embodiments of the present disclosure. As shown in Figure 9, the picture is divided into 20 tiles (5 tile columns and 4 tile rows), 12 on the left side each covering one slice of a 4x4 CTU, and 8 on the right side each covering two vertically stacked slices of a 2x2 CTU, resulting in a total of 28 slices of various dimensions and 28 sub-pictures (each slice is a sub-picture).

[0075]

[0102] According to some disclosed embodiments, sub-picture split information is signaled in a sequence parameter set (SPS). Figure 10 shows an example Table 1 illustrating an example SPS syntax for sub-picture split according to some embodiments of the present disclosure.

[0076]

[0103] In Table 1, the syntax element sps_num_subpics_minus1 plus 1 specifies the number of subpictures in a picture, the syntax elements subpic_ctu_top_left_x[i] and subpic_ctu_top_left_y[i] specify the position of the top-left CTU of the i-th subpicture in CtbSizeY units, and the syntax elements subpic_width_minus1[i] plus 1 and subpic_height_minus1[i] plus 1 specify the width and height, respectively, of the i-th subpicture in CtbSizeY units. The semantics of these syntax elements are as follows:

[0077]

[0104] A subpics_present_flag equal to 1 specifies that the subpicture parameters are present in the SPS RBSP syntax, and a subpics_present_flag equal to 0 specifies that the subpicture parameters are not present in the SPS RBSP syntax.

[0078]

[0105] sps_num_subpics_minus1 plus 1 specifies the number of subpictures. The syntax element sps_num_subpics_minus1 is in the range of 0 to 254. If not present, the value of the syntax element sps_num_subpics_minus1 is inferred to be equal to 0.

[0079]

[0106] subpic_ctu_top_left_x[i] specifies the horizontal location of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0080]

[0107] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0081]

[0108] subpic_width_minus1[i] plus 1 specifies the width of the i-th subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1.

[0082]

[0109] subpic_height_minus1[i] plus 1 specifies the height of the i-th subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1.

[0083]

[0110] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each coded picture in the Coded Layer Video Sequence (CLVS) is treated as a picture in the decoding process that excludes an in-loop filtering operation. The syntax element subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each coded picture in the CLVS is not treated as a picture in the decoding process that excludes an in-loop filtering operation. If absent, the value of the syntax element subpic_treated_as_pic_flag[i] is inferred to be equal to 0.

[0084]

[0111] The syntax element loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations can be performed across the i-th subpicture boundary in each coded picture in the CLVS. The syntax element loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not performed across the i-th subpicture boundary in each coded picture in the CLVS. If not present, the value of the syntax element loop_filter_across_subpic_enabled_pic_flag[i] equal to 1 is inferred.

[0085]

[0112] According to the disclosed embodiments of the present disclosure, each sub-picture may be assigned an identifier. Sub-picture identifier information may be signaled in a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header (PH). Figure 11 shows an example Table 2 illustrating an example SPS syntax of a sub-picture identifier according to some embodiments of the present disclosure. Figure 12 shows an example Table 3 illustrating an example PPS syntax of a sub-picture identifier according to some embodiments of the present disclosure. Figure 13 shows an example Table 4 illustrating an example PH syntax of a sub-picture identifier according to some embodiments of the present disclosure.

[0086]

[0113] As shown in Tables 2 to 4, the syntax element sps_subpic_id_present_flag indicates whether subpicture ID mapping is in the SPS, the syntax elements sps_subpic_id_signaling_present_flag, pps_subpic_id_signaling_present_flag, and ph_subpic_id_signaling_present_flag indicate whether subpicture ID mapping is signaled in the SPS, PPS, or PH, respectively, and the syntax elements sps_subpic_id_len_minus1 plus 1, syntax element pps_subpic_id_len_minus1 plus 1, and syntax element ph_subpic_id_len_minus1 plus 1 specify the number of bits used to present syntax elements sps_subpic_id[i], pps_subpic_id[i], and ph_subpic_id[i], respectively, which are the subpicture IDs signaled in the SPS, PPS, and PH, respectively.

[0087]

[0114] The semantics of the above syntax elements and associated bitstream conformance requirements are described below.

[0088]

[0115] The syntax element sps_subpic_id_present_flag equal to 1 specifies that subpicture ID mapping is present in the SPS, and the syntax element sps_subpic_id_present_flag equal to 0 specifies that subpicture ID mapping is not present in the SPS.

[0089]

[0116] The syntax element sps_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the SPS, and the syntax element sps_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the SPS. If absent, the value of the syntax element sps_subpic_id_signaling_present_flag is inferred to be equal to 0.

[0090]

[0117] sps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sps_subpic_id[i]. The value of the syntax element sps_subpic_id_len_minus1 may be in the range of 0 to 15.

[0091]

[0118] sps_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element sps_subpic_id[i] is sps_subpic_id_len_minus1 + 1 bits. If absent and if the syntax element sps_subpic_id_present_flag is equal to 0, the value of the syntax element sps_subpic_id[i] is inferred to be equal to i for each i in the range 0 to sps_num_subpics_minus1.

[0092]

[0119] The syntax element pps_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the PPS. The syntax element pps_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the PPS. If the syntax element sps_subpic_id_present_flag is 0 or if the syntax element sps_subpic_id_signaling_present_flag is equal to 1, the syntax element pps_subpic_id_signaling_present_flag may be equal to 0.

[0093]

[0120] pps_num_subpics_minus1 plus 1 specifies the number of subpictures in the coded picture that refer to the PPS. It may be a bitstream conformance requirement that the value of the syntax element pps_num_subpic_minus1 be equal to the syntax element sps_num_subpics_minus1.

[0094]

[0121] pps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element pps_subpic_id[i]. The value of the syntax element pps_subpic_id_len_minus1 is in the range of 0 to 15. It may be a bitstream conformance requirement that the value of the syntax element pps_subpic_id_len_minus1 is the same for all PPSs referenced by coded pictures in the CLVS.

[0095]

[0122] pps_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element pps_subpic_id[i] is pps_subpic_id_len_minus1+1 bits.

[0096]

[0123] The syntax element ph_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the PH, and the syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the PH.

[0097]

[0124] ph_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element ph_subpic_id[i]. The value of the syntax element pic_subpic_id_len_minus1 may be in the range of 0 to 15. It may be a bitstream conformance requirement that the value of the syntax element ph_subpic_id_len_minus1 is the same for all PHs referenced by coded pictures in the CLVS.

[0098]

[0125] ph_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element ph_subpic_id[i] is ph_subpic_id_len_minus1+1 bits.

[0099]

[0126] After parsing these syntax elements related to subpicture IDs, the subpicture ID list SubpicIdList is derived using the following syntax (1): for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? Syntax(1) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (ph_subpic_id_signaling_present_flag?ph_subpic_id[i]:pps_subpic_id[i])):i

[0100]

[0127] However, there are some problems with the above signaling of subpicture split. First, according to the semantics, both syntax elements sps_subpic_id_present_flag and sps_subpic_id_signaling_present_flag specify whether the subpicture ID is in the SPS or not. According to Table 2, the subpicture id information is signaled in the SPS only if both of these two syntax elements are true. So there is redundancy in the signaling. Second, even if sps_subpics_id_present_flag is true, the subpicture ID can still be signaled in the PPS or PH if pps_subpic_id_signalling_present_flag or ph_subpic_id_signalling_present_flag is true. Third, according to syntax (1), if sps_subpic_id_present_flag is not true, a default ID equal to the subpicture index is assigned to each subpicture. If sps_subpic_id_present_flag is true, then SubpicIdList[i] is derived as sps_subpic_id[i], ph_subpic_id[i], or pps_subpic_id[i]. However, if sps_subpic_id_present_flag is true, then the subpicture ID may not be in the SPS, PPS, or PH, and therefore, in that case, an undefined value of the syntax element pps_subpic_id is specified in SubpicIdList. In syntax (1), if the syntax element sps_subpic_id_present_flag is true, then the syntax elements sps_subpic_id_signaling_present_flag and ph_subpic_id_signaling_present_flag are both false, and the syntax element pps_subpic_id[i] is specified in SubpicIdList[i] regardless of the value of the syntax element pps_subpic_id_signaling_present_flag. If the syntax element ps_subpic_id_signaling_present_flag is false, the syntax element pps_subpic_id[i] is undefined.

[0101]

[0128] Furthermore, as shown in Table 1 of Figure 10, if the syntax element subpics_present_flag is true, the number of subpictures is signaled first, followed by the top-left position, width and height of each subpicture and two control flags subpic_treated_as_pic_flag and loop_filter_across_subpic_enabled_flag. The top-left position, width and height and these two control flags are also signaled if there is only one subpicture (if the syntax element sps_num_subpics_minus1 is equal to 0). However, if there is only one subpicture in a picture, there is no need to indicate these things, since the subpicture is equal to the picture, and therefore the signaled information can be derived from the picture itself.

[0102]

[0129] Furthermore, a sub-picture is obtained by splitting a picture, and a picture is formed by merging all sub-pictures. The position and size of the last sub-picture can be derived from the size of the whole picture and the positions and sizes of all previous sub-pictures. Therefore, there is no need to signal the position, width and height information of the last sub-picture.

[0103]

[0130] Furthermore, as shown in Table 2 of Figure 11, the syntax element sps_subpic_id_present_flag is always signaled regardless of the value of the syntax element subpics_present_flag. Therefore, in the above signaling method, a subpicture identifier may still be signaled even if there is no subpicture, which is meaningless.

[0104]

[0131] The present disclosure provides a signaling method for solving the above problems, and several exemplary embodiments are described in detail below.

[0105]

[0132] In some embodiments of the present disclosure, it is possible to force signaling of the sub-picture ID in the picture header if the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false, thereby avoiding cases where the sub-picture ID is not defined.

[0106]

[0133] For example, a bitstream conformance constraint can be imposed in two ways: In the first way, the semantics for the bitstream conformance constraint are as follows (emphasis in italics): ph_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the PH. Syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the PH. If syntax element sps_subpic_id_present_flag is equal to 1, syntax element sps_subpic_id_signaling_present_flag is equal to 0, and syntax element pps_subpic_id_signaling_present_flag is equal to 0, then a value of 1 for the syntax element ph_subpic_id_signaling_present_flag may be a bitstream conformance requirement.

[0107]

[0134] In the second method, the semantics for the bitstream conformance constraints are as follows (emphasis in italics): ph_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the PH. Syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the PH. If the syntax element sps_subpic_id_present_flag is equal to 1, the syntax element sps_subpic_id_signaling_present_flag is equal to 0, and the syntax element pps_subpic_id_signaling_present_flag is equal to 0 in all of the PPSs referenced by a coded picture in the CLVS, then it may be a bitstream conformance requirement that of all PHs referenced by a coded picture in the CLVS references, there is at least one PH whose value of the syntax element ph_subpic_id_signaling_present_flag is equal to 1.

[0108]

[0135] FIG. 14 is a schematic diagram illustrating an example bitstream adaptation constraint of this second method, according to some embodiments of the present disclosure.

[0109]

[0136] The semantics of the syntax element sps_subpic_id_present_flag is not clearly defined in the current VVC draft and can be changed as follows (emphasis in italics):

[0110]

[0137] The syntax element sps_subpic_id_present_flag equal to 1 specifies that the subpicture ID mapping is in the SPS, PPS, or PH. The syntax element sps_subpic_id_present_flag equal to 0 specifies that the subpicture ID mapping is not in the SPS, PPS, or PH.

[0111]

[0138] This may ensure that the syntax element ph_subpic_id is signaled when the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false. This syntax is shown in Tables 2-4 of Figures 11-13.

[0112]

[0139] In the above embodiment, if the sub-picture ID present flag is true (syntax element sps_subpic_id_present_flag=1), the sub-picture ID to be used is signaled in the bitstream (in the SPS, PPS or PH) and no inference rules are required. If the sub-picture ID present flag is true (syntax element sps_subpic_id_present_flag=1), then forcing signaling of the sub-picture ID in one of the SPS, PPS or PH may be better than deriving the sub-picture ID using inference rules without signaling the sub-picture ID in the bitstream, since the syntax element sps_subpic_id_present_flag indicates that a sub-picture ID is present.

[0113]

[0140] As another example, Figure 15 illustrates an example Table 5 illustrating another example PH syntax for a sub-picture identifier, according to some embodiments of the present disclosure. Table 5 illustrates a modification (shown in box 1501 and highlighted in italics) of the PH syntax shown in Table 4. With reference to Table 5, if the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false, then signaling of the syntax element ph_subpic_id is forced by inferring that the syntax element ph_sub_pic_id_signaling_present_flag is true.

[0114]

[0141] The syntax element ph_subpic_id_signaling_present_flag can have two alternative semantics (emphasis in italics):

[0115]

[0142] The first set of semantics includes (emphasis in italics): The syntax element ph_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the PH. The syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the PH. If absent, the value of the syntax element ph_subpic_id_signaling_present_flag is inferred to be 1.

[0116]

[0143] The second set of semantics includes (emphasis in italics): The syntax element ph_subpic_id_signaling_present_flag equal to 1 specifies that subpicture ID mapping is signaled in the PH, and the syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that subpicture ID mapping is not signaled in the PH. Otherwise, if the syntax element sps_subpic_id_present_flag is equal to 1 and the syntax element sps_subpic_id_signaling_present_flag is equal to 0, the value of the syntax element ph_subpic_id_signaling_present_flag is inferred to be 1.

[0117]

[0144] The semantics of the syntax element sps_subpic_id_present_flag can be changed as follows (emphasis in italics):

[0118]

[0145] The syntax element sps_subpic_id_present_flag equal to 1 specifies that the subpicture ID mapping is in the SPS, PPS, or PH. The syntax element sps_subpic_id_present_flag equal to 0 specifies that the subpicture ID mapping is not in the SPS, PPS, or PH.

[0119]

[0146] If the syntax element sps_subpic_id_present_flag is true, the subpicture ID inference rule is not needed in some embodiments. Thus, if the subpicture ID present flag is true (syntax element sps_subpic_id_present_flag=1) but the subpicture ID is not signaled in the SPS or PPS (syntax element sps_subpic_id_signaling_present_flag=0 and syntax element pps_subpic_id_signaling_present_flag=0), the signaling of the syntax element ph_subpic_id_signaling_present_flag is skipped. This can save one bit.

[0120]

[0147] As another example, if the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false (i.e. the subpicture ID is not signaled in the SPS or PPS), then signaling of the syntax element ph_subpic_id is forced, which in this case means that the subpicture ID is signaled in the PH. If the syntax element pps_subpic_id is signaled (syntax element sps_subpic_id_present_flag is true, syntax element sps_subpic_id_signaling_present_flag is false, and syntax element pps_subpic_id_signaling_present_flag is true), then the syntax element ph_subpic_id cannot be signaled. FIG. 16 shows an example Table 6 illustrating another example PH syntax for a sub-picture identifier (emphasis in box 1601 highlighted in italics), according to some embodiments of the disclosure.

[0121]

[0148] The list SubpicIdList[i] is derived according to syntax (2) as follows: for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? Syntax(2) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (pps_subpic_id_signaling_present_flag?pps_subpic_id[i]:ph_subpic_id[i])):i

[0122]

[0149] If the syntax element sps_subpic_id_present_flag is true, then the subpicture ID inference rule is not needed in some embodiments. The syntax element ph_subpic_id_signaling_present_flag can be removed. Thus, one bit is saved if the syntax element sps_subpic_id_present_flag is equal to 1, the syntax element sps_subpic_id_signaling_present_flag is equal to 0, and the syntax element pps_subpic_id_signaling_present_flag is equal to 1.

[0123]

[0150] If the sub-picture ID is already signaled in the PPS, some embodiments may give the encoder the option to override the sub-picture ID in the PPS by signaling it again in the PH, which is a more flexible way for the encoder.

[0124]

[0151] In some embodiments of this disclosure, inference rules are provided to ensure that the subpicture ID list SubpicIdList can be derived.

[0125]

[0152] As an example, if a subpicture ID is not signaled in the SPS, PPS or PH, an inference rule is provided for the syntax element pps_subpic_id to derive the subpicture ID list SubpicIdList using a default value of the syntax element pps_subpic_id that is inferred by the inference rule.

[0126]

[0153] The semantics of the syntax element pps_subpic_id are as follows (emphasis in italics): pps_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element pps_subpic_id[i] is pps_subpic_id_len_minus1 + 1 bits. If absent, the value of the syntax element pps_subpic_id[i] is inferred to be i for each i in the range 0 to pps_num_subpics_minus1.

[0127]

[0154] As another example, an inference rule is given in the derivation process of SubpicIdList: If the syntax element sps_subpic_id_present_flag is true and the syntax elements sps_subpic_id_signaling_present_flag, pps_subpic_id_signaling_present_flag, and ph_subpic_id_signaling_present_flag are all false, then a default value is assigned to SubpicIdList[i].

[0128]

[0155] The derivation of SubpicIdList follows syntax (3) as follows (emphasis in italics): for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? Syntax(3) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (ph_subpic_id_signaling_present_flag?ph_subpic_id[i]: (pps_subpic_id_signaling_present_flag?pps_subpic_id[i]:i))):i

[0129]

[0156] In some embodiments, by imposing inference rules on any pps_subpic_id, it is possible to ensure that the SubpicIdList can be derived even without signaling any subpicture IDs in the bitstream, thus saving bits devoted to signaling subpicture IDs.

[0130]

[0157] In some embodiments of the present disclosure, the derivation rules for the subpicture ID list SubpicIdList may be modified to give higher priority to subpicture IDs signaled in the PPS than to the PH, so that the syntax element pps_subpic_id_signaling_present_flag is checked before the syntax element ph_subpic_id_signaling_present_flag.

[0131]

[0158] As an example, the derivation rule for SubpicIdList follows syntax (4) shown below (emphasis in italics), where an inference rule is given for the syntax element ph_subpic_id. for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? Syntax(4) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (pps_subpic_id_signaling_present_flag?pps_subpic_id[i]:ph_subpic_id[i])):i

[0132]

[0159] The semantics according to syntax (4) (emphasis in italics) are as follows: ph_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element ph_subpic_id[i] is ph_subpic_id_len_minus1+1 bits. If absent, the value of the syntax element ph_subpic_id[i] is inferred to be i for each i in the range from 0 to ph_num_subpics_minus1.

[0133]

[0160] As another example, the derivation rule SubpicIDList follows syntax (5) shown below (emphasis in italics), and in this example there are no additional inference rules for the syntax element ph_subpic_id. for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? Syntax(5) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (pps_subpic_id_signaling_present_flag?pps_subpic_id[i]: (ph_subpic_id_signaling_present_flag?ph_subpic_id[i]:i))):i

[0134]

[0161] Inference rules are provided for the derivation process of syntax elements ph_subpic_id or SubpicIdList to ensure that the SubpicIdList can be correctly derived without signaling any subpicture IDs in the bitstream, thus saving bits devoted to subpicture signaling if a default subpicture ID inferred by the inference rules is used.

[0135]

[0162] In some embodiments of the present disclosure, redundant information signaled regarding sub-pictures when the number of sub-pictures is equal to one may be removed.

[0136]

[0163] As an example, the SPS syntax is shown in Table 7A of Figure 17A (highlighted in italics with emphasis in boxes 1701-1702) or Table 7B of Figure 17B (highlighted in italics with emphasis in boxes 1711-1712). It will be understood that Tables 7A and 7B are equivalent. The semantics that follow the syntax in Tables 7A and 7B (highlighted in italics) are shown below.

[0137]

[0164] subpic_ctu_top_left_x[i] specifies the horizontal location of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0138]

[0165] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0139]

[0166] subpic_width_minus1[i] plus 1 specifies the width of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1, where "Ceil()" is the function to round up to the nearest integer. Thus, Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0140]

[0167] subpic_height_minus1[i] plus 1 specifies the height of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1, where "Ceil()" is the function to round up to the nearest integer. Thus, Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0141]

[0168] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each coded picture in the CLVS is treated as a picture in the decoding process that excludes an in-loop filtering operation. The syntax element subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each coded picture in the CLVS is not treated as a picture in the decoding process that excludes an in-loop filtering operation. Otherwise, if the syntax element subpics_present_flag is equal to 1 and the syntax element sps_num_subpics_minus1 is equal to 0, the value of the syntax element subpic_treated_as_pic_flag[i] is inferred to be equal to 1, otherwise the value of the syntax element subpic_treated_as_pic_flag[i] is inferred to be equal to 0.

[0142]

[0169] The syntax element loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that an in-loop filtering operation can be performed across the i-th subpicture boundary in each coded picture in the CLVS. The syntax element loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that an in-loop filtering operation is not performed across the i-th subpicture boundary in each coded picture in the CLVS. Otherwise, if the syntax element subpics_present_flag is equal to 1 and the syntax element sps_num_subpics_minus1 is equal to 0, the value of the syntax element loop_filter_across_subpic_enabled_flag[i] is inferred to be equal to 0, otherwise the value of the syntax element loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0143]

[0170] As another example, the SPS syntax is shown in Table 8 of Figure 18 (highlighted in italics with emphasis in boxes 1801-1802). The semantics that follow the syntax in Table 8 (highlighted in italics) are shown below.

[0144]

[0171] subpic_ctu_top_left_x[i] specifies the horizontal location of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0145]

[0172] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0146]

[0173] subpic_width_minus1[i] plus 1 specifies the width of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil((pic_width_max_in_luma_samples÷CtbSizeY)-1, where "Ceil()" is the function to round up to the nearest integer. Thus, Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0147]

[0174] subpic_height_minus1[i] plus 1 specifies the height of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1, where "Ceil()" is the function to round up to the nearest integer. Thus, Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0148]

[0175] In some embodiments of the present disclosure, the position and / or size information of the last subpicture may be skipped and derived from the size of the full picture and the sizes and positions of all previous subpictures. Figure 19 shows an example Table 9 illustrating another example SPS syntax, according to some embodiments of the present disclosure. In Table 9 (highlighted in italics with emphasis in boxes 1901-1902), the width and height of the last subpicture, which is the subpicture with index equal to syntax element sps_num_subpics_minus1, is skipped.

[0149]

[0176] The width and height of the last subpicture are derived from the width and height of the entire picture and the top left position of the last subpicture.

[0150]

[0177] The following semantics follow those of Table 9 (emphasis in italics):

[0151]

[0178] subpic_ctu_top_left_x[i] specifies the horizontal location of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0152]

[0179] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If absent, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0153]

[0180] subpic_width_minus1[i] plus 1 specifies the width of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil((pic_width_max_in_luma_samples)÷CtbSizeY)-1-(i==sps_num_subpics_minus1?subpic_ctu_top_left_x[sps_num_subpics_minus1]:0), where "Ceil()" is the function to round up to the nearest integer. That is, Ceil(pic_width_max_in_luma_samples÷CtbSizeY) is equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division. "sps_num_subpics_minus1" is the number of subpictures in the picture. For the last subpicture in a picture, i is inferred to be equal to sps_num_subpics_minus1, in which case subpic_width_minus1[i] is inferred to be equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_x[sps_num_subpics_minus1]. If there is only one subpicture in a picture, i can only be 0 and subpic_ctu_top_left_x[0] is 0. Therefore, subpic_width_minus1[i] is inferred to be equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_x[0] or (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1.

[0154]

[0181] subpic_height_minus1[i] plus 1 specifies the height of the ith subpicture in CtbSizeY units. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_height_minus1[sps_num_subpics_minus1] is inferred to be equal to (Ceil(pic_height_max_in_luma_samples)÷CtbSizeY)-1-(i==sps_num_subpics_minus1?subpic_ctu_top_left_y[i]:0), where "Ceil()" is the function to round up to the nearest integer. That is, Ceil(pic_height_max_in_luma_samples÷CtbSizeY) is equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division. "sps_num_subpics_minus1" is the number of subpictures in the picture. For the last subpicture in a picture, i is inferred to be equal to sps_num_subpics_minus1, in which case subpic_height_minus1[i] is inferred to be equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_y[sps_num_subpics_minus1]. If there is only one subpicture in a picture, i can only be 0 and subpic_ctu_top_left_y[0] is 0. Therefore, subpic_width_minus1[i] is inferred to be equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_y[0] or (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1.

[0155]

[0182] In some embodiments of the present disclosure, the sub-picture ID is signaled only if there is a sub-picture. Figure 20 shows an example Table 10 illustrating another example SPS syntax (emphasis in boxes 2001-2002 highlighted in italics) according to some embodiments of the present disclosure.

[0156]

[0183] FIG. 21 shows a flowchart of an exemplary video processing method 2100 according to some embodiments of the present disclosure. The method 2100 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 2100. In some embodiments, the method 2100 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).

[0157]

[0184] In step 2101, it may be determined whether sub-picture ID mapping is in the bitstream. In some embodiments, the method 2100 may include signaling a flag indicating whether sub-picture ID mapping is in the bitstream. For example, the flag may be sps_subpic_id_present_flag shown in Table 2 of FIG. 11, Table 4 of FIG. 13, Table 5 of FIG. 15, or Table 6 of FIG. 16.

[0158]

[0185] In step 2103, it may be determined whether one or more sub-picture IDs are signaled in the first syntax or the second syntax. In step 2105, in response to determining that there is a sub-picture ID mapping and determining that the one or more sub-picture IDs are not signaled in the first syntax and the second syntax, the one or more sub-picture IDs are signaled in a third syntax. The first syntax, the second syntax, or the third syntax is one of SPS, PPS, and PH. For example, the first syntax, the second syntax, and the third syntax are SPS, PPS, and PH, respectively. Thus, if there is a sub-picture ID mapping (e.g., sps_subpic_id_present_flag=1), the sub-picture IDs may be forced to be signaled in SPS, PPS, or PH.

[0159]

[0186] In some embodiments, method 2100 may include signaling a first flag (e.g., sps_subpic_id_signaling_present_flag shown in Table 2 of FIG. 11, Table 5 of FIG. 15, or Table 6 of FIG. 16) indicating that one or more sub-picture IDs are signaled in a first syntax (e.g., SPS shown in Table 2 of FIG. 11). In some embodiments, method 2100 may include signaling a second flag (e.g., pps_subpic_id_signaling_present_flag shown in Table 3 of FIG. 12, Table 5 of FIG. 15, or Table 6 of FIG. 16) indicating that one or more sub-picture IDs are signaled in a second syntax (e.g., PPS shown in Table 3 of FIG. 12). In some embodiments, the method 2100 may include signaling a third flag (e.g., ph_subpic_id_signaling_present_flag shown in Table 5 of FIG. 15) indicating that one or more subpicture IDs are signaled within a third syntax (e.g., PH shown in Table 5 of FIG. 15).

[0160]

[0187] In some embodiments, the method 2100 may include determining whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled in the third syntax, and in response to the bitstream not including the third flag, signaling the one or more sub-picture IDs in the third syntax. For example, the third syntax may be PH, and the third flag may be ph_subpic_id_signaling_present_flag. If ph_subpic_id_signaling_present_flag is not signaled in PH, then ph_subpic_id_signaling_present_flag may be inferred to be 1, and one or more sub-picture IDs are signaled in PH.

[0161]

[0188] In some embodiments, the method 2100 may include signaling one or more sub-picture IDs in the second syntax and the third syntax (e.g., Table 5 of FIG. 15 ) in response to determining that the one or more sub-picture IDs are not signaled in the first syntax.

[0162]

[0189] 22 illustrates a flowchart of an exemplary video processing method 2200 according to some embodiments of the present disclosure. The method 2200 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 2200. In some embodiments, the method 2200 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).

[0163]

[0190] At step 2201, it may be determined whether one or more sub-picture IDs are signaled in at least one of an SPS, a PH, or a PPS. In some embodiments, method 2200 may include determining whether one or more sub-picture IDs are signaled in a PH before determining whether one or more sub-picture IDs are signaled in a PPS. In some embodiments, method 2200 may include determining whether one or more sub-picture IDs are signaled in a PPS before determining whether one or more sub-picture IDs are signaled in a PH.

[0164]

[0191] At step 2203, in response to determining that one or more sub-picture IDs are not signaled in the SPS, PH, and PPS, it may be determined that the one or more sub-picture IDs have default values.

[0165]

[0192] 23 shows a flowchart of an exemplary video processing method 2300 according to some embodiments of the present disclosure. The method 2300 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 2300. In some embodiments, the method 2300 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).

[0166]

[0193] In step 2301, it may be determined whether the number of subpictures of the coded picture is equal to 1. For example, it may be determined whether sps_num_subpic_minus1 is greater than 0, as shown in Table 8 of FIG.

[0167]

[0194] In step 2303, in response to determining that the number of sub-pictures is equal to 1, the sub-pictures of the coded picture may be treated as pictures in the decoding process. For example, in response to determining that the number of sub-pictures is equal to 1, a flag subpic_treated_as_pic_flag[i] may be inferred to be equal to 1. In some embodiments, in response to determining that the number of sub-pictures is equal to 1, an in-loop filtering operation may be omitted.

[0168]

[0195] FIG. 24 illustrates a flowchart of an exemplary video processing method 2400 according to some embodiments of the present disclosure. The method 2400 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 2400. In some embodiments, the method 2400 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).

[0169]

[0196] In step 2401, it may be determined whether the subpicture is the last subpicture of the picture. In step 2403, in response to the subpicture being determined to be the last subpicture, position or size information of the subpicture may be derived from the size of the picture and the size and position of the previous subpicture of the picture.

[0170]

[0197] FIG. 25 illustrates a flowchart of an exemplary video processing method 2500 according to some embodiments of the present disclosure. The method 2500 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 2500. In some embodiments, the method 2500 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).

[0171]

[0198] In step 2501, it may be determined whether a subpicture is present in the picture. For example, this determination may be made based on a flag (e.g., subpics_present_flag shown in Table 10 of FIG. 20).

[0172]

[0199] In step 2503, in response to determining that one or more sub-pictures are in the picture, a first flag may be signaled indicating whether a sub-picture ID mapping is in the SPS. For example, the first flag may be sps_subpic_id_present_flag shown in Table 10 of FIG.

[0173]

[0200] In some embodiments, the method 2500 may include signaling a second flag indicating whether the sub-picture ID mapping is signaled in the SPS in response to the first flag indicating that the sub-picture ID mapping is in the SPS. In response to the second flag indicating that the sub-picture ID mapping is signaled in the SPS, a sub-picture ID of one or more sub-pictures may be signaled in the SPS. For example, the second flag may be sps_subpic_id_signaling_present_flag shown in Table 10 of FIG. 20. If sps_subpic_id_signaling_present_flag is true, sps_subpic_id[i] may be signaled.

[0174]

[0201] The embodiments may be further described using the following clauses: 1. A video processing method comprising: determining whether a sub-picture ID mapping is in the bitstream; determining whether one or more sub-picture IDs are signaled in the first syntax or the second syntax; and signaling the one or more sub-picture IDs in a third syntax in response to determining that there is a sub-picture ID mapping and determining that the one or more sub-picture IDs are not signaled in the first syntax and the second syntax. A video processing method comprising: 2. The method according to clause 1, wherein the first syntax, the second syntax, or the third syntax is one of a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH). 3. signaling a first flag indicating that one or more sub-picture IDs are signaled within a first syntax; or signaling a second flag indicating that one or more sub-picture IDs are signaled in the second syntax. 3. The method of clauses 1 and 2, further comprising: 4. Signaling a third flag indicating that one or more sub-picture IDs are signaled within the third syntax. 4. The method of any one of clauses 1 to 3, further comprising: 5. Determining whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled within a third syntax; and signaling one or more sub-picture IDs in a third syntax in response to the bitstream not including the third flag. 5. The method of any one of clauses 1 to 4, further comprising: 6. In response to determining that the one or more sub-picture IDs are not signaled in the first syntax, signaling the one or more sub-picture IDs in the second syntax and in the third syntax. 6. The method of any one of clauses 1 to 5, further comprising: 7. Signaling a fourth flag indicating whether a sub-picture ID mapping is in the bitstream. 7. The method of any one of clauses 1 to 6, further comprising: 8. A video processing device, at least one memory for storing instructions; and at least one processor, the at least one processor comprising: determining whether a sub-picture ID mapping is in the bitstream; determining whether one or more sub-picture IDs are signaled in the first syntax or the second syntax; and signaling the one or more sub-picture IDs in a third syntax in response to determining that there is a sub-picture ID mapping and determining that the one or more sub-picture IDs are not signaled in the first syntax and the second syntax. a video processing device configured to execute instructions to cause the device to 9. The device of clause 8, wherein the first syntax, the second syntax, or the third syntax is one of a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH). 10. At least one processor: signaling a first flag indicating that one or more sub-picture IDs are signaled in the first syntax; or signaling a second flag indicating that one or more sub-picture IDs are signaled in the second syntax. 10. An apparatus as described in clauses 8 and 9, configured to execute instructions to cause the apparatus to 11. At least one processor: signaling a third flag indicating that one or more sub-picture IDs are signaled within the third syntax 11. The apparatus of any one of clauses 8 to 10, configured to execute instructions to cause the apparatus to perform the 12. At least one processor: determining whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled within a third syntax; and signaling one or more sub-picture IDs in a third syntax in response to the bitstream not including the third flag. 12. The apparatus of any one of clauses 8 to 11, configured to execute instructions to cause the apparatus to perform the 13. At least one processor: signaling the one or more sub-picture IDs in the second syntax and in the third syntax in response to determining that the one or more sub-picture IDs are not signaled in the first syntax. 13. An apparatus as described in any one of clauses 8 to 12, configured to execute instructions to cause the apparatus to perform the 14. At least one processor: signaling a fourth flag indicating whether a sub-picture ID mapping is present in the bitstream; 14. An apparatus as described in any one of clauses 8 to 13, configured to execute instructions to cause the apparatus to perform the 15. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: determining whether a sub-picture ID mapping is in the bitstream; determining whether one or more sub-picture IDs are signaled in the first syntax or the second syntax; and signaling the one or more sub-picture IDs in a third syntax in response to determining that there is a sub-picture ID mapping and determining that the one or more sub-picture IDs are not signaled in the first syntax and the second syntax. A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including: 16. The non-transitory computer-readable storage medium of clause 15, wherein the first syntax, the second syntax, or the third syntax is one of a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH). 17. A set of instructions is signaling a first flag indicating that one or more sub-picture IDs are signaled in the first syntax; or signaling a second flag indicating that one or more sub-picture IDs are signaled in the second syntax. 17. A non-transitory computer-readable storage medium as described in clauses 15 and 16, executable by one or more processing devices to cause a video processing device to perform the steps of: 18. A set of instructions is signaling a third flag indicating that one or more sub-picture IDs are signaled within the third syntax The non-transitory computer-readable storage medium of any one of clauses 15 to 17, executable by one or more processing devices to cause a video processing device to perform the above. 19. A set of instructions is determining whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled within a third syntax; and signaling one or more sub-picture IDs in a third syntax in response to the bitstream not including the third flag. The non-transitory computer-readable storage medium of any one of clauses 15 to 18, executable by one or more processing devices to cause a video processing device to perform the above. 20. A set of instructions is signaling the one or more sub-picture IDs in the second syntax and in the third syntax in response to determining that the one or more sub-picture IDs are not signaled in the first syntax. The non-transitory computer-readable storage medium of any one of clauses 15 to 19, executable by one or more processing devices to cause a video processing device to perform the above. 21. A set of instructions is signaling a fourth flag indicating whether a sub-picture ID mapping is present in the bitstream; The non-transitory computer-readable storage medium of any one of clauses 15 to 20, executable by one or more processing devices to cause a video processing device to perform the above. 22. A video processing method comprising: determining whether one or more sub-picture IDs are signaled in at least one of a sequence parameter set (SPS), a picture header (PH), or a picture parameter set (PPS); and determining, in response to determining that the one or more sub-picture IDs are not signaled in the SPS, the PH, and the PPS, that the one or more sub-picture IDs have default values. A video processing method comprising: 23. The method of clause 22, wherein determining whether one or more sub-picture IDs are signaled in the PH is performed before determining whether one or more sub-picture IDs are signaled in the PPS. 24. The method of clause 22, wherein determining whether one or more sub-picture IDs are signaled in the PPS is performed before determining whether one or more sub-picture IDs are signaled in the PH. 25. A video processing method comprising: determining whether the number of sub-pictures of the coded picture is equal to 1; and in response to determining that the number of sub-pictures is equal to one, treating the sub-pictures of the coded picture as pictures in the decoding process. A video processing method comprising: 26. In response to determining that the number of sub-pictures is equal to one, excluding an in-loop filtering operation. 26. The method of claim 25, further comprising: 27. A video processing method comprising: determining whether the subpicture is the last subpicture of the picture; and deriving position or size information of the subpicture from the size of the picture and the size and position of a previous subpicture of the picture in response to the subpicture being determined to be the last subpicture; A video processing method comprising: 28. A video processing method comprising: Determining whether a sub-picture is within a picture; and signaling a first flag indicating whether a sub-picture ID mapping is in a sequence parameter set (SPS) in response to determining that one or more sub-pictures are in the picture; A video processing method comprising: 29. In response to the first flag indicating that the sub-picture ID mapping is in the SPS, signaling a second flag indicating whether the sub-picture ID mapping is signaled in the SPS. 29. The method of claim 28, further comprising: 30. In response to the second flag indicating that sub-picture ID mapping is signaled in the SPS, signaling a sub-picture ID of one or more sub-pictures in the SPS. 30. The method of claim 29, further comprising: 31. A video processing device, at least one memory for storing instructions; and at least one processor, the at least one processor comprising: determining whether one or more sub-picture IDs are signaled in at least one of a sequence parameter set (SPS), a picture header (PH), or a picture parameter set (PPS); and determining, in response to determining that the one or more sub-picture IDs are not signaled in the SPS, the PH, and the PPS, that the one or more sub-picture IDs have default values. a video processing device configured to execute instructions to cause the device to At least one processor: Determining whether one or more sub-picture IDs are signaled in a PH before determining whether one or more sub-picture IDs are signaled in a PPS. 32. An apparatus as described in clause 31 configured to execute instructions to cause the apparatus to At least one processor: Determining whether one or more sub-picture IDs are signaled in a PPS before determining whether one or more sub-picture IDs are signaled in a PH. 32. An apparatus as described in clause 31 configured to execute instructions to cause the apparatus to 34. A video processing device, at least one memory for storing instructions; and at least one processor, the at least one processor comprising: determining whether the number of sub-pictures of the coded picture is equal to 1; and in response to determining that the number of sub-pictures is equal to one, treating the sub-pictures of the coded picture as pictures in the decoding process. a video processing device configured to execute instructions to cause the device to At least one processor: excluding an in-loop filtering operation in response to determining that the number of sub-pictures is equal to one; 35. An apparatus as described in clause 34, configured to execute instructions to cause the apparatus to 36. A video processing device, at least one memory for storing instructions; and at least one processor, the at least one processor comprising: determining whether the subpicture is the last subpicture of the picture; and deriving position or size information of the subpicture from the size of the picture and the size and position of a previous subpicture of the picture in response to the subpicture being determined to be the last subpicture; a video processing device configured to execute instructions to cause the device to 37. A video processing device, at least one memory for storing instructions; and at least one processor, the at least one processor comprising: Determining whether a sub-picture is within a picture; and signaling a first flag indicating whether a sub-picture ID mapping is in a sequence parameter set (SPS) in response to determining that one or more sub-pictures are in the picture; a video processing device configured to execute instructions to cause the device to 38. At least one processor: signaling a second flag indicating whether sub-picture ID mapping is signaled in the SPS in response to the first flag indicating that the sub-picture ID mapping is in the SPS. 38. An apparatus as described in clause 37, configured to execute instructions to cause the apparatus to 39. At least one processor: signaling, in response to the second flag indicating that the sub-picture ID mapping is signaled in the SPS, a sub-picture ID of the one or more sub-pictures in the SPS. 39. An apparatus as described in clause 38, configured to execute instructions to cause the apparatus to 40. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: determining whether one or more sub-picture IDs are signaled in at least one of a sequence parameter set (SPS), a picture header (PH), or a picture parameter set (PPS); and determining, in response to determining that the one or more sub-picture IDs are not signaled in the SPS, the PH, and the PPS, that the one or more sub-picture IDs have default values. A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including: 41. A set of instructions is Determining whether one or more sub-picture IDs are signaled in a PH before determining whether one or more sub-picture IDs are signaled in a PPS. 41. A non-transitory computer-readable storage medium as described in clause 40, executable by one or more processing devices to cause a video processing device to perform the steps of: 42. A set of instructions is Determining whether one or more sub-picture IDs are signaled in a PPS before determining whether one or more sub-picture IDs are signaled in a PH. 41. A non-transitory computer-readable storage medium as described in clause 40, executable by one or more processing devices to cause a video processing device to perform the steps of: 43. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: determining whether the number of sub-pictures of the coded picture is equal to 1; and in response to determining that the number of sub-pictures is equal to one, treating the sub-pictures of the coded picture as pictures in the decoding process. A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including: 44. A set of instructions is excluding an in-loop filtering operation in response to determining that the number of sub-pictures is equal to one; 44. A non-transitory computer-readable storage medium as described in clause 43, executable by one or more processing devices to cause a video processing device to perform the steps of: 45. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: determining whether the subpicture is the last subpicture of the picture; and deriving position or size information of the subpicture from the size of the picture and the size and position of a previous subpicture of the picture in response to the subpicture being determined to be the last subpicture; A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including: 46. ​​A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: Determining whether a sub-picture is within a picture; and signaling a first flag indicating whether a sub-picture ID mapping is in a sequence parameter set (SPS) in response to determining that one or more sub-pictures are in the picture; A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including: 47. A set of instructions is signaling a second flag indicating whether sub-picture ID mapping is signaled in the SPS in response to the first flag indicating that the sub-picture ID mapping is in the SPS. 47. A non-transitory computer-readable storage medium as described in clause 46, executable by one or more processing devices to cause a video processing device to perform the steps of: 48. A set of instructions is signaling, in response to the second flag indicating that the sub-picture ID mapping is signaled in the SPS, a sub-picture ID of the one or more sub-pictures in the SPS. 48. A non-transitory computer-readable storage medium as described in clause 47, executable by one or more processing devices to cause a video processing device to perform the steps of: 49. A video processing method comprising: determining whether the bitstream contains sub-picture information according to a sub-picture information present flag signaled in the bitstream; and In response to the bitstream including subpicture information, The number of subpictures in the picture, Target subpicture width, height, position and identifier (ID) mapping, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag Signaling at least one of the following in the bitstream: A video processing method comprising: 50. The method of clause 49, wherein the signaling of at least one of the width, height, and position of the target subpicture is based on the number of subpictures in the picture. 51. If there are at least two subpictures in a picture, signal at least one of subpic_treated_as_pic_flag, loop_filter_across_subpic_enabled_flag, and the width, height, and position of the target subpicture. Further comprising: If there is only one sub-picture in a picture, the signaling of at least one of the width, height and position of the target sub-picture is skipped; subpic_treated_as_pic_flag indicates whether subpictures of each coded picture in the coding layer video sequence (CLVS) are treated as pictures in the decoding process excluding the in-loop filtering operation; and The method of clause 50, wherein loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across subpicture boundaries are enabled across subpicture boundaries of each coded picture in the CLVS. 52. If the target subpicture width is not signaled, determining the value of the target subpicture width as the picture width; and If the target subpicture height is not signaled, determine the value of the target subpicture height as the picture height. 52. The method of claim 51, further comprising: 53. If the target subpicture width is not signaled, determining the value of the target subpicture width in coding tree block (CTB) size units as the picture width in CTB size units; and If the target subpicture height is not signaled, determine the value of the target subpicture height in CTB size units as the picture height in CTB size units. 53. The method of claim 52, further comprising: 54. determining that subpic_treated_as_pic_flag has a value of 1 if subpic_treated_as_pic_flag is not signaled in the bitstream; and If loop_filter_across_subpic_enabled_flag is not signaled in the bitstream, determine that loop_filter_across_subpic_enabled_flag has a value of 0. 52. The method of claim 51, further comprising: 55. Skip signaling of at least one of the width and height of a target subpicture if the target subpicture is the last subpicture in a picture. 50. The method of claim 49, further comprising: 56. If the target subpicture width is not signaled, determining the value of the target subpicture width in coding tree block (CTB) size units as the picture width in CTB size units minus the horizontal position of the top-left coding tree unit (CTU) of the target subpicture in CTB size units, or determining the value of the target subpicture width as the picture width minus the horizontal position of the top-left coding tree unit (CTU) of the target subpicture; and If the target subpicture height is not signaled, determine the target subpicture height value in CTB size units as the picture height in CTB size units minus the vertical position of the top-left CTU of the target subpicture in CTB size units, or determine the target subpicture height value as the picture height minus the vertical position of the top-left CTU of the target subpicture. 56. The method of claim 55, further comprising: 57. Signaling of target subpicture ID mapping: signaling a first flag in the bitstream; and signaling a target sub-picture ID mapping within the first data unit or the second data unit in response to the first flag being equal to one; Further comprising: 49. The method according to claim 49, wherein a first flag equal to 0 indicates that the ID mapping of the target sub-picture is not signaled in the bitstream. 58. In response to the first flag being equal to 1 and the target sub-picture ID mapping not being signaled in the first data unit, signaling the target sub-picture ID mapping in the second data unit; or skipping signaling the target sub-picture ID mapping in the second data unit in response to the first flag being equal to 0 or the target sub-picture ID mapping being signaled in the first data unit. 58. The method of claim 57, further comprising: 59. The method of clause 58, wherein each of the first data unit and the second data unit is one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header (PH). 60. A video processing device, at least one memory for storing instructions; and at least one processor, the at least one processor comprising: determining whether the bitstream contains sub-picture information according to a sub-picture information present flag signaled in the bitstream; and In response to the bitstream including subpicture information, The number of subpictures in the picture, Target subpicture width, height, position and identifier (ID) mapping, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag Signaling at least one of the following in the bitstream: a video processing device configured to execute instructions to cause the device to 61. The apparatus of clause 60, wherein the signaling of at least one of the width, height, and position of the target subpicture is based on the number of subpictures in the picture. 62. At least one processor: If there are at least two subpictures in the picture, signal at least one of subpic_treated_as_pic_flag, loop_filter_across_subpic_enabled_flag, and the width, height, and position of the target subpicture. configured to execute instructions to cause the device to If there is only one sub-picture in a picture, the signaling of at least one of the width, height and position of the target sub-picture is skipped; subpic_treated_as_pic_flag indicates whether subpictures of each coded picture in the coding layer video sequence (CLVS) are treated as pictures in the decoding process excluding the in-loop filtering operation; and The apparatus of claim 61, wherein loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across subpicture boundaries are enabled across subpicture boundaries of each coded picture in the CLVS. At least one processor: if the target sub-picture width is not signaled, determining the value of the target sub-picture width as the picture width; and If the target subpicture height is not signaled, determine the value of the target subpicture height as the picture height. 63. An apparatus as described in clause 62 configured to execute instructions to cause the apparatus to 64. At least one processor: determining the value of the target subpicture width in coding tree block (CTB) size units as the picture width in CTB size units if the target subpicture width is not signaled; and If the target subpicture height is not signaled, determine the value of the target subpicture height in CTB size units as the picture height in CTB size units. 64. An apparatus as described in clause 63 configured to execute instructions to cause the apparatus to 65. At least one processor: determining that subpic_treated_as_pic_flag has a value of 1 if subpic_treated_as_pic_flag is not signaled in the bitstream; and If loop_filter_across_subpic_enabled_flag is not signaled in the bitstream, determine that loop_filter_across_subpic_enabled_flag has a value of 0. 63. An apparatus as described in clause 62 configured to execute instructions to cause the apparatus to 66. At least one processor: skipping signaling of at least one of the width and height of the target subpicture if the target subpicture is the last subpicture in a picture 61. An apparatus as described in clause 60 configured to execute instructions to cause the apparatus to 67. At least one processor: if the target subpicture width is not signaled, determining the value of the target subpicture width in coding tree block (CTB) size units as the picture width in CTB size units minus the horizontal position of the top-left coding tree unit (CTU) of the target subpicture in CTB size units, or determining the value of the target subpicture width as the picture width minus the horizontal position of the top-left coding tree unit (CTU) of the target subpicture; and If the target subpicture height is not signaled, determine the target subpicture height value in CTB size units as the picture height in CTB size units minus the vertical position of the top-left CTU of the target subpicture in CTB size units, or determine the target subpicture height value as the picture height minus the vertical position of the top-left CTU of the target subpicture. 66. An apparatus as described in clause 66 configured to execute instructions to cause the apparatus to 68. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: determining whether the bitstream contains sub-picture information according to a sub-picture information present flag signaled in the bitstream; and In response to the bitstream including subpicture information, The number of subpictures in the picture, Target subpicture width, height, position and identifier (ID) mapping, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag Signaling at least one of the following in the bitstream: A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including:

[0175]

[0202] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes instructions that can be executed by a device (such as the encoder and decoder of the present disclosure) to perform the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0176]

[0203] It should be noted that relational terms herein, such as "first" and "second," are merely used to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the words "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and that the element or elements following any of these words are not meant to be a definitive listing of such elements or elements, or to be limited to only the listed element or elements.

[0177]

[0204] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a database can include A or B, then the database can include A or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if it is stated that a database can include A, B, or C, then the database can include A, B, or C, or A and B, A and C, or B and C, or A and B and C, unless specifically stated otherwise or impracticable.

[0178]

[0205] It is understood that the above-described embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the methods of the present disclosure. The computational units and other functional units described in the present disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple ones of the above-described modules / units can be combined into one module / unit, and each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0179]

[0206] In the above specification, the embodiments have been described with reference to many specific details that may vary from implementation to implementation. Certain adaptations and modifications of the above-described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as examples only, with the true scope and spirit of the invention being indicated by the appended claims. It is also intended that the sequences of steps depicted in the figures are for illustrative purposes only, and are not intended to be limited to any particular sequence of steps. Thus, one skilled in the art can appreciate that these steps may be performed in different orders while performing the same method.

[0180]

[0207]

[0023] In the drawings and specification, illustrative embodiments have been disclosed. However, many variations and modifications to these embodiments may be made. Thus, although specific terminology is employed, it is used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. 1. A video processing method implemented in a decoder, comprising the steps of: determining whether the bitstream contains sub-picture information according to a sub-picture information present flag signaled in the bitstream; In response to the bitstream including the sub-picture information, The number of subpictures in the picture, Target subpicture width, height, position and identifier (ID) mapping; subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag in said bitstream; skipping decoding at least one of the width and the height of the target subpicture if the target subpicture is the last subpicture in the picture. A video processing method comprising:

2. The method of claim 1 , wherein decoding at least one of the width, the height, and the position of the target subpicture is based on the number of the subpictures in the picture.

3. if there are at least two subpictures in the picture, decoding at least one of the subpic_treated_as_pic_flag, the loop_filter_across_subpic_enabled_flag, and the width, the height, and the position of the target subpicture; Further comprising: if there is only one sub-picture in the picture, decoding the at least one of the width, the height and the position of the target sub-picture is skipped; the subpic_treated_as_pic_flag indicates whether a subpicture of each coded picture in a coded layer video sequence (CLVS) is treated as a picture in the decoding process excluding the in-loop filtering operation; and The method of claim 2 , wherein the loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across subpicture boundaries are enabled across subpicture boundaries for each coded picture in the CLVS.

4. if the width of the target subpicture is not signaled, determining the value of the width of the target subpicture as the width of the picture; and determining the value of the height of the target subpicture as the height of the picture if the height of the target subpicture is not signaled; The method of claim 3 further comprising:

5. if the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in coding tree block (CTB) size units as the width of the picture in CTB size units; and determining the value of the height of the target subpicture in CTB size units as the height of the picture in CTB size units if the height of the target subpicture is not signaled; The method of claim 4 further comprising:

6. determining that the subpic_treated_as_pic_flag has a value of 1 if the subpic_treated_as_pic_flag is not signaled in the bitstream; and determining that the loop_filter_across_subpic_enabled_flag has a value of 0 if the loop_filter_across_subpic_enabled_flag is not signaled in the bitstream; The method of claim 3 further comprising:

7. if the width of the target subpicture is not signaled, determining the value of the width of the target subpicture in coding tree block (CTB) size units as the width of the picture in CTB size units minus the horizontal position of a top-left coding tree unit (CTU) of the target subpicture in CTB size units, or determining the value of the width of the target subpicture as the width of the picture minus the horizontal position of a top-left coding tree unit (CTU) of the target subpicture; and if the height of the target subpicture is not signaled, determining the value of the height of the target subpicture in CTB size units as the height of the picture in CTB size units minus the vertical position of the top-left CTU of the target subpicture in CTB size units, or determining the value of the height of the target subpicture as the height of the picture minus the vertical position of the top-left CTU of the target subpicture. The method of claim 1 further comprising:

8. Decoding the ID mapping of the target sub-picture comprises: decoding a first flag in the bitstream; and decoding the ID mapping of the target sub-picture within a first data unit or a second data unit in response to the first flag being equal to one; Further comprising: The method of claim 1 , wherein the first flag equal to 0 indicates that the ID mapping of the target subpicture is not signaled in the bitstream.

9. in response to the first flag being equal to one and the ID mapping of the target sub-picture not being signaled in the first data unit, decoding the ID mapping of the target sub-picture in the second data unit; or skipping decoding the ID mapping of the target sub-picture in the second data unit in response to the first flag being equal to 0 or the ID mapping of the target sub-picture being signaled in the first data unit. The method of claim 8 further comprising:

10. The method of claim 9 , wherein each of the first data unit and the second data unit is one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header (PH).

11. A video processing device implemented in an encoder, comprising: at least one memory for storing instructions; at least one processor, the at least one processor comprising: determining whether the bitstream contains sub-picture information according to a sub-picture information present flag signaled in the bitstream; In response to the bitstream including the sub-picture information, The number of subpictures in the picture, Target subpicture width, height, position and identifier (ID) mapping; subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag signaling in the bitstream at least one of skipping signaling of at least one of the width and the height of the target subpicture if the target subpicture is the last subpicture in the picture. a video processing device configured to execute the instructions to cause the device to

12. The apparatus of claim 11 , wherein the signaling of at least one of the width, the height, and the position of the target subpicture is based on the number of the subpictures in the picture.

13. The at least one processor signaling at least one of the subpic_treated_as_pic_flag, the loop_filter_across_subpic_enabled_flag, and the width, the height, and the position of the target subpicture if there are at least two subpictures in the picture; configured to execute the instructions to cause the device to if there is only one sub-picture in the picture, the signaling of the at least one of the width, the height and the position of the target sub-picture is skipped; the subpic_treated_as_pic_flag indicates whether a subpicture of each coded picture in a coded layer video sequence (CLVS) is treated as a picture in the decoding process excluding the in-loop filtering operation; and The apparatus of claim 12 , wherein the loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across subpicture boundaries are enabled across subpicture boundaries of each coded picture in the CLVS.

14. The at least one processor if the width of the target subpicture is not signaled, determining the value of the width of the target subpicture as the width of the picture; and determining the value of the height of the target subpicture as the height of the picture if the height of the target subpicture is not signaled; 14. The device of claim 13, configured to execute the instructions to cause the device to:

15. The at least one processor if the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in coding tree block (CTB) size units as the width of the picture in CTB size units; and determining the value of the height of the target subpicture in CTB size units as the height of the picture in CTB size units if the height of the target subpicture is not signaled; 15. The device of claim 14, configured to execute the instructions to cause the device to:

16. The at least one processor determining that the subpic_treated_as_pic_flag has a value of 1 if the subpic_treated_as_pic_flag is not signaled in the bitstream; and determining that the loop_filter_across_subpic_enabled_flag has a value of 0 if the loop_filter_across_subpic_enabled_flag is not signaled in the bitstream; 14. The device of claim 13, configured to execute the instructions to cause the device to:

17. The at least one processor if the width of the target subpicture is not signaled, determining the value of the width of the target subpicture in coding tree block (CTB) size units as the width of the picture in CTB size units minus the horizontal position of a top-left coding tree unit (CTU) of the target subpicture in CTB size units, or determining the value of the width of the target subpicture as the width of the picture minus the horizontal position of a top-left coding tree unit (CTU) of the target subpicture; and if the height of the target subpicture is not signaled, determining the value of the height of the target subpicture in CTB size units as the height of the picture in CTB size units minus the vertical position of the top-left CTU of the target subpicture in CTB size units, or determining the value of the height of the target subpicture as the height of the picture minus the vertical position of the top-left CTU of the target subpicture. The device of claim 11 , configured to execute the instructions to cause the device to:

18. A method for storing a video bitstream, the method comprising: Receiving a video sequence; encoding one or more pictures of the video sequence; generating a bitstream based on said encoding; and storing the bitstream in a non-transitory computer readable storage medium; and said encoding comprises: determining whether the bitstream contains sub-picture information according to a sub-picture information present flag signaled in the bitstream; In response to the bitstream including the sub-picture information, The number of subpictures in the picture, Target subpicture width, height, position and identifier (ID) mapping; subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag signaling in the bitstream at least one of skipping signaling of at least one of the width and the height of the target subpicture if the target subpicture is the last subpicture in the picture. A method comprising:

Citation Information

Patent Citations

  • Tile partitions including subtiles in video coding

    JP2021528003A

  • Intra-prediction-based image encoding / decoding method and apparatus

    JP2022515992A

  • Video coding with subpicture, slice, and tile support

    JP2022553599A

  • Tile partitions with sub-tiles in video coding

    WO2019243539A1